Dynamic spectrum allocation method and system, storage medium, device and program product

By combining target reinforcement learning models and resource allocation models in the 5G NTN satellite communication system, spectrum resources are dynamically allocated, solving the problems of low spectrum resource utilization efficiency and insufficient adaptability, and realizing efficient utilization and flexible management of spectrum resources.

CN120935790APending Publication Date: 2025-11-11ZTE CORP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511092515.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

In existing 5G NTN satellite communication systems, spectrum resources are allocated statically, resulting in low spectrum resource utilization efficiency, weak dynamic adaptability, and a waste of spectrum resources.

Method used

By acquiring raw spectrum data and terminal demand data, and dynamically combining target reinforcement learning models and target resource allocation models, target spectrum allocation parameters are determined, including spectrum allocation probability, power parameters, and dynamic value coefficients, thereby achieving efficient allocation of spectrum resources.

Benefits of technology

It improves the utilization rate of spectrum resources, avoids the waste of spectrum resources, enhances the dynamic adaptability of spectrum resource allocation, and can meet the real-time needs of complex and ever-changing network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120935790A_ABST
    Figure CN120935790A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a dynamic spectrum allocation method and system, a storage medium, a device and a program product, and the method comprises the steps: obtaining original spectrum data and terminal demand data, and determining target spectrum data corresponding to the original spectrum data; processing the target spectrum data and the terminal demand data through a target reinforcement learning model to obtain a target spectrum allocation parameter corresponding to the target spectrum data; the target spectrum allocation parameters comprise a target spectrum allocation probability, a target power parameter and a target dynamic value coefficient; and determining a target spectrum allocation result through the target resource allocation model according to the target spectrum data, the terminal demand data and the target spectrum allocation parameters. Therefore, efficient allocation of the spectrum resources can be realized, the utilization rate of the spectrum resources is improved, waste of the spectrum resources is avoided, spectrum resource allocation can be realized based on spectrum data of an actual scene and terminal requirements, and the dynamic adaptability of spectrum resource allocation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communications, and more specifically, to a dynamic spectrum allocation method, system, storage medium, device, and program product. Background Technology

[0002] With the rapid popularization of 5G technology and the continuous growth of global communication demand, non-terrestrial networks (NTNs) have become a key component of 5G systems. 5G NTN technology, by integrating satellite communication with terrestrial 5G networks, can achieve broadband communication coverage in remote areas, oceans, aviation, and other scenarios where terrestrial networks are difficult to cover, thereby providing seamless global communication services.

[0003] Satellite communication plays a crucial role in 5G NTN technology. In related technologies, satellite communication systems typically employ a static allocation method for spectrum resources, meaning spectrum resources are pre-allocated to specific users or services. This method suffers from low spectrum resource utilization efficiency, weak dynamic adaptability, and a waste of spectrum resources. Summary of the Invention

[0004] This application provides a dynamic spectrum allocation method, system, storage medium, device, and program product to at least solve the problems of low spectrum resource utilization efficiency, weak dynamic adaptability, and spectrum resource waste in related technologies.

[0005] According to one embodiment of this application, a dynamic spectrum allocation method is provided, comprising:

[0006] Acquire raw spectrum data and terminal demand data, and determine the target spectrum data corresponding to the raw spectrum data;

[0007] The target spectrum data and the terminal demand data are processed using a target reinforcement learning model to obtain target spectrum allocation parameters corresponding to the target spectrum data; the target spectrum allocation parameters include target spectrum allocation probability, target power parameter, and target dynamic value coefficient.

[0008] The target spectrum allocation result is determined using the target resource allocation model based on the target spectrum data, the terminal demand data, and the target spectrum allocation parameters.

[0009] According to another embodiment of this application, a dynamic spectrum allocation system is provided, comprising:

[0010] Target satellite, target base station, user terminal, and spectrum management platform;

[0011] The target satellite and the target base station acquire raw spectrum data, and the user terminal acquires terminal demand data.

[0012] The spectrum management platform determines the target spectrum data corresponding to the original spectrum data;

[0013] The spectrum management platform processes the target spectrum data and the terminal demand data using a target reinforcement learning model to obtain target spectrum allocation parameters corresponding to the target spectrum data; the target spectrum allocation parameters include target spectrum allocation probability, target power parameter, and target dynamic value coefficient.

[0014] The spectrum management platform determines the target spectrum allocation result based on the target spectrum data, the terminal demand data, and the target spectrum allocation parameters through a target resource allocation model.

[0015] According to yet another embodiment of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, and the computer program is configured to execute the steps in any of the above embodiments of the dynamic spectrum allocation method when it is run.

[0016] According to yet another embodiment of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above embodiments of the dynamic spectrum allocation method.

[0017] According to yet another embodiment of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps in any of the above-described methods for dynamic spectrum allocation.

[0018] This application embodiment acquires raw spectrum data and terminal demand data, and determines the target spectrum data corresponding to the raw spectrum data. Using a target reinforcement learning model, the target spectrum data and terminal demand data are processed to obtain target spectrum allocation parameters corresponding to the target spectrum data. These target spectrum allocation parameters include target spectrum allocation probability, target power parameters, and target dynamic value coefficients. Using a target resource allocation model, the target spectrum allocation result is determined based on the target spectrum data, terminal demand data, and target spectrum allocation parameters. This application embodiment, through the dynamic combination of the target reinforcement learning model and the target resource allocation model, achieves efficient spectrum resource allocation, improves spectrum resource utilization, avoids spectrum resource waste, and enables spectrum resource allocation based on actual scenario spectrum data and terminal demands, thus improving the dynamic adaptability of spectrum resource allocation. Attached Figure Description

[0019] Figure 1 This is a hardware structure block diagram of a computer terminal used in an embodiment of the method of this application;

[0020] Figure 2 This is a schematic diagram of the communication network architecture of a satellite communication system according to an embodiment of this application;

[0021] Figure 3 This is a flowchart of a dynamic spectrum allocation method according to an embodiment of this application;

[0022] Figure 4 This is a flowchart of another dynamic spectrum allocation method according to an embodiment of this application;

[0023] Figure 5 This is a schematic diagram illustrating a final reward and an immediate reward according to an embodiment of this application;

[0024] Figure 6 This is a schematic diagram of a PPO2 algorithm model according to an embodiment of this application;

[0025] Figure 7 This is a schematic diagram illustrating the execution flow of a target resource allocation model according to an embodiment of this application;

[0026] Figure 8 This is a schematic diagram of the interaction process of a dynamic spectrum allocation method according to an embodiment of this application;

[0027] Figure 9 This is a schematic diagram illustrating the execution flow of a dynamic spectrum allocation method according to an embodiment of this application;

[0028] Figure 10 This is a schematic diagram of a dynamic spectrum allocation system according to an embodiment of this application;

[0029] Figure 11 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0030] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples. It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0031] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation portals for users to choose to authorize or refuse.

[0032] Furthermore, the technical solution involved in this application, which analyzes user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.) and uses artificial intelligence technology to make automated decisions, and makes decisions that have a significant impact on personal rights based on the results of automated decisions, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decisions; if the user chooses to reject, the process will proceed to the expert decision-making process.

[0033] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.

[0034] In 5G NTN satellite communication applications, satellite communication systems play a crucial role. Low-Earth orbit (LEO) satellite constellations can cover a wide geographical area and communicate with ground user terminals and satellite ground stations. In related technologies, satellite communication systems typically allocate spectrum resources statically. This allocation method lacks intelligent dynamic adjustment mechanisms and economic incentive strategies, failing to dynamically optimize based on real-time communication needs. This results in low utilization efficiency of spectrum resources in different time periods and regions, leading to wasted spectrum resources. For example, during peak tourist seasons, spectrum resources are in short supply, causing communication congestion; while during off-seasons, a large amount of spectrum resources are idle and wasted. Furthermore, the spectrum resource management in related technologies lacks flexibility and intelligence, making it difficult to adapt to complex and changing communication environments. Its low dynamic adaptability fails to meet the demands of modern communication systems for efficient and flexible spectrum management.

[0035] To address the aforementioned issues, embodiments of this application provide a dynamic spectrum allocation method, system, storage medium, device, and program product. The dynamic spectrum allocation system acquires raw spectrum data and terminal demand data, and determines the target spectrum data corresponding to the raw spectrum data. Through a target reinforcement learning model, the target spectrum data and terminal demand data are processed to obtain target spectrum allocation parameters corresponding to the target spectrum data. These target spectrum allocation parameters include target spectrum allocation probability, target power parameters, and target dynamic value coefficients. Through a target resource allocation model, the target spectrum allocation result is determined based on the target spectrum data, terminal demand data, and target spectrum allocation parameters. Thus, embodiments of this application, through the dynamic combination of a target reinforcement learning model and a target resource allocation model, can achieve efficient spectrum resource allocation, improve spectrum resource utilization, avoid spectrum resource waste, and allocate spectrum resources based on spectrum data and terminal demands in real-world scenarios, improving the dynamic adaptability of spectrum resource allocation and meeting real-time requirements in complex and ever-changing network environments.

[0036] The methods and embodiments provided in this application can be executed on a mobile terminal, computer terminal, or similar computing device (or electronic device, etc.). Taking running on a computer terminal as an example, Figure 1 This is a hardware structure block diagram of a computer terminal used in an embodiment of the method of this application. Figure 1 As shown, a computer terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor (MCU) or a programmable gate array (FPGA)) and a memory 104 for storing data are also shown. The computer terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the computer terminal described above. For example, the computer terminal may also include components that are more complex than those described above. Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0037] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to a computer terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0038] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the computer terminal. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0039] Figure 2 This is a schematic diagram of the communication network architecture of a satellite communication system according to an embodiment of this application. Figure 2 As shown, in this satellite communication system, the user equipment (UE) can communicate with the target satellite to provide various communication services; the target base station located on the ground can communicate with the target satellite to exchange data. This application does not limit the specific network type in the satellite communication system.

[0040] Figure 3 This is a flowchart illustrating a dynamic spectrum allocation method according to an embodiment of this application. This dynamic spectrum allocation method can be applied to the aforementioned computer terminal or to a dynamic spectrum allocation system for the aforementioned satellite communication system. Specifically, it can be applied to the spectrum management platform within the dynamic spectrum allocation system; however, this embodiment does not limit its application in this regard. Figure 3 As shown, this dynamic spectrum method includes the following steps:

[0041] Step S301: Obtain the original spectrum data and terminal requirement data, and determine the target spectrum data corresponding to the original spectrum data.

[0042] In this embodiment, the raw spectrum data can refer to the original, unprocessed initial spectrum data collected, which may include signal strength, spectrum occupancy, signal-to-noise ratio (SNR), bandwidth, frequency offset, user location, user mobility information, traffic volume, quality of service (QoS) parameters, device status information, service type, and environmental information. Of course, the raw spectrum data may also include other data, which can be set based on actual needs; this embodiment does not limit this. Terminal requirement data can refer to the actual user requirement data uploaded by the user terminal, which may specifically include bandwidth requirements, server quality requirements, and cost requirements. Target spectrum data can refer to the spectrum data obtained after processing the raw spectrum data, which can be directly processed by subsequent models.

[0043] Spectrum sensing is the foundation of dynamic spectrum allocation. A dynamic spectrum allocation system can monitor spectrum usage in real time, providing accurate data support for subsequent spectrum allocation decisions. Specifically, the spectrum management platform can collect raw spectrum data through spectrum sensors deployed on target satellites and ground-based target base stations. These sensors cover frequency bands such as Ka (K-band Above) and Q (Quad), as well as the mainstream frequency bands of terrestrial 5G networks, enabling real-time collection of various types of spectrum data. In addition, the spectrum management platform can also collect terminal demand data uploaded by user terminals in real time. The spectrum management platform may include a data processing center for data processing. Based on this data processing center, the platform can process the raw spectrum data to obtain target spectrum data that subsequent models can directly process.

[0044] Step S302: The target spectrum data and terminal demand data are processed through the target reinforcement learning model to obtain the target spectrum allocation parameters corresponding to the target spectrum data; the target spectrum allocation parameters include the target spectrum allocation probability, the target power parameter, and the target dynamic value coefficient.

[0045] In this embodiment, the target reinforcement learning model can refer to a model used for spectrum allocation prediction, specifically a proximal policy optimization algorithm (PPO2), a soft actor-critic (SAC) algorithm, or a deep deterministic policy gradient algorithm (DDPG), etc. The following description uses the PPO2 algorithm as the target reinforcement learning model. The PPO2 algorithm is a proximal policy optimization algorithm with good stability and convergence. It updates the policy network parameters through multiple iterations and uses the truncation between the agent policy and the old policy to limit the update pace, thus ensuring the stability of the policy update. Of course, in practical scenarios, other algorithm models can also be used for this target reinforcement learning model, and this embodiment does not limit this.

[0046] The input data for the target reinforcement learning model can be target spectrum data and terminal demand data, and the output data can be target spectrum allocation parameters. The target spectrum allocation parameters corresponding to the target spectrum data can include target spectrum allocation probabilities, target power parameters, and target dynamic value coefficients. The target spectrum allocation probability refers to the allocation probability of each frequency band for each user; the target power parameter refers to the power range of each frequency band, such as the maximum allowed transmit power; and the target dynamic value coefficient can be used to guide the resource allocation of subsequent target resource allocation models, balancing the resource needs of different network nodes and ensuring efficient utilization of spectrum resources.

[0047] Specifically, after the target spectrum data is determined, the spectrum management platform can process the target spectrum data through a target reinforcement learning model to obtain the target spectrum allocation parameters. Based on the target spectrum data and terminal demand data, dynamic prediction of spectrum resource allocation can be achieved, which can ensure the flexibility and dynamism of spectrum resource allocation.

[0048] Step S303: Using the target resource allocation model, determine the target spectrum allocation result based on the target spectrum data, terminal demand data, and target spectrum allocation parameters.

[0049] In this embodiment, the target resource allocation model can be used for dynamic and efficient allocation of spectrum resources. Specifically, it can refer to auction algorithm models, proportional fairness algorithm models, and genetic algorithm models, etc. The following description uses an auction algorithm model as an example of the target resource allocation model. Of course, other algorithm models can also be used, and this embodiment does not limit this. The target spectrum allocation result can refer to the spectrum resource allocation method based on the target spectrum data. Specifically, after the spectrum management platform determines the target spectrum allocation parameters corresponding to the target spectrum data through the target reinforcement learning model, it can further determine the target spectrum allocation result by combining the target spectrum data, terminal demand data, and target spectrum allocation parameters through the target resource allocation model. In this way, through the target resource allocation model, the spectrum management platform can achieve reasonable and efficient allocation of spectrum resources based on user needs, avoiding spectrum resource waste.

[0050] In this embodiment, the spectrum management platform acquires raw spectrum data and terminal demand data, and determines the target spectrum data corresponding to the raw spectrum data. Using a target reinforcement learning model, the target spectrum data and terminal demand data are processed to obtain target spectrum allocation parameters corresponding to the target spectrum data. These parameters include the target spectrum allocation probability, target power parameter, and target dynamic value coefficient. Using a target resource allocation model, the target spectrum allocation result is determined based on the target spectrum data, terminal demand data, and target spectrum allocation parameters. This embodiment, through the dynamic combination of the target reinforcement learning model and the target resource allocation model, achieves efficient spectrum resource allocation, improves spectrum resource utilization, avoids spectrum resource waste, and enables spectrum resource allocation based on actual scenario spectrum data and terminal demands, thus improving the dynamic adaptability of spectrum resource allocation.

[0051] Based on the above embodiments, Figure 4 Here is a flowchart of another dynamic spectrum allocation method according to an embodiment of this application, such as... Figure 4 As shown, the dynamic spectrum allocation method includes the following steps:

[0052] Step S401: Collect raw spectrum data according to a preset time period using a spectrum sensor; the spectrum sensor includes a satellite spectrum sensor and a ground spectrum sensor; acquire terminal demand data sent by the user terminal; the terminal demand data is generated by the user terminal in response to the user's interactive operation.

[0053] The dynamic spectrum allocation system in this application embodiment may include user terminals, target base stations, target satellites, and a spectrum management platform. The target satellites may refer to a satellite system comprising a low-Earth orbit (LEO) satellite constellation with an orbital altitude between 500 and 1200 kilometers. The constellation size can be adjusted according to coverage requirements to achieve seamless global coverage. Each satellite is equipped with a multi-beam antenna and a software-defined radio (SDR) payload, providing flexible spectrum configuration capabilities. The multi-beam antenna enables simultaneous coverage of different areas, improving spectrum reuse; the SDR payload can dynamically adjust frequency band allocation and transmit power via commands from the intelligent spectrum management platform.

[0054] The target base station can refer to terrestrial base stations and their core networks. Target base stations are widely distributed and, based on communication technology and spectrum resources, can provide high-speed, low-latency communication services. The spectrum management platform can adopt a distributed architecture, including a data processing center and an algorithm engine. The data processing center and algorithm engine in the spectrum management platform can be deployed in the target base stations or in cloud servers; this application embodiment does not limit this. The spectrum management platform possesses powerful computing capabilities, can run complex dynamic spectrum allocation algorithms, and has high reliability and security to ensure the stable operation of the communication system. The user terminal can refer to a dual-mode terminal compatible with both satellite and terrestrial networks (such as satellite phones or emergency communication equipment), possessing automatic switching capabilities between satellite and terrestrial networks to ensure communication continuity.

[0055] In this embodiment, a satellite spectrum sensor is deployed in the target satellite, and a ground spectrum sensor is deployed in the target base station. The satellite spectrum sensor in the target satellite collects raw spectrum data of the Ka / Q band in real time and transmits this raw spectrum data to the data processing center of the ground station via an inter-satellite link. The ground spectrum sensor in the target base station monitors raw spectrum data of the 5G New Radio (NR) band in real time and transmits this raw spectrum data to the data processing center via optical fiber. In addition, the user terminal can respond to the user's interactive operation (such as touch input or voice input), generate terminal demand data corresponding to the interactive operation, and report the terminal demand data to the spectrum management platform. Of course, the terminal demand data can also be collected by the spectrum sensor, and this embodiment does not limit this. In this way, the spectrum management platform, through the spectrum sensor, realizes real-time monitoring and dynamic collection of raw spectrum data and terminal demand data, which can ensure the real-time and dynamic nature of spectrum resource allocation and meet the spectrum allocation needs of actual scenarios.

[0056] Step S402: The original spectrum data is parsed and processed to extract candidate spectrum data from the original spectrum data; the candidate spectrum data is filtered and features are extracted to obtain the target spectrum data corresponding to the original spectrum data.

[0057] In this embodiment, the data processing center of the spectrum management platform can be integrated into the target base station, specifically using a Graphics Processing Unit (GPU) cluster for data processing. The data processing center first parses the raw spectrum data, extracting the spectrum data and auxiliary information, such as timestamps and frequency band identifiers, to obtain candidate spectrum data. Then, the data processing center performs data cleaning and feature extraction on the candidate spectrum data. Data cleaning includes removing missing and outlier values, while feature extraction refers to calculating features such as spectrum occupancy and signal-to-noise ratio based on the candidate spectrum data. In this way, the spectrum management platform, through data preprocessing of the raw spectrum data by the data processing center, can improve the accuracy and reliability of the target spectrum data, thereby enhancing the rationality of subsequent spectrum allocation decisions.

[0058] Step S403: Based on the target spectrum data and historical spectrum data, construct a dynamic decision model corresponding to the target reinforcement learning model.

[0059] In this embodiment, historical spectrum data can refer to spectrum data within a historical period. The dynamic decision model can refer to a decision framework used for reinforcement learning, specifically a Markov Decision Process (MDP) model, a Partially Observable Markov Decision Process (POMDP) ​​model, or Stochastic Dynamic Programming (SDP), etc. The following description uses an MDP model as an example of the dynamic decision model; however, other models can also be used, and this embodiment does not limit this. This dynamic decision model can be part of a target reinforcement learning model, enabling the target reinforcement learning model to perform reinforcement learning based on this dynamic decision model, achieving dynamic and reasonable prediction of the target spectrum data.

[0060] In this embodiment, the dynamic spectrum allocation problem in the 5G NTN scenario can be transformed into a dynamic decision problem. Taking the MDP model as an example, the dynamic decision model can be represented as a tuple.<S,A,L,R,γ> Where S is the target state space; A is the target action space; L is the target state transition function; and R is the target reward function, where r n =R(s) n =s,an =a,s n+1 =s′) represents the immediate reward obtained after taking action a in state s and transitioning to state s′ at time n; γ∈[0,1) is a discount factor that can be used to distinguish between long-term and short-term rewards. This discount factor can be a preset value and can be set according to actual needs. For example, the discount factor can be set to 0.9 so that the algorithm considers both short-term rewards and long-term benefits when considering future rewards.

[0061] In one exemplary embodiment, the dynamic decision model in step S403 can be constructed as follows:

[0062] Based on the target spectrum data, determine the target state space corresponding to the dynamic decision-making model; the target state space includes: spectrum state parameters, terminal demand data, satellite state parameters, historical spectrum parameters, and environmental parameters; determine the target action space corresponding to the dynamic decision-making model; the target action space includes spectrum allocation probability, power parameters, and dynamic value coefficients; determine the target reward function, target state transition function, and discount factor corresponding to the dynamic decision-making model; construct the dynamic decision-making model based on the target state space, target action space, target reward function, target state transition function, and discount factor.

[0063] In this embodiment, based on the tuples of the above dynamic decision-making model, the spectrum management platform can construct a dynamic decision-making model corresponding to the target reinforcement learning model based on the target state space, target action space, target reward function, target state transition function, and discount factor. Specifically, for the application scenario of 5G NTN satellite communication spectrum allocation, the target state space and target action space of the dynamic decision-making model in this embodiment adopt new definitions. Compared with the fixed-dimensional state space and action space used in related technologies, the target state space in this embodiment includes spectrum state parameters, terminal demand data, satellite state parameters, historical spectrum parameters, and environmental parameters, which can comprehensively and accurately reflect the dynamic spectrum environment and system operation status. Among them, the spectrum state parameters, satellite state parameters, and environmental parameters can be determined based on the target spectrum data. The spectrum state parameters can specifically include frequency band occupancy rate and signal-to-noise ratio, etc. The satellite state parameters can include satellite position, Doppler offset, and satellite power, etc. The environmental parameters can include environmental status and weather attenuation coefficient, etc.; the historical spectrum parameters can be determined based on historical spectrum data. For example, the target state space can be represented by the following formula (1):

[0064]

[0065] In the above formula (1), s t F represents the target state space at time t; tIt represents the occupancy rate of each frequency band, with a value range of 0 to 1. As a spectrum status parameter, it can be measured in real time by a spectrum sensor. This indicates the bandwidth currently requested by user u, which can be determined based on terminal demand data; The QoS level for user u can be determined based on terminal demand data, specifically based on core network signaling; Pos t These are the satellite's three-dimensional orbital coordinates; H represents the satellite's remaining available transmission power. t-T:t This refers to historical spectral efficiency data for the past T steps; Δf t This refers to the Doppler offset, also known as the user movement parameter, etc.; ξ t This represents the weather attenuation coefficient, which is used to characterize the degree of signal strength attenuation under severe weather conditions. The vector space representing the field of real numbers.

[0066] In this embodiment, the target state space of the dynamic decision-making model integrates multi-dimensional information such as spectrum status, user demand, satellite status, historical data, and environmental parameters. This enables the target reinforcement learning model to more comprehensively perceive the complex state of the spectrum sharing environment, and to perceive in real time the impact of dynamic factors such as satellite orbit changes and user movement on channel quality. This improves the dynamic adaptability of spectrum resource allocation and reduces the fluctuation range of the average user transmission rate, ensuring the rationality and accuracy of spectrum allocation prediction and reducing the resource mismatch rate.

[0067] The target action space in this embodiment may include spectrum allocation probability, power parameters, and dynamic value coefficients. Spectrum allocation probability, also known as spectrum allocation priority, refers to the probability distribution of allocating different spectrums to different users, and can be represented by softmax probability. Power parameters can refer to the transmit power range of each frequency band, specifically the maximum allowable transmit power of each frequency band. Dynamic value coefficients can be used in subsequent target resource allocation models for resource allocation. For example, the target action space can be represented by the following formula (2):

[0068]

[0069] In the above formula (2), a t Represents the target action space; p b.u This represents the probability of spectrum allocation. α represents the maximum permissible transmit power for frequency band b. tThe dynamic value coefficient, also known as the dynamic pricing coefficient or dynamic priority coefficient, represents the target action space in this application embodiment. By introducing spectrum allocation priority and interference avoidance parameters such as power parameters, the flexibility and accuracy of spectrum allocation can be ensured to meet the actual usage needs of different users. In addition, by determining the dynamic value coefficient, the dynamic interaction between the target reinforcement learning model and the target resource allocation model can be ensured, thereby ensuring the accuracy and efficiency of spectrum allocation and improving spectrum utilization.

[0070] In one exemplary embodiment, the objective reward function of the dynamic decision model can be determined as follows:

[0071] Based on the bandwidth data and spectrum allocation results at the target time, determine the spectrum utilization rate at the target time; based on the actual interference intensity and preset interference intensity of each frequency band at the target time, determine the interference level parameters at the target time; based on the service quality level and number of users at the target time, determine the fairness index parameters at the target time; the service quality level is determined based on the actual allocated bandwidth, user demand bandwidth, and user weight coefficient; based on the service quality level, service quality threshold, actual transmit power, and power threshold at the target time, determine the anomaly penalty parameters at the target time; based on the spectrum utilization rate, interference level parameters, fairness index parameters, and anomaly penalty parameters, determine the target reward function.

[0072] In this embodiment, the target reward function can be used to evaluate whether the representation of each state is excellent, and can be used to train the target reinforcement learning model to guide the efficient and accurate allocation of spectrum resources. Compared with the reward function in related technologies that focuses solely on spectrum utilization, the target reward function in this embodiment is determined comprehensively based on spectrum utilization, interference level parameters, fairness index parameters, and anomaly penalty parameters, thus achieving balanced optimization of multiple objectives in the spectrum allocation process. Specifically, spectrum utilization can be determined based on the bandwidth data at the target time and the spectrum allocation results; interference level parameters can be determined based on the actual interference intensity of each frequency band and the preset interference intensity (e.g., the maximum allowable interference intensity); fairness index parameters can be calculated based on the spectrum allocation differences between users, specifically based on service quality parameters and user quality; anomaly penalty parameters can be determined based on whether the power limit is exceeded or the service quality is met. If the power limit is exceeded or the QoS is not met, it is set to a preset normal value; otherwise, it is set to 0. For example, the target reward function can be expressed by the following formula (3):

[0073] r t =ω1·U t +ω2·(1-I t )+ω3·E t -ω4·C t (3)

[0074] In the above formula (3), r t This represents the target reward function value; ω1, ω2, ω3, and ω4 are weight values, which can be pre-configured or determined based on actual business needs. t For spectral efficiency, the value ranges from 0 to 1. This U t It can be calculated using the following formula (4):

[0075]

[0076] In the above formula (4), W represents the spectrum allocation result obtained from the target resource allocation model. b W represents the bandwidth of the b-th frequency band. total The total available bandwidth, i.e. In the above formula (3), I t The interference level parameter can be calculated using the following formula (5):

[0077]

[0078] In the above formula (5), The actual interference intensity is represented by I, which is the real-time measurement value of the spectrum sensor in each frequency band, and represents the actual interference intensity of the b-th frequency band at time t. th The preset interference strength, i.e., the maximum allowable interference strength threshold, can be a pre-configured fixed threshold parameter. In the above formula (3), E t The fairness indicator parameters can be determined using the following formula (6):

[0079]

[0080] In the above formula (5), This represents the quality of service level of user u at time t, which can be specifically expressed through... The calculation involves determining the actual allocated bandwidth, which can refer to the allocated bandwidth for user u calculated by the target resource allocation model, and the user-requested bandwidth, which can refer to the user-demanded bandwidth in the terminal demand data. The user weight coefficient, or user priority coefficient, can be determined based on user type. For example, the user weight coefficient for a regular user is 1, and the user weight coefficient for an emergency user is 3. In the above formula (3), C t The exception penalty parameter can be determined using the following formula (7):

[0081]

[0082] In the above formula (7), the anomaly penalty parameter Ct It can be used to constrain violation penalties, QoS min The pre-configured Quality of Service (QoS) threshold, i.e. the minimum QoS threshold required by the system, can be configured according to the type of communication service (e.g., normal service or emergency communication) or by the administrator of the spectrum management platform, to ensure basic communication quality for users. This indicates the actual transmit power of frequency band b; The maximum permissible transmit power for frequency band b is the output of the PPO2 algorithm. If the power exceeds the limit or QoS is not met, an abnormal penalty parameter C is applied. t Set to the default normal value; otherwise, set to 0.

[0083] In this embodiment, the objective reward function comprehensively incorporates spectrum utilization, interference level, and fairness indicators, while also introducing penalty terms to constrain corresponding negative behaviors. This can more accurately guide the algorithm towards improving the overall performance of spectrum allocation, solve the problem of increased interference or unfairness to users caused by prioritizing spectrum utilization in spectrum allocation of related technologies, improve the efficiency and flexibility of spectrum allocation, meet the actual needs of users, and enhance the user experience.

[0084] In one exemplary embodiment, the target state transition function in the dynamic decision model can be determined as follows:

[0085] Based on historical spectrum data and historical action data within a historical period, the historical mapping relationship between historical states, historical actions, and the next historical state is determined; based on the historical mapping relationship, the state transition frequency corresponding to different historical actions is determined; based on the state transition frequency, the state transition probability within the historical period is determined, and the target state transition function is obtained.

[0086] In this embodiment, the target state transition function can also be called the target state transition probability L(s′|s,a), which is used to predict the impact of the current action on the future spectrum state and provide a probability distribution prediction of the future state. Specifically, the spectrum management platform can determine the historical mapping relationship of historical state-historical executed action-next historical state (i.e., state-action-next state) based on historical spectrum data and historical action data. Then, based on this historical mapping relationship, it can count the transition frequency of historical periods (e.g., within the past T periods) to construct a state transition probability matrix and obtain the target state transition function.

[0087] For example, assuming the spectrum occupancy change satisfies the first-order Markov property, the target state transition function can be obtained by training a Markov model based on historical spectrum data and historical action data. During the PPO2 algorithm training phase, the target state transition function can be used to calculate the expected reward of the state value function. During the decision-making phase, it can be used through Monte Carlo sampling or approximate inference to provide a probability distribution prediction of future states for the policy network, thereby achieving causal chain modeling of "current action - future state - long-term reward". Of course, in practical applications, this target state transition function can also be implicitly learned using deep neural networks, such as recurrent neural networks; this application does not limit this approach.

[0088] In this embodiment, the spectrum management platform can construct a dynamic decision-making model based on the target state space, target action space, target reward function, target state transition function, and discount factor. The target reinforcement learning model then uses this dynamic decision-making model to predict spectrum resource allocation, enabling flexible adjustment of spectrum resource allocation according to real-time communication needs and improving frequency utilization. For example, in scenarios with large fluctuations in user density, such as remote areas and tourist attractions, this dynamic spectrum allocation method can significantly improve spectrum utilization. During peak user density periods, spectrum resources can be quickly allocated to the demand area to meet communication needs; during off-peak periods, idle spectrum resources can be promptly recovered to avoid waste.

[0089] Step S404: Determine the real-time dynamic parameters of the target reinforcement learning model.

[0090] In one exemplary embodiment, the real-time dynamic parameters can be determined in the following manner:

[0091] Within the initial time period, the initial parameters of the target reinforcement learning model are determined and used as real-time dynamic parameters. Outside the initial time period, the real-time dynamic parameters of the target reinforcement learning model are determined based on the initial parameters and the fluctuation parameters of the spectral data.

[0092] In this embodiment, the target reinforcement learning model may include a dynamic decision model, a policy network, and a value network. The dynamic decision model provides the environmental framework for reinforcement learning, the policy network performs probability prediction, and the value network performs value calculation. Before starting the spectrum allocation task, the spectrum management platform can initialize the target reinforcement learning model by setting its initial parameters, such as the learning rate η and discount factor γ. During the spectrum allocation task, real-time dynamic parameters can be determined based on these initial parameters. "Various model parameters" can refer to various model parameters within the target reinforcement learning model.

[0093] Within the initial time period, the target reinforcement learning model can use the initial parameters as real-time dynamic parameters, enabling rapid exploration of the parameter space in the early stages of algorithm training. Taking the learning rate η as an example, this learning rate η can be used to control the step size of model parameter updates, and the initial value of η can be 0.001 within the initial time period. The specific duration of this initial time period can be set based on actual needs, and this embodiment does not limit it. As the spectrum allocation task continues to execute, the model parameters can be dynamically adjusted during model training. That is, within non-output time periods, the target reinforcement learning model can improve the convergence accuracy of the model by determining real-time dynamic parameters, ensuring the accuracy of spectrum resource allocation.

[0094] Specifically, during non-initial time periods, real-time dynamic parameters can be determined based on initial parameters and spectral data fluctuation parameters, which may include user density fluctuations, signal strength fluctuations, or signal-to-noise ratio fluctuations. Taking the real-time dynamic parameter as the dynamic learning rate as an example, the dynamic learning rate can be calculated using the following formula (8):

[0095]

[0096] In the above formula (8), η base The base learning rate is the initial parameter of the learning rate; σ demand The fluctuation parameters of the spectrum data, specifically user volatility or user density standard deviation, can be calculated using historical spectrum data, specifically according to formula (9). In the above formula (9), D t The total bandwidth requested by users in time slot t can be obtained based on terminal demand data. It reflects the overall demand for spectrum resources from all users in the current time slot, such as the bandwidth demand d declared by a user terminal when submitting a communication request. u D t =∑d u μ D The average bandwidth requirement over the past T time slots can be calculated using historical bandwidth requirement data obtained through a sliding window. Specifically, it can be calculated using the formula (10) above, which is used to characterize the long-term trend of user demand and provide a benchmark reference for dynamic learning rate adjustment. The dynamic learning rate η is then obtained. t The post-input policy network performs parameter updates in the policy gradient. Furthermore, when σ... demand In the event of a sudden surge, such as a disaster, the target reward function can simultaneously increase the weight ω3 corresponding to the fairness indicator parameter to ensure that base stations in disaster areas are given priority in spectrum resource allocation and to meet the needs of emergency services; η t At lower speeds, the model will narrow the action exploration range, improving the accuracy of spectrum allocation.

[0097] It should be noted that real-time dynamic parameters, in addition to the dynamic learning rate, may also include dynamic batch size and dynamic threshold, and other calculation methods can be used accordingly. This application does not limit these methods. In related technologies, model training and decision-making processes typically use parameters with a fixed decay rate, which cannot perceive changes in the communication environment and can lead to policy divergence in scenarios with fluctuating demand. However, in this application embodiment, the target reinforcement learning model, by determining real-time dynamic parameters, can achieve dynamic optimization of the model's hyperparameters, which can improve the convergence speed and stability of the algorithm in complex communication environments, improve the effectiveness and accuracy of spectrum allocation, and enhance dynamic adaptability.

[0098] In one exemplary embodiment, the spectrum allocation method may further include the following steps:

[0099] If the spectrum allocation task meets the preset termination conditions, determine the action trajectory corresponding to the spectrum allocation task; based on the action trajectory and the spectrum utilization rate, interference level parameters and fairness index parameters of the spectrum allocation task at different target times, determine the final reward corresponding to the spectrum allocation task; based on the final reward, update the network parameters of the target reinforcement learning model.

[0100] In this embodiment, the preset termination condition can refer to a pre-set model prediction termination condition, specifically reaching a preset allocation period (or a preset number of iterations) or a preset spectrum utilization target. When the target reinforcement learning model reaches the preset termination condition during the spectrum allocation task, the dynamic spectrum allocation system enters the final state. At this time, the spectrum management platform can determine the action trajectory of the spectrum allocation task, and then calculate the final reward (or final performance) based on the action trajectory, and update the network parameters of the target reinforcement learning model based on the final reward.

[0101] Specifically, in the implementation of this application, a final long-term planning mechanism is introduced during the execution of the objective reinforcement learning model to achieve a more comprehensive evaluation of long-term rewards. In the final state, the environment will provide a final reward M. T The final reward comprehensively reflects the final effect of this spectrum allocation task, such as the overall spectrum utilization rate and the QoS compliance status of all users at the end of the task. The long-term reward of the objective reinforcement learning model can be calculated by the following formula (11):

[0102]

[0103] The objective reinforcement learning model updates the long-term reward at time t using formula (11), where T is the time step from time t to the final state. Related technologies suffer from a lack of a final objective in their spectrum allocation methods, causing the model optimization direction to deviate from actual needs. However, in this embodiment, the final reward M is calculated... T The final reward can be defined as the average comprehensive score within the round, which can encourage the agent to optimize the immediate reward while taking into account the stability of the entire communication cycle. In the above long-term reward calculation process, by introducing the final reward, the problem of missing final objectives in related technologies can be solved, ensuring that the reinforcement learning objective function is combined with the needs of satellite communication services. The meaning of long-term reward can be changed from "pure accumulation of immediate rewards" to "fusion of immediate rewards and final global performance", further ensuring that the model optimization direction is consistent with actual needs. For example, the final reward can be calculated by the following formula (12):

[0104]

[0105] In the above formula (12), U t For spectrum utilization, I t Interference level parameter E t The fairness index parameters are calculated using formulas (4), (5), and (6) mentioned above, and will not be repeated here in this embodiment. ω5, ω6, and ω7 are weighting coefficients, which can be configured based on actual needs, and are not limited in this embodiment.

[0106] For example, Figure 5 This is a schematic diagram illustrating a final reward and an immediate reward according to an embodiment of this application. Figure 5 As shown, the final reward mechanism in this embodiment incorporates the end-game reward into the long-term return calculation, enabling the algorithm to consider the immediate reward r at each time step. t At the same time, it also considers optimizing the final result, avoiding suboptimal decisions due to excessive focus on short-term rewards, thereby guiding the policy network to generate allocation strategies that are more conducive to the efficient use of spectrum resources in the long term. The goal of the algorithm is to find the optimal spectrum allocation strategy π. * S→A, maximizing the expected cumulative reward, this optimal spectrum allocation strategy can be specifically expressed by the following formula (13):

[0107]

[0108] The action trajectory τ = (s0, a0, r0, s1, a1, r1, ...) is jointly generated by the target reinforcement learning model and the target resource allocation model. Specifically, it refers to the state-action-reward sequence (s0, a0, r0, ...) generated by the target reinforcement learning model and the target resource allocation model during the interaction process, recording the complete decision-making process of the spectrum allocation task from the initial state to the current state. This action trajectory can serve as process data (or "memory carrier," training data, interaction data, etc.) for the collaborative optimization of the target reinforcement learning model and the target resource allocation model. It is both the core data of the experience pool and the basis for long-term reward evaluation, ensuring that the algorithm learns from historical experience and achieves adaptive capability to dynamic environments. When t = T, the immediate reward r t For the final performance M T The calculation is based on the spectrum utilization rate, fairness index parameters, and interference level parameters at the end of the spectrum allocation task, forming a composite optimization objective of immediate rewards and final global performance, thereby improving the accuracy of spectrum resource allocation.

[0109] In this application, an embodiment of the target reinforcement learning model is introduced with a final reward mechanism. The global performance indicators (such as spectrum utilization and fairness parameters) at the time of task termination are explicitly embedded into the long-term reward function, which can improve spectrum utilization, training stability and target achievement rate. It effectively avoids policy optimization bias caused by reward sparsity. The hierarchical design of final reward and immediate reward binds short-term action optimization with long-term global goals, avoiding the problem of agents getting trapped in local optima in related technologies.

[0110] Step S405: Based on the dynamic decision-making model and real-time dynamic parameters, the target spectrum data is processed through the target reinforcement learning model to obtain the target spectrum allocation parameters.

[0111] In this embodiment, the target reinforcement learning model, based on a dynamic decision model and real-time dynamic parameters, can perform predictive processing on target spectrum data to obtain target spectrum allocation parameters, including target spectrum allocation probability, target power parameters, and target dynamic value coefficients.

[0112] Specifically, the target reinforcement learning model can perform in-depth analysis of the target spectrum data to predict future spectrum demand. This spectrum allocation task can be based on time series analysis and machine learning models, comprehensively considering potential spectrum usage trends in different time periods and regions. For example, in a tourist city, spectrum demand surges when entertainment activities increase in the evening, and decreases after people rest late at night. The target reinforcement learning model can capture these patterns by analyzing historical spectrum data and predict future spectrum demand in different time periods and regions accordingly.

[0113] Taking the PPO2 algorithm as an example, this study focuses on the objective reinforcement learning model. Figure 6 This is a schematic diagram of a PPO2 algorithm model according to an embodiment of this application. Figure 6 As shown, the PPO2 algorithm adopts a dual-network architecture of policy network and value network. The specific structure of the policy network and value network can adopt a multi-layer neural network. The policy network is used to generate the probability distribution of spectrum allocation actions based on the current spectrum state. The multi-layer neural network learns and transforms the input target spectrum data and finally outputs the probability value corresponding to each possible action. The value network is used to evaluate the value of the current spectrum state, that is, the estimated value of the long-term cumulative reward that may be obtained after taking the corresponding action in this state. During the training process, the PPO2 algorithm uses the collected spectrum data to train the network and continuously adjusts the network parameters so that the policy network can generate better spectrum allocation strategies and the value network can more accurately evaluate the value of the spectrum state. The two networks work together to provide a scientific basis for spectrum allocation decisions. The objective function of the PPO2 algorithm includes policy loss and value function loss and introduces a clipping mechanism. Based on the deep reinforcement learning framework of the PPO2 algorithm, its objective function can be expressed as the following formula (14):

[0114]

[0115] Where ρ t (θ) is the strategy ratio, which can be calculated using formula (15), A θ Let be the advantage function, representing the state s t Take action a t Compared to the average strategy, ε is the pruning threshold, which can generally be set to 0.2. Of course, the objective function of this PPO2 algorithm can also take other forms, and this application does not limit it.

[0116] like Figure 6As shown, in the dual-stream attention mechanism network architecture of the PPO2 algorithm, the policy network consists of two identical neural networks, Pi and Old Pi (Pi'). The policy network first extracts spatiotemporal features from the input target spectrum data to capture its temporal and spatial variation patterns, and then generates actions such as spectrum allocation decisions. Its output is a probability distribution, representing the probability of taking different spectrum allocation actions in the current state. For example, it provides the probability value of allocating a specific frequency band to different users, providing a basis for the spectrum allocation decision of the subsequent target resource allocation model. The policy network learns and transforms the input state features through multi-layer neural networks, and finally outputs the probability value corresponding to each possible action; the value network evaluates the value of the input state and judges the quality of the spectrum state. The actions generated by the policy network are applied to the spectrum execution devices (i.e., target terminals, such as satellite SDR payloads and 5G base stations) in the actual environment for interaction. Based on the effect of the action execution, the environment monitors the reward signals fed back by spectrum sensors, which include spectrum utilization, interference level, QoS compliance rate, etc. The reward signals are used to optimize the evaluation of the value network after calculating the reward function. The optimization of the value network further guides the policy network to make more reasonable spectrum allocation decisions. The value network, by evaluating the long-term value of the current spectrum state (e.g., future spectrum utilization, user fairness, etc.), and the cumulative reward value network's evaluation results, guides the policy network's action selection, providing a quantifiable reference for the quality of states. Based on this reference, the policy network adjusts the probability distribution of action selection, for example, assigning higher probabilities to spectrum allocation actions corresponding to high-value states, thus guiding the algorithm to prioritize strategies that bring long-term benefits. The evaluation results of the value network serve as an important input for the policy network's gradient update, jointly driving iterative optimization along with the reward function. This shifts action selection from solely focusing on immediate rewards (e.g., current transmission rate in related technologies) to considering end-game objectives (e.g., global spectrum efficiency at the end of the task), forming a closed-loop mechanism of state evaluation, policy optimization, and long-term value enhancement. This makes the policy network more inclined to choose spectrum allocation actions with higher long-term value. The collaborative work of these two networks provides a scientific basis for spectrum allocation decisions.

[0117] The PPO2 algorithm's objective reward function R includes positive rewards (spectrum utilization, interference level parameters, fairness index parameters) and negative rewards (QoS violation penalties and power overruns). The reward function R(s) evaluated by the value network is compared with the old policy π' of the policy network to calculate the truncated loss. This limits the policy update magnitude to ensure training stability, thus forming a closed-loop process of continuous optimization. Action outputs are transmitted to the execution device via satellite ground station commands. The execution effect is monitored in real-time by spectrum sensors and fed back to the experience pool. The experience pool stores a four-tuple of data: state, action, reward, and next state, providing samples for algorithm iteration and forming a complete closed loop of perception-decision-execution-optimization, improving the dynamic adaptability of spectrum resource allocation.

[0118] Step S406: Normalize the target spectrum allocation probability to obtain the target allocation priority for each frequency band.

[0119] In this embodiment, the target reinforcement learning model performs prediction processing on the target spectrum data to obtain the target spectrum allocation parameters corresponding to the target spectrum data. Then, the spectrum management platform can perform data conversion processing on the target spectrum allocation parameters to obtain the data that the target resource allocation model can process, thereby realizing the dynamic interaction between the target reinforcement learning model and the target resource allocation model.

[0120] In one exemplary embodiment, step S406 can be implemented in the following manner:

[0121] By using the target normalization function, the target spectrum allocation probability is mapped to obtain the target allocation priority corresponding to each frequency band; the sum of the target allocation priorities corresponding to each frequency band is the target constant.

[0122] In this embodiment, the target normalization function can be used to normalize the target spectrum allocation probability. The target allocation priority refers to the allocation priority of each frequency band to different users, which determines the allocation direction of spectrum resources in the target resource allocation model. The target constant can be a constant 1, meaning the sum of the target allocation priorities of each frequency band to different users is 1.

[0123] Specifically, after the target reinforcement learning model outputs the target spectrum allocation parameters, a cascaded target resource allocation model is performed for secondary optimization, achieving synergistic optimization of algorithmic decision-making and value allocation. This is done for the target spectrum allocation probability p. b.u The spectrum management platform can map the target spectrum allocation probability using a target normalization function to obtain the target allocation priority. For example, the normalization of the target spectrum allocation probability can be achieved using the following formula (16):

[0124]

[0125] In the above formula (16), Assign priorities to targets, p b,u The target spectrum allocation probability is output by the target reinforcement learning model. In this embodiment, the spectrum management platform can convert the original Softmax probability value into a comparable relative competitive advantage value between frequency bands through the target normalization function, thereby obtaining the target allocation priority. This allows the target spectrum allocation probabilities of different frequency bands to participate in spectrum allocation on a unified scale, realizing the data connection between the target reinforcement learning model and the target resource allocation model.

[0126] Step S407: Map the target power parameters to obtain the target constraint power corresponding to the target power parameters.

[0127] In one exemplary embodiment, step S407 can be implemented in the following manner:

[0128] By using the target mapping function, the target power parameters are mapped to the safe power range of the target terminal, thus obtaining the target constrained power corresponding to the target power parameters.

[0129] In this embodiment of the application, due to the physical constraints of the target satellite power amplifier, the safe operating range is typically [P]. min P sat To ensure that the target satellite operates within a safe operating range, the spectrum management platform can use a target mapping function to map the target power parameters output by the target reinforcement learning model. A mapping transformation is performed to obtain the target constrained power.

[0130] Specifically, the spectrum management platform can use the sigmoid function, a nonlinear mapping function, as the target mapping function, and apply it to the target power parameters output by the target reinforcement learning model. Converting the power to a safe, executable value for the target satellite yields the target constrained power, ensuring that the transmission power is controlled within a safe threshold and avoiding hardware overload or interference overflow that might occur from direct power output. For example, the target constrained power can be determined using the following formula (17):

[0131]

[0132] In the above formulas (17) and (18), where σ(x) is the Sigmoid function, P min For minimum transmission power, P sat The target power is saturated. The target mapping function can be a nonlinear mapping, whose smooth transition characteristics can prevent communication quality fluctuations caused by power abrupt changes, ensuring the safety and reliability of power control in practical scenarios.

[0133] In this embodiment, the spectrum management platform normalizes the target spectrum allocation probability to obtain the target allocation priority, which determines the subsequent spectrum resource allocation direction and ensures the accuracy of spectrum resource allocation. The target dynamic value coefficient can be used to generate a dynamic benchmark value. The target power parameter and the dynamic benchmark value can jointly constitute the core variables of the user valuation function. The dynamic spectrum allocation system optimizes the allocation according to the dual objectives of communication quality and value efficiency, giving priority to allocating spectrum resources to users with high valuations and meeting service quality requirements, thereby achieving an adaptive balance between communication quality and spectrum resource allocation efficiency.

[0134] Step S408: Calculate the user valuation for each user in different frequency bands based on the target allocation priority, target spectrum data, and terminal demand data.

[0135] In one exemplary embodiment, step S408 can be implemented in the following manner:

[0136] The target channel gain is determined based on satellite status parameters, environmental parameters, and user mobility parameters; the target signal-to-noise ratio is calculated based on the target channel gain, actual transmit power, actual bandwidth, and noise power spectral density; and user estimates for each user in different frequency bands are calculated based on the target signal-to-noise ratio, target allocation priority, and user requested bandwidth.

[0137] In this embodiment, user valuation can refer to the estimated value of different users in different frequency bands. This user valuation can be calculated based on target allocation priority, target spectrum data, and terminal demand data. The user valuation v u (b) It can be calculated using the following formulas (19), (20) and (21):

[0138]

[0139] In the above formulas (19), (20), and (21), v u (b) Estimated user value for user u in frequency band b. The target signal-to-noise ratio (SNR) is the target SNR of user u in frequency band b. This target SNR comes from the target state space and can be calculated based on real-time measurements. W is the actual transmission power. b Where N is the actual bandwidth and N0 is the noise power spectral density. ξ is the target channel gain; k is a constant, dist is the distance between the user and the satellite, which can be determined based on the satellite state parameters. t Δf is the weather attenuation coefficient in the environmental parameters. t This allows users to move parameters or perform Doppler offsets.

[0140] As shown in formula (21) above, the target channel gain is explicitly derived based on real-time parameters in the target state space, rather than directly using a single measurement value. This enables dynamic modeling of channel quality and allows for the integration of multi-dimensional real-time data such as satellite state parameters, environmental parameters, and user mobility parameters to accurately reflect the dynamic changes in channel gain, avoiding the one-sidedness of relying solely on spectrum sensor measurements. Furthermore, in formula (20) above, by incorporating parameters such as actual transmit power, noise power spectral density, and actual bandwidth into the calculation of the target signal-to-noise ratio, the user estimate v is made more accurate. u (b) It can simultaneously reflect the physical layer communication quality (e.g., signal strength) and resource occupation cost, integrate multi-source dynamic information to improve the quality of modeling channels, support the accuracy of spectrum estimation and algorithm adaptability, and provide a scientific and reasonable spectrum physical value assessment benchmark for target resource allocation models.

[0141] Step S409: Determine the dynamic benchmark value of each frequency band based on user valuation, historical value data, and target dynamic value coefficient.

[0142] In this embodiment, the dynamic benchmark value (or dynamic benchmark price) serves as a core parameter in the target resource allocation model, used to balance market stability and real-time environmental adaptability in spectrum allocation. The dynamic benchmark value is essentially a dynamically adjusted reference value for spectrum resources, reflecting both the average value of historical transactions and undergoing real-time correction based on current channel conditions (e.g., signal-to-noise ratio, user demand), ensuring that spectrum allocation conforms to market principles while rapidly responding to changes in the communication environment. The dynamic benchmark value for each frequency band b can be expressed as price. b The spectrum management platform can use a target resource allocation model to allocate resources based on user valuation (v). u (b) Historical value data and target dynamic value coefficient are obtained through α t The calculation can be performed using the following formula (22):

[0143]

[0144] In the above formula (22), MA b This represents the historical average value of frequency band b. The target resource allocation model can be achieved through the target dynamic value coefficient α. t To achieve a balance between responsiveness to real-time demands and historical stability, when the target dynamic value coefficient α... t When the value approaches 1, the dynamic benchmark value focuses on the current user value and responds to real-time channel conditions to adapt to sudden interference; when the target dynamic value coefficient α... tWhen the value approaches 1, the dynamic benchmark value focuses on the historical average value, maintaining historical stability and reducing user fluctuations. This avoids the problem in related technologies where static allocation cannot adapt to traffic fluctuations. The embodiments of this application calculate the dynamic benchmark mechanism using user valuation, historical value data, and a target dynamic value coefficient. This integrates real-time value assessment with historical market trends, ensuring that the dynamic benchmark price reflects both the current use value of the spectrum (e.g., the communication quality requirements of high-priority services) and conforms to long-term market patterns, avoiding frequent value fluctuations.

[0145] Step S410: Determine the actual valuation of the user in different frequency bands based on terminal demand data and dynamic benchmark value.

[0146] In one exemplary embodiment, step S410 can be implemented in the following manner:

[0147] The actual valuation of a user in different frequency bands is determined based on the dynamic benchmark value, the user's preset value factor, and the user's business priority weight; the user's business priority weight is determined based on the urgency of the user's business type.

[0148] In this embodiment, the actual valuation can refer to a valuation determined based on the user's actual needs, or it can be called the actual quote. User u can submit terminal data according to their own business needs, which may include a user-preset value factor β. u The target resource allocation model can calculate the actual valuation of a user based on dynamic benchmark value, user-preset value factors, and user business priority weights. For example, this actual valuation can be calculated using the following formula (23):

[0149]

[0150] In the above formula (23), bid u (b) represents the actual estimated value of user u in frequency band b. The user service priority weight can be determined based on the urgency of the user's service type. For example, the user service priority weight for urgent service types can be 2.0, while the user service priority weight for ordinary service types can be 1.0. This service priority weight can be dynamically assigned based on the user's real-time needs using the PPO2 algorithm. β u The value factor preset for users can be determined based on terminal demand data. It is a value reported by users based on actual needs. The value range of the user-preset value factor can be from 0.8 to 1.2. This application embodiment does not limit this.

[0151] In this embodiment, the target resource allocation model can determine the actual valuation of a user in different frequency bands based on dynamic benchmark value, user-preset value factor, and user service priority weight. This enables the actual valuation of the user to reflect the intensity of the user's real demand and ensures the efficiency of spectrum resource allocation.

[0152] Subsequently, during the spectrum allocation process, if the actual estimated value of user u is less than the dynamic benchmark value, i.e., the bid... u (b) <price b If the user's valuation of frequency band b is lower than the system benchmark, no allocation may be made. However, spectrum resource allocation may be made in scenarios where there is surplus spectrum resources and / or in situations where benchmark communication needs to be guaranteed. If multiple users have actual valuations for the same frequency band, the target resource allocation model may prioritize allocating to the user with the highest ratio of actual valuation to dynamic benchmark value to ensure maximum spectrum utilization efficiency.

[0153] Step S411: Based on user valuation, dynamic benchmark value, actual valuation, and target constraint power, determine the target spectrum allocation result through the target resource allocation model.

[0154] In one exemplary embodiment, step S411 can be implemented in the following manner:

[0155] Based on the target constrained power, determine the target power constraint conditions corresponding to the target resource allocation model; determine the bandwidth constraint conditions and frequency band allocation constraint conditions corresponding to the target resource allocation model; based on the target power constraint conditions, bandwidth constraint conditions, and frequency band allocation constraint conditions, calculate the total spectrum resource estimate based on user estimates, dynamic benchmark values, and actual estimates; if the total spectrum resource estimate meets the preset conditions, determine the target spectrum allocation result corresponding to the target spectrum data; the target spectrum allocation result includes the target spectrum allocation matrix, the target value vector, and the target transmit power.

[0156] In this embodiment, after determining the user valuation, dynamic benchmark value, actual valuation, and target constraint power, the spectrum management platform can determine the target spectrum allocation result corresponding to the target spectrum data through the target resource allocation model. In the target resource allocation model, the spectrum allocation task can be represented as a spectrum allocation optimization model, specifically the Hungarian algorithm model, simulated annealing algorithm model, or particle swarm optimization algorithm model, etc. When the total valuation of the spectrum resources reaches its maximum, the spectrum allocation optimization model can determine that the total valuation of the spectrum resources meets the preset conditions. The objective function of the spectrum allocation optimization model can be expressed by the following formula (24):

[0157]

[0158] In the above formula (24), the objective function aims to maximize the total spectrum value (i.e., the total estimated value of spectrum resources). b,u x is a binary variable representing whether frequency band b has been allocated to user u; if so, then x... b,u If x is 1, then x is 1; otherwise, x is 1. b,u The objective function's constraints can include target power constraints, bandwidth constraints, and frequency band allocation constraints. The target power constraint is... That is, the actual transmit power of frequency band b is less than the maximum allowable transmit power in the target constraint power. In this way, the power control of the target terminal is realized based on the target power parameter output by the target reinforcement learning model, so as to ensure that the transmit power of the target satellite is within the safe range. This represents the minimum bandwidth requirement for user u. Bandwidth constraints are used to ensure that the user's minimum bandwidth requirement is met, guaranteeing the user's basic Quality of Service (QoS) parameters and ensuring the user's communication needs. Frequency band allocation constraints can mean that each frequency band b can be allocated to at most one user, i.e., x. b,u A value ≤1 represents an exclusive constraint on frequency band resources, ensuring that spectrum resources are not repeatedly allocated to multiple users, thereby avoiding mutual interference and resource waste between frequency bands. Thus, the objective resource allocation model, through this objective function, ensures that limited spectrum resources are allocated to suitable and needy users, thereby improving overall spectrum resource utilization efficiency and overall system performance.

[0159] In this embodiment, the spectrum management platform combines a target reinforcement learning model with a target resource allocation model. Users submit terminal demand data based on their communication needs. The target reinforcement learning model determines the target spectrum allocation parameters based on the target spectrum data. The target resource allocation model generates a dynamic benchmark value and the user's actual valuation. The target resource allocation model can sort the actual valuations from high to low and allocate spectrum resources based on multiple constraints, which can improve spectrum utilization.

[0160] Based on the above objective function, in order to enhance market stability, a value smoothing component can be introduced into the objective function of the objective resource allocation model. The objective function based on value smoothing can be expressed by the following formula (25):

[0161]

[0162]

[0163] In the above formula (25), the objective function introduces a value efficiency weight λ and a value bias to achieve dual-objective optimization of communication quality and spectrum utilization efficiency. The value efficiency weight λ ranges from 0 to 1, and can use a default value (e.g., 0.3) or be dynamically adjusted according to the actual scenario through a target reinforcement learning model. For example, in emergency communication scenarios, λ tends to 0 to ensure basic service quality, while in daily commercial scenarios, λ can tend to 0.3 to balance spectrum utilization and overall spectrum value. The objective function also introduces a value smoothing term λ·(bid). u (b)-price b This approach enhances allocation robustness and is suitable for scenarios with rapidly changing satellite channel conditions. Specifically, the target resource allocation model can be implemented using the Hungarian algorithm. This algorithm models spectrum allocation as a bipartite graph maximum weight matching problem and uses the aforementioned objective function to solve for the optimal allocation scheme. The dynamic benchmark price is used as a reference. b As a dynamic benchmark for spectrum value, it reflects historical transaction patterns and adapts to real-time channel conditions. The value efficiency weight λ can be optimized online by the target reinforcement learning model algorithm, enabling the spectrum management platform to adaptively switch between different modes that prioritize communication quality and spectrum utilization efficiency. This avoids the limitations of relying solely on communication indicators for spectrum resource allocation in related technologies and can improve the utilization efficiency of spectrum resources.

[0164] During the value payment phase, the spectrum management platform can adopt the Vickrey-Clarke-Groves (VCG) honest pricing payment mechanism. For user u who has obtained spectrum resource allocation, the payment value of the user can be expressed by the following formula (26):

[0165]

[0166] In the above formula (26), the payment value of user u is equal to the reduction in the total value of other users caused by their participation in spectrum allocation. Of course, other methods of value payment can also be used in actual scenarios, and this application embodiment does not limit this.

[0167] Step S412: Adjust the spectrum parameters of the target terminal according to the target spectrum allocation result, and collect the channel quality data of the target terminal; generate a target feedback signal according to the target spectrum allocation result and the channel quality data; update the target reinforcement learning model and the target resource allocation model according to the target feedback signal.

[0168] In this embodiment of the application, the target spectrum allocation result may include a target spectrum allocation matrix, a target value vector, and a target transmit power. The target spectrum allocation matrix... It can be a binary matrix representing whether frequency band b is allocated to user u; the target value vector can refer to the user's actual payment value vector. Channel quality data can refer to the channel quality after the target spectrum allocation result is executed, or it can be called the target spectrum parameters of the next state, which can be used as the input of the next state of the target reinforcement learning model.

[0169] Specifically, after obtaining the target spectrum allocation result through the target reinforcement learning model and the target resource allocation model, the spectrum management platform can collect channel quality data after the target spectrum allocation result is executed using spectrum sensors and other means. Then, it can generate a target feedback signal based on the target spectrum allocation result and the channel quality data, and update the target reinforcement learning model and the target resource allocation model based on this target feedback signal. For example, based on this target feedback signal, the actual payment value vector of this spectrum allocation task can be added to the historical value data, which can be used to update the historical average value MA in the target resource allocation model. b This ensures the real-time calculation of dynamic benchmark value. Based on the target feedback signal, the policy parameters in the target reinforcement learning model can be adjusted and optimized. Specifically, these parameters can be stored in the experience pool of the PPO2 algorithm to assist in algorithm optimization. For example, if network congestion occurs in certain areas after a certain allocation, the target reinforcement learning model can be adjusted accordingly, allocating more spectrum resources to these congested areas in subsequent spectrum resource allocation processes. In this way, the spectrum management platform generates target feedback signals based on the target spectrum allocation results and channel quality data to optimize and update the target reinforcement learning model and the target resource allocation model. This allows the algorithm to continuously learn and improve, ensuring accurate and efficient subsequent prediction and allocation, and meeting actual communication needs.

[0170] Based on the above embodiments, Figure 7 This is a schematic diagram illustrating the execution flow of a target resource allocation model according to an embodiment of this application. Figure 7As shown, the output of the target reinforcement learning model is the target spectrum allocation probability, target power parameter, and target dynamic value coefficient. The target resource allocation model first transforms the target spectrum allocation probability and target power parameter. Specifically, it maps the target spectrum allocation probability using a target normalization function to obtain the target allocation priority, ensuring fairness in spectrum resource allocation. It also maps the target power parameter to a safe power range using a target mapping function to obtain the target constrained power. Next, the target resource allocation model calculates user estimates based on the target allocation priority, target spectrum data, and terminal demand data. Then, it combines user estimates, historical value data, and the target dynamic value coefficient to determine the dynamic benchmark value for each frequency band. Finally, it determines the actual user estimates in different frequency bands based on terminal demand data and the dynamic benchmark value. Then, using the Hungarian algorithm as the solution engine, it optimizes the allocation based on an objective function that includes communication quality and spectrum utilization efficiency, satisfying constraints such as target power constraints (power limitation), bandwidth constraints (QoS guarantee), and frequency band allocation constraints (frequency band exclusivity), generating the target spectrum allocation result. After executing the target spectrum allocation, channel quality data can be collected. Based on this data and the target spectrum allocation results, a target feedback signal is generated. This feedback signal then assists in the iterative optimization of the target reinforcement learning model and the target resource allocation model. Feedback is then fed back to the algorithm for iterative optimization, ensuring that spectrum resources are optimally allocated through market mechanisms.

[0171] Based on the above embodiments, Figure 8 This is a schematic diagram illustrating the interactive process of a dynamic spectrum allocation method according to an embodiment of this application. Figure 8 As shown, this dynamic spectrum allocation method is applied to a dynamic spectrum allocation system, which includes a target base station, a target satellite, a data processing center, a spectrum allocation algorithm engine, and user terminals. The target base station is a ground-based base station, including a ground spectrum sensor, and the target satellite includes a satellite spectrum sensor. The target base station and target satellite can act as execution devices to execute the target spectrum allocation results and dynamically adjust spectrum parameters. The data processing center and the spectrum allocation algorithm engine can form a spectrum management platform. The spectrum allocation algorithm engine can include a target reinforcement learning model (i.e., the PPO2 algorithm layer), a target resource allocation model (i.e., the auction algorithm layer), and an experience pool. This experience pool can be a distributed data storage module used to store state-action-reward-next state interaction data, supporting the online learning of the PPO2 algorithm.

[0172] exist Figure 8The overall architecture of the dynamic spectrum allocation method is divided into an input layer, a data processing layer, a PPO2 algorithm layer, an auction algorithm layer, and an output and feedback layer. The input layer collects raw spectrum data through spectrum sensors (including ground-based and satellite spectrum sensors) and simultaneously collects terminal demand data from user terminals. The data processing layer cleans, filters, and extracts features from the raw spectrum data through a data processing center, obtaining target spectrum data which is then input to the PPO2 algorithm layer. The PPO2 algorithm layer generates spectrum allocation probabilities through a policy network, evaluates state values ​​through a value network, and iteratively trains and optimizes the policy using an experience pool, ultimately outputting the target spectrum allocation parameters. The auction algorithm layer receives the target spectrum allocation parameters output by the PPO2 algorithm layer, performs normalization, power mapping, and dynamic benchmark value generation, determines the actual user valuation based on the terminal demand data, and then uses the Hungarian algorithm solver embedded in the auction algorithm to perform bipartite graph maximum weight matching to calculate the target spectrum allocation result. The output and feedback layer can perform spectrum allocation. Specifically, it can send the target spectrum allocation result to the execution device through the satellite command link or 5G core network to adjust the frequency band allocation and transmit power. Furthermore, it can collect channel quality data (such as signal-to-noise ratio, latency, etc.) and generate a target feedback signal in combination with the target spectrum allocation result. This signal is then fed back to the spectrum management platform in real time through the reverse link to update the experience pool of the PPO2 algorithm, store training data, trigger policy updates for the PPO2 algorithm, and simultaneously update the auction algorithm. This forms a closed-loop system of "perception-decision-execution-optimization" to achieve dynamic and intelligent management of spectrum resources.

[0173] It should be noted that in emergency communication scenarios, the spectrum management platform can switch to emergency mode, directly allocating idle frequency bands to user terminals in specific areas (such as disaster areas) through a target resource allocation model, without performing a value calculation process. Simultaneously, the target reinforcement learning model dynamically adjusts the weights of the target reward function, prioritizing QoS indicators (e.g., increasing the weight ω3 corresponding to the fairness indicator parameter). In daily commercial scenarios, the spectrum management platform achieves efficient allocation of spectrum resources through dynamic benchmark value and actual user valuations.

[0174] Based on the above embodiments, Figure 9 This is a schematic diagram illustrating the execution flow of a dynamic spectrum allocation method according to an embodiment of this application. Figure 9As shown, raw spectrum data is collected through a spectrum sensor, and terminal demand data is collected through a user terminal. After data preprocessing, target spectrum data is obtained. This target spectrum data and terminal demand data are input into a target reinforcement learning model. The target reinforcement learning model determines real-time dynamic parameters by constructing a dynamic decision model, performs spectrum classification prediction, and obtains target spectrum allocation parameters. The target spectrum classification parameters output by the target reinforcement learning model are used in a target resource allocation model. After data transformation, user valuation calculation, dynamic benchmark value calculation, and user actual valuation calculation, the target spectrum allocation result is determined based on the Hungarian algorithm. Thus, in this embodiment, the dynamic spectrum allocation method ensures the dynamic adaptability of spectrum resource classification and spectrum utilization efficiency through the dynamic combination of the target reinforcement learning model and the target resource allocation model. It can achieve real-time adaptation to the highly dynamic satellite environment, such as dynamically adjusting the allocation strategy when the satellite orbit changes, and can also achieve multi-objective balance, such as prioritizing QoS in emergency scenarios. After executing the target spectrum allocation result, the model combination (target reinforcement learning model and target resource allocation model) is updated based on the target feedback signal including channel quality data and the target spectrum allocation result. This completes the entire process of "data input - algorithm decision - allocation execution - feedback optimization". The overall closed-loop process enables the model combination to dynamically iterate according to the actual execution strategy and communication quality, continuously improving spectrum allocation efficiency and system robustness.

[0175] The dynamic spectrum allocation method in this application combines a target reinforcement learning model and a target resource allocation model to achieve dynamic management and efficient utilization of spectrum resources. First, this dynamic spectrum allocation method effectively solves the resource waste problem caused by static spectrum allocation in related technologies, especially in remote areas and scenarios with large fluctuations in user density, significantly improving spectrum resource utilization efficiency. Second, the multi-mode adaptive mechanism (routine commercial allocation / emergency direct allocation) can quickly respond to changes in network demand, reduce resource allocation latency, shorten the response time of critical services, significantly improve the efficiency and fairness of satellite network spectrum resource allocation, effectively increase spectrum utilization, and significantly improve communication system performance and user experience.

[0176] Furthermore, this dynamic spectrum allocation method, based on this model combination, can achieve cross-network standard interference avoidance and resource coordination in the air-space-ground integrated communication scenario, ensuring that limited spectrum is allocated to users with "high demand and reasonable actual valuation," thereby improving overall utilization efficiency. The model combination achieves scenario-based evolution through a "learning-allocation-feedback-optimization" closed-loop mechanism combined with multi-objective balance design. The model combination can dynamically adjust the strategy network parameters based on the target spectrum allocation results and channel quality data, forming a continuous optimization loop. This allows the system to accurately match actual communication needs in complex scenarios such as satellite orbit changes and disaster response. Specifically, the improved PPO2 algorithm generates target spectrum allocation probabilities through spatiotemporal feature extraction, directly serving as the target allocation priority for the auction algorithm. The dynamic benchmark mechanism fed back from the auction mechanism, along with the target spectrum allocation results, in turn optimizes the target reward function of the strategy network, forming a closed-loop optimization. This deep collaboration enables an adaptive balance between optimal communication quality and optimal spectrum utilization efficiency, improving the effective allocation rate of spectrum resources and user fairness indicators. In multi-user concurrent scenarios, it can ensure the service quality for edge users.

[0177] In practical applications, in daily communication scenarios, the spectrum management platform achieves fair allocation of spectrum resources by using dynamic benchmark value and actual user valuation, thereby improving the resource reuse efficiency between 5G base stations and satellites. In emergency communication scenarios, portable ground stations can be quickly deployed, the spectrum management platform switches to "emergency mode," the PPO2 algorithm prioritizes QoS for users in disaster areas, and the auction algorithm directly allocates idle frequency bands to ensure the real-time transmission of rescue instructions.

[0178] In this embodiment, the dynamic spectrum allocation method can be applied to various scenarios. For example, in the scenario of low-Earth orbit satellite constellation and terrestrial network collaboration, it can realize dynamic access and efficient allocation of spectrum resources during satellite movement, solving the problem that traditional static allocation cannot adapt to the high dynamic characteristics of satellites. When satellites cover high-density areas, the spectrum management platform can guide strategy optimization, automatically increase the allocation weight of edge users, optimize spectrum utilization, effectively reduce co-channel interference in urban environments, and realize dynamic resource allocation in high-dynamic scenarios.

[0179] For integrated air-space-ground communication scenarios, addressing the collaborative needs of low-Earth orbit satellites and terrestrial networks, the spectrum management platform leverages algorithmic coordination to balance the resource demands of different network nodes, avoiding the efficiency and fairness imbalances inherent in traditional static allocation. In emergency communication and disaster response scenarios, by dynamically adjusting service priority weights (e.g., Emergency Communication Weight 2.0), adaptively adjusting the learning rate based on user volatility, and prioritizing QoS indicators through a target reward function and final reward, the platform ensures uninterrupted emergency communication even when satellite nodes malfunction. The target resource allocation model switches to an emergency direct allocation mode, skipping the conventional value calculation process and directly allocating idle frequency bands based on the target spectrum allocation parameters output by the target reinforcement learning model, while combining VCG value payment rules to suppress non-critical service usage. In this way, the spectrum management platform can significantly reduce resource scheduling latency in disaster scenarios, improving the reliability of critical services such as rescue, command, and dispatch.

[0180] For drone communication and satellite IoT scenarios, the target state space can be compatible with parameters such as drone movement trajectory and IoT device density. Real-time drone services are dynamically scheduled based on actual estimates, avoiding spectrum fragmentation caused by IoT device access and improving the response speed of high-latency-sensitive drone services, thus meeting diverse communication needs. Of course, the dynamic spectrum allocation method in this embodiment can also be applied to other scenarios, and the specific choice can be based on actual needs; this embodiment does not limit this application.

[0181] It should be understood that in the various embodiments of this application, the sequence number of each process and step does not imply the order of execution. The execution order of each process and step should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0182] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0183] This embodiment also provides a dynamic spectrum allocation device for implementing the above embodiments and implementation methods; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated. The dynamic spectrum allocation device includes:

[0184] The acquisition module is used to acquire raw spectrum data and terminal demand data, and determine the target spectrum data corresponding to the raw spectrum data;

[0185] The processing module is used to process the target spectrum data and terminal demand data through a target reinforcement learning model to obtain the target spectrum allocation parameters corresponding to the target spectrum data; the target spectrum allocation parameters include the target spectrum allocation probability, the target power parameter, and the target dynamic value coefficient;

[0186] The determination module is used to determine the target spectrum allocation result based on the target spectrum data, terminal demand data, and target spectrum allocation parameters using the target resource allocation model.

[0187] In one exemplary embodiment, the processing module is specifically used for:

[0188] Based on target spectrum data, terminal demand data, and historical spectrum data, a dynamic decision-making model corresponding to the target reinforcement learning model is constructed.

[0189] Determine the real-time dynamic parameters of the target reinforcement learning model;

[0190] Based on a dynamic decision-making model and real-time dynamic parameters, the target spectrum data is processed through a target reinforcement learning model to obtain target spectrum allocation parameters.

[0191] In one exemplary embodiment, the processing module is specifically used for:

[0192] Based on the target spectrum data, the target state space corresponding to the dynamic decision-making model is determined; the target state space includes: spectrum state parameters, terminal demand data, satellite state parameters, historical spectrum parameters, and environmental parameters;

[0193] Determine the target action space corresponding to the dynamic decision-making model; the target action space includes spectrum allocation probability, power parameters, and dynamic value coefficients.

[0194] Determine the objective reward function, objective state transition function, and discount factor corresponding to the dynamic decision-making model;

[0195] A dynamic decision-making model is constructed based on the target state space, target action space, target reward function, target state transition function, and discount factor.

[0196] In one exemplary embodiment, the processing module is specifically used for:

[0197] Based on the bandwidth data and spectrum allocation results at the target time, determine the spectrum utilization rate at the target time;

[0198] Based on the actual interference intensity of each frequency band at the target time and the preset interference intensity, determine the interference level parameters at the target time;

[0199] The fairness index parameters for the target time are determined based on the service quality level and the number of users at the target time; the service quality level is determined based on the actual allocated bandwidth, the bandwidth demanded by users, and the user weight coefficient.

[0200] Based on the service quality level, service quality threshold, actual transmit power, and power threshold at the target time, determine the abnormal penalty parameters at the target time;

[0201] The target reward function is determined based on spectrum utilization, interference level parameters, fairness index parameters, and anomaly penalty parameters.

[0202] In one exemplary embodiment, the processing module is specifically used for:

[0203] Based on historical spectrum data and historical action data within the historical period, determine the historical mapping relationship between historical state, historical executed action, and the next historical state;

[0204] Based on the historical mapping relationship, determine the state transition frequency corresponding to different historical execution actions;

[0205] Based on the state transition frequency, the state transition probability within the historical period is determined, and the target state transition function is obtained.

[0206] In one exemplary embodiment, the processing module is specifically used for:

[0207] Within the initial time period, the initial parameters of the target reinforcement learning model are determined, and these initial parameters are used as real-time dynamic parameters.

[0208] Within non-initial time periods, the real-time dynamic parameters of the target reinforcement learning model are determined based on the initial parameters and the fluctuation parameters of the spectral data.

[0209] In one exemplary embodiment, the apparatus is further configured to:

[0210] If the spectrum allocation task meets the preset termination conditions, determine the action trajectory corresponding to the spectrum allocation task.

[0211] Based on the action trajectory and the spectrum utilization rate, interference level parameters, and fairness index parameters of the spectrum allocation task at different target times, the final reward corresponding to the spectrum allocation task is determined.

[0212] Update the network parameters of the objective reinforcement learning model based on the final reward.

[0213] In one exemplary embodiment, the determining module is specifically used for:

[0214] The target spectrum allocation probability is normalized to obtain the target allocation priority for each frequency band;

[0215] The target power parameters are mapped to obtain the target constraint power corresponding to the target power parameters.

[0216] Based on the target allocation priority, target spectrum data, and terminal demand data, calculate the user valuation of each user in different frequency bands;

[0217] Based on user valuations, historical value data, and target dynamic value coefficients, determine the dynamic benchmark value for each frequency band;

[0218] Based on terminal demand data and dynamic benchmark value, the actual valuation of users in different frequency bands is determined;

[0219] Based on user valuation, dynamic benchmark value, actual valuation, and target constraint power, the target spectrum allocation result is determined through the target resource allocation model.

[0220] In one exemplary embodiment, the determining module is specifically used for:

[0221] By using the target normalization function, the target spectrum allocation probability is mapped to obtain the target allocation priority corresponding to each frequency band; the sum of the target allocation priorities corresponding to each frequency band is the target constant.

[0222] In one exemplary embodiment, the determining module is specifically used for:

[0223] By using the target mapping function, the target power parameters are mapped to the safe power range of the target terminal, thus obtaining the target constrained power corresponding to the target power parameters.

[0224] In one exemplary embodiment, the determining module is specifically used for:

[0225] The target channel gain is determined based on satellite status parameters, environmental parameters, and user mobility parameters.

[0226] Calculate the target signal-to-noise ratio based on the target channel gain, actual transmit power, actual bandwidth, and noise power spectral density;

[0227] Based on the target signal-to-noise ratio, target allocation priority, and user request bandwidth, calculate the user valuation for each user in different frequency bands.

[0228] In one exemplary embodiment, the determining module is specifically used for:

[0229] The actual valuation of a user in different frequency bands is determined based on the dynamic benchmark value, the user's preset value factor, and the user's business priority weight; the user's business priority weight is determined based on the urgency of the user's business type.

[0230] In one exemplary embodiment, the determining module is specifically used for:

[0231] Based on the target constrained power, determine the target power constraint conditions corresponding to the target resource allocation model;

[0232] Determine the bandwidth constraints and frequency band allocation constraints corresponding to the target resource allocation model;

[0233] Based on the target power constraints, bandwidth constraints, and frequency band allocation constraints, the total spectrum resource valuation is calculated based on user estimates, dynamic benchmark values, and actual estimates.

[0234] If the total estimated value of spectrum resources meets the preset conditions, the target spectrum allocation result corresponding to the target spectrum data is determined; the target spectrum allocation result includes the target spectrum allocation matrix, the target value vector, and the target transmit power.

[0235] In one exemplary embodiment, the apparatus is further configured to:

[0236] The target terminal's spectrum parameters are adjusted based on the target spectrum allocation results, and the target terminal's channel quality data is collected.

[0237] Based on the target spectrum allocation results and channel quality data, a target feedback signal is generated;

[0238] Based on the target feedback signal, update the target reinforcement learning model and the target resource allocation model.

[0239] Figure 10 This is a schematic diagram of a dynamic spectrum allocation system according to an embodiment of this application. Figure 10 As shown, the dynamic spectrum allocation system 100 includes a target satellite 1001, a target base station 1002, a user terminal 1003, and a spectrum management platform 1004.

[0240] The target satellite 1001 and the target base station 1002 acquire raw spectrum data, and the user terminal 1003 acquires terminal demand data.

[0241] The spectrum management platform 1004 determines the target spectrum data corresponding to the original spectrum data;

[0242] The spectrum management platform 1004 processes target spectrum data and terminal demand data through a target reinforcement learning model to obtain target spectrum allocation parameters corresponding to the target spectrum data; the target spectrum allocation parameters include target spectrum allocation probability, target power parameter and target dynamic value coefficient;

[0243] The spectrum management platform 1004 determines the target spectrum allocation result based on the target spectrum data, terminal demand data, and target spectrum allocation parameters through the target resource allocation model.

[0244] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.

[0245] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when run.

[0246] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), computer hard disk, magnetic disk, or optical disk.

[0247] Figure 11 This is a structural block diagram of an electronic device according to an embodiment of this application, such as... Figure 11 As shown, embodiments of this application also provide an electronic device 110, including a memory 1101 and a processor 1102. The memory 1101 stores a computer program, and the processor 1102 is configured to run the computer program to perform the steps in any of the above-described embodiments of the dynamic spectrum allocation method.

[0248] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0249] Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.

[0250] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.

[0251] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. A dynamic spectrum allocation method, characterized in that, include: Acquire raw spectrum data and terminal demand data, and determine the target spectrum data corresponding to the raw spectrum data; The target spectrum data and the terminal demand data are processed using a target reinforcement learning model to obtain target spectrum allocation parameters corresponding to the target spectrum data; the target spectrum allocation parameters include target spectrum allocation probability, target power parameter, and target dynamic value coefficient. The target spectrum allocation result is determined using the target resource allocation model based on the target spectrum data, the terminal demand data, and the target spectrum allocation parameters.

2. The method according to claim 1, characterized in that, The step of processing the target spectrum data and the terminal demand data using a target reinforcement learning model to obtain the target spectrum allocation parameters corresponding to the target spectrum data includes: Based on the target spectrum data, the terminal demand data, and historical spectrum data, a dynamic decision model corresponding to the target reinforcement learning model is constructed. Determine the real-time dynamic parameters of the target reinforcement learning model; Based on the dynamic decision-making model and the real-time dynamic parameters, the target spectrum data is processed by the target reinforcement learning model to obtain the target spectrum allocation parameters.

3. The method according to claim 2, characterized in that, The step of constructing a dynamic decision model corresponding to the target reinforcement learning model based on the target spectrum data, the terminal demand data, and historical spectrum data includes: Based on the target spectrum data, the target state space corresponding to the dynamic decision-making model is determined; the target state space includes: spectrum state parameters, terminal demand data, satellite state parameters, historical spectrum parameters, and environmental parameters; Determine the target action space corresponding to the dynamic decision-making model; the target action space includes spectrum allocation probability, power parameters, and dynamic value coefficients. Determine the target reward function, target state transition function, and discount factor corresponding to the dynamic decision-making model; The dynamic decision-making model is constructed based on the target state space, the target action space, the target reward function, the target state transition function, and the discount factor.

4. The method according to claim 2, characterized in that, The construction of the dynamic decision model corresponding to the target reinforcement learning model includes determining the target reward function corresponding to the dynamic decision model, wherein determining the target reward function corresponding to the dynamic decision model includes: Based on the bandwidth data and spectrum allocation results at the target time, the spectrum utilization rate at the target time is determined; The interference level parameters for the target time are determined based on the actual interference intensity of each frequency band at the target time and the preset interference intensity. The fairness index parameters for the target time are determined based on the service quality level and the number of users at the target time; the service quality level is determined based on the actual allocated bandwidth, the bandwidth demanded by users, and the user weight coefficient. Based on the service quality level, service quality threshold, actual transmit power, and power threshold at the target time, the abnormality penalty parameters for the target time are determined. The target reward function is determined based on the spectrum utilization rate, the interference level parameter, the fairness index parameter, and the anomaly penalty parameter.

5. The method according to claim 2, characterized in that, The construction of the dynamic decision model corresponding to the target reinforcement learning model includes determining the target state transition function corresponding to the dynamic decision model, wherein determining the target state transition function corresponding to the dynamic decision model includes: Based on historical spectrum data and historical action data within the historical period, determine the historical mapping relationship between historical states, historical actions, and the next historical state; Based on the historical mapping relationship, determine the state transition frequency corresponding to different historical execution actions; Based on the state transition frequency, the state transition probability within the historical period is determined, and the target state transition function is obtained.

6. The method according to claim 2, characterized in that, Determining the real-time dynamic parameters of the target reinforcement learning model includes: Within the initial time period, the initial parameters of the target reinforcement learning model are determined, and the initial parameters are used as the real-time dynamic parameters; During non-initial time periods, the real-time dynamic parameters of the target reinforcement learning model are determined based on the initial parameters and the spectral data fluctuation parameters.

7. The method according to claim 2, characterized in that, The method further includes: If the spectrum allocation task meets the preset termination conditions, the action trajectory corresponding to the spectrum allocation task is determined. Based on the action trajectory and the spectrum utilization rate, interference level parameters, and fairness index parameters of the spectrum allocation task at different target times, the final reward corresponding to the spectrum allocation task is determined. The network parameters of the target reinforcement learning model are updated based on the final reward.

8. The method according to any one of claims 1 to 7, characterized in that, The step of determining the target spectrum allocation result through the target resource allocation model, based on the target spectrum data, the terminal demand data, and the target spectrum allocation parameters, includes: The target spectrum allocation probability is normalized to obtain the target allocation priority for each frequency band; The target power parameters are mapped to obtain the target constraint power corresponding to the target power parameters; Based on the target allocation priority, the target spectrum data, and the terminal demand data, calculate the user valuation for each user in different frequency bands; Based on the user valuation, historical value data, and the target dynamic value coefficient, the dynamic benchmark value of each frequency band is determined; Based on the terminal demand data and the dynamic benchmark value, the actual estimated value of the user in different frequency bands is determined; Based on the user valuation, the dynamic benchmark value, the actual valuation, and the target constraint power, the target spectrum allocation result is determined through the target resource allocation model.

9. The method according to claim 8, characterized in that, The normalization process for the target spectrum allocation probability to obtain the target allocation priority for each frequency band includes: The target spectrum allocation probability is mapped using a target normalization function to obtain the target allocation priority corresponding to each frequency band; the sum of the target allocation priorities corresponding to each frequency band is the target constant.

10. The method according to claim 8, characterized in that, The mapping process for the target power parameters to obtain the target constraint power corresponding to the target power parameters includes: By using a target mapping function, the target power parameter is mapped to the safe power range of the target terminal, thereby obtaining the target constraint power corresponding to the target power parameter.

11. The method according to claim 8, characterized in that, The step of calculating the user valuation for each user in different frequency bands based on the target allocation priority, the target spectrum data, and the terminal demand data includes: The target channel gain is determined based on satellite status parameters, environmental parameters, and user mobility parameters. The target signal-to-noise ratio is calculated based on the target channel gain, actual transmit power, actual bandwidth, and noise power spectral density. Based on the target signal-to-noise ratio, the target allocation priority, and the user request bandwidth, calculate the user estimate for each user in different frequency bands.

12. The method according to claim 8, characterized in that, The step of determining the actual estimated value of a user in different frequency bands based on the terminal demand data and the dynamic benchmark value includes: The actual valuation of the user in different frequency bands is determined based on the dynamic benchmark value, the user preset value factor, and the user service priority weight; the user service priority weight is determined according to the urgency of the user service type.

13. The method according to claim 8, characterized in that, The step of determining the target spectrum allocation result based on the user valuation, the dynamic benchmark value, the actual valuation, and the target constraint power, through the target resource allocation model, includes: Based on the target constraint power, determine the target power constraint conditions corresponding to the target resource allocation model; Determine the bandwidth constraints and frequency band allocation constraints corresponding to the target resource allocation model; Based on the target power constraint, the bandwidth constraint, and the frequency band allocation constraint, the total spectrum resource estimate is calculated based on the user estimate, the dynamic benchmark value, and the actual estimate. If the total estimated value of the spectrum resources meets the preset conditions, the target spectrum allocation result corresponding to the target spectrum data is determined; the target spectrum allocation result includes the target spectrum allocation matrix, the target value vector, and the target transmit power.

14. The method according to claim 1, characterized in that, The method further includes: The target terminal's spectrum parameters are adjusted based on the target spectrum allocation results, and the channel quality data of the target terminal is collected. Based on the target spectrum allocation result and the channel quality data, a target feedback signal is generated; Based on the target feedback signal, update the target reinforcement learning model and the target resource allocation model.

15. A dynamic spectrum allocation system, characterized in that, This includes target satellites, target base stations, user terminals, and a spectrum management platform; The target satellite and the target base station acquire raw spectrum data, and the user terminal acquires terminal demand data. The spectrum management platform determines the target spectrum data corresponding to the original spectrum data; The spectrum management platform processes the target spectrum data and the terminal demand data using a target reinforcement learning model to obtain target spectrum allocation parameters corresponding to the target spectrum data; the target spectrum allocation parameters include target spectrum allocation probability, target power parameter, and target dynamic value coefficient. The spectrum management platform determines the target spectrum allocation result based on the target spectrum data, the terminal demand data, and the target spectrum allocation parameters through a target resource allocation model.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the dynamic spectrum allocation method according to any one of claims 1 to 14.

17. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the dynamic spectrum allocation method according to any one of claims 1 to 16.

18. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the dynamic spectrum allocation method according to any one of claims 1 to 14.

Citation Information

Cited By

  • Spectrum resource dynamic allocation method and device based on multi-protocol fusion, computer equipment and storage medium

    CN121357546A