A method, apparatus, network device and computer storage medium for allocating

By optimizing spectrum resource allocation using the pre-trained A3C algorithm and Actor-Critic model, the problem of low spectrum resource utilization is solved, achieving more efficient spectrum resource management and reducing conflicts, thereby improving the spectrum resource utilization of wireless systems.

CN115843111BActive Publication Date: 2026-02-03CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111051872.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-08
Publication Date
2026-02-03
Estimated Expiration
2041-09-08

AI Technical Summary

Technical Problem

Existing spectrum allocation methods result in low spectrum utilization, especially in unlicensed bands for industrial, scientific, and medical applications where resources are congested and idle bands frequently appear.

Method used

The pre-trained A3C algorithm is used to receive data packet transmission requests from terminal devices. Combined with the spectrum resource information of the wireless system where the terminal device is located, spectrum resources are determined for data packets from idle spectrum resources. The spectrum resources are allocated according to their priority and conflict status. The Actor-Critic model is used to update network parameters to optimize spectrum resource allocation.

Benefits of technology

It improves the accuracy and utilization of spectrum resources, reduces the probability of conflicts between cognitive users, and enhances the spectrum resource utilization efficiency of wireless systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115843111B_ABST
    Figure CN115843111B_ABST
Patent Text Reader

Abstract

The application discloses a kind of allocation methods, the method includes: receiving terminal equipment for the sending request of data packet, data packet and the information of the frequency spectrum resource in the wireless system where terminal equipment is located is input to pre-trained A3C algorithm, determine the frequency spectrum resource from the idle frequency spectrum resource of wireless system for data packet, and the frequency spectrum resource is allocated to terminal equipment.The embodiment of the application also discloses a kind of allocation device, network equipment and computer storage medium, improve the allocation accuracy of frequency spectrum resource in wireless system, and then improve the utilization of frequency spectrum resource in wireless system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to spectrum resource allocation technology, and more particularly to an allocation method, apparatus, network device, and computer storage medium. Background Technology

[0002] Currently, unlicensed frequency bands for industrial, scientific, and medical applications are extremely congested, while a large number of frequency bands are often idle. This indicates that existing spectrum resource allocation methods suffer from the technical problem of low spectrum resource utilization. Summary of the Invention

[0003] In view of this, the present invention provides an allocation method, apparatus, network device, and computer storage medium to solve the technical problem of low utilization rate of spectrum resources in the prior art.

[0004] The technical solution of this invention is implemented as follows:

[0005] In a first aspect, embodiments of the present invention provide an allocation method, comprising:

[0006] Receive terminal device's request to send data packets;

[0007] The data packet and the spectrum resource information of the wireless system in which the terminal device is located are input into the pre-trained A3C algorithm to determine the spectrum resource for the data packet from the idle spectrum resources of the wireless system;

[0008] The spectrum resources are allocated to the terminal device; wherein the spectrum resources are used by the terminal device to access and send the data packets;

[0009] The trained A3C algorithm was obtained using the following method:

[0010] Obtain information about the data packets to be trained and the spectrum resources in the wireless system where the data packets to be trained are located;

[0011] The data packet to be trained and the spectrum resource information of the wireless system in which the data packet to be trained are located are input into the preset A3C algorithm for training, so as to obtain the trained A3C algorithm.

[0012] In the above method, after inputting the data packet and the spectrum resource information of the wireless system in which the terminal device is located into the pre-trained A3C algorithm to determine the spectrum resource for the data packet from the idle spectrum resources of the wireless system, the method further includes:

[0013] When the spectrum resource does not conflict with the first spectrum resource, the spectrum resource is allocated to the terminal device; wherein, the first spectrum resource includes spectrum resources determined for data packets of other terminal devices;

[0014] When the spectrum resource conflicts with the first spectrum resource, it is determined whether to allocate the spectrum resource to the terminal device based on the priority of the spectrum resource recorded in advance.

[0015] In the above method, the method further includes:

[0016] When the reward function value in the trained A3C algorithm is a first preset threshold, it is determined that the spectrum resource conflicts with the first spectrum resource.

[0017] When the reward function value in the trained A3C algorithm is the second preset threshold, it is determined that the spectrum resource does not conflict with the first spectrum resource.

[0018] In the above method, determining whether to allocate the spectrum resource to the data packet based on the pre-recorded priority of the spectrum resource includes:

[0019] The priority of the terminal device and the priorities of the other terminal devices are determined from the pre-recorded priorities of the spectrum resources.

[0020] When the priority of the terminal device is higher than that of the other terminal devices, the spectrum resources are allocated to the terminal device, and the priority of the terminal device is incremented by one.

[0021] When the priority of the terminal device is lower than or equal to the priority of the other terminal devices, the process returns to the step of inputting the data packet and the spectrum resource information of the wireless system in which the terminal device is located into the pre-trained A3C algorithm, determining the spectrum resource for the data packet from the idle spectrum resources of the wireless system, and incrementing the priority of the other terminal devices by one.

[0022] In the above method, the step of inputting the data packet to be trained and the spectrum resource information of the wireless system in which the data packet to be trained is located into a preset A3C algorithm for training to obtain the trained A3C algorithm includes:

[0023] Based on the data packet to be trained and the spectrum resource information of the wireless system in which the data packet to be trained is located, the Actor network in the A3C algorithm is trained according to the reward function in the preset A3C algorithm to update the network parameters of the Actor network in the A3C algorithm.

[0024] Furthermore, based on the pre-defined advantage function in the A3C algorithm, the Critic network in the A3C algorithm is trained to update the network parameters of the Critic network in the A3C algorithm, thereby obtaining the trained A3C algorithm.

[0025] In the above method, the value of the reward function is positively correlated with the percentage of the noise ratio of the terminal device when using spectrum resources relative to the total noise ratio of the wireless system.

[0026] In the above method, the reward function is as follows:

[0027]

[0028] Where R represents the function value of the reward function, a and b are constants, and p i This represents the percentage of the signal-to-noise ratio of spectrum resource i to the total signal-to-noise ratio of the wireless system.

[0029] In a second aspect, the present invention provides a dispensing device, comprising:

[0030] The receiving module is used to receive the terminal device's request to send data packets;

[0031] The determination module is used to input the data packet and the spectrum resource information of the wireless system in which the terminal device is located into the pre-trained A3C algorithm to determine the spectrum resource for the data packet from the idle spectrum resources of the wireless system;

[0032] An allocation module is used to allocate the spectrum resources to the terminal device; wherein the spectrum resources are used by the terminal device to access and send the data packets;

[0033] The device is further configured to obtain the trained A3C algorithm in the following manner:

[0034] Obtain information about the data packets to be trained and the spectrum resources in the wireless system where the data packets to be trained are located;

[0035] The data packet to be trained and the spectrum resource information of the wireless system in which the data packet to be trained are located are input into the preset A3C algorithm for training, so as to obtain the trained A3C algorithm.

[0036] In the above-described device, the device is also used for:

[0037] After inputting the data packet and the spectrum resource information of the wireless system in which the terminal device is located into the pre-trained A3C algorithm, and determining the spectrum resource for the data packet from the idle spectrum resources of the wireless system, the spectrum resource is allocated to the terminal device when there is no conflict between the spectrum resource and a first spectrum resource; wherein, the first spectrum resource includes the spectrum resource determined for the data packets of other terminal devices;

[0038] When the spectrum resource conflicts with the first spectrum resource, it is determined whether to allocate the spectrum resource to the terminal device based on the priority of the spectrum resource recorded in advance.

[0039] In the above-described device, the device is also used for:

[0040] When the reward function value in the trained A3C algorithm is a first preset threshold, it is determined that the spectrum resource conflicts with the first spectrum resource.

[0041] When the reward function value in the trained A3C algorithm is the second preset threshold, it is determined that the spectrum resource does not conflict with the first spectrum resource.

[0042] In the above-described apparatus, determining whether to allocate the spectrum resources to the data packet based on the pre-recorded priority of the spectrum resources includes:

[0043] The priority of the terminal device and the priorities of the other terminal devices are determined from the pre-recorded priorities of the spectrum resources.

[0044] When the priority of the terminal device is higher than that of the other terminal devices, the spectrum resources are allocated to the terminal device, and the priority of the terminal device is incremented by one.

[0045] When the priority of the terminal device is lower than or equal to the priority of the other terminal devices, the process returns to the step of inputting the data packet and the spectrum resource information of the wireless system in which the terminal device is located into the pre-trained A3C algorithm, determining the spectrum resource for the data packet from the idle spectrum resources of the wireless system, and incrementing the priority of the other terminal devices by one.

[0046] In the above-described device, the rotor inputs the data packet to be trained and the spectrum resource information of the wireless system in which the data packet is located into a preset A3C algorithm for training, resulting in a trained A3C algorithm that includes:

[0047] Based on the data packet to be trained and the spectrum resource information of the wireless system in which the data packet to be trained is located, the Actor network in the A3C algorithm is trained according to the reward function in the preset A3C algorithm to update the network parameters of the Actor network in the A3C algorithm.

[0048] Furthermore, based on the pre-defined advantage function in the A3C algorithm, the Critic network in the A3C algorithm is trained to update the network parameters of the Critic network in the A3C algorithm, thereby obtaining the trained A3C algorithm.

[0049] In the above-described device, the value of the reward function is positively correlated with the percentage of the noise ratio of the terminal device when using spectrum resources relative to the total noise ratio of the wireless system.

[0050] In the above-described device, the reward function is as follows:

[0051]

[0052] Where R represents the function value of the reward function, a and b are constants, and p i This represents the percentage of the signal-to-noise ratio of spectrum resource i to the total signal-to-noise ratio of the wireless system.

[0053] Thirdly, embodiments of the present invention also provide a network device, the network device comprising: a processor and a storage medium storing processor-executable instructions, the storage medium performing operations via a communication bus dependent on the processor, and when the instructions are executed by the processor, performing the allocation method described in one or more of the above embodiments.

[0054] Fourthly, embodiments of the present invention provide a computer storage medium storing executable instructions, wherein when the executable instructions are executed by one or more processors, the processors execute the allocation method described in one or more of the above embodiments.

[0055] The present invention provides an allocation method, apparatus, network device, and computer storage medium. The method includes: receiving a data packet transmission request from a terminal device; inputting the data packet and information about spectrum resources in the wireless system where the terminal device is located into a pre-trained A3C algorithm; determining spectrum resources for the data packet from idle spectrum resources in the wireless system; and allocating the spectrum resources to the terminal device. The spectrum resources are used by the terminal device to access and transmit data packets. The pre-trained A3C algorithm is obtained by: acquiring the data packet to be trained and information about spectrum resources in the wireless system where the data packet is located; inputting the data packet to be trained and information about spectrum resources in the wireless system where the data packet is located into the pre-trained A3C algorithm for training; and obtaining the trained A3C algorithm. In other words, in this embodiment of the invention, the method... The trained A3C algorithm is used to allocate spectrum resources for data packets that terminal devices need to send. Here, the trained A3C algorithm can allocate appropriate spectrum resources from idle spectrum resources for each terminal device that needs to send data packets. Since the trained A3C algorithm is obtained by training on the data packets to be trained and the spectrum resource information of the wireless system in which the data packets are located, the A3C algorithm can be trained using the data packets to be trained from different terminal devices and the different spectrum resource information of the wireless system in which the data packets to be trained from different terminal devices are located. This optimizes the original A3C algorithm, enabling the trained A3C algorithm to allocate appropriate spectrum resources for data packets from different terminal devices, thereby improving the accuracy of spectrum resource allocation and thus improving the utilization rate of spectrum resources in the wireless system. Attached Figure Description

[0056] Figure 1 This is a flowchart illustrating an optional allocation method in an embodiment of the present invention;

[0057] Figure 2 This is a schematic diagram of the cognitive radio network structure in related technologies;

[0058] Figure 3 This is a schematic diagram of the spectrum environment in a cognitive radio system in related technologies;

[0059] Figure 4 This is a schematic diagram of the structure of the Actor-Critic model in related technologies;

[0060] Figure 5 This is a network architecture diagram of the Actor-Critic model in related technologies;

[0061] Figure 6 This is a schematic diagram of an optional architecture for training the A3C algorithm, provided as an embodiment of the present invention.

[0062] Figure 7 A flowchart illustrating an example of an optional A3C algorithm training method provided in an embodiment of the present invention;

[0063] Figure 8 A schematic diagram of an optional dispensing device provided in an embodiment of the present invention;

[0064] Figure 9 This is a schematic diagram of the structure of an optional network device provided in an embodiment of the present invention. Detailed Implementation

[0065] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0066] Example 1

[0067] This invention provides a method for allocating spectrum resources. Figure 1 This is a flowchart illustrating an optional allocation method in an embodiment of the present invention, such as... Figure 1 As shown, the allocation method may include:

[0068] S101: Receive a request from the terminal device to send a data packet;

[0069] In recent decades, communication technology has flourished globally, with the number of communication users increasing year by year and people's demand for communication gradually growing. However, the speed and bandwidth of wireless communication are related, and the limited availability of wireless resources has become a key factor restricting the development of wireless communication.

[0070] Unlicensed frequency bands for industrial, scientific, and medical use are extremely congested, while a large number of primary user frequency bands are frequently idle. Therefore, the rational utilization of spectrum resources is a crucial breakthrough in addressing the current spectrum scarcity. The development of Cognitive Radio (CR) technology has become key to improving spectrum utilization and realizing the "secondary use" of spectrum resources.

[0071] As is well known, the spectrum environment of cognitive radio is dynamic. In a real wireless communication environment, there may be multiple primary users and multiple cognitive users (equivalent to the terminal devices mentioned above and other terminal devices described below). When cognitive users select an idle channel without prior knowledge of the environment, they try to avoid conflicts with other cognitive users. Figure 2 This is a schematic diagram of the cognitive radio network structure in related technologies, such as... Figure 2As shown, the cognitive radio network structure includes: cognitive users, primary users, and network devices. It is assumed that there are K cognitive users, N primary users, and N licensed channels. Primary users have priority in using spectrum resources at any time period, and cognitive users can only access the network when primary users are not occupying the spectrum.

[0072] Figure 3 This is a schematic diagram of the spectrum environment in a cognitive radio system in related technologies, such as... Figure 3 As shown, channel 1 corresponds to primary user 1, channel 2 to primary user 2, channel 3 to primary user 3, and channel N to primary user N. The boxes filled with cross lines represent spectrum resources occupied by primary users. Cognitive users and primary users are in the same wireless system. Cognitive users need to seize spectrum opportunities when channels are idle and intelligently access and use those channels. Each round of exploration by a cognitive user is divided into equal-length time slots. Users maintain time slot synchronization, and data packets are divided into lengths that can be transmitted within a single time slot. In each time slot, the channel occupancy status of the primary user is represented by y. p This indicates that when the primary user needs to send data packets, y p =1, when the primary user has no need to send data packets. p =0, while all cognitive users always have the need to send data packets. Each cognitive user has the ability to learn and make decisions independently, and uses algorithms to select and try to access channels with a higher signal-to-noise ratio. At the end of each time slot, each cognitive user will receive a binary observation acknowledgment character (ACK) and the corresponding channel's reward value.

[0073] As can be seen, in a wireless system, the primary user does not always occupy the corresponding channel (equivalent to the aforementioned spectrum resources). Therefore, there are idle spectrum resources in the wireless system. This embodiment of the invention provides an optional spectrum resource allocation method, mainly for allocating spectrum resources for the terminal devices of cognitive users. When the terminal device needs to send a data packet, the terminal device sends a data packet transmission request to the network device. At this time, the network device receives the data packet transmission request from the terminal device and needs to allocate spectrum resources for the data packet to be sent by the terminal device.

[0074] S102: Input the information of the data packet and the spectrum resources of the wireless system in which the terminal device is located into the pre-trained A3C algorithm to determine the spectrum resources for the data packet from the idle spectrum resources of the wireless system;

[0075] Specifically, after obtaining the data packet that the terminal device needs to send, the data packet and the spectrum resource information of the wireless system in which the terminal device is located are both input into the A3C algorithm. That is to say, the embodiment of the present invention uses the A3C algorithm to determine the spectrum resource for the data packet from the idle spectrum resources of the wireless system. The Asynchronous Advantage Actor-Critic (A3C) algorithm is based on the Actor-Critic model framework.

[0076] Figure 4 This is a schematic diagram of the structure of the Actor-Critic model in related technologies, such as... Figure 4 As shown, the Actor-Critic model is a combination of the value function method and the policy gradient method. The Actor-Critic model can both overcome the problem that policy gradients struggle to determine the quality of actions and expand the application scope of value function-based methods. The Actor-Critic model's core consists of two parts: the Actor (policy network) action module and the Critic (value function) evaluation module.

[0077] Figure 5 The network architecture diagram of the Actor-Critic model in related technologies, such as Figure 5 As shown, the agent obtains an action and executes it under the Actor network, obtains a reward function from the interaction with the environment and gets a new state. At this time, the Critic network evaluates the action and calculates the temporal difference error (TD) of the action as the evaluation standard. Then, the TD is passed back to the Actor network. The Actor network will adjust itself in time according to the passed TD error. The probability of correct actions will increase in the subsequent iterations.

[0078] In order to allocate appropriate spectrum resources to cognitive users and improve the utilization rate of spectrum resources in the wireless system, in this embodiment of the invention, it is necessary to train a preset A3C algorithm. The trained A3C algorithm is obtained by training in the following manner:

[0079] Obtain information about the data packets to be trained and the spectrum resources in the wireless system where the data packets are located;

[0080] The data packets to be trained and the spectrum resource information of the wireless system in which the data packets are located are input into the preset A3C algorithm for training, and a trained A3C algorithm is obtained.

[0081] Specifically, training samples need to be obtained first. The training samples can include the data packets to be trained and the spectrum resources of the wireless system in which the data packets are located. It should be noted that during the sampling of training samples, multiple network environments are sampled for each cognitive user to increase the diversity of training samples. This makes the trained A3C algorithm more optimized and helps to improve the utilization rate of spectrum resources in the wireless system.

[0082] Figure 6 This invention provides an optional architecture diagram for training the A3C algorithm, as shown in the embodiment of the invention. Figure 6 As shown, a multi-threaded approach is used to allow simultaneous interaction and learning between multiple threads and the environment. Each thread summarizes and stores the learning results in a public network. Each thread periodically retrieves the collective learning results from the public network to guide its subsequent learning interactions with the environment.

[0083] Figure 6 The public network in the model is mainly a public neural network model. This neural network includes the functions of an Actor network and a Critic network. There are n threads (workers) under it. Each thread has the same network structure as the public neural network. Each thread will independently interact with the environment to obtain experience data.

[0084] In this process, each thread interacts with the environment to obtain a certain amount of data, and then calculates the gradient of the loss function of its own neural network. However, these gradients do not update the neural network within their own thread; instead, they update the shared neural network. In other words, n threads independently use accumulated gradients to update the parameters of the shared neural network model. Every so often, each thread updates its own neural network parameters to match the parameters of the shared neural network, thus guiding subsequent interactions with the environment.

[0085] It is evident that the network model in the common part is the model to be learned, while the network model in the thread is mainly used for interaction with the environment. These models in the thread can help the thread interact with the environment better, obtain high-quality data, and help the model converge faster.

[0086] It should be noted that the Actor-Critic model introduces an advantage function as the basis for evaluation, and the formula is defined as follows:

[0087] A(s,a)=Q(s,a)-V(s) (1)

[0088] The above formula expresses the merits or demerits of action a in state s. V(s) represents the average expected value. If the value of action a is higher than the average expected value, the advantage function is positive; otherwise, it is negative.

[0089] For the Critic network, its loss function is obtained using the squared error method:

[0090] L(θ v )=(r t +γV(S t+1 )-V(S t )) 2 (2)

[0091] Where L represents the loss function of the Critic network, t represents the time slot, γ represents the discount factor of the wireless system, r represents the reward function, and θ v These are the parameters of the Critic network, and their update is based on the following formula:

[0092]

[0093] Where α is the learning rate, and the time difference error TD = r t +γV(S t+1 )-V(S t To evaluate the quality of action a in state s, the loss function of the Actor network is calculated as follows:

[0094] L(θ π )=-∑log(π(s t ,a t ))*TD (4)

[0095] Where, θ π The parameters of the action network are updated using the following formula:

[0096]

[0097] By continuously updating θ v and θ π This optimizes the network, ultimately leading to the best action. Compared to the Actor-Critic model, the A3C algorithm goes a step further in sampling, using n-step sampling to accelerate convergence. Its advantage function expression changes as follows:

[0098] A(s,t)=r t +γr t+1 +...γ n-1 r t+n-1 +γ n V(s')-V(s) (6)

[0099] The Actor network parameters change as follows:

[0100]

[0101] The entropy term H of the strategy π is added, where c is a coefficient, which can prevent premature entry into a suboptimal strategy.

[0102] Thus, a more optimized A3C algorithm is obtained through the above training method. The spectrum resources determined by the trained A3C algorithm are more suitable, thereby improving the utilization rate of spectrum resources in the wireless system.

[0103] S103: Allocate spectrum resources to terminal devices;

[0104] Among them, spectrum resources are used for terminal devices to access and send data packets.

[0105] Finally, after the spectrum resources are determined for the data packets using the A3C algorithm, the determined spectrum resources are allocated to the terminal devices, enabling the terminal devices to access the spectrum resources and send data packets through the accessed spectrum resources.

[0106] In addition, after the spectrum resources for the data packet are determined, the determined spectrum resources can be directly allocated to the terminal device. Alternatively, the allocation of the determined spectrum resources to the terminal device can be determined according to preset rules. In an optional embodiment, after S102, the above method may further include:

[0107] When there is no conflict between the spectrum resource and the first spectrum resource, the spectrum resource will be allocated to the terminal device;

[0108] When a spectrum resource conflicts with the first spectrum resource, the system determines whether to allocate the spectrum resource to the terminal device based on the pre-recorded priority of the spectrum resource.

[0109] The first spectrum resource includes the spectrum resources determined for data packets from other terminal devices. In other words, it is first determined whether there is a conflict between the determined spectrum resources and the spectrum resources determined for data packets from other terminal devices. If there is no conflict, it means that the determined spectrum resources have not been contested by multiple cognitive users. Therefore, the determined spectrum resources are directly allocated to the terminal device, so that the terminal device can access the spectrum resources and send the data packets through the accessed spectrum resources.

[0110] Furthermore, if the determined spectrum resources conflict with those determined for data packets from other terminal devices, it indicates that the determined spectrum resources are being contested by multiple cognitive users. Therefore, in order to allocate the determined spectrum resources to more suitable cognitive users, the priority of the spectrum resources is pre-recorded. That is, each cognitive user has a corresponding access priority for each spectrum resource. Thus, the priority of the spectrum resources can be used to determine whether to allocate the spectrum resources to the terminal device.

[0111] To determine whether the identified spectrum resources conflict with spectrum resources identified for data packets from other terminal devices, in an optional embodiment, the method may further include:

[0112] When the reward function value in the trained A3C algorithm is the first preset threshold, it is determined that the spectrum resource conflicts with the first spectrum resource.

[0113] When the reward function value in the trained A3C algorithm is the second preset threshold, it is determined that there is no conflict between the spectrum resource and the first spectrum resource.

[0114] Here, the function value of the reward function in the trained A3C algorithm is used to determine whether there is a conflict between the spectrum resources determined for the data packets of this terminal device and the spectrum resources determined for the data packets of other terminal devices. Specifically, during the execution of actions by the Actor network in the trained A3C algorithm, it continuously interacts with the licensed channel and obtains the function value of the reward function given by the network environment. The function value of the reward function can reflect whether there is a conflict between the spectrum resources determined for the data packets of this terminal device and the spectrum resources determined for the data packets of other terminal devices.

[0115] Specifically, when the reward function value is a first preset threshold, the spectrum resources determined by the data packets for this terminal device conflict with the spectrum resources determined by the data packets for other terminal devices; when the reward function value is a second preset threshold, the spectrum resources determined by the data packets for this terminal device do not conflict with the spectrum resources determined by the data packets for other terminal devices.

[0116] In practical applications, during the cognitive user channel selection process, when multiple cognitive users simultaneously select the same channel, leading to conflicts, the system latency increases and the algorithm struggles to converge if these users remain in a contention phase. To effectively improve network performance and reduce conflicts caused by cognitive users selecting the same channel, this invention proposes a user priority memory method. A simple neural network memory is established to record the channel selection information for each time slot, and each channel is prioritized based on the number of times a cognitive user has selected it. When multiple cognitive users select the same channel, the user with the most historical channel selections is prioritized by calling the priority information in the memory.

[0117] At the end of each time slot, cognitive users are prioritized and their priorities are stored in the memory bank. The priorities w of all cognitive users are initialized. i,j =0. If cognitive user j successfully accesses channel i and receives a positive reward during this time slot, its priority for channel i is incremented by one; otherwise, the priority remains unchanged. When multiple cognitive users select channel i, the right to use it is determined by their priority. The cognitive user with the higher priority accesses and uses channel i, while all other cognitive users exit and wait for the next time slot to reselect. Here, the priority w of cognitive user j for channel i is... i,j It can be expressed by the following formula:

[0118]

[0119] To determine whether to allocate the identified spectrum resources to the terminal device, in one optional embodiment, determining whether to allocate the spectrum resources to the data packet based on the priority of the pre-recorded spectrum resources includes:

[0120] The priority of the terminal device is determined from the pre-recorded priority of spectrum resources, and the priority of other terminal devices is determined.

[0121] When the priority of a terminal device is higher than that of other terminal devices, spectrum resources are allocated to the terminal device, and the priority of the terminal device is incremented by one.

[0122] When the priority of a terminal device is lower than or equal to the priority of other terminal devices, the algorithm returns to execute the input of the data packet and the spectrum resource information of the wireless system in which the terminal device is located into the pre-trained A3C algorithm, determines the spectrum resource for the data packet from the idle spectrum resources of the wireless system, and increments the priority of other terminal devices by one.

[0123] Specifically, the pre-recorded priority of spectrum resources includes the priority of each terminal device for that spectrum resource. Then, the priority of the terminal device and the priorities of other conflicting terminal devices can be determined from the pre-recorded priority of spectrum resources. Then, the relationship between the priorities of the terminal device and the priorities of other terminal devices is compared. When the priority of the terminal device is higher than the priorities of other terminal devices, it means that the terminal device has accessed the spectrum resource more times than other terminal devices. Therefore, the spectrum resource is allocated to the terminal device, and the priority of the terminal device should be updated by incrementing the priority of the terminal device by one.

[0124] When the priority of a terminal device is lower than or equal to that of other terminal devices, it means that the terminal device has accessed the spectrum resource less often than other terminal devices. Therefore, the determined spectrum resource is allocated to other terminal devices with higher priority. The algorithm then returns to the pre-trained A3C algorithm, inputting the data packet and the spectrum resource information of the wireless system in which the terminal device is located. The algorithm then determines the spectrum resource for the data packet from the idle spectrum resources of the wireless system, i.e., reallocates the spectrum resource for the data packet of the terminal device. At the same time, the priorities of other terminal devices should be updated by incrementing their priorities by one.

[0125] For a preset A3C algorithm, in order to train the A3C algorithm, in one optional embodiment, the data packet to be trained and the spectrum resource information of the wireless system in which the data packet is located are input into the preset A3C algorithm for training to obtain a trained A3C algorithm, including:

[0126] Based on the data packets to be trained and the spectrum resources in the wireless system where the data packets are located, the Actor network in the A3C algorithm is trained according to the reward function in the preset A3C algorithm in order to update the network parameters of the Actor network in the A3C algorithm.

[0127] Furthermore, based on the pre-defined advantage function in the A3C algorithm, the Critic network in the A3C algorithm is trained to update the network parameters of the Critic network in the A3C algorithm, thus obtaining the trained A3C algorithm.

[0128] Specifically, in each sampling, the data packet to be trained and the spectrum resource information of the wireless system in which the data packet is located are obtained as a training sample and input into the preset A3C algorithm. The Actor network is trained and its network parameters are updated using the reward function in the preset A3C algorithm. Similarly, the Critic network is trained and its network parameters are updated using the same reward function in the preset A3C algorithm, thus obtaining a trained A3C algorithm. Likewise, the next sampled training sample is obtained, and the A3C algorithm is trained in the same way to update its network parameters. This training process is iterated until the required number of iterations is reached, resulting in a trained A3C algorithm.

[0129] In addition, regarding the reward function, in one optional embodiment, the function value of the reward function is positively correlated with the percentage of the noise ratio of the terminal device when using spectrum resources relative to the total noise ratio of the wireless system.

[0130] In this way, the original reward function is proportional to the percentage of the noise ratio of the spectrum resources used by the terminal device relative to the total noise ratio of the wireless system. This allows the reward function to reflect the noise of the spectrum resources used by the terminal device, thus using the noise of the spectrum resources used by the terminal device as feedback to train the A3C algorithm and optimize it. In an optional embodiment, the above reward function is as follows:

[0131]

[0132] Where R represents the value of the reward function, a and b are constants, and p i This represents the percentage of the signal-to-noise ratio of spectrum resource i to the total signal-to-noise ratio of the wireless system.

[0133] In other words, the noise ratio of spectrum resources is introduced into the reward function to explore the impact of the noise ratio of spectrum resources on cognitive users' selection of spectrum resources.

[0134] It should be noted that when a cognitive user and other cognitive users conflict due to selecting the same channel, the reward function is -3; when a cognitive user and other cognitive users do not select the same channel, the reward function is 1; in the reward function, "a·(p i -b)” indicates the impact of cognitive user usage of channel i on the reward function, p i Let be the percentage of the signal-to-noise ratio of channel i to the total signal-to-noise ratio of the system, and a and b are constants.

[0135] The following examples illustrate the training method of the A3C algorithm described in one or more of the above embodiments.

[0136] Figure 7A flowchart illustrating an example of an optional A3C algorithm training method provided in an embodiment of the present invention is shown below. Figure 7 As shown, the above training method may include:

[0137] S701: Initialize network parameters.

[0138] Specifically, the parameters θ of the Actor network and Critic network are initialized. π and θ v The discount factor γ, learning rate α and β, and exploration rate ε of the wireless system are set respectively. A state is arbitrarily selected from the state set as the current cognitive user's active state s, i.e., a channel is arbitrarily selected as the active channel.

[0139] S702: The Actor network executes actions.

[0140] Specifically, the Actor network interacts with the environment by inputting a state vector, selects the channel with the larger Q value for execution through an ε-greedy strategy, and adjusts the probability of an action in a timely manner through the action advantage function fed back by the Critic network.

[0141] S703: Critic Network Action Evaluation.

[0142] Specifically, the Critic network calculates the corresponding Q value based on the action output by the Actor network through the Q network and stores it in the Q network. At the same time, it calculates the advantage function of the action and feeds it back to the Actor network as a reference for action execution.

[0143] S704: Calculate the return R.

[0144] Specifically, during the execution of actions, the Actor network continuously interacts with the authorized channel and obtains reward values ​​from the environment. The signal-to-noise ratio (SNR) of the channel is introduced into the reward function R to explore the impact of the average SNR of the channel on the cognitive user's channel selection.

[0145] S705: Determine the priority of cognitive users.

[0146] Specifically, when multiple cognitive users have a conflict, the information in the user priority memory is called to prioritize the cognitive users in selecting channels, and the cognitive user with the higher priority will have the right to use the channel.

[0147] S706: Calculate the average capacity of the system.

[0148] The system capacity is defined as follows:

[0149] C = Blog2(1 + SNR) (10)

[0150] The average capacity of a system is defined as:

[0151]

[0152] Where c(i) is the system capacity when a cognitive user occupies channel i, and m is the number of times the statistical average capacity is calculated. It is assumed that a cognitive user will successfully access and use the channel and obtain system capacity only when a single cognitive user selects the channel. If two or more cognitive users select the same channel at the same time, all cognitive users will exit the current channel and wait for the next time slot to select again, while obtaining a channel capacity of 0.

[0153] S707: Q value update.

[0154] After the action is executed, the state is updated to a new state s', and the Q value is updated through the Critic network and stored in the Q neural network.

[0155] S708: Update the network parameters for Actor and Critic.

[0156] At the end of each iteration, the network parameters of the Actor network and the Critic network need to be updated using gradients.

[0157] S709: Update global network parameters.

[0158] At the end of each iteration, each thread updates the parameters of its own neural network to the parameters of the common neural network, thereby guiding subsequent environmental interactions.

[0159] In other words, in this example, the cognitive user inputs the channel state set S into the Actor network for processing. The Actor network provides the cognitive user with rewards for channel access actions and environmental feedback, and uses these actions as input to the Critic network. The Critic network calculates the Q-value of the action through the Q-network and stores it in the Q-neural network. Simultaneously, it calculates the dominance function of the action and feeds this dominance function back to the Actor network. The Actor network updates its parameters to increase the probability of actions with larger dominance functions being selected. At the end of each time slot, the cognitive user's channel selection information is recorded in the user priority memory. Through continuous updates to the parameters of the Actor and Critic networks, the Critic network evaluates all channel access actions output by the Actor network, collectively guiding the cognitive user to select an available channel. The A3C algorithm introduces asynchronous processing, which accelerates algorithm convergence while reducing the probability of conflicts between cognitive users.

[0160] Through the above examples, the main purpose of the dynamic spectrum access design based on the A3C algorithm is to obtain the mapping relationship between each state and action, enabling cognitive users to make optimal action decisions based on previously learned experience and knowledge during subsequent channel selection. A3C guides cognitive users to select the optimal channel access through an Actor network and a Critic network. Cognitive users input channel states into the Actor network, which provides rewards for the cognitive user's channel access actions and environmental feedback. The actions output by the Actor network are evaluated by the Critic's value function network, and the calculated Q-value is stored in a Q-neural network. Simultaneously, a user priority memory method is introduced to reduce conflicts caused by multiple cognitive users selecting the same channel. By continuously updating the parameters of the Actor and Critic networks, the Critic network evaluates all channel access actions output by the Actor network, increasing the probability of superior actions being selected. The A3C algorithm introduces asynchronous processing, which accelerates algorithm convergence while minimizing interference from cognitive users to the primary user during communication, and simultaneously improving the average capacity of the system.

[0161] As can be seen, in view of the problem that the real spectrum environment is complex, changeable and difficult to model, this invention proposes dynamic spectrum access based on the A3C algorithm, which guides cognitive users to select accurate idle channels by introducing two neural networks, Actor and Critic.

[0162] Cognitive users interact with the channel environment through the Actor network. Each access action receives a reward value from the channel. The signal-to-noise ratio of the channel is incorporated into the reward function. The Critic network calculates the Q value corresponding to each channel action and stores it in the Q network. At the same time, it evaluates the action through an internal dominance function, increasing the probability of the cognitive user accurately accessing the channel. By continuously updating the parameters of the Actor network and the Critic network, the optimal action is found to accurately access an idle channel.

[0163] The A3C algorithm introduces the concept of asynchronous training, which can effectively reduce the probability of conflicts between cognitive users while greatly improving the speed of system spectrum resource allocation.

[0164] This example addresses the challenge of modeling complex and unknown spectrum environments by introducing the A3C algorithm to guide cognitive users in intelligently selecting and accessing idle spectrum. The A3C algorithm combines reinforcement learning and deep learning, endowing cognitive users with the ability to self-learn about the spectrum environment. Through Actor and Critic networks, it guides cognitive users in channel selection, while asynchronous processing accelerates convergence, achieving efficient and rapid allocation of spectrum resources.

[0165] Compared to traditional spectrum allocation algorithms, the A3C algorithm guides cognitive users to continuously update their neural networks through interaction with the spectrum environment. This effectively improves the accuracy of spectrum allocation and reduces the probability of cognitive users accessing idle spectrum. Furthermore, the A3C algorithm effectively addresses the limitations of traditional algorithms when dealing with large numbers of cognitive users and licensed spectrum, significantly increasing the speed of spectrum allocation.

[0166] The present invention provides an allocation method comprising: receiving a data packet transmission request from a terminal device; inputting the data packet and information on spectrum resources in the wireless system where the terminal device is located into a pre-trained A3C algorithm; determining spectrum resources for the data packet from the idle spectrum resources of the wireless system; and allocating the spectrum resources to the terminal device. The spectrum resources are used by the terminal device to access and transmit data packets. The pre-trained A3C algorithm is obtained by: acquiring the data packet to be trained and information on the spectrum resources in the wireless system where the data packet is located; inputting the data packet to be trained and information on the spectrum resources in the wireless system where the data packet is located into the pre-trained A3C algorithm for training; and obtaining the trained A3C algorithm. In other words, in this embodiment of the invention, the trained A3C algorithm... The algorithm allocates spectrum resources to the data packets that terminal devices need to send. Here, the trained A3C algorithm can allocate appropriate spectrum resources from the idle spectrum resources for each terminal device that needs to send data packets. Since the trained A3C algorithm is trained using the data packets to be trained and the spectrum resource information of the wireless system in which the data packets are located, the A3C algorithm can be trained using the data packets to be trained from different terminal devices and the different spectrum resource information of the wireless system in which the data packets to be trained from different terminal devices are located. This optimizes the original A3C algorithm, enabling the trained A3C algorithm to allocate appropriate spectrum resources for the data packets of different terminal devices, thereby improving the accuracy of spectrum resource allocation and thus improving the utilization rate of spectrum resources in the wireless system.

[0167] Example 2

[0168] Based on the same inventive concept, embodiments of the present invention also provide a dispensing device. Figure 8 This is a schematic diagram of an optional dispensing device provided in an embodiment of the present invention, as shown below. Figure 8 As shown, the dispensing device may include:

[0169] The receiving module 81 is used to receive the terminal device's request to send data packets;

[0170] The determination module 82 is used to input the information of the spectrum resources in the wireless system where the data packet and the terminal device are located into the pre-trained A3C algorithm to determine the spectrum resources for the data packet from the idle spectrum resources of the wireless system.

[0171] The allocation module 83 is used to allocate spectrum resources to terminal devices; wherein the spectrum resources are used by the terminal devices to access and send data packets;

[0172] The device is also used to obtain the trained A3C algorithm in the following manner:

[0173] Obtain information about the data packets to be trained and the spectrum resources in the wireless system where the data packets are located;

[0174] The data packets to be trained and the spectrum resource information of the wireless system in which the data packets are located are input into the preset A3C algorithm for training, and a trained A3C algorithm is obtained.

[0175] In an optional embodiment, the above-described apparatus is further used for:

[0176] The information on the spectrum resources in the wireless system where the data packet and the terminal device are located is input into the pre-trained A3C algorithm. After determining the spectrum resources for the data packet from the idle spectrum resources of the wireless system, the spectrum resources are allocated to the terminal device when there is no conflict between the spectrum resources and the first spectrum resources. The first spectrum resources include the spectrum resources determined for the data packets of other terminal devices.

[0177] When a spectrum resource conflicts with a primary spectrum resource, the system determines whether to allocate the spectrum resource to the terminal device based on the pre-recorded priority of the spectrum resource.

[0178] In an optional embodiment, the above-described apparatus is further used for:

[0179] When the reward function value in the trained A3C algorithm is the first preset threshold, it is determined that the spectrum resource conflicts with the first spectrum resource.

[0180] When the reward function value in the trained A3C algorithm is the second preset threshold, it is determined that there is no conflict between the spectrum resource and the first spectrum resource.

[0181] In an optional embodiment, the above-described apparatus determines whether to allocate spectrum resources to data packets based on the pre-recorded priority of the spectrum resources, including:

[0182] The priority of the terminal device is determined from the pre-recorded priority of spectrum resources, and the priority of other terminal devices is determined.

[0183] When the priority of a terminal device is higher than that of other terminal devices, spectrum resources are allocated to the terminal device, and the priority of the terminal device is incremented by one.

[0184] When the priority of a terminal device is lower than or equal to the priority of other terminal devices, the algorithm returns to the previous state and inputs the data packet and the spectrum resource information of the wireless system in which the terminal device is located into the pre-trained A3C algorithm. The algorithm then determines the spectrum resource for the data packet from the idle spectrum resources of the wireless system and increments the priority of other terminal devices by one.

[0185] In an optional embodiment, the above-mentioned device inputs the data packet to be trained and the spectrum resource information of the wireless system in which the data packet to be trained is located into a preset A3C algorithm for training, and the trained A3C algorithm includes:

[0186] Based on the data packets to be trained and the spectrum resources in the wireless system where the data packets are located, the Actor network in the A3C algorithm is trained according to the reward function in the preset A3C algorithm in order to update the network parameters of the Actor network in the A3C algorithm.

[0187] Furthermore, based on the pre-defined advantage function in the A3C algorithm, the Critic network in the A3C algorithm is trained to update the network parameters of the Critic network in the A3C algorithm, thus obtaining the trained A3C algorithm.

[0188] In one alternative embodiment, the value of the reward function is positively correlated with the percentage of the noise ratio of the terminal device when using spectrum resources relative to the total noise ratio of the wireless system.

[0189] In an alternative embodiment, the reward function described above is as shown in formula (9).

[0190] In practical applications, the receiving module 81, determining module 82 and allocating module 83 can be implemented by a processor located on the network device, specifically a central processing unit (CPU), microprocessor unit (MPU), digital signal processor (DSP) or field programmable gate array (FPGA).

[0191] Figure 9 A schematic diagram of an optional network device provided in an embodiment of the present invention, such as... Figure 9 As shown, an embodiment of the present invention provides a network device 900, including:

[0192] The processor 91 and the storage medium 92 storing instructions executable by the processor 91 are provided. The storage medium 92 performs operations in dependence on the processor 91 via a communication bus 93. When the instructions are executed by the processor 91, the allocation method described in Embodiment 1 is executed.

[0193] It should be noted that in practical applications, the various components in the terminal are coupled together via the communication bus 93. It can be understood that the communication bus 93 is used to achieve communication between these components. In addition to the data bus, the communication bus 93 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 9 The general labeled all buses as communication bus 93.

[0194] This invention provides a computer storage medium storing executable instructions. When the executable instructions are executed by one or more processors, the processors execute the allocation method described in Embodiment 1.

[0195] The computer-readable storage medium can be a magnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.

[0196] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0197] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0198] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0199] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0200] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention.

Claims

1. An allocation method, characterized in that, include: Receive terminal device's request to send data packets; The data packet and the spectrum resource information of the wireless system in which the terminal device is located are input into the pre-trained A3C algorithm to determine the spectrum resource for the data packet from the idle spectrum resources of the wireless system; The spectrum resources are allocated to the terminal device; wherein the spectrum resources are used by the terminal device to access and send the data packets; The trained A3C algorithm was obtained using the following method: Obtain information about the data packets to be trained and the spectrum resources in the wireless system where the data packets to be trained are located; The data packet to be trained and the spectrum resource information of the wireless system in which the data packet to be trained is located are input into the preset A3C algorithm for training, so as to obtain the trained A3C algorithm. The method further includes, after inputting the data packet and the spectrum resource information of the wireless system in which the terminal device is located into a pre-trained A3C algorithm to determine the spectrum resource for the data packet from the idle spectrum resources of the wireless system, the method further includes: When the spectrum resource does not conflict with the first spectrum resource, the spectrum resource is allocated to the terminal device; wherein, the first spectrum resource includes spectrum resources determined for data packets of other terminal devices; When the spectrum resource conflicts with the first spectrum resource, it is determined whether to allocate the spectrum resource to the terminal device based on the priority of the spectrum resource recorded in advance. The method further includes: When the reward function value in the trained A3C algorithm is a first preset threshold, it is determined that the spectrum resource conflicts with the first spectrum resource. When the reward function value in the trained A3C algorithm is the second preset threshold, it is determined that the spectrum resource does not conflict with the first spectrum resource.

2. The method according to claim 1, characterized in that, The step of determining whether to allocate the spectrum resources to the data packet based on the pre-recorded priority of the spectrum resources includes: The priority of the terminal device and the priorities of the other terminal devices are determined from the pre-recorded priorities of the spectrum resources. When the priority of the terminal device is higher than that of the other terminal devices, the spectrum resources are allocated to the terminal device, and the priority of the terminal device is incremented by one. When the priority of the terminal device is lower than or equal to the priority of the other terminal devices, the process returns to the step of inputting the data packet and the spectrum resource information of the wireless system in which the terminal device is located into the pre-trained A3C algorithm, determining the spectrum resource for the data packet from the idle spectrum resources of the wireless system, and incrementing the priority of the other terminal devices by one.

3. The method according to claim 1 or 2, characterized in that, The step of inputting the data packet to be trained and the spectrum resource information of the wireless system in which the data packet to be trained is located into a preset A3C algorithm for training to obtain the trained A3C algorithm includes: Based on the data packet to be trained and the spectrum resource information of the wireless system in which the data packet to be trained is located, the Actor network in the A3C algorithm is trained according to the reward function in the preset A3C algorithm to update the network parameters of the Actor network in the A3C algorithm. Furthermore, based on the pre-defined advantage function in the A3C algorithm, the Critic network in the A3C algorithm is trained to update the network parameters of the Critic network in the A3C algorithm, thereby obtaining the trained A3C algorithm.

4. The method according to claim 3, characterized in that, The value of the reward function is positively correlated with the percentage of the noise ratio of the terminal device when using spectrum resources relative to the total noise ratio of the wireless system.

5. The method according to claim 4, characterized in that, The reward function is as follows: Where R represents the function value of the reward function, a and b are constants, and p i This represents the percentage of the signal-to-noise ratio of spectrum resource i to the total signal-to-noise ratio of the wireless system.

6. A dispensing device, characterized in that, include: The receiving module is used to receive the terminal device's request to send data packets; The determination module is used to input the data packet and the spectrum resource information of the wireless system in which the terminal device is located into the pre-trained A3C algorithm to determine the spectrum resource for the data packet from the idle spectrum resources of the wireless system; An allocation module is used to allocate the spectrum resources to the terminal device; wherein the spectrum resources are used by the terminal device to access and send the data packets; The device is further configured to obtain the trained A3C algorithm in the following manner: Obtain information about the data packets to be trained and the spectrum resources in the wireless system where the data packets to be trained are located; The data packet to be trained and the spectrum resource information of the wireless system in which the data packet to be trained is located are input into the preset A3C algorithm for training, so as to obtain the trained A3C algorithm. The device is also used for: After inputting the data packet and the spectrum resource information of the wireless system in which the terminal device is located into the pre-trained A3C algorithm, and determining the spectrum resource for the data packet from the idle spectrum resources of the wireless system, the spectrum resource is allocated to the terminal device when there is no conflict between the spectrum resource and the first spectrum resource; wherein, the first spectrum resource includes the spectrum resource determined for the data packets of other terminal devices; When the spectrum resource conflicts with the first spectrum resource, it is determined whether to allocate the spectrum resource to the terminal device based on the priority of the spectrum resource recorded in advance. The device is also used for: When the reward function value in the trained A3C algorithm is a first preset threshold, it is determined that the spectrum resource conflicts with the first spectrum resource. When the reward function value in the trained A3C algorithm is the second preset threshold, it is determined that the spectrum resource does not conflict with the first spectrum resource.

7. A network device, characterized in that, The network device includes: A processor and a storage medium storing processor-executable instructions, the storage medium performing operations via a communication bus dependent on the processor, wherein when the instructions are executed by the processor, the allocation method described in any one of claims 1 to 5 is performed.

8. A computer storage medium, characterized in that, The system stores executable instructions that, when executed by one or more processors, enable the allocation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Cognitive radio dynamic spectrum distribution method integrated with system overall transmission speed and distribution equity

    CN103491550A

  • Edge container resource allocation method based on deep reinforcement learning

    CN111885137A

  • Frequency spectrum resource management and distribution method based on federated learning

    CN113038616A