Resource scheduling method and apparatus, and device, medium and product
By scheduling resources based on channel state information in non-cellular access networks using discrete differential evolution and game theory algorithms, the problems of small coverage and poor communication quality in NAFD are solved, thereby maximizing system spectrum efficiency and effectively utilizing resources.
Patent Information
- Application Number
- PCT/CN2024/106221
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2024-07-18
- Publication Date
- 2025-12-04
AI Technical Summary
In Network Assisted Free-Duplex (NAFD), centralized processing results in small coverage and poor communication quality. The AP's operating mode cannot be flexibly expanded, leading to resource waste and interference. Furthermore, centralized processing of all data in the CPU results in high backhaul overhead.
Based on a pre-established network-assisted free-duplex architecture under a non-cellular access network, a discrete differential evolution algorithm is used to determine the association between edge distribution units and access points. A game theory algorithm is combined to determine the working mode of the access points, and a reinforcement learning algorithm is used to schedule data transmission resources to maximize the overall spectral efficiency of the system.
It improved the coverage and communication quality of the communication system, reduced resource waste and interference, optimized backhaul overhead, and achieved effective utilization of spectrum resources.
Smart Images

Figure CN2024106221_04122025_PF_FP_ABST
Abstract
Description
Resource scheduling methods, devices, equipment, media and products
[0001] This application claims priority to Chinese Patent Application No. 202410707417.5, filed with the Chinese Patent Office on May 31, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of communication technology, and in particular to a resource scheduling method, apparatus, device, medium and product. Background Technology
[0003] Future mobile communication networks need to meet service requirements such as global coverage, ultra-high-speed transmission, ultra-low latency, and ultra-reliable secure connections. In the development of communication technology, massive MIMO (Multi-Input Multi-Output) is an important technology that breaks through the two-dimensional resource limitations of "time-spectrum" and expands resources in the three-dimensional "space-time-frequency" space. To improve user quality, base stations in cellular networks are deployed more densely to reduce the distance between them and users, forming small cell networks. However, this introduces drawbacks such as edge hardening and increased inter-user interference. Cellular networks break away from the traditional cell concept of cellular networks, using distributed access points (APs) instead of base stations. Compared to traditional small cells, cellular massive MIMO can significantly improve system performance, achieving up to 95% throughput per user, and is unaffected by the spatial correlation of shadow fading. Cellular massive MIMO systems have fully centralized, semi-distributed, and fully distributed architectures, with fully distributed architectures reducing computational complexity and providing scalability. Previous research has shown that cellular networks can be implemented in a Cloud Radio Access Network (C-RAN) architecture, where baseband processing is moved from the AP to a central processor in the "cloud".
[0004] Time Division Duplex (TDD) and Frequency Division Duplex (FDD) are fixed-duplex methods widely used in 4G, while 5G and 6G require more flexible duplex modes to significantly improve system throughput. In-Band Full-Duplex (IBFD) technology allows base stations to simultaneously transmit and receive data within the same frequency band, offering a potential opportunity to improve the spectral efficiency of wireless networks, but also introducing self-interference issues. Simultaneously, Co-frequency Co-time Full Duplex (CCFD) can theoretically double the system's spectral efficiency, but in practice, the signal transmitted by the base station's downlink antenna can cause significant interference to the uplink antenna, resulting in performance loss. The Coordinated Multi-Point for In-band Full Duplex (CoMP flex) system uses two spatially separated and coordinated half-duplex (HD) base stations to simulate a full-duplex (FD) base station. Because the uplink and downlink antennas are geographically separated, the system's self-interference is greatly reduced, while also reducing the transceiver design complexity caused by interference cancellation.
[0005] However, with increasingly dense base station deployments, eliminating cross-link interference remains a challenge for architectures such as IBFD, CCFD, and CoMPflex. Cellular-free networks can effectively solve the problem of inter-cell link interference, among which Network-Assisted Full Duplex (NAFD) holds great potential. In NAFD, access points (APs) can operate in duplex modes such as CCFD, mixed duplex, and full duplex. In the same time slot, different APs can choose different modes for uplink reception and downlink transmission, thus satisfying simultaneous signal reception and transmission without the internal AP interference found in CCFD. However, traditional NAFD suffers from a lack of flexibility in expansion. The centralized processing of APs in NAFD leads to small coverage and poor communication quality, and the inflexible expansion of AP operating modes results in resource waste and interference.
[0006] Besides considering duplex mode selection, resolving the user association problem can also improve the system's spectral efficiency under limited resources. Related technologies propose the Dynamic Cooperation Clustering (DCC) method to construct a scalable, cellular-free massive MIMO architecture. The Hungarian algorithm can associate AP clusters with users based solely on AP location information. Although it has low complexity and minimal backhaul overhead, it cannot guarantee consistently optimal performance. Other user association algorithms also suffer from high complexity, over-reliance on user state or system performance, and cannot consistently achieve optimal performance.
[0007] Summary of the Invention
[0008] This application provides a resource scheduling method, apparatus, device, medium, and product to solve the technical problems in related technologies, such as the small coverage and poor communication quality caused by centralized processing of APs in NAFD, the inability of APs to flexibly expand their working modes leading to resource waste and interference, and the large backhaul consumption caused by centralized processing of all data in the CPU.
[0009] According to one aspect of this application, a resource scheduling method is provided, comprising:
[0010] Based on the pre-established channel model of the network-assisted free-duplex architecture under the non-cellular access network, the channel state information between each user equipment and the edge distribution unit, the uplink user equipment and the downlink user equipment, and the uplink access point and the downlink access point is obtained.
[0011] With the goal of maximizing the overall system spectral efficiency, a pre-configured discrete differential evolution algorithm is used to determine the association between edge distribution units and access points, and antenna resources are scheduled based on this association.
[0012] Furthermore, a game theory algorithm is used to determine the operating mode of each access point, so as to schedule time-frequency domain resources and spatial domain resources based on the operating mode.
[0013] Furthermore, the association status between edge distribution units and user equipment is determined based on backhaul capacity constraints and reinforcement learning algorithms, so as to schedule data transmission resources based on the association status; wherein, the total spectral efficiency of the system is determined based on the channel state information.
[0014] According to another aspect of this application, a resource scheduling apparatus is provided, comprising:
[0015] The information acquisition module is used to obtain channel state information between each user equipment and edge distribution unit, uplink user equipment and downlink user equipment, and uplink access point and downlink access point based on the pre-established channel model of the network-assisted free duplex architecture under the non-cellular access network.
[0016] The resource scheduling module, aiming to maximize the overall system spectral efficiency, uses a pre-configured discrete differential evolution algorithm to determine the association between edge distribution units and access points, and schedules antenna resources based on this association.
[0017] Furthermore, a game theory algorithm is used to determine the operating mode of each access point, so as to schedule time-frequency domain resources and spatial domain resources based on the operating mode.
[0018] Furthermore, the association status between edge distribution units and user equipment is determined based on backhaul capacity constraints and reinforcement learning algorithms, so as to schedule data transmission resources based on the association status; wherein, the total spectral efficiency of the system is determined based on the channel state information.
[0019] According to another aspect of this application, an electronic device is provided, the electronic device comprising:
[0020] At least one processor; and
[0021] A memory communicatively connected to the at least one processor; wherein,
[0022] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the resource scheduling method described in any embodiment of this application.
[0023] According to another aspect of this application, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the resource scheduling method described in any embodiment of this application.
[0024] According to another aspect of this application, a computer program product is provided, the computer program product including a computer program that, when executed by a processor, implements the resource scheduling method described in any embodiment of this application.
[0025] The technical solution of this application aims to maximize the overall system spectral efficiency. Based on the semi-distributed, semi-centralized characteristics of non-cellular access networks, it uses a discrete differential evolution algorithm to cluster access points, achieving semi-distributed, semi-centralized signal processing. This strikes a balance between scalability and system spectral efficiency, allowing for better management and scheduling of antenna resources, and improving the coverage and communication quality of the communication system. Simultaneously, leveraging the network-assisted free duplex feature's ability to freely schedule AP duplex modes, and employing game theory algorithms to determine the operating mode of each AP, it effectively reduces downlink interference to uplink in simultaneous full-duplex on the same frequency, and effectively utilizes spectrum resources, avoiding resource waste and interference. Furthermore, under backhaul capacity constraints, the EDU meets smaller capacity constraints by disconnecting from the UE and selects appropriate AP clusters to provide services to the UE, ensuring user service quality and reducing backhaul overhead.
[0026] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1a is a schematic diagram of a system model configuration of a non-cellular access network based on a network-assisted free-duplex architecture provided in an embodiment of this application;
[0029] Figure 1b is a schematic diagram of interference in a non-cellular access network based on a network-assisted free-duplex architecture according to an embodiment of this application;
[0030] Figure 2 is a flowchart of a resource scheduling method provided in an embodiment of this application;
[0031] Figure 3 is a flowchart of another resource scheduling method provided in an embodiment of this application;
[0032] Figure 4 is a performance curve comparison between the EDU-AP association algorithm and the traditional K-means clustering method under different AP numbers and EDU numbers provided in the embodiments of this application;
[0033] Figure 5 is a comparison of the performance curves of two different game theory-based AP mode selection algorithms with greedy algorithms, exhaustive algorithms, and random duplex algorithms under different AP numbers, according to an embodiment of this application.
[0034] Figure 6a is a spectral efficiency diagram of the system after EDU-UE association by the proposed algorithm under different capacity constraints, according to an embodiment of this application.
[0035] Figure 6b is a diagram showing the maximum number of associated nodes in the system after the proposed algorithm performs EDU-UE association under different capacity constraints, according to an embodiment of this application.
[0036] Figure 7 is a schematic diagram of a resource scheduling device provided in an embodiment of this application;
[0037] Figure 8 is a structural block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0038] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0039] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0040] Figure 1a is a schematic diagram of a system model configuration for a cellular-free access network based on a network-assisted free-duplex architecture according to an embodiment of this application. As shown in Figure 1a, the Central Processing Unit (CPU) can manage data from multiple Edge Distributed Units (EDUs), and each EDU can be connected to multiple Access Points (APs). Each AP can communicate with multiple User Equipments (UEs). During communication, the AP can receive uplink data sent by each UE and send the received uplink data to the EDU. The EDU processes the uplink data sent by multiple APs and sends the processed uplink data to the CPU. By processing the data before sending it to the CPU, the complexity of data processing can be reduced, and CPU resources can be saved.
[0041] Figure 1b is a schematic diagram of interference in a non-cellular access network based on a network-assisted free-duplex architecture according to an embodiment of this application. In this system, the interference of uplink user equipment to downlink user equipment, and the interference of downlink AP to uplink AP, are shown in Figure 1b.
[0042] It should be noted that the resource scheduling method proposed in this application is implemented in a non-cellular access network based on a network-assisted free-duplex architecture, as shown in Figures 1a and 1b. The edge distribution unit in the embodiments of this application can also be referred to as an edge distributed unit.
[0043] This application proposes a Network-Assisted Free-Duplex (NA-FD) architecture in a Cellular Free-Duplex Radio Access Network (CF-RAN). In this architecture, users operate in half-duplex mode, and access points (APs) support both full-duplex and half-duplex modes. During uplink transmission, the AP receives signals transmitted by the user and sends them to the Edge Distributed Unit (EDU). The EDU demodulates the signals and sends the data to the Central Processing Unit (CPU) for merging. Simultaneously, the CPU sends downlink data to the EDU, which pre-encodes the data and then transmits it to the downlink user equipment via the associated AP. This application aims to maximize system spectral efficiency by implementing reasonable resource scheduling to address the technical problems in related technologies where the inflexible expansion of AP operating modes in NAFD leads to low system performance, and the centralized processing of all data in the CPU results in high backhaul overhead.
[0044] In one embodiment, Figure 2 is a flowchart of a resource scheduling method provided by an embodiment of this application. This embodiment is applicable to the allocation of joint uplink and downlink resources in a network-assisted free-duplex non-cellular access network. The method can be executed by a resource scheduling device, which can be implemented in hardware and / or software and can be configured in an edge distribution unit. As shown in Figure 2, the method includes:
[0045] S110. Based on the pre-established channel model of the network-assisted free-duplex architecture under the non-cellular access network, obtain the channel state information between each user equipment and the edge distribution unit, the uplink user equipment and the downlink user equipment, and the uplink access point and the downlink access point.
[0046] The channel model is used to represent signal attenuation, delay, distortion, and noise interference during transmission. Channel state information may include, but is not limited to, the following: channel vector between edge distribution unit and downlink user equipment, interference channel matrix between uplink access point and downlink access point, channel vector between uplink user equipment and edge distribution unit, and channel information between downlink user equipment and uplink user equipment.
[0047] In this embodiment, based on the channel model, the channel vector between each edge distribution unit and each downlink user associated with it, the interference channel matrix between each downlink access point associated with each edge distribution unit and all uplink access points associated with other edge distribution units, the channel vector between each uplink user equipment and its associated edge distribution unit, and the channel information between all uplink user equipment and downlink user equipment are obtained through channel estimation.
[0048] S120. With the goal of maximizing the overall spectral efficiency of the system, a pre-configured discrete differential evolution algorithm is used to determine the association between edge distribution units and access points, and antenna resources are scheduled based on the association. Additionally, a game theory algorithm is used to determine the operating mode of each access point, and time-frequency domain resources and spatial domain resources are scheduled based on the operating mode. Furthermore, the association state between edge distribution units and user equipment is determined based on backhaul capacity constraints and reinforcement learning algorithms, and data transmission resources are scheduled based on the association state.
[0049] The overall system spectral efficiency is determined based on channel state information. Overall system spectral efficiency refers to the ratio between the total amount of information that can be transmitted within a given spectral bandwidth and the spectral bandwidth itself; it is generally used to measure the spectral resource utilization efficiency of a communication system. In one embodiment, the overall system spectral efficiency is the uplink overall spectral efficiency R. U Downlink total spectral efficiency R D The sum;
[0050] Among them, the uplink total spectral efficiency Where, r U,j Let J be the signal-to-interference-plus-noise ratio (SIR) of the j-th uplink user equipment; j≤J, where J is the total number of uplink user equipment in the non-cellular access network;
[0051] The downlink total spectral efficiency Where, r D,k Let K be the signal-to-interference-plus-noise ratio (SIR) of the k-th downlink user equipment; k ≤ K, where K is the total number of downlink user equipment in the non-cellular access network.
[0052] Correspondingly, the objective of maximizing the overall system spectral efficiency is to satisfy... Among them, {q U,l q D,l}∈{0,1} represents the working mode of the l-th access point, where U represents uplink and D represents downlink; q U,l When q is 1, it indicates that the working mode of the l-th access point is uplink mode; D,l When R is 1, it indicates that the working mode of the l-th access point is downlink mode; U,x R represents the actual backhaul capacity between the uplink central processing unit and the x-th edge distribution unit; D,x This represents the actual backhaul capacity between the downlink central processing unit and the x-th edge distribution unit; C x This represents the backhaul capacity constraint from the edge distribution unit x to the central processing unit; α D,x,k α represents the association state between the x-th edge distribution unit and the k-th downlink user equipment. U,x,j This represents the association state between the x-th edge distribution unit and the j-th uplink user equipment; for example, in α D,x,k When α equals 1, it indicates that the x-th edge distribution unit and the k-th downlink user equipment are associated, meaning that the x-th edge distribution unit is used to process the data transmitted by the k-th downlink user equipment; in α D,x,k When α equals 0, it indicates that the x-th edge distribution unit and the k-th downlink user equipment are not associated, meaning the x-th edge distribution unit does not need to process the data transmitted by the k-th downlink user equipment; U,x,j When α equals 1, it indicates that the x-th edge distribution unit and the j-th uplink user equipment are associated, meaning that the x-th edge distribution unit is used to process the data transmitted by the j-th uplink user equipment; in α U,x,j When the value is 0, it means that the x-th edge distribution unit and the j-th uplink user equipment are not associated, that is, the x-th edge distribution unit does not need to process the data transmitted by the j-th uplink user equipment.
[0053] Generally, differential evolution algorithms have good global search capabilities and adaptability to the parameter space, making them more suitable for solving continuous optimization problems. In this embodiment, based on the known physical location of the access point, a discrete differential evolution algorithm (EDU) with an adaptive multiple mutation strategy can be used to solve the discrete optimization problem in AP, thereby obtaining an optimal solution with improved performance.
[0054] The operating mode is used to characterize the data transmission direction supported by the access point in the wireless network. In an embodiment, the operating mode may include uplink mode and / or downlink mode. Of course, if an access point supports full-duplex, its operating modes include both uplink and downlink modes; if an access point supports half-duplex, its operating modes include either uplink or downlink mode. In a network-assisted free-duplex architecture, the access point can flexibly perform full-duplex operation in the space-time-frequency domain.
[0055] The game theory algorithms include evolutionary game theory algorithms and coalition game theory algorithms. In the embodiments, both evolutionary game theory algorithms and coalition game theory algorithms can be used to select the working mode of the access point.
[0056] In this embodiment, the total spectral efficiency of the system can be used as the utility function in the game theory algorithm. If an access point has a higher actual utility value when supporting uplink mode, then the access point is included as a member of the uplink region; if an access point has a higher actual utility value when supporting downlink mode, then the access point is included as a member of the downlink region.
[0057] The backhaul capacity constraint characterizes the maximum backhaul capacity between each edge distribution unit and the central processing unit. In this embodiment, the backhaul capacity constraint refers to the maximum amount of data that the resources between the edge distribution unit and the central processing unit can handle when the edge distribution unit sends data feedback to the central processing unit. For example, the reinforcement learning algorithm may include, but is not limited to, one of the following: distributed Q-learning algorithm and deep Q-network (DQN) algorithm. In this embodiment, the association state between the edge distribution unit and the user equipment is used to indicate whether all access points associated with the edge distribution unit provide services to uplink or downlink user equipment. If the association state between the edge distribution unit and one of the uplink user equipment is associated, then all access points associated with that edge distribution unit provide services to that uplink user equipment; if the association state between the edge distribution unit and one of the downlink user equipment is associated, then all access points associated with that edge distribution unit provide services to that downlink user equipment.
[0058] In this embodiment, the association states between the edge distribution unit and the downlink user equipment, and the association states between the edge distribution unit and the uplink user equipment are taken as the current states, and the action with the highest expected reward value is selected based on a greedy strategy. Then, the reward function is configured based on backhaul capacity constraints and maximizing the total system spectral efficiency. Then, the state, action, reward, and next state at each time step are stored in an array in the replay buffer. The expected reward value is updated by randomly sampling the batch size from the replay buffer until the expected reward value is optimal, thus obtaining the association states between the edge distribution unit and the downlink user equipment, and the association states between the edge distribution unit and the uplink user equipment.
[0059] In traditional NA-FD architectures in cellular unrestricted access networks (NA-FD), the access points (APs) centrally process the transmitted signals in the CPU, achieving good spectral efficiency but lacking scalability. While local precoding and demodulation by the APs can increase scalability, the lack of cooperation between APs degrades system performance. The NA-FD architecture in this application allows for precoding and receiver design of the data distribution units (EDUs), and partial cooperation between EDUs, achieving a partially centralized and partially distributed signal processing, thus striking a balance between scalability and system performance.
[0060] Centralized architectures offer good performance but lack scalability, while fully distributed architectures suffer from poor performance due to a lack of collaboration. This application addresses the collaboration issue between APs by determining the association between the EDU and APs, achieving semi-distributed, semi-centralized signal processing and reducing the complexity of centralized transceiver design. APs have a fixed number and orientation of antennas; the EDU can improve the coverage and communication quality of the communication system by scheduling antenna resources from associated APs. Simultaneously, co-frequency full-duplex systems can significantly increase the system's spectral efficiency and enhance time-frequency resource utilization, but also introduce strong inter-antenna interference. In the NA-FD system of this application, by selecting the AP's duplex mode, downlink interference to uplink in co-frequency full-duplex systems is reduced, and time-frequency and spatial resources are rationally scheduled to avoid resource waste. The EDU and CPU have limited backhaul capacity; therefore, the EDU needs to disconnect from the UE to meet this backhaul capacity constraint while ensuring user service quality. Thus, by selecting associated UEs and allocating data transmission resources to them, the system's spectral efficiency is maximized while satisfying the backhaul capacity constraint.
[0061] The technical solution of this embodiment aims to maximize the overall system spectral efficiency. Based on the semi-distributed, semi-centralized characteristics of non-cellular access networks, it uses a discrete differential evolution algorithm to cluster access points, achieving semi-distributed, semi-centralized signal processing. This strikes a balance between scalability and system spectral efficiency, allowing for better management and scheduling of antenna resources, and improving the coverage and communication quality of the communication system. Simultaneously, leveraging the network-assisted free-duplex feature's ability to freely schedule AP duplex modes, and employing game theory algorithms to determine the operating mode of each AP, it effectively reduces downlink interference to uplink in simultaneous full-duplex communication on the same frequency, and effectively utilizes spectrum resources, avoiding resource waste and interference. Furthermore, under backhaul capacity constraints, the EDU meets smaller capacity constraints by disconnecting from the UE and selects appropriate AP clusters to provide services to the UE, ensuring user service quality and reducing backhaul overhead.
[0062] In one embodiment, Figure 3 is a flowchart of another resource scheduling method provided by an embodiment of this application. This embodiment further refines the process of estimating channel state information, determining the association between EDU and AP, determining the working mode of each AP, and determining the association state between EDU and UE, based on the above embodiments. As shown in Figure 3, the method includes:
[0063] S210. Based on the channel model, the channel vector between each edge distribution unit and its associated downlink user equipment is obtained through channel estimation.
[0064] Assume the channel vector between the x-th edge distribution unit and the k-th downlink user equipment is denoted as... Where, N x M represents the number of access points associated with the edge distribution unit x; M represents the number of antennas contained in each access point. The Nth edge distribution unit x is associated with x The channel vector between the access point and the k-th downlink user equipment; It refers to a pre-configured set of complex numbers.
[0065] S220. Obtain the interference channel matrix between the downlink access point associated with each edge distribution unit and the uplink access point associated with other edge distribution units through channel estimation.
[0066] Suppose the interference channel matrix between the downlink access point associated with the x-th edge distribution unit and the uplink access point associated with the x'-th edge distribution unit is denoted as . in, Let IAI be the interference channel matrix between the nth downlink access point associated with the xth edge distribution unit and all uplink access points associated with the x'th edge distribution unit; IAI is the interference from the downlink access point to the uplink access point; n is less than or equal to N. x In one embodiment, if the distance d between the uplink access point associated with edge distribution unit x and the downlink access points associated with other edge distribution units x′ is... i,j Less than the preset distance threshold d min If the interference is not properly addressed, then the interference cancellation principle is applied. This involves assigning other edge distribution units x′ to a set of other edge distribution units that cooperate with edge distribution unit x, in order to cancel interference.
[0067] S230. Obtain the channel information between each uplink user equipment and downlink user equipment through channel estimation.
[0068] Assume that both uplink and downlink user equipment have 1 antenna. Let h be the channel information between the j-th uplink user equipment and the k-th downlink user equipment. k,j In other words, the channel information between the uplink and downlink user equipment is a scalar. It is assumed that the channel follows flat fading, meaning that the channel coefficients remain unchanged during the channel coherence time.
[0069] S240. Obtain the channel vector between each uplink user equipment and its associated edge distribution unit through channel estimation.
[0070] Suppose that the channel vector between the j-th uplink user equipment and the x-th edge distribution unit is denoted as .
[0071] S250, the objective problem of maximizing the overall spectral efficiency of the system.
[0072] In this embodiment, with the goal of maximizing the overall system spectral efficiency, the following issues are addressed in the network-assisted free-duplex architecture: the association between edge distribution units and access points, access point mode selection, and user selection.
[0073] And, at the same time, q D,l q U,l ∈{0,1},
[0074] Among them, {q U,l q D,l}∈{0,1} represents the working mode of the l-th access point; q U,l =1 indicates that the l-th AP is operating in uplink mode; q D,l =1 indicates that the l-th AP is operating in downlink mode; and q is satisfied.U,l +q D,l =1 condition. Let represent the uplink and downlink mode vectors of the AP associated with the x-th edge distribution unit, respectively. Convert these vectors to diagonal matrices. This can more clearly represent the AP working mode associated with the xth edge distribution unit.
[0075] in, This refers to the total uplink spectral efficiency;
[0076] Where, r U,j p is the uplink signal-to-interference-plus-noise ratio (SIR) of the j-th uplink user equipment; j v is the transmit power of uplink user equipment j. x,j Let x be the reception vector between the edge distribution unit x and the uplink user equipment j, i.e. N edu The total number of EDUs included in this architecture; its numerator is the useful signal power of uplink user equipment j, and its denominator is the interference signal power of uplink user equipment j; It is the conjugate transpose of the receive vectors of edge distribution unit x and uplink user equipment j.
[0077] In this equation, the first term on the right-hand side represents the interference signal power of other uplink user equipment j'; the second term represents the interference signal power of downlink user equipment on an EDU that can perform interference cancellation (i.e., can cooperate) to uplink user equipment j; the third term represents the interference signal power of downlink user equipment on an EDU that cannot perform interference (i.e., cannot cooperate) to uplink user equipment j; and the fourth term represents the channel noise interference signal power. Here, j' represents other uplink user equipment besides uplink user equipment j. ω represents the residual interference channel matrix between edge distribution units x and x' after interference cancellation. k,x′ For the precoding of downlink user equipment k to edge distribution unit x', N IAI,x This is the set of other edge distribution units that can collaborate with edge distribution unit x; This is the set of other edge distribution units that cannot cooperate with edge distribution unit x; Let σ be the interference channel matrix between the downlink access point associated with the xth edge distribution unit and the uplink access point associated with the x'th edge distribution unit, where σ is the channel noise.
[0078] This represents the actual backhaul capacity between the uplink central processing unit and the x-th edge distribution unit;
[0079] The backhaul formula contains the same parts as each part of the aforementioned interference signal power formula. The difference lies in that the backhaul is the user rate transmitted between a single EDU and the CPU, therefore it does not need to be summed over the EDU, i.e., there is no summation over the EDU.
[0080] This represents the total downlink spectral efficiency.
[0081] Where, r D,k The formula for downlink signal-to-interference-plus-noise ratio (SINIRR) is as follows: the numerator is the useful signal power of downlink user equipment k, the denominator is as follows: the first term is the interference signal power of other downlink user equipment, the second term is the interference signal power of uplink user equipment to the downlink user equipment, and the third term is the channel noise power.
[0082] This represents the actual backhaul capacity between the downlink central processing unit and the xth edge distribution unit.
[0083] C x This represents the backhaul capacity constraint from the edge distribution unit x to the central processing unit. The binary variable q... U,l and q D,l Indicates the mode selection for access point l, α D,x,k and α U,x,j This represents the association state between edge distribution units and user equipment. Due to the integer binary nature of the variables, this optimization problem is a nonlinear integer programming problem without concavity or convexity. Therefore, discrete difference evolutionary algorithms with various strategies can be used to determine the association state between edge distribution units and access points. Then, game theory algorithms are used to solve the access point mode selection problem, thereby scheduling resources for the network-assisted free-duplex system to maximize the system's spectral efficiency. This can be understood as AP duplex mode selection effectively reducing downlink-to-downlink interference in simultaneous full-duplex operation on the same frequency, and simultaneously full-duplex operation on the same frequency can effectively utilize spectrum resources, avoiding resource waste and interference.
[0084] In this embodiment, the precoding from the edge distribution unit x to the downlink user equipment k is calculated, i.e. in, It is the conjugate transpose of the channel vector between the x-th edge distribution unit and the k-th downlink user equipment; The downlink signal sent is in, Information to be sent to downlink user equipment k.
[0085] The received signal at user equipment k in the downlink is:
[0086] in p j Let j be the transmit power of user equipment. For uplink information, Gaussian noise; h k,j This refers to the channel information between the j-th uplink user equipment and the k-th downlink user equipment.
[0087] In the uplink, the edge distribution units can cooperate to eliminate interference between access points. However, due to channel estimation errors, this interference cannot be completely eliminated. Considering the remaining self-interference, the uplink signal received by the edge distribution units is:
[0088] Where N IAI,x Let x be the set of other edge distribution units that can cooperate with edge distribution unit x; let d be the distance between the uplink access point associated with edge distribution unit x and the downlink access points associated with other edge distribution units x′. i,j Less than the preset distance threshold d min ,Right now Then, other edge distribution units x′ are assigned to other edge distribution units that cooperate with edge distribution unit x for interference cancellation. That is, interference between edge distribution unit x and each edge distribution unit x′ in other edge distribution unit sets is cancelled, and the remaining interference channel matrix between IAI is obtained after interference cancellation. This represents the residual interference power resulting from incomplete interference cancellation. This is the set of other edge distribution units that cannot cooperate with edge distribution unit x. It is Gaussian noise. It is the receive vector between the edge distribution unit x and the uplink user equipment j; MN x ×MN x The identity matrix; This represents the relative distance between the first receiving AP (i.e., uplink AP) of edge distribution unit x and the first transmitting AP (i.e., downlink AP) of edge distribution unit x′; This represents the relative distance between the second receiving AP (i.e., uplink AP) of edge distribution unit x and the second transmitting AP (i.e., downlink AP) of edge distribution unit x′.
[0089] S260. Initialize the population based on the total number of access points and the total number of edge distribution units in the non-cellular access network.
[0090] Each individual in the population is represented by a continuous sequence of length equal to the total number of access points; each element in the continuous sequence represents the association state between an access point and an edge distribution unit.
[0091] The population contains a total of N p There are 10 individuals, each represented by a continuous sequence b of length L equal to the total number of access points. Each element of the continuous sequence b can be derived from the solution space [b...]. min b max Take a random value from ], i.e., b i,j =b min +rand(0,1)·(b max -b min );b i,j This represents the association state between the j-th access point in the i-th sequence of the continuous sequence b and the edge distribution unit; using (b max -b min +1) / N edu Divide the continuous sequence b uniformly into N. edu Intervals yield discrete sequences b i,j =x indicates that the j-th access point is associated with edge distribution unit x. In actual operation, the difference in the number of access points associated with different edge distribution units in each individual does not exceed 1.
[0092] S270. Determine the fitness of each individual based on the discrete sequence corresponding to the continuous sequence.
[0093] Based on discrete sequence Calculate the fitness of each individual. The fitness function is defined as the sum of the distances between access points in each association. Where N... x and N x′ Let d represent the sets of access points associated with edge distribution units x and x′, respectively. i,j This represents the distance between the i-th access point associated with edge distribution unit x and the j-th access point associated with other edge distribution units x′.
[0094] S280. Divide the population into multiple subpopulations according to fitness, and use a mutation strategy corresponding to the fitness of each subpopulation to perform mutation operation on each individual in the corresponding subpopulation to obtain the corresponding mutant individual.
[0095] The mutation strategy can be an adaptive multiple mutation strategy. In this embodiment, employing an adaptive multiple mutation strategy can prevent the population from lacking diversity, getting trapped in local optima, or increasing convergence time and reducing search efficiency. For example, the adaptive multiple mutation strategy may include: DE / best / 1 mutation strategy, DE / rand-to-best / 1 mutation strategy, and DE / current-to-best / 1 mutation strategy, etc. In this strategy, the population is divided into three subpopulations according to fitness, and each subpopulation uses a different mutation strategy.
[0096] For the subpopulation with the worst fitness, the DE / rand / l mutation strategy is adopted:
[0097] For subpopulations with moderate fitness, the DE / best / 2 mutation strategy is used:
[0098] For the subpopulation with the best fitness, the DE / best-ass rand / 2 mutation strategy is adopted:
[0099] Among them, b min The individuals with the best fitness are r1, r2, r3, and r4, which are [1, N]. p A random set of distinct integers within the range (0, 1), where a1, a2, and a3 are random numbers within the range (0, 1) such that a1 + a2 + a3 = 1; F represents the adaptive mutation probability, F = F0·2 λ Where F0 is the initial parameter and λ represents the adaptive factor. iteration is the maximum number of iterations, and iter represents the current iteration.
[0100] S290. Perform crossover operations on each individual and its corresponding mutant individual to generate the corresponding experimental individual and form the corresponding experimental population.
[0101] Each individual b i The corresponding mutant individual m i Cross-generating experimental individuals t i The crossover process is as follows: Wherein, CR represents the crossover rate.
[0102] S2100 follows the greedy principle to select individuals from the experimental group to obtain the next generation group.
[0103] Based on the discrete differential evolution algorithm, and following a greedy criterion, individuals are selected from the experimental population t to serve as the next generation population. The selection method is as follows:
[0104] S2110. Use the pre-configured boundary absorption method to check the boundary conditions and process the individuals outside the boundary. Return the steps to determine the fitness of each individual based on the discrete sequence corresponding to the continuous sequence until the current iteration number reaches the maximum iteration number. Output the optimal individual, which represents the final association between the edge distribution unit and the access point.
[0105] Due to mutation and crossover operations, some elements within an individual may exceed a given boundary. Therefore, it is necessary to perform boundary condition checks and process individuals outside the boundary; this can be achieved using the boundary absorption method. Perform boundary condition checks and process individuals outside the boundary.
[0106] The implementation process of the discrete differential evolution algorithm, from fitness calculation to boundary handling, continues until the maximum number of iterations is reached, at which point the optimal individual is output. This indicates the final association between the edge distribution unit and the access point.
[0107] S2120, taking the set of access points as participants and the total spectral efficiency of the system as a utility function.
[0108] After resolving the association issue between edge distribution units and access points, Q can be determined. U,x and Q D,x This application utilizes game theory algorithms to select the working mode of the access point. Game theory algorithms emphasize cooperation among participants, prioritizing overall optimal performance over individual gains.
[0109] The working modes of access points are divided into uplink mode and downlink mode. The corresponding set of access points consists of a subset of access points working in uplink mode and a subset of access points working in downlink mode.
[0110] Access point set As a participant, and taking the total system spectral efficiency as a utility function; wherein, the access point's operating mode selection is divided into uplink and downlink regions, and the current partition is denoted as And satisfy in, This is the uplink area; This refers to the downlink region. It should be noted that, from a half-duplex perspective, if two half-duplex APs are located in the same physical location, then the AP supports full-duplex.
[0111] S2130, based on the total spectral efficiency of the system operating in uplink and downlink modes, and using game theory algorithms, partition switching is performed until the coalition partition finally converges to the Nash stable partition, so as to determine the operating mode of each access point by maximizing the final set of access points corresponding to the total spectral efficiency of the system.
[0112] During the initialization phase, access points are randomly assigned to different regions. If the i-th access point is assigned to... Compared to being assigned to To achieve higher efficiency, switching the i-th access point is called... Members; when they are from Switch to and At that time, the current partition Transform into a new partition Represented as If and only if satisfying The above switching occurs when the condition {m, m′} ∈ {up, down} is met;
[0113] After randomly initializing the access points into regions, access point i is selected in a predetermined order, and the current alliance is saved. Then, a temporary alliance is selected, allowing switching between uplink and downlink alliances. After selecting an alliance, the utility function is calculated. If the overall system spectral efficiency increases, an alliance switching operation is performed on the access point, and the current uplink and downlink regions are updated accordingly. If the overall system spectral efficiency decreases, the count parameter tracking consecutive non-switching operations is incremented by 1. Furthermore, to prevent premature convergence and avoid getting trapped in local optima, access point j can be randomly selected for alliance switching operations in each iteration based on the mutation probability nμ, until the count parameter reaches ten times the number of access points. After a finite number of switching operations, the alliance partition eventually converges to a Nash stable partition, denoted as . Then, the working mode of each access point in the final set of access points corresponding to the consortium partition is used as the working mode of the corresponding access point.
[0114] S2140. Initialize the state space of each edge distribution unit based on the association state between the edge distribution unit and the downlink user equipment, and the association state between the edge distribution unit and the uplink user equipment.
[0115] In the embodiment, after resolving the EDU-AP association and AP operating mode issues, Q U,x and Q D,x Both the dimensions and values can be determined. Specifically, the process of determining the EDU-AP association involves determining N. x The process of determining Q U,x and Q D,x The process of determining the dimensions of AP; the process of determining the working mode of AP is to determine Q. U,x and Q D,x The process of determining the numerical value. The EDU-UE correlation problem is solved using a distributed Q-learning algorithm based on experience replay to determine α. D,x,k and α U,x,jThe value of . Among them, EDU-AP association is equivalent to clustering APs. By increasing the scalability of the system through precoding at the EDU and receiver design, the system can achieve semi-distributed and semi-centralized processing, thereby better managing and optimizing resource allocation and increasing the service range; EDU-UE association is to select the corresponding AP cluster to provide services to users, ensuring the quality of service for users and reducing backhaul overhead.
[0116] Q-learning is a classic algorithm in reinforcement learning, mainly composed of states, actions, and rewards. We define S = {s1, s2, ..., s}. n As a set of states, A = {a1, a2, ..., a} n As a set of operations, the agent can learn by taking actions a∈A from the current environment and receiving feedback from the environment. The environment provides a reward value r. t The form (s, a) provides immediate feedback, and using the Bellman equation, we can obtain the expected reward Q value; that is, the larger the Q value, the better the situation. Among them, S t+1 and A t+1 S represents the state and action at time t+1, respectively. t and A t Let them represent the state and action at time t, respectively. This represents the expectation operator. The discount factor γ (0≤γ≤1) determines the importance of future rewards; a higher γ value indicates a greater emphasis on future rewards.
[0117] In the distributed Q-learning algorithm, each edge distribution unit is treated as an independent agent, which is trained independently and generates its own Q-table. The state space of each edge distribution unit is defined as S. x ={α U,x,1 , …, α U,x,J α D,x,1 , …, α D,x,K ), used to represent the association status between the edge distribution unit x and the uplink and downlink user equipment; α D,x,k α represents the association between edge distribution unit x and downlink user equipment k. U,x,j This represents the association between edge distribution unit x and uplink user equipment j; where k≤K, j≤J; J is the total number of uplink user equipment in the non-cellular access network; and K is the total number of downlink user equipment in the non-cellular access network. (Using 1-α) U,x,j or 1-α D,x,k The operation enables the switching of associated states.
[0118] S2150: The dynamic decay greedy strategy is used to optimize the overall spectral efficiency of the system as the objective function, and the action with the highest expected reward value is selected.
[0119] The dynamically decaying greedy strategy is used to dynamically decay the probability of action selection based on the greedy strategy. The probability of random action selection based on the greedy strategy is denoted as ∈(iter); correspondingly, the probability of action selection based on past experience is denoted as 1-∈(iter). The dynamically decaying ∈ greedy strategy optimizes the objective function. The agent makes decisions based on past experience with a probability of 1-∈(iter), selecting the action with the highest expected reward Q value. As training progresses, ∈(iter) decreases, thus achieving a more refined optimization process, which can be defined as: ∈0 is the initial value of ∈, φ is a parameter that controls the decay rate. The higher the value of φ, the slower the decay rate during training. |action| represents the size of the action set.
[0120] S2160, configure a reward function based on backhaul capacity constraints and the objective of maximizing the overall system spectral efficiency.
[0121] Based on backhaul capacity constraints and maximizing the overall system spectral efficiency, the reward function is configured as follows: Where, χ t χ indicates whether the edge distribution unit satisfies the capacity constraint at time t. t =1 indicates that the capacity constraint is met, χ t =0 indicates non-compliance; R sum,r =R U,t +R D,t R represents the total spectral efficiency of the system at time t. sum,all =R U,all +R D,all R represents the spectral efficiency under full connectivity of edge distributed units and users; x,t This represents the actual return capacity at time t.
[0122] S2170. Store the state, action, reward, and next state at each moment as a set of experience values in the replay buffer.
[0123] The core idea of experience replay is to store the agent's experience in a buffer, where each time step's state, action, reward, and next state are grouped as a set of experience values, denoted as (s). t ,a t r t s t+1 These data are stored in an array in the replay buffer; the size of the array is the replay size. When the array is full, the oldest data will be replaced by the newest data. Once enough data is stored in the replay buffer, the replay update process begins.
[0124] S2180: Randomly sample experience values of batch size from the replay buffer to update the expected reward value until the expected reward value is optimal.
[0125] During the replay update, experience values of a random batch size are drawn from the replay buffer to update the expected reward Q value until the expected reward Q value is optimal, thus obtaining α. D,x,k and α U,x,j The value of this allows the agent to optimize its strategy by repeatedly learning from past experiences. Furthermore, sharing experience buffers among different edge distribution units can enhance collaboration.
[0126] This application proposes a network-assisted free-duplex non-cellular wireless access network. In this network, access points can flexibly select duplex modes and are scalable, significantly improving system performance and providing services to more users. Simultaneously, this application addresses the EDU-AP association problem, reducing the CPU signal processing burden. The EDU performs transceiver design locally, meaning precoding and receiver computations are performed within the EDU. Furthermore, multiple APs integrate information for the same user at the EDU, enhancing user service quality and reducing computational complexity. The use of a discrete differential evolution algorithm achieves better performance than the traditional K-means clustering algorithm. This application also addresses the AP mode selection problem, allowing APs to operate in either uplink receive or downlink transmit modes. A flexible duplex strategy eliminates cross-link interference in full-duplex mode, and game theory algorithms achieve good performance. Finally, this application addresses the EDU-UE association problem, considering user selection under backhaul capacity constraints to maximize system spectral efficiency. A distributed Q-learning algorithm based on experience replay is employed, allowing each EDU to independently and in parallel generate its own Q-table, while experience replay enhances their collaboration.
[0127] Figure 4 is a performance comparison chart of the EDU-AP association algorithm and the traditional K-means clustering method under different AP and EDU numbers provided in the embodiments of this application. As shown in Figure 4, the performance comparison chart of the EDU-AP association algorithm proposed in the embodiments of this application and the traditional K-means clustering method is given as the number of APs and EDUs changes. It can be clearly seen that the algorithm proposed in the embodiments of this application can achieve better system spectral efficiency.
[0128] Figure 5 is a performance comparison chart of two different game theory-based AP mode selection algorithms (Evolution_game and Coalitional_game) with greedy algorithms, exhaustive algorithms (i.e., Exhaustion), and random duplex algorithms provided in this application embodiment under different AP numbers. As shown in Figure 5, the performance comparison chart of the AP mode selection scheme proposed in this application embodiment with other algorithms is given under different AP numbers. It can be clearly seen that the algorithm proposed in this application embodiment has performance close to that of the exhaustive algorithm, achieving good system performance with lower computational complexity.
[0129] Figure 6a shows the spectral efficiency of the system after EDU-UE association by the proposed algorithm under different capacity constraints, according to an embodiment of this application. Figure 6b shows the maximum number of associated nodes in the system after EDU-UE association by the proposed algorithm under different capacity constraints, according to an embodiment of this application. As shown in Figures 6a and 6b, the performance and maximum number of associated nodes of the EDU-UE association method proposed in this application are presented under different backhaul capacity constraints. It can be clearly seen that when the backhaul capacity constraint is reduced, the EDU meets the smaller capacity limit by increasing the number of disconnected users, while ensuring that the spectral efficiency of the system decreases slowly, thereby maintaining satisfactory system performance.
[0130] In one embodiment, FIG7 is a schematic diagram of a resource scheduling device provided in an embodiment of this application. As shown in FIG7, the device includes: an information acquisition module 310 and a resource scheduling module 320.
[0131] The information acquisition module 310 is used to obtain channel state information between each user equipment and edge distribution unit, uplink user equipment and downlink user equipment, and uplink access point and downlink access point based on the pre-established channel model of the network-assisted free duplex architecture under the non-cellular access network.
[0132] The resource scheduling module 320 is used to determine the association between edge distribution units and access points using a pre-configured discrete differential evolution algorithm with the goal of maximizing the overall system spectral efficiency, and to schedule antenna resources based on the association. It also uses a game theory algorithm to determine the operating mode of each access point, and to schedule time-frequency domain resources and spatial domain resources based on the operating mode. Furthermore, it determines the association state between edge distribution units and user equipment based on backhaul capacity constraints and reinforcement learning algorithms, and to schedule data transmission resources based on the association state. The overall system spectral efficiency is determined based on channel state information.
[0133] In one embodiment, the information acquisition module 310 is specifically used for:
[0134] Based on the channel model, the channel vector between each edge distribution unit and its associated downlink user equipment is obtained through channel estimation.
[0135] The interference channel matrix between the downlink access point associated with each edge distribution unit and the uplink access points associated with other edge distribution units.
[0136] Channel information between each uplink user equipment and downlink user equipment
[0137] The channel vector between each uplink user equipment and its associated edge distribution unit.
[0138] In one embodiment, the total system spectral efficiency is the uplink total spectral efficiency R. U Downlink total spectral efficiency R D The sum;
[0139] Among them, the uplink total spectral efficiency Where, r U,j Let J be the signal-to-interference-plus-noise ratio (SIR) of the j-th uplink user equipment; j≤J, where J is the total number of uplink user equipment in the non-cellular access network;
[0140] The downlink total spectral efficiency Where, r D,k Let K be the signal-to-interference-plus-noise ratio (SIR) of the k-th downlink user equipment; k ≤ K, where K is the total number of downlink user equipment in the non-cellular access network.
[0141] Correspondingly, the objective of maximizing the overall system spectral efficiency is to satisfy... Among them, {q U,l q D,l}∈{0,1} represents the working mode of the l-th access point; R U,x R represents the actual backhaul capacity between the uplink central processing unit and the x-th edge distribution unit; D,x This represents the actual backhaul capacity between the downlink central processing unit and the x-th edge distribution unit; C x This represents the backhaul capacity constraint from the edge distribution unit x to the central processing unit; α D,x,k This represents the association state between the x-th edge distribution unit and the k-th downlink user equipment; α U,x,j This indicates the association status between the x-th edge distribution unit and the j-th uplink user equipment; the actual backhaul capacity is determined based on the channel state information.
[0142] In one embodiment, the Discrete Differential Evolutionary Algorithm (DDE) is obtained by adding crossover and mutation operations to the Discrete Differential Algorithm; the pre-configured DDE is used to determine the association between edge distribution units and access points, specifically for:
[0143] The population is initialized based on the total number of access points and the total number of edge distribution units in the non-cellular access network; wherein, each individual in the population is represented by a continuous sequence of length equal to the total number of access points; wherein, each element in the continuous sequence represents the association state between an access point and an edge distribution unit;
[0144] The fitness of each individual is determined based on the discrete sequence corresponding to the continuous sequence;
[0145] The population is divided into multiple subpopulations according to the fitness, and each individual in the corresponding subpopulation is mutated using a mutation strategy corresponding to the fitness of each subpopulation to obtain the corresponding mutated individual.
[0146] Each individual is cross-crossed with its corresponding mutant individual to generate the corresponding experimental individual and form the corresponding experimental population;
[0147] Following a greedy criterion, individuals are selected from the experimental group to obtain the next generation group;
[0148] A pre-configured boundary absorption method is used to check boundary conditions and process individuals outside the boundary.
[0149] Return to the steps of determining the fitness of each individual based on the discrete sequence corresponding to the continuous sequence, until the current iteration number reaches the maximum iteration number, and output the optimal individual, representing the final association between the edge distribution unit and the access point.
[0150] In one embodiment, if the distance d between the uplink access point associated with edge distribution unit x and the downlink access points associated with other edge distribution units x′ is... i,j Less than the preset distance threshold d min Then, other edge distribution units x′ are assigned to other edge distribution units that cooperate with edge distribution unit x in order to eliminate interference.
[0151] In one embodiment, a game theory algorithm is used to determine the operating mode of each access point, specifically for:
[0152] The set of access points is taken as participants, and the total spectral efficiency of the system is taken as the utility function; wherein, the working mode of the access points is divided into uplink mode and downlink mode, and the corresponding set of access points is composed of the subset of access points working in uplink mode and the subset of access points working in downlink mode.
[0153] Based on the overall system spectral efficiency operating in uplink and downlink modes, and using game theory algorithms to switch partitions until the coalition partitions finally converge to the Nash stable partition, the operating mode of each access point is determined by maximizing the final set of access points corresponding to the overall system spectral efficiency.
[0154] In one embodiment, the association state between edge distribution units and user equipment is determined based on backhaul capacity constraints and reinforcement learning algorithms, specifically for:
[0155] The state space of each edge distribution unit is initialized based on the association state between the edge distribution unit and the downlink user equipment, and the association state between the edge distribution unit and the uplink user equipment.
[0156] The optimization employs a dynamic decay greedy strategy, taking the total spectral efficiency of the system as the objective function, to select the action with the highest expected reward value; wherein, the dynamic decay greedy strategy is used to dynamically decay the strategy selection probability based on the greedy strategy.
[0157] A reward function is configured based on backhaul capacity constraints and the objective of maximizing the overall system spectral efficiency.
[0158] The state, action, reward, and next state at each moment are stored as a set of experience values in an array in the replay buffer;
[0159] The expected reward value is updated by randomly selecting a batch of experience values from the replay buffer until the expected reward value is optimal.
[0160] The resource scheduling device provided in this application embodiment can execute the resource scheduling method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of the execution method.
[0161] In one embodiment, FIG8 is a structural block diagram of an electronic device provided in an embodiment of this application. As shown in FIG8, a structural schematic diagram of an electronic device 10 that can be used to implement an embodiment of this application is illustrated. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the application described and / or claimed herein. The electronic device may be the aforementioned EDU.
[0162] As shown in Figure 8, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer programs stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0163] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0164] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as resource scheduling methods.
[0165] In some embodiments, the resource scheduling method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the resource scheduling method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the resource scheduling method by any other suitable means (e.g., by means of firmware).
[0166] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0167] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0168] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0169] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0170] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0171] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0172] This application also provides a computer program product, including a computer program that, when executed by a processor, can implement the resource scheduling method provided in any embodiment of this application.
[0173] In the implementation of the computer program product, computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer through any type of network—including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0174] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.
[0175] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A resource scheduling method, comprising: obtaining channel state information between each user equipment and an edge distribution unit, between uplink user equipment and downlink user equipment, and between uplink access points and downlink access points based on a pre-established channel model of a network-assisted free duplex architecture under a cell-free access network; determining an association between edge distribution units and access points using a pre-configured discrete differential evolution algorithm to maximize system total spectral efficiency, scheduling antenna resources based on the association, and determining an operation mode of each access point using a game theory algorithm to schedule time-frequency domain resources and space domain resources based on the operation mode; and determining an association state between edge distribution units and user equipment based on backhaul capacity constraints and a reinforcement learning algorithm to schedule data transmission resources based on the association state, wherein the system total spectral efficiency is determined based on the channel state information.
2. The method of claim 1, wherein, The obtaining of the channel state information between each user equipment and an edge distribution unit, between uplink user equipment and downlink user equipment, and between uplink access points and downlink access points based on the pre-established channel model of the network-assisted free duplex architecture under the cell-free access network comprises: obtaining a channel vector between each edge distribution unit and its associated downlink user equipment through channel estimation based on the channel model, an interference channel matrix between a downlink access point associated with each edge distribution unit and an uplink access point associated with other edge distribution units, channel information between each uplink user equipment and downlink user equipment, and a channel vector between each uplink user equipment and its associated edge distribution unit.
3. The method of claim 1, wherein, The total spectral efficiency of the system is the sum of the uplink total spectral efficiency R U and the downlink total spectral efficiency R D . The uplink total spectral efficiency wherein r U,j is the signal-to-interference-and-noise ratio of the jth uplink user equipment; j≤J, where J is the total number of uplink user equipments in the cell-free access network. The downlink total spectral efficiency wherein r D,k is the signal-to-interference-and-noise ratio of the kth downlink user equipment; k≤K, where K is the total number of downlink user equipments in the cell-free access network. Correspondingly, the maximum system spectral efficiency is taken as the target, and the following conditions are met wherein {q U,l , q D,l} ∈ {0, 1} denotes the operating mode of the lth access point; R U,x denotes the uplink central processing unit and the downlink central processing unit, respectively. actual backhaul capacity of the xth edge distribution unit; R D,x represents the actual backhaul capacity of the xth edge distribution unit to the downlink central processing unit; C x represents the backhaul capacity constraint of the edge distribution unit x to the central processing unit; a D,x,k represents the association status between the xth edge distribution unit and the kth downlink user equipment; a U,x,j represents the association status between the xth edge distribution unit and the jth uplink user equipment; the actual backhaul capacity is determined based on the channel status information.
4. The method of claim 1, wherein, The discrete differential evolution algorithm is obtained by adding a crossover operation and a mutation operation to a discrete differential algorithm;The determination of the association between edge distribution units and access points using the pre-configured discrete differential evolution algorithm comprises: initializing a population based on a total number of access points and a total number of edge distribution units in the cell-free access network;Each individual in the population is represented by a continuous sequence with a length of the total number of access points;Each element in the continuous sequence represents an association state between an access point and an edge distribution unit; determining a fitness of each individual according to a discrete sequence corresponding to the continuous sequence; dividing the population into a plurality of sub-populations according to the fitness, and performing a mutation operation on each individual in a corresponding sub-population using a mutation strategy corresponding to the fitness of the sub-population to obtain a corresponding mutated individual; performing a crossover operation on each individual and its corresponding mutated individual to generate a corresponding trial individual and form a corresponding trial population; selecting individuals from the trial population in accordance with a greedy rule to obtain a next generation population; performing a boundary condition check using a pre-configured boundary absorption method and processing individuals outside the boundary; returning to the step of determining the fitness of each individual according to the discrete sequence corresponding to the continuous sequence until the current iteration number reaches a maximum iteration number, and outputting an optimal individual representing the final association between edge distribution units and access points.
5. The method of claim 2, wherein, If the distance d between the uplink access point associated with the edge distribution unit x and the downlink access point associated with the other edge distribution unit x' is smaller than a pre-set distance threshold d i,j d < d min then the other edge distribution unit x' is divided into the set of other edge distribution units that cooperate with the edge distribution unit x for interference cancellation.
6. The method of claim 1, wherein, The determination of the operation mode of each access point using the game theory algorithm comprises: The access point set is taken as a participant, and the system total spectrum efficiency is taken as an utility function; wherein, the working mode of the access point is divided into an uplink mode and a downlink mode, and a subset of access points working in the uplink mode and a subset of access points working in the downlink mode form the corresponding access point set; The system total spectrum efficiency based on the working in the uplink mode and the downlink mode and a game theory algorithm are used for partition switching until the alliance partition converges to a Nash stable partition, so as to determine the working mode of each access point by maximizing the final access point set corresponding to the system total spectrum efficiency.
7. The method of claim 1, wherein, The association state of the edge distribution unit and the user equipment is determined based on the backhaul capacity constraint and the reinforcement learning algorithm, including: The state space of each edge distribution unit is initialized based on the association state between the edge distribution unit and the downlink user equipment and the association state between the edge distribution unit and the uplink user equipment; A dynamic decay greedy strategy is used to optimize the system total spectrum efficiency as a target function, and an action with the highest expected reward value is selected; wherein, the dynamic decay greedy strategy is used to dynamically decay the strategy selection probability of action selection according to the greedy strategy; A reward function is configured based on the backhaul capacity constraint and the maximization of the system total spectrum efficiency as a target; Each state, action, reward and next state at each time is taken as a set of experience values, and is stored in a replay buffer; Experience values of a batch size are randomly extracted from the replay buffer to update the expected reward value until the expected reward value is optimal, and the association state of the edge distribution unit and the user equipment is obtained.
8. A resource scheduling apparatus, comprising: an information acquisition module, configured to obtain channel state information between each user equipment and edge distribution unit, uplink user equipment and downlink user equipment, and uplink access point and downlink access point based on a pre-established channel model of a network-assisted full-duplex architecture under a cell-free access network; a resource scheduling module, configured to maximize the system total spectrum efficiency as a target, determine the association between the edge distribution unit and the access point by using a pre-configured discrete differential evolution algorithm, schedule antenna resources based on the association, and determine the working mode of each access point by using a game theory algorithm, schedule time-frequency domain resources and space domain resources based on the working mode, and determine the association state of the edge distribution unit and the user equipment based on the backhaul capacity constraint and the reinforcement learning algorithm, and schedule data transmission resources based on the association state; wherein, the system total spectrum efficiency is determined based on the channel state information.
9. An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the resource scheduling method of any one of claims 1-7. 10. A computer readable storage medium storing computer instructions for causing a processor to implement the resource scheduling method of any one of claims 1-7 when executed.
11. A computer program product comprising a computer program for implementing the resource scheduling method of any one of claims 1-7 when executed by a processor.
Citation Information
Patent Citations
Network-assisted full duplex system energy efficiency optimization method and system
CN114885423A
Method and device for determining energy efficiency of communication system, equipment and storage medium
CN116321393A
Method for optimizing pre-allocation-optimization duplex mode of network-assisted full duplex system
CN117395688A
Energy efficiency optimization method and system for network-assisted full-duplex system
WO2023185077A1
Cited By
Joint task unloading and resource allocation method in low-altitude mobile edge computing system
CN122054238A