Method and apparatus for downlink resource allocation with swipt assistance

By constructing a resource allocation model based on neural networks, spectrum, power, and power offloading ratio are allocated to machine user equipment in SWIPT-assisted H2H/M2M coexisting cellular networks, solving the problem of poor resource allocation for machine user equipment and achieving high energy efficiency and high QoS resource allocation results.

CN115915454BActive Publication Date: 2026-04-17BEIJING INFORMATION SCI & TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING INFORMATION SCI & TECH UNIV
Filing Date
2022-10-09
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In SWIPT-assisted H2H/M2M coexisting cellular networks, machine user equipment receives poor resource allocation, resulting in poor performance and difficulty in meeting network energy efficiency and QoS requirements.

Method used

By acquiring communication link status observations between tolerant machine user equipment and machine equipment gateways, and utilizing a resource allocation model based on neural networks, resource allocation strategies are selected for tolerant machine user equipment, including the allocation of spectrum, power, and power split ratio, to optimize resource allocation and meet QoS constraints.

Benefits of technology

It achieves high energy efficiency and high QoS resource allocation, improves the overall performance of M2M devices, and meets the service quality needs of different types of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115915454B_ABST
    Figure CN115915454B_ABST
Patent Text Reader

Abstract

The application discloses a SWIPT-assisted downlink resource allocation method and device. The method comprises the following steps: obtaining a state observation value of a current environment state of a communication link between a tolerant machine user equipment and a machine equipment gateway; based on the state observation value of the current environment state, a resource allocation model constructed based on a neural network is used to select a resource allocation strategy for the tolerant machine user equipment; and based on the selected resource allocation strategy, the downlink resource is allocated to the tolerant machine user equipment; wherein the tolerant machine user equipment is a machine user equipment which needs to complete a transmission task of periodically generated load. The application solves the technical problems that machine user equipment with energy limited characteristics will generate excessive energy consumption when establishing a communication link, and various Internet of Things equipment resource allocation is uneven, and especially, machine user equipment with relatively low QoS demand obtains poor resource allocation and thus has poor performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communications, and more specifically, to a SWIPT-assisted downlink resource allocation method and apparatus. Background Technology

[0002] The rapid development of Internet of Things (IoT) technology has given rise to data-driven fields such as interconnected industries, intelligent transportation systems, and smart cities, which are driving human society to gradually become digital.

[0003] Machine-to-machine (M2M) communication, with its ability to operate without direct human intervention and through near-instantaneous wireless connectivity, is a major driver of digital transformation. With the exponential growth in the number of connected machines, M2M communication will offer new possibilities for serving citizens in scenarios such as home, entertainment, healthcare, and work, and is expected to become a key driver in realizing the vision of a fully connected world.

[0004] Unlike traditional person-to-person (H2H) communication, machine-to-machine (M2M) communication services possess unique characteristics in terms of network energy efficiency, data transmission volume, and high-reliability transmission in some mission-critical applications. While low-power wide-area networks (LPWANs) have been developed specifically for M2M communication in the Internet of Things (IoT), providing extremely low-power, long-distance transmission for M2M devices, these networks can only offer very low transmission rates, making it difficult to meet the QoS requirements of most mission-critical M2M services. Therefore, cellular networks, due to their greater bandwidth and stronger network scalability, can meet the access needs of various services and provide high data transmission rates, and are considered a key factor in deploying M2M communication. However, cellular networks themselves are too energy-intensive and costly for many M2M communication applications.

[0005] To mitigate the negative impact of cellular networks on M2M communication, Wireless Powered Communication (SWIPT) is a promising technology that can significantly improve the energy efficiency of M2M devices in H2H / M2M coexisting cellular networks. While traditional energy harvesting technologies allow mobile devices or base stations to extract energy from natural environments such as wind or sunlight, their efficiency is severely constrained by geographical location and weather conditions. SWIPT, as a radio frequency-based energy harvesting technology, allows application devices to obtain relatively controllable and stable energy. Typically, this technology converts interference signals into electrical energy, meaning that strong interference environments, while reducing system throughput, also result in greater energy harvesting. Furthermore, SWIPT-assisted H2H / M2M coexisting cellular networks face several challenges in meeting diverse Quality of Service (QoS) requirements, such as intra-layer and inter-layer interference control, spectrum subband selection, and the trade-off between system information transmission rate and energy harvesting.

[0006] There is currently no effective solution to the above problems. Summary of the Invention

[0007] This application provides a SWIPT-assisted downlink resource allocation method and apparatus to at least solve the technical problem of poor performance due to poor resource allocation for machine user equipment.

[0008] According to one aspect of the embodiments of this application, a SWIPT-assisted downlink resource allocation method is provided, comprising: obtaining state observations of the current environmental state of the communication link between a tolerant machine user equipment and a machine device gateway; selecting a resource allocation strategy for the tolerant machine user equipment based on the state observations of the current environmental state using a resource allocation model constructed based on a neural network; and allocating downlink resources to the tolerant machine user equipment based on the selected resource allocation strategy; wherein the tolerant machine user equipment is a machine user equipment that needs to complete the transmission task of periodically generated payloads.

[0009] According to another aspect of the embodiments of this application, a SWIPT-assisted downlink resource allocation apparatus is also provided, comprising: an acquisition module configured to acquire state observations of the current environmental state of the communication link between a tolerant machine user equipment and a machine device gateway; a selection module configured to select a resource allocation strategy for the tolerant machine user equipment based on the state observations of the current environmental state and using a resource allocation model constructed based on a neural network; and an allocation module configured to allocate downlink resources to the tolerant machine user equipment based on the selected resource allocation strategy; wherein the machine user equipment is required to complete the transmission task of periodically generated payloads.

[0010] In this embodiment, based on the state observations of the current environmental state, a resource allocation model based on a neural network is used to select a resource allocation strategy for tolerant machine user equipment, thereby solving the technical problem that machine user equipment receives poor resource allocation and thus performs poorly. Attached Figure Description

[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0012] Figure 1 This is a flowchart of a SWIPT-assisted downlink resource allocation method according to an embodiment of this application;

[0013] Figure 2 This is a flowchart of a method for constructing a downlink resource allocation model according to an embodiment of this application;

[0014] Figure 3This is a flowchart of another SWIPT-assisted downlink resource allocation method according to an embodiment of this application;

[0015] Figure 4 This is a flowchart of another SWIPT-assisted downlink resource allocation method according to an embodiment of this application;

[0016] Figure 5 This is a schematic diagram of a SWIPT-assisted downlink resource allocation device according to an embodiment of this application;

[0017] Figure 6 This is a schematic diagram of the structure of a SWIPT-assisted downlink resource allocation system according to an embodiment of this application;

[0018] Figure 7 This is a schematic diagram comparing the changes in the number of M2M devices and the total energy efficiency of M2M devices according to embodiments of this application;

[0019] Figure 8 This is a schematic diagram comparing the change in the number of tolerant M2M devices with the QoS requirement satisfaction rate of H2H users according to an embodiment of this application;

[0020] Figure 9 This is a schematic diagram comparing the change in the number of tolerant M2M devices with the probability of critical M2M link interruption according to an embodiment of this application;

[0021] Figure 10 This is a schematic diagram comparing the change in the number of tolerant M2M devices with the load transmission completion rate according to an embodiment of this application. Detailed Implementation

[0022] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0023] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0024] Example 1

[0025] According to an embodiment of this application, a SWIPT-assisted downlink resource allocation method is provided, such as... Figure 1 As shown, the method includes:

[0026] Step S102: Obtain the current environmental status observation value of the communication link between the tolerant machine user equipment and the machine equipment gateway.

[0027] Before obtaining the current environmental status observations of the communication link between the tolerant machine user equipment and the machine equipment gateway, it is necessary to classify the machine user equipment into tolerant machine user equipment and critical machine user equipment based on the service type of the machine user equipment. Among them, critical machine user equipment is machine user equipment whose transmission reliability requirements are higher than the preset threshold, and tolerant machine user equipment is machine user equipment that needs to complete the transmission task of periodically generated payloads.

[0028] In addition to considering existing conventional H2H (Human-to-Human) user equipment (also known as human user equipment or H2H users), this embodiment further subdivides machine-type communication devices (MTCDs) into tolerant machine user equipment (also known as tolerant M2M devices or tolerant M2M users) and critical machine user equipment (also known as critical M2M devices or critical M2M users). This allows for the selection of different resource allocation strategies based on the service characteristics of different machine user equipment, thereby enabling more rational allocation of resources to machine user equipment.

[0029] In one example, the machine user equipment (MTCD) may be equipped with Simultaneous Wireless Information and Power Transfer (SWIPT) to enable it to acquire power from the radio frequency environment and simultaneously decode information. By configuring SWIPT, the MTCD can simultaneously acquire power and decode information.

[0030] Step S104: Based on the state observations of the current environmental state, a resource allocation strategy is selected for the tolerant machine user equipment using a resource allocation model built on a neural network.

[0031] In one example, it can be like this: Figure 2 As shown, a resource allocation model is constructed based on the following method:

[0032] Step S1042: Construct the state function.

[0033] The state function is a set of state observations. For example, the state function can be constructed based on the channel gain information of the communication link in each spectrum subband, the interference power experienced by the communication link in each spectrum subband, the remaining payload and transmission time of the communication link, the current iteration number, and a greedy factor representing the current environmental exploration rate of the communication link.

[0034] Step S1044: Construct the action function.

[0035] The action function is a set of downlink spectrum resources, transmit power level, and power shunting ratio.

[0036] Step S1046: Construct the reward function.

[0037] The reward function is constructed based on the balance between resource allocation optimization objectives and QoS constraints.

[0038] In some examples, the reward function can be constructed as follows.

[0039] First, based on the remaining payload and transmission time of tolerant machine user equipment (MUE), the penalty term of the reward function provided by the tolerant MUE is determined. Next, based on the outage probability of critical MUE, the penalty term of the reward function provided by the critical MUE is determined. Then, based on the signal-to-interference-plus-noise ratio (SIR) of human user equipment (HUE), the penalty term of the reward function is determined. Finally, based on the total energy efficiency of the communication links of all MUEs, the penalty term of the reward function provided by the tolerant MUE, the penalty term of the reward function provided by the critical MUE, and the penalty term of the reward function of the human user equipment, a reward function is constructed.

[0040] In this embodiment, for QoS constraints, the threshold for H2H user equipment signal-to-interference-plus-noise ratio, the probability of interruption of critical M2M communication links, and the effective load transmission probability of M2M communication links are explicitly modeled and solved, and the corresponding reward function penalty terms are obtained, thereby making the reward value of the reward function more accurate and improving the rationality of resource allocation.

[0041] In other examples, the reward function can be constructed as follows.

[0042] First, set Quality of Service (QoS) constraints. For example, set the QoS constraint for human user equipment to be that the signal-to-interference-plus-noise ratio (SINR) is greater than a set minimum threshold; set the QoS constraint for tolerant machine user equipment to set the success rate of transmitting a payload of a preset size V within a time constraint T to be higher than a set success rate threshold; and set the QoS constraint for critical machine user equipment to set the outage probability to be no higher than a set outage threshold.

[0043] Next, the total energy efficiency (EE) achieved by tolerant machine user equipment (M2M) and critical machine user equipment (CMA) is used as the resource allocation optimization target. This embodiment aims to maximize the EE of the M2M communication link and allocates joint spectrum, transmit power, and power shunting ratio (PS ratio) while considering user QoS constraints.

[0044] Finally, a reward function is constructed based on the resource allocation optimization objective and QoS constraints.

[0045] Step S1048: Construct an experience replay pool, wherein the experience replay pool is used to store data and provide data for training the neural network.

[0046] An experience replay pool is constructed to store training data for training the resource allocation model. The training data includes state observations, reward values, and selected resource allocation strategies for the current and next time steps. The neural network comprises a training network and a target network. The training network is trained in each iteration using data randomly drawn from the experience replay pool and employing stochastic gradient descent. The target network is a fixed neural network whose parameters are updated periodically to reflect the parameters of the training network at the current time step.

[0047] Step S1049: Construct and train a resource allocation model based on the state function, action function, reward function, and experience replay pool.

[0048] Step S106: Based on the selected resource allocation strategy, allocate downlink resources to the tolerant machine user equipment.

[0049] After obtaining the current environmental status observations, or after allocating downlink resources, the data related to resource allocation can be put into the experience replay pool corresponding to each communication link for use as training data.

[0050] For example, calculate the total energy efficiency value of all machine user equipment and the QoS level of each machine user equipment and human user equipment at the current moment. Put the total energy efficiency value, QoS level, interference power received by the communication link in each spectrum sub-band, selected resource allocation strategy, state observation value, reward output by the reward function, and prior information received by the machine device gateway into the experience replay pool of each communication link, and use it as training data for training the resource allocation model.

[0051] This embodiment addresses the energy resource allocation problem in SWIPT-assisted H2H / M2M (Machine-to-Machine) coexistence cellular networks. By allocating appropriate resources (i.e., spectrum, power, and power splitting ratio) to tolerant machine user equipment, high energy efficiency (EE) is achieved while ensuring QoS requirements.

[0052] This embodiment provides a stable training process based on behavior tracking states, and the shared reward function enables agents to work collaboratively in a distributed manner. Furthermore, this embodiment provides a precise mathematical expression for QoS constraints, thereby reducing the computational complexity of the reward function.

[0053] Furthermore, to support different QoS requirements, this embodiment models the EE optimization problem of multiple M2M communication links as a multi-agent problem, where multiple tolerant M2M communication links share the spectrum occupied by critical M2M and H2H communication links. Moreover, this embodiment designs state functions, action functions, and reward functions to enable the agents to adapt to the dynamic network environment, thereby achieving the desired optimization effect.

[0054] Example 2

[0055] According to embodiments of this application, a SWIPT-assisted QoS-based downlink resource allocation method is also provided, such as... Figure 3 As shown, the method includes:

[0056] Step S302: Initialize the resource allocation system.

[0057] Based on pre-allocated spectrum subbands and transmit power, prior information such as channel conditions and SINR or outage probability of the spectrum subbands occupied by H2H users and critical M2M users is obtained. More specifically, initialization refers to determining the initial state of each tolerant M2M link at time t.

[0058] Step S304: Build the training neural network and the target neural network for each agent.

[0059] Treat each tolerant M2M link as an agent, and build a training neural network Q(s, a; ω) and a target neural network Q for each agent. - (s, a; ω) - ), where ω and ω - These represent the parameters of the training network and the target network, respectively. The input layer of the network is s, which is the state observation value of each tolerable link. The output layer corresponds to the estimated reward value for each action selection a, denoted as the Q value.

[0060] An experience replay pool D is established for each agent to store the state observations S of its own link at the current time t and the next time t+1. t and S t+1 Action selection A at current time t t And the reward value R received in the next moment. t+1 , with topology (S t A t R t+1 S t+1 ) represents the data stored in D at time t.

[0061] Step S306: Design the state function, action function, and reward function for the agent.

[0062] Each agent establishes a link communication energy efficiency optimization model based on prior information and action choices made according to local state observations. Specifically, this includes defining constraints on transmit power, power split ratio, and QoS constraints.

[0063] The QoS constraints are modeled according to the actual characteristics of the services as follows: H2H services are usually based on voice and Internet communication traffic, which places high demands on data transmission rate. Therefore, the QoS constraint for H2H services is modeled as a SINR value greater than a set minimum threshold. Most M2M services are event-driven, and their traffic is usually generated periodically at different frequencies according to the needs of the M2M service itself. Therefore, the QoS constraint for tolerant M2M services is modeled as the transmission success rate of data packets of size V within a time constraint T. In addition, for critical M2M services, which often have strict requirements for latency and transmission reliability, the QoS constraint for critical M2M services is modeled as an interruption probability that must not exceed a set interruption threshold.

[0064] Step S308: Train the training network Q and allocate network resources.

[0065] Based on the link communication energy efficiency optimization model, the tolerant M2M link continuously randomly extracts data from a large amount of data stored in the experience replay pool to train the training network Q. The error between the parameters of the training network and the target network is reduced by stochastic gradient descent. Finally, after the resource allocation model training converges, a resource allocation scheme that satisfies the constraints and QoS constraints can be made.

[0066] Furthermore, to stabilize the training environment, this implementation employs a low-dimensional mapping to high-dimensional behavior tracking method, enabling each tolerant M2M link to understand the current state of its learned experience when faced with training data randomly selected from the experience replay pool. Moreover, to improve overall energy efficiency, this embodiment shares a reward function to encourage each tolerant M2M link to explore the environment cooperatively and make intelligent decisions that maximize overall performance improvement.

[0067] In one embodiment, the specific process of resource allocation may include: the tolerant M2M link selects the spectrum subband with the best channel conditions for multiplexing based on the received prior information, and allocates appropriate transmit power and power split ratio to the tolerant machine devices in the link. At each moment (1ms), the channel conditions of the spectrum subband in the network will change due to channel fading. In addition, the resource allocation selection made by each tolerant M2M link at each moment will also affect the QoS indicators of various types of users in the network. Tolerant M2M uses these changed QoS metrics—namely, SINR of H2H users, outage probability of critical M2M users, remaining payload and transmission time of tolerant M2M users, and achievable energy efficiency of all M2M users—to determine the merits of the resource allocation scheme at any given moment. It stores the current channel conditions, QoS metrics, interference, resource selection scheme, and achievable energy efficiency in an experience replay pool as training data for network Q in each tolerant M2M link. As environmental data changes, the training data sample increases. By randomly sampling data from the experience replay pool for training, data correlations can be effectively eliminated.

[0068] For the tasks of spectrum subband selection, transmit power, and power split ratio allocation for tolerant M2M links, a distributed implementation scheme is adopted. At each time t, each tolerant M2M link selects the spectrum subband with the best channel conditions and an appropriate transmit power and power split ratio based on the aforementioned prior information and its own training network, and communicates using this resource allocation scheme. The overall energy efficiency of all M2M devices and the QoS indicators of other users in the network are then obtained, referred to as the reward function. All tolerant M2M links in the network share the same reward function; therefore, the resource allocation method proposed in this application encourages cooperative behavior among links to achieve the goal of maximizing the overall energy efficiency of M2M.

[0069] The resource allocation task further includes: at each time t, based on the prior information received by the machine device gateway m in the cluster, updating the available environment data set of the tolerant M2M links m and n, and designing its state set as follows:

[0070]

[0071] in, It contains channel gain information for links m and n in each spectral subband. This describes the magnitude of interference power experienced on each spectral subband of links m and n, V. m,n T m,n Representing the remaining payload and transmission time for links m and n respectively, {QoS h} h∈H , {QoS s}s∈S Let be binary variables, representing the QoS indicators for H2H users and critical M2M users respectively. That is, if the QoS constraints for all H2H users are satisfied, then {QoS} h} h∈H =1, otherwise 0, {QoS s} s∈S The same binary variable is used to represent whether the QoS constraints of all critical M2M links are met.

[0072] Furthermore, e represents the current iteration number, ∈ is the greedy factor, and represents the current environment exploration rate of links m and n. In this embodiment, these two items are added to the state set to track the environment state when the iteration number e and exploration probability ∈ are in the training process. During the training process, the resource allocation strategy of each tolerant M2M link will change with the resource allocation strategy of other links. Therefore, from the perspective of a single link, the network environment is constantly changing. Moreover, since the experience is randomly extracted from the experience replay pool, the training data extracted at the current moment does not reflect the network environment at this time, but is very likely to be outdated data. This makes it difficult for each link to know the state of the experience and the strategy selection of each link when facing the training data. The value function that truly represents the strategy selection of each link is a high-dimensional neural network parameter, which cannot be used as experience data for each link to learn. Therefore, the design of adding the iteration number e and the greedy factor ∈ to the experience data is made to map the high-dimensional neural network parameter, so that the tolerant M2M link can track the environment state when facing the selected experience data and achieve the purpose of stable training.

[0073] Based on the above environmental information, the tolerant M2M links m and n at time t utilize a greedy strategy: randomly selecting the spectrum sub-band, transmit power, and power split ratio with probability ∈ , or selecting the optimal action selection strategy given by the currently trained network with probability 1- ∈ . The action (resource) set is defined as:

[0074]

[0075] in Indicates the selection of spectrum subbands. These represent the selection of transmit power and power shunt ratio, respectively. L represents the maximum transmit power, and Z represents the power and power shunt ratio division levels, respectively.

[0076] In this embodiment, it is assumed that the available spectrum subbands in the network are indexed by K = H∪S, where H = {1, 2...H} represents HUE, and S = {1, 2...S} represents Critical Machine User Equipment (CMTCD). Furthermore, M = {1, 2...M} represents MTCG, and each Tolerant Machine User Equipment (TMTCD) under the control of an MTCG is represented by N = {1, 2...N}, where all TMTCDs and CMTCDs are equipped with SWIPT technology, denoted by the set D, i.e., D = S∪N.

[0077] Tolerant M2M link m,n selects actions based on time t When establishing a communication link, based on spectrum subband selection, the SINR of users with high QoS requirements is calculated in the following two ways:

[0078]

[0079]

[0080] In the above formula, the first term of the denominator This indicates the interference experienced by H2H users or critical M2M users in spectrum subband k. This refers to the cross-cluster interference and intra-cluster interference experienced when reusing the k-th spectrum subband. The second term σ is a binary variable representing the occupancy of spectrum subbands. 2 P represents the power of additive white Gaussian noise. BS,h and P m,s ρ represents the transmit power allocated by the base station and the machine equipment gateway to H2H users and critical M2M users, respectively. m,s P represents the power shunt ratio allocated by the machine equipment gateway to critical M2M users. BS,h This indicates the transmit power allocated by the base station to H2H users. This indicates the channel gain between the base station and the H2H user. This indicates the channel gain between the machine / equipment gateway and the H2H user. This represents the channel gain between the gateway and critical M2M users within the same cluster. After calculating the SINR for H2H users and critical M2M users, the following formula is used:

[0081]

[0082]

[0083] To determine whether the QoS requirements of H2H users and critical M2M users are guaranteed, respectively. and These represent the SINR thresholds for H2H services and key M2M services, respectively. This refers to the maximum tolerable interruption probability. If the QoS constraints of the H2H service are met, then the environmental state information {QoS} h} h∈H =1, otherwise {QoS h} h∈H =0, correspondingly, {QoS h} s∈S The same indicator is used to represent the QoS satisfaction of all critical M2M users.

[0084] For a tolerant M2M link itself, its SINR value is expressed by the following formula:

[0085]

[0086] in, The interference experienced by the m-th and n-th tolerant M2M links can be categorized into two cases based on the different occupants of the reused spectrum subbands:

[0087]

[0088]

[0089] Among them, P m,s This indicates the transmit power allocated by the machine equipment gateway to critical M2M users. This represents the channel gain of a tolerant M2M link m,n on spectrum subband k. Indicates the spectrum subband allocation status indicator. This represents the transmit power of the m′-th machine device gateway in spectrum subband k. This represents the channel gain of the m′-th machine-to-machine gateway and the tolerant M2M device n in spectrum subband k. Indicates the spectrum subband allocation status indicator. This indicates that the tolerant M2M links m and n communicate on spectrum subband k. This indicates that the tolerant M2M links m and n do not communicate on spectrum subband k. This represents the transmit power allocated by machine device gateway m to tolerant M2M user n, where M represents the total number of machine device gateways and N represents the total number of tolerant M2M devices within the control range of a machine device gateway.

[0090] Considering that the information processed by tolerant M2M services is often data packets generated periodically at different frequencies according to their own service characteristics, their QoS is modeled as the success rate of data packet transmission within a time budget T:

[0091]

[0092] in, This represents the information transmission rate of links m and n on spectrum subband k at different times t, where B is the bandwidth of each spectrum subband, and V represents the periodically generated M2M payload in the above formula, in bits. T This is the channel coherence time.

[0093] The energy efficiency (EE) of all SWIPT-assisted M2M devices can be expressed as the ratio of spectral efficiency to energy consumption, and all M2M devices are represented by D = S∪N:

[0094]

[0095] in

[0096]

[0097] This represents the achievable spectral efficiency of all M2M devices;

[0098]

[0099] This represents the energy consumption of all M2M devices. Where P... C This represents the power consumption of all M2M links.

[0100]

[0101] Let θ represent the energy harvesting capacity of the M2M device, where θ∈[0,1] is the energy conversion efficiency.

[0102] In conclusion, the M2M link energy efficiency optimization model of this application is summarized as follows:

[0103]

[0104]

[0105]

[0106]

[0107]

[0108]

[0109]

[0110]

[0111] in and These represent the power shunting ratio transmit power allocation and spectrum subband allocation strategies, respectively. Conditions (1a), (1b), and (1c) represent the QoS constraints for H2H users, critical M2M users, and tolerant M2M users, respectively. Condition (1d) specifies that the power shunting ratio of M2M devices using SWIPT is not greater than 1, and (1e) ... It is a binary variable, the tolerant M2M link m, n is allocated a spectrum subband k, which is represented as 1 otherwise as 0, in (1f) The upper limit of the transmit power of the tolerant M2M link is specified, and (1g) is used to constrain each tolerant M2M link to select at most one spectrum subband for communication.

[0112] The channel gain model for each communication and interference link (where g represents the channel gain) is summarized as follows:

[0113] use Calculate the channel gain of the M2M communication link m and n, formed between the machine gateway m and the tolerant machine n within the same cluster, on the spectrum subband k. Where X m,n β m,n The terms representing path loss and shadowing fading are called large-scale fading. These two parameters are frequency-independent and remain constant over a long period of time. This indicates frequency-dependent fast fading.

[0114] use Calculate the channel gain of the M2M communication link m between the machine device gateway m and the critical machine device s within the same cluster, in the spectrum subband k.

[0115] use Calculate the channel gain of the H2H communication link BS between base station BS and user equipment h in spectrum subband k.

[0116] use Calculate the interference channel gain of base station BS and tolerant M2M link m, n on spectrum subband k.

[0117] use Calculate the interference channel gain of machine device gateway m′ and tolerant M2M link m,n in spectrum subband k between different clusters.

[0118] use Calculate the interference channel gain of machine device gateway m′ and critical M2M link m,s in spectrum subband k between different clusters.

[0119] use Interference channel gain of computer equipment gateway m and human type user equipment h in spectrum subband k.

[0120] In order to ensure that multiple tolerant M2M links can move towards the goal of maximizing overall energy efficiency during training while also ensuring the QoS constraints of various types of users, this application incorporates the overall EE and QoS constraints into the design of the reward function.

[0121] Obviously, the SINR value of H2H users at each time t can be easily obtained from the power and interference received by the user. However, the QoS constraint of M2M users is a probability value. At each time, each agent needs to randomly generate a large number of simulated channels to obtain the required probability value, which consumes a lot of computing resources and slows down the convergence speed of the algorithm. In order to solve this problem, this application transforms such QoS constraints into precise explicit expressions.

[0122] First, use U n This indicates the remaining transmission time for a tolerant M2M user while the load is not yet fully transmitted; once the transmission is complete, the U... n This is set to a constant value. Therefore, at time t, the rewards and penalties associated with tolerant M2M users are set as follows:

[0123]

[0124] Through this transformation, each agent can simultaneously consider the impact of remaining transmission time and load transmission rate during the training process.

[0125] Secondly, this application uses theoretical analysis and mathematical transformations to obtain precise values ​​for the outage probability of critical M2M devices. When the s-th critical M2M device shares the k-th spectrum subband with tolerant M2M devices in different clusters, the outage probability can be replaced as follows:

[0126]

[0127] According to the lemma: Suppose z1, ..., z n The mean is For an independent exponentially distributed random variable, the formula can be obtained:

[0128]

[0129] Where c is a positive constant. Therefore, based on Rayleigh fading characteristics and the above equation, the interruption probability expression can be rewritten as:

[0130]

[0131] in Depending on the source of interference, it can be categorized as m or m′, representing inter-cluster interference and intra-cluster interference, respectively. Furthermore, It can also include both m and m′, indicating that the interference received includes both inter-cluster interference and intra-cluster interference. It is worth noting that when the interference received is only inter-cluster interference, the cumulative terms in the above equation... Only keep Π n Item. Accordingly, Let be the interference experienced by the s-th CMTCD in the spectral subband k, and let be the interference experienced by the s-th CMTCD in the spectral subband k.

[0132]

[0133] In addition, when When both m and m′ are included, the calculation of the interruption probability requires continuous multiplication of all terms in the above formula, resulting in a large computational resource overhead.

[0134] To further accelerate algorithm convergence, the above expression is transformed into the following form:

[0135]

[0136] The above formula allows for the precise outage probability value of each critical M2M device with relatively low computational resource consumption. To further differentiate the effectiveness of selected resource allocation strategies during training, this application further sets... The QoS satisfaction of critical M2M devices within time slot t is expressed as follows:

[0137]

[0138] Where ps is the interrupt probability value obtained using the interrupt probability expression derived above, and similarly... Let QoS satisfaction of 2H users within time slot t of H be represented:

[0139]

[0140] In summary, this application sets the global reward function as follows:

[0141]

[0142] μ is a positive real number to balance rewards and penalties. By designing that all agents share the same reward function, agents can cooperate to explore the location environment and learn resource allocation strategies that maximize overall energy efficiency. At the same time, each penalty term ensures that agents consider the QoS constraints of various users when allocating resources.

[0143] After each iteration, each agent randomly draws a small batch of data from the experience replay pool to train the training network Q, and updates the training network parameters ω using stochastic gradient descent to approximate the target network parameters ω. -It then passes its own network parameters to the target network after a period of iteration.

[0144] The target network is represented as:

[0145]

[0146] Where 0≤γ≤1 is the discount factor, the magnitude of which determines the importance of the current reward / penalty value during the iteration process; the smaller the value, the more important the current reward / penalty value is, and vice versa. S t+1 Let a' represent the state observation at the next moment, and let a' represent the state S. t+1 The optimal action choice made below, ω - This represents the target network parameters.

[0147] In this embodiment, the loss function is expressed as:

[0148] L(θ)=E[(Q - -Q(S t A t ;ω)) 2 ]

[0149] Where L(θ) represents the loss function, E represents the expected value, Q represents the training network value function, and S... t Let A represent the state observation at time t. t State S t The action selection is given below, where ω represents the parameters of the network being trained.

[0150] In this embodiment of the application, SWIPT technology is provided for M2M devices in H2H / M2M coexisting cellular networks. Under the complex interference environment caused by spectrum sharing, high EE performance is achieved and QoS requirements of all types of users are ensured by allocating spectrum, transmit power and power split ratio resources to tolerant M2M links.

[0151] Furthermore, in this embodiment of the application, in addition to traditional H2H services, the SWIPT-assisted H2H / M2M coexistence cellular network also distinguishes M2M service types into critical services and tolerable services. M2M devices exist in clusters and establish communication links with the local machine device gateway. All M2M devices are equipped with SWIPT technology, enabling them to perform energy harvesting and information decoding simultaneously. OFDM spectrum subbands are pre-allocated to H2H users and critical M2M users in the network, while the remaining tolerable M2M users reuse the same spectrum resources.

[0152] In the resource allocation method proposed in this application, the energy efficiency optimization problem is modeled as a multi-agent problem. Each tolerant M2M link is treated as an agent, and a set of system state, action, and reward function designs is proposed. Through the designed behavior-tracking-based state set and reward function, each agent can more efficiently select appropriate spectrum sub-bands, transmit power, and power split ratio resources from the action set in a distributed cooperative manner during interactions with the unknown network environment, thereby optimizing the overall M2M energy efficiency. Furthermore, this application further reduces the computational complexity of the model training process by expressing the QoS requirements (QoS constraints) of each user type using precise mathematical formulas.

[0153] Example 3

[0154] According to embodiments of this application, another SWIPT-assisted QoS-based downlink resource allocation method is also provided, such as... Figure 4 As shown, the method includes:

[0155] Step S402, Initialization.

[0156] The resource allocation system is initialized, and based on the pre-allocated spectrum subbands and transmit power, the channel conditions of the spectrum subbands occupied by H2H links and critical M2M links, as well as prior information such as SINR or outage probability, are obtained.

[0157] Step S404: Build a neural network and create an experience replay pool for each agent.

[0158] Each tolerant M2M link is treated as an agent, and a training neural network and a target neural network are built. An experience replay pool is established for each agent.

[0159] In this embodiment, the resource allocation problem is first modeled as a multi-agent intelligent decision-making problem, where each tolerant M2M link is considered an agent to interact with the unknown network environment and optimize resource allocation strategies. Then, a resource allocation method is used to construct a training network, a target network, and an experience replay pool for each agent, where the experience replay pool stores data used to train the training network.

[0160] Step S406: Design the state function, action function, and reward function for the agent.

[0161] Each agent establishes a link communication energy efficiency optimization model based on the prior information and the action selection made according to the local state observation, and stores the M2M system energy efficiency, various QoS indicators and interference data returned by the link energy efficiency optimization model into the experience replay pool.

[0162] Designing state functions, action functions, and reward functions for an agent involves the following steps:

[0163] S3-1, a low-dimensional mapping to high-dimensional trajectory tracking method is designed in the state set: the iteration number and greedy factor are added to the state set so that the agent can track the environment state when it is at the iteration number e and the exploration probability ∈ during the training process when extracting training data. This solves the problem of non-stationarity of the training environment caused by the multi-agent setting.

[0164] S3-2, the agent establishes a link communication energy efficiency optimization model based on state observations and selected actions. Specifically, this includes defining constraints on transmit power, power splitting ratio, and QoS requirements. The QoS requirements are modeled according to the actual characteristics of the service as follows: H2H services are typically based on voice and internet communication traffic, which places high demands on data transmission rates. Therefore, the QoS requirement for H2H services is modeled as a SINR value greater than a set minimum threshold. Most M2M services are event-driven, and their traffic is typically generated periodically at different frequencies based on the needs of the M2M service itself. Therefore, the QoS requirement for tolerant M2M services is modeled as the transmission success rate of data packets of size V within a time constraint T. Furthermore, for critical M2M services, which often have strict latency and transmission reliability requirements, the QoS requirement for critical M2M services is modeled as an interruption probability not exceeding a set threshold.

[0165] S3-3, the reward function is designed as a combination of main reward terms and penalty terms, and all agents share the same reward function. This encourages cooperation among agents to explore the optimal resource allocation strategy while also taking into account the QoS requirements of various users. In addition, in order to reduce computational resource consumption and speed up convergence, the various QoS requirements are precisely expressed mathematically in each penalty term of the reward function.

[0166] S3-4: Calculate the total EE of the current time slot network M2M devices and the QoS level of various users through the link optimization model. Store the calculated EE, QoS level, interference of the tolerant M2M link, selected action, state observation, reward and prior information received by the machine device gateway into the experience replay pool for use as training data for training the network.

[0167] Step S408: Train the resource allocation model and allocate resources.

[0168] Based on the link communication energy efficiency optimization model and the experience replay pool, the agent continuously randomly extracts data from the large amount of data stored in the experience replay pool to train the training network Q. The error between the parameters of the training network and the target network is reduced by stochastic gradient descent. Finally, after the model training converges, a resource allocation scheme that meets the constraints and QoS requirements can be made.

[0169] By randomly sampling batches of data from the experience replay pool to train the training network, it is possible to learn from both current and past experiences simultaneously, effectively eliminating data correlations. By asynchronously updating the target network parameters, the target network parameters remain unchanged over a period of time, thus making the algorithm update more stable.

[0170] SWIPT technology employs two receiver designs: "time switching" and "power splitting." "Time switching" alternates between the duration of information decoding and energy harvesting, while "power splitting" divides the received energy into information decoding and energy harvesting portions using a power split ratio. This embodiment uses the "power splitting" design, therefore power split ratio allocation is introduced into the action set design of the resource allocation method. In other embodiments, "time switching" can also be used, replacing power split ratio allocation with time allocation.

[0171] In existing H2H / M2M coexisting cellular networks, the complex network structure and variable resource allocation tasks lead to poor performance of traditional resource allocation schemes. Furthermore, these schemes neglect the diverse service types of machine user equipment and interference caused by spectrum sharing. Additionally, network models considering energy consumption (EH) do not utilize swapping to obtain stable and reliable energy. In practical applications, existing technologies fail to accurately model the QoS requirements of each user type mathematically, often resulting in poor performance for user equipment with low QoS demands. Moreover, traditional algorithms suffer from high computational complexity and poor performance when solving optimization problems with nonlinear constraints, while the centralized training of the proposed intelligent algorithm exhibits slow convergence.

[0172] Compared with existing intelligent resource allocation methods, the embodiments of this application introduce SWIPT-assisted M2M devices within the network structure to obtain stable and reliable power from the radio frequency environment.

[0173] Furthermore, in implementing the resource allocation method, a multi-agent cooperation and distributed execution approach is chosen. Since the resource allocation scheme of multi-agent distributed execution can take into account the situation of each agent and make the most suitable resource allocation strategy, while the resource allocation strategy made by single-agent centralized execution tends to be applicable to all agents in the network, this embodiment greatly reduces the service load of the base station, and the resource allocation strategy of each agent can be updated in each iteration, which greatly improves the training convergence speed and achieves better performance.

[0174] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0175] Example 4

[0176] According to embodiments of this application, a SWIPT-assisted QoS-based downlink resource allocation device is also provided, such as... Figure 5 As shown, the device includes an acquisition module 542, a selection module 544, and an allocation module 546.

[0177] Based on the service type of the machine user equipment, the machine user equipment is divided into tolerant machine user equipment and critical machine user equipment; wherein, the critical machine user equipment is machine user equipment whose transmission reliability requirements are higher than a preset threshold, and the tolerant machine user equipment is machine user equipment that needs to complete the transmission task of periodically generated payloads.

[0178] The acquisition module 542 is configured to acquire status observations of the current environmental state of the communication link between the tolerant machine user equipment and the machine equipment gateway.

[0179] Selection module 544 is configured to select a resource allocation strategy for the tolerant machine user equipment based on the state observations of the current environmental state and using a resource allocation model built on a neural network.

[0180] The allocation module 546 allocates downlink resources to the tolerant machine user equipment based on the selected resource allocation strategy.

[0181] The downlink resource allocation device in this embodiment can implement the downlink resource allocation method in the above embodiments, therefore, it will not be described again here.

[0182] Example 5

[0183] This application provides a SWIPT-assisted QoS-based downlink resource allocation system, such as... Figure 6 As shown, it includes: a base station (BS) 52, a human user equipment (HUE) 60, a machine equipment gateway (MTCG) 54, and machine user equipment, wherein the machine user equipment includes critical machine user equipment (CMTCD) 56 and tolerable machine user equipment (TMTCD) 58.

[0184] The machine user equipment is equipped with a SWIPT. SWIPT is used to embed M2M communication, giving the machine user equipment the ability to obtain stable and reliable power from the radio frequency environment, as well as perform information decoding functions simultaneously.

[0185] The spectrum subband K = H∪S represents that each orthogonal spectrum subband is pre-allocated to a human user equipment or a critical machine user equipment, where H = {1, 2, ..., H} and S = {1, 2, ..., S} represents critical machine user equipment; in addition, M = {1, 2, ..., M} represents the number of machine device gateways or clusters; and tolerable machine user equipment is represented by N = {1, 2, ..., N}.

[0186] Base station 52 is used to establish H2H communication links with human user equipment 60, pre-allocate spectrum subbands and transmission power, and broadcast prior information such as channel gain information and QoS indicators of all H2H links to machine equipment gateway 54 within the coverage area.

[0187] The machine equipment gateway 54 is used to form a cluster with machine user equipment and establish multiple M2M communication links, and pre-allocate spectrum sub-bands, transmission power and power split ratio for key M2M links. One machine equipment gateway and several machine user equipment form a cluster.

[0188] The Machine Equipment Gateway (MTCG) 54 can be the downlink resource allocation device in the above embodiments, and can implement the downlink resource allocation method in the above embodiments. Therefore, it will not be described again here.

[0189] Example 6

[0190] Embodiments of this application also provide a storage medium configured to store program code for executing the downlink resource allocation method in the above embodiments.

[0191] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0192] Simulation test

[0193] This application sets the following parameters: base station coverage radius of 500m, machine-to-device (M2M) gateway coverage radius of 30m, 2 H2H user devices, 2 machine-to-device gateways, each machine-to-device gateway's coverage area includes one critical M2M device, and the number of tolerable M2M devices within the coverage area increases from 0 to 5. The system bandwidth is 4MHz, and the path loss is modeled as 128 + 37.6 log... 10 d, where d is in km, is modeled as a log-normal distribution with a standard deviation of 8 dB for shadow fading, and as Rayleigh fading for fast fading. The ambient noise power σ... 2 The transmit power P allocated by the base station to H2H user equipment is -114dBm. BS,h The transmit power P allocated by the machine equipment gateway to critical M2M user equipment is 30dBm. m,s It is 23dBm, and the power shunt ratio ρ m,s The maximum transmit power allocated to tolerant M2M user equipment is 0.8. The transmit power is 15dBm, the power shunt ratio segmentation level L and Z are 10 and 5 respectively, and the SINR threshold for H2H service is... The SINR threshold for critical M2M services is 7dB. The interruption probability is 5dB. The tolerance-based M2M service payload transmission time constraint T is 100ms, and the payload size V is 3×1024 bytes. Furthermore, considering the practical effect of finite-precision digital signal processing, the received SINR value for each user is specified to not exceed 30dB. The parameters required for the algorithm are set as follows: discount factor γ = 0.5, constant μ = 1 / 50, number of iterations 3000, and the greedy factor ∈ linearly decays from 1 to 0.01 in the first 80% of iterations. The neural network for each tolerance-based M2M link consists of 3 fully connected hidden layers, with the number of neurons in each layer equal to the number of selectable actions, i.e., K×L×Z = 200. The ReLU function is used as the activation function, and the RMSProp optimization algorithm updates the network parameters with a learning rate of 0.001.

[0194] The spectrum subband, transmit power, and power split ratio allocation method proposed in this application is named the Multi-Agent SWIPT-Assisted Adaptive Spectrum-Power-Ratio Allocation Method (MA-SWIPT-ASPRA) based on its characteristics, and is compared with three efficient intelligent resource allocation methods: (1) Single-Agent SWIPT-Assisted Adaptive Spectrum-Power-Ratio Allocation Method (SA-SWIPT-ASPRA), which is simply a change from the resource allocation method published in this application, which implements resource allocation decisions by distributing multiple M2M links simultaneously to a single M2M link, which implements resource allocation decisions asynchronously; (2) Non-SWIPT-Assisted Adaptive Spectrum-Power-Ratio Allocation Method (MA-Non-SWIPT-ASPRA), which is simply a change from the resource allocation method published in this application, which removes the SWIPT function; (3) Q-Learning-Based SWIPT-Assisted Spectrum-Power-Ratio Allocation Method (QL-SWIPT-ASPRA), which is the most classic intelligent resource allocation scheme based on reinforcement learning.

[0195] Figure 7 The graph illustrates the change in the total energy efficiency of M2M devices as the number of connected M2M devices in the system increases. As shown in the graph, the total energy efficiency of M2M devices first increases and then decreases with the increase in the number of M2M devices. The reasons for this are as follows:

[0196] When only critical M2M devices exist in the network, i.e., the number of M2M devices is 2, the total energy efficiency achieved by both the SWIPT and non-SWIPT schemes is almost the same. This is because there is no interference caused by spectrum sharing in the network at this time. The critical M2M devices have strong channel gain due to their proximity to the local machine device gateway, thus obtaining a high SINR value. In addition, in the SWIPT scheme, the critical M2M devices use most of their energy for information decoding and only a small portion of their energy for collection. The large order of magnitude difference makes the energy efficiency performance almost indistinguishable.

[0197] When the number of M2M devices in the network increases from 2 to 6, the total energy efficiency achieved by each scheme increases accordingly and reaches its maximum value when the number of M2M devices is 6. This is because spectrum reuse can improve spectrum efficiency and thus further improve network energy efficiency. When the number of M2M devices is 6, there are 4 tolerant M2M links that reuse the spectrum resources occupied by 2 H2H links and 2 critical M2M links. Each scheme can allocate a pre-occupied spectrum subband to each tolerant M2M link without generating additional intra-cluster interference. Therefore, the network spectrum efficiency reaches its maximum at this time, and the highest energy efficiency value is achieved.

[0198] When the number of M2M devices in the network increases from 6 to 12, the energy efficiency achieved by each scheme decreases. This is because excessive spectrum reuse will generate both intra-cluster and inter-cluster interference, resulting in widespread interference links, significantly increased power consumption, and thus poor energy efficiency.

[0199] Looking at the enabling conditions of each scheme, in the scheme without SWIPT, the energy efficiency drops rapidly when the number of M2M devices exceeds 6. The other SWIPT schemes, however, can maintain relatively high energy efficiency until the number of M2M devices exceeds 8. This is because, compared to the scheme without SWIPT, the SWIPT scheme can convert a large amount of interference power into energy harvesting while slightly reducing spectral efficiency. Therefore, based on the trade-off between spectral efficiency and power consumption, the SWIPT scheme can exhibit better performance under certain interference environments. This performance degradation becomes less pronounced as the number of M2M devices increases, because further increases in the number of M2M devices severely impair spectral efficiency, thus continuously reducing the potential for further performance degradation.

[0200] From the performance of each scheme, the resource allocation scheme proposed in this application achieves the highest energy efficiency. Furthermore, when the number of M2M devices increases from 6 to 8, the energy efficiency achieved by this application shows almost no decrease, indicating that this scheme performs better in balancing spectral efficiency and energy consumption. This is because this application enables distributed execution of the training process through a multi-agent setup, fully considering the conditions of each agent at each time step and making the most suitable resource allocation strategy for each agent. The slightly worse performance of the centralized training single-agent scheme is because its resource allocation scheme is more general and can be applied to each agent, to some extent ignoring the individual situation of each agent at each time step. The poor performance of the Q-learning scheme is due to the complexity of the network environment, the large number of state and action sets, and the low computational efficiency of the traditional reinforcement learning table lookup method, which ignores some high-performing resource allocation schemes during the iteration process.

[0201] Figure 8The figure illustrates the change in H2H user QoS requirement satisfaction rate as the number of tolerant M2M devices in the system increases. As shown in the figure, the QoS requirement satisfaction rate of H2H users decreases with the increase in the number of tolerant M2M devices. This is because increasing the number of tolerant M2M devices increases the number of links reusing the spectrum of the H2H link, thus increasing the interference power at the H2H link receiver and reducing the SINR value at the receiver. Furthermore, the QoS requirement satisfaction rate of H2H users in the non-SWIPT scheme responds more drastically to the increase in the number of tolerant M2M devices, with performance rapidly declining even when only a few M2M devices are present. There are two reasons for this phenomenon. First, in the non-SWIPT scheme, tolerant M2M links prefer to share the same spectrum resources as H2H links because the M2M receivers of tolerant links are mostly far from the base station, thus experiencing less interference and achieving higher energy efficiency. Similarly, critical M2M links within the management range of the machine equipment gateway can also achieve high energy efficiency when no other users reuse their occupied spectrum subbands. Second, in the SWIPT scheme, tolerant M2M links prefer to share the same spectrum resources as critical M2M links because SWIPT can convert received interference into energy, thereby improving energy efficiency. Among all schemes, the resource allocation method proposed in this application achieves the best H2H user QoS requirement satisfaction rate, demonstrating that this method performs better in guaranteeing user QoS requirements.

[0202] Figure 9 This illustrates the change in the probability of critical M2M link outages as the number of tolerant M2M devices in the system increases. From... Figure 9 It is evident that as the number of tolerant M2M devices increases, the probability of critical M2M link outages also rises. This is based on the fact that more spectrum access leads to a higher probability of outages. Furthermore, the no-swipt solution performs best, for reasons including... Figure 2 In addition to the aforementioned spectrum reuse priority, the higher SINR value is achieved because critical M2M devices in the SWIPT scheme dedicate all received energy to information decoding, thus reducing the probability of outages. However, the SWIPT scheme sacrifices some link transmission reliability to achieve higher energy efficiency. The resource allocation method proposed in this application performs best in the SWIPT scheme, demonstrating the high reliability of the proposed scheme.

[0203] Reference Figure 4This paper describes the change in the success rate of tolerant M2M user payload transmission as the number of tolerant M2M devices in the system increases. As shown in the figure, the payload transmission success rate decreases with the increase in the number of tolerant M2M devices. This is because excessive spectrum reuse leads to more interference and power consumption, thus reducing the capacity of each link for payload transmission. Furthermore, the worst performance of the non-SWIPT scheme is due to the fact that, without energy harvesting capabilities, each tolerant M2M link can only maintain a high level of energy efficiency by selecting a lower transmit power, but this reduces capacity. The SWIPT scheme, with energy harvesting capabilities, allows each tolerant M2M link to accept a certain amount of interference to achieve higher energy efficiency, thus choosing a higher transmit power, and the link capacity increases accordingly. The resource allocation method proposed in this application exhibits the best performance among all schemes, further confirming the effectiveness of the proposed method.

[0204] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for downlink resource allocation with SWIPT assistance, characterized in that, include: Obtain status observations of the current environmental state of the communication link between the tolerant machine user equipment and the machine equipment gateway; Based on the state observations of the current environmental state, a resource allocation strategy is selected for the tolerant machine user equipment using a resource allocation model built on a neural network. Based on the selected resource allocation strategy, downlink resources are allocated to the tolerant machine user equipment; The tolerant machine user equipment is a machine user equipment that needs to complete the transmission task of periodically generated payloads. The resource allocation model is constructed using the following methods: a state function is constructed, which is a set of state observations; an action function is constructed, which is a set of downlink spectrum resources, transmit power levels, and power shunting ratios; a reward function is constructed based on the balance between resource allocation optimization objectives and QoS constraints; and the resource allocation model is constructed based on the state function, the action function, and the reward function. The construction of the reward function includes: Based on the remaining load and transmission time of the tolerant machine user equipment, the penalty term of the reward function provided by the tolerant machine user equipment is determined; based on the outage probability of the critical machine user equipment, the penalty term of the reward function provided by the critical machine user equipment is determined; based on the signal-to-interference-plus-noise ratio of the human user equipment, the penalty term of the reward function is determined; the reward function is constructed based on the total energy efficiency value of the communication links of all machine user equipment, the penalty term of the reward function provided by the tolerant machine user equipment, the penalty term of the reward function provided by the critical machine user equipment, and the penalty term of the reward function of the human user equipment; or Set QoS constraints; use the total energy efficiency achieved by the tolerant machine user equipment and the critical machine user equipment as the resource allocation optimization target; construct the reward function based on the resource allocation optimization target and the QoS constraints.

2. The method of claim 1, wherein, Before acquiring state observations of the current environmental state of the communication link between the tolerant machine user equipment and the machine equipment gateway, the method further includes: Based on the service type of the machine user equipment, the machine user equipment is divided into tolerable machine user equipment and critical machine user equipment; Among them, the critical machine user equipment is machine user equipment whose transmission reliability requirements are higher than a preset threshold.

3. The method of claim 1, wherein, The state function is constructed based on the channel gain information of the communication link in each spectrum sub-band, the interference power received by the communication link in each spectrum sub-band, the remaining payload and transmission time of the communication link, the current iteration number, and a greedy factor representing the current environment exploration rate of the communication link.

4. The method of claim 1, wherein, Setting QoS constraints includes: The QoS constraint for human user equipment is set to a signal-to-interference-plus-noise ratio (SINR) greater than a set minimum threshold. Setting the QoS constraint condition of the tolerant machine user equipment as a time constraint The inner size is a preset size The transmission success rate of the load is higher than a set success rate threshold The QoS constraint for critical machine user equipment is set to ensure that the interruption probability does not exceed a set interruption threshold.

5. The method of claim 1, wherein, After acquiring the state observations of the current environmental state of the communication link between the tolerant machine user equipment and the machine equipment gateway, the method further includes: An experience replay pool is constructed to store training data for training the resource allocation model, wherein the training data includes state observations, reward values, and selected resource allocation strategies at the current and next time steps. The neural network includes a training network and a target network. The training network is trained in each iteration using data randomly drawn from the experience replay pool and employing stochastic gradient descent. The target network is a fixed neural network whose network parameters are updated periodically to match the parameters of the training network at the current time.

6. The method according to any one of claims 2 to 5, characterized in that, The machine user equipment is equipped with a wireless communication power-carrying device (SWIPT) to enable the machine user equipment to obtain power from the radio frequency environment and simultaneously perform information decoding.

7. A device for downlink resource allocation with SWIPT assistance, characterized in that, include: The acquisition module is configured to acquire status observations of the current environmental state of the communication link between the tolerant machine user equipment and the machine equipment gateway; The selection module is configured to select a resource allocation strategy for the tolerant machine user equipment based on the state observations of the current environmental state and using a resource allocation model built on a neural network. The allocation module allocates downlink resources to the tolerant machine user equipment based on the selected resource allocation strategy. The tolerant machine user equipment is a machine user equipment that needs to complete the transmission task of periodically generated payloads. The resource allocation model is constructed using the following methods: a state function is constructed, which is a set of state observations; an action function is constructed, which is a set of downlink spectrum resources, transmit power levels, and power shunting ratios; a reward function is constructed based on the balance between resource allocation optimization objectives and QoS constraints; and the resource allocation model is constructed based on the state function, the action function, and the reward function. The construction of the reward function includes: Based on the remaining load and transmission time of the tolerant machine user equipment, the penalty term of the reward function provided by the tolerant machine user equipment is determined; based on the outage probability of the critical machine user equipment, the penalty term of the reward function provided by the critical machine user equipment is determined; based on the signal-to-interference-plus-noise ratio of the human user equipment, the penalty term of the reward function is determined; the reward function is constructed based on the total energy efficiency value of the communication links of all machine user equipment, the penalty term of the reward function provided by the tolerant machine user equipment, the penalty term of the reward function provided by the critical machine user equipment, and the penalty term of the reward function of the human user equipment; or Set QoS constraints; use the total energy efficiency achieved by the tolerant machine user equipment and the critical machine user equipment as the resource allocation optimization target; construct the reward function based on the resource allocation optimization target and the QoS constraints.

Citation Information

Patent Citations

  • Terminal direct communication resource management scheme based on energy collection

    CN107087305A

  • Uplink resource allocation system and method based on QoS (Quality of Service)

    CN114422986A