D2D communication resource allocation method and system in different scenes
By optimizing D2D communication resource allocation through the Q-learning algorithm, the problem of insufficient spectrum resource utilization in cellular networks is solved, and efficient resource allocation and throughput improvement are achieved in multiple scenarios.
Patent Information
- Application Number
- CN202511481094.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2025-12-12
AI Technical Summary
Existing technologies have failed to effectively optimize the utilization of limited spectrum resources in cellular networks across various scenarios, resulting in insufficient allocation of communication resources, especially inefficient in complex multi-scenario environments.
The Q-learning algorithm is adopted to calculate the signal-to-noise ratio and channel quality indicators based on the transmit power, interference noise power and device location, construct an adaptive reward function, and optimize the allocation of D2D communication resources through Q-tables to adapt to network conditions in different scenarios.
It improves the efficiency of D2D communication resource allocation, enhances the balance between system throughput and energy consumption, and adapts to communication needs in complex scenarios.
Smart Images

Figure CN121126434A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of communication, and particularly relates to a D2D communication resource allocation method and system in different scenarios. BACKGROUND
[0002] The increasing number of wireless users and emerging wireless services have led to a surge in mobile traffic in wireless cellular networks. This increase in traffic load leads to a large number of devices connecting, resulting in an overload of the cellular network. The demand for data traffic is expected to grow exponentially. Device-to-device (D2D) communication refers to the use of physical proximity of communication devices to extend cellular coverage, mainly in sparse deployment scenarios. D2D communication allows devices close to each other to communicate directly, exchange information, or relay data. This technology is applied in different scenarios, such as public safety, commercial and social networks, location-based services, and network offloading. At the same time, D2D devices can also act as gateways, aggregating and forwarding data from sensors to the cellular base station (BS). Allowing direct communication between two devices with minimal involvement of the base station (BS) improves the communication quality of existing networks. And it can be easily combined with other advanced technologies, thereby further improving the performance of the wireless communication system.
[0003] With the increasing number of users and mobile devices, the demand for higher data rates in wireless communication systems continues to grow, putting tremendous pressure on traditional cellular networks. The existing technology focuses on improving throughput and interference management, without addressing the problem of millimeter waves or complex multi-scenario. Therefore, there are limitations in the optimal use of limited spectrum resources. SUMMARY
[0004] In order to solve the problem of insufficient communication resource allocation in multi-scenario state in the prior art, the present application provides a D2D communication resource allocation method in different scenarios.
[0005] In order to achieve the above-mentioned purpose, the present application provides the following technical scheme: A D2D communication resource allocation method in different scenarios, comprising the following steps: Obtaining the transmit power, interference noise power, interference level and real-time position coordinates of the device in the communication scenario, calculating the signal-to-noise ratio (SINR) value based on the transmit power and interference noise power, obtaining the channel quality indicator based on the signal-to-noise ratio (SINR) value; defining the action in Q-learning according to the SINR value; constructing an adaptive reward function based on the throughput and energy consumption with the goal of maximizing the throughput of the system; defining the state in Q-learning according to the real-time position coordinates of the device, the channel quality indicator and the interference level; Based on the action and the three-dimensional state space, an initialized communication Q-table is obtained. The selected action is executed to perform a state transition. A new state observation value is obtained based on the transitioned state. An immediate reward is calculated based on the transitioned state using the adaptive reward function. The Q-value corresponding to the current state is updated based on the new state observation value and the immediate reward. A convergence condition is set, and the Q-value is repeatedly calculated iteratively until convergence. The converged Q-table is then output. The corresponding action is queried from the converged Q-table, and communication resources for the current scenario are allocated based on the queried action.
[0006] Preferably, when acquiring the transmit power and interference noise power in a communication scenario, and calculating the signal-to-noise ratio (SINR) based on the transmit power and interference noise power, the communication scenario includes a stable communication scenario and a changing communication scenario; the stable communication scenario Specifically, it is calculated using the following formula: ; The changing communication scenario Specifically, it is calculated using the following formula: ; in, It is the first The transmit power of a pair of D2D communication devices. It is the first Channel gain of a D2D link, It is the first The transmit antenna gain of a D2D communication device pair It is the first Path loss of a D2D link, This is the interference power in the first scenario. This represents the interference power in the second scenario. It is noise power.
[0007] Preferably, the adaptive reward function ; in, , and For adaptive weights, For the throughput generated by the action, Interference caused by the action The transmission power consumed for the action.
[0008] Preferably, the dynamic adjustment logic of the adaptive weights specifically involves: setting a threshold for adaptive adjustment of the SINR value. ; Set interference threshold for adaptive adjustment ; Set an adaptive adjustment threshold for the difference between the total D2D power and the maximum allowable power. .
[0009] Preferably, the execution of the selected action is specifically achieved through random selection or - A greedy strategy selects transmit power and spectrum resources.
[0010] This invention also provides a D2D communication resource allocation system for different scenarios, specifically including: The data acquisition module is used to acquire the transmit power, interference noise power, interference level, and real-time location coordinates of the device in the communication scenario; calculate the signal-to-noise ratio (SINR) based on the transmit power and interference noise power; obtain the channel quality index based on the SINR; define the actions in Q-learning based on the SINR value; construct an adaptive reward function based on throughput and energy consumption with the goal of maximizing system throughput; and define the state in Q-learning based on the real-time location coordinates of the device, the channel quality index, and the interference level.
[0011] The model processing module obtains an initialized communication Q-table based on the action and the three-dimensional state space, executes the selected action to perform a state transition, and obtains a new state observation value based on the transitioned state; calculates an immediate reward based on the transitioned state using the adaptive reward function; updates the Q-value corresponding to the current state based on the new state observation value and the immediate reward; sets a convergence condition, repeatedly iterates to calculate the Q-value until convergence, and outputs the converged Q-table.
[0012] The resource allocation module is used to query the corresponding action according to the converged Q table, and allocate communication resources for the current scenario based on the queried action.
[0013] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps described in the method for allocating D2D communication resources in different scenarios.
[0014] The present invention also provides a computer-readable storage medium storing a computer program, which, when loaded by a processor, is capable of executing the steps described in the method for allocating D2D communication resources in different scenarios.
[0015] The D2D communication resource allocation method provided by this invention in different scenarios has the following beneficial effects: This invention defines actions and states in Q-learning based on relevant parameters of the communication scenario. An initialized communication Q-table is obtained based on the actions and a 3D state space. A selected action is executed to update the state. An immediate reward is calculated based on the updated state using an adaptive reward function, which dynamically enhances throughput and expands bandwidth to adapt to the communication needs of different scenarios. The Q-value corresponding to the current state is updated according to the new state observation and the immediate reward. A convergence condition is set, and the Q-value is iteratively calculated until convergence, outputting the converged Q-table. Communication resources for the current scenario are allocated based on the retrieved actions. Compared with other traditional methods, this approach has advantages in resource allocation across multiple scenarios, improving the efficiency of resource allocation in D2D communication in complex scenarios. Attached Figure Description
[0016] To more clearly illustrate the embodiments and design schemes of the present invention, the accompanying drawings required for this embodiment will be briefly described below. The drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart of a D2D communication resource allocation method under different scenarios according to the present invention.
[0018] Figure 2 This is a schematic diagram of the first scene layout according to an embodiment of the present invention.
[0019] Figure 3 This is a schematic diagram of the second scene layout according to an embodiment of the present invention.
[0020] Figure 4 This is a flowchart of D2D communication resource allocation based on Q-learning in an embodiment of the present invention.
[0021] Figure 5 This is a system model for D2D wireless cellular communication in an embodiment of the present invention.
[0022] Figure 6 This is a graph showing the relationship between the number of iterations in the first scenario and the system throughput in an embodiment of the present invention.
[0023] Figure 7 This is a graph showing the relationship between the logarithm and power amplitude in the first scenario of this invention.
[0024] Figure 8 This is a graph showing the relationship between the number of iterations in the second scenario and the system throughput in an embodiment of the present invention.
[0025] Figure 9 This is a graph showing the relationship between the number of iterations and the logarithm of successful matches in the second scenario according to an embodiment of the present invention.
[0026] Figure 10 This is a graph showing the relationship between the logarithm and power amplitude during the first simulation in the second scenario of this embodiment of the invention.
[0027] Figure 11 This is a graph showing the relationship between the logarithm and power amplitude during the last simulation in the second scenario of this embodiment of the invention.
[0028] Figure 12 This is a graph showing the relationship between the number of iterations of Q-learning PA and the average reward in an embodiment of the present invention. Detailed Implementation
[0029] To enable those skilled in the art to better understand and implement the technical solutions of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be construed as limiting the scope of protection of the present invention.
[0030] Example Consider two single-cell cellular networks. Taking a stable communication scenario as the first scenario (a common underlying cellular network using D2D pairs in 5G protocols), consider a single-cell 5G cellular network environment with a radius of approximately 250 meters. The BS (Base Station) is located in the cell center, and devices are randomly distributed throughout the cell. Assume that mobile users (CUEs), D2D transmitters (D2D Tx), and receivers (D2D Rx) are all within the coverage area of the central BS. Consider allocating 30 D2D pairs and 20 CUEs. For example... Figure 2 As shown, this simulation utilizes a common underlying cellular network with D2D pairs in a 5G protocol. The system aims to reduce device energy consumption while extending terminal communication reliability. In this scenario, the device's real-time location coordinates, channel quality indicators, and interference levels reflect relatively stable network conditions with minimal parameter variations, making the system's trade-off between throughput and energy efficiency quite straightforward.
[0031] Using a variable communication scenario as the second scenario (a cellular network using the 5G mmW band for D2D communication), we simulate a communication-constrained environment. Consider a single-cell autonomous network environment with a radius of approximately 200 meters, where the BS is excluded from the cell due to impairment, and devices are randomly distributed within the cell. Assume that mobile users (CUEs), D2D transmitters (D2DTx), and receivers (D2D Rx) are all within the coverage area of the central BS. Consider 30 pairs of D2D devices and 20 CUEs allocated, and approximately 35 obstacles simulated in the cell based on a Poisson point process (PPP) random distribution. In this scenario, the system relies entirely on D2D communication, and the base station remains relative, assisting in establishing these connections. In this case, the minimum function of the BS will be device discovery and allocating available energy to appropriate devices; the remaining operations are performed by the devices themselves.Figure 3 As shown in the diagram, in this scenario, the device's location coordinates, channel quality, and interference levels change drastically, requiring the Q-learning algorithm to adjust its state and actions more frequently to adapt to the constantly changing environment.
[0032] This invention provides a method for allocating D2D communication resources in different scenarios, such as... Figure 4 As shown, the specific steps include: The communication system model is based on a single-cell architecture, comprising a central base station (BS) and two types of user equipment: M cellular user equipment (CUE) and N pairs of D2D communication devices. The two user sets are defined as CUE = {C1, C2, ..., CM} and D = {D1, D2, ..., DN}, where CM and DN represent the maximum number of cellular user-D2D communication pairs supported by the system. Cellular users transmit data through the base station and may also participate in the cooperative transmission process of D2D communication as potential relay nodes. All user equipment (including cellular and D2D) is randomly distributed within the cell. A D2D transmitter (Tx) and a D2D receiver (Rx) constitute a D2D communication pair. To achieve efficient spectrum resource utilization, the system adopts a mechanism of spectrum sharing between cellular and D2D users, assuming that the number of available resource blocks (RBs) is equal to the number of cellular users, i.e., the resource block set is B = {B1, B2, ..., BM}. In this system, a single resource block serves as the smallest unit of spectrum resource allocation. Each D2D pair can occupy only one resource block at a time, but the same resource block can be dynamically multiplexed by multiple D2D pairs. Meanwhile, maintaining system QoS interference can be cross-layer or co-layer. If there are both D2D connections and traditional communication connections between the CUE and the base station (BS), it can lead to inter-device interference. Specifically, the D2D receiver will be affected by all operating CUE-BS channels. Conversely, the base station will be affected by transmitted signals from the D2D transmitter.
[0033] Step 1: Communication scenario initialization.
[0034] Step 2: Determine the first step using the Euclidean formula. The distance between D2D pairs. The Rayleigh fading path loss model, expressed in dB, is given by:
[0035] ; in, Indicates frequency; Indicates distance.
[0036] Path loss (PL) is typically expressed in decibels (dB), while channel gain is a dimensionless ratio. To convert path loss to channel gain, use the following formula:
[0037] ; Calculate the signal-to-noise ratio of the first scene and the second scene respectively. The first scene The formula is: ; The SINR formula for the second scenario is: ; in, , , It is the first The transmit power of each D2D pair It is the first Channel gain of a D2D link, It is the first Transmit antenna gain of a D2D pair It is the first Path loss of a D2D link, This is the interference power in the first scenario. Interference power in the second scenario. This refers to noise power, specifically: ;
[0038] in, It is the bandwidth used for each transmission frequency. It is a variable called the noise figure, which represents the degradation of the signal-to-noise ratio and is used in conjunction with the sensitivity of the radio receiver.
[0039] Based on Shannon's formula, and combined with The throughput is calculated for each link as follows: ; Total throughput It is the sum of the throughput of all links, where the throughput of each link is determined by its bandwidth. and SINR Decision. Total throughput for:
[0040] ; in, The parameters, It depends on the receiver in different scenarios. , For the receiver in the second scenario , This indicates the total number of D2D devices.
[0041] The interference value should be kept below the threshold, for the first The first D2D link and the first There are one interference source, and its interference value is: ; ; in, Alternatively, n can be used to represent the total interference calculation for 1 to n interference sources.
[0042] To improve system performance, interference constraints were considered. Interference induced by the D2D transmitter was kept less than or equal to a certain threshold, while maintaining... Value higher than minimum Threshold.
[0043] Antenna gain is calculated based on the transmission angle. The total antenna gain is the product of the transmit gain and the receive gain (the linear values are multiplied, and the corresponding decibel values are added together).
[0044] .
[0045] Step 3: Based on different scenarios The values define the parameters for actions, states, and rewards in Q-learning. By dynamically incorporating SINR into the Q-learning framework, the system can adapt to channel variations and achieve synergistic optimization of interference suppression and throughput improvement.
[0046] The SINR value obtained in step 2 defines the three-dimensional state space of Q-learning. This three-dimensional state space includes the device's real-time location coordinates, the Channel Quality Index (CQI), and the interference level. The SINR value is the core indicator for measuring channel quality. It reflects the current channel attenuation and interference level, directly affecting the representation of the state.
[0047] The three-dimensional state space (device location, channel quality, interference level) and action space (power allocation strategy) defined in step 2 are mapped together using Q-learning. In step 3, the real-time network state is mapped to the state in the Q-table, triggering the execution of corresponding actions to achieve adaptive resource allocation.
[0048] An adaptive reward function is constructed based on a weighted fusion of two key performance indicators: throughput and energy consumption. By dynamically adjusting the weights, it adapts to network conditions and optimization objectives under different scenarios, thereby enabling the Q-learning algorithm to adaptively optimize system performance. The adaptive reward function differs significantly from the reward function in the conventional Q-learning algorithm, primarily in the following ways: it dynamically adjusts the weights according to network conditions in different scenarios (e.g., throughput, energy consumption, interference levels); it achieves dynamic trade-offs among multiple objectives through weighted fusion of throughput, energy consumption, and other indicators; and it can self-adjust according to network conditions under different scenarios, while the reward function of conventional Q-learning typically remains unchanged throughout the entire training process.
[0049] The reward function combines throughput and energy consumption parameters, weighted by specific parameters, and adaptively adjusts for different scenarios and environments. This balances optimization objectives, such as prioritizing throughput while also considering energy consumption limitations.
[0050] Through this adaptive reward mechanism, the system can dynamically adjust its resource allocation strategy based on real-time conditions (such as channel quality and interference), thereby improving overall network efficiency and resource utilization.
[0051] Step 4: Allocate D2D communication resources based on Q-learning.
[0052] S41: Initialize a Q-table containing multi-dimensional state-action pairs, where the state space consists of dynamic parameters such as the device's real-time location coordinates, channel quality index (CQI), and interference level. The iteration process first observes the current state and then selects an operation based on the observed current state result.
[0053] For the initialization of Q-table: denoted as , indicating the state Select action The expected cumulative reward.
[0054] ; Actions can be performed in different ways, including random or strategic. - The expression for the greedy strategy is as follows:
[0055] ; in, The action that indicates the next state. Indicates the probability of exploration. This indicates the number of actions with the highest Q value (there may be more than one).
[0056] S42: After the selected action is performed, the environmental feedback includes an immediate reward signal and a new state observation. Action Cause state transition → and return instant rewards .
[0057] ; ; ; in, Indicate the next state From probability distribution The distribution, extracted from the data, describes the state from which action a is taken. Transition to state The transition probability, Current state Next action of value, The learning rate controls the extent to which new information updates old values. To perform the action The immediate reward obtained afterward, with the timing difference error being a defined error term, This is the discount factor.
[0058] Indicates the power allocation under Q-learning. The transmit power of each D2D transmitter is as follows: ; in, .
[0059] S43: Update the Q-value of the current state-action pair, which includes the reward obtained and the expected future reward. The Q-value update mechanism integrates the current reward and potential future gains, controls the update speed using the learning rate parameter, and balances short-term benefits with long-term network stability through a discount factor. Repeat this loop until the algorithm converges, and output the converged Q-table.
[0060] Step 5: Query the corresponding action based on the converged Q-table to obtain the optimal resource mapping strategy for the current D2D communication, thereby realizing adaptive spectrum efficiency optimization and cross-layer interference coordination in complex network environments.
[0061] To address the power allocation problem in D2D communication, a dynamic optimization framework based on Q-learning (Q-learning PA) is proposed and systematically compared with four power allocation schemes (PL PA, Equ PA, RanPA, and Thr PA). The Path Loss Power Allocation (PL PA) scheme dynamically adjusts the transmit power based on the signal transmission path loss, employing an adaptive compensation mechanism positively correlated with transmission distance to maintain basic communication quality by compensating for signal attenuation. Equal Power Allocation (EquPA) distributes total power equally among all active devices, characterized by its simplicity and stable allocation. Random Power Allocation (RanPA) randomly allocates power to devices without considering network conditions or specific device requirements, resulting in low implementation cost but potentially degrading overall network performance. Throughput Power Allocation (Thr PA) aims to maximize instantaneous throughput. In contrast, the Q-Learning algorithm achieves precise perception of the dynamic characteristics of millimeter-wave channels by defining a three-dimensional state space (including path loss value, interference intensity index, and channel quality level), constructing an adaptive reward function (integrating spectral efficiency and energy consumption weighted indices), and designing an ε-attenuation exploration strategy. The proposed Q-learning algorithm is compared and analyzed with the four methods mentioned above. The practicality and efficiency of Q-Learning-based D2D communication power allocation in a given system model are demonstrated in terms of performance indicators and system parameters.
[0062] The effectiveness of the algorithm was verified through two typical test scenarios: A quantitative analysis was conducted on the simulation results for Scenario 1 (standard 5G cellular network). In the first scenario, for each power allocation model, the number of iterations performed during the simulation process refers to a total of 50 training cycles, with each cycle containing 450 iterations, for a total of 22,500 iteration calculations. Figure 6 This paper compares the system throughput of the Q-Learning power allocation strategy with four other power modulation methods. The output is aggregated or averaged every five iterations to smooth the results. Simulation results show that the capacity of all five power allocation strategies exhibits a fluctuating distribution as the number of iterations increases. Compared with existing technologies, the Q-Learning method significantly increases system throughput, forming a quasi-periodic oscillation after 20 iterations with an average of approximately 157.5 Mbps. The stable average values of Equ PA, Ran PA, PL PA, and Thr PA are approximately 134.3 Mbps, 133.6 Mbps, 137.2 Mbps, and 153.3 Mbps, respectively.
[0063] Figure 7The diagram illustrates the relationship between the number of pairs and the power amplitude in the first scenario, representing the average energy provided by the system to each of the 25 pairs during the simulation. Simulation results show that the proposed Q-learning PA closely follows Thr PA in the 7-17 pair range, and both strategies maintain the power level within a relatively stable range of fluctuations in this range. With the total system energy remaining constant, the energy delivered per pair by Equ PA is also constant. The overall power amplitude of Q-learning PA is higher than that of Equ PA, and its overall fluctuation is second only to Equ PA. It maintains a stable power level in the later stages, keeping the power amplitude between 0.128W and 0.132W, thus ensuring signal reliability while saving energy. The power amplitude under PL PA gradually increases with the number of pairs, with a significant increase after 15 pairs.
[0064] Table 4 summarizes the performance of various power allocation techniques in the first scenario. The maximum data rate of Q-learning PA in the later stages of iteration was 161.97 Mbps, while Throughput PA achieved 157.26 Mbps and Random PA 130.12 Mbps. Therefore, for the considered D2D system model, the Q-learning PA algorithm provides approximately 3% more output than Throughput PA and approximately 24% more output than Random PA. Regarding power amplitude, Q-learning shows an output of 0.126–0.131 W in the 5–25 pair interval, which is smaller than the amplitude of Throughput PA and Path Loss PA outputs. The analysis shows that the Q-learning strategy provides more and more stable network capacity than Throughput PA and Random PA. Although Q-learning does not consistently achieve high capacity output, Q-learning PA gradually converges during iteration, indicating its reliability under standard 5G cellular network conditions.
[0065] Table 4 Simulation Results Data for the First Scenario During network outages caused by BS functionality impairment, the second scenario execution does not consider the Thr PA scenario. Figure 8The relationship between the number of iterations and system throughput in the second scenario is shown. Simulation results indicate that the overall system throughput is reduced by approximately 20.6-56 Mbps compared to the first scenario. The stable mean of the Q-Learning PA is approximately 122.2 Mbps, while the stable mean values of the Equ PA, Ran PA, and PL PA are approximately 79.8 Mbps, 77.6 Mbps, and 116.6 Mbps, respectively. In the simulated scenario, when a point is generated inside an obstacle, that point is removed from the available devices, thus reducing the number of successfully linked pairs. Therefore, after creating the shortest distance pairs, N is 25, and thereafter, if some points are located inside obstacles, the program excludes them from the network along with their partners. Figure 9 This shows the instabilities that occur during execution, with an average number of 18-21 pairs.
[0066] The amplitude is relatively stable and outputs 0.187W at 16 pairs, while Equ PA, PL PA, and Ran PA exhibit data power of 0.168W, 0.167W, and 0.157W, respectively. The four strategies show significant divergence after a sharp drop. Q-Learning PA and PL PA show significant improvements before the sharp drop at 15 and 16 pairs, indicating that they respond to adaptive power allocation under complex network conditions.
[0067] Figure 10 The relationship between logarithm and power amplitude in the first simulation of the second scenario is shown. Simulation results indicate that the Q-Learning PA has a relatively stable amplitude in the range of 1-16 pairs, with an output of 0.187W at 16 pairs, while the Equ PA, PL PA, and Ran PA exhibit data power of 0.168W, 0.167W, and 0.157W, respectively. The four strategies show significant divergence after a sharp drop. The Q-Learning PA and PL PA show significant improvements before the sharp drop at 15 and 16 pairs, indicating that they respond to adaptive power allocation under complex network conditions.
[0068] Figure 11 The relationship between the logarithm and power amplitude in the last simulation run under the second scenario is presented. Simulation results show that, unlike the first simulation run, the Q-Learning PA exhibits more stable power amplitude in the 1-18 pair range compared to the first simulation run. At 18 pairs, the outputs of the Q-Learning PA and the other three PAs are 0.182W, 0.146W, 0.174W, and 0.143W, respectively. The overall output drops sharply after 18 pairs. Compared to other methods considered, the proposed method demonstrates its ability to provide optimal power allocation due to its higher energy efficiency.
[0069] Table 5 summarizes the performance of various power allocation techniques in the second scenario. Although the overall system throughput decreased, Q-learning PA still achieved a maximum data rate of 117.18 Mbps in the later stages of iteration, while Path Loss PA achieved 105.65 Mbps and Random PA achieved 70.78 Mbps. Therefore, for the considered D2D system model, the Q-learning PA algorithm provides approximately 11% more output than the Path Loss PA algorithm, and approximately 66% more output than the Random PA algorithm. Regarding power amplitude, in the first execution simulation scenario, Q-learning PA showed a significant power boost in pairs 15-16 before the sharp drop. Similarly, in the second execution simulation scenario, a significant power boost was observed in pairs 17-18, indicating that Q-learning PA has adaptive power management capabilities in complex network conditions. Comparison with the results of the first scenario shows that Q-learning is more suitable for autonomous D2D communication in environments with impaired communication.
[0070] Table 5. Simulation Results Data for the Second Scenario The typical convergence characteristic of Q-learning algorithms is that the iterative process to achieve a stable match means that the convergence time may be longer, such as... Figure 12 As shown, the average reward value increases sharply from -33 to -5 in 0-5 iterations. Q-learning achieves an 80% performance improvement in only about 5 iterations, indicating that it can quickly capture the time-varying characteristics of D2D channels, corresponding to the agent's transition from random exploration to... The algorithm initially establishes an effective power mapping strategy. During iterations 6-30, the reward growth rate slows but remains positive, indicating that the algorithm has entered a local optimization phase, gradually adjusting the power allocation scheme to balance throughput improvement and power loss. During iterations 31-50, the reward value stabilizes in the range [-2, 1.1], with fluctuations less than ±0.5, demonstrating that the algorithm can maintain a stable strategy and reach convergence even during sudden channel changes.
[0071] This invention also provides a D2D communication resource allocation system for different scenarios, specifically including: The data acquisition module is used to acquire transmit power, interference noise power, interference level, and real-time location coordinates of the device in the communication scenario. It calculates the signal-to-noise ratio (SINR) based on the transmit power and interference noise power, and obtains the channel quality index based on the SINR. It defines the actions in Q-learning based on the SINR value. With the goal of maximizing the system throughput, it constructs an adaptive reward function based on throughput and energy consumption. It defines the state in Q-learning based on the real-time location coordinates of the device, the channel quality index, and the interference level.
[0072] The model processing module obtains an initialized communication Q-table based on the action and the 3D state space, executes the selected action to perform a state transition, and obtains a new state observation value based on the transitioned state; calculates the immediate reward based on the transitioned state using an adaptive reward function; updates the Q value corresponding to the current state based on the new state observation value and the immediate reward; sets the convergence condition, iteratively calculates the Q value until convergence, and outputs the converged Q-table.
[0073] The resource allocation module is used to query the corresponding action based on the converged Q table, and allocate communication resources for the current scenario based on the queried action.
[0074] The modules in the D2D communication resource allocation system described above for different scenarios can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0075] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps in an embodiment of a D2D communication resource allocation method under different scenarios. Specific implementation methods can be found in the method embodiments, and will not be repeated here.
[0076] Furthermore, the present invention also provides a non-transitory computer-readable storage medium containing instructions on which a computer program is stored. For example, a memory containing instructions that can be executed by a processor of a computer device to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc. When the computer program is executed by the processor, it can implement the steps in an embodiment of the D2D communication resource allocation method in a different scenario. Specific implementation methods can be found in the method embodiments, which will not be repeated here.
[0077] Those skilled in the art will understand that embodiments of the present invention can provide methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0078] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0079] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0080] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0081] It should be noted that the specific embodiments described above enable those skilled in the art to more fully understand the present invention, but do not limit the present invention in any way. Therefore, although the present invention has been described in detail in this specification and embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the present invention; and all technical solutions and improvements that do not depart from the spirit and scope of the present invention are covered within the protection scope of the present invention patent. No reference numerals in the claims should be construed as limiting the scope of the claims. Any simple variations or equivalent substitutions of technical solutions that can be readily obtained by those skilled in the art within the scope of the technology disclosed in the present invention are within the protection scope of the present invention.
Claims
1. A method for allocating D2D communication resources in different scenarios, characterized in that, Includes the following steps: The system acquires the transmit power, interference noise power, interference level, and real-time device location coordinates in the communication scenario. Based on the transmit power and interference noise power, it calculates the signal-to-noise ratio (SINR) and obtains the channel quality index (CQI). Actions in Q-learning are defined based on the SINR value. An adaptive reward function is constructed with a weighted average of throughput and energy consumption, aiming to maximize system throughput. States in Q-learning are defined based on the device's real-time location coordinates, CQI, and interference level. Based on the action and the three-dimensional state space, an initialized communication Q-table is obtained. The selected action is executed to perform a state transition. A new state observation value is obtained based on the transitioned state. An immediate reward is calculated based on the transitioned state using the adaptive reward function. The Q-value corresponding to the current state is updated based on the new state observation value and the immediate reward. A convergence condition is set, and the Q-value is repeatedly calculated iteratively until convergence. The converged Q-table is then output. The corresponding action is queried from the converged Q-table, and communication resources for the current scenario are allocated based on the queried action.
2. The method for allocating D2D communication resources in different scenarios according to claim 1, characterized in that, When acquiring the transmit power and interference noise power in a communication scenario, and calculating the signal-to-noise ratio (SINR) based on the transmit power and interference noise power, the communication scenario includes a stable communication scenario and a changing communication scenario; the stable communication scenario... Specifically, it is calculated using the following formula: ; The changing communication scenario Specifically, it is calculated using the following formula: ; in, It is the first The transmit power of a pair of D2D communication devices. It is the first Channel gain of a D2D link, It is the first The transmit antenna gain of a pair of D2D communication devices It is the first Path loss of a D2D link, This is the interference power in the first scenario. This represents the interference power in the second scenario. It is noise power.
3. The method for allocating D2D communication resources in different scenarios according to claim 1, characterized in that, The adaptive reward function ; in, , and For adaptive weights, For the throughput generated by the action, Interference caused by the action The transmission power consumed for the action.
4. The method for allocating D2D communication resources in different scenarios according to claim 1, characterized in that, The specific logic for the dynamic adjustment of the adaptive weights is as follows: A threshold is set for the SINR value for adaptive adjustment. ; Set interference threshold for adaptive adjustment ; Set a threshold for the difference between the total D2D power and the maximum allowable power for adaptive adjustment. .
5. The method for allocating D2D communication resources in different scenarios according to claim 1, characterized in that, The execution of the selected action is specifically achieved through random selection or... - A greedy strategy selects transmit power and spectrum resources.
6. A D2D communication resource allocation system for different scenarios, characterized in that, include: The data acquisition module is used to acquire the transmit power, interference noise power, interference level, and real-time location coordinates of the device in the communication scenario; calculate the signal-to-noise ratio (SINR) based on the transmit power and interference noise power; obtain the channel quality index based on the SINR; define the actions in Q-learning based on the SINR value; construct an adaptive reward function based on throughput and energy consumption with the goal of maximizing system throughput; and define the state in Q-learning based on the real-time location coordinates of the device, the channel quality index, and the interference level. The model processing module obtains an initialized communication Q-table based on the action and the three-dimensional state space, executes the selected action to perform a state transition, and obtains a new state observation value based on the transitioned state; calculates an immediate reward based on the transitioned state using the adaptive reward function; updates the Q-value corresponding to the current state based on the new state observation value and the immediate reward; sets a convergence condition, iteratively calculates the Q-value until convergence, and outputs the converged Q-table. The resource allocation module is used to query the corresponding action according to the converged Q table, and allocate communication resources for the current scenario based on the queried action.
7. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is loaded by the processor, it is able to perform the steps of the method according to any one of claims 1 to 5.