D2D-assisted resource allocation method and system for ultra-dense Internet of Things

By building network models and conflict hypergraph models in a D2D-assisted ultra-dense Internet of Things environment, combined with federated multi-agent deep reinforcement learning, the problem of resource interference between devices is solved, and network throughput and spectrum resource utilization is improved.

CN119071927BActive Publication Date: 2025-05-16CHONGQING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411275221.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2025-05-16
Estimated Expiration
2044-09-12

AI Technical Summary

Technical Problem

In a D2D-assisted ultra-density Internet of Things environment, the resource interference between devices is severe, resulting in limited network throughput and spectrum resource utilization, and the existing technology is difficult to effectively solve this problem.

Method used

By constructing a network model, communication model and conflict model of D2D assisted hyper-secret Internet of Things, the conflict model is transformed into a conflict hyper-graph model using group theory and hypergraph theory. Based on this, the spectrum resource allocation problem is constructed, and the spectrum resource allocation strategy is solved through federal multi-agent deep reinforcement learning.

Benefits of technology

It effectively avoids resource interference between IoED, improves network throughput, optimizes the utilization rate of spectrum resources, and achieves better performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119071927B_ABST
    Figure CN119071927B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of resource management of Internet of Things, and specifically discloses a resource allocation method and system for D2D-assisted ultra-dense Internet of Things. In order to avoid resource conflicts between IoEDs in D2D-assisted UD-IoE, the conflict relationship between IoEDs is first analyzed; then, a conflict hypergraph model that can simultaneously and intuitively reflect these conflict relationships is established by using clique theory and hypergraph theory; based on the conflict hypergraph model, spectrum resource allocation is described as a mixed integer linear programming (MILP) problem; in order to solve the mixed integer linear programming problem, a multi-agent DRL (deep reinforcement learning) framework is constructed for resource allocation in the entire D2D-assisted UD-IoE, and the reward function of each agent is intended to improve the data transmission rate while alleviating resource conflicts. Simulation experiments show that in D2D-assisted UD-IoE, the method and system can achieve higher network throughput and better performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Internet of Things resource management, and in particular to a resource allocation method and system for a D2D-assisted ultra-dense Internet of Things. Background Art

[0002] With the development of wireless communication technology, the Internet of Everything (IoE) is rapidly expanding, covering various industries such as smart cities, smart transportation, and smart manufacturing. At the same time, with the growing demand for seamless connectivity and comprehensive integration in various industries, the number of devices is experiencing rapid growth. According to the latest report from suppliers, there will be 13.2 billion Internet of Everything connections in the wireless infrastructure in 2022, and with an annual growth rate of 18%, it will reach 34.7 billion by 2028. This means that the density of IoED (Internet of Everything Device, IoE Device) will rise significantly, resulting in an ultra-high density IoE (UDIoE) environment, which brings challenges such as limited network throughput and spectrum resource utilization. To address these issues, device-to-device (D2D) communication has attracted great research interest as a key technology to support future ubiquitous communication. This technology supports direct communication between two IoEDs by sharing resources between them, and many studies have used it to effectively improve performance, including resource utilization and throughput. However, D2D communication can cause severe interference between IoEDs when transmitting concurrently through the same spectrum in an ultra-dense environment. Therefore, in D2D-assisted UD-IoE, it is crucial to find an effective resource allocation method to improve network performance and prevent significant interference.

[0003] Many existing works have studied the use of D2D technology to improve network performance. Some literature adopts a resource sharing scalable framework for D2D communication and maximizes system throughput, spectrum utilization, and device access rate by using many-to-many resource sharing in dense scenarios. Some literature proposes a new dual-time scale resource allocation strategy under cooperative D2D communication system, using D2D transmitters as relay stations to assist dense cellular devices to improve the transmission service quality. In addition, in order to maximize the utilization of idle or excess resources and reduce the total network overhead, some literature proposes an adaptive capacity task offloading scheme using multi-hop D2D communication, some literature uses D2D technology to achieve multi-directional task offloading in large-scale edge networks, and some literature designs a three-layer partial offloading framework for mobile edge computing supporting D2D to offload all computationally intensive tasks to resource-rich devices. The above literature improves network performance through various D2D communication methods, including improving spectrum efficiency and increasing network throughput. However, it can be noted that these schemes only perform resource sharing and auxiliary communication, and lack research on multi-user interference between devices.

[0004] In order to circumvent the above limitations, many works consider using various methods and techniques to reduce or mitigate interference between devices, while improving resource utilization and avoiding severe multi-user interference. Some literature analyzes the statistical information of D2D and inter-cell interference and proposes a new statistical rate upgrade technology, which uses allocation information to improve user rates while suppressing interference. Some literature proposes a statistical resource allocation strategy suitable for ultra-dense heterogeneous networks, which mitigates multi-user interference by cognitively characterizing multi-user interference instead of using real-time interference data. Some literature introduces a mobile edge device interference-aware allocation game method, which incorporates communication interference into the system cost model to reduce interference. Some literature proposes a framework based on matching theory, in which the interference between D2D devices is modeled as the receiver coverage probability under same-layer / cross-layer interference, and the Poisson process in random geometry is used as a constraint to achieve interference reduction. In addition, some studies reduce interference between devices based on optimal theory to allocate power, such as a new uplink interference avoidance resource allocation framework designed by some literature, which combines D2D devices and cellular devices to reuse resources and optimize their power to reduce interference. There are literatures that allocate unsent data to more available resource blocks, and then optimize the transmission power of D2D transmitters to reduce interference under normal communication conditions. In fact, in large-scale networks containing a large number of devices, representing communication relationships, device interference, and topology information is crucial for optimizing resource allocation in modern communication systems. However, it is difficult for the above methods to simultaneously consider and model this information.

[0005] As a powerful modeling tool, graph theory can consider multiple types of information at the same time and is also widely adopted by more researchers. At the same time, the resource allocation problem in wireless networks can usually be modeled as a graph coloring problem. They mainly focus on designing algorithms to achieve the coloring of graphs with multiple colors. It is worth noting that the graph coloring problem in dense environments is a problem worth studying because it affects the number of colors and thus the performance. Researchers specializing in graph coloring have also studied similar problems. However, due to the density of D2D-assisted UD-IoE, there is a large amount of overlapping communication range coverage between IoEDs (i.e., there are a large number of IoEDs within the communication range of an IoED). In this case, serious overlapping interference will occur between IoEDs (i.e., large-scale multi-user relationships in the graph model), resulting in a rapid increase in the complexity of coloring and computational load. Therefore, how to effectively avoid resource interference between IoEDs in D2D-assisted UD-IoE remains an unresolved problem. Summary of the invention

[0006] The present invention provides a resource allocation method and system for a D2D-assisted ultra-dense Internet of Things, and solves the technical problem of how to effectively avoid resource interference between IoEDs in a D2D-assisted UD-IoE.

[0007] In order to solve the above technical problems, the present invention provides a D2D-assisted ultra-dense Internet of Things resource allocation method and system, comprising the steps of:

[0008] S1. Construct a network model of D2D-assisted ultra-dense IoT;

[0009] The network model consists of a D2D-assisted UD-IoE layer and a baseband unit pool, namely BBU. D2D refers to device-to-device, and UD-IoE refers to ultra-dense Internet of Things. In the D2D-assisted UD-IoE layer, there are multiple remote control radio heads, namely RRHs, and N mobile Internet of Things devices, namely IoEDs, which communicate through half-duplex channels and have a single antenna. The N IoEDs include N T transmitters and N R receivers; IoED uses D2D communication mode to form a self-organizing network, and IoED senses the surrounding environment information through D2D links; each IoED has a transmission range R n , R n Represents the service area where the nth IoED communicates within a circle with a communication radius R; RRH is connected to the BBU pool through a high-speed front-end link and is responsible for providing basic coverage and auxiliary access; the BBU pool collects all environmental information from the IoED through the RRH, and then allocates resources to the transmitter through the RRH; the total spectrum resources in the BBU pool are divided into K orthogonal resource blocks, i.e., RBs; the total time in the communication network is divided into T time slots;

[0010] S2. Constructing a communication model of D2D-assisted ultra-dense Internet of Things based on the network model;

[0011] S3, constructing a conflict model of D2D-assisted ultra-dense Internet of Things based on the network model and the communication model;

[0012] S4. transforming the conflict model into a conflict hypergraph model based on clique theory and hypergraph theory;

[0013] S5. Based on the conflict hypergraph simplified model and the communication model, a spectrum resource allocation problem is constructed with the goal of maximizing the network throughput of the D2D-assisted ultra-dense Internet of Things;

[0014] S6. Solve the spectrum resource allocation problem based on federated multi-agent deep reinforcement learning to obtain a spectrum resource allocation strategy.

[0015] Further, in step S2, the communication model includes:

[0016] Resource allocation matrix at time slot t n=1,2,…,N,k=1,2,…,K,the element in the nth row and kth column The values ​​are as follows:

[0017]

[0018] The single interference noise rate of the nth IoED allocated with the kth RB, ie, RB k, in time slot t is

[0019]

[0020] in, and The nth and The transmit power of each IoED, and They correspond to the nth and nth The power gain of the IoED channel, σ 2 represents the noise power;

[0021] Data rate of the nth IoED allocated to the kth RB in time slot t

[0022]

[0023] Where B is the bandwidth;

[0024] Data Rate Not less than the corresponding minimum transmission rate

[0025]

[0026] in, Indicates any.

[0027] Furthermore, in step S3, the conflict model is constructed as follows: in is a finite vertex set, namely the IoED set, which is divided into the head vertex subset V H and the tail vertex subset V T , the head vertex represents the transmitter, and the tail vertex represents the receiver; ε={e1,e2,...,e M} is the edge set, i.e., the communication relationship set between IoEDs;

[0028] The relationship between vertices is represented by the adjacency matrix G A ={0,±1} N×N To indicate that the nth row The meaning of the column elements are as follows:

[0029]

[0030] Among them, v n is the head vertex, representing the adjacency matrix G A The nth row in is the tail vertex, indicating the List;

[0031] IoEDs only collide with each other, and for conflicts with secondary neighbors: if the nth and nth emitters are 2nd-order neighbors of each other, and they use the same RB at the same time, then the nth and nth An emitter conflicts with other emitters, and the conflicting emitters form an edge.

[0032] Furthermore, in step S4, the conflict hypergraph model is constructed as in is the vertex set, is a superedge set;

[0033] The conflict hypergraph model at time slot t is composed of the incidence matrix Represents that the elements are represented as:

[0034]

[0035] Among them, h(v,e) = 1 means that vertex v is associated with hyperedge e, that is, hyperedge e contains vertex v. represents the number of vertices, represents the number of hyperedges;

[0036] According to the corresponding features of the receiver, i.e. the tail vertex in the adjacency matrix genetic algorithm, a hyperedge containing the head vertex adjacent to the tail vertex is constructed for each column of the tail vertex, and the conflict hypergraph model is simplified using clique, maximum clique and hypergraph theory;

[0037] The association matrix of the simplified conflict hypergraph model is expressed as Define the color-hyperedge relationship matrix ψ t =(i) M×K It is expressed as:

[0038]

[0039] The overall conflict level of D2D-assisted ultra-dense IoT, i.e., D2D-assisted UD-IoE, is defined as:

[0040] Φ t =log(max(ψ t ,1)),

[0041] Among them, max(ψ t ,1) represents taking the matrix ψ t The maximum value of the elements in and 1, if the matrix Φ t If the element value in is 0, there is no conflict in the conflict hypergraph.

[0042] Furthermore, in step S5, the spectrum resource allocation problem is constructed as follows:

[0043]

[0044] st

[0045] C1: ||Φ t ||1=0,

[0046]

[0047]

[0048]

[0049]

[0050] Among them, max means maximization, st means need to be satisfied, || ||1 means 1 norm, C1 to C5 represent different constraints, P max Indicates the maximum transmit power of the transmitter.

[0051] Furthermore, the step S6 specifically includes:

[0052] A multi-agent deep reinforcement learning framework is constructed in the interactive environment of D2D-assisted UD-IoE, and deep reinforcement learning and federated learning methods are combined to solve the spectrum resource allocation problem;

[0053] The multi-agent deep reinforcement learning framework is modeled as a Markov decision process, where each IoED is regarded as an agent and the goal of each agent is to maximize its reward by accumulating experience in interacting with the environment;

[0054] The state space of a Markov decision process Defined as: the state set s of all agents in the entire network at time slot t t ={s1(t),s2(t),…,s N (t)}, where the state s of the nth agent, i.e. agent n, at time slot t is n (t) As feedback to agent n, state s n (t) is described as follows:

[0055]

[0056] in, and are the unique identifiers of the agent n acting as the transmitter and the agent n acting as the receiver, respectively. represents the mutual interference information of agent n, and the subscript n,: represents [(H t ) T H t ], the superscript T indicates the matrix transpose, a t-1 represents the actions performed by all agents in time slot t-1;

[0057] Action Space of Markov Decision Process It is defined as follows: Each agent will decide to allocate RBs and transmit power to the corresponding transmitter in each time slot; that is, the actions performed by agent n will include the transmit power and RB selection The transmit power is discrete as N P class Action Space The dimension is K×N P , the joint action at time slot t is expressed as a n (t) represents the resource allocation action of agent n in time slot t;

[0058] The reward function of the Markov decision process is defined as: t , the agent obtains an immediate reward for cooperation, which is determined by the objective function of the spectrum resource allocation problem; for agent n, the immediate reward r for performing an action at time slot t t n It is expressed as:

[0059]

[0060] Among them, t =((H t Φ t )⊙K t )I is the device interference vector, [Ψ t ] u Denotes the interference vector Ψ t The u-th element of t represents the phase shift matrix of time slot t, H t represents the channel gain matrix of time slot t, represents the resource allocation matrix, I = {1} K is a K-dimensional column vector with all elements equal to 1; ζ>0 indicates a penalty parameter that determines the size of the penalty. represents the rate of agent n at time slot t, Indicates the minimum rate of agent n.

[0061] Furthermore, in the Markov decision process, the state value function of state s is expressed as:

[0062]

[0063] in, represents the expectation, v represents the index of the time step, γ v is the exponential form of the discount factor γ, r t+v is the immediate reward at time step t+v, s t =s represents the state s at time slot t t for s;

[0064] Current Status t Select s, current action a t The action value function when selecting a is expressed as:

[0065]

[0066] The optimal strategy π for choosing the optimal action * (a|s) by maximizing Q π (s,a) to obtain, that is:

[0067]

[0068] Furthermore, for agent n, it uses a dual-depth Q network as its network model, and the dual-depth Q network includes parameters θ n The main Q network and parameters are The target Q network;

[0069] Action-value function Q π (s,a) uses the parameter θ n The main Q network is approximated as:

[0070] Q π (s,a;θ n )≈Q π (s,a);

[0071] Q π (s,a;θ n ) is expressed as:

[0072]

[0073] Among them, Qπ(s,a;θ n ) represents the Q value estimated by the main Q network of agent n when taking action a in state s under strategy π, Vπ(s; θ n ,β) represents the state value function value estimated by the main Q network of agent n in state s, A π (s,a;θn ,α) represents the advantage function value of agent n taking action a in state s under strategy π, A π (s,a';θ n ,α) represents the advantage function value of agent n taking different actions a′ in state s under strategy π, |A represents the total number of different actions, α and β represent hyperparameters;

[0074] Agent n performs the optimal action a at time slot t n The corresponding target value It is expressed as:

[0075]

[0076] Among them, r t n represents agent n performing the optimal action a at time slot t n The instant rewards you receive, represents the state of agent n at time slot t+1, The action with the largest Q value estimated by the main Q network of agent n is taken as the optimal action a n , Indicates the next state Take the best action a n The Q value estimated by the post-main Q network; The target Q network of agent n is represented by the state The best action under The estimated Q value.

[0077] Furthermore, in federated learning, each agent trains a local network model based on the mini-batch sampling in its replay buffer. The federated learning cloud server obtains the global network model by aggregating the local network models of all agents and feeds it back to all agents, who will download the same global network model to update their local network models.

[0078] For agent n, from the replay buffer The sample set For training, the loss function for:

[0079]

[0080] in, The master Q network of agent n estimates that agent n is in state at time slot t Next action The Q value, represents the sample set from agent n The i-th experience sample drawn;

[0081] Adaptive moment estimation method is used to update the parameters θ of the main Q network n ;

[0082] When updating the parameters θ of the main Q network n When the parameters of the target Q network are fixed Only in the main Q network parameter θ n Updated Target It was updated after that.

[0083] The present invention also provides a D2D-assisted ultra-dense Internet of Things resource allocation system, the key of which is that an intelligent agent is provided, and the intelligent agent is used to execute the resource allocation method.

[0084] The resource allocation method and system of the D2D-assisted ultra-dense Internet of Things provided by the present invention, in order to avoid resource conflicts between IoEDs in the D2D-assisted UD-IoE, first analyzes the conflict relationship between IoEDs; then, using clique theory and hypergraph theory, establishes a conflict hypergraph model that can simultaneously and intuitively reflect these conflict relationships; in order to reduce the complexity of determining the relationship between multiple IoEDs, the conflict hypergraph model is simplified by using clique, maximum clique and hypergraph theory; based on the simplified conflict hypergraph model, the conflict-free resource allocation problem is transformed into a point-strong coloring problem of the hypergraph, the coloring conflict of the hypergraph is analyzed, and a calculation method for the degree of color conflict is designed; then, in order to improve the network throughput, the spectrum resource allocation is described as a mixed integer linear programming (MILP) problem; in order to solve the mixed integer linear programming problem, a multi-agent DRL (deep reinforcement learning) framework is constructed for resource allocation in the entire D2D-assisted UD-IoE, and the reward function of each agent is intended to improve the data transmission rate while alleviating resource conflicts. Simulation experiments show that in the D2D-assisted UD-IoE, the method and system can achieve higher network throughput and better performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] Figure 1 This is an example diagram of a network model of a D2D-assisted ultra-dense Internet of Things provided by an embodiment of the present invention;

[0086] Figure 2 1 is an example diagram of a communication model of a D2D-assisted ultra-dense Internet of Things provided by an embodiment of the present invention, wherein (a) represents a communication network topology, and (b) represents an adjacency matrix of (a);

[0087] Figure 3 1 is an example diagram of a conflict hypergraph model of a D2D-assisted ultra-dense Internet of Things provided by an embodiment of the present invention, wherein (a) represents a conflict hypergraph model, and (b) represents an association matrix of (a);

[0088] Figure 4is an example diagram of a simplified conflict hypergraph model provided by an embodiment of the present invention, wherein (a) represents a simplified conflict hypergraph model, and (b) represents an association matrix of (a);

[0089] Figure 5 The different frequencies f provided by the embodiment of the present invention round Cumulative reward curve under federated learning;

[0090] Figure 6 is a cumulative reward curve diagram under different iteration numbers provided by an embodiment of the present invention;

[0091] Figure 7 is a histogram of network throughput of four algorithms provided by the embodiments of the present invention under different numbers of devices;

[0092] Figure 8 It is a cumulative distribution function relationship diagram of normalized network throughput of four algorithms provided in the embodiments of the present invention. DETAILED DESCRIPTION

[0093] The following specifically illustrates the implementation mode of the present invention in conjunction with the accompanying drawings. The embodiments are provided for illustrative purposes only and cannot be understood as limiting the present invention. The accompanying drawings are provided for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.

[0094] The embodiment of the present invention provides a D2D-assisted ultra-dense Internet of Things resource allocation method, comprising the steps of:

[0095] S1. Construct a network model of D2D-assisted ultra-dense IoT;

[0096] S2, building a communication model of D2D-assisted ultra-dense IoT based on the network model;

[0097] S3, build the conflict model of D2D-assisted ultra-dense Internet of Things based on network model and communication model;

[0098] S4. Based on clique theory and hypergraph theory, the conflict model is transformed into a conflict hypergraph model;

[0099] S5. Based on the conflict hypergraph simplified model and communication model, the spectrum resource allocation problem is constructed with the goal of maximizing the network throughput of D2D-assisted ultra-dense IoT.

[0100] S6. Solve the spectrum resource allocation problem based on federated multi-agent deep reinforcement learning and obtain the spectrum resource allocation strategy.

[0101] The above steps are described in detail below.

[0102] (1) Step S1: Constructing a network model for D2D-assisted ultra-dense IoT

[0103] like Figure 1 As shown, this embodiment considers a cloud architecture of D2D-assisted ultra-dense IoT, which is mainly composed of a D2D-assisted UD-IoE layer and a baseband unit (BBU) pool. In the D2D-assisted UD-IoE layer, there are multiple remote radio heads (RRHs), N IoEDs communicating through half-duplex channels and having a single antenna, which can function as a transmitter or a receiver.

[0104] These IoEDs use the D2D communication mode to form a self-organizing network. Accordingly, IoEDs can perceive the surrounding environment information through D2D links and expand their perception range. To avoid complexity, it is assumed that each IoED has a limited transmission range R n , which represents the service area of ​​the nth IoED communicating within a circle of communication radius R. RRH is connected to the BBU pool through a high-speed front-end link and is responsible for providing basic coverage and auxiliary access. The BBU pool consists of BBU (baseband unit) and computing server, and collects all environmental information from IoED through RRH. Subsequently, it allocates resources to transmitters through RRH. The total spectrum resources in the BBU pool are divided into K orthogonal resource blocks (RB). The total time in the communication network can be divided into T time slots. This example defines N IoEDs including N T transmitters and N R receivers. Set A n ={m∈N|d m,n ≤R} represents all IoEDs within the transmission range of the nth IoED, where d m,n is the distance between the mth IoED and the nth IoED.

[0105] (2) Step S2: Constructing a communication model for D2D-assisted ultra-dense IoT

[0106] In order to express the relationship between IoED and RB at time slot t, this example defines the resource allocation matrix at time slot t: To record it, the value of the element in the nth row and the kth column is as follows:

[0107]

[0108] The single interference noise rate (SINR) of the nth IoED allocated with the kth RB, ie, RB k, in time slot t is as follows:

[0109]

[0110] in, and The nth and The transmit power of each IoED, and They correspond to the nth and nth The power gain of the channel of an IoED. 2 Represents the noise power.

[0111] The data rate of the nth IoED allocated to the kth RB in time slot t can be expressed as:

[0112]

[0113] Where B is the bandwidth.

[0114] To meet the basic communication prerequisites of IoED, ensure the communication rate Exceeding the minimum transmission rate of IoED Right now:

[0115]

[0116] (3) S3: Constructing a conflict model for D2D-assisted ultra-dense IoT

[0117] In the case of half-duplex communication, according to the transmission distance A of the nth IoED n and IoED functions (the nth receiver can only detect A n The transmission relationship between IoEDs is determined by the transmitter within the network, which is also the only way to establish the communication relationship between IoEDs. Then, the communication network is modeled and represented using directed graph theory.

[0118] Given a directed graph in is a finite vertex set (i.e., IoED set), which is divided into head vertex subsets V H and the tail vertex subset V T The head vertex represents the transmitter, and the tail vertex represents the receiver. ε={e1,e2,...,e M} is the edge set (i.e., the communication relationship set between IoEDs). In addition, is defined as the set of j-th neighboring points of vertex n ( Represents a vertex The set of first-order neighbor vertices of represents the set of j-1th order neighbor vertices of vertex n), the connection relationship between vertices is called adjacency relationship, and the relationship between edges and vertices is called association relationship.

[0119] Generally speaking, the relationship between vertices can be expressed using the adjacency matrix G. A ={0,±1} N×N To represent it, its elements are defined as follows:

[0120]

[0121] Among them, v n is the head vertex, representing the adjacency matrix G A The nth row in is the tail vertex, indicating the In half-duplex communication mode, IoED can only act as a transmitter or a receiver. Figure 2 An example of a communication network topology (a) and its adjacency matrix (b). Figure 2 The adjacency matrix G shown in (b) A In the example, rows and columns will not contain both 1 and -1.

[0122] For the transmission data between IoEDs, some IoEDs will broadcast their messages as transmitters of effective broadcast, while other IoEDs will continue to listen. Only IoEDs with different functions from the nth IoED are retained. Considering that RBs are assigned to transmitters, conflicts only occur between transmitters, and the relationship between them is established with the receiver as the center. Therefore, conflicts between IoEDs are analyzed with the receiver as the center. In addition, since IoED is half-duplex, transmission and reception cannot occur at the same time (that is, there will be no conflict between the transmitter and the receiver), so transmission conflicts mainly occur in the following situations:

[0123] Conflict with next neighbor: If the nth and emitters are 2nd-order neighbors of each other (i.e., and ), and they use the same RB at the same time, then the nth and emitters conflict with other emitters, and the conflicting emitters form an edge. Figure 2 In (a), IoED 11 and IoED 13 Share RB and transmit data to IoED 12 , and then they conflict with each other.

[0124] In the directed graph model, there are two cases in D2D-assisted UD-IoE where there is no resource conflict, so they are not considered:

[0125] ① The vertex is not associated with any edge, which means it is an isolated vertex. This may indicate that the IoED has stopped working or it has no corresponding receiver or transmitter in the D2D-assisted UD-IoE;

[0126] ②The vertex is associated with only one edge. This may mean that the nth receiver is There is only one emitter in a , which means that it will not conflict with other emitters when sharing resources. Figure 2 (a) IoED 14 and IoED 15 ,In this case, arbitrary resources can be allocated to the IoED.

[0127] (4) Step S4: Convert the conflict model into a conflict hypergraph model

[0128] Based on the nature of the conflict with the secondary neighbor, the nth receiver The IoEDs in the All IoEDs in a cluster can build a fully connected subgraph, usually called a clique, following the principles of general graph theory. In the clique, each pair of IoEDs is connected by an edge, representing the conflict relationship between them. However, as the number of IoEDs increases, the complexity of determining the relationship between IoEDs increases rapidly, which may lead to a lot of computational time and resource waste. To address this problem, hypergraphs are introduced as a powerful tool to analyze the relationships between multiple vertices simultaneously to analyze the conflict relationships between multiple IoEDs in order to reduce the computational complexity.

[0129] For example, in a clique with N IoEDs, the relationship between IoEDs is determined by whether they are associated with the same edge, and then the degree of relationship determination between IoEDs is represented by N(N-1) / 2, showing a quadratic growth. In a hyperedge with N IoEDs, the relationship between IoEDs can be determined by whether they are associated with the same hyperedge, so the degree of relationship determination is N, which is a constant coefficient growth. Therefore, converting the clique into a hyperedge to represent the relationship between IoEDs will effectively reduce the computational complexity.

[0130] This example analyzes the conflict relationship between multiple IoEDs by constructing a conflict hypergraph based on the similarity between groups and hyperedges. First, a set of definitions and theorems are given:

[0131] Definition 1 (clique). Suppose an undirected graph Expressed as One of the groups is A subset of That is, for any vertex in the cluster S, there is an edge e that enables a one-to-one correspondence between any vertex in the cluster S.

[0132] Definition 2 (Maximum clique). If isomorphic to the complete graph, A subgraph of yes The maximal group in and (any vertex i in the subgraph and any vertex j among all vertices in the graph except this subgraph) such that there is no edge

[0133] Definition 3 (Hypergraph). Denoted as A hypergraph of is the vertex set, is a hyperedge set, a hyperedge yes A subset of , where the vertices in a hyperedge have the same relationship as any other vertex.

[0134] Usually, the conflict hypergraph model at time slot t can be represented by the association matrix Represents that the elements are represented as:

[0135]

[0136] Where h(v,e) = 1 means that vertex v is associated with hyperedge e, that is, hyperedge e contains vertex v. For any vertex v, u∈e, v and u are adjacent, and the set of adjacent vertices of v is denoted by N(v). represents the number of vertices, Represents the number of hyperedges.

[0137] Definition 4 (Strong coloring of vertices in a hypergraph). A strong coloring of a point is through the function φ(), so that for each hyperedge All vertices in e have different colors. Formally, if e = {v1,v2,...,v n} is a hyperedge, then φ(v1)≠φ(v2)≠...≠φ(v n ).

[0138] Theorem 1. For a graph is a maximal group, then the maximal group based on this maximal group must have

[0139] Theorem 2. For a graph Vertex Set Divided into two subsets in and Must have make Right now make At this time, with the figure In and The corresponding maximum cliques cannot be repeated.

[0140] exist In , all IoEDs adjacent to the nth receiver exhibit conflicting relationships with each other, forming a clique. This interaction can be effectively represented by hyperedges by building the definitions of cliques and hyperedges.

[0141] Method for building a hypergraph: According to the corresponding features of the receiver (i.e., tail vertex) in the adjacent matrix genetic algorithm, each column of tail vertices can construct a hyperedge containing the head vertex adjacent to the tail vertex. For example, Figure 2 In (a), IoED1 is the center. Including IoED2, IoED3, IoED8 and IoED 10 ,Right now Figure 2 In column v1 of (b), the vertex of the element that represents the communication relationship with v1 is 1, so a clique can be established. Then Figure 3 In (a), it is represented as hypergraph e1.

[0142] Based on this method, the communication network can generate a conflict hypergraph. Figure 3 The corresponding conflict hypergraph model in (a) shows Figure 2 The conflict relationship between the IoEDs in (a) is shown in the following figure. The correlation matrix of the conflict hypergraph model is Figure 3 as shown in (b).

[0143] According to the definitions of cliques and maximal cliques and Theorem 1, some cliques are inevitably included in the maximal clique, namely:

[0144]

[0145] Therefore, there are some hyperedges (i.e., sub-hyperedges) in the conflict hypergraph that are contained in other hyperedges in the conflict hypergraph, i.e. For example, hyperedge e4 is included in Figure 3 In addition, according to Theorem 2, it is known that there will be no duplication between the largest groups, so the simplification of the group can effectively reduce the number of hyperedges. Figure 4 is a simplified conflict hypergraph model example diagram provided by an embodiment of the present invention. According to formula (7), Figure 3 (a) can be simplified to Figure 4 (a), where Figure 4 The corresponding association matrix of the conflict hypergraph model shown in (a) is expressed as H t s ,like Figure 4 as shown in (b).

[0146] In A n 1In , all IoEDs adjacent to the nth receiver cannot be assigned the same RB. Therefore, based on the definition of strong vertex coloring in a hypergraph, the conflict avoidance resource allocation problem in D2D-assisted UD-IoE can be generalized to the problem of strong vertex coloring in a conflict hypergraph.

[0147] To clarify the use of color in the hyperedges of time slot t (i.e., The use of RB in the color-hyperedge relationship matrix ψ t =(i) M×K is represented as follows:

[0148]

[0149] The value of the element ψ(m,k)=i in the mth row and kth column indicates that the number of kth RBs used in the mth superedge is i times, and M represents the number of superedges.

[0150] To represent the conflicts between nodes, we use the color-hyperedge relationship matrix Φ at time slot t t A conflict degree matrix ψ is proposed t To quantify the overall conflict level of D2D-assisted UD-IoE, it is defined as follows:

[0151] Φ t =log(max(ψ t ,1)) (9)

[0152] Among them, max(ψ t ,1) represents taking the matrix ψ t The maximum value of the elements in and 1. If the matrix Φ t If the element value in is 0, there is no conflict in the conflict hypergraph.

[0153] (5) Step S5: Constructing the optimization problem

[0154] In order to improve the overall network throughput of D2D-assisted UD-IoE by optimizing spectrum resource allocation, the resource allocation problem can be described based on rate, power and conflict degree:

[0155]

[0156] Among them, max means maximization, st means need to be satisfied, ||||1 means 1 norm, constraint C1 is used to ensure that there is no overlapping interference in D2D-assisted UD-IoE, constraint C2 guarantees the minimum transmission rate requirement of IoED, and constraint C3 stipulates that there are at most N channels with different channels. TThe transmitters reuse RBs in an orthogonal manner. Constraint C4 means that each IoED transmitter can use at most one RB. Constraint C5 means that the transmitter must not exceed its maximum transmit power P when in operation. max Equation (10) is a mixed integer nonlinear programming problem. Traditional convex optimization methods involve splitting the mixed integer nonlinear programming problem into several sub-problems, which is well known to be complex and inefficient.

[0157] (6) S6: Solving the optimization problem

[0158] In order to solve the mixed integer nonlinear programming problem (10), this example constructs a multi-agent DRL (deep reinforcement learning) framework in the interactive environment of D2D-assisted UD-IoE, combines the DRL method with the FL (federated learning) method, and proposes a problem-solving algorithm called FD3QN (resource allocation method based on federated multi-agent deep reinforcement learning) in this example.

[0159] DRL can be modeled as a Markov decision process (MDP), where the interaction of an agent with its environment is defined by states, actions, and rewards. A typical MDP can be represented by a tuple with five elements, namely in represents the state space, represents the action space, represents the probability transition distribution, represents the immediate reward, and γ∈[0,1] is the discount factor. In particular, each agent (the nth agent, n=1,2,…,N T ) obtains the state at each time slot (time slot t, t = 1, 2, ..., T) by observing its surrounding environment Then perform the action at time slot t Probability transition distribution defines the probability that the current state is changed by the executed action and then transferred to a new state. Since the environment considered in this example is a D2D-assisted UDIoE system, which is model-free, no transition probability is required. In addition, the biggest advantage of DRL is its model-free scheme, which does not pre-form a state transition model to learn, but interacts with the environment. n (t) Execute action a n (t) Immediate Rewards for Results is provided to each agent. In the multi-agent DRL framework, the goal of each agent is to maximize its reward by accumulating experience in interacting with the environment. In this example, according to formula problem (10), a resource allocation model that maximizes network throughput is established, and on this basis, a resource allocation algorithm based on FD3QN is proposed. Here, each transmitter and receiver will be regarded as an agent (that is, there are n agents in total) to explore this D2D-assisted UD-IoE environment at the same time. The key components of the multi-agent DRL framework are described in detail as follows.

[0160] 1) State space: At time slot t, the state s n (t) is used as feedback to the nth agent (i.e., agent n), reflecting the impact of its resource allocation on the data transmission rate. In addition, s t ={s1(t),s2(t),…,s N (t)} represents the state set of all agents in the entire network at time slot t. Therefore, the state s of agent n at time slot t n (t) is described as follows:

[0161]

[0162] in, and Represent the unique identifiers of agent n as the transmitter and the receiver respectively. represents the mutual interference information of agent n (subscript n,: represents [(H t ) T H t ], the superscript T indicates matrix transposition), a t-1 represents the actions performed by all agents in time slot t-1.

[0163] 2) Action Space: In the multi-agent DRL framework, each agent will decide which RB and how much transmit power should be allocated to the corresponding transmitter at each time slot. Therefore, the action performed by the nth agent will include the transmit power and RB selection In order to reduce the difficulty of processing continuous values ​​such as transmit power, it is necessary to discretize the transmit power into N P Level, that is According to the constraints, each agent only needs to be allocated a single RB for data transmission. Then, the action space The dimension is K×N P , each action corresponds to a specific combination of RB selection and power control. The joint actions of the t+1 .

[0164] 3) Reward function: When executing at When , the agent receives an immediate reward for cooperation, which is determined by the objective function of the resource allocation problem. This example has two goals: to maximize the transmission rate and to reduce the impact of mutual interference within a certain time constraint T. Here, in order to more accurately quantify the impact of mutual interference of each agent, this example defines the device interference vector Ψ t , the formula is as follows:

[0165] Ψ t =((H t Φ t )⊙K t )I, (12)

[0166] in, Φ t The phase shift matrix (Phase Shift Matrix) of time slot t is usually used to adjust the phase of each signal to optimize the transmission performance. t represents the channel gain matrix (Channel GainMatrix) of time slot t, which reflects the channel state information between nodes in the network at time slot t. Since each agent as a transmitter only uses one RB, Resource Allocation Matrix (1 means allocating resources to the agent, 0 means not allocating resources) Perform Hadamard product, I = {1} K is a K-dimensional column vector with all elements equal to 1. Therefore, for agent n, the immediate reward It is expressed as:

[0167]

[0168] Among them, ζ>0 represents the penalty parameter that determines the penalty size, [Ψ t ] u Denotes the interference vector Ψ t The u-th element of represents the rate of agent n at time slot t, Indicates the minimum rate of agent n.

[0169] 4) Strategy and value function: For each agent, its strategy π(a|s) is an action selection strategy that optimizes long-term rewards. The state value function of state s, that is, the expected return at the beginning of state s and after the strategy π, can be expressed as:

[0170]

[0171] Indicates the expectation, v represents the index of the time step, the range of v is 0 to infinity, and in calculating the infinite return in the future, v = 0 represents the return of the current time step t, v = 1 represents the return of the next time step, and so on. v is the exponential form of the discount factor γ, which represents the importance of the return at time step t+v relative to the current time step t. As v increases, γ v The value of will gradually decrease, indicating that the importance of future returns decreases. t+v That is, the immediate reward at time step t+v.

[0172] Then, when action a is selected at time slot t in state s, the action-value function (also denoted as Q-value) represents the expected return under the guidance of policy π, which can be expressed as:

[0173]

[0174] Obviously, the optimal strategy π for choosing the optimal action * (a|s) can be obtained by maximizing Q π (s,a) to obtain, that is:

[0175]

[0176] When the optimal action-value function of all agents is explored with ∈, the greedy policy is improved by randomly selecting an action with probability ∈ and selecting a greedy action with probability 1-∈.

[0177] In the dual duel deep Q network (D3QN), for agent n, there is a parameter θ n The main Q network and parameters are The target Q network, where θ n and are the parameters of the neural network. Then, the action value function Q of agent n π (s,a) using parameter θ n The main Q network is approximated as:

[0178] Q π (s,a;θ n )≈Q π (s,a) (17)

[0179] For the action-value function Q of the nth agent (i.e. agent n) π (s,a;θ n ), which can be obtained by combining the state value and action advantage and then subtracting the average advantage of the action. Therefore, based on equations (14) and (15), the action value function of agent n is expressed by the average advantage method:

[0180]

[0181] Among them, Q π (s,a;θ n ) represents the Q value of taking action a in state s under policy π estimated by the main Q network of agent n, V π (s;θ n ,β) represents the state value function value estimated by the main Q network of agent n in state s, A π (s,a;θ n ,α) represents the advantage function value of agent n taking action a in state s under strategy π, A π (s,a';θ n ,α) represents the advantage function value of agent n taking different actions a′ in state s under strategy π, |A| represents the total number of different actions, and α and β represent hyperparameters (parameters pre-set during the training process).

[0182] Due to the use of the double Q network, agent n performs the optimal action a at time slot t n The corresponding target value for:

[0183]

[0184] Among them, r t n represents agent n performing the optimal action a at time slot t n The instant rewards you receive, represents the state of agent n at time slot t+1 (the next time slot) (referred to as the next state), The action with the largest Q value estimated by the main Q network of agent n is taken as the optimal action a n (In a given state What action should the agent take? n To get the biggest cumulative reward), Indicates the next state Take the best action a n The Q value estimated by the post-main Q network; The target Q network of agent n is represented by the state The best action under The estimated Q value.

[0185] Federated learning (FL) is a distributed machine learning model training method, where the training process is performed on a single agent and the training data is stored locally. Specifically, in the considered D2D-assisted UD-IoE, each agent trains a local D3QN model based on mini-batch sampling in its replay buffer. The training data of each agent is utilized locally instead of being uploaded to the FL cloud server to avoid privacy leakage. The FL cloud server obtains the global D3QN model by aggregating all local D3QN models and feeds it back to all agents, which will download the same global D3QN model to update their local models. The model training process generally includes the following four periods.

[0186] 1) Local update: Each agent will update from the replay buffer Sample a small batch of I samples from the , expressed as a sample set For agent n, its loss function The formula is:

[0187]

[0188] in, represents the main Q network of agent n (whose parameters are θ n ) The estimated agent n in time slot t is in state Next action First, through the inner layer Operation finds the state in the next time slot Next, take the action that maximizes the Q value, and then use this action to calculate the target Q value. represents the sample set from agent n The i-th experience sample drawn.

[0189] Therefore, the parameters of the main Q network of each agent can be updated by the adaptive moment estimation (ADAM) method. For the parameters θ of the main Q network of agent n n , its parameter update can be expressed as:

[0190]

[0191] Where η represents the learning rate, t represents the current time slot, and t+1 refers to the next time slot. Then, when updating the parameters θ of the main Q network n When the parameters of the target Q network are fixed Only in the main Q network parameter θ n Updated Target The parameters of the target Q network are updated using a soft update factor τ≤1, which can be expressed as:

[0192]

[0193] Here i indicates the current update parameter round, i+1 is the next round.

[0194] 2) Local model upload: After local update, agent n uploads the parameters of the main Q network and the target Q network Upload to FL cloud server.

[0195] 3) Global Aggregation: After each round of joint communication, when all uploaded local D3QN models are received, the FL cloud server aggregates the models through a typical FL algorithm called Federated Averaging (FedAvg). In this process, it will produce an updated global D3QN model, which can be expressed as:

[0196]

[0197]

[0198] Among them, p n represents the weight of agent n, λ represents the current round of updating the global model parameters, and λ+1 refers to the next round.

[0199] Here, this example assumes that all agents must participate in each round of federated communication. In this way, distributed participants can collaboratively train the global model.

[0200] 4) Global model download: FL cloud server broadcasts global D3QN model parameters to all agents Each agent then updates its local D3QN model by downloading the global D3QN model parameters, which can be expressed as:

[0201]

[0202]

[0203] Based on the above method, this embodiment further provides a D2D-assisted ultra-dense Internet of Things resource allocation system, which is provided with an intelligent agent, and the intelligent agent is used to implement steps S1 to S6 in the above method.

[0204] This example uses Minconda 3 and Python 3.11.7 for simulation, and all algorithms are implemented on a PC with one CPU (Intel Xeon W5-3425 12C24T 3.2-4.6GHz) and two GPUs (NVIDIA RTX 4090 24G). The PC has 32GB of memory. In this simulation, multiple devices are randomly deployed in an area of ​​800×800 square meters. Then, this example assumes that the communication radius of each device is 200m. For reference, the main simulation parameters are summarized in Tables 1 and 2.

[0205] Table 1

[0206]

[0207] Table 2

[0208]

[0209] During training, the agent randomly selects an action to explore with probability ∈ through the ∈-greedy strategy to try new decision paths. ∈ start is the value of ∈ at the beginning of the ∈-greedy strategy, indicating the probability of the agent exploring unknown actions in the early stages of learning. ∈ dec Represents the decay rate of the ∈ value, which is used to adjust the speed at which the ∈ value decreases over time, that is, how ∈ gradually decreases with each time step or each training cycle. ∈ min is the minimum ∈ value in the ∈-greedy strategy, ensuring that even after long-term learning, there is still a certain probability of exploration, preventing the agent from stopping exploration completely.

[0210] In order to evaluate the effect of the resource allocation algorithm (FD3QN algorithm) used in this example, this example simulates the following methods for performance comparison.

[0211] 1) Multi-Agent Deep Q Network (MADQN): In this algorithm, each device as a transmitter will be regarded as an agent and utilize the locally collected information to update its own DQN model for resource allocation.

[0212] 2) Single Agent Deep Q Network (SADQN): In this algorithm, there is only one transmitter as an agent responsible for updating its DQN model using the information it obtains locally for resource allocation. Other agents replicate the DQN model under the guidance of the RRH.

[0213] 3) Baseline: This example adopts this method and randomly selects RBs to transmit power for the entire network in each time slot t.

[0214] Figure 5The performance of the proposed federated multi-agent deep reinforcement learning framework in a simulation environment (with 60 devices) is shown, where moving average smoothing is applied to the curve to obtain a smoother curve. Figure 5 The relationship between the cumulative reward and the number of FL transmission rounds is shown. First, as the number of devices increases, the cumulative reward increases simultaneously. The proposed FD3QN-based resource allocation scheme is particularly able to alleviate mutual interference and achieve higher network throughput. It is worth noting that as the frequency of joint communication decreases, that is, as FROUND increases, the cumulative reward will be more unstable. After each FL communication is completed, the performance of the D3QN model decreases and then immediately decreases. Figure 5 As shown in Figure 1, the dotted lines of the corresponding colors are the upper and lower bounds of the original data. Therefore, it is not difficult to see that as the frequency of FL communication increases, the cumulative reward after convergence is more stable.

[0215] Figure 6 The cumulative reward of each algorithm fluctuates as the amount of data increases (the number of devices is 60). It can be observed that the proposed FD3QN algorithm and MADQN algorithm show similar trends. The proposed FD3QN-based resource allocation scheme achieves higher rewards compared to other comparison schemes. In the approach of this example, agents with higher learning levels share valuable experience with others, allowing them to learn from a stronger knowledge base. This reduces the exploration time and speeds up the training process.

[0216] To demonstrate the advantages of the FD3QN algorithm, Figure 7 The network throughput achieved by different algorithms is compared, where the x-axis represents five different device counts. Figure 7 As shown in Figure 2, in D2D-assisted UD-IoE, the network throughput increases with the number of devices. When the number of devices is small, the performance of the proposed FD3QN algorithm is similar to that of the MADQN algorithm, but it has a higher network throughput as the number of devices increases. Figure 6 It shows that SADQN converges slower than the proposed algorithm and MADQN algorithm. The network throughput of the SADQN algorithm is much lower than that of the proposed algorithm and MADQN algorithm, such as Figure 7 As shown. To this end, the proposed FD3QN algorithm can be combined with the proposed conflict hypergraph model to dynamically allocate RBSs and control the transmit power, and the proposed FD3QN algorithm can achieve higher network throughput.

[0217] also, Figure 8The relationship between the cumulative distribution function (CDF) of the normalized network throughput during the 50ms dynamic process and different algorithms is shown. Here, the number of devices for D2D-assisted UD-IoE is 60. Since the SADQN algorithm shares a DRL model with all devices as transmitters, its performance is similar to the baseline algorithm. For this reason, the performance of the proposed FD3QN algorithm is much better than other solutions.

[0218] In summary, the resource allocation method and system for the D2D-assisted ultra-dense Internet of Things provided by the embodiment of the present invention, in order to avoid resource conflicts between IoEDs in the D2D-assisted UD-IoE, first analyzes the conflict relationship between IoEDs; then, using the clique theory and the hypergraph theory, establishes a conflict hypergraph model that can simultaneously and intuitively reflect these conflict relationships; in order to reduce the complexity of determining the relationship between multiple IoEDs, the clique, the maximum clique and the hypergraph theory are used to simplify the conflict hypergraph model; based on the simplified conflict hypergraph model, the conflict-free resource allocation problem is transformed into a point-strong coloring problem of the hypergraph, the coloring conflict of the hypergraph is analyzed, and a method for calculating the degree of color conflict is designed; then, in order to improve the network throughput, the spectrum resource allocation is described as a mixed integer linear programming (MILP) problem; in order to solve the mixed integer linear programming problem, a multi-agent DRL (deep reinforcement learning) framework is constructed for resource allocation in the entire D2D-assisted UD-IoE, and the reward function of each agent is designed to improve the data transmission rate while alleviating resource conflicts. Simulation experiments show that in D2D-assisted UD-IoE, the method and system can achieve higher network throughput and better performance.

[0219] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. A D2D-assisted ultra-dense Internet of Things resource allocation method, characterized in that: Includes steps: S1. Construct a network model of D2D-assisted ultra-dense IoT; The network model consists of a D2D-assisted UD-IoE layer and a BBU pool. D2D refers to device-to-device, and UD-IoE refers to ultra-dense Internet of Things. In the D2D-assisted UD-IoE layer, there are multiple remote radio heads RRH and N mobile Internet of Things devices IoED that communicate through a half-duplex channel and have a single antenna. The N IoEDs include N T transmitters and N R receivers; IoED uses D2D communication mode to form a self-organizing network. IoED senses the surrounding environment information through D2D links. Each IoED has a transmission range R n , R n It represents the service area where the nth IoED communicates within the circle of the communication radius R; RRH is connected to the BBU pool through a high-speed front-end link and is responsible for providing basic coverage and auxiliary access; The BBU pool collects all environmental information from the IoED through the RRH and subsequently allocates resources to the transmitter through the RRH; The total spectrum resources in the BBU pool are divided into K orthogonal resource blocks RB; The total time in the communication network is divided into T time slots; S2. Constructing a communication model of D2D-assisted ultra-dense Internet of Things based on the network model; S3, constructing a conflict model of D2D-assisted ultra-dense Internet of Things based on the network model and the communication model; S4. transforming the conflict model into a conflict hypergraph model based on clique theory and hypergraph theory; S5. Based on the conflict hypergraph simplified model and the communication model, a spectrum resource allocation problem is constructed with the goal of maximizing the network throughput of the D2D-assisted ultra-dense Internet of Things; S6. Solve the spectrum resource allocation problem based on federated multi-agent deep reinforcement learning to obtain a spectrum resource allocation strategy.

2. The D2D-assisted ultra-dense Internet of Things resource allocation method according to claim 1, characterized in that: In step S2, the communication model includes: Resource allocation matrix at time slot t The element in the nth row and kth column The values ​​are as follows: Single interference noise rate of the nth IoED allocated with the kth RB in time slot t in, and The nth and The transmit power of each IoED, and They correspond to the nth and kth RBs respectively. The power gain of the IoED channel, σ 2 represents the noise power; Data rate of the nth IoED allocated to the kth RB in time slot t Where B is the bandwidth; Data Rate Not less than the corresponding minimum transmission rate in, Indicates any.

3. The resource allocation method for D2D-assisted ultra-dense Internet of Things according to claim 2, characterized in that: In step S3, the conflict model is constructed as in is the IoED set, which is divided into the head vertex subset V H and the tail vertex subset V T , the head vertex represents the transmitter, and the tail vertex represents the receiver; ε={e1,e2,…,e M } is the communication relationship set between IoEDs; The relationship between vertices is represented by the adjacency matrix G A ={0,±1} N×N To indicate that the nth row The meaning of the column elements are as follows: Among them, v n is the head vertex, representing the adjacency matrix G A The nth row in is the tail vertex, indicating the List; IoEDs only collide with each other, and for conflicts with secondary neighbors: if the nth and nth emitters are 2nd-order neighbors of each other, and they use the same RB at the same time, then the nth and nth An emitter conflicts with other emitters, and the conflicting emitters form an edge.

4. The resource allocation method for D2D-assisted ultra-dense Internet of Things according to claim 3 is characterized by: In step S4, the conflict hypergraph model is constructed as in is the vertex set, is a superedge set; The conflict hypergraph model at time slot t is composed of the incidence matrix Represents that the elements are represented as: Among them, h(v,e) = 1 means that vertex v is associated with hyperedge e. represents the number of vertices, represents the number of hyperedges; According to the corresponding features of the tail vertices in the adjacency matrix genetic algorithm, a hyperedge containing the head vertices adjacent to the tail vertex is constructed for each column of the tail vertices, and the conflict hypergraph model is simplified using clique, maximum clique and hypergraph theory; The association matrix of the simplified conflict hypergraph model is expressed as Define the color-hyperedge relationship matrix ψ t =(i) M×K It is expressed as: The overall conflict level of D2D-assisted ultra-dense IoT is defined as: F t =log(max(ψ t ,1)), Among them, max(ψ t ,1) represents taking the matrix ψ t The maximum value of the elements in and 1, if the matrix Φ t If the element value in is 0, there is no conflict in the conflict hypergraph.

5. The resource allocation method for D2D-assisted ultra-dense Internet of Things according to claim 4, characterized in that: In step S5, the spectrum resource allocation problem is constructed as follows: Among them, max means maximization, st means need to be satisfied, || ||1 means 1 norm, C1 to C5 represent different constraints, P max Indicates the maximum transmit power of the transmitter.

6. The resource allocation method for D2D-assisted ultra-dense Internet of Things according to claim 5, characterized in that: The step S6 specifically includes: A multi-agent deep reinforcement learning framework is constructed in the interactive environment of D2D-assisted UD-IoE, and deep reinforcement learning and federated learning methods are combined to solve the spectrum resource allocation problem; The multi-agent deep reinforcement learning framework is modeled as a Markov decision process, where each IoED is regarded as an agent and the goal of each agent is to maximize its reward by accumulating experience in interacting with the environment; The state space of a Markov decision process Defined as: the state set s of all agents in the entire network at time slot t t ={s1(t),s2(t),…,s N (t)}, where agent n is in state s at time slot t n (t) As feedback to agent n, state s n (t) is described as follows: in, and denotes the unique identifier of the agent n as the transmitter and the agent n as the receiver, respectively, [(H t ) T H t ] n,: represents the mutual interference information of agent n, and the subscript n,: represents [(H t ) T H t ], the superscript T indicates the matrix transpose, a t-1 represents the actions performed by all agents in time slot t-1; Action Space of Markov Decision Process Defined as: The action performed by agent n includes the transmission power and RB selection The transmit power is discrete as N P class Action Space The dimension is K×N P , the joint action at time slot t is expressed as a n (t) represents the resource allocation action of agent n in time slot t; The reward function of the Markov decision process is defined as: t , the agent obtains an immediate reward for cooperation, which is determined by the objective function of the spectrum resource allocation problem; for agent n, the immediate reward r for performing an action at time slot t t n It is expressed as: Among them, t =((H t Φ t )⊙K t )I is the device interference vector, [Ψ t ] u Denotes the interference vector Ψ t The u-th element of t represents the phase shift matrix of time slot t, H t represents the channel gain matrix of time slot t, represents the resource allocation matrix, I = {1} K is a K-dimensional column vector with all elements equal to 1; ζ>0 indicates a penalty parameter that determines the size of the penalty. represents the rate of agent n at time slot t, Indicates the minimum rate of agent n.

7. The D2D-assisted ultra-dense Internet of Things resource allocation method according to claim 6, characterized in that: In the Markov decision process, the state value function of state s is expressed as: in, represents the expectation, v represents the index of the time step, γ v is the exponential form of the discount factor γ, r t+v is the immediate reward at time step t+v, s t =s represents the state s at time slot t t for s; Current Status t Select s, current action a t The action value function when selecting a is expressed as: The optimal strategy π for choosing the optimal action * (a|s) is obtained by the following formula:

8. The D2D-assisted ultra-dense Internet of Things resource allocation method according to claim 7, characterized in that: For agent n, it uses a double-depth Q network as its network model, which includes parameters θ n The main Q network and parameters are The target Q network; Action-value function Q π (s,a) uses the parameter θ n The main Q network is approximated as: Q π (s,a;θ n )≈Q π (s,a); Q π (s,a;θ n ) is expressed as: Among them, Q π (s,a;θ n ) represents the Q value estimated by the main Q network of agent n when taking action a in state s under policy π, V π (s;θ n ,β) represents the state value function value estimated by the main Q network of agent n in state s, A π (s,a;θ n ,α) represents the advantage function value of agent n taking action a in state s under strategy π, A π (s,a';θ n ,α) represents the advantage function value of agent n taking different actions a′ in state s under strategy π, |A| represents the total number of different actions, and α and β represent hyperparameters; Agent n performs the optimal action a at time slot t n The corresponding target value It is expressed as: Among them, r t n represents agent n performing the optimal action a at time slot t n The instant rewards you receive, represents the state of agent n at time slot t+1, The action with the largest Q value estimated by the main Q network of agent n is taken as the optimal action a n , Indicates the next state Take the best action a n The Q value estimated by the post-main Q network; The target Q network of agent n is represented by the state The best action under The estimated Q value.

9. The resource allocation method for D2D-assisted ultra-dense Internet of Things according to claim 8, characterized in that: In federated learning, each agent trains a local network model based on the mini-batch sampling in its replay buffer. The federated learning cloud server aggregates the local network models of all agents to obtain a global network model and feeds it back to all agents. These agents will download the same global network model to update their local network models. For agent n, from the replay buffer The sample set For training, the loss function for: in, The master Q network of agent n estimates that agent n is in state at time slot t Next action The Q value, represents the sample set from agent n The i-th experience sample drawn; Adaptive moment estimation method is used to update the parameters θ of the main Q network n ; When updating the parameters θ of the main Q network n When the parameters of the target Q network are fixed Only in the main Q network parameter θ n Updated Target It was updated after that.

10. A D2D-assisted ultra-dense Internet of Things resource allocation system, characterized by: An intelligent agent is provided, and the intelligent agent is used to execute the resource allocation method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Ultra-dense Internet of Things resource allocation method and system based on GCN-DDPG

    CN117640417A

  • Resource allocation method supporting 6G NGMA IoE network

    CN118175646A