A strategy optimization method for a cell-free massive MIMO system

By building a user association and energy consumption model and combining it with a graph attention reinforcement learning algorithm to optimize content caching and resource allocation in the CF-mMIMO network, we solved the caching and resource allocation problems caused by dynamic user changes, improved network speed and energy efficiency, and adapted to diverse needs.

CN119095076BActive Publication Date: 2025-10-21CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411149227.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2025-10-21
Estimated Expiration
2044-08-21

AI Technical Summary

Technical Problem

The dynamic changes in user demands in existing CF-mMIMO networks affect AP cache deployment and resource allocation, resulting in increased content acquisition latency and poor service quality. Traditional strategies are difficult to adapt to dynamic network environments.

Method used

A user association, downlink signal and system energy consumption model is constructed, and a multi-agent reinforcement learning algorithm based on graph attention is used to optimize content caching, user association and power allocation strategies, which are abstracted into a partially observable Markov decision process model to achieve autonomous decision-making and resource management.

Benefits of technology

It improves network speed and energy efficiency, meets different business needs, adapts to diverse content caching and dynamic network environments, and optimizes interference control during content delivery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119095076B_ABST
    Figure CN119095076B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of strategy optimization algorithm of cell-free massive MIMO system, including the construction user association model, user association model is used to model the association between mobile device MD and access point AP in each time slot;Downlink signal model is constructed, downlink signal model is used to model the network reachable rate of cell-free massive MIMO system in each time slot;System energy consumption model is constructed, system energy consumption model is used to model the total energy consumption of all access points AP in each time slot to provide service;According to the user association model, downlink signal model and system energy consumption model constructed, target optimization problem model is constructed, and the optimal strategy is obtained by using the multi-agent reinforcement learning algorithm based on graph attention to solve.It is reduced to the energy consumption of the present application, improves user service quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of mobile communication technology, and in particular relates to a strategy optimization method for a non-cellular large-scale MIMO system. Background Art

[0002] With the rapid development of the Internet of Things (IoT) and edge computing technologies, global communication traffic has exploded, posing new technical challenges to wireless network architectures. Due to the inherent structured boundaries and centralized management mechanisms of traditional cellular networks, they are no longer able to efficiently meet the current complex and high-density communication service demands. Consequently, the wireless communications field has begun to research and explore new network architectures, with cell-free massive MIMO (CF-mMIMO) being considered a promising solution. Unlike the fixed boundaries of traditional cellular MIMO systems, CF-mMIMO provides services to users through a large number of distributed nodes, achieving more uniform network coverage. Its decentralized nature significantly improves the flexibility and adaptability of communication networks.

[0003] Most existing technologies target traditional CF-mMIMO network scenarios and consistent user service requirements, and fail to fully consider the coupling relationship between content caching deployment, user association, and resource allocation in multi-user CF-mMIMO network environments. In actual network scenarios, diverse user service requirements, decentralized resource deployment, and dynamic network environments make resource management in CF-mMIMO networks extremely complex. First, diverse user service requirements lead to significant spatial differences in content caching and resource allocation among different AP nodes, posing challenges to efficient collaboration between AP nodes. Second, the decentralized resource deployment of CF-mMIMO networks leads to uneven resource distribution, affecting management performance between APs and network resource utilization efficiency. Furthermore, the dynamic nature of the network environment can also lead to deep uncertainty in network status and resource availability. Traditional user association and resource allocation strategies are no longer able to adapt to changes in the highly dynamic network environment.

[0004] In summary, the existing technical problem is that the dynamically changing user demands in CF-mMIMO content caching networks affect AP cache deployment and resource allocation. However, the wireless AP cache capacity is limited and cannot store all the services that users may request. Therefore, missing content needs to be obtained from the fronthaul link or backhaul link, which will increase the user content acquisition delay and poor service quality. Summary of the Invention

[0005] To address the problems in the background technology, the present invention proposes a policy optimization method for a cell-free massive MIMO system to learn optimal content caching, user association, and power allocation strategies. The system comprises: N access points (APs) with cache resources and M mobile devices (MDs), and is characterized by:

[0006] S1: Build a user association model, which is used to model the association relationship between the mobile device MD and the access point AP in each time slot;

[0007] S2: Build a downlink signal model, where the downlink signal model is used to model the network achievable rate of the non-cellular massive MIMO system in each time slot;

[0008] S3: Build a system energy consumption model, where the system energy consumption model is used to model the total energy consumption of all access points (APs) providing content services in each time slot.

[0009] S4: Construct a target optimization problem model based on the constructed user association model, downlink signal model, and system energy consumption model;

[0010] S5: The target optimization problem model is constructed as a partially observable Markov decision process model, and the graph attention-based multi-agent reinforcement learning algorithm is used to solve it to obtain the association strategy between the access point AP and the mobile device MD, the content caching strategy of the access point AP, and the power allocation strategy of the access point AP.

[0011] Preferably, the user association model includes:

[0012]

[0013]

[0014]

[0015] in, Indicates time slot With the i-th mobile device MD i The associated access point AP set, represents a set of time slots; Indicates that the i-th mobile device MD in time slot t i and the jth access point AP j the relationship between Representing a collection The number of access points (APs) in the network; Represents the set of all mobile device MDs, Represents the set of all access points AP.

[0016] Preferably, the downlink signal model includes:

[0017]

[0018]

[0019] g ij (t)=(d ij / d0)-αh ij (t)

[0020]

[0021] Among them, R sum (t) represents the network achievable rate of the non-cellular massive MIMO system at time slot t; R i (t) represents the i-th mobile device MD at time slot t i The received signal rate; Indicates the time slot t and the i-th mobile device MD i Associated access point AP set; represents the set of all mobile devices MD that have service requirements in time slot t; v ij (t) represents the i-th mobile device MD at time slot t i and the jth access point AP j The relationship between P ij (t) represents the jth access point AP at time slot t j Assigned to the i-th mobile device MD i transmission power; represents the jth access point AP j Maximum transmission power; g ij (t0 represents the jth access point AP j To the i-th mobile device MD i Downlink channel; represents the set of all APs in service at time slot t; Indicates the jth access point AP in time slot t j To the i′th mobile device MD i′ The estimated channel gain P i′j (t) represents the jth access point AP at time slot t j Assigned to the i′th mobile device MD i′ The transmission power of i′j (t) represents the jth access point AP at time slot t j and the i′th mobile device MD i′ The relationship between i (t) represents the i-th mobile device MD at time slot t i Received interference signal; d ijrepresents the jth access point AP j and the i-th mobile device MD i The actual distance between them; d0 represents the reference distance; α is the path attenuation factor; h ij (t) indicates that it obeys the complex Gaussian distribution small-scale fading.

[0022] Preferably, the system energy consumption model includes:

[0023]

[0024]

[0025]

[0026]

[0027]

[0028]

[0029] Where P(t) represents the total energy consumption of all serving APs in time slot t, represents the set of all APs in service at time slot t; P j (t) represents the jth access point AP at time slot t j Total energy consumption; Indicates the jth access point AP in time slot t j Downlink content transmission power consumption; P ij (t) represents the jth access point AP at time slot t j Assigned to the i-th mobile device MD i transmission power; Indicates the jth access point AP in time slot t j The MD collection of the service; Indicates the jth access point AP in time slot t j Energy consumption for updating or replacing service content; P FL The energy consumption required to transmit unit content on the fronthaul link; represents the jth access point AP j The set of service contents cached in time slot t, represents the jth access point AP j The maximum content cache capacity that can be cached; F represents the number of service content types in the network; represents the jth access point AP j The set of service contents cached in time slot t-1; Indicates the jth access point AP in time slot t jEnergy consumption of content forwarding; pen represents the jth access point AP j Get the energy consumption factor of an uncached content; Indicates that in time slot t, the access point AP j The collection of all content requests of all associated mobile device MDs; Indicates that the i-th mobile device MD in time slot t i The content of the service requested, Represents a collection of service content in the network.

[0030] Preferably, the target optimization problem model P1 includes:

[0031]

[0032]

[0033]

[0034]

[0035]

[0036]

[0037] in, is the content caching strategy of access point AP in time slot t, represents the jth access point AP in time slot t j The set of service contents cached in time slot t, where N represents the number of access points AP; represents the association strategy between the access point AP and the mobile device MD in time slot t, Indicates time slot With the i-th mobile device MD i Associated access point AP set; The power allocation strategy of the access point AP. For AP j The power allocation set, P ij (t) At time slot t, the jth access point AP j Assigned to the i-th mobile device MD i The transmission power; T represents the number of time slots; represents mathematical expectation; R sum (t) represents the network achievable rate of the non-cellular massive MIMO system at time slot t; P(t) represents the total energy consumption of all serving APs at time slot t; v ij (t) represents the i-th mobile device MD at time slot t i and the jth access point AP jthe relationship between represents the jth access point AP j The maximum power value; Indicates AP j The MD collection of the service; represents the jth access point AP j The maximum content cache capacity that can be cached; Represents a collection of mobile devices; Indicates an access point AP set.

[0038] Preferably, constructing the target optimization problem model as a partially observable Markov decision process model includes:

[0039] The target optimization problem model is converted into a Dec-POMDP model with N access points AP, where each AP represents an agent and is represented by a tuple Represents the global network environment status of the non-cellular massive MIMO system; represents the jth access point AP j The local observation space of represents the jth access point AP j The action space, R is the reward function, γ∈[0,1) represents the discount factor;

[0040] At time slot t, the environmental state Defined as:

[0041]

[0042]

[0043] l1(t),l2(t),...,l M (t)}

[0044] in, represents the jth access point AP in time slot t j Content cache status; k fj (t) = 1 indicates the jth access point AP j Content f has been cached at time slot t, otherwise k fj (t) = 0; g ij (t) represents the jth access point AP in time slot t j With the i-th mobile device MD i The channel gain between i (t) represents the ith mobile device MD i location information;

[0045] At time slot t, local observation Defined as:

[0046]

[0047] At time slot t, the action space Defined as:

[0048]

[0049] in, represents the jth access point AP j The set of service contents cached in time slot t; represents the jth access point AP j The set of mobile devices MD associated at time slot t; Indicates the jth access point AP in time slot t j The power allocation set of

[0050] At time slot t, the reward function r(t)∈R is defined as:

[0051]

[0052] Among them, R sum (t) represents the network achievable rate of the non-cellular massive MIMO system at time slot t; P(t) represents the total energy consumption of all serving APs at time slot t.

[0053] Preferably, the multi-agent reinforcement learning algorithm based on graph attention is used to solve the problem, comprising: a local action value network, a graph attention module and a hybrid module;

[0054] The local action value network configures a deep Q network composed of multi-layer perceptrons for each agent. At time slot t, the agent AP j Receive local observation value o j (t), and choose action a j (t), o j (t) and a j (t) Input deep Q network output local action value Q j (o j (t),a j (t));

[0055] The graph attention module first inputs the environment state s(t) into the MLP encoder and encodes s(t) into local potential representation vectors h1(t),h2(t),...,h N (t), where h j (t) represents the jth access point AP j feature representation; GAT is then used to adaptively capture the correlation between agents to obtain the agent APj The eigenvector of h′ j (t); then the agent AP j The eigenvector of h′ j (t) Input MLP as agent AP j The local action value generates weight w j (t);

[0056] The hybrid module is based on the local action value Q j (o j (t),a j (t)) and the local action value w j (t) Calculate the joint action value

[0057] The reinforcement learning model is trained by minimizing the loss function, namely:

[0058]

[0059] Among them, θ represents the parameters of the evaluation network, X represents the number of mini-batch samples randomly sampled from the experience replay pool, x represents the sample number, and y tot =r+γmax a′ Q tot (s′,a′;θ - ), r represents the reward, a and a′ represent the action, s and s′ represent the environment state; θ represents the parameters of the target network;

[0060] The target optimization problem model is solved by the trained reinforcement learning model to obtain the association strategy between the access point AP and the mobile device MD, the content caching strategy of the access point AP, and the power allocation strategy of the access point AP.

[0061] The present invention has at least the following beneficial effects

[0062] To address dynamic, time-varying network environments and incomplete network state observations, this paper abstracts the aforementioned joint optimization problem into a decentralized partially observable Markov decision process (Dec-POMDP). It then designs autonomous decision-making for content cache deployment, user association, and transmission power control. Taking into account the diverse content caching requirements and wide-area, differentiated network spatial characteristics in CF-mMIMO scenarios, a graph attention network is employed to learn and capture network spatial characteristics, enabling adaptive interference control during content delivery and meeting diverse service requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1This is a diagram of the CF-mMIMO multi-user multi-content caching network scenario of the present invention;

[0064] Figure 2 It is the GMJOC algorithm block diagram of the present invention;

[0065] Figure 3 It is the convergence performance curve of the GMJOC algorithm of the present invention;

[0066] Figure 4 This is a graph showing how the network rate of the GMJOC algorithm of the present invention changes as the number of training rounds increases;

[0067] Figure 5 This is a curve diagram of the change in power consumption of the GMJOC algorithm of the present invention as the number of training rounds increases. DETAILED DESCRIPTION

[0068] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0069] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.

[0070] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0071] See also Figure 1 and Figure 2The present invention provides a strategy optimization method for a non-cellular massive MIMO system, wherein the system comprises: N access points AP with cache resources and M mobile devices MD, including:

[0072] S1: Build a user association model, which is used to model the association relationship between the mobile device MD and the access point AP in each time slot;

[0073] S2: Build a downlink signal model, where the downlink signal model is used to model the network achievable rate of the non-cellular massive MIMO system in each time slot;

[0074] S3: Build a system energy consumption model, where the system energy consumption model is used to model the total energy consumption of all access points (APs) providing content services in each time slot.

[0075] S4: Construct a target optimization problem model based on the constructed user association model, downlink signal model, and system energy consumption model;

[0076] S5: The target optimization problem model is constructed as a partially observable Markov decision process model, and the graph attention-based multi-agent reinforcement learning algorithm is used to solve it to obtain the association strategy between the access point AP and the mobile device MD, the content caching strategy of the access point AP, and the power allocation strategy of the access point AP.

[0077] In this embodiment, a typical CF-mMIMO network with multiple users and multiple content caches is considered, such as Figure 1 As shown, the network contains N APs and M mobile devices (MDs), and the AP and MD sets are defined as and Among them, different MDs have differentiated content requirements, that is, different MDs generate different content requests and are associated with different APs based on the current network status (distance to each AP, channel status, content deployment status); different APs cache corresponding service content based on content requirements within the service range and are connected to the central processing unit (CPU) through optical fiber links. Assume that there are F types of service content of equal size in the network, define Represents the set of all service contents. The network operates in discrete time slots, and the time slot set is defined as

[0078] In this embodiment, a user association model is constructed. The user association model is used to model the association relationship between the mobile device MD and the access point AP in each time slot. The construction process is as follows:

[0079] In any time slot MD chooses to associate with different APs based on the distance, channel status, and content deployment status between them. Indicates that in time slot t and MD i Associated AP set, use Indicates MD i With AP j The relationship between them is defined as follows:

[0080]

[0081] To ensure that all MDs can be served, In addition, not all APs are providing content services to MD. The set of all APs in service at time slot t is and

[0082] In this embodiment, a downlink signal model is constructed. The downlink signal model is used to model the network achievable rate of the non-cellular massive MIMO system in each time slot. The construction process is as follows:

[0083] After the MD initiates a content request to its associated AP, if the AP has cached the content, the AP will transmit the content to the corresponding MD via the downlink wireless channel. j To MD i The downlink channel can be expressed as:

[0084] g ij (t) = d ij / d0) -α h ij (t)

[0085] Among them, g ij (t) represents the jth access point AP j To the i-th mobile device MD i Downlink channel; d ij Indicates AP j and MD i The actual distance between them, d0 represents the reference distance, usually d0 = 1m, α is the path attenuation factor, h ij (t) indicates that it obeys the complex Gaussian distribution small-scale fading.

[0086] Assume AP j Transfer to MD i The symbol is q i (t), where represents the mathematical expectation, then AP j The conjugate beamforming transmission signal with complete channel state information is xj (t), expressed as:

[0087]

[0088] in, Indicates the jth access point AP in time slot t j Assigned to the i-th mobile device MD i transmission power; Indicates AP j The maximum transmission power, For AP j and MD i The estimated channel gain at time slot t is, For AP j The MD collection of the service.

[0089] Further MD can be obtained i The received signal model r at time slot t i (t) is as follows:

[0090]

[0091] Among them, w i (t) is MD i The noise signal at time slot t, Indicates the set of APs that do not provide services to MDi among all service APs. Denotes the set of all MDs with service requirements in time slot t.

[0092] According to Shannon's formula, we get MD i The received signal rate is:

[0093]

[0094] in, represents the estimated channel gain; w i It represents the noise signal received by MDi in time slot t.

[0095] The network achievable rate is further obtained as

[0096]

[0097] in, represents the set of all MDs with service requirements in time slot t, R i (t) represents the i-th mobile device MD at time slot t i The received signal rate, R sum (t) represents the network achievable rate of the non-cellular massive MIMO system at time slot t.

[0098] In this embodiment, a system energy consumption model is constructed. The system energy consumption model is used to model the total energy consumption of all access points APs providing services in each time slot. The construction process is as follows:

[0099] APs are usually equipped with limited storage resources and can only store part of the service content, while CPUs have sufficient storage resources and can cache all network content. In CF-mMIMO network scenarios, different APs need to cache different service content according to the requirements of the associated MD. Due to the limitation of AP storage capacity, a single AP cannot cache all content. represents the jth access point AP j The maximum content cache capacity that can be cached; represents the jth access point AP j The set of service contents cached in time slot t; and F represents the number of service contents in the network.

[0100] Assume that the total number of contents in the network is F equal to 100, numbered {1, 2, ..., 100}, and the content set is Indicates that in each time slot t, each AP caches 30 of them, using F j (t) indicates that each user MD generates a request, using f i (t) indicates that the access point AP j The set of all content requests for all associated mobile device MDs is defined as Indicates AP j The MD set of the service. For each time slot t, if Indicates no hit, AP needs to obtain from CPU If the content is missing from the cache, the corresponding energy consumption will be greater than directly obtaining it from the AP. The AP will continuously adjust the cache strategy to improve the hit rate and reduce energy consumption.

[0101] If the AP associated with the MD has cached the service content requested by the MD, the AP directly transmits the cached service content to the MD via the downlink wireless transmission link. If the AP does not have cached the service content requested by the MD, the AP first downloads the corresponding service content from the CPU and then transmits it to the MD via the downlink wireless transmission link.

[0102] During the CF-mMIMO content cache deployment and distribution process, energy consumption mainly includes three parts:

[0103] 1) AP downlink content transmission power consumption:

[0104] The AP provides services to its associated MD. Different APs use different transmission powers for their associated MDs according to their customized power transmission strategies. In time slot t, the AP j Downlink content transmission power consumption Defined as:

[0105]

[0106] in, represents the set of all MDs with service requirements in time slot t, P ij (t) represents the jth access point AP at time slot t j Assigned to the i-th mobile device MD i transmission power;

[0107] 2) Energy consumption of AP service content update or replacement

[0108] To adapt to the dynamic changes of MD content requests, AP should cache, update or replace service content according to demand. j The set of service contents cached, updated or replaced in time slot t is Further AP can be obtained j Energy consumption for updating or replacing service content in time slot t for:

[0109]

[0110] Among them, P FL The energy consumption required to transmit unit content on the forward link (FL); represents the jth access point AP j The set of service contents cached in time slot t; represents the jth access point AP j The set of service contents cached in time slot t-1.

[0111] 3) AP service content forwarding energy consumption

[0112] If the AP does not cache the service content requested by the MD, then the AP needs to obtain the corresponding service content from the CPU and forward it to the MD. In this case, the MD needs to pass through the AP in time slot t. j The collection of content obtained from the CPU is Indicates that in time slot t, the access point AP j The collection of all content requests of all associated mobile device MDs; represents the jth access point AP j The set of service contents cached in time slot t can be used to define AP j The content forwarding energy consumption is:

[0113]

[0114] Among them, σ pen Indicates the energy consumption factor of obtaining uncached content at the AP.

[0115] In summary, we can get AP j The energy consumption in time slot t is:

[0116]

[0117] In summary, the total energy consumption of all serving APs is:

[0118]

[0119] in, Indicates the set of APs providing services to the MD.

[0120] In this embodiment, a target optimization problem model is constructed based on the constructed user association model, downlink signal model, and system energy consumption model. The process is as follows:

[0121] AP storage resources are limited, and content caching, user association, and power allocation are coupled. This paper aims to achieve higher network speeds while maintaining a certain power consumption. The AP content caching, user association, and power allocation issues are modeled as follows:

[0122]

[0123]

[0124]

[0125]

[0126]

[0127]

[0128] in, is the content caching strategy of access point AP in time slot t, represents the jth access point AP in time slot t j The set of service contents cached in time slot t, where N represents the number of access points AP; represents the association strategy between the access point AP and the mobile device MD in time slot t, Indicates time slot With the i-th mobile device MD i Associated access point AP set; The power allocation strategy of the access point AP. For AP j The power allocation set, P ij (t) At time slot t, the jth access point AP j Assigned to the i-th mobile device MD i The transmission power; T represents the number of time slots; represents mathematical expectation; R sum (t) represents the network achievable rate of the non-cellular massive MIMO system at time slot t; P(t) represents the total energy consumption of all serving APs at time slot t; v ij (t) represents the i-th mobile device MD at time slot t i and the jth access point AP j the relationship between represents the jth access point AP j The maximum power value; Indicates AP j The MD collection of the service; represents the jth access point AP j The maximum content cache capacity that can be cached; Represents a collection of mobile devices; Indicates an access point AP set.

[0129] In this embodiment, the target optimization problem model is constructed as a partially observable Markov decision process model, and is solved using a multi-agent reinforcement learning algorithm based on graph attention. The solution process is as follows:

[0130] This paper designs a graph attention multi-agent deep reinforcement learning based joint optimization of content caching, user association and resource allocation for CF-mMIMO (GMJOC). Using access points as agents, the algorithm learns content caching, user association, and power allocation strategies. Taking into account the spatial correlation between different APs, a multi-head attention mechanism is introduced during network parameter aggregation to exploit the correlation between agents.

[0131] The optimization problem P1 is transformed into a partially observable Markov decision process model (Dec-POMDP) ​​with N APs, where each AP represents an agent and is represented by a tuple <S,{O j} j∈N ,{Aj} j∈N ,R,γ to describe, Indicates the CF-mMIMO global network environment status, It is AP j The local observation space of It is AP j The action space is R, R is the reward function, and γ∈[0,1) represents the discount factor. At time slot t, each AP agent receives a local observation and select an action After that, the joint actions of all agents (such as content deployment, user association, resource allocation) can be obtained by using a(t)∈A=A1×…×A N After the agent performs the joint action, the environment returns a global reward r(t) = R(s(t), a(t)) and transfers the state to the next state s(t+1), where is the environment state of CF-mMIMO at time slot t. Next, we define the environment state, local observation, action and reward function.

[0132] 1) State Space

[0133] At time slot t, the environment state includes all AP content caches, channel state information, and user location state information. Therefore, the environment state Defined as

[0134]

[0135]

[0136] l1(t),l2(t),...,l M (t)}

[0137] in, represents the jth access point AP in time slot t j Content cache status; k fj (t) = 1 indicates the jth access point AP j Content f has been cached at time slot t, otherwise k fj (t) = 0; g ij (t) represents the jth access point AP in time slot t j With the i-th mobile device MD i The channel gain between i (t) represents the ith mobile device MD i location information.

[0138] 2) Local observation

[0139] In a partially observable CF-mMIMO environment, any AP can only observe the content cache status of its current time slot t and the user location information in the network. j The local observation at time slot t is expressed as:

[0140]

[0141] 3) Action Space

[0142] According to the optimization problem P1, AP j The variables that can be optimized are content caching, user association and power allocation; based on this, the intelligent AP is defined j Action a at time slot t j (t) is:

[0143]

[0144] in, represents the jth access point AP j The set of service contents cached in time slot t; represents the jth access point AP j The set of mobile devices MD associated at time slot t; Indicates the jth access point AP in time slot t j The power allocation set of

[0145] 4) Reward Function

[0146] All agents perform a joint action a(t) = {a1(t), a2(t), ..., a N (t)}, the environment will return a global reward r(t) to evaluate the joint action. According to the optimization problem P1, the reward function is defined as:

[0147]

[0148] In partially observable environments, the agent AP j Receive the local observation value o j (t), and according to its local strategy π j Select action a j (t). Definition Represents the joint strategy of all agents. The ultimate goal of GMJOC is to learn a strategy for joint optimization of content caching, user association, and resource allocation to maximize the decaying cumulative global reward h represents the time slot set An element in , h=0,1,2.... Therefore, the joint action value function can be defined as:

[0149]

[0150] in, is the expected operation, the action value function Q π (s(t), a(t)) represents the expected decaying cumulative global reward with s(t) as the initial state, π as the initial joint strategy and a(t) as the initial action. The optimal joint strategy π * Is to make Q π (s(t),a(t)) is the largest joint strategy.

[0151] This paper constructs the multi-agent environment as an undirected graph in, is a set of nodes, each node represents an agent AP, ε is a set of edges representing the connectivity between nodes, and AP j There is an edge between it and the nearest AP. In the following, due to this one-to-one correspondence, AP, agent and node all represent the same object. Then, this paper adopts a value decomposition network based on graph attention to decompose the joint action value function into a combination of local action value function and attention-based relationship. The graph attention mechanism can mine the spatial correlation between APs and calculate the weight of the local action value function of each agent. This paper adopts the architecture of centralized training with decentralized execution (CTDE) for learning. In the centralized training stage, any agent learns and updates the network parameters of the agent in combination with the state information of other agents in the network environment. In the execution stage, each AP only needs to select its content caching, user association and resource allocation actions based on its local observation information, without obtaining global state information. The GMJOC algorithm framework proposed in this paper is as follows Figure 2 As shown in Figure 3, it consists of three independent modules: 1) the agent’s local action-value network, 2) the graph attention module, and 3) the hybrid module.

[0152] 1) Local action value network: each agent is equipped with a deep Q-network (DQN) composed of a multilayer perceptron (MLP). Figure 2 As shown. In time slot t, the agent AP j Receive local observation value o j (t), output a local action value function Q j (o j (t),a j (t)).

[0153] 2) Graph attention module, such as Figure 2As shown, the environment state s(t) is first input into the MLP encoder, which encodes s(t) into local potential representation vectors h1(t),h2(t),...,h N (t), where h j (t) is the feature representation of each node. Then, GAT is used to adaptively capture the correlation between agents.

[0154] In CF-mMIMO, any node (such as AP j ) has a set of neighbor nodes determined by the edge set ε If there is an edge between nodes, they are neighbors. In GAT, nodes AP j Its adjacent nodes The attention coefficient between them can be expressed by e j,j =att(W·h j (t),W·h j′ (t)) is calculated, where att(·) is the self-attention mechanism and W is the corresponding learnable weight matrix. Attention coefficient e j,j′ Indicates the neighboring node AP j′ Characteristics of node AP j In this paper, only the nodes The first-order neighbor nodes of (including node j). In order to facilitate comparison between different nodes, the attention coefficients need to be normalized using the softmax function, as shown below:

[0155]

[0156] Among them, α j,j ′ represents the attention coefficient after normalization;

[0157] In order to stabilize the learning process, this paper adopts a multi-head attention mechanism. After obtaining the normalized attention coefficient, the node AP with L independent attention mechanisms j The output feature representation vector h j (t) can be given by:

[0158]

[0159] Among them, σ is a nonlinear function, || represents the concatenation operation, and l represents the sequence number of the attention mechanism. Then, the MLP network is based on h′ j (t) as input and is the agent AP j The local action value function generates weights w j (t).

[0160] 3) Hybrid module. According to the above analysis, the graph attention weight of the hybrid module can be obtained Then the joint action value function Qtot Can be broken down into:

[0161]

[0162] The reinforcement learning model is trained by minimizing the loss function to learn the corresponding action selection strategy π, that is:

[0163]

[0164] Among them, θ represents the parameters of the evaluation network, X represents the number of mini-batch samples randomly sampled from the experience replay pool, x represents the sample number, and y tot =r+γmax a′ Q tot (s′,a′;θ - ), r represents reward, a and a′ represent actions, s and s′ represent environmental states; θ - Represents the parameters of the target network;

[0165] The target optimization problem model is solved by the trained reinforcement learning model to obtain the association strategy between the access point AP and the mobile device MD, the content caching strategy of the access point AP, and the power allocation strategy of the access point AP.

[0166] from Figure 3 It can be seen that as the number of training rounds increases, the average reward of the GMJOC algorithm continues to increase, and finally stabilizes around the 400th round. At this time, the joint action strategy of the AP does not change much, and the reward value obtained is about 4, indicating that the agent is constantly optimizing its own cache strategy, user association strategy, and power allocation strategy. Figure 4 It can be seen that the system network rate is increasing and eventually converges, while Figure 5 The system power consumption in the training process decreases with the increase of training rounds, so the system energy efficiency gradually increases, which proves the effectiveness of the GMJOC algorithm.

[0167] In summary, this invention addresses the dynamic, time-varying network environment and incomplete network state observations by abstracting the aforementioned joint optimization problem into a decentralized partially observable Markov decision process (Dec-POMDP). It then designs autonomous decision-making for content cache deployment, user association, and transmission power control. Taking into account the diverse content caching requirements and wide-area, differentiated network spatial characteristics in CF-mMIMO scenarios, a graph attention network is employed to learn and capture network spatial characteristics, enabling adaptive interference control during content delivery and meeting diverse service requirements.

[0168] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A policy optimization method for a non-cellular massive MIMO system, the system comprising: N access points AP with cache resources and M mobile devices MD, characterized by including: S1: Build a user association model, which is used to model the association relationship between the mobile device MD and the access point AP in each time slot; The user association model includes: and in, Indicates time slot With the i-th mobile device MD i The associated access point AP set, represents a set of time slots; Indicates that the i-th mobile device MD in time slot t i and the jth access point AP j the relationship between Representing a collection The number of access points (APs) in the network; Represents the set of all mobile device MDs, Represents the set of all access points AP; S2: Build a downlink signal model, where the downlink signal model is used to model the network achievable rate of the non-cellular massive MIMO system in each time slot; The downlink signal model includes: g ij (t)=(d ij / d0) -α h ij (t) Among them, R sum (t) represents the network achievable rate of the non-cellular massive MIMO system at time slot t; R i (t) represents the i-th mobile device MD at time slot t i The received signal rate; Indicates the time slot t and the i-th mobile device MD i Associated access point AP set; represents the set of all mobile devices MD that have service requirements in time slot t; v ij (t) represents the i-th mobile device MD at time slot t i and the jth access point AP j The relationship between P ij (t) represents the jth access point AP at time slot t j Assigned to the i-th mobile device MD i transmission power; represents the jth access point AP j Maximum transmission power; g ij (t) represents the jth access point AP j To the i-th mobile device MD i Downlink channel; represents the set of all APs in service at time slot t; Indicates the jth access point AP in time slot t j To the i′th mobile device MD i′ The estimated channel gain P i′j (t) represents the jth access point AP at time slot t j Assigned to the i′th mobile device MD i′ The transmission power of i′j (t) represents the jth access point AP at time slot t j and the i′th mobile device MD i′ The relationship between i (t) represents the i-th mobile device MD at time slot t i Received interference signal; d ij represents the jth access point AP j and the i-th mobile device MD i The actual distance between them; d0 represents the reference distance; α is the path attenuation factor; h ij (t) indicates that it obeys the complex Gaussian distribution Small-scale fading; S3: Build a system energy consumption model, where the system energy consumption model is used to model the total energy consumption of all access points (APs) providing content services in each time slot. The system energy consumption model includes: Where P(t) represents the total energy consumption of all serving APs in time slot t, represents the set of all APs in service at time slot t; P j (t) represents the jth access point AP at time slot t j Total energy consumption; Indicates the jth access point AP in time slot t j Downlink content transmission power consumption; P ij (t) represents the jth access point AP at time slot t j Assigned to the i-th mobile device MD i transmission power; Indicates the jth access point AP in time slot t j The MD collection of the service; Indicates the jth access point AP in time slot t j Energy consumption for updating or replacing service content; P FL The energy consumption required to transmit unit content on the fronthaul link; represents the jth access point AP j The set of service contents cached in time slot t, represents the jth access point AP j The maximum content cache capacity that can be cached; F represents the number of service content types in the network; represents the jth access point AP j The set of service contents cached in time slot t-1; Indicates the jth access point AP in time slot t j Energy consumption of content forwarding; pen represents the jth access point AP j Get the energy consumption factor of an uncached content; Indicates that in time slot t, the access point AP j The collection of all content requests of all associated mobile device MDs; Indicates that the i-th mobile device MD in time slot t i The content of the service requested, Represents a collection of service contents in the network; S4: Construct a target optimization problem model based on the constructed user association model, downlink signal model, and system energy consumption model; The target optimization problem model P1 includes: in, is the content caching strategy of access point AP in time slot t, represents the jth access point AP in time slot t j The set of service contents cached in time slot t, where N represents the number of access points AP; represents the association strategy between the access point AP and the mobile device MD in time slot t, Indicates time slot With the i-th mobile device MD i Associated access point AP set; The power allocation strategy of the access point AP. For AP j The power allocation set, P ij (t) At time slot t, the jth access point AP j Assigned to the i-th mobile device MD i The transmission power; T represents the number of time slots; represents mathematical expectation; R sum (t) represents the network achievable rate of the non-cellular massive MIMO system at time slot t; P(t) represents the total energy consumption of all serving APs at time slot t; v ij (t) represents the i-th mobile device MD at time slot t i and the jth access point AP j the relationship between represents the jth access point AP j The maximum power value; Indicates AP j The MD collection of the service; represents the jth access point AP j The maximum content cache capacity that can be cached; Represents a collection of mobile devices; Represents the access point AP set; S5: The target optimization problem model is constructed as a partially observable Markov decision process model. A multi-agent reinforcement learning algorithm based on graph attention is used to solve it, resulting in the association strategy between access points (APs) and mobile devices (MDs), the content caching strategy of the APs, and the power allocation strategy of the APs. The target optimization problem model is constructed as a partially observable Markov decision process model, which includes: The target optimization problem model is converted into a Dec-POMDP model with N access points AP, where each AP represents an agent and is represented by a tuple Represents the global network environment status of the non-cellular massive MIMO system; represents the jth access point AP j The local observation space of represents the jth access point AP j The action space, R is the reward function, γ∈[0,1) represents the discount factor; At time slot t, the environmental state Defined as: in, represents the jth access point AP in time slot t j Content cache status; k fj (t) = 1 indicates the jth access point AP j Content f has been cached at time slot t, otherwise k fj (t) = 0; g ij (t) represents the jth access point AP in time slot t j With the i-th mobile device MD i The channel gain between i (t) represents the ith mobile device MD i location information; At time slot t, local observation Defined as: At time slot t, the action space Defined as: in, represents the jth access point AP j The set of service contents cached in time slot t; represents the jth access point AP j The set of mobile devices MD associated at time slot t; Indicates the jth access point AP in time slot t j The power allocation set of At time slot t, the reward function r(t)∈R is defined as: Among them, R sum (t) represents the network achievable rate of the non-cellular massive MIMO system at time slot t; P(t) represents the total energy consumption of all serving APs at time slot t; The multi-agent reinforcement learning algorithm based on graph attention is used to solve the problem, including: a local action value network, a graph attention module and a hybrid module; The local action value network configures a deep Q network composed of multi-layer perceptrons for each agent. At time slot t, the agent AP j Receive local observation value o j (t), and choose action a j (t), o j (t) and a j (t) Input deep Q network output local action value Q j (o j (t),a j (t)); The graph attention module first inputs the environment state s(t) into the MLP encoder and encodes s(t) into local potential representation vectors h1(t),h2(t),...,h N (t), where h j (t) represents the jth access point AP j feature representation; GAT is then used to adaptively capture the correlation between agents to obtain the agent AP j The eigenvector of h′ j (t); then the agent AP j The eigenvector of h′ j (t) Input MLP as agent AP j The local action value generates weight w j (t); The hybrid module is based on the local action value Q j (o j (t),a j (t)) and the local action value w j (t) Calculate the joint action value The reinforcement learning model is trained by minimizing the loss function, namely: Among them, θ represents the parameters of the evaluation network, X represents the number of mini-batch samples randomly sampled from the experience replay pool, x represents the sample number, and y tot =r+γmax a′ Q tot (s′,a′;θ - ), r represents reward, a and a′ represent actions, s and s′ represent environmental states; θ - Represents the parameters of the target network; The target optimization problem model is solved by the trained reinforcement learning model to obtain the association strategy between the access point AP and the mobile device MD, the content caching strategy of the access point AP, and the power allocation strategy of the access point AP.