Heterogeneous network multi-spectrum cooperation method and device and medium

By adopting a multi-agent multi-spectral collaboration method based on reinforcement learning in heterogeneous networks, dynamically allocating frequency band resources, the problem of difficulty in collaborating multi-spectral resources in the existing technology is solved, and seamless integration of multiple types of services and high-quality user experience is achieved.

CN119946884AActive Publication Date: 2025-05-06BEIJING UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510109943.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-04
Filing Date
2025-01-23
Publication Date
2025-05-06
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively coordinate multi-spectrum resources in heterogeneous networks, especially when meeting the high requirements for delay, rate, reliability and power consumption of multiple services such as Internet of Vehicles and immersive communications, resulting in poor user experience.

Method used

Using multi-agent multi-spectral collaboration method based on reinforcement learning, multi-agent multi-spectral collaboration method is used to dynamically allocate frequency band resources to optimize user experience through real-time learning of agents, experience playback and target reinforcement learning.

Benefits of technology

It realizes seamless integration of multiple types of services in heterogeneous networks, meets the user experience needs of different services, and improves the overall performance and efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946884A_ABST
    Figure CN119946884A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous network multi-band cooperation method, a heterogeneous network multi-band cooperation device and a medium. The method comprises the following steps: inputting user measurement information into a plurality of agents of a multi-agent multi-band collaborative algorithm, executing a real-time reinforcement learning unit, outputting a Q value to an experience playback unit, and simultaneously triggering a timer 1; if the timer 1 is not stopped, the experience replay unit continues iteration; if the timer 1 stops, outputting the result of the experience replay unit to the target reinforcement learning unit, training the Q value, and triggering the timer 2 at the same time; if the timer 2 stops, the target reinforcement learning unit outputs the obtained Q value to the real-time reinforcement learning unit; if the timer 2 is not stopped, the target reinforcement learning unit continuously iterates; a plurality of agents of the multi-agent multi-spectrum-band collaborative algorithm output a plurality of Q values, and each Q value corresponds to a frequency band resource allocation result of one or more users. The invention provides technical support for various types of services, communication terminals and communication networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of communication technology and artificial intelligence, especially to the fields of mobile communications and Internet of Vehicles services, and more specifically to a heterogeneous network multi-spectrum collaboration method, device and medium. Background Art

[0002] With the continuous changes in wireless communication scenarios, wireless communication services and communication indicators, especially with the further development of artificial intelligence, building an intelligent endogenous wireless network has gradually become a research hotspot. At present, the typical networking mode of mobile communication systems is heterogeneous networking, which is reflected in the different types of terminals and base stations, and the existence of multiple frequency bands, multiple service types and different communication indicators of multiple services in each coverage area. Compared with 4G, heterogeneous network networking is particularly prominent in the 5G era. For multi-band networking, a typical feature is to use low frequency bands for wide coverage and provide basic connections, and to use relatively high frequency bands for capacity expansion and increase transmission rates. For 5G systems, millimeter wave bands (for example, 28GHz, 49GHz bands, etc.) are introduced to perform high-speed transmission based on their large bandwidth characteristics. Future 6G systems, new wifi technologies, etc., also further consider using higher frequency bands and larger bandwidths to maximize the rate of the system.

[0003] Traditional multi-band networking is generally based on relatively low frequency bands, without considering millimeter wave bands or even higher frequency bands. In addition, it does not consider multi-spectrum or multi-band coordination methods under the inherent needs of network intelligence, especially new business types, including Internet of Vehicles business, immersive communication and other needs, which have different requirements for latency, rate, reliability, power consumption and other aspects from previous businesses to ensure the QoE of the business. Summary of the invention

[0004] The present invention is provided to solve the above problems existing in the prior art. Therefore, a multi-spectrum coordination method, device and medium for heterogeneous networks are needed, which are applicable to a heterogeneous system, which includes services with different communication requirements such as latency, rate, reliability, etc., in order to ensure a good user experience and provide network support for various types of services and communication terminals.

[0005] According to a first aspect of the present invention, a heterogeneous network multi-spectrum coordination method is provided, characterized in that the method comprises:

[0006] Acquiring user information, wherein the user information includes channel quality and service characteristics of each user;

[0007] There are multiple pieces of user information, each of which is input into an intelligent agent, which includes a real-time learning unit, an experience replay unit, a target reinforcement learning unit and a neural network unit. When the intelligent agent obtains the user information, it executes the real-time reinforcement learning unit, outputs the Q value to the experience replay unit, and triggers timer 1 at the same time;

[0008] If timer 1 is not stopped, the experience replay unit continues to iterate;

[0009] If timer 1 stops, the result of the experience replay unit is output to the target reinforcement learning unit to train the Q value and trigger timer 2 at the same time;

[0010] If timer 2 stops, the target reinforcement learning unit outputs the obtained Q value to the real-time reinforcement learning unit;

[0011] If timer 2 does not stop, the target reinforcement learning unit continues to iterate;

[0012] Multiple agents output multiple Q values, each Q value corresponds to the frequency band resource allocation result of one or several users;

[0013] Start timer 3;

[0014] If timer 3 is not stopped, the iteration of the multi-agent multi-spectrum coordination step is continued;

[0015] If the timer 3 stops, the frequency band resource allocation result of each user is input into the access control unit to perform spectrum resource allocation and user scheduling for multiple users.

[0016] Furthermore, after obtaining the user information, the method further includes:

[0017] Categorize user information;

[0018] Determine whether it is a new user;

[0019] If it is a new user, determine whether the user information is the same as the existing user information;

[0020] If the user information of the new user is the same as the classified user information, no new agent is allocated to the new user;

[0021] If the user information of the new user is different from one or more types of user information that have been classified, a new type of user information is classified for the new user, and a new agent is assigned.

[0022] Furthermore, the multi-agent multi-spectrum coordination step includes:

[0023] The target SINR list is used as a state set. In the state set, a user is in a state, which means that the traffic of the user's request for the corresponding target SINR is satisfied. Among them, s i Represents the given target SINRβ i , the state space is represented by S = {Θ, β1, β2, ..., β i};

[0024] Defining Action Sets where a i Represents a frequency band, N F is the number of frequency bands, the action space is redefined as:

[0025] A={f1,f2,...,f K}, K≤N F

[0026] Among them, f K represents the frequency band in the redefined action space;

[0027] The reward generation process is as follows:

[0028] At time t, the agent takes action f(t)∈A in state β(t)∈S;

[0029] In this interaction with the environment, the agent receives an immediate reward The system transfers to a new state β(t+1)∈S;

[0030] The immediate reward is expressed as:

[0031]

[0032] Among them, ε represents negative reward, and represents the positive reward of agent i and user i.

[0033] Furthermore, the experience replay unit is based on a fully connected feedforward multi-layer perception (MLP) neural network, and the environmental data is represented as:

[0034] e i (t) = (β i (t),f i (t),R i (t),s i (t+1))

[0035] Among them, e i (t) can be expressed as the environmental data of agent i or terminal i at time t, β i (t) represents the target SINR of agent i or terminal i at time t, Ri (t) represents the reward value of reinforcement learning of agent i or terminal i at time t, f i (t) represents the frequency band allocated to agent i or terminal i at time t, s i (t+1) represents the target SINR of agent i or terminal i at time t+1.

[0036] Furthermore, when the experience replay unit performs experience replay, the memory data of each agent is D i (t) = {e i (1),e i (2),...,e i (t)}.

[0037] Furthermore, the specific steps of implementing experience replay by the experience replay unit include:

[0038] Experience replay memory initialization;

[0039] Initialize θ in the MLP network i , set a random value;

[0040] Initialization in MLP network Make

[0041] During the working period of timer 1,

[0042] At any time t, choose a random action that has the largest reward;

[0043] Update the user's transmit power and the reward from the user's target rate;

[0044] Store the results in the experience replay unit; update θ by gradient descent i ;

[0045] Every N iterations, update Make

[0046] If timer 1 expires, the implementation of the experience replay unit is terminated.

[0047] Furthermore, the neural network unit includes two independent MLP networks, one of which implements the action function to estimate the corresponding Q value Q i (β i ,f i ,θ i ), another MLP network realizes target action estimation Among them, θ i and Represent the current and historical parameters respectively. In the time range corresponding to timer 2, θ i Update using the following formula:

[0048]

[0049] Among them, L(θ i ) represents the cost function, E represents the expectation, R i (β,f) represents the reward function, and Q i represents the Q value, γ represents an iteration coefficient, and max() represents the maximum value function.

[0050] Furthermore, in the real-time reinforcement learning unit and the target reinforcement learning unit, the appropriate frequency band resources are selected by maximizing the sum of the expected discounted rewards. The corresponding objective function is expressed as

[0051]

[0052] in, Indicates that the user's transmission power is not less than zero; represents the positive reward of agent i or user i; r i,t represents the rate of user or agent i at time t; represents user or agent i at time t and frequency band f k The rate of time;

[0053] The transmission power of user or agent i is

[0054]

[0055] Among them, G i represents the channel gain of the user or agent represented by user i;

[0056] The network is optimized by optimizing the behavior of each user type to maximize the user's cumulative future rewards. The future reward is the optimal Q-value function of reinforcement learning, expressed as:

[0057]

[0058] Among them, f i represents action, π represents frequency selection strategy, R i (β i ,f i ) indicates that agent i is in state β i Next, execute action f i The reward, γ t Represents the coefficient at the current time t.

[0059] According to a second aspect of the present invention, a multi-agent multi-band collaboration device is provided, the device comprising:

[0060] A data acquisition module is configured to acquire user information, wherein the user information includes channel quality and service characteristics of each user, and the user information is multiple;

[0061] Agent module, each of the user information is input into an agent, the agent module includes a real-time learning unit, an experience replay unit, a target reinforcement learning unit and a neural network unit, and the agent is configured as follows:

[0062] When user information is obtained, the real-time reinforcement learning unit is executed, the Q value is output to the experience replay unit, and timer 1 is triggered at the same time;

[0063] If timer 1 is not stopped, the experience replay unit continues to iterate;

[0064] If timer 1 stops, the result of the experience replay unit is output to the target reinforcement learning unit to train the Q value and trigger timer 2 at the same time;

[0065] If timer 2 stops, the target reinforcement learning unit outputs the obtained Q value to the real-time reinforcement learning unit;

[0066] If timer 2 does not stop, the target reinforcement learning unit continues to iterate;

[0067] Multiple agents output multiple Q values, each Q value corresponds to the frequency band resource allocation result of one or several users;

[0068] Start timer 3;

[0069] If timer 3 is not stopped, the iteration of the multi-agent multi-spectrum coordination step is continued;

[0070] If the timer 3 stops, the frequency band resource allocation result of each user is input into the access control unit to perform spectrum resource allocation and user scheduling for multiple users.

[0071] According to a third aspect of the present invention, there is provided a readable storage medium storing one or more programs, wherein the one or more programs can be executed by one or more processors to implement the method as described above.

[0072] The present invention has at least the following beneficial effects:

[0073] The present invention provides an algorithm based on reinforcement learning for multi-spectrum collaboration, and at the same time, introduces transfer learning to reduce the complexity of the algorithm. In the algorithm, experience replay is introduced to improve the efficiency of reinforcement learning. The method designed by the present invention provides a lightweight multi-spectrum collaboration for multi-type services and multi-type terminals in heterogeneous network multi-spectrum networking scenarios, which can realize the seamless integration of different types of services and meet the user experience of different types of services. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 A schematic diagram of a scenario of heterogeneous multi-spectrum networking is shown;

[0075] Figure 2 A flowchart of implementing multi-spectrum coordination by a network access point according to an embodiment of the present invention is shown;

[0076] Figure 3 A schematic diagram of the structure of a multi-agent multi-spectrum collaborative algorithm according to an embodiment of the present invention is shown;

[0077] Figure 4 A schematic diagram of experience retransmission according to an embodiment of the present invention;

[0078] Figure 5 A schematic diagram of transfer learning according to an embodiment of the present invention is shown;

[0079] Figure 6 A structural diagram of a multi-type service scheduling device according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0080] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. The embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings and specific embodiments, but are not intended to limit the present invention. For the various steps described herein, if there is no necessity for a causal relationship between each other, the order in which they are described as examples herein should not be regarded as a limitation, and those skilled in the art should know that they can be adjusted in order, as long as the logic between them is not destroyed, resulting in the inability to implement the entire process.

[0081] An embodiment of the present invention provides a heterogeneous network multi-spectrum coordination method, which corresponds to a system. A device that provides wireless data services and service scheduling for mobile terminals or users is called a network access point, which can be a base station in a cell, or a device with multi-spectrum coordination, or a wifi access device. There can be a relative moving speed between the network access point and the user, the network access point can have a moving speed, and the user can have a moving speed. Secondly, the area covered by a network access point is called a coverage area, and a coverage area contains multiple mobile terminals, and the multiple mobile terminals and the network access point are relatively static or moving. Furthermore, the multiple mobile terminals covered by a network access point can change over time, and the number is variable. The wireless communication technology in the present invention is based on the 5G standard, or an evolved version of the 5G standard. The mobile terminal and the network access point in the present invention are implemented based on the above-mentioned wireless communication technology.

[0082] This embodiment provides an application scenario, which includes a network access point and multiple mobile terminals. Specifically, Figure 1 A schematic diagram of a multi-agent multi-spectrum collaborative algorithm system in a heterogeneous network according to an embodiment of the present invention is shown.

[0083] like Figure 1 As shown, the wireless network 140 includes a network access point 100, a first terminal 101, a second terminal 102, and a third terminal 103. The link between the first terminal 101 and the network access point 100 includes a first link 130 and a second link 131, wherein the first link 130 and the second link 131 may be in different frequency bands. The link between the second terminal 102 and the network access point 100 includes a third link 120 and a fourth link 121, wherein the third link 120 and the fourth link 121 are in different frequency bands. The link between the third terminal 103 and the network access point 100 includes a fifth link 110 and a sixth link 111, wherein the fifth link 110 and the sixth link 111 are in different frequency bands.

[0084] exist Figure 1 In the embodiment, the fifth link 110, the third link 120 and the first link 130 are used for basic coverage of the first terminal 101, the second terminal 102 and the third terminal 103, and can also be used for service transmission. The sixth link 111, the fourth link 121 and the second link 131 are used for service transmission of the first terminal 101, the second terminal 102 and the third terminal 103 to improve the service rate, and are not used for network coverage.

[0085] In an example of the present invention, the base station 100 acts as a network access point to provide data services to terminals within a coverage area and perform multi-spectrum coordination.

[0086] like Figure 2A flowchart of a network access point implementing multi-spectrum coordination according to an embodiment of the present invention is shown, which specifically includes the following steps:

[0087] Step S210, the user connects to the network;

[0088] Step S211, measuring and collecting statistics of each user's information, and classifying the user information.

[0089] In this embodiment, the user information includes the channel quality and service characteristics of each user. Specifically, the channel quality can be expressed as the signal to interference and noise ratio (SINR); the service characteristics include the data packet size, rate requirements, quality model based on MOS (mean opinion core) and service type.

[0090] Exemplarily, MOS is expressed as

[0091] M D =alog 10 (br i (1-p e2e ))

[0092] Among them, M D Represents the MOS model of common data services, p e2e represents the end-to-end packet loss rate, a and b are two constants for specific scenarios and services, and r i Represents the rate of user or terminal i.

[0093] In this embodiment, a is set to 1.3619 and b is set to 0.678. It can be understood that the specific values ​​of a and b as above are merely examples and are not limitations of the present invention.

[0094] For example, MOS can also be expressed as

[0095] M V =c(1+exp(d(klog 10 (r i )+ph))) -1

[0096] Among them, M V The MOS model representing the video service, c, d, k, p and h are obtained by fitting the real video data.

[0097] In this embodiment, c is set to 6.6431, d is set to -0.1344, k is set to 10.4, p is set to 28.7221, and h is set to 30.4264. It should be noted that the specific values ​​of the above parameters are only examples and are not intended to limit the present invention.

[0098] Step S212: the classified user information is sent to the multi-agent multi-spectrum collaborative algorithm module.

[0099] It should be noted that the multi-agent multi-spectrum collaborative algorithm module includes multiple agents, which are used to implement the multi-agent multi-spectrum collaborative steps. Each agent corresponds to a number of users, and the several users have the same or similar user information. Each agent includes a real-time reinforcement learning unit, a target reinforcement learning unit, an experience replay unit, and a neural network unit.

[0100] Specifically, after the agent receives the corresponding user information, it executes the real-time reinforcement learning unit, outputs the Q value to the experience replay unit, and triggers timer 1;

[0101] If timer 1 is not stopped, the experience replay unit continues to iterate;

[0102] If timer 1 stops, the result of the experience replay unit is output to the target reinforcement learning unit to train the Q value and trigger timer 2 at the same time;

[0103] If timer 2 stops, the target reinforcement learning unit outputs the obtained Q value to the real-time reinforcement learning unit;

[0104] If timer 2 does not stop, the target reinforcement learning unit continues to iterate;

[0105] The multiple agents of the multi-agent multi-spectrum band collaborative algorithm output multiple Q values, and each Q value corresponds to the frequency band resource allocation result of one or several users.

[0106] In this embodiment, the multi-agent multi-band coordination step is based on reinforcement learning, and the construction of its problem space includes state space, action space and reward space. In the technology of the present invention, resource allocation is performed based on SINR, and the target SINR list is used as a state set. In the state set, a user is in a state, which means that the traffic of the user's request for the corresponding target SINR is satisfied. Therefore, let Among them, s i Represents the given target SINRβ i Since the target SINR is a discrete, finite value, the state set S is also a finite discrete set. The user may not find any frequency band with its target SINR, and the state of this condition is marked as Θ. Further, the state space can be rewritten as S = {Θ, β1, β2, ..., β i}.

[0107] The action is to select a specific frequency band from all possible frequency bands, defining the action set where a i Represents a frequency band, N Fis the number of frequency bands. Note that if the frequency band does not need to be changed, the actions at adjacent times t and t+1 may be the same. Therefore, the action space is redefined as

[0108] A={f1,f2,...,f K},K≤N F

[0109] Among them, f K represents the frequency band in the redefined action space.

[0110] The reward generation process is as follows:

[0111] At time t, the agent takes action f(t)∈A in state β(t)∈S;

[0112] In this interaction with the environment, the agent receives an immediate reward The system moves to a new state β(t+1)∈S.

[0113] According to the user experience model (i.e., the quality model based on MOS (mean opinions core)), MOS is used as a positive reward, indicating

[0114]

[0115] Among them, ε represents negative reward, and represents the positive reward of agent i and user i.

[0116] In some embodiments, an agent corresponds to a type of user information, corresponding to a number of users, r i It also represents the rate of the i-th user, that is, the rate of agent i.

[0117] In some embodiments, multiple agents of the multi-agent multi-spectral band collaborative algorithm output multiple Q values, each Q value corresponds to the frequency band resource allocation result of one or several users, and each Q value corresponds to one or more users represented by a type of user information.

[0118] In some embodiments, in the real-time reinforcement learning unit and the target reinforcement learning unit, after repeated iterations, the optimal strategy is obtained, that is, to select appropriate frequency band resources by maximizing the sum of expected discounted rewards. The corresponding objective function is expressed as

[0119]

[0120] in, Indicates that the user's transmission power is not less than zero. It can be further obtained that the transmission power of user type i is

[0121]

[0122] Among them, G i represents the channel gain of the user or agent represented by user i or agent i. Since MOS is always positive, maximizing the overall MOS and discounted rewards can be achieved by maximizing the individual MOS of each user type.

[0123] The objective function reflects that different types of services share the same quality indicators. Both video services and data services can use MOS as a unified measurement standard. This property enables seamless integration of different types of services. Maximizing the overall network MOS can be achieved by optimizing the behavior of each user type to maximize the user's cumulative future rewards. The future reward is the optimal Q-value function of reinforcement learning, expressed as

[0124]

[0125] Among them, f i represents action, π represents frequency selection strategy, R i (β i ,f i ) indicates that agent i is in state β i Next, execute action f i The reward, γ t Represents the coefficient at the current time t.

[0126] The objective function Reflecting the goal of pursuing high throughput, agent i will choose the available frequency band that provides the highest throughput at time t.

[0127] Step S213, performing multi-user scheduling and resource allocation according to the result of the multi-agent multi-spectrum cooperative algorithm.

[0128] Specifically, after the result of the multi-agent multi-spectrum collaborative algorithm has been obtained, start timer 3;

[0129] If timer 3 is not stopped, the multi-agent multi-spectrum collaborative algorithm iteration continues;

[0130] If the timer 3 stops, the frequency band resource allocation result of each user is input into the access control unit to perform spectrum resource allocation and user scheduling for multiple users.

[0131] Step S214: Count system performance, including user rate and MOS.

[0132] Figure 3The schematic diagram of the multi-agent multi-spectrum collaborative algorithm is shown, specifically including: 300 represents the multi-agent multi-spectrum collaborative algorithm, 301 represents one or more user information input by the system, 302 represents a type of user information to be input, and initializes the real-time reinforcement learning unit for an agent 370. An agent 370 includes a real-time reinforcement learning unit 303 and a target reinforcement learning unit 304. 303 outputs the result to the experience replay unit 305 and starts the timer 341. If 341 stops timing, the obtained result is sent to the target reinforcement learning unit 304, and the timer 342 is started. If 341 has not stopped, the iteration of the reward function in the experience replay unit continues. If the timer 342 stops, the target reinforcement learning unit outputs the result to the real-time reinforcement learning unit. If the timer 342 does not stop, the Q value iteration in the target reinforcement learning unit continues.

[0133] In some embodiments, the experience replay unit is based on a fully connected feedforward multi-layer perception (MLP) neural network, and the environment data is represented as

[0134] e i (t) = (β i (t),f i (t),R i (t),s i (t+1))

[0135] Among them, e i (t) can be expressed as the environmental data of agent i or terminal i at time t, β i (t) represents the target SINR of agent i or terminal i at time t, R i (t) represents the reward value of reinforcement learning of agent i or terminal i at time t, f i (t) represents the frequency band allocated to agent i or terminal i at time t, s i (t+1) represents the target SINR of agent i or terminal i at time t+1;

[0136] When the experience replay unit performs experience replay, the memory data of each agent is D i (t) = {e i (1),e i (2),...,e i (t)}.

[0137] The neural network unit includes two independent MLP networks, one of which implements the action function estimation Q i (β i ,f i ,θ i ), another MLP network realizes target action estimation Among them, θ i and Represent the current and historical parameters respectively. In the time range corresponding to timer 2, θ i Update using the following formula:

[0138]

[0139] Among them, L(θ i ) represents the cost function, E represents the expectation, R i (β,f) represents the reward function, and Q i represents the Q value, γ represents an iteration coefficient, and max() represents the maximum value function.

[0140] Figure 4 A schematic diagram of the experience replay process is shown, including:

[0141] Step S400, initializing the experience replay memory;

[0142] Step S401, initialize θ in the MLP network i , set a random value;

[0143] Step S402, initialization in the MLP network Make

[0144] Step S403, determine whether timer 1 has stopped timing,

[0145] Step S404, if the result of step S403 is no, at any time t, select a random action that has the largest reward;

[0146] Step S405, updating the user's transmission power and updating the reward brought by the user's target rate;

[0147] Step S406, store the result in the experience replay unit; update θ every mini-bath i ;

[0148] Step S407, every N iterations, update Make

[0149] Step S408: If the judgment result of step S403 is yes, end the implementation of the experience replay unit and output the result.

[0150] In some embodiments, Figure 5 The transfer learning process diagram is shown, including:

[0151] Step S500, classifying user information;

[0152] Step S501, determining whether it is a new user;

[0153] Step S502: if the result of step S501 is a new user, determine whether the user information is the same as the existing user information;

[0154] Step S503, if the result of step S502 is that the user information of the new user is different from one or more types of user information that have been classified, classify a new type of user information for the new user and assign an agent of the multi-agent multi-band collaborative algorithm;

[0155] Step S504: if the result of the determination in step S502 is that the user information of the new user is the same as the classified user information, no intelligent agent of the multi-band cooperative algorithm body is allocated to the new user;

[0156] If the result of step S501 is that the user is not a new user, the process returns to step S500.

[0157] The embodiment of the present invention provides a multi-agent multi-band collaboration device, such as Figure 6 As shown, the device 600 includes:

[0158] The data acquisition module 601 is configured to acquire user information, wherein the user information includes channel quality and service characteristics of each user, and the user information is multiple;

[0159] Agent module 602, each of the user information is input into an agent, the agent module includes a real-time learning unit 6021, an experience replay unit 6022, a target reinforcement learning unit 6023 and a neural network unit 6024, and the agent 602 is configured as follows:

[0160] When user information is obtained, the real-time reinforcement learning unit is executed, the Q value is output to the experience replay unit, and timer 1 is triggered at the same time;

[0161] If timer 1 is not stopped, the experience replay unit continues to iterate;

[0162] If timer 1 stops, the result of the experience replay unit is output to the target reinforcement learning unit to train the Q value and trigger timer 2 at the same time;

[0163] If timer 2 stops, the target reinforcement learning unit outputs the obtained Q value to the real-time reinforcement learning unit;

[0164] If timer 2 does not stop, the target reinforcement learning unit continues to iterate;

[0165] Multiple agents output multiple Q values, each Q value corresponds to the frequency band resource allocation result of one or several users;

[0166] Start timer 3;

[0167] If timer 3 is not stopped, the iteration of the multi-agent multi-spectrum coordination step is continued;

[0168] If the timer 3 stops, the frequency band resource allocation result of each user is input into the access control unit to perform spectrum resource allocation and user scheduling for multiple users.

[0169] In some embodiments, the apparatus further comprises a user classification module, wherein the user classification module is configured to:

[0170] Categorize user information;

[0171] Determine whether it is a new user;

[0172] If it is a new user, determine whether the user information is the same as the existing user information;

[0173] If the user information of the new user is the same as the classified user information, no new agent is allocated to the new user;

[0174] If the user information of the new user is different from one or more types of user information that have been classified, a new type of user information is classified for the new user, and a new agent is assigned.

[0175] In some embodiments, the agent module is further configured to:

[0176] The target SINR list is used as a state set. In the state set, a user is in a state, which means that the traffic of the user's request for the corresponding target SINR is satisfied. Among them, s i Represents the given target SINRβ i , the state space is represented by S = {Θ, β1, β2, ..., β i};

[0177] Defining Action Sets where a i Represents a frequency band, N F is the number of frequency bands, the action space is redefined as:

[0178] A={f1,f2,...,f K},K≤N F

[0179] Among them, f K represents the frequency band in the redefined action space;

[0180] The reward generation process is as follows:

[0181] At time t, the agent takes action f(t)∈A in state β(t)∈S;

[0182] In this interaction with the environment, the agent receives an immediate reward The system transfers to a new state β(t+1)∈S;

[0183] The immediate reward is expressed as:

[0184]

[0185] Among them, ε represents negative reward, and represents the positive reward for agent i or user i.

[0186] In some embodiments, the experience replay unit is based on a fully connected feedforward multi-layer perception (MLP) neural network, and the environment data is represented as

[0187] e i (t) = (β i (t),f i (t),R i (t),s i (t+1))

[0188] Among them, e i (t) can be expressed as the environmental data of agent i or terminal i at time t, β i (t) represents the target SINR of agent i or terminal i at time t, R i (t) represents the reward value of reinforcement learning of agent i or terminal i at time t, f i (t) represents the frequency band allocated to agent i or terminal i at time t, s i (t+1) represents the target SINR of agent i or terminal i at time t+1;

[0189] In some embodiments, when the experience replay unit performs experience replay, the memory data of each agent is D i (t) = {e i (1),e i (2),...,e i (t)}.

[0190] In some embodiments, the specific steps of implementing experience replay by the experience replay unit include:

[0191] Experience replay memory initialization;

[0192] Initialize θ in the MLP networki , set a random value;

[0193] Initialization in MLP network Make

[0194] During the working period of timer 1,

[0195] At any time t, choose a random action that has the largest reward;

[0196] Update the user's transmit power and the reward from the user's target rate;

[0197] Store the results in the experience replay unit; update θ by gradient descent i ;

[0198] Every N iterations, update Make

[0199] If timer 1 expires, the implementation of the experience replay unit is terminated.

[0200] In some embodiments, the neural network unit includes two independent MLP networks, one of which implements the action function estimation Q i (β i ,f i ,θ i ), another MLP network realizes target action estimation Among them, θ i and Represent the current and historical parameters respectively. In the time range corresponding to timer 2, θ i Update using the following formula:

[0201]

[0202] Among them, L(θ i ) represents the cost function, E represents the expectation, R i (β,f) represents the reward function, and Q i represents the Q value, γ represents an iteration coefficient, and max() represents the maximum value function.

[0203] In some embodiments, in the real-time reinforcement learning unit and the target reinforcement learning unit, the appropriate frequency band resources are selected by maximizing the sum of the expected discounted rewards. The corresponding objective function is expressed as

[0204]

[0205] in, Indicates that the user's transmission power is not less than zero, and the transmission power of user type i is

[0206]

[0207] Among them, G i represents the channel gain of the user or agent represented by user i;

[0208] The network is optimized by optimizing the behavior of each user type to maximize the user's cumulative future rewards. The future reward is the optimal Q-value function of reinforcement learning, expressed as:

[0209]

[0210] Among them, f i represents action, π represents frequency selection strategy, R i (β i ,f i ) indicates that agent i is in state β i Next, execute action f i The reward, γ t Represents the coefficient at the current time t.

[0211] It should be noted that the device proposed in this embodiment and the method previously described belong to the same technical idea, have the same limited working principle, and can achieve the same beneficial effects, which will not be repeated here.

[0212] An embodiment of the present invention provides a readable storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the multi-type service scheduling method as described in any of the above embodiments.

[0213] In addition, although exemplary embodiments have been described herein, the scope includes any and all embodiments based on the present invention with equivalent elements, modifications, omissions, combinations (e.g., various embodiments intersecting schemes), adaptations or changes. The elements in the claims will be interpreted broadly based on the language adopted in the claims, and are not limited to the examples described in this specification or during the implementation of this application, and the examples will be interpreted as non-exclusive. Therefore, this specification and examples are intended to be considered as examples only, and the true scope and spirit are indicated by the following claims and the full scope of their equivalents.

[0214] The above description is intended to be illustrative rather than restrictive. For example, the above examples (or one or more of them) can be used in combination with each other. For example, those of ordinary skill in the art can use other embodiments when reading the above description. In addition, in the above-mentioned specific embodiments, various features can be grouped together to simplify the present invention. This should not be interpreted as a feature of an invention that is not claimed for protection being necessary for any claim. On the contrary, the subject matter of the present invention may be less than all the features of the embodiments of a specific invention. Thus, the following claims are incorporated into the specific embodiments as examples or embodiments, wherein each claim is independently used as a separate embodiment, and it is considered that these embodiments can be combined with each other in various combinations or arrangements. The scope of the present invention should be determined with reference to the full scope of the equivalent forms of the attached claims and these claims.

Claims

1. A heterogeneous network multi-spectrum collaboration method, characterized in that: The method comprises: Acquiring user information, wherein the user information includes channel quality and service characteristics of each user; There are multiple pieces of user information, each of which is input into an intelligent agent, which includes a real-time learning unit, an experience replay unit, a target reinforcement learning unit and a neural network unit. When the intelligent agent obtains the user information, it executes the real-time reinforcement learning unit, outputs the Q value to the experience replay unit, and triggers timer 1 at the same time; If timer 1 is not stopped, the experience replay unit continues to iterate; If timer 1 stops, the result of the experience replay unit is output to the target reinforcement learning unit to train the Q value and trigger timer 2 at the same time; If timer 2 stops, the target reinforcement learning unit outputs the obtained Q value to the real-time reinforcement learning unit; If timer 2 does not stop, the target reinforcement learning unit continues to iterate; Multiple agents output multiple Q values, each Q value corresponds to the frequency band resource allocation result of one or several users; Start timer 3; If timer 3 is not stopped, the iteration of the multi-agent multi-spectrum coordination step is continued; If the timer 3 stops, the frequency band resource allocation result of each user is input into the access control unit to perform spectrum resource allocation and user scheduling for multiple users.

2. The method according to claim 1, characterized in that After obtaining the user information, the method further includes: Categorize user information; Determine whether it is a new user; If it is a new user, determine whether the user information is the same as the existing user information; If the user information of the new user is the same as the classified user information, no new agent is allocated to the new user; If the user information of the new user is different from one or more types of user information that have been classified, a new type of user information is classified for the new user, and a new agent is assigned.

3. The method according to claim 2, characterized in that The multi-agent multi-spectrum coordination step includes: The target SINR list is used as a state set. In the state set, a user is in a state, which means that the traffic of the user's request for the corresponding target SINR is satisfied. Among them, s i Represents the given target SINRβ i , the state space is represented by S = {Θ, β1, β2, ..., β i }; Defining Action Sets where a i Represents a frequency band, N F is the number of frequency bands, the action space is redefined as: A={f1,f2,...,f K },K≤N F Among them, f K represents the frequency band in the redefined action space; K represents the total number of frequency bands in the newly defined action space; The reward generation process is as follows: At time t, the agent takes action f(t)∈A in state β(t)∈S; In this interaction with the environment, the agent receives an immediate reward The system transfers to a new state β(t+1)∈S; The immediate reward is expressed as: Among them, ε represents negative reward, and represents the positive reward of agent i and user i.

4. The method according to claim 1, characterized in that: The experience replay unit is based on a fully connected feedforward multi-layer perception (MLP) neural network, and the environmental data is represented as: e i (t)=(β i (t),f i (t),R i (t),s i (t+1)) Among them, e i (t) represents the environmental data of agent i or terminal i at time t, β i (t) represents the target SINR of agent i or terminal i at time t, R i (t) represents the reward value of reinforcement learning of agent i or terminal i at time t, f i (t) represents the frequency band allocated to agent i or terminal i at time t, s i (t+1) represents the target SINR of agent i or terminal i at time t+1.

5. The method according to claim 4, characterized in that When the experience replay unit performs experience replay, the memory data of each agent is D i (t) = {e i (1),e i (2),...,e i (t)}.

6. The method according to claim 4, characterized in that The specific steps of implementing experience replay by the experience replay unit include: Experience replay memory initialization; Initialize θ in the MLP network i , set a random value; Initialization in MLP network Make During the working period of timer 1, At any time t, choose a random action that has the largest reward; Update the user's transmit power and the reward from the user's target rate; Store the results in the experience replay unit; update θ by gradient descent i ; Every N iterations, update Make If timer 1 expires, the implementation of the experience replay unit is terminated.

7. The method according to claim 6, characterized in that The neural network unit includes two independent MLP networks, one of which implements the action function estimation Q i (β i ,f i ,θ i ), another MLP network realizes target action estimation Among them, θ i and Represent the current and historical parameters respectively. In the time range corresponding to timer 2, θ i Update using the following formula: Among them, L(θ i ) represents the cost function, E represents the expectation, R i (β,f) represents the reward function, and Q i represents the Q value, γ represents an iteration coefficient, and max() represents the maximum value function.

8. The method according to claim 3, characterized in that In the real-time reinforcement learning unit and the target reinforcement learning unit, the appropriate frequency band resources are selected by maximizing the sum of expected discounted rewards. The corresponding objective function is expressed as in, Indicates that the user's transmission power is not less than zero; represents the positive reward of agent i or user i; r i,t represents the rate of user or agent i at time t; represents user or agent i at time t and frequency band f k The rate of User i transmission power P i for Among them, G i represents the channel gain of the user or agent represented by user i; The network is optimized by optimizing the behavior of each user type to maximize the user's cumulative future rewards. The future reward is the optimal Q-value function of reinforcement learning, expressed as: Among them, f i represents action, π represents frequency selection strategy, R i (β i ,f i ) indicates that agent i is in state β i Next, execute action f i The reward, γ t Represents the coefficient at the current time t.

9. A heterogeneous network multi-spectrum coordination device, characterized in that: The device comprises: A data acquisition module is configured to acquire user information, wherein the user information includes channel quality and service characteristics of each user, and the user information is multiple; Agent module, each of the user information is input into an agent, the agent module includes a real-time learning unit, an experience replay unit, a target reinforcement learning unit and a neural network unit, and the agent is configured as follows: When user information is obtained, the real-time reinforcement learning unit is executed, the Q value is output to the experience replay unit, and timer 1 is triggered at the same time; If timer 1 is not stopped, the experience replay unit continues to iterate; If timer 1 stops, the result of the experience replay unit is output to the target reinforcement learning unit to train the Q value and trigger timer 2 at the same time; If timer 2 stops, the target reinforcement learning unit outputs the obtained Q value to the real-time reinforcement learning unit; If timer 2 does not stop, the target reinforcement learning unit continues to iterate; Multiple agents output multiple Q values, each Q value corresponds to the frequency band resource allocation result of one or several users; Start timer 3; If timer 3 is not stopped, the iteration of the multi-agent multi-spectrum coordination step is continued; If the timer 3 stops, the frequency band resource allocation result of each user is input into the access control unit to perform spectrum resource allocation and user scheduling for multiple users. 10 . A non-transitory computer-readable storage medium storing instructions, which, when executed by a processor, perform the method according to claim 1 .

Citation Information

Patent Citations

  • D2D communication network slice allocation method based on deep reinforcement learning

    CN113163451A

  • Wireless positioning device and method based on multi-band CSI cooperation

    CN113938823A

  • Wireless positioning method and device based on multi-band CSI cooperation

    CN115086864A