Beam switching method and device based on channel prediction
Through channel prediction and reinforcement learning combined with RIS passive relay beam switching method, the problem of unstable beam switching in the Internet of Vehicles is solved, and the saving of spectrum resources and the stability of communication links is achieved.
Patent Information
- Application Number
- CN202310020833.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-06
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-01-06
AI Technical Summary
In the Internet of Vehicles, frequent beam switching causes channel blockage and excessive switching, affecting the stability of the communication link, and how to achieve fast and accurate beam switching to reduce spectrum resource consumption and channel blockage.
A beam switching method based on channel prediction is adopted, combined with reinforcement learning Q-value network and RIS passive relay, by obtaining the state information of the Internet of Vehicles scenario, the actions are output using the ε-greedy method, the actions are executed and stored in the experience pool for training, and the beam switching decision is optimized.
It improves the stability of the communication link, reduces the number of excessive switching times caused by channel blockage, saves spectrum resources, and adapts to the complex mobile environment in the Internet of Vehicles.
Smart Images

Figure CN116209021B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication technology, and in particular to a beam switching method based on channel prediction and a beam switching device based on channel prediction. Background Art
[0002] With the rapid development of wireless communication technology, millimeter wave technology, with its large bandwidth and high frequency characteristics, has gradually become a hot topic of research. However, millimeter waves also have serious propagation losses. With the development of large-scale multiple-input multiple-output (MIMO) antenna technology, millimeter wave MIMO technology can compensate for the propagation losses caused by millimeter waves, fully utilize airspace resources, send directional beams, and achieve higher system capacity and spectrum efficiency.
[0003] In the Internet of Vehicles, considering the limited spectrum resources, telepresence integration technology is widely used in this scenario. This technology integrates perception and communication, and the two share spectrum resources, focusing on providing lower latency and higher reliability for the Internet of Vehicles. Due to the complexity of the Internet of Vehicles scenario, signal transmission is easily blocked by ground obstacles or moving obstacles, which will cause signal interruption. In addition, most vehicles are in a high-speed moving state, which causes the positions of vehicles and base stations to constantly change. Especially in the scenario of multiple base stations and multiple vehicles, high mobility will lead to frequent beam switching, and channel blocking will also affect the frequency of beam switching. How to achieve fast and accurate beam switching and reduce the overhead of beam switching is a major challenge in current research. Summary of the Invention
[0004] The present invention aims to address, at least to some extent, one of the technical problems in the aforementioned technologies. To this end, one objective of the present invention is to propose a beam switching method based on channel prediction that comprehensively considers both communication and perception functions, conserving spectrum resources to a certain extent. Furthermore, by introducing RIS passive relays, this method addresses the issue of channel susceptibility to congestion in connected vehicles (IoVs) and reduces the number of excessive handoffs caused by channel congestion, thereby improving the stability of the communication link.
[0005] The second object of the present invention is to provide a beam switching device based on channel prediction.
[0006] To achieve the above-mentioned purpose, the first embodiment of the present invention proposes a beam switching method based on channel prediction, including obtaining the status information of the current time slot in the vehicle network scenario in each time slot, wherein the status information includes the vehicle receiving signal, channel blocking situation, base station echo signal, communication perception performance, communication rate between the vehicle and the communication base station, and estimated communication rate between the vehicle and other candidate base stations; inputting the status information into the reinforcement learning Q value network to obtain the Q values corresponding to all actions of the reinforcement learning Q value network, and using the ε-greedy method to output the corresponding action in the current Q value, wherein the action includes the switching indication of each vehicle, the transmission beam vector from the base station to the vehicle, and each RIS The phase shift matrix of the device and the power transmitted by each base station to each vehicle; when the action is performed, the state of the current environment will be updated to obtain a new state information as the state information of the next time slot, and a corresponding reward value will be obtained each time an action is performed; the state information of the current time slot, the action, the reward value and the state information of the next time slot are stored in an experience pool, so as to obtain the state information of the current time slot, the action, the reward value and the state information of the next time slot corresponding to each time slot in the experience pool; data is sampled from the experience pool, so as to train the reinforcement learning Q value network according to the sampled data to obtain a trained Q value network, so as to perform beam switching according to the trained Q value network.
[0007] According to the beam switching method based on channel prediction of an embodiment of the present invention, first, in each time slot, the state information of the current time slot in the vehicle network scenario is obtained, wherein the state information includes the vehicle receiving signal, the channel blocking situation, the base station echo signal, the communication perception performance, the communication rate between the vehicle and the communication base station, and the estimated communication rate between the vehicle and other candidate base stations; the state information is input into the reinforcement learning Q value network to obtain the Q values corresponding to all actions of the reinforcement learning Q value network, and the ε-greedy method is used to output the corresponding action in the current Q value, wherein the action includes the switching indication of each vehicle, the transmission beam vector from the base station to the vehicle, the phase shift matrix of each RIS device and the power transmitted by each base station to each vehicle; the action is executed, and the state of the current environment is updated to obtain a new state information to As the state information of the next time slot, and a corresponding reward value will be obtained after each action is performed; the state information, action, reward value of the current time slot and the state information of the next time slot are stored in the experience pool, so that the state information, action, reward value of the current time slot and the state information of the next time slot corresponding to each time slot can be obtained in the experience pool; then, data is sampled from the experience pool to train the reinforcement learning Q-value network according to the sampled data to obtain a trained Q-value network, so that beam switching can be performed according to the trained Q-value network; thus, the two functions of communication and perception are comprehensively considered, spectrum resources are saved to a certain extent, and RIS passive relays are introduced to solve the problem that channels in the Internet of Vehicles are easily blocked, while also reducing the number of excessive switching caused by channel blocking, thereby improving the stability of the communication link.
[0008] In addition, the beam switching method based on channel prediction proposed in the above embodiment of the present invention may also have the following additional technical features:
[0009] Optionally, the Internet of Vehicles scenario includes several base stations, vehicles and RIS devices, the positions of the several base stations and RIS devices are fixed, the vehicles move randomly in the Internet of Vehicles scenario, and the central controller, as an intelligent agent of reinforcement learning, masters the information of all base stations.
[0010] Optionally, at the beginning of each iteration, the reinforcement learning agent is initialized so that the base station randomly forms vehicle power allocation, transmit signal vector, transmit beam matrix and phase shift matrix of RIS device.
[0011] Optionally, within each time slot, the status information of the current time slot in the Internet of Vehicles scenario is obtained, including: the vehicle receiving signal, channel blocking conditions, base station echo signal, communication perception performance, communication rate between the vehicle and the communication base station, and estimated communication rate between the vehicle and other candidate base stations.
[0012] Optionally, if the switching indication of the vehicle is 1, the vehicle undergoes base station switching; if the switching indication of the vehicle is 0, the vehicle does not undergo base station switching.
[0013] Optionally, the state information is updated iteratively in each time slot according to a preset termination condition, and the state information, action, reward value of the current time slot and the state information of the next time slot corresponding to each time slot in the iterative process are stored in the experience pool.
[0014] To achieve the above-mentioned purpose, the second embodiment of the present invention proposes a beam switching device based on channel prediction, including: an acquisition module, used to obtain the status information of the current time slot in the vehicle network scenario in each time slot, wherein the status information includes the vehicle receiving signal, channel blocking situation, base station echo signal, communication perception performance, communication rate between the vehicle and the communication base station, and estimated communication rate between the vehicle and other candidate base stations; a processing module, used to input the status information into the reinforcement learning Q value network to obtain the Q values corresponding to all actions of the reinforcement learning Q value network, and use the ε-greedy method to output the corresponding action in the current Q value, wherein the action includes the switching indication of each vehicle, the transmission beam vector from the base station to the vehicle, and the phase shift of each RIS device. matrix and the power transmitted by each base station to each vehicle; a state update module, used to execute the action, at which time the state of the current environment will be updated to obtain a new state information as the state information of the next time slot, and a corresponding reward value will be obtained each time an action is executed; a storage module, used to store the state information of the current time slot, the action, the reward value and the state information of the next time slot in the experience pool, so as to obtain the state information of the current time slot, the action, the reward value and the state information of the next time slot corresponding to each time slot in the experience pool; a training switching module, used to sample data from the experience pool, so as to train the reinforcement learning Q value network according to the sampled data to obtain a trained Q value network, so as to perform beam switching according to the trained Q value network.
[0015] According to the beam switching device based on channel prediction of an embodiment of the present invention, the acquisition module obtains the status information of the current time slot in the vehicle network scenario in each time slot, wherein the status information includes the vehicle receiving signal, channel blocking situation, base station echo signal, communication perception performance, communication rate between the vehicle and the communication base station, and estimated communication rate between the vehicle and other candidate base stations; the processing module inputs the status information into the reinforcement learning Q value network to obtain the Q values corresponding to all actions of the reinforcement learning Q value network, and uses the ε-greedy method to output the corresponding action in the current Q value, wherein the action includes the switching indication of each vehicle, the transmission beam vector from the base station to the vehicle, the phase shift matrix of each RIS device, and the power transmitted by each base station to each vehicle; the state update module executes the action, and the state of the current environment will be updated to obtain a new state The information is used as the state information of the next time slot, and a corresponding reward value will be obtained after each action is performed; the storage module stores the state information, action, reward value of the current time slot and the state information of the next time slot in the experience pool, so as to obtain the state information, action, reward value of the current time slot and the state information of the next time slot corresponding to each time slot in the experience pool; the training switching module samples data from the experience pool so as to train the reinforcement learning Q-value network according to the sampled data to obtain a trained Q-value network, so as to perform beam switching according to the trained Q-value network; thus, the two functions of communication and perception are comprehensively considered, spectrum resources are saved to a certain extent, and RIS passive relay is introduced to solve the problem that channels in the Internet of Vehicles are easily blocked, and at the same time, the number of excessive switches caused by channel blocking can be reduced, thereby improving the stability of the communication link.
[0016] In addition, the beam switching device based on channel prediction proposed in the above embodiment of the present invention may also have the following additional technical features:
[0017] Optionally, the Internet of Vehicles scenario includes several base stations, vehicles and RIS devices, the positions of the several base stations and RIS devices are fixed, the vehicles move randomly in the Internet of Vehicles scenario, and the central controller, as an intelligent agent of reinforcement learning, masters the information of all base stations.
[0018] Optionally, at the beginning of each iteration, the reinforcement learning agent is initialized so that the base station randomly forms vehicle power allocation, transmit signal vector, transmit beam matrix and phase shift matrix of RIS device.
[0019] Optionally, within each time slot, the status information of the current time slot in the Internet of Vehicles scenario is obtained, including: the vehicle receiving signal, channel blocking conditions, base station echo signal, communication perception performance, communication rate between the vehicle and the communication base station, and estimated communication rate between the vehicle and other candidate base stations. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 1 is a flow chart of a beam switching method based on channel prediction according to an embodiment of the present invention;
[0021] Figure 2 A schematic diagram of a communication scenario of a RIS-assisted Internet of Vehicles according to an embodiment of the present invention;
[0022] Figure 3 FIG. 4 is a block diagram of a beam switching device based on channel prediction according to an embodiment of the present invention. DETAILED DESCRIPTION
[0023] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.
[0024] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0025] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0026] Figure 1 FIG. 1 is a flow chart of a beam switching method based on channel prediction according to an embodiment of the present invention. Figure 1 As shown, the beam switching method based on channel prediction includes the following steps:
[0027] S101, in each time slot, obtain the status information of the current time slot in the Internet of Vehicles scenario, wherein the status information includes the vehicle receiving signal, channel blocking situation, base station echo signal, communication perception performance, communication rate between the vehicle and the communication base station, and estimated communication rate between the vehicle and other candidate base stations.
[0028] As an example, a vehicle networking scenario includes several base stations, vehicles, and RIS devices. The positions of the several base stations and RIS devices are fixed, and the vehicles move randomly in the vehicle networking scenario. The central controller, as an intelligent agent for reinforcement learning, has information about all base stations.
[0029] That is to say, if Figure 2As shown in the figure, there are S base stations, R RIS devices and V vehicles in the Internet of Vehicles scenario, where the positions of the base stations and RIS are fixed. The base stations are evenly placed in the scenario, and the RIS is generally placed at a location where the road is easily congested. The vehicles move randomly in the scenario, and each vehicle is equipped with a single omnidirectional antenna. The data between the base stations is shared, and the central controller serves as the intelligent agent of the reinforcement learning DDQN algorithm. The base station transmits and receives signals through the antenna array. During the signal transmission process, the base station sends an ISAC signal to the vehicle. At the same time, when the direct link is blocked, the RIS device is used as a relay to rebuild the virtual Los link with the help of the RIS device. The base station controls the phase shift matrix of the RIS device. The RIS device performs a certain phase shift on the received signal before transmitting the signal. After receiving the signal, the vehicle feeds back the echo signal to the base station as a perception signal to obtain the vehicle's motion information. The channel matrices between the base station and the vehicle, between the base station and the RIS device, and between the RIS device and the vehicle can be expressed as h sv ,G sr ,h rv .
[0030] As an embodiment, at the beginning of each iteration, the reinforcement learning agent is initialized so that the base station randomly forms the vehicle power allocation, transmit signal vector, transmit beam matrix and phase shift matrix of the RIS device.
[0031] That is, the reinforcement learning agent is initialized, and the base station randomly forms the vehicle power distribution P, P = [p1, p2, ... p v ], the transmitted signal vector x, the transmitted beam matrix w and the phase shift matrix Θ of RIS.
[0032] As an embodiment, within each time slot, the status information of the current time slot in the Internet of Vehicles scenario is obtained, including: vehicle reception signal, channel blocking situation, base station echo signal, communication perception performance, communication rate between the vehicle and the communication base station, and estimated communication rate between the vehicle and other candidate base stations.
[0033] As a specific embodiment, different channels are first modeled. The base station sends a synaesthesia integrated signal to the vehicle. Without considering the RIS relay, the signal received by the v-th vehicle from the s-th base station can be expressed as:
[0034]
[0035] in, It represents the channel matrix between the base station and the vehicle, H represents the transposed conjugate of a matrix, represents the antenna array gain, where N tIt represents the number of antennas at the base station transmitter, p sv It represents the power sent from the sth base station to the vth vehicle, α sv represents the path loss coefficient, represents the beamforming vector of the s-th base station, Indicates the Doppler shift of the signal from the base station to the user end. and They represent the transmission and reception steering vectors between the base station and the vehicle, respectively, x sv represents the transmitted signal, and n sv Represents a noise signal, satisfying In addition, a(θ) can be expressed by the following formula:
[0036]
[0037] Among them, N t and δ represent the number of antennas and the normalized antenna spacing at the transmitting end, respectively.
[0038] Due to the presence of the relay RIS device, the base station can also send signals to vehicles through the RIS device. At this time, the signal received by the v-th vehicle transmitted by the r-th RIS relay can be expressed as:
[0039]
[0040] in, It represents the channel matrix between the base station and RIS. represents the channel matrix between RIS and vehicle, M represents the number of elements of RIS equipment, Indicates the Doppler frequency shift of the signal from the RIS device to the user end. represents the phase shift matrix of the r-th RIS device.
[0041] Combining the above two channel conditions, the received signal at the vehicle end can be expressed as:
[0042]
[0043] Since the channel is blocked, the present invention introduces the blocking coefficient γ under three channels: sv ,γ sr ,γ rv Modeling the channel blocking situation, the signal received by the vehicle can be expressed as follows:
[0044]
[0045] Let γ svr ={γ sv ,γ sr ,γrv}, indicating the congestion situation of the current channel, where γ sv ,γ sr ,γ rv It is expressed as a value from 0 to 1, which represents the channel blocking index. γ∈(0,0.1) means that the channel is completely blocked, γ∈(0.9,1) means that the channel is not blocked, and in other cases it means that the channel is in a semi-blocked state.
[0046] Based on the above signals, the communication rate between the vth vehicle and the sth base station can be obtained as:
[0047]
[0048] After receiving the signal from the base station, the vehicle will feed back an echo signal to the base station. The echo signal of the vth vehicle received by the sth base station can be expressed as:
[0049]
[0050] in, represents the baseband signal in the echo signal, β sv Reflection coefficient, μ sv represents the Doppler shift of the echo signal, τ sv Indicates the delay of the received echo signal, z sv Represents a noise signal.
[0051] Based on the echo signal, the observed time delay can be obtained and Doppler shift The formula can be expressed as:
[0052]
[0053] Then the observed echo signal can be obtained as:
[0054]
[0055] Therefore, the angle perception error of the vth vehicle obtained by the sth base station can be expressed by the following formula:
[0056]
[0057] Among them, θ sv Represents the angular offset of the vehicle relative to the base station.
[0058] In addition, when the base station connected to a vehicle cannot meet the current communication requirements, the communication rate between other base stations and the vehicle is estimated. To represent the estimated communication rate set.
[0059] In summary, the environment state space can be expressed as:
[0060]
[0061] in, y sv It represents the signal received by the vth vehicle from the sth base station, γ svr ={γ sv ,γ sr ,γ rv} respectively represent the channel blocking coefficients between base station s and vehicle v, base station s and RISr, and RISr and vehicle v, C sv It represents the communication rate between the vth vehicle and the sth base station, r sv Represents the echo signal from the vth vehicle to the sth base station, CRLB sv It represents the angle error of the vth vehicle perceived by the sth base station, D ij Denotes the estimated communication rate between the candidate i-th base station and the j-th vehicle.
[0062] S102: Input the state information into the reinforcement learning Q-value network to obtain the Q-values corresponding to all actions of the reinforcement learning Q-value network, and use the ε-greedy method to output the corresponding action in the current Q-value. The action includes the switching instruction of each vehicle, the transmission beam vector from the base station to the vehicle, the phase shift matrix of each RIS device, and the power transmitted by each base station to each vehicle.
[0063] That is to say, the state information is input into the Q network, and the Q value output corresponding to all actions of all Q networks is obtained. The ε-greedy method is used to output the corresponding action a in the current Q value. The action space can be expressed as follows:
[0064]
[0065] Among them, ε v Indicates the switching instruction of the vth vehicle; w sv represents the transmission beam vector sent by the s-th base station to the v-th vehicle, Θ r represents the phase shift matrix of the rth RIS; p sv represents the power transmitted from the s-th base station to the v-th vehicle.
[0066] As an embodiment, if the switching indication of the vehicle is 1, the vehicle undergoes base station switching; if the switching indication of the vehicle is 0, the vehicle does not undergo base station switching.
[0067] That is, when ε v =1, it means that the vth vehicle has a base station switch, and when ε v=0 indicates that no switching occurs.
[0068] S103, executing the action. At this time, the state of the current environment will be updated to obtain a new state information as the state information of the next time slot, and each time an action is executed, a corresponding reward value will be obtained.
[0069] As a specific embodiment, when action a is executed, the agent obtains a new state s' and a reward r. The reward mainly considers the synaesthesia performance and link connectivity. In this invention, the synaesthesia efficiency and connectivity probability are used to calculate the size of the reward. The reward can be expressed by the following formula:
[0070]
[0071] E sv It represents the synaesthesia efficiency of the vth vehicle obtained by the sth base station, and the formula is as follows:
[0072]
[0073] Among them, η represents the trade-off parameter between communication and perception functions.
[0074] β represents the probability connectivity coefficient, is the link connectivity probability between base station s and vehicle v, and the calculation formula is as follows:
[0075]
[0076] Among them, C con Indicates the communication rate threshold when the link is fully connected.
[0077] S104, storing the state information, action, reward value of the current time slot and the state information of the next time slot in the experience pool, so as to obtain the state information, action, reward value and state information of the current time slot corresponding to each time slot in the experience pool.
[0078] As an embodiment, the state information is updated iteratively in each time slot according to a preset termination condition, and the state information, action, reward value of the current time slot and the state information of the next time slot corresponding to each time slot in the iterative process are stored in the experience pool.
[0079] As a specific embodiment, {s, a, r, s', is_end} is stored in the experience pool, the current state s = s' is updated, and the above steps are continued to store the state space, action space, reward value and state space at the next moment in each time slot, where is_ indicates whether the preset termination condition is met.
[0080] S105 , sampling data from the experience pool, so as to train the reinforcement learning Q value network according to the sampled data, so as to obtain a trained Q value network, so as to perform beam switching according to the trained Q value network.
[0081] As an example, sampling is performed from the experience replay pool to calculate the current target Q value:
[0082]
[0083] Where θ and θ' are the parameters of the evaluation Q-network and the target Q-network, respectively.
[0084] Furthermore, using the mean square error loss function Update all parameters θ of the Q network through gradient backpropagation of the neural network.
[0085] After a certain number of iterations, the parameter θ′=θ of the target Q network is updated. If the preset termination condition is met, the current iteration ends; otherwise, the process goes to S102.
[0086] Repeat the above steps until convergence, and output the current action set so that corresponding beam switching can be performed according to the action set.
[0087] In summary, the beam switching method based on channel prediction according to the embodiment of the present invention comprehensively considers both communication and perception functions, saves spectrum resources to a certain extent, and introduces RIS passive relays, which solves the problem of channel susceptibility to blockage in the Internet of Vehicles. At the same time, it can also reduce the number of excessive switching caused by channel blockage; and is no longer limited to the current research that only considers fixed obstacles. Instead, an algorithm is proposed to predict channel blockage conditions, including mobile obstacles, and a reinforcement learning algorithm is used to predict beam switching decisions. Overall, this invention improves the stability of the communication link and alleviates the problem of excessive switching caused by channel interruption to a certain extent.
[0088] In order to implement the above embodiment, Figure 3 As shown, an embodiment of the present invention proposes a beam switching device based on channel prediction, including an acquisition module 10, a processing module 20, a state update module 30, a storage module 40 and a training switching module 50.
[0089] Among them, the acquisition module 10 is used to obtain the state information of the current time slot in the vehicle network scenario in each time slot, wherein the state information includes the vehicle receiving signal, channel blocking situation, base station echo signal, communication perception performance, communication rate between the vehicle and the communication base station, and estimated communication rate between the vehicle and other candidate base stations; the processing module 20 is used to input the state information into the reinforcement learning Q value network to obtain the Q value corresponding to all actions of the reinforcement learning Q value network, and use the ε-greedy method to output the corresponding action in the current Q value, wherein the action includes the switching indication of each vehicle, the transmission beam vector from the base station to the vehicle, the phase shift matrix of each RIS device, and the transmission vector of each base station. to the power of each vehicle; the state update module 30 is used to execute actions to obtain corresponding reward values, and update the state information of the current time slot according to the reward value to obtain the state information of the next time slot; the storage module 40 is used to store the state information, action, reward value of the current time slot and the state information of the next time slot in the experience pool, so as to obtain the state information, action, reward value of the current time slot and the state information of the next time slot corresponding to each time slot in the experience pool; the training switching module 50 is used to sample data from the experience pool, so as to train the reinforcement learning Q-value network according to the sampled data to obtain a trained Q-value network, so as to perform beam switching according to the trained Q-value network.
[0090] It should be noted that the above description and examples of the beam switching method based on channel prediction are also applicable to the beam switching device based on channel prediction in this embodiment, and will not be elaborated here.
[0091] According to the beam switching device based on channel prediction of an embodiment of the present invention, the acquisition module obtains the status information of the current time slot in the vehicle network scenario in each time slot, wherein the status information includes the vehicle receiving signal, channel blocking situation, base station echo signal, communication perception performance, communication rate between the vehicle and the communication base station, and estimated communication rate between the vehicle and other candidate base stations; the processing module inputs the status information into the reinforcement learning Q value network to obtain the Q values corresponding to all actions of the reinforcement learning Q value network, and uses the ε-greedy method to output the corresponding action in the current Q value, wherein the action includes the switching indication of each vehicle, the transmission beam vector from the base station to the vehicle, the phase shift matrix of each RIS device, and the power transmitted by each base station to each vehicle; the state update module executes the action, and the state of the current environment will be updated to obtain a new state The information is used as the state information of the next time slot, and a corresponding reward value will be obtained after each action is performed; the storage module stores the state information, action, reward value of the current time slot and the state information of the next time slot in the experience pool, so as to obtain the state information, action, reward value of the current time slot and the state information of the next time slot corresponding to each time slot in the experience pool; the training switching module samples data from the experience pool so as to train the reinforcement learning Q-value network according to the sampled data to obtain a trained Q-value network, so as to perform beam switching according to the trained Q-value network; thus, the two functions of communication and perception are comprehensively considered, spectrum resources are saved to a certain extent, and RIS passive relay is introduced to solve the problem that channels in the Internet of Vehicles are easily blocked, and at the same time, the number of excessive switches caused by channel blocking can be reduced, thereby improving the stability of the communication link.
[0092] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0093] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0094] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0095] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0096] It should be noted that in the claims, any reference signs placed between parentheses shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claim. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by one and the same item of hardware. The use of the words first, second, third etc. does not indicate any order. These words may be interpreted as names.
[0097] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0098] Obviously, those skilled in the art may make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if such modifications and variations fall within the scope of the claims and their equivalents, the present invention is intended to include such modifications and variations.
[0099] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0100] In the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," "connect," "fixed," etc. should be understood broadly. For example, they may refer to fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; internal communication between two components or interaction between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0101] In the present invention, unless otherwise expressly specified or limited, when a first feature is "above" or "below" a second feature, it may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediary. Furthermore, when a first feature is "above," "above," or "above" a second feature, it may mean that the first feature is directly above or diagonally above the second feature, or simply means that the first feature is at a higher level than the second feature. When a first feature is "below," "below," or "below" a second feature, it may mean that the first feature is directly below or diagonally below the second feature, or simply means that the first feature is at a lower level than the second feature.
[0102] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms should not be understood as necessarily referring to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0103] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are exemplary and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and deform the above embodiments within the scope of the present invention.
Claims
1. A beam switching method based on channel prediction, characterized in that: The following steps are involved: In each time slot, the state information of the current time slot in the vehicle network scenario is obtained, wherein the state information includes the vehicle receiving signal, channel blocking status, base station echo signal, communication perception performance, communication rate between the vehicle and the communication base station, and estimated communication rate between the vehicle and other candidate base stations; Inputting the state information into a reinforcement learning Q-value network to obtain Q-values corresponding to all actions of the reinforcement learning Q-value network, and outputting corresponding actions in the current Q-value using an ε-greedy method, wherein the actions include a handover instruction for each vehicle, a transmit beam vector from the base station to the vehicle, a phase shift matrix for each RIS device, and a power transmitted from each base station to each vehicle; When the action is executed, the state of the current environment will be updated to obtain a new state information as the state information of the next time slot, and each action will be rewarded accordingly. Among them, the reward value is calculated using two indicators: synaesthesia efficiency and connectivity probability. The reward value can be expressed by the following formula: E sv It represents the synaesthesia efficiency of the vth vehicle obtained by the sth base station, and the formula is as follows: Among them, η represents the trade-off parameter between communication and perception, C sv represents the communication rate between the vth vehicle and the sth base station, CRLB sv represents the angle perception error of the vth vehicle obtained by the sth base station; β represents the probability connectivity coefficient, is the link connectivity probability between base station s and vehicle v, and the calculation formula is as follows: Among them, C con The communication rate threshold when the link is fully connected; The state information of the current time slot, the action, the reward value, and the state information of the next time slot are stored in an experience pool, so as to obtain the state information of the current time slot, the action, the reward value, and the state information of the next time slot corresponding to each time slot in the experience pool; Data sampling is performed from the experience pool so as to train the reinforcement learning Q-value network according to the sampled data to obtain a trained Q-value network, so as to perform beam switching according to the trained Q-value network.
2. The beam switching method based on channel prediction according to claim 1, wherein: The IoV scenario includes several base stations, vehicles, and RIS devices. The positions of the several base stations and RIS devices are fixed, and the vehicles move randomly in the IoV scenario. The central controller, as an intelligent agent for reinforcement learning, has information about all base stations.
3. The beam switching method based on channel prediction according to claim 2, wherein: At the beginning of each iteration, the reinforcement learning agent is initialized so that the base station randomly forms the vehicle power allocation, transmit signal vector, transmit beam matrix and phase shift matrix of the RIS device.
4. The beam switching method based on channel prediction according to claim 3, wherein: In each time slot, obtain the status information of the current time slot in the Internet of Vehicles scenario, including: The vehicle receiving signal, channel blocking situation, base station echo signal, communication perception performance, communication rate between the vehicle and the communication base station, and estimated communication rate between the vehicle and other candidate base stations.
5. The beam switching method based on channel prediction according to claim 4, wherein: If the switching indication of the vehicle is 1, the vehicle undergoes base station switching; if the switching indication of the vehicle is 0, the vehicle does not undergo base station switching.
6. The beam switching method based on channel prediction according to claim 5, wherein: The state information is updated and iterated in each time slot according to the preset termination condition, and the state information, action, reward value of the current time slot and the state information of the next time slot corresponding to each time slot in the iterative process are stored in the experience pool.
7. A beam switching device based on channel prediction, characterized in that: include: An acquisition module is configured to acquire, within each time slot, state information of the current time slot in the IoV scenario, wherein the state information includes vehicle reception signals, channel blocking conditions, base station echo signals, communication perception performance, communication rates between the vehicle and the communication base station, and estimated communication rates between the vehicle and other candidate base stations; a processing module, configured to input the state information into a reinforcement learning Q-value network to obtain Q-values corresponding to all actions of the reinforcement learning Q-value network, and output corresponding actions in the current Q-value using an ε-greedy method, wherein the actions include a handover instruction for each vehicle, a transmission beam vector from the base station to the vehicle, a phase shift matrix for each RIS device, and a power transmitted from each base station to each vehicle; The state update module is used to execute the above actions. At this time, the state of the current environment will be updated to obtain a new state information as the state information of the next time slot, and each action will be rewarded after it is completed. Among them, the reward value is calculated using two indicators: synaesthesia efficiency and connectivity probability. The reward value can be expressed by the following formula: E sv It represents the synaesthesia efficiency of the vth vehicle obtained by the sth base station, and the formula is as follows: Among them, η represents the trade-off parameter between communication and perception, C sv represents the communication rate between the vth vehicle and the sth base station, CRLB sv represents the angle perception error of the vth vehicle obtained by the sth base station; β represents the probability connectivity coefficient, is the link connectivity probability between base station s and vehicle v, and the calculation formula is as follows: Among them, C con The communication rate threshold when the link is fully connected; a storage module, configured to store the state information of the current time slot, the action, the reward value, and the state information of the next time slot in an experience pool, so as to obtain the state information of the current time slot, the action, the reward value, and the state information of the next time slot corresponding to each time slot in the experience pool; A training switching module is used to sample data from the experience pool so as to train the reinforcement learning Q value network according to the sampled data to obtain a trained Q value network, so as to perform beam switching according to the trained Q value network.
8. The beam switching device based on channel prediction according to claim 7, characterized in that: The IoV scenario includes several base stations, vehicles, and RIS devices. The positions of the several base stations and RIS devices are fixed, and the vehicles move randomly in the IoV scenario. The central controller, as an intelligent agent for reinforcement learning, has information about all base stations.
9. The beam switching device based on channel prediction according to claim 8, wherein: At the beginning of each iteration, the reinforcement learning agent is initialized so that the base station randomly forms the vehicle power allocation, transmit signal vector, transmit beam matrix and phase shift matrix of the RIS device.
10. The beam switching device based on channel prediction according to claim 9, characterized in that: In each time slot, obtain the status information of the current time slot in the Internet of Vehicles scenario, including: The vehicle receiving signal, channel blocking situation, base station echo signal, communication perception performance, communication rate between the vehicle and the communication base station, and estimated communication rate between the vehicle and other candidate base stations.
Citation Information
Patent Citations
Millimeter wave vehicle networking joint beam allocation and relay selection method
CN113709701A
Beam management method and device
CN115276725A