Two-stage beam tracking methods, apparatus and electronic devices

CN121508579BActive Publication Date: 2026-08-14CHINA MOBILE GROUP DESIGN INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-16
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0003]然而,当前采用滤波算法实现波束追踪的问题在于应用场景单一,仅适用于线性的、高斯场景,但在长时间的追踪中,由于噪声的积累,该算法很容易失效,需要重新进行波束训练;此外,基于深度学习的波束追踪面临数据集获取困难等问题,并且由于实际通信环境可能发生改变,也可能导致数据集失效等问题

Benefits of technology

[0032]如以下将详细描述的,根据本公开实施例的通信与感知一体化场景下的双阶段波束追踪方法、装置、电子设备、计算机可读存储介质以及计算机程序产品,在通信与感知一体化(ISAC)场景下,将波束追踪分为了动态强化学习的波束覆盖和深度强化学习的波束追踪两个阶段。在波束覆盖阶段使用诸如动态Q-learning算法进行马尔可夫决策过程的建模,更好地适应环境的动态变化,以及随着车辆的运行,不同波束提供的增益在不断变化的情况;在波束追踪阶段使用诸如DQN算法进行马尔可夫决策过程建模,利用神经网络解决连续状态问题,可以实现对强化学习算法性能的大幅提升,并将优化目标定为使通信可达速率最大化,对路测单元RSU与车辆之间的通信链路状态进行检测,以使得通信系统获得最大增益。可见,将RSU与车辆之间的感知信息用于辅助波束追踪,通过强化学习和感知信息的结合,提升波束追踪算法的适用性,并且利用强化学习的特殊性,无需获取庞大的训练集,可以通过部分离线、部分线上的方式,在实际测量中不断改善自身策略;同时考虑到动态的场景,利用动态强化学习的思想使波束追踪算法的有效时间变长,降低了波束追踪算法失效的频率,减少通信过程中对波束训练的需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121508579B_ABST
    Figure CN121508579B_ABST
Patent Text Reader

Abstract

This disclosure provides a two-stage beam tracking method, apparatus, and electronic device for integrated communication and sensing scenarios, relating to the field of vehicle-to-everything (V2X) cooperative communication technology. Specifically, in integrated communication and sensing scenarios, beam tracking is divided into two stages: dynamic reinforcement learning-based beam coverage and deep reinforcement learning-based beam tracking. The beam coverage stage better adapts to dynamic environmental changes and the constantly changing gain provided by different beams as the vehicle operates. The beam tracking stage solves the continuous state problem, significantly improving the performance of reinforcement learning algorithms, and setting the optimization objective as maximizing the achievable communication rate. This disclosure, through the aforementioned two-stage beam tracking method, significantly improves tracking accuracy, maintains beam tracking performance over long periods, and can withstand low signal-to-noise ratio communication environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of vehicle-road cooperative communication technology, and in particular to a two-stage beam tracking method, device and electronic device in a communication and sensing integrated scenario. Background Technology

[0002] With the development of vehicle-to-everything (V2X) technology, for the specific needs of vehicle-road cooperation in the communication field, the simplest way to achieve beam tracking in the industry is to select several beams around the main beam to establish a beam set and perform small-scale beam training. For example, some current research methods use tracking algorithms such as Kalman filtering and particle filtering to achieve beam tracking. Based on this, it is possible to estimate state parameters such as angle and velocity, while also incorporating channel gain into the tracking scope to maximize the recovery of channel information, thereby making the tracking data more consistent with the channel scenario. In addition, the industry has also proposed beam tracking methods based on deep learning, which can obtain accurate optimization results through large datasets.

[0003] However, the current method of using filtering algorithms for beam tracking has the problem of limited application scenarios. It is only applicable to linear and Gaussian scenarios. However, during long-term tracking, the algorithm is prone to failure due to the accumulation of noise, requiring retraining of the beam. In addition, deep learning-based beam tracking faces problems such as difficulty in acquiring datasets, and the dataset may also become invalid due to changes in the actual communication environment. Summary of the Invention

[0004] This disclosure is made in view of the above-mentioned problems. This disclosure provides a two-stage beam tracking method, apparatus, and electronic device for integrated communication and sensing scenarios, and correspondingly provides a computer-readable storage medium.

[0005] According to one aspect of this disclosure, a two-stage beam tracking method for an integrated communication and sensing scenario is provided, comprising:

[0006] Based on the preset beam coverage range related to vehicle status, the number of antennas and beam pointing angle for beam coverage in the integrated communication and sensing scenario are determined.

[0007] The beamforming vector in the codebook is set using the number of antennas and the beam pointing angle;

[0008] The first model of the Markov decision process is performed based on the completed codebook, and the dynamic reinforcement learning algorithm is used to solve it to obtain the first beam covering the entire vehicle body.

[0009] Beamforming is performed using all antennas to transmit a second beam with a width smaller than that of the first beam within the coverage area of ​​the first beam.

[0010] A second model of the Markov decision process is performed based on the alignment angle offset between the second beam and the first beam, and a deep reinforcement learning algorithm is used to track the second beam so that the roadside unit and the vehicle's antenna can be docked.

[0011] Furthermore, according to the two-stage beam tracking method in the integrated communication and sensing scenario of one aspect of this disclosure, the method for determining the number of antennas includes: using the beam coverage range set to twice the vehicle length, the beamwidth, and a given beam alignment angle to obtain the number of antennas.

[0012] Furthermore, according to the two-stage beam tracking method in the integrated communication and sensing scenario of one aspect of this disclosure, the beam pointing angle is set in the following ways:

[0013] Based on the obtained number of antennas, two adjacent beams are overlapped so that each beam covers half of the beamwidth, and the included angle between the two adjacent beams is obtained.

[0014] The beam pointing angle set in the codebook is set using the included angle.

[0015] Furthermore, according to the two-stage beam tracking method in the integrated communication and sensing scenario of one aspect of this disclosure, the first modeling method of the Markov decision process includes: setting the set of codeword corresponding indexes in the codebook during the beam coverage stage as the state space of reinforcement learning, setting the reinforcement learning behavior as the addition and subtraction operation of beam indexes, and using the signal-to-noise ratio of the echo signal as the reward mechanism.

[0016] Furthermore, according to the two-stage beam tracking method in the integrated communication and sensing scenario of one aspect of this disclosure, the second modeling method of the Markov decision process includes: setting the number of second beams that can be covered in the codebook during the beam coverage stage and the corresponding index as the behavior space, setting the angular offset between the second beam alignment position and the first beam alignment position as the state space, and using the communication achievable rate as the reward mechanism.

[0017] Furthermore, according to a two-stage beam tracking method in a communication and sensing integrated scenario according to one aspect of this disclosure, the achievable communication rate is determined based on the downlink communication link status between the roadside unit and the vehicle.

[0018] According to another aspect of this disclosure, a two-stage beam tracking device for an integrated communication and sensing scenario is provided, comprising:

[0019] The beam coverage condition configuration module is used to determine the number of antennas and beam pointing angles for beam coverage in the integrated communication and sensing scenario based on the preset beam coverage range related to the vehicle status.

[0020] The codebook setting module is used to set the beamforming vector in the codebook using the number of antennas and the beam pointing angle.

[0021] The beam coverage execution module is used to perform the first modeling of the Markov decision process based on the set codebook, and to solve it using a dynamic reinforcement learning algorithm to obtain the first beam covering the entire vehicle body.

[0022] A fine beamforming module is used to beamform using all antennas and transmit a second beam with a width smaller than the first beam within the coverage area of ​​the first beam.

[0023] The beam tracking execution module is used for the second modeling of the Markov decision process based on the alignment angle offset between the second beam and the first beam, and uses a deep reinforcement learning algorithm to track the second beam so that the roadside unit can dock with the vehicle's antenna.

[0024] Furthermore, according to one aspect of this disclosure, in a two-stage beam tracking device for an integrated communication and sensing scenario, the beam coverage condition configuration module includes an antenna number calculation unit for obtaining the number of antennas using the beam coverage range set to twice the vehicle length, the beamwidth, and a given beam alignment angle.

[0025] Furthermore, according to a dual-stage beam tracking device in a communication and sensing integrated scenario according to one aspect of this disclosure, the beam coverage condition configuration module further includes a beam pointing angle setting unit, which is used to overlap two adjacent beams based on the obtained number of antennas, so that the two beams cover half of the beam width respectively, and obtain the included angle between the two adjacent beams; and use the included angle to set the beam pointing angle set in the codebook.

[0026] Furthermore, according to one aspect of this disclosure, in a two-stage beam tracking device for a communication and sensing integrated scenario, the beam coverage execution module includes a first modeling unit, which is used to set the codeword corresponding index set in the codebook during the beam coverage stage as the state space of reinforcement learning, set the reinforcement learning behavior as an operation of adding or subtracting beam indices, and use the signal-to-noise ratio of the echo signal as a reward mechanism.

[0027] Furthermore, according to one aspect of this disclosure, a two-stage beam tracking device in a communication and sensing integrated scenario includes a beam tracking execution module comprising a second modeling unit, which sets the number of second beams that can be covered in the codebook during the beam coverage stage and their corresponding indices as the behavior space, sets the angular offset between the second beam alignment position and the first beam alignment position as the state space, and uses the achievable communication rate as a reward mechanism.

[0028] Furthermore, according to one aspect of this disclosure, a two-stage beam tracking device in a communication and sensing integrated scenario determines the achievable communication rate based on the downlink communication link status between the roadside unit and the vehicle.

[0029] According to another aspect of this disclosure, an electronic device is provided, comprising: a memory for storing computer-readable instructions; and a processor for executing the computer-readable instructions, causing the electronic device to perform the two-stage beam tracking method in a communication and sensing integrated scenario as described above.

[0030] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing computer-readable instructions that, when executed by a processor, cause the processor to perform the two-stage beam tracking method in the integrated communication and sensing scenario as described above.

[0031] According to yet another aspect of this disclosure, a computer program product is provided, which, when executed by a computer, performs the method described in the first aspect or any possible implementation thereof. In one possible design of the fifth aspect, the relevant program involved in the product may be stored wholly or partially on a memory packaged with a processor, or partially or entirely on a storage medium not packaged with a processor.

[0032] As will be described in detail below, the two-stage beam tracking method, apparatus, electronic device, computer-readable storage medium, and computer program product in the integrated communication and sensing (ISAC) scenario according to embodiments of this disclosure divide beam tracking into two stages: dynamic reinforcement learning beam coverage and deep reinforcement learning beam tracking. In the beam coverage stage, algorithms such as dynamic Q-learning are used to model Markov decision processes, better adapting to dynamic environmental changes and the constantly changing gain provided by different beams as the vehicle operates. In the beam tracking stage, algorithms such as DQN are used to model Markov decision processes, utilizing neural networks to solve continuous state problems, which can significantly improve the performance of reinforcement learning algorithms. The optimization objective is set to maximize the achievable communication rate, and the communication link status between the roadside unit (RSU) and the vehicle is detected to ensure the communication system obtains maximum gain. It is evident that using the perception information between the RSU and the vehicle to assist beam tracking, and combining reinforcement learning with perception information, enhances the applicability of the beam tracking algorithm. Furthermore, leveraging the unique characteristics of reinforcement learning, it eliminates the need for a large training set and can continuously improve its strategy through partial offline and partial online methods in actual measurements. Simultaneously, considering dynamic scenarios, the idea of ​​dynamic reinforcement learning extends the effective time of the beam tracking algorithm, reduces the frequency of beam tracking algorithm failure, and decreases the need for beam training during communication.

[0033] Furthermore, in some preferred embodiments of this disclosure, the transmitted beam codebook is specifically designed to avoid a "sudden drop" in beam coverage, ensuring that the beam always provides good coverage for the vehicle. During the beam tracking phase, the antenna position falls within the beam range selected after solving the beam coverage phase, thus increasing the convergence of the algorithm.

[0034] In summary, this disclosure presents a two-stage beam tracking method based on reinforcement learning in the ISAC scenario, which is more applicable to dynamic scenarios under ISAC and can significantly improve tracking accuracy. At the same time, it can maintain beam tracking performance for a long time and can withstand communication environments with low signal-to-noise ratio. Therefore, it has broad application prospects in communication and sensing integration scenarios and provides a highly adaptable beam tracking solution for rapidly changing communication environments.

[0035] It should be understood that both the foregoing general description and the following detailed description are exemplary and intended to provide further illustration of the claimed technology. Attached Figure Description

[0036] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0037] Figure 1 This is a schematic diagram of the main process of the two-stage beam tracking method in the integrated communication and sensing scenario of this disclosure embodiment.

[0038] Figure 2 This is a schematic flowchart illustrating the process of determining the number of antennas and the beam pointing angle according to an embodiment of this disclosure.

[0039] Figure 3 This is a schematic diagram of the beam antenna directions in the first and second stages of an embodiment of this disclosure.

[0040] Figure 4 This is a flowchart illustrating the beam tracking algorithm of an embodiment of this disclosure.

[0041] Figure 5 This is a functional block diagram of a two-stage beam tracking device in a communication and sensing integrated scenario according to an embodiment of this disclosure.

[0042] Figure 6 This is a hardware block diagram of an electronic device according to an embodiment of the present disclosure.

[0043] Figure 7 This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present disclosure. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of this disclosure more apparent, exemplary embodiments according to this disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments of this disclosure. It should be understood that this disclosure is not limited to the exemplary embodiments described herein.

[0045] First, refer to Figures 1 to 4 This invention describes a two-stage beam tracking method in a communication and sensing integrated scenario according to embodiments of the present disclosure. Figure 1 This is a schematic diagram of the main flow of a two-stage beam tracking method in a communication and sensing integrated scenario according to an embodiment of the present disclosure. Figure 2 This is a further illustration of a flowchart for determining the number of antennas and the beam pointing angle according to an embodiment of the present disclosure. Figure 3 This is a schematic diagram further illustrating the beam antenna directions of the first and second stages according to embodiments of the present disclosure. Figure 4This is a further schematic diagram illustrating the flow of a beam tracking algorithm according to an embodiment of the present disclosure.

[0046] Reference Figures 1 to 4 A detailed description of a two-stage beam tracking method in a communication and sensing integrated scenario according to embodiments of this disclosure includes:

[0047] Step S1: Based on the preset beam coverage range related to vehicle status, determine the number of antennas and beam pointing angle for beam coverage in the integrated communication and sensing scenario.

[0048] Step S2: Using the number of antennas and the beam pointing angle, set the beamforming vector in the codebook;

[0049] Step S3: Based on the completed codebook, perform the first modeling of the Markov decision process and solve it using a dynamic reinforcement learning algorithm to obtain the first beam covering the entire vehicle body.

[0050] Step S4: Use all antennas to perform beamforming, and transmit a second beam with a width smaller than the first beam within the coverage area of ​​the first beam;

[0051] Step S5: Based on the alignment angle offset between the second beam and the first beam, perform a second modeling of the Markov decision process, and use a deep reinforcement learning algorithm to track the second beam so that the roadside unit and the vehicle's antenna can be docked.

[0052] Based on the above embodiments, it can be pointed out that in the present disclosure, the first three steps can be set as the first stage, the main purpose of which is to use a relatively wide ISAC beam for vehicle coverage; the latter two steps can be set as the second stage, the main purpose of which is to perform fine beam tracking based on the first stage.

[0053] Therefore, this specification will elaborate on the two aforementioned stages separately.

[0054] Phase 1: ISAC Beam Coverage

[0055] In the first-stage beam alignment, this disclosure chooses to utilize a dynamic reinforcement learning algorithm to implement ISAC beam tracking. When constructing the reinforcement learning problem, the scenario can be modeled as a Markov decision process. In the first stage, the primary objective of beam tracking is to ensure that the roadside unit's beam covers the vehicle, rather than maximizing communication quality. Simultaneously, applying different algorithms with different optimization objectives at different stages can address issues such as the signal-to-noise ratio exposed in the hierarchical codebook, thereby maintaining good performance across all stages.

[0056] Specifically, the purpose of beam selection in the first stage is to ensure that the beam transmitted by the roadside unit covers the entire vehicle body. Therefore, in the beamforming of the first stage, only some antennas participate to obtain a larger beamwidth. Since beam coverage determines the performance of beam tracking in the subsequent second stage, and in order to reduce the dimensionality of the state space of reinforcement learning to ensure fast algorithm convergence, this disclosure optimizes the transmission beam design in the first stage in some embodiments. The design logic is described below.

[0057] Beam tracking occurs after beam training. Through beam training and subsequent communication, the ISAC RSU can obtain vehicle status information, such as vehicle length and distance between the RSU and the vehicle. Furthermore, considering the premise of ensuring a beam sufficient to cover the entire vehicle at all times, and minimizing beamwidth to achieve a higher echo signal-to-noise ratio, beams at adjacent angles overlap by the length of the vehicle body.

[0058] Therefore, a general concept for calculating the number of antennas and the beam pointing angle is presented, such as... Figure 2 As shown, it includes:

[0059] Step S11: Using the beam coverage area, beamwidth, and given beam alignment angle set to twice the vehicle length, obtain the number of antennas;

[0060] Step S12: Based on the obtained number of antennas, two adjacent beams are overlapped so that each beam covers half of the beamwidth, and the included angle between the two adjacent beams is obtained.

[0061] Step S13: Use the included angle to set the beam pointing angle set in the codebook.

[0062] Specifically, according to the definition of beamwidth, when the RSU expects an alignment angle θ k At that time, the beamwidth of the array is:

[0063]

[0064] Where k0 is the beamwidth factor, c is the speed of light, and N t,k For pointing to θ k The corresponding number of antennas is used to control the beamwidth, r0 is the distance between two adjacent antennas in the RSU transmit antenna array, which is set to 1 / 2λ in this disclosure, f c Let be the carrier frequency. Equation (1) gives the transmit beamwidth. With the number of antennas N t,k and the beam alignment angle θ kThe relationship is that, given the pointing angle and the number of antennas, the beamwidth is fixed. By considering the beamwidth and distance, the coverage area of ​​the beam can be determined for a given number of antennas used for beamforming.

[0065]

[0066] Equation (2) indicates that the beam coverage range is related to the number of beamforming antennas N. t,k and alignment angle θ k The relationship between beam tracking and beam quality is complex. In beam tracking problems, a "dropout" phenomenon is prone to occur. As the beam used in the previous stage gradually loses its gain over time, the communication quality of the channel degrades as it moves further away from the beamwidth coverage area. Specifically, after reaching the 3dB bandwidth, the beam gain drops rapidly, leading to brief periods of poor communication quality, with coverage failures occurring in areas not covered by the 3dB bandwidth. Therefore, the RSU aims to transmit a beam sufficient to cover the entire vehicle body at any given time. Considering the above issues, in the first stage, the beam must cover twice the length of the vehicle body to achieve full beam coverage. Therefore, assuming the vehicle body length is x... v Width is y v The longest state of the vehicle is Therefore, in this embodiment, Δd is set... k =2z v Given θ k and Δd k In the case of ideal conditions, the number of antennas can be obtained by using equations (1) and (2).

[0067]

[0068] Obviously, the number of antennas calculated in equation (3) is not necessarily an integer, and considering the upper limit of the number of antennas equipped on the RSU, the number of antennas can be expressed as:

[0069]

[0070] The RSU can perform beamforming by calculating the number of antennas using equation (4), thereby enabling the transmission of a beam that covers a range twice the length of the vehicle body.

[0071] Thus, by overlapping two adjacent beams, each covering half of the beamwidth, the angle between the two adjacent beams is determined. This ensures that at any given time, only one beam completely covers the entire vehicle, while minimizing the beamwidth to achieve sufficient beam gain.

[0072] Based on this, the set of beam pointing angles in the codebook is set as follows:

[0073]

[0074] Where θ0 and θ K Let θ0 and θ2 be the minimum and maximum angles that the RSU can serve, respectively, where θ0 < θ2. K The codebook setting in equation (5), combined with the antenna setting in equation (4), can realize the beamforming vector setting for the first-stage beam coverage. The set of indexes corresponding to the codewords in the codebook is set as the state space in the first-stage beam coverage enhancement algorithm. In the first phase, this disclosure sets the behavior of reinforcement learning as an increment / decrement operation on the beam index. Therefore, the behavior space for the first phase is: Since the goal of beam coverage in the first stage is to capture the vehicle as much as possible, the signal-to-noise ratio of the ISAC echo signal can be calculated given the beam size:

[0075]

[0076] Therefore, the signal-to-noise ratio of the received echo is used as the reward R1 = SNR(f) for the first stage of beam coverage. n Thus, the modeling of the first-stage beam coverage Markov decision process has been achieved.

[0077] For solving the first-stage Markov decision process, this disclosure chooses to use a dynamic Q-learning algorithm. Because of the use of coarse beams, the state space is not large during the beam coverage stage, and the fast acquisition in the first stage facilitates fast beam selection in the second stage. Dynamic reinforcement learning algorithms take into account the dynamic nature of the environment and are applicable in dynamic environments compared to static reinforcement learning algorithms. Since the state in the first stage is set as the beam in this disclosure, and the gain provided by different beams changes continuously as the vehicle moves, this disclosure also considers static reinforcement learning algorithms unsuitable for deployment.

[0078] Finally, it can be added that the first-stage ISAC beam coverage algorithm, in practical operation, can be referred to Table 1 below. The input stage of the algorithm models environmental changes. In the environmental changes, M... ni With n i The elements are selected from a set without specifying their order. This representation is intended to reflect that environmental changes are random and unpredictable. In the second-stage beam tracking described in step 10, since there are no longer any communication beams within the current beam coverage area that can guarantee communication quality, the first-stage beam coverage can no longer completely capture the vehicle. The algorithm considers that the environment has been updated and therefore changes the environment flag, meaning that the environmental changes will be considered in step 9 of the Q-function update. Meanwhile, if This means the vehicle is about to leave the service range of the current beam, but the algorithm is still in the previous stage, proving that an environmental update is needed. Here, 0 < ε < 1 represents an adjustment factor, indicating the algorithm's sensitivity to environmental updates. When ε is close to 1, it shows the algorithm is relatively optimistic about beam selection, only switching at critical points; when ε decreases, it indicates the algorithm considers its update speed to be slower than environmental changes, requiring more frequent environmental updates.

[0079] The design of 'i' in step 5 of the algorithm takes into account that although the system environment is changing, there may be a period of relative calm during which environmental changes are slow and insufficient to affect the algorithm. Compared to the traditional Q-learning algorithm, the dynamic reinforcement learning algorithm takes environmental changes into account and sets criteria for algorithm updates, thus making it suitable for Markov decision processes with changing environments.

[0080] Table 1 ISAC Beam Coverage Algorithm

[0081]

[0082]

[0083] Phase Two: Beam Tracking

[0084] After successfully achieving beam coverage in the first phase, the RSU captures the vehicle using a relatively wide beam. However, at this point, the beam gain is limited due to the relatively small number of antennas introduced. Systems equipped with massive MIMO prefer to use all antennas for communication to obtain maximum beam gain. Therefore, the second phase of the beam tracking process begins.

[0085] In the second stage of beam tracking, the purpose of beam tracking changes. To maximize the communication signal-to-noise ratio, all transmit antennas of the RSU are used for beamforming, resulting in a very narrow communication beam and maximum beam gain.

[0086] The second-stage beam tracking task involves docking the RSU with the vehicle's antenna array and directing the RSU's transmit beam towards the antenna on the vehicle. Therefore, in the second stage, this disclosure no longer utilizes the ISAC waveform; instead, the RSU transmits communication signals to achieve maximum communication performance. Since the ISAC beam of the first stage covers the entire vehicle body, the subsequent second-stage communication beam will fall within the coverage area of ​​the first-stage ISAC beam.

[0087] The Markov decision process modeling in the second stage also differs from that in the first stage. By estimating the physical state of the echoes after successful coverage in the first stage, the vehicle's center position θ can be obtained. vIn the first stage, the beam is transmitted according to the direction in codebook F1. Therefore, there is an angular offset Δθ between the vehicle's center position and the position where the beam is aligned in the first stage. v This disclosure utilizes angular offset Δθ v The second stage of beam tracking has several advantages. Firstly, due to the beam coverage design, the antenna positions in the second stage will fall within the beam range selected in the first stage. This means the beam angle in the second stage is limited to the beam coverage angle of the first stage, significantly reducing the set of possible beam selections. This lowers the dimensionality of the state and behavior space in reinforcement learning problems, facilitating the operation and convergence of the reinforcement learning algorithm. Secondly, the distance the vehicle travels within the beam coverage area in the first stage is much shorter than the distance outside the beam coverage area. This means that environmental changes are always controllable in the second stage. In the second stage, to ensure communication quality, all antennas are used for beamforming, resulting in an extremely narrow beam. The constraints imposed in the first stage significantly reduce the algorithm's complexity, effectively improving algorithm efficiency, beam tracking time, and convergence.

[0088] In the second stage, the behavior space is set to the number of fine beams that can be covered in the codebook of beam coverage in the first stage and their corresponding indices. The DFT codebook formed by Nt and the antenna can divide the angle space into 2π / N. t Equal division, and the coverage angle in the first stage The number of codebooks that can be covered is:

[0089]

[0090] The behavioral space for the second stage is then: It is important to note that the elements in the behavior space do not represent the index of the transmitted beam, but rather the sequence number of the fine beam under beam coverage in the first stage of beam coverage. For example... Figure 3 As shown, the codeword in the second stage falls within the effective width of the beam covered by the beam in the first stage.

[0091] Because the number of beams used by the system is more limited in the second-stage beam tracking, it is preferable to directly select the optimal beam, rather than attempting to obtain information about environmental changes by switching beams between different indices, as is done with beam coverage. Setting the behavior space as beam indexing allows the system to directly point to the antenna position. Furthermore, in the second stage, the state is set as an angle rather than a beam, which transforms the environment from beam gain to a one-dimensional angle, making the environment less sensitive to changes in the system than in the first-stage beam coverage. Therefore, the state in the second stage is set as the angular offset Δθ between the beam alignment position in the second stage and the beam alignment position in the first stage. v .

[0092] Obviously, Δθ v Since the variable is continuous, the dimension of its state space is infinitely large. In this case, the dimension of the Q-table in Q-learning also becomes infinitely large, causing the algorithm to never be able to access the Q-table, thus Q-learning cannot solve problems under continuous states. This disclosure chooses to use the DQN algorithm to solve the beam tracking problem in the second stage. DQN, by utilizing a neural network instead of the Q-table, can significantly improve the performance of reinforcement learning algorithms. The continuous state problem that Q-learning cannot solve can be addressed through the approximate approximation capability of neural networks, enabling behavioral selection under continuous states. Therefore, the state space of the second stage is:

[0093] In the second-stage beam tracking, the optimization objective becomes maximizing the achievable communication rate. For this, the ISAC echo signal-to-noise ratio measurement from the first stage is no longer applicable. Based on vehicle modeling and target expansion considerations, the information provided by the ISAC echo is always imperfect; therefore, the second stage requires detecting the communication link status between the RSU and the vehicle, rather than utilizing the echo signal-to-noise ratio. Thus, in the second stage, the downlink communication signal received by the vehicle can be represented as:

[0094]

[0095] in, For RSU transmit antenna gain, The channel coefficients representing the LOS path are as follows: α represents the path loss of the LOS path. ref This represents the path loss per unit length. Therefore, obtaining the propagation path implies the LOS path channel coefficient α. n It is known. Therefore, the signal-to-noise ratio (SNR) of communication can be expressed as:

[0096]

[0097] in This represents the beamforming gain factor. When the beam is perfectly aligned... otherwise From equation (9), it can be seen that the achievable communication rate can be expressed as:

[0098]

[0099] Therefore, the reward for the second-stage beam tracking is set as the communication achievable rate R2 = R n This is to maximize the gain of the communication system. At this point, the second phase of Markov decision process modeling is complete.

[0100] Similarly, the second-stage ISAC beam tracking algorithm can be referenced in Table 2 below. In step 1, the DQN algorithm is initialized. Since the prerequisite for algorithm 2 to run is that algorithm 1 has successfully captured the vehicle, the system already possesses the successfully selected ISAC beam f from algorithm 1 during the initialization phase. n This also defines the environment in which Algorithm 2 is implemented. Steps 6-8 involve selecting a... i This implements the selection of the transmit beam, the transmission of the downlink communication beam, and the state S obtained from the transmit beam in Algorithm 1. i+1 And obtain feedback r through channel state detection i The parameter update of DQN is implemented in steps 10-14. Here, ∈-greedy is used in both Algorithm 1 and Algorithm 2; the greedy algorithm aims to balance the exploration of the environment with achieving the optimal effect in the current state. In reinforcement learning, the system allows the agent to explore, even if this leads to beam tracking failure. However, due to the large beam coverage in Algorithm 1 and the fine beam tracking design in Algorithm 2, the system's detection range is greatly reduced, and if the overhead of the large beam coverage in the first stage is ignored, very few beams are used for each beam tracking attempt. After some exploration, the system can develop an overall understanding of the environment covered by the ISAC beam, which makes beam selection more stable. Even if beam tracking fails, the cost of exhaustive beam training is greatly reduced because the first-stage beam coverage limits the range of codebook selection in the second stage.

[0101] Table 2 ISAC Beam Tracking Algorithm

[0102]

[0103]

[0104] Specifically, the process of the beam tracking algorithm can be as follows: Figure 4 As shown. DQN approximates the Q-value using a Deep Neural Network (DNN). Compared to Q-learning, DQN's behavior value function is q(s,a;η), which incorporates the influence of the weights η. For the purpose of training the DNN, the DQN algorithm first initializes the Q-function using random weights η0. Then, in the k-th iteration, the approximation of the Q-value is updated towards the target:

[0105]

[0106] The weight η in equation (11) k Then, it is updated using the stochastic gradient descent method. η is updated by minimizing the mean squared error.k :

[0107]

[0108] Ultimately, the weights are updated:

[0109]

[0110] When the weights η are updated, the objective function It is also changing, and this change makes training unstable. Therefore, experience replay and having the target network calculate the target Q-value are introduced to solve this problem. First, the target q-function in equation (13) is replaced with in The update is only performed every C iterations of the algorithm.

[0111] The core idea of ​​experience replay is to allow the agent to store its past experiences, i.e., state transitions, and use them when training the DNN. This process allows the agent to randomly sample independent mini-batches of data, thereby improving the network's learning ability.

[0112] The above describes a two-stage beam tracking method in an integrated communication and sensing scenario according to embodiments of the present disclosure. The following will further describe a two-stage beam tracking apparatus in an integrated communication and sensing scenario for implementing the above control method. Figure 5 This is a functional block diagram of a two-stage beam tracking device in a communication and sensing integrated scenario according to an embodiment of the present disclosure.

[0113] like Figure 5 As shown, the dual-stage beam tracking device 500 in a communication and sensing integrated scenario according to an embodiment of this disclosure includes:

[0114] The beam coverage condition configuration module 501 is used to determine the number of antennas and beam pointing angles for beam coverage in the integrated communication and sensing scenario based on the preset beam coverage range related to the vehicle status.

[0115] The codebook setting module 502 is used to set the beamforming vector in the codebook using the number of antennas and the beam pointing angle.

[0116] The beam coverage execution module 503 is used to perform the first modeling of the Markov decision process based on the set codebook, and to solve it using a dynamic reinforcement learning algorithm to obtain the first beam covering the entire vehicle body.

[0117] The fine beamforming module 504 is used to perform beamforming using all antennas and transmit a second beam with a width smaller than that of the first beam within the coverage area of ​​the first beam.

[0118] The beam tracking execution module 505 is used to perform a second modeling of the Markov decision process based on the alignment angle offset between the second beam and the first beam, and to use a deep reinforcement learning algorithm to track the second beam so that the roadside unit and the vehicle's antenna can be docked.

[0119] Furthermore, the beam coverage condition configuration module includes an antenna quantity calculation unit, used to obtain the number of antennas using the beam coverage range set to twice the vehicle length, the beamwidth, and a given beam alignment angle.

[0120] Furthermore, the beam coverage condition configuration module also includes a beam pointing angle setting unit, which is used to overlap two adjacent beams based on the obtained number of antennas, so that the two beams cover half of the beam width respectively, and obtain the included angle between the two adjacent beams; and use the included angle to set the beam pointing angle set in the codebook.

[0121] Furthermore, the beam coverage execution module includes a first modeling unit, which is used to set the set of codeword corresponding indexes in the codebook during the beam coverage stage as the state space of reinforcement learning, set the reinforcement learning behavior as adding or subtracting beam indices, and use the signal-to-noise ratio of the echo signal as the reward mechanism.

[0122] Furthermore, the beam tracking execution module includes a second modeling unit, which is used to set the number of second beams that can be covered in the codebook during the beam coverage phase and their corresponding indices as the behavior space, set the angular offset between the second beam alignment position and the first beam alignment position as the state space, and use the achievable communication rate as the reward mechanism.

[0123] Furthermore, the achievable communication rate is determined based on the downlink communication link status between the roadside unit and the vehicle.

[0124] Figure 6 This is a hardware block diagram illustrating an electronic device 600 according to an embodiment of the present disclosure. The electronic device according to an embodiment of the present disclosure includes at least a processor and a memory for storing computer-readable instructions. When the computer-readable instructions are loaded and executed by the processor, the processor performs the two-stage beam tracking method in an integrated communication and sensing scenario as described above.

[0125] Figure 6The illustrated electronic device 600 specifically includes a central processing unit (CPU) 601, a graphics processing unit (GPU) 602, and a memory 603. These units are interconnected via a bus 604. The CPU 601 and / or GPU 602 can function as the aforementioned processor, and the memory 603 can function as the aforementioned memory storing computer-readable instructions. Furthermore, the electronic device 600 may also include a communication unit 605, a storage unit 606, an output unit 607, an input unit 608, and an external device 609, all of which are also connected to the bus 604.

[0126] Figure 7 This is a schematic diagram illustrating a computer-readable storage medium according to an embodiment of the present disclosure. Figure 7 As shown, a computer-readable storage medium 700 according to an embodiment of the present disclosure stores computer-readable instructions 701 thereon. When the computer-readable instructions 701 are executed by a processor, the two-stage beam tracking method in a communication and sensing integrated scenario according to an embodiment of the present disclosure, as described with reference to the above figures, is performed. The computer-readable storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, optical disk, magnetic disk, etc.

[0127] Finally, it can be added that, according to an embodiment of this disclosure, a computer program product (which may include the aforementioned apparatus) runs on a terminal device, causing the terminal device to execute the two-stage beam tracking method in the communication and sensing integrated scenario described in the foregoing embodiments or equivalent implementations. As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that all or part of the steps in the above implementation methods can be implemented using software plus the necessary general-purpose hardware platform. Based on this understanding, the aforementioned computer program product may include, but is not limited to, an App.

[0128] The above description, with reference to the accompanying drawings, outlines a two-stage beam tracking method, apparatus, electronic device, computer-readable storage medium, and computer program product in an integrated communication and sensing scenario according to embodiments of the present disclosure. The main concept is that, in an integrated communication and sensing (ISAC) scenario, beam tracking is divided into two stages: dynamic reinforcement learning-based beam coverage and deep reinforcement learning-based beam tracking. In the beam coverage stage, a dynamic Q-learning algorithm is used to model the Markov decision process, better adapting to dynamic environmental changes and the constantly changing gain provided by different beams as the vehicle operates. In the beam tracking stage, a DQN algorithm is used to model the Markov decision process, utilizing neural networks to solve continuous state problems, which significantly improves the performance of reinforcement learning algorithms. The optimization objective is set to maximize the achievable communication rate, and the communication link status between the roadside unit (RSU) and the vehicle is detected to ensure the communication system obtains maximum gain. It is evident that using the perception information between the RSU and the vehicle to assist beam tracking, and combining reinforcement learning with perception information, enhances the applicability of the beam tracking algorithm. Furthermore, leveraging the unique characteristics of reinforcement learning, it eliminates the need for a large training set and can continuously improve its strategy through partial offline and partial online methods in actual measurements. Simultaneously, considering dynamic scenarios, the idea of ​​dynamic reinforcement learning extends the effective time of the beam tracking algorithm, reduces the frequency of beam tracking algorithm failure, and decreases the need for beam training during communication.

[0129] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0130] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0131] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0132] Additionally, as used herein, the “or” used in a list of items beginning with “at least one” indicates a separate list, such that a list of, for example, “at least one of A, B, or C” means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word “exemplary” does not imply that the described example is preferred or better than other examples.

[0133] It should also be noted that in the systems and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.

[0134] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.

[0135] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0136] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A two-stage beam tracking method for an integrated communication and sensing scenario, characterized in that, The tracking method includes: Based on the preset beam coverage range related to vehicle status, the number of antennas and beam pointing angle for beam coverage in the integrated communication and sensing scenario are determined. The beamforming vector in the codebook is set using the number of antennas and the beam pointing angle; The first model of the Markov decision process is performed based on the completed codebook, and the solution is obtained by using a dynamic reinforcement learning algorithm to obtain the first beam covering the entire vehicle body; wherein, the dynamic reinforcement learning algorithm changes the environmental signs when the environment changes, and updates the Q function based on the changed environmental signs; Beamforming is performed using all antennas to transmit a second beam with a width smaller than that of the first beam within the coverage area of ​​the first beam. A second model of the Markov decision process is performed based on the alignment angle offset between the second beam and the first beam, and a deep reinforcement learning algorithm is used to track the second beam so that the roadside unit and the vehicle's antenna can be docked.

2. The two-stage beam tracking method in the integrated communication and sensing scenario as described in claim 1, characterized in that, The number of antennas can be determined by using the beam coverage area, beamwidth, and given beam alignment angle, which are set to twice the vehicle length.

3. The two-stage beam tracking method in the integrated communication and sensing scenario as described in claim 1, characterized in that, The beam pointing angle can be set in the following ways: Based on the obtained number of antennas, two adjacent beams are overlapped so that each beam covers half of the beamwidth, and the included angle between the two adjacent beams is obtained. The beam pointing angle set in the codebook is set using the included angle.

4. The two-stage beam tracking method in the integrated communication and sensing scenario as described in claim 1, characterized in that, The first modeling approach for Markov decision processes includes: setting the set of codeword-corresponding indices in the codebook during the beam coverage phase as the state space of reinforcement learning, setting the reinforcement learning behavior as incrementing or decrementing beam indices, and using the signal-to-noise ratio of the echo signal as the reward mechanism.

5. The two-stage beam tracking method for integrated communication and sensing scenarios as described in any one of claims 1 to 4, characterized in that, The second modeling approach for the Markov decision process includes: setting the number of second beams that can be covered in the codebook during the beam coverage phase and their corresponding indices as the behavior space, setting the angular offset between the second beam alignment position and the first beam alignment position as the state space, and using the achievable communication rate as the reward mechanism.

6. The two-stage beam tracking method in the integrated communication and sensing scenario as described in claim 5, characterized in that, The achievable communication rate is determined based on the downlink communication link status between the roadside unit and the vehicle.

7. A two-stage beam tracking device for an integrated communication and sensing scenario, characterized in that, The device includes: The beam coverage condition configuration module is used to determine the number of antennas and beam pointing angles for beam coverage in the integrated communication and sensing scenario based on the preset beam coverage range related to the vehicle status. The codebook setting module is used to set the beamforming vector in the codebook using the number of antennas and the beam pointing angle. The beam coverage execution module is used to perform the first modeling of the Markov decision process based on the set codebook, and to solve it using a dynamic reinforcement learning algorithm to obtain the first beam covering the entire vehicle body; wherein, the dynamic reinforcement learning algorithm changes the environmental markers when the environment changes, and updates the Q function based on the changed environmental markers; A fine beamforming module is used to beamform using all antennas and transmit a second beam with a width smaller than the first beam within the coverage area of ​​the first beam. The beam tracking execution module is used for the second modeling of the Markov decision process based on the alignment angle offset between the second beam and the first beam, and uses a deep reinforcement learning algorithm to track the second beam so that the roadside unit can dock with the vehicle's antenna.

8. An electronic device, characterized in that, include: Memory, used to store computer-readable instructions; as well as A processor for executing the computer-readable instructions, causing the electronic device to perform the two-stage beam tracking method in a communication and sensing integrated scenario as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium for storing computer-readable instructions, characterized in that, When the computer-readable instructions are executed by a processor, the processor performs the two-stage beam tracking method for an integrated communication and sensing scenario as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the two-stage beam tracking method for the integrated communication and sensing scenario as described in any one of claims 1 to 6.