Perception-aided predictive handover method in high-speed railway tunnel scenario
By combining a clamped antenna system with a leaky cable in a bistatic architecture and a simulated radio frequency multiplexing waveform in high-speed railway tunnels, a perception-assisted predictive handover method was designed. This method solves the problems of handover decision lag and ping-pong effect in high-speed railway tunnels, and achieves the fusion of high-reliability communication and high-precision location perception, thereby improving the handover success rate and system stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JILIN UNIVERSITY
- Filing Date
- 2026-03-19
- Publication Date
- 2026-06-23
Smart Images

Figure CN122269392A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wireless communication and intelligent transportation technology, and specifically relates to a perception-assisted predictive handover method in a communication and perception integrated system based on a clamping antenna system and a leaky cable in a high-speed railway tunnel scenario. Background Technology
[0002] In recent years, with the rapid development of 5G / 6G mobile communication technology and the deep evolution of intelligent transportation systems, high-speed railway communication, as a key component of the Internet of Things (IoT) network, has increasingly urgent needs for ultra-reliable low-latency connections. High-speed railway tunnel scenarios face severe challenges such as strong multipath fading, significant signal attenuation, and limited coverage geometry, placing extremely high demands on the design of communication systems.
[0003] In tunnel wireless coverage, leaky cables are currently the mainstream technology, achieving controlled energy leakage through periodic slots. However, the static omnidirectional radiation mode of leaky cables cannot concentrate power to a specific receiver, resulting in significant path loss and limited capacity. As an emerging reconfigurable antenna technology, clamped antenna systems form activatable radiating nodes by flexibly clamping dielectric particles on a dielectric waveguide, enabling dynamic beam direction optimization and providing a powerful solution to overcome the inherent limitations of leaky cables. However, existing research on clamped antenna systems is limited to static or low-speed movement scenarios, lacking consideration for the severe dual-selectivity channel effects introduced by high-speed movement.
[0004] In waveform design, the high mobility of high-speed railways introduces severe dual-selectivity channel effects. Most research on communication-sensing integration employs multiple-input multiple-output orthogonal frequency division multiplexing (OFDM) technology to improve spectral efficiency, but severe Doppler broadening can disrupt subcarrier orthogonality, causing significant inter-carrier interference. Orthogonal time-frequency spatial modulation addresses this issue through delay-Doppler processing, but requires high pilot overhead and complexity. Analogous radio frequency multiplexing (ARFDM) achieves complete path separation and full diversity by configuring chirp parameters, providing reliable performance with lower complexity. Its chirp structure facilitates joint delay-Doppler estimation, providing accurate position estimation with Cramer-Rao lower bound quantization accuracy. However, integrating clamped antenna systems with ARFDM in high-speed mobile scenarios remains a gap.
[0005] More critically, frequent handovers are a core bottleneck in high-speed railway tunnel communication. When trains travel at 350 km / h and the distance between access points is several hundred meters, the handover interval can be as short as 0.26 to 1 second, far exceeding typical cellular scenarios. Traditional reactive handover mechanisms only initiate handover after the received signal reference power degrades, resulting in poor performance due to the extremely compressed decision window. Existing methods include precoding-based schemes, dual-connectivity architectures, QoS-aware scheduling, and resource allocation and handover optimization based on deep reinforcement learning, but these methods mostly address each aspect in isolation. Integrating real-time sensing capabilities, antenna configuration, and power allocation into handover optimization introduces a sequential coupling problem of mixed discrete-continuous variables, which traditional convex optimization or iterative optimization techniques struggle to handle effectively. On one hand, handover decisions are discrete variables, tightly coupled with continuous antenna position and power allocation variables; on the other hand, the problem is inherently sequential, with a handover action at a certain moment reshaping the feasible domain of all subsequent time slots, and the system operates under partially observable conditions, making precise future trajectories unavailable.
[0006] Therefore, in the high-speed railway tunnel scenario, how to effectively integrate the dynamic focusing capability of the clamped antenna system, the distributed sensing and receiving capability of the leaky cable, and the anti-dual selective channel capability of the simulated radio frequency multiplexing waveform, how to use sensing information to achieve predictive handover rather than passive reactive handover, and how to jointly optimize the intercoupled variables such as handover decision-making, antenna configuration, and power allocation have become key technical problems that need to be solved by those skilled in the art. Summary of the Invention
[0007] The technical problem this invention aims to solve is that in high-speed railway tunnel scenarios, the high-speed operation of trains leads to significant Doppler shift and extremely short cell dwell time. Traditional reactive handover mechanisms rely on real-time signal quality triggering, resulting in handover decision lag, frequent communication interruptions, and low handover success rates. At the same time, the enclosed environment of the tunnel causes severe multipath propagation, and existing systems lack the ability to accurately predict train positions and plan handover timing in advance, further increasing the probability of handover failure and the "ping-pong effect," making it difficult to meet the service quality requirements of reliable communication in high-speed railways. This invention provides a perception-assisted predictive handover method for high-speed railway tunnel scenarios.
[0008] The specific technical solution of the present invention is as follows:
[0009] A perception-assisted predictive handover method for high-speed railway tunnel scenarios includes the following steps:
[0010] S1. Construct a bistatic communication and sensing integrated architecture based on a clamped antenna system and a leaky cable within a high-speed railway tunnel. A dielectric waveguide is deployed longitudinally at the tunnel roof. Multiple dielectric particles are clamped onto the waveguide to form a dynamically reconfigurable clamped antenna array, serving as the transmitting end and enabling concentrated power focusing towards the moving train. A leaky cable is deployed parallel to the waveguide as the sensing receiving end, receiving the reflected echo signal from the target through periodic slots. Access points are deployed on the train roof as communication receiving ends. Multiple access points are deployed at equal intervals along the tunnel to form a cellular coverage structure.
[0011] S2. Using RF multiplexing as a unified communication and sensing waveform, a signal model under dual-selective channels is established. Based on discrete affine Fourier transform, multi-chirped waveforms are constructed. By optimizing the chirped parameters, complete separation of different propagation paths in the discrete affine Fourier domain is achieved, realizing full diversity gain under high Doppler broadening environment. Downlink and uplink communication channel models are established, and the gain factor of clamped antenna configuration is introduced to quantify the impact of antenna array configuration on communication performance. A bistatic sensing channel model is established, and the time delay-Doppler characteristics of the echo signal are analyzed.
[0012] S3. Based on the sensing signal model, derive the position estimation performance index, use the chirped structure of the simulated radio frequency multiplexing waveform to realize the joint delay-Doppler parameter estimation, construct the Fisher information matrix and derive the Cramer-Rao lower bound as the theoretical limit of position estimation accuracy; establish a position prediction model based on current sensing observations, train speed and antenna configuration to provide forward-looking position information for handover decision-making.
[0013] S4. Design a three-state predictive handover mechanism based on sensing information, defining three handover states: camping, preparation, and handover. Use location prediction derived from the Cramer-Rao lower bound to drive forward-looking state transitions; set handover quality indicators to determine the correctness of handover, including dual criteria based on communication rate improvement and sensing accuracy enhancement; set handover necessity conditions based on coverage distance and minimum service quality; introduce a handover cooling mechanism to prevent the ping-pong effect.
[0014] S5. Propose a perception-assisted dual depth. The network algorithm jointly optimizes handover decision-making, clamped antenna configuration, and power allocation, modeling the optimization problem as a partially observable Markov decision process. It constructs an observation vector containing channel features, sensing features, location features, serving access point features, next access point features, handover state features, and configuration features. A composite action space is designed, including handover actions, antenna configuration actions, and power allocation actions. A multi-component reward function is designed, comprising basic communication rewards, handover decision rewards, and constraint violation penalties. A multi-head encoder extracts different features, a long short-term memory network captures temporal dependencies, and a dual network structure is used to calculate each action type separately. The value is improved, and an auxiliary Cramer-Rhodes lower bound prediction task is introduced to enhance perceptual feature learning.
[0015] Further, in step S2, the simulated radio frequency multiplexing waveform is constructed with multi-chirped basis functions based on discrete affine Fourier transform. The chirped parameters are jointly determined by the maximum normalized Doppler shift and the spacing factor. A chirped periodic prefix is added before transmission to counteract multipath propagation and force channel periodicity. The downlink communication channel model comprehensively considers the time-varying distance from each clamped antenna to the roof access point, the Doppler shift, and the phase accumulation during propagation within the waveguide. The sum of squares of the power coupling coefficients of all antennas satisfies the normalized energy conservation condition. An effective signal-to-noise modulation is introduced by a clamped antenna configuration gain factor consisting of the product of array gain, power allocation efficiency gain, and proximity gain. Compared to the previous model, the bistatic sensing channel model simultaneously detects three access point targets: the current carriage, the previous carriage, and the next carriage. The detection signal propagates through a bistatic path of "transmitting waveguide → target access point → leaky cable". High-speed movement causes different access points to produce distinctly different Doppler characteristics, and this Doppler diversity effect enhances the ability to distinguish multiple targets. The position estimation performance index derives the Cramer-Rao lower bound through the Fisher information matrix, and the lower bound of the root mean square error of position estimation is determined by the trace of the inverse of the Fisher information matrix. Based on the current sensing observation and the prediction of the Cramer-Rao lower bound for future moments, it provides forward-looking position information for handover.
[0016] Furthermore, in step S5, a joint optimization model is established with the objective of maximizing the cumulative communication rate and ensuring timely and beneficial handover. This model specifically includes the following joint optimization model:
[0017]
[0018] in, and These represent instantaneous communication rate and sensing rate, respectively; , and These indicate correct switching, incorrect switching, and missed necessary switching, respectively. This indicates the speed improvement resulting from the handover; This represents a power constraint condition that ensures the total radiated power does not exceed the waveguide input power. This indicates the constraints on the antenna location range. This represents the minimum spacing constraint between adjacent antennas to avoid electromagnetic coupling effects. This indicates the constraint on the number of active antennas; This represents the minimum quality of service constraint. This represents the perception accuracy constraint, ensuring that the position estimation accuracy meets the handover prediction requirements; Indicates the coverage connection constraint; These represent discrete switching action constraints, corresponding to the three states of dwell, preparation, and switching, respectively.
[0019] Furthermore, in step S5, the perception-assisted dual depth... The network algorithm uses a partially observable Markov decision process to model the handover optimization scenario, including six basic elements: state space, observation space, composite action space, transition probability, multi-component reward function, and discount factor. The observation space is a 22-dimensional observation vector composed of seven sets of concatenated features: channel features, perception features, location features, serving access point features, next access point features, handover state features, and configuration features. The composite action space includes three components: handover action, antenna configuration action, and power allocation action. The multi-component reward is divided into basic communication reward, handover decision reward, and constraint violation penalty. The handover decision reward distinguishes five cases: correct handover, incorrect handover, handover failure, necessary handover omission, and correct stay. The algorithm abstracts three elements: state, action, and reward, allowing the agent to continuously adjust its strategy based on reward signals through trial and error interaction in the environment.
[0020] Furthermore, in step S5, the perception-assisted dual depth... The network algorithm first inputs the observation vector into three parallel encoders to extract channel, sensing, and context features, respectively. After concatenation, these features are normalized by layers and processed by a long short-term memory network to capture temporal dependencies. Then, three independent dual network heads are used to... The value is decomposed into state value and advantage function, which output the handover, antenna configuration, and power allocation respectively. Value; simultaneously, an auxiliary Cramer-Robb lower bound prediction task is introduced to enhance perceptual feature learning; priority experience replay and dual The learning method trains the network and updates the network parameters by calculating the gradient of the loss function.
[0021] Beneficial effects:
[0022] As can be seen from the above technical solution, compared with the prior art, the present invention discloses a perception-assisted predictive handover method in a high-speed railway tunnel scenario, which has the following beneficial effects:
[0023] This invention transforms passively triggered reactive handover into proactively planned predictive handover through a three-state forward-looking handover mechanism driven by Cramer-Rao lower bound prediction. This effectively solves the handover failure problem caused by decision lag in extremely short dwell times, significantly improving the handover success rate in high-speed scenarios and greatly reducing the probability of communication interruption. Furthermore, the bistatic architecture of the clamped antenna system and leaky cable, combined with a unified waveform mimicking radio frequency multiplexing, simultaneously achieves highly reliable communication transmission and high-precision position sensing on the same hardware platform. These two aspects mutually enhance each other, eliminating the need for additional independent sensing equipment, thereby reducing system cost and complexity and achieving deep integration of communication and sensing. Further, sensing-assisted dual-depth... The network algorithm performs end-to-end joint optimization of handover decisions, antenna configuration, and power allocation under partially observable environments. It exhibits strong adaptability to complex time-varying channels, effectively preventing the ping-pong effect and ensuring long-term stable system operation. This technology can be widely applied to high-speed mobile communication and sensing integrated systems in confined spaces such as high-speed railway tunnels, subways, and mines. It deeply integrates 6G communication and sensing integration with next-generation information technologies such as intelligent transportation, possessing significant theoretical and practical value. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0025] Figure 1 This is an overall flowchart of the perception-assisted predictive handover method provided by the present invention.
[0026] Figure 2 This is a schematic diagram of the integrated system model of clamped antenna system and leaky cable bistatic communication and sensing system in high-speed railway tunnel scenario provided by the present invention;
[0027] Figure 3 This is a schematic diagram of the system frame structure provided by the present invention;
[0028] Figure 4 This invention provides a perception-assisted dual-depth system. Network architecture diagram; Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] like Figures 1 to 4 As shown in the figure, this invention discloses a perception-assisted predictive handover method in a high-speed railway tunnel scenario, comprising five core steps: system architecture construction, waveform design and channel modeling, perception-assisted position estimation and prediction, three-state predictive handover mechanism design, and joint optimization based on deep reinforcement learning. The following detailed descriptions of each step are provided through multiple embodiments.
[0031] Example 1: Overall System Structure of the Invention
[0032] This embodiment corresponds to step S1, which details the system deployment scheme of the dual-base communication and sensing integrated architecture in high-speed railway tunnels.
[0033] Establish a three-dimensional Cartesian coordinate system inside the tunnel. The axis runs along the length of the tunnel. The axis is in the horizontal direction. The axis is vertical. The system operates under the narrowband assumption, with a carrier frequency of... free space wavelength , wave number .
[0034] The transmitting dielectric waveguide is placed longitudinally along the tunnel at a height of At this location, the horizontal coordinate is By clamping at different positions on the waveguide A reconfigurable antenna array is formed by several dielectric particles, the first... The three-dimensional coordinates of the clamping antenna are , The effective refractive index of a dielectric waveguide is The waveguide wavelength is As the signal propagates within the waveguide, phase accumulation occurs, from the feed point to the... The cumulative phase of the clamped antenna is The core advantage of the clamped antenna system lies in the fact that the spatial configuration of the antenna array can be reconstructed within millisecond timescales by simply adjusting the clamping position of the dielectric particles, without changing the hardware structure, thus providing a physical basis for dynamic beam tracking in high-speed mobile scenarios.
[0035] The leaky cable is deployed parallel to the transmitting waveguide. Axis coordinates are The height is Leaking cable passes through Electromagnetic spatial sampling is achieved using periodic slots, with a slot spacing of [missing information]. , No. The spatial position of each slot is , Unlike dielectric waveguides that use flexible clamping antennas to concentrate detection power, leaky cables utilize densely and periodically distributed slots to efficiently capture reflected echo signals, eliminating the sensing blind spots of traditional discrete antenna arrays and providing dense spatial sampling to improve sensing resolution.
[0036] Access points are deployed on the train roof as communication receivers. Multiple access points are deployed at equal intervals along the tunnel, forming a set of access points. Adjacent spacing , No. The access points are located at Coverage radius of each access point Adjacent coverage areas overlap to ensure seamless connectivity. Trains travel at speeds... along Traveling in the positive direction of the axis, at the initial moment The time car reference point is located at ,time The longitudinal displacement is .
[0037] Through this deployment, the clamped antenna array simultaneously undertakes the tasks of transmitting communication data and radiating sensing and detection signals, while the leaky cable is dedicated to receiving sensing echoes. The rooftop access point serves as both a communication terminal and a sensing target. This bistatic configuration fully utilizes the concentrated power focusing capability of the transmitter and the distributed sensing and receiving capability of the receiver, forming complementary advantages.
[0038] Example 2: Simulated RF Multiplexing Waveform Design of the Present Invention
[0039] This embodiment corresponds to the waveform design section in step S2, and details the construction method of the simulated radio frequency multiplexing waveform and the principle of achieving full diversity in a dual selective channel.
[0040] Affine radio frequency multiplexing (RFD) is a multi-chirped waveform modulation technique based on discrete affine Fourier transform. Bandwidth is considered. Symbol duration The simulated radio frequency multiplexing symbol, in which For the number of chirped subcarriers, Let be the subcarrier spacing. For discrete affine Fourier domain sign vectors, after... The point-inverse discrete affine Fourier transform, when converted to the time domain, gives the basis functions as follows:
[0041]
[0042] in and For chirp parameters, The chirp slope is determined. The inverse discrete affine Fourier transform can be expressed in matrix form as follows: ,in The discrete affine Fourier transform matrix. For the normalized DFT matrix, the diagonal matrix .
[0043] To ensure full diversity in a dual-selective channel and avoid overlap of different paths in the discrete affine Fourier domain, the chirp parameter... Set as ,in The integer part of the maximally normalized Doppler frequency shift. To counteract the interval factor of fractional Doppler. Parameters It can be set to any irrational number or sufficiently smaller than The rational number of the . This parameter design ensures that different propagation paths are naturally separated in the discrete affine Fourier domain under high Doppler broadening conditions, thereby achieving full diversity gain.
[0044] Before transmission, a chirped periodic prefix is added before the simulated radio frequency multiplexing symbol to combat multipath propagation and force channel periodicity in the time domain. Represented as , ,in This is an integer not less than the maximum number of channel delay taps. The continuous-time transmitted signal is... .
[0045] Example 3: Communication Channel Modeling of the Present Invention
[0046] This embodiment corresponds to the communication channel modeling part in step S2, and details the channel models of downlink and uplink communication links and the method for constructing the gain factor of the clamping antenna configuration.
[0047] In the downlink communication link, assuming the high-speed train travels at a speed of... along When traveling in the positive direction of the axis, the roof connection point is at time The position is . No. The time-varying distance from the clamped antenna to the access point is: ,in The corresponding Doppler frequency shift is Signal propagation within a waveguide results in phase accumulation, and the propagation vector within the waveguide... Represented as ,in For the first The power coupling coefficient of the clamped antenna satisfies , denoted as the effective refractive index of the dielectric waveguide. This vector fully characterizes the amplitude and phase modulation effect of the waveguide propagation on the signals at each clamped antenna.
[0048] When a signal is coupled from the waveguide into free space, the free space propagation vector... The Each element is represented as ,in Taking into account the effects of antenna gain and free-space path loss, The wave number is [wavenumber]. The received signal at the access point is represented as [significant wavenumber]. This expression integrates the combined effects of waveguide propagation, free-space propagation, the Doppler effect, and the input signal.
[0049] time The instantaneous received signal-to-noise ratio is ,in For the first The total phase of the clamped antenna path Total transmission power, This refers to noise power. Within the coherent processing range. Within, the average signal-to-noise ratio is According to Shannon's capacity formula, the upper limit of the communication rate is... This communication rate expression establishes a quantitative relationship between system physical parameters and communication performance, providing a clear objective function for subsequent joint optimization.
[0050] To accurately characterize the impact of clamp antenna configuration on communication performance, a configuration gain factor is introduced. Array gain As the number of active antennas increases linearly, Power allocation efficiency gain By normalizing entropy Characterization, in which Position gain Distance by power weighting Modeling. The effective signal-to-noise ratio is... .
[0051] In the uplink communication link, the leaky cable acts as the uplink signal receiver, from the access point to the... The time-varying distance of each slot is Upward Doppler frequency shift is Time-varying uplink channel vector The Each element is represented as The propagation vector within the leaky cable is... ,in The effective refractive index of the leaky cable. The uplink signal received by the base station is... .time The instantaneous uplink signal-to-noise ratio is The expression contains The coherent cumulative effect at each receiving point. The uplink communication rate is... .
[0052] Example 4: Performance Indicators of Sensing Channel Modeling and Location Estimation
[0053] This embodiment corresponds to step S3, which details the method for establishing the bistatic sensing channel model and the derivation of the position estimation performance index based on the Fisher information matrix.
[0054] The system needs to detect three key targets simultaneously: the current carriage's access point. The connection point of the previous carriage Access point of the next carriage At that moment The positions of the three access points relative to the tunnel coordinate system are as follows: , , ,in Let be the longitudinal distance between the centers of adjacent carriages. The vector of real-valued position parameters to be estimated is . .
[0055] The detection signal emitted by the waveguide illuminates each access point, generating an echo signal. The radiated signal received by each access point from the waveguide is Among them, from the waveguide The clamping antenna to the first Time-varying free-space channel vectors of each access point The The elements are
[0056]
[0057] in For the first The access point and the first The time-varying distance between the clamped antennas This is the downlink Doppler frequency shift. The equivalent radar cross section of each access point is: The reflection coefficient is modeled as Under the assumptions of the Swerling-I model ,in This represents the average reflection intensity.
[0058] Leaky cables capture echo signals through densely distributed slots. From the... The access point to the first The echo path distance of each slot is , The echo path Doppler frequency shift is the total two-way Doppler frequency shift. From the first Time-varying echo channel vectors from each access point to each slot of the leaky cable The The elements are
[0059]
[0060] Considering the propagation effect within the cable, the total sensed signal received by the leaky cable is:
[0061]
[0062] in , This is additive white Gaussian noise. This expression reflects the superposition effect of the echoes from the three access points.
[0063] When the chirping parameter When the design conditions are met, the path indices of different objectives are naturally separated in the discrete affine Fourier domain, achieving effective multi-objective differentiation.
[0064] To quantitatively assess the accuracy limit of position estimation, Fisher's information matrix theory is employed. The system transmits within the coherent processing interval. A series of probe symbols, Fisher information matrix The element is defined as
[0065]
[0066] in Let be the desired received signal vector. This matrix comprehensively considers the joint information of spatial gradient and Doppler effect. The lower bound of the root mean square error of the position estimation is .
[0067]
[0068] The smaller this lower bound, the higher the accuracy of the position estimation. Simultaneously, a sensing rate is defined to quantify the system's ability to acquire sensing information.
[0069]
[0070] in For a moment The equivalent sensing channel matrix, This represents the number of snapshots within the coherent processing interval. The Doppler diversity effect introduced by high-speed motion results in drastically different Doppler characteristics at different access points. (Current carriage...) The Doppler shift is close to zero in the previous carriage. A negative Doppler frequency shift is generated in the rear carriage. The positive Doppler frequency shift is generated. This Doppler diversity effect makes the three access points spatially close and can also be effectively distinguished through path separation in the discrete affine Fourier domain, thereby enhancing the sensing rate.
[0071] Based on current sensing observations Train speed and the position of the clamping antenna Predicting future moments Crame-Lower World This predictive model provides forward-looking location information for handover decisions, enabling the system to proactively anticipate train position changes before signal degradation.
[0072] Example 5: Design of a Three-State Predictive Handover Mechanism
[0073] This embodiment corresponds to step S4, which details the three-state predictive handover mechanism based on sensing information, including handover state definition, quality judgment, necessity conditions, and cooling anti-ping-pong mechanism.
[0074] Constrained by the unidirectional travel characteristics of trains, cross-zone handover decisions are limited to forward transitions. The handover action is defined as follows: ,in Maintain the current service access point. Initiate preparations for handover to the next access point. The system performs a handover to the next forward access point. Sensing capabilities provide crucial information for handover decisions; the Cramer-Rao lower bound, a key indicator of positioning accuracy, is related to communication channel quality. The sensing capability, mimicking radio frequency multiplexing waveforms, allows the system to predict train position before signal degradation and proactively initiate handover. The state enables "build first, disconnect later" resource pre-allocation. This is the core innovation of this invention, which distinguishes it from the traditional reactive switching mechanism.
[0075] Switching quality indicators The determination rule is: when the communication rate of the new service access point... Subtract the old access point communication rate A value greater than zero indicates a correct switch; when the rate loss is within the tolerance threshold... The lower bound of the inner and Clamer-Loh improved Exceeding the threshold In some cases, the handover is considered correct; in others, it is considered incorrect. The design of this dual criterion is that even if the communication rate is not directly improved, if the sensing accuracy is significantly improved, the handover is still considered beneficial because higher sensing accuracy will provide more accurate location predictions in subsequent time slots, thereby indirectly improving communication performance.
[0076] The necessity of switching is determined by the coverage constraint. ,in The distance to the current service access point. To cover the boundary coefficient, This is an indicator function. Handover is deemed necessary when the train's distance from the current service access point exceeds the coverage boundary threshold or the communication rate falls below the minimum quality of service requirement. The handover cooling mechanism is configured with a set cooling duration. Implemented, switching execution occurs during the cooldown countdown. During this process, new handover actions are prevented to avoid the ping-pong effect caused by signal fluctuations. This mechanism ensures the stability of handover decisions and avoids communication service interruptions caused by frequent back-and-forth handovers in overlapping coverage areas. This problem has characteristics such as sequential decision-making, prediction-dependent rewards, mixed discrete-continuous actions, and tight coupling between perception and communication, which cannot be handled by conventional convex optimization or iterative optimization methods and require a deep reinforcement learning framework for solution.
[0077] Example 6: Modeling the Joint Optimization Problem
[0078] This embodiment corresponds to the first half of step S5, detailing the mathematical modeling of the joint optimization problem and the construction method of a partially observable Markov decision process.
[0079] Optimization variables include handover decisions Antenna clamping position Number of activated antennas and power allocation factor The objective function is
[0080]
[0081] in and For instantaneous communication and sensing speed, and These indicate the occurrence of correct and incorrect switching, respectively. Indicates the necessary switch that was missed. The rate improvement resulting from the switching. The constraints include eight items: For power constraints ; Antenna position range constraints ; Minimum spacing constraint for adjacent antennas ; To activate antenna quantity constraints ; Minimum service quality constraint ; Constraints on perception accuracy ; To cover connection constraints ; Constraints for Discrete Switching Actions This problem has characteristics such as mixed discrete-continuous variables, sequential coupling, prediction-reward dependence, and tight coupling between perception and communication. Conventional convex optimization or iterative optimization methods cannot effectively handle it, and a deep reinforcement learning framework is required to solve it.
[0082] The optimization problem is modeled as a partially observable Markov decision process, with tuples... Observation vector It consists of seven sets of features, with a total dimension of 22. Channel features. Includes normalized communication rate and sensing rate. Sensing characteristics. Includes the normalized current Cramer-Rhodes lower bound, the predicted Cramer-Rhodes lower bound, the minimum Cramer-Rhodes lower bound within the coverage area, the Cramer-Rhodes lower bound improvement, and a switching necessity indicator. Location and movement characteristics. Service access point characteristics Includes normalized distance, coverage indication, communication rate, and relative location. Next Access Point Characteristics Includes normalized distance, coverage indication, communication rate, and rate difference. Handover state characteristics. ,in It corresponds to three states: resident, ready, and switch. Configuration features. Includes the normalized antenna count and the service access point index. The complete observation vector is... This 22-dimensional observation vector encompasses comprehensive information such as channel quality, sensing accuracy, location information, access point status, handover status, and system configuration, providing ample decision-making basis for deep reinforcement learning agents.
[0083] The composite action space contains three components: switching action. Antenna configuration action Jointly select the number and position pattern of the clamping antennas to activate, base number Power distribution action Choose from predefined allocation patterns, including uniform allocation, centralized allocation, pre-weighted allocation, and post-weighted allocation, cardinality. ,time The total action is a tuple .
[0084] Multiple Rewards It consists of three parts. Basic communication reward ),in and These are the weights for communication and sensing rates, respectively. Switching decision reward. Distinguish between five scenarios: Correct switching yields positive rewards Incorrect switching is penalized. Switching failure is subject to a fixed penalty. Necessary switching omissions will be penalized. Correct residency reward is .in and These represent positive and negative rate changes, respectively. Constraint violation penalty. The penalties for violating four constraints—minimum rate, Cramer-Rhodes lower bound, coverage connectivity, and minimum antenna spacing—are respectively set as follows: , , and These represent the penalty coefficients for each item.
[0085] Example 7: Perception-Assisted Dual Depth Network Architecture and Training
[0086] This embodiment corresponds to the latter half of step S5, and details the perception-assisted dual depth. Network architecture design and training methods.
[0087] The observation vectors are processed by three parallel encoders to extract different types of features. Channel Feature Encoder It is a two-layer fully connected network that processes communication rate and sensing rate information. (Sensing feature encoder) It is a two-layer fully connected network that processes perception-related information such as the Cramer-Rhodes lower bound. Context feature encoder. It is a two-layer fully connected network, in which It includes contextual information such as location, access point, handover status, and configuration. The three sets of encoded features are concatenated and then normalized to obtain a fused feature vector. The design of the multi-head encoder enables the network to learn independent representations for different types of features, avoiding interference between features.
[0088] Long Short-Term Memory (LSTM) network layers process fused features to capture temporal dependencies, and their dynamic equations include forgetting gates. Input gate Candidate cell status Cell state update and output gate Hidden state ,in for function, This is an element-wise multiplication. The introduction of LSTM layers enables the network to utilize temporal patterns in historical observation sequences, which is of great significance for understanding the continuous changing trends of train positions in high-speed railway scenarios.
[0089] For each action type Computation using independent dual network heads Value. Dual networks will The value is decomposed into two parts: state value and advantage function. The value flow is... The dominant flow is . The value is obtained by combining two streams. The advantage of this dual structure lies in the fact that the separation of the state value function and the advantage function enables the network to learn more efficiently which states are inherently valuable, as well as the relative merits of each action in a given state.
[0090] Assist Crame-Roll lower bound prediction head ,in Ensure non-negative output. It is a two-layer fully connected network. The role of this auxiliary task is to enhance the network's ability to learn perceptually relevant features by predicting the Cramer-Rhodes lower bound. Hidden states encode richer perceptual information.
[0091] Use dual depth The network framework is trained using a priori experience replay approach. The total loss function is... The timing difference loss is adopted. Loss guarantees training stability ,in Weights are assigned to prioritize playback based on importance. Target Values are double Learning computation , For the target network. Aggregation. Values take three action heads average value Cramer-Rao lower bound loss prediction as an auxiliary task. .
[0092] Adaptive Greedy strategy This design increases exploration at higher speeds to cope with more rapidly changing conditions. This allows the agent to maintain a higher exploration rate while operating at high speeds because the environment changes more quickly, requiring more exploration to discover the optimal strategy. The target network employs a soft update strategy. ,in Ensure the stability of the target network.
[0093] The complete training process is as follows: Initialize and evaluate network parameters and target network parameters Initialize the priority experience replay buffer Set the initial exploration rate For each training round, reset the environment to obtain initial observations. Initialize the hidden state of the Long Short-Term Memory network. Set to switch cooldown timers At each time step, perceptual features are extracted and calculated. Check the necessity of switching. For each action type, with probability Randomly select an action, otherwise select to make The action with the highest value is executed as a compound action, and the reward and the next observation are observed and stored in the experience replay buffer. When the buffer has sufficient samples, mini-batch samples are sampled according to priority, the temporal difference objective and total loss are calculated, and the network parameters are updated. Every [period]... Step-by-step soft update of the target network; exploration rate decays at the end of each round. Repeat until finished. There are [number] training rounds. During the deployment phase, the inference complexity at each time step is [value]. It is suitable for real-time decision-making needs with millisecond-level decision intervals in high-speed railway scenarios.
[0094] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. Embodiments 1 to 7 progressively elucidate the complete technical solution of the present invention, from system architecture, waveform design, communication channel modeling, sensing channel modeling and location estimation, handover mechanism, optimization problem modeling to deep reinforcement learning network architecture and training.
[0095] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A perception-assisted predictive handover method for high-speed railway tunnel scenarios, comprising the following steps: S1. Construct a bistatic communication and sensing integrated architecture based on a clamped antenna system and a leaky cable inside a high-speed railway tunnel. Deploy a dielectric waveguide longitudinally at the top of the tunnel. Form a dynamically reconfigurable clamped antenna array by clamping multiple dielectric particles on the waveguide as the transmitting end to achieve concentrated power focusing towards the moving train. Deploy a leaky cable at a position parallel to the waveguide as the sensing receiving end to receive the target reflected echo signal through periodic slots. Access points are deployed on the roof of the train as communication receivers; multiple access points are deployed at equal intervals along the tunnel to form a cellular coverage structure. S2. Using RF multiplexing as a unified communication and sensing waveform, a signal model under dual selective channels is established. Based on discrete affine Fourier transform, multi-chirped waveforms are constructed. By optimizing the chirped parameters, complete separation of different propagation paths in the discrete affine Fourier domain is achieved, realizing full diversity gain under high Doppler broadening environment. Downlink and uplink communication channel models are established, and the gain factor of clamped antenna configuration is introduced to quantify the impact of antenna array configuration on communication performance. A bistatic sensing channel model was established to analyze the time delay-Doppler characteristics of the echo signal; S3. Based on the sensing signal model, derive the position estimation performance index, use the chirped structure of the simulated radio frequency multiplexing waveform to realize the joint delay-Doppler parameter estimation, construct the Fisher information matrix and derive the Cramer-Rao lower bound as the theoretical limit of the position estimation accuracy. Establish a location prediction model based on current sensing observations, train speed, and antenna configuration to provide forward-looking location information for handover decisions; S4. Design a three-state predictive handover mechanism based on sensing information, defining three handover states: camping, preparation, and handover. Use location prediction derived from the Cramer-Rao lower bound to drive forward-looking state transitions; set handover quality indicators to determine the correctness of handover, including dual criteria based on communication rate improvement and sensing accuracy enhancement; set handover necessity conditions based on coverage distance and minimum service quality; introduce a handover cooling mechanism to prevent the ping-pong effect. S5. Propose a perception-assisted dual depth. The network algorithm jointly optimizes handover decision-making, clamped antenna configuration, and power allocation, modeling the optimization problem as a partially observable Markov decision process. It constructs an observation vector containing channel features, sensing features, location features, serving access point features, next access point features, handover state features, and configuration features. A composite action space is designed, including handover actions, antenna configuration actions, and power allocation actions. A multi-component reward function is designed, comprising basic communication rewards, handover decision rewards, and constraint violation penalties. A multi-head encoder extracts different features, a long short-term memory network captures temporal dependencies, and a dual network structure is used to calculate each action type separately. The value is improved, and an auxiliary Cramer-Rhodes lower bound prediction task is introduced to enhance perceptual feature learning.
2. The perception-assisted predictive handover method in a high-speed railway tunnel scenario according to claim 1, characterized in that, In step S2, the RF multiplexing waveform is constructed with multi-chirped basis functions based on discrete affine Fourier transform. The chirped parameters are determined by the maximum normalized Doppler shift and the spacing factor. A chirped periodic prefix is added before transmission to counteract multipath propagation and force channel periodicity. The downlink communication channel model comprehensively considers the time-varying distance from each clamping antenna to the roof access point, the Doppler shift, and the phase accumulation of propagation within the waveguide. The sum of squares of the power coupling coefficients of all antennas satisfies the normalized energy conservation condition. The clamping antenna configuration gain factor, which is composed of the product of array gain, power allocation efficiency gain, and position proximity gain, is introduced to modulate the effective signal-to-noise ratio. The bistatic sensing channel model simultaneously detects three access point targets: the current carriage, the preceding carriage, and the following carriage. The detection signal propagates through a bistatic path: "transmitting waveguide → target access point → leaky cable". High-speed movement causes distinctly different Doppler characteristics at different access points, and this Doppler diversity effect enhances the ability to distinguish multiple targets. The position estimation performance index derives the Cramer-Rao lower bound through the Fisher information matrix, and the lower bound of the root mean square error of position estimation is determined by the trace of the inverse of the Fisher information matrix. Based on the current sensing observations and the prediction of the Cramer-Rao lower bound for future moments, forward-looking position information is provided for handover.
3. The perception-assisted predictive handover method in a high-speed railway tunnel scenario according to claim 1, characterized in that, In step S5, a joint optimization model is established with the goal of maximizing the cumulative communication rate and ensuring timely and beneficial handover. This model includes the following joint optimization model: ; in, and These represent instantaneous communication rate and sensing rate, respectively. , and These indicate correct switching, incorrect switching, and missed necessary switching, respectively. This indicates the speed improvement brought about by the handover; This represents a power constraint condition that ensures the total radiated power does not exceed the waveguide input power. This indicates the constraints on the antenna location range. This represents the minimum spacing constraint between adjacent antennas to avoid electromagnetic coupling effects. This indicates the constraint on the number of active antennas; This represents the minimum quality of service constraint. This represents the perception accuracy constraint, ensuring that the position estimation accuracy meets the handover prediction requirements; Indicates the coverage connection constraint; These represent discrete switching action constraints, corresponding to the three states of dwell, preparation, and switching, respectively.
4. The perception-assisted predictive handover method in a high-speed railway tunnel scenario according to claim 1, characterized in that, In step S5, the perception-assisted dual depth The network algorithm uses a partially observable Markov decision process to model the handover optimization scenario, including six basic elements: state space, observation space, composite action space, transition probability, multi-component reward function, and discount factor. The observation space is a 22-dimensional observation vector composed of seven sets of concatenated features: channel features, perception features, location features, serving access point features, next access point features, handover state features, and configuration features. The composite action space includes three components: handover action, antenna configuration action, and power allocation action. The multi-component reward is divided into basic communication reward, handover decision reward, and constraint violation penalty. The handover decision reward distinguishes five cases: correct handover, incorrect handover, handover failure, necessary handover omission, and correct stay. The algorithm abstracts three elements: state, action, and reward, allowing the agent to continuously adjust its strategy based on reward signals through trial and error interaction in the environment.
5. The perception-assisted predictive handover method in a high-speed railway tunnel scenario according to claim 1, characterized in that, In step S5, the perception-assisted dual depth The network algorithm first inputs the observation vector into three parallel encoders to extract channel, sensing, and context features, respectively. After concatenation, these features are normalized by layers and processed by a long short-term memory network to capture temporal dependencies. Then, three independent dual network heads are used to... The value is decomposed into state value and advantage function, which output the handover, antenna configuration, and power allocation respectively. Value; at the same time, an auxiliary Cramer-Robb lower bound prediction task is introduced to enhance perceptual feature learning; Employing priority experience replay and dual The learning method trains the network and updates the network parameters by calculating the gradient of the loss function.