A physical layer frame synchronization method based on multi-scale feature extraction and residual attention mechanism

CN122845370APending Publication Date: 2026-09-29XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611127498.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]然而,该方法主要依赖单一尺度的一维卷积结构对接收信号进行整体特征建模,网络感受野主要通过卷积层堆叠获得,对于不同前导序列长度、不同同步不确定区间以及不同多径扩展条件的适应能力仍然有限;同时,该方法采用全序列回归并结合峰值检测的方式完成同步判决,其输出结果易受噪声扰动和局部伪峰影响,对判决门限和后处理策略具有一定依赖性

Benefits of technology

本发明可应用于 5G/6G 基站、物联网网关、车联网终端、卫星通信接收机等无线通信设备的基带接收链路中,用于在复杂信道条件下确定无线数据帧的起始位置。针对城市楼宇反射、高速移动、低信噪比传输、载波频偏、相位噪声以及 IQ 不平衡等因素叠加导致的帧同步困难问题,本发明通过多级技术手段的协同演进,提升了帧起始位置检测的性能。其一,通过前导序列编码(步骤3)引入本地前导结构先验信息,并配合残差注意力机制(步骤4)对特征通道进行自适应加权。该协同过程实现了特征层面的精准聚焦,使网络能够从源头上有效强化与帧同步相关的有效信息,并抑制由多径衰落、噪声干扰及射频损伤导致的伪峰。其二,在多任务判决阶段(步骤5),峰值检测分支集成了多尺度并行特征提取结构。通过不同尺度感受野的并行处理,网络既能捕获应对突变干扰的采样点级局部特征,亦可兼顾应对时延扩展的符号级全局上下文特征,从而实现多维度的特征表征。其三,通过上述“前导先验引导”与“多尺度并行感知”的特征增强机制,结合位置回归分支提供的连续性平滑约束,实现了对帧起始位置的高精度协同判决。该方法提升了接收机在复杂信道环境下同步定位的准确性与鲁棒性。。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845370A_ABST
    Figure CN122845370A_ABST
Patent Text Reader

Abstract

The application discloses a physical layer frame synchronization method based on multi-scale feature extraction and residual attention mechanism, which comprises the following steps: step 1, converting original IQ signals and local preambles into standardized tensors that can be processed by a neural network; step 2, simultaneously capturing local detailed features and global context features of the signals through a parallel multi-branch convolution structure; step 3, encoding ideal preamble structure information into a fixed-dimension feature vector; step 4, obtaining weighted output features; step 5, through the cooperative work of two branches of peak detection and position regression; step 6, through a combined loss function, simultaneously constraining detection accuracy, estimation accuracy and positioning deviation in the training stage. The application improves the accuracy and robustness of wireless data frame start position detection, thereby improving the reliability of subsequent channel estimation, equalization processing and data demodulation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication technology, specifically to a physical layer frame synchronization method based on multi-scale feature extraction and residual attention mechanism. Background Technology

[0002] In burst-mode wireless communication systems, frame synchronization is a crucial step in the receiver's baseband processing. Its purpose is to accurately determine the start time of a data frame from the received signal, even in the presence of channel distortion and non-ideal factors in the RF front-end. Existing frame synchronization techniques typically rely on the good autocorrelation characteristics of a specific preamble sequence under ideal channel conditions. The frame start point is determined by calculating the cross-correlation function between the received signal and the local preamble sequence, and based on the position of the correlation peak. However, in real-world wireless transmission environments, multipath fading can cause correlation peak broadening or even splitting, carrier frequency offset can lead to attenuation and position shift of the correlation peak, and non-ideal factors in the RF front-end can increase the correlation sidelobe level. When these multiple non-ideal factors are combined, the accuracy of traditional cross-correlation-based frame synchronization methods decreases, with carrier frequency offset having a particularly significant impact on synchronization performance degradation.

[0003] In recent years, deep learning-based technologies have been increasingly introduced into the physical layer processing of wireless communications to improve the robustness of frame synchronization under non-ideal channel conditions. A relatively close existing technique was proposed by Kalade et al., who disclosed a physical layer frame synchronization method based on a fully convolutional neural network in their paper "Training Deep Filters for Physical-Layer Frame Synchronization" (IEEE Open Journal of the Communications Society, 2022). This method constructs an end-to-end deep convolutional network model, performs one-dimensional convolution processing on the in-phase and quadrature components of the received signal, directly regresses the synchronization decision result with the same length as the input sequence, and determines the frame start time by detecting the peak position in the output sequence. During the training phase, the method uses a mean squared error loss function for parameter optimization and is trained and validated in a simulated channel environment containing additive noise, carrier phase offset, carrier frequency offset, and specific multipath fading conditions. This method completes frame synchronization decision in an end-to-end manner and shows certain performance advantages over traditional incoherent cross-correlation methods under certain simulation conditions.

[0004] However, this method primarily relies on a single-scale one-dimensional convolutional structure to model the overall features of the received signal. The network's receptive field is mainly obtained through stacking convolutional layers, resulting in limited adaptability to different preamble lengths, synchronization uncertainty intervals, and multipath propagation conditions. Furthermore, this method employs full-sequence regression combined with peak detection for synchronization decision-making, making its output susceptible to noise disturbances and local spurious peaks, and dependent on decision thresholds and post-processing strategies. Moreover, the method primarily uses a purely data-driven approach in its network structure and training objective design, failing to explicitly introduce or utilize prior information about the known structure of the preamble, thus hindering the full realization of the preamble's correlation and decision constraint role in frame synchronization. Under conditions of strong noise or short preamble sequences, there is still room for improvement in its synchronization decision accuracy and stability. Summary of the Invention

[0005] In order to overcome the shortcomings of the existing technology, the present invention aims to provide a physical layer frame synchronization method based on multi-scale feature extraction and residual attention mechanism, so as to improve the accuracy and robustness of the receiver in detecting the frame start position under complex channel conditions.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A physical layer frame synchronization method based on multi-scale feature extraction and residual attention mechanism includes the following steps: Step 1: Received signal preprocessing and preamble template construction: Convert the original IQ signal and local preamble into a normalized tensor that can be processed by the neural network to obtain the input sequence. This provides the input basis for subsequent processing, ensuring that the model can learn the manifestation of channel impairments such as multipath and frequency offset in the IQ components; Step 2: Multi-scale feature extraction; By using a parallel multi-branch convolutional structure, the local detail features and global context features of the received baseband signal are captured simultaneously, enabling the network to acquire multi-scale feature representations under complex channel conditions. ; Step 3: Preamble Sequence Encoder; Encodes the structural information of the local preamble into a fixed-dimensional feature vector. This guides the network to focus on the structural features of the target leader; Step 4: Residual Attention Module; using the preceding feature vector With the multi-scale feature representation Feature map after fusion As input, residual connections are used to avoid feature degradation, and a channel attention mechanism is employed to highlight effective feature channels related to synchronization decisions, suppress redundant information and noise features. In complex scenarios with multiple impairments such as multipath interference, frequency offset, and noise, this helps the network focus on key information and obtain weighted output features. ; Step 5: Multi-task output header; with the shared fusion feature As a parallel input, it works in conjunction with an independent peak detection head and a regression output head; the peak detection head... Perform high, medium, and low-scale feature extraction and attention fusion to output peak prediction. The method detects abrupt peaks at the start of a frame; the regression output head is processed through a convolution, batch normalization, ReLU, and sigmoid cascade structure. Output regression prediction output The frame start position is continuously estimated in the form of a probability distribution; finally, the peak prediction is output. With the regression prediction output A fusion decision is made to obtain the optimal estimate of the frame start position. ; Step 6: Combine loss function; By combining loss function, the detection accuracy, estimation accuracy and positioning error are constrained simultaneously during the training phase. This ensures that the model can learn the optimal feature representation and decision strategy when facing complex and unknown channel environments. This enables highly reliable positioning of the start position of wireless data packet frames in complex wireless communication environments and provides an accurate synchronization reference for symbol timing recovery, channel estimation, equalization processing and data demodulation in the communication baseband receiver.

[0007] Steps 1 to 6 above together constitute the complete process of the frame synchronization method of the present invention. Its final output is a probability distribution sequence of the frame start position. By performing peak detection on this sequence, the estimated value of the frame start position can be obtained, thereby achieving high-accuracy frame synchronization detection under complex channel conditions. The above six steps are interconnected and progressively advance to form an end-to-end frame synchronization processing flow. Through the synergistic effect of multi-scale feature extraction, preamble guidance, attention enhancement and multi-task joint decision, it effectively copes with complex channel conditions such as multipath fading, carrier frequency offset, radio frequency impairment and noise interference, and significantly improves the detection accuracy and robustness of frame synchronization.

[0008] Step 1 specifically involves: Received signal preprocessing: The receiver acquires the radio frequency received signal containing the local preamble and data portion, and obtains the baseband complex sampling sequence after down-conversion and analog-to-digital conversion; the baseband complex sampling is then subjected to root-raised cosine filtering to suppress out-of-band interference and achieve matched filtering, resulting in the filtered baseband signal; and a basic feature map is constructed through a basic feature extraction layer. ; The filtered baseband signal is decomposed into in-phase components (I) and quadrature components (Q), and stacked along the time dimension to form a two-dimensional input feature matrix of size (2, L), which serves as the input to the subsequent neural network, where L is the sequence length. Preamble template construction: The known preamble symbol sequence (ZC sequence) undergoes the same shaping filtering process as the transmitter, i.e., first root-raised cosine transmission filtering, then root-raised cosine reception filtering (matched filtering), to obtain a preamble template that is homomorphic to the received signal; the real and imaginary parts of the obtained preamble template are decomposed to form a two-dimensional preamble feature matrix with the same dimension as the input feature matrix, which is used to characterize the structural information of the preamble code; the preamble template can be adapted to preamble sequences of arbitrary length and modulation scheme; The baseband feature matrix output from step 1 is denoted as the input sequence. ,in Let the sequence length be denoted as the constructed two-dimensional preceding feature matrix as the prior template. ,in This is the length of the leading template.

[0009] Step 2 specifically involves: To capture multi-scale temporal features in the received signal, a parallel multi-branch convolutional feature extraction structure is constructed, which contains three processing paths with different temporal resolutions: First branch (details branch): In the input sequence We directly use small-scale convolution kernels (kernel size 3, dilation rate 2) for feature extraction, which preserves the short-term transient information of the signal and is used to capture key local peak features during the synchronization process. Second branch (mesoscale branch): First, process the input sequence... We perform a 2x downsampling and then use a medium-scale convolutional kernel (kernel size 5, dilation rate 2) to extract features. This expands the receptive field while taking into account detailed information, in order to cope with the time delay spread caused by multipath propagation. Third branch (global branch): For the input sequence... A 4x downsampling was performed, and then a large-scale convolutional kernel (kernel size 7, dilation rate 2) was used to extract long-term dependent features to capture the overall envelope structure of the signal. The output features of each branch are interpolated (upsampled) to restore the original time length, and then concatenated along the feature channel dimension to form a multi-scale feature representation. This multi-scale feature representation achieves a deep fusion of fine synchronization features at the sampling point level and macroscopic channel features at the symbol level at the physical level, providing physical feature basis for anti-interference frame synchronization decision under complex channels.

[0010] This process can be represented as: in, The input sequence is the output of step 1. , , These represent the convolution operations in the three branches, respectively. Indicates k-fold downsampling. Indicates k times upsampling, This represents channel-dimensional concatenation. Through the multi-scale feature extraction described above, the network can simultaneously capture local details and global context in the signal, thereby achieving robust synchronous feature representation under complex channel conditions.

[0011] Step 3 specifically involves: The leader sequence, as prior information in the synchronization process, contains stable and distinguishable temporal structural features. To make full use of this prior information, this method sets up a leader sequence encoder to map leader sequences of different lengths into embedded feature vectors of fixed dimensions, which are used to participate in subsequent synchronization feature fusion and decision-making. The leader sequence encoder consists of a convolutional feature extraction layer and a temporal dimension aggregation layer. It extracts hierarchical features of the leader sequence through progressive convolution operations and aggregates the temporal dimension features, thereby uniformly mapping the variable-length leader sequence into a fixed-dimensional embedded feature vector. Its expression is as follows: Composed of two convolutional layers and global average pooling, it can effectively compress the temporal structure of the leader sequence: = in To represent the embedded features of the leading sequence, the temporal aggregation unit uses global pooling to compress the leading sequence features, thereby obtaining a compact representation that includes key temporal pattern information. The generated leading features and the features output by the multi-scale feature extraction module The feature map is obtained by fusing information across feature channels through a feature transformation unit, and the fusion is performed at the channel dimension. This enhances the network's ability to perceive the target's leading structure, thereby improving the accuracy and robustness of synchronous positioning under complex channel conditions such as multipath, noise, and frequency offset.

[0012] Step 4 specifically involves: To enhance feature representation capabilities and improve the network's discrimination performance under complex channel conditions, this method introduces a residual attention module during feature processing. The residual attention module introduces residual connections to achieve cross-layer feature transfer and superposition, effectively mitigating feature degradation while enhancing the retention of key information. The residual module employs a pre-activation structure, first performing a nonlinear transformation on the input features, then fusing them with the original features through residual connections to enhance the stability and expressive power of the feature mapping. Simultaneously, a channel attention mechanism is introduced to adaptively adjust the importance of different feature channels. The calculation process of the channel attention mechanism can be represented as follows: in, For the input feature map ( For the number of channels, (for sequence length) This represents the channel statistics obtained after global average pooling. and These are the 1×1 convolutional layer weight matrices for dimensionality reduction and dimensionality enhancement, respectively. It is the ReLU activation function. It is the Sigmoid activation function. For the generated channel weight coefficients, This represents element-wise multiplication along the channel dimension. The weighted output features; Through the above operations, the input features are first extracted with global statistical information through global pooling, then channel weight coefficients are generated by the feature transformation unit, and after being normalized by the Sigmoid activation function, they are applied to the original feature channels, thereby strengthening the effective features related to the preamble structure and synchronization decision, and suppressing redundant or noisy features. Through the synergistic effect of the above residual connections and channel attention, the network can better learn the key features related to the frame start position, and improve the synchronization reliability under multipath fading and hardware impairment conditions.

[0013] Step 5 specifically involves: To simultaneously improve the accuracy and stability of frame start position detection, this method employs a multi-task output structure at the network output. This structure uses the weighted output features from the residual attention module. As shared input features, it includes two parallel output branches, used for the detection decision of the frame start position and the position estimation, respectively: The first output branch is a detection decision branch, used for detecting and deciding the start of the frame; this branch calculates the weighted output features from the input. One-dimensional convolutional feature mapping is performed, and the output is constrained by a normalized activation function (such as the sigmoid function) to obtain the detection confidence sequence corresponding to each time position, denoted as the peak detection output. This is used to characterize the probability that the location is the start point of a frame, thereby achieving reliable detection of the start point of a frame. Figure 1 The detection and decision formula is output in the middle; The second output branch is a continuous estimation branch, used for continuous estimation of the frame start position; this branch also receives the weighted output features. The location confidence curve that varies over time is generated through convolutional network layers, and this curve is denoted as the location regression output. A smooth model of the frame start position is performed. This continuous estimation branch can refine the detection results by utilizing the continuity information between adjacent time points, thereby improving the accuracy of the frame start position estimation. Through the collaborative decision of the outputs of the two branches, high-precision synchronous positioning of the start position of the wireless radio frequency received signal frame is finally achieved.

[0014] The second output branch can utilize the continuity information between adjacent time points to refine the detection results, thereby improving the accuracy of frame start position estimation. Through the collaborative work of the above multiple output branches, the network can accurately identify abrupt peaks in the received signal and obtain smooth and stable frame start position estimation results, thus improving synchronization performance and model generalization ability under complex channel conditions.

[0015] Step 6 specifically involves: To adapt to the joint optimization requirements of multi-task output structures, this method employs a combined loss function during training to apply weighted constraints to different output branches. This overall loss function is also the core optimization objective for evaluating the feature extraction performance of preceding steps: steps 1 to 5 of the preceding steps are responsible for multi-scale extraction and attention enhancement of the received signal, ultimately predicting the start position of the output frame; the combined loss function of step 6 is responsible for calculating the error between these predictions and the true labels. During network training, this error is backpropagated to each preceding feature extraction module through the backpropagation algorithm, guiding the iterative update of their convolutional kernel parameters and attention weights. Through this joint constraint, the overall loss function is specifically expressed as: in, To detect the decision loss term, which is used to constrain the accuracy of the frame start position detection results; This is the location estimation loss term, used to measure the error between the output location confidence curve and the target location distribution; This is a positioning penalty term used to further enhance the positioning accuracy of the frame start position.

[0016] The detection decision loss term employs a weighted classification loss function to mitigate the imbalance in positive and negative sample distribution during frame start-of-frame location detection. The location estimation loss term uses mean squared error to constrain the smoothness and consistency of the regression output. The localization penalty term strengthens the penalty for frame start-of-frame deviation. The weight coefficients of each loss term are set according to actual training requirements. Through the design of the above combined loss function, joint optimization of multi-task outputs is achieved, thereby improving the synchronization accuracy and robustness of the model under complex noise, multipath, and frequency offset conditions.

[0017] A physical layer frame synchronization system based on multi-scale feature extraction and residual attention mechanism includes a signal preprocessing module, a multi-scale feature extraction module, a preamble coding module, a residual attention module, a multi-task decision module, and a training optimization module. Signal preprocessing module: This module is used to perform step 1. It receives the radio frequency signal, performs down-conversion, analog-to-digital conversion, and root-raised cosine filtering. It then decomposes the baseband signal into in-phase and quadrature components and stacks them in time to form a two-dimensional input feature matrix. At the same time, it performs the same shaping filtering on the known preamble sequence to generate a two-dimensional preamble feature matrix, providing standardized input for subsequent networks. Multi-scale feature extraction module: used to perform step 2. This module contains three parallel processing branches: the detail branch extracts short-term synchronous features using small-scale convolutional kernels on the original input sequence; the mesoscale branch first downsamples the input by 2x, then extracts features using medium-scale convolutional kernels; the global branch first downsamples the input by 4x, then captures long-term dependent features using large-scale convolutional kernels. The outputs of each branch are upsampled and aligned, then concatenated along the channel dimension to form a multi-scale feature representation. Preamble Encoding Module: This module performs step 3. It maps local preamble sequences of arbitrary length into fixed-dimensional embedded feature vectors using multi-layer one-dimensional convolutions and global average pooling. These vectors are then channel-fused with multi-scale features to inject prior information about the preamble structure into the network. Residual Attention Module: Used to perform step 4. This module achieves direct feature transfer and superposition fusion through residual connections to avoid feature degradation; at the same time, it introduces a channel attention mechanism, uses global average pooling to extract channel statistics, generates channel weights through dimensionality reduction-up transformation and Sigmoid activation, and adaptively weights the feature channels to strengthen effective features related to synchronous decision and suppress redundancy and noise; Multi-task decision module: used to execute step 5. This module sets up two parallel output branches: the first branch is the peak detection branch, which outputs the detection confidence at each time position; the second branch is the position regression branch, which outputs a smooth position confidence curve. The outputs of the two branches are fused to generate the final frame start position probability sequence; Training optimization module: This module is used to perform step 6. During the training phase, it employs a combined loss function. Joint optimization of network parameters, where To detect the loss of the judgment, To estimate the loss for location, To locate the penalty term, achieve collaborative optimization of multi-task outputs, and improve the model's generalization performance under complex channels; Interrelationships between modules: The two-dimensional input features output by the signal preprocessing module are simultaneously fed into the multi-scale feature extraction module and the preamble encoding module; the multi-scale features output by the multi-scale feature extraction module and the preamble features output by the preamble encoding module are fused along the channel dimension and then fed into the residual attention module; the enhanced features output by the residual attention module serve as the input to the multi-task decision module to generate a frame start position probability sequence; the training optimization module performs joint parameter optimization on the above modules (except for signal preprocessing) during the offline phase. After training is completed, the signal preprocessing module, multi-scale feature extraction module, preamble encoding module, residual attention module, and multi-task decision module together constitute the online inference link, realizing end-to-end frame synchronization detection.

[0018] This invention can be applied to the baseband processing link of wireless communication receiving devices such as base stations, terminals, and IoT gateways. Under complex channel conditions with the superposition of multipath fading, carrier frequency offset, low signal-to-noise ratio, and non-ideal RF front-end factors, it can improve the accuracy and robustness of wireless data frame start position detection, thereby enhancing the reliability of subsequent channel estimation, equalization processing, and data demodulation.

[0019] The beneficial effects of this invention are: This invention can be applied to the baseband receiving link of wireless communication devices such as 5G / 6G base stations, IoT gateways, vehicle-to-everything (V2X) terminals, and satellite communication receivers to determine the start position of wireless data frames under complex channel conditions. Addressing the difficulties in frame synchronization caused by factors such as urban building reflections, high-speed movement, low signal-to-noise ratio transmission, carrier frequency offset, phase noise, and IQ imbalance, this invention improves the performance of frame start position detection through the collaborative evolution of multiple technical means. First, preamble sequence encoding (step 3) introduces prior information about the local preamble structure, and a residual attention mechanism (step 4) is used to adaptively weight the feature channels. This collaborative process achieves precise focus at the feature level, enabling the network to effectively strengthen relevant information related to frame synchronization from the source and suppress spurious peaks caused by multipath fading, noise interference, and radio frequency impairments. Second, in the multi-task decision stage (step 5), the peak detection branch integrates a multi-scale parallel feature extraction structure. By processing receptive fields at different scales in parallel, the network can capture both local features at the sampling point level to cope with sudden interference and symbol-level global context features to cope with delay spread, thus achieving multi-dimensional feature representation. Thirdly, through the aforementioned feature enhancement mechanisms of "preamble guidance" and "multi-scale parallel sensing," combined with the continuity and smoothness constraints provided by the position regression branch, high-precision collaborative decision-making on the frame start position is achieved. This method improves the accuracy and robustness of receiver synchronization and positioning in complex channel environments.

[0020] In complex wireless communication scenarios, this invention can still achieve stable frame start position detection within the allowable synchronization tolerance window, and its synchronization success rate is significantly better than traditional cross-correlation detection methods and basic fully convolutional neural network methods. Therefore, this invention can reduce the error propagation caused by frame synchronization errors to subsequent symbol timing recovery, channel estimation, equalization processing, and data demodulation, thereby improving the success rate of wireless data packet reception. It is suitable for practical application scenarios such as weak signal communication, high-speed mobile communication, multipath communication in dense urban areas, and bursty short packet communication in the Internet of Things.

[0021] For various channel conditions, including those with only radio frequency (RF) impairment, only multipath fading, and both multipath and RF impairment, this invention demonstrates consistent and stable performance improvement, avoiding the problem of drastic performance degradation of traditional methods under non-ideal hardware conditions, and enhancing the engineering applicability of the frame synchronization algorithm in practical wireless communication systems.

[0022] This invention adopts a data-driven approach to directly learn frame synchronization features from received signals, eliminating the need to establish precise channel models or radio frequency impairment parameter models. This reduces reliance on prior information, lowers the complexity of receiver algorithm design and parameter configuration, and facilitates deployment in different communication systems and channel environments. Attached Figure Description

[0023] Figure 1 This is a diagram illustrating the overall framework of a wireless frame synchronization method and system based on a multi-scale residual attention network according to the present invention.

[0024] Figure 2 This is a schematic diagram showing the synchronization success rate of the three frame synchronization methods as a function of signal-to-noise ratio under the condition that only radio frequency non-ideal factors exist; Figure 3 This is a schematic diagram showing the synchronization success rate of the three frame synchronization methods as a function of signal-to-noise ratio under the condition that only a multipath channel exists. Figure 4 This diagram illustrates the synchronization success rate of the three frame synchronization methods as a function of the signal-to-noise ratio under the conditions of simultaneous radio frequency non-ideal factors and multipath channels. Detailed Implementation

[0025] The present invention will now be described in further detail with reference to the accompanying drawings.

[0026] In practical wireless communication systems, receivers need to quickly determine the start position of data frames under non-ideal conditions such as noise, multipath fading, carrier frequency offset, phase noise, and IQ imbalance. Inaccurate frame synchronization directly affects subsequent channel estimation, symbol timing recovery, and data demodulation, leading to increased bit error rate or even packet reception failure. Therefore, this invention is not only an improved neural network structure method but can also serve as a frame synchronization unit in the baseband processing module of a wireless communication receiver to improve the packet reception success rate in complex channel environments.

[0027] In this embodiment, the present invention provides a physical layer frame synchronization method based on multi-scale feature extraction and residual attention mechanism; the method mainly includes four stages. The specific correspondence is as follows: Signal reception and preprocessing stage: This stage is used to acquire baseband signals and construct preamble templates.

[0028] Preamble feature extraction stage: This stage is used to encode the locally known preamble into a fixed-dimensional feature vector.

[0029] The frame synchronization feature extraction and decision stage based on neural networks is the core of the method, which includes multi-scale feature extraction, residual attention enhancement, and multi-task joint decision.

[0030] Synchronization position output stage: This stage performs peak detection on the probability sequence of the multi-task output head to determine the frame start position.

[0031] To enable those skilled in the art to more clearly understand the technical solution of the present invention, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings. This embodiment uses a ZC sequence as the preamble and a BPSK modulation signal with a length of 64 symbols as an example for illustration, but the present invention is not limited thereto, and parameters such as the preamble length, modulation method, and symbol length can be adjusted according to the actual application scenario.

[0032] In this embodiment, the present invention provides a physical layer frame synchronization method based on multi-scale feature extraction and residual attention mechanism. The method mainly includes a signal reception and preprocessing stage, a preamble feature extraction stage, a neural network-based frame synchronization feature extraction and decision stage, and a synchronization position output stage.

[0033] In the signal reception and preprocessing stage, the receiver acquires the radio frequency signal containing the preamble and data portion. This signal undergoes down-conversion and analog-to-digital conversion to obtain a baseband complex sampling sequence. Then, this baseband sequence is subjected to root-raised cosine filtering to suppress out-of-band interference and achieve matched filtering. A root-raised cosine filter with a roll-off factor of 0.35, a span of 6 symbols, and 8 sampling points per symbol is used for filtering. The filtered signal is input to the subsequent processing module in the form of two channels, one for the real part and one for the imaginary part, forming a fixed-length input sequence of 3360 sampling points.

[0034] In the preamble feature extraction stage, the known local preamble template is encoded to obtain a high-dimensional feature representation. This local preamble template is then input into the preamble encoding network. This network sequentially passes through multiple one-dimensional convolutional layers, batch normalization layers, and activation function layers for feature extraction. Finally, it is mapped to a fixed-dimensional feature vector through global pooling and fully connected layers. This feature vector represents the structural information of the preamble and can be used to guide subsequent frame synchronization decisions. This method is adaptable to preamble sequences of different lengths.

[0035] In the neural network-based frame synchronization feature extraction and decision stage, the preprocessed received signal is input into the backbone feature extraction network. This network first extracts preliminary spatiotemporal features through multiple one-dimensional convolutional layers, and then enters multiple residual connection blocks. After the residual blocks, channel attention modules and spatial attention modules are applied sequentially to weight and enhance the feature map. Specifically, the channel attention module generates channel weights through global average pooling and multiple convolutions; the spatial attention module generates spatial weights by concatenating the mean and maximum values ​​of the channel dimensions and then performing convolution. Then, the feature vector obtained from the preamble encoding is expanded along the sequence dimension and concatenated with the feature map output by the backbone network, and then fused through a 1×1 convolutional layer to achieve preamble-aware feature enhancement. Next, the fused features are input into the multi-scale peak detection module. This module extracts high-resolution, medium-resolution, and low-resolution features in parallel, upsamples and aligns them, concatenates them, and then fuses them through attention-weighted layers. Finally, it outputs a peak score sequence through a convolutional layer. Simultaneously, a regression path and a peak refinement path are branched from the fused features to generate smooth regression scores and refined peak scores, respectively. Finally, the peak score and refined peak score are concatenated along the channel dimension and then fused by convolution to obtain the final frame start position probability sequence.

[0036] In this embodiment, the parameters of all convolutional layers, batch normalization layers, activation function layers, and attention modules of the aforementioned neural network are obtained through offline training. Supervised learning can be employed during the training phase, using synthetic or experimental datasets containing known frame start positions for optimization. A combined loss function can be used during training, including a classification loss for peak sequences, a mean squared error loss for regression sequences, and a positional bias penalty term. The training process employs the AdamW optimizer, cosine annealing learning rate scheduling, and mixed-precision computing techniques. After training, the model parameters are fixed for online inference.

[0037] Finally, in the synchronization result output and application stage, peak detection is performed on the final probability sequence output by the neural network, and the position of the maximum value is selected as the estimated frame start position. The obtained frame start position is used for subsequent symbol timing recovery, channel estimation, data demodulation, and other modules to achieve complete data packet reception.

[0038] like Figure 1The diagram shows the overall architecture of the frame synchronization network of this invention. The input layer contains two inputs: one is a known preamble sequence, serving as prior knowledge; the other is the actual IQ signal received by the receiver, which contains noise, multipath propagation, frequency offset, and other impairments, serving as the data to be synchronized. The preamble input module extracts basic features through convolution, batch normalization, and ReLU activation, compresses the feature dimension through global pooling, and then generates a fixed-dimensional preamble feature vector through feature projection, completing the encoding of prior knowledge. The temporal feature extraction module performs multi-scale feature extraction on the received data. It performs preliminary feature extraction through convolution and Dropout, and then uses four residual blocks combined with dilated convolutions with dilation rates of 1, 2, 4, and 8 to simultaneously capture local details and global contextual features of the signal. The attention-enhanced fusion module fuses leading features with data features: the channel attention branch generates channel weights through average pooling, convolution, and sigmoid, highlighting effective feature channels related to synchronization; the spatial attention branch generates spatial weights through concatenation of max pooling and average pooling, followed by convolution and sigmoid, focusing on key temporal positions in the signal containing leading information; the two attention branches weight the features separately, then fuse them by adding the residuals, and finally complete feature integration through 1×1 convolution. The multi-task output layer contains two parallel branches: the peak detection head extracts high, medium, and low-scale features from the fused features, and outputs peak prediction output after attention fusion. It is used to detect abrupt peaks at the start of a frame; the regression output head outputs regression prediction output through convolution, batch normalization, ReLU, and sigmoid structures. The continuous estimate of the frame start position is given in the form of a probability distribution; the optimal estimate of the frame start position is obtained by fusing the outputs of the two branches.

[0039] like Figure 2 As shown, under conditions where only radio frequency non-ideal factors exist (including carrier frequency offset, phase noise, IQ imbalance, etc.): the synchronization success rate of the traditional cross-correlation method is at most about 40%; the FCN method almost completely fails when the signal-to-noise ratio is below -5 dB (success rate is only 7.2% at -10 dB), and gradually improves after the signal-to-noise ratio is above 0 dB, reaching a maximum of 90.5% (20 dB); the method of the present invention achieves a success rate of 33.9% at a signal-to-noise ratio of -10 dB, 55.5% at -6 dB, 64.2% at -4 dB, 73.9% at -2 dB, 83.5% at 0 dB, 97.3% at 4 dB, and 99.2% at 6 dB, and the success rate stabilizes at 100% when the signal-to-noise ratio is above 8 dB. Experimental results show that the method of the present invention outperforms the traditional cross-correlation method and the FCN method in the range of -9 dB to 20 dB.

[0040] like Figure 3As shown, under the condition that only multipath channels exist (including multipath delay spread and inter-symbol interference), the synchronization success rate of the traditional cross-correlation method is up to about 78%; the success rate of the FCN method is less than 20% when the signal-to-noise ratio is below 0 dB, and gradually increases after the signal-to-noise ratio is above 0 dB, reaching a maximum of about 96%; the success rate of the method of this invention is 61.1% when the signal-to-noise ratio is 0 dB, and stabilizes above 96% when the signal-to-noise ratio is above 12 dB.

[0041] like Figure 4 As shown, under conditions of simultaneous radio frequency non-ideal factors and multipath channels, the highest synchronization success rate of the traditional cross-correlation method is approximately 30%; the highest success rate of the FCN method is 96.9%; and the method of this invention remains stable at over 98% when the signal-to-noise ratio is higher than 12 dB. These results demonstrate that the method of this invention, through preamble prior guidance and attention mechanisms, can still achieve highly reliable synchronization under strong interference environments.

[0042] Specific application examples: This invention can be applied to cellular mobile communication systems in urban environments (such as 5G / 6G base stations or IoT gateway receivers). In actual deployment scenarios, wireless signal transmission faces severe multipath fading (caused by building reflection and scattering), carrier frequency offset (caused by high-speed terminal movement or crystal oscillator errors), and non-ideal factors in the radio frequency front-end (such as IQ imbalance, phase noise, etc.). The coupled superposition of these channel impairments significantly degrades the performance of traditional cross-correlation-based frame synchronization methods, severely affecting the reliability of data transmission.

[0043] In this embodiment, the communication receiver employs the frame synchronization method based on multi-scale feature extraction and residual attention mechanism described in this invention. The specific implementation process is as follows: Figure 1 The network architecture shown is completely consistent. Follow these steps in sequence: Step 1 (Signal Reception and Preprocessing): The receiver performs down-conversion, analog-to-digital conversion, and root-raised cosine filtering on the RF signal to obtain the baseband complex sampling sequence, and decomposes it into I / Q dual-channel input feature matrices. (correspond Figure 1 (Input data); simultaneously, the locally known ZC sequence preamble is subjected to the same shaping filter to generate a two-dimensional preamble feature matrix. (correspond Figure 1 (Preamble input) Step 2 (Multi-scale Feature Extraction): Input Features It is sent to the multi-scale feature extraction module (corresponding to) Figure 1The "Temporal Feature Extraction Module" contains three parallel branches: the detail branch directly processes the original sequence using a small-scale convolutional kernel (3×1, dilation rate 2) to extract short-term synchronous features; the mesoscale branch first downsamples by 2x and then uses a 5×1 convolutional kernel to balance details and receptive field; the global branch first downsamples by 4x and then uses a 7×1 convolutional kernel to capture long-term dependent features. The outputs of each branch are upsampled, aligned, and then concatenated to form a multi-scale feature representation. ; Step 3 (Preamble Sequence Encoding): Local Preamble Features It is fed into the preamble coding network (corresponding to) Figure 1 The network consists of a preamble input module ("preamble input module"). It comprises multiple layers of one-dimensional convolutions and global average pooling, mapping a preamble sequence of arbitrary length to a fixed-dimensional embedded feature vector. This is used for subsequent fusion guidance; Step 4 (Residual Attention Fusion): Integrate multi-scale features With the expanded leading feature vector After concatenation along the channel dimension, input the residual attention module (corresponding to) Figure 1 The module first extracts deep features from multiple residual blocks (dilation rates of 1, 2, 4, and 8), and then weights them through channel attention (global average pooling + 1×1 convolution + sigmoid) and spatial attention (concatenation of channel mean / maximum values ​​+ 7×1 convolution + sigmoid) to obtain enhanced features. ; Step 5 (Multi-task output): Enhance features Sent to the multi-task output layer (corresponding to) Figure 1 The "peak detection head" and "regression output head" are used in the middle. The peak detection branch outputs a confidence sequence of the frame start position through a multi-scale detector. The regression branch outputs a smooth location confidence curve through the convolutional layer. The results from the two branches are fused using a 1×1 convolution to obtain the final frame start position probability sequence; Step 6 (Training Optimization): During the offline training phase, use the combined loss function. Joint optimization of network parameters, where For peak detection , For the MSE loss of the regression branch, To locate the penalty term. After training, the model parameters are stored in the receiver baseband processing unit for online inference.

[0044] Simulation verification: In a typical urban microcellular channel environment (multipath delay spread ≤ 2 μs, carrier frequency offset ± 2 kHz, IQ imbalance amplitude deviation ± 2%), the method described in this invention achieves an average positioning error of 5.1 sampling points at a signal-to-noise ratio (SNR) of 0 dB, a reduction of approximately 2.3 sampling points compared to the 7.4 sampling points of the traditional cross-correlation method. Within a tolerance window of ± 2 sampling points, the synchronization success rate is 57.0%, an improvement of approximately 35.0 percentage points compared to the 22.0% of the FCN baseline method. Furthermore, under various channel conditions—including only RF impairment, only multipath fading, and both—this invention demonstrates consistent and stable synchronization performance improvement, maintaining a synchronization success rate of approximately 10.7% even in extremely low SNR scenarios of -10 dB.

[0045] This invention does not rely on accurate channel models or impairment parameter estimation. It only requires integrating a pre-trained model into the receiver baseband processing unit. It has the advantages of simple engineering implementation and strong adaptability, and can be widely used in frame synchronization modules of wireless communication devices such as base stations, terminals, and IoT gateways.

Claims

1. A physical layer frame synchronization method based on multi-scale feature extraction and residual attention mechanism, characterized in that, It includes the following steps: Step 1: Convert the original IQ signal and local preamble into a normalized tensor that can be processed by a neural network to obtain the input sequence. ; Step 2: Based on the input sequence By employing a parallel multi-branch convolutional structure, both local detail features and global contextual features of the received baseband signal are captured simultaneously, thus obtaining multi-scale feature representations. ; Step 3: Encode the structural information of the local preamble into a fixed-dimensional feature vector. This guides the network to focus on the structural features of the target leader; Step 4: Using the aforementioned leading feature vector With the multi-scale feature representation Feature map after fusion As input, residual connections are used to avoid feature degradation, and a channel attention mechanism is employed to highlight effective feature channels related to synchronization decisions, suppress redundant information and noise features. In complex scenarios with multiple impairments such as multipath interference, frequency offset, and noise, this helps the network focus on key information and obtain weighted output features. ; Step 5: Using the shared fusion features As a parallel input, it works in conjunction with an independent peak detection head and a regression output head; the peak detection head... Perform high, medium, and low-scale feature extraction and attention fusion to output peak prediction. The method detects abrupt peaks at the start of a frame; the regression output head is processed through a convolution, batch normalization, ReLU, and sigmoid cascade structure. Output regression prediction output The frame start position is continuously estimated in the form of a probability distribution; finally, the peak prediction is output. With the regression prediction output A fusion decision is made to obtain the optimal estimate of the frame start position. ; Step 6: By combining loss functions, the detection accuracy, estimation precision, and positioning bias are constrained simultaneously during the training phase. This ensures that the model can learn the optimal feature representation and decision strategy when facing complex and unknown channel environments. This enables highly reliable positioning of the start position of wireless data packet frames in complex wireless communication environments and provides an accurate synchronization reference for symbol timing recovery, channel estimation, equalization processing, and data demodulation in the communication baseband receiver.

2. The physical layer frame synchronization method based on multi-scale feature extraction and residual attention mechanism according to claim 1, characterized in that, Step 1 specifically involves: The receiver acquires the radio frequency received signal containing the local preamble and data portion, and obtains the baseband complex sampling sequence after down-conversion and analog-to-digital conversion. The baseband complex sampling is then subjected to root-raised cosine filtering to obtain the filtered baseband signal. Finally, a basic feature map is constructed through a basic feature extraction layer. ; The filtered baseband signal is decomposed into in-phase components (I) and quadrature components (Q), and stacked along the time dimension to form a two-dimensional input feature matrix of size (2, L), which serves as the input to the subsequent neural network, where L is the sequence length. The known preamble sequence is subjected to the same shaping filtering process as the transmitting end, that is, first root-raised cosine transmission filtering is performed, and then root-raised cosine reception filtering is performed to obtain a preamble template that is homomorphic to the received signal; the real and imaginary parts of the obtained preamble template are decomposed to form a two-dimensional preamble feature matrix consistent with the input feature dimension, which is used to characterize the structural information of the preamble code; The preamble template can be adapted to preamble sequences of any length and modulation scheme; The baseband feature matrix output from step 1 is denoted as the input sequence. ,in Let the sequence length be denoted as the constructed two-dimensional preceding feature matrix as the prior template. ,in This is the length of the leading template.

3. The physical layer frame synchronization method based on multi-scale feature extraction and residual attention mechanism according to claim 2, characterized in that, Step 2 specifically involves: For baseband signals containing physical channel impairments, a parallel multi-branch convolutional feature extraction structure is constructed to capture multi-scale temporal features in the received signal. This multi-branch convolutional feature extraction structure includes three processing paths with different physical temporal resolutions: First branch: Feature extraction is performed directly on the input sequence using small-scale convolutional kernels, based on the baseband sampling period. To preserve short-term transient information of the signal for the highest physical time resolution, it is used to capture key local peak features and fine carrier phase changes during synchronization; The second branch: First, the input sequence is downsampled by a factor of 2 to amplify the physical time resolution to [value missing]. Then, medium-scale convolution kernels are used to extract features, which expands the receptive field while taking into account detailed information, in order to cover and extract the delay spread features caused by multipath effects in the wireless channel. The third branch: downsamples the input sequence by a factor of 4, increasing the physical time resolution to [value missing]. Then, large-scale convolution kernels are used to extract long-term dependent features to capture the overall macroscopic energy envelope structure of the signal after experiencing long-term channel fading and carrier frequency offset. The output features of each branch are interpolated to restore the original time length and then spliced ​​together along the feature channel dimension to form a multi-scale feature representation.

4. The physical layer frame synchronization method based on multi-scale feature extraction and residual attention mechanism according to claim 3, characterized in that, This process can be represented as: in, The input sequence is the output of step 1. , , These represent the convolution operations in the three branches, respectively. Indicates k-fold downsampling. Indicates k times upsampling, This indicates concatenation of channel dimensions.

5. The physical layer frame synchronization method based on multi-scale feature extraction and residual attention mechanism according to claim 4, characterized in that, Step 3 specifically involves: The leader sequence, as prior information in the synchronization process, contains stable and distinguishable temporal structural features. A leader sequence encoder is set up to map leader sequences of different lengths into embedded feature vectors of fixed dimensions, which are used to participate in subsequent synchronization feature fusion and decision-making. The leader sequence encoder consists of a convolutional feature extraction layer and a temporal dimension aggregation layer. It extracts hierarchical features of the leader sequence through progressive convolution operations and aggregates the temporal dimension features. Its expression is as follows: = in To represent the embedded features of the leading sequence, the temporal aggregation unit uses global pooling to compress the leading sequence features, thereby obtaining a compact representation that includes key temporal pattern information. The generated leading features and the features output by the multi-scale feature extraction module The feature map is obtained by fusing information across feature channels through a feature transformation unit, and the fusion is performed at the channel dimension. .

6. The physical layer frame synchronization method based on multi-scale feature extraction and residual attention mechanism according to claim 5, characterized in that, Step 4 specifically involves: A residual attention module is introduced during the feature processing to enhance the fusion features of the leading perception output in step 3. The residual attention module achieves direct transmission and superposition fusion of features through the residual connection structure. The residual module adopts a pre-activation structure, that is, the input features are first subjected to nonlinear transformation, and then fused with the original features through the residual connection. At the same time, a channel attention mechanism is introduced to adaptively adjust the importance of different feature channels. The computational process of the channel attention mechanism is represented as follows: in, The fusion features of the leading perception output from step 3. For the number of channels, For sequence length, This represents the channel statistics obtained after global average pooling. and These are the 1×1 convolutional layer weight matrices for dimensionality reduction and dimensionality enhancement, respectively. It is the ReLU activation function. It is the Sigmoid activation function. For the generated channel weight coefficients, This represents element-wise multiplication along the channel dimension. The weighted output features.

7. The physical layer frame synchronization method based on multi-scale feature extraction and residual attention mechanism according to claim 6, characterized in that, Step 5 specifically involves: The weighted output features of the residual attention module As shared input features, it includes two parallel output branches, used for the detection decision of the frame start position and the position estimation, respectively: The first output branch is a detection decision branch, used for detecting and deciding the start of the frame; this branch calculates the weighted output features from the input. One-dimensional convolutional feature mapping is performed, and the output is constrained by a normalized activation function to obtain the detection confidence sequence corresponding to each time position, denoted as the peak detection output. ; The second output branch is a continuous estimation branch, used for continuous estimation of the frame start position; this branch also receives the weighted output features. The location confidence curve that varies over time is generated through convolutional network layers, and this curve is denoted as the location regression output. Smooth modeling is performed on the frame start position.

8. The physical layer frame synchronization method based on multi-scale feature extraction and residual attention mechanism according to claim 7, characterized in that, Step 6 specifically involves: During training, a combined loss function is used to apply weighted constraints to different output branches. The overall loss function is expressed as follows: in, To detect the decision loss term, which is used to constrain the accuracy of the frame start position detection results; This is the location estimation loss term, used to measure the error between the output location confidence curve and the target location distribution; This is a positioning penalty term used to further enhance the positioning accuracy of the frame start position.

9. A physical layer frame synchronization system based on multi-scale feature extraction and residual attention mechanism for implementing the method of any one of claims 1-8, characterized in that, It includes a signal preprocessing module, a multi-scale feature extraction module, a leader coding module, a residual attention module, a multi-task decision module, and a training optimization module; The signal preprocessing module receives the radio frequency signal, performs down-conversion, analog-to-digital conversion, and root-raised cosine filtering, decomposes the baseband signal into in-phase and quadrature components, and stacks them in the time dimension to form a two-dimensional input feature matrix; at the same time, the known preamble sequence is subjected to the same shaping filtering process to generate a two-dimensional preamble feature matrix, which provides standardized input for subsequent networks; The multi-scale feature extraction module contains three parallel processing branches: the detail branch extracts short-term synchronous features using small-scale convolutional kernels on the original input sequence; The mesoscale branch first downsamples the input by a factor of 2, and then uses a mesoscale convolution kernel to extract features; The global branch first downsamples the input by a factor of 4, and then uses a large-scale convolutional kernel to capture long-term dependent features; The outputs of each branch are upsampled and aligned, then concatenated along the channel dimension to form a multi-scale feature representation. The preamble encoding module maps local preamble sequences of arbitrary length into fixed-dimensional embedded feature vectors through multi-layer one-dimensional convolution and global average pooling. These vectors are then channel-fused with multi-scale features to inject prior information about the preamble structure into the network. The residual attention module achieves direct feature propagation and superposition fusion through residual connections, avoiding feature degradation. At the same time, it introduces a channel attention mechanism, which uses global average pooling to extract channel statistics, generates channel weights through dimensionality reduction-dimensionality increase transformation and Sigmoid activation, and adaptively weights the feature channels to strengthen effective features related to synchronous decision and suppress redundancy and noise. The multi-task decision module has two parallel output branches: the first branch is the peak detection branch, which outputs the detection confidence at each time position; the second branch is the position regression branch, which outputs a smooth position confidence curve; the outputs of the two branches are fused to generate the final frame start position probability sequence. The training optimization module uses a combined loss function during the training phase. Joint optimization of network parameters, where To detect the loss of the judgment, To estimate the loss for location, To locate the penalty item.