Carrier center frequency estimation method based on adaptive learning rate fine tuning

Through the carrier center frequency estimation method with fine-tuning adaptive learning rate, combined with the symmetry characteristics of Mask R-CNN and STFT image signal segments, high-level features are frozen, and instance segmentation network is trained using the adaptive learning rate optimizer strategy, which solves the problem of insufficient accuracy and speed of carrier center frequency estimation in satellite communications, and realizes efficient frequency estimation.

CN120281620APending Publication Date: 2025-07-08BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510417557.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing carrier center frequency estimation methods have problems with insufficient estimation accuracy and speed in satellite communications, especially in low signal-to-noise ratio and variable wireless channel environments, which are difficult to achieve high-time efficiency and low complexity frequency estimation.

Method used

Using an adaptive learning rate fine-tuning method, combining the symmetry characteristics of Mask R-CNN and STFT image signal segments, high-level features are frozen, and instance segmentation networks are trained using the adaptive learning rate optimizer strategy to improve the training efficiency and feature extraction capabilities of deep networks, and pixel-level mask segmentation is performed through instance segmentation networks to reduce dependence on complex time-frequency domain analysis.

Benefits of technology

Fast and accurate carrier center frequency estimation is achieved in a low signal-to-noise environment, which improves estimation accuracy and reduces memory overhead, and adapts to the high refresh rate and high-precision signal monitoring requirements of satellite communications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281620A_ABST
    Figure CN120281620A_ABST
Patent Text Reader

Abstract

The invention discloses a carrier center frequency estimation method based on adaptive learning rate fine tuning, and belongs to the technical field of carrier center frequency estimation in satellite communication. The method is based on an instance segmentation network Mask R-CNN, combines symmetry features of STFT image signal segments, finely adjusts bottom layer parts in an instance segmentation backbone network and a feature extraction network, freezes high-level features, adopts an adaptive learning rate optimizer strategy in a training process to ensure that each layer is updated and kept stable, and finally forms a lightweight instance segmentation network. The training efficiency and the key feature extraction capability of the deep network are improved, and rapid and accurate carrier center frequency estimation in a low signal-to-noise ratio environment can be realized. Training can be rapidly completed according to different communication scenes so as to adapt to the environment; the estimation precision of the signal carrier center frequency can be improved under the condition of low signal-to-noise ratio; the memory overhead of carrier center frequency estimation can be effectively reduced; and the dependence on artificial feature extraction and complex time-frequency domain joint analysis is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a carrier center frequency estimation method based on adaptive learning rate fine-tuning, belonging to the technical field of carrier center frequency estimation in satellite communication. Background Art

[0002] With the evolution of 5G communication networks towards massive multiple-input multiple-output antenna technology (Multiple-Input, Multiple-Output, MIMO) and space-ground integrated scenarios, low bit error and high timeliness have become the core challenges in the design of communication systems. Especially in the non-terrestrial network defined by 3GPP Release 17, the severe Doppler frequency shift between satellites and between satellites and ground terminals requires real-time high-precision frequency tracking. At the same time, in scenarios such as dynamic spectrum sharing and ultra-reliable low-latency communication, the accurate estimation of the carrier center frequency directly affects beamforming efficiency, multi-user interference suppression ability, and end-to-end transmission reliability.

[0003] Currently, the main carrier center frequency estimation methods include the following two categories:

[0004] Traditional carrier center frequency estimation methods mainly rely on spectrum analysis techniques, and analyze the spectral characteristics of signals through methods such as Fast Fourier Transformer (FFT) and Maximum Likelihood Estimation (MLE). The above methods can relatively accurately estimate the frequency characteristics of high signal-to-noise ratio signals. However, traditional methods usually require stable signals and a narrow spectral range. In actual satellite communication systems, with the increase in spectrum and the complexity of the environment, the estimation accuracy of the above methods gradually decreases, and it is difficult to cope with the changing wireless channel environment.

[0005] Band estimation methods based on intelligent algorithms utilize modern machine learning techniques such as deep learning and convolutional neural networks to comprehensively extract complex features in signals and achieve high-precision carrier center frequency estimation. For satellite communication systems, under the influence of Doppler frequency shift, the carrier center frequency changes frequently, and strong timeliness is required for carrier center frequency estimation. At the same time, in practical applications, since satellite communication systems include satellites with various orbital altitudes and ground base stations, the usage scenarios need to be frequently switched. The training complexity of deep learning methods is relatively high, and it takes a long time to re-train in large batches. At the same time, the spectral characteristics of signals are blurred under low signal-to-noise ratio conditions, which significantly affects the estimation accuracy.

[0006] Therefore, in the context of the real-time requirements of satellite communication and the characteristics of low signal-to-noise ratio in existing carrier center frequency estimation technologies, there are obvious deficiencies in estimation accuracy and estimation speed. It is necessary to provide an advanced method to improve the estimation accuracy of carrier center frequency estimation technology in the field of satellite communication in non-terrestrial networks, shorten the signal observation time, and at the same time alleviate the problem of resource occupation caused by high computational complexity and large amount of training data in deep networks. Summary of the Invention

[0007] Aiming at the deficiencies of the existing technology and the transmission characteristics of satellite communication under high timeliness requirements and strong noise environments, the main purpose of the present invention is to provide a carrier center frequency estimation method based on adaptive learning rate fine-tuning. Based on the instance segmentation network Mask R-CNN, combined with the symmetry characteristics of STFT image signal segments, fine-tune the underlying parts of the instance segmentation backbone network ResNet50 and the feature extraction network (Feature Pyramid Networks, FPN), freeze the high-level features, and adopt an adaptive learning rate optimizer strategy during training to ensure the stability of layer updates. Finally, a lightweight instance segmentation network is formed to improve the training efficiency of the deep network and the key feature extraction ability, and can achieve fast and accurate carrier center frequency estimation in a low signal-to-noise ratio environment.

[0008] The object of the present invention is achieved by the following technical solutions:

[0009] A carrier center frequency estimation method based on adaptive learning rate fine-tuning of the present invention includes the following steps:

[0010] Step 1: Establish a satellite communication system model and generate signal data samples;

[0011] Among them, the transmitting end sends a modulated random signal, and the channel is affected by noise interference;

[0012] At the receiving end, the signal is sampled to obtain a received signal model of length K, as shown in Equation (1):

[0013]

[0014] where, r k is the symbol sequence of the modulated signal, f c is the carrier center frequency, is the Gaussian additive noise, with a mean of 0 and a variance of is the root raised cosine shaping filter function, as shown in Equation (2):

[0015]

[0016] where, T sis the sampling period, T is the symbol period, and α is the filter roll-off coefficient;

[0017] The signal-to-noise ratio SNR of the signal is shown in Equation (3):

[0018]

[0019] Step 2: Use the Short Time Fourier Transformer (STFT) and the selected symmetric window function to characterize the signal in the time-frequency domain, as shown in Equation (4): as follows:

[0020]

[0021] where is the sliding window function, ζ is the window length, w is the time index, and h is the frequency index;

[0022] The model uses the STFT transform to generate the time-frequency amplitude spectrum of the signal, and converts the time-frequency amplitude spectrum of the signal to the logarithmic scale for processing. The STFT image transform is shown in Equation (5):

[0023]

[0024] where is the amplitude value corresponding to each pixel, determined by the time index w and the frequency index h;

[0025] Finally, the STFT image is formed whose information includes the time axis, i.e., the width W of the image matrix, and the image frequency axis, i.e., the height H of the image matrix, in pixels, as shown in Equation (6):

[0026]

[0027] Label the images in the COCO (Common Objects in Context) manner to form a network training set, validation set, and test set;

[0028] Step 3: In the instance segmentation backbone network, freeze the high-level layers of ResNet50 and FPN, and fine-tune the low-level layers. The generated feature fusion image passes through the Region Proposed Network (RPN) to obtain candidate regions. After RoI Align processing, the classification probability result P, the bounding box regression result Δ, and the mask prediction result M are obtained for updating the network parameters used for training. Specifically, it includes the following sub-steps:

[0029] Step 3.1: Use the STFT image as the input of the backbone network, i.e., F0 = I. In the residual module stage i of ResNet50, the convolutional layer functions are combined and defined as Conv i , and its output size is W i ×H i , where W i and H i are the width and height of the feature image in the stage i layer. The convolutional output of the residual module stage i in the ResNet50 network is shown in Equation (7):

[0030]

[0031] where, is the output function of the max-pooling layer, b i is the bias of the convolutional layer and λ i is the learnable parameter weight, which adapts and changes during the training process;

[0032] After being processed by four residual modules, ResNet50 inputs the feature map to the P5 layer of the FPN network as the input image of the FPN network, i.e., The FPN network propagates the image features from the bottom up and uses the upsampling function to uniformly adjust the dimensions of the feature maps of each layer of the FPN to W j-1 ×H j-1 to obtain the fused feature map vector of the P j layer as shown in Equation (8):

[0033]

[0034] where, The learnable weight parameter ξ i is used to balance the influence of low-level features and high-level features on the fused image;

[0035] The output layer of the FPN fuses the features of the P2 - P5 layers to obtain the final output feature map shown in Equation (9)

[0036]

[0037] where, Adjust the dimension of the feature map of the P5 layer in the FPN to

[0038] During the training process, due to the freezing operation on the high-level part of the backbone network, the parameters λ3, λ4, ξ3, and ξ4 do not change with training after being preset;

[0039] Step 3.2: The RPN network sets K anchor boxes at each pixel point (i, j) on the feature map, and the coordinates of each anchor box defined on the feature map S are E k =(x k , y k , σ k , θ k );

[0040] Among them, x k , y k are the abscissa and ordinate of the upper left corner point of the anchor box respectively, σ k , θ k are the width and height of the anchor box respectively, and k is the index of the anchor box;

[0041] The output of the classification branch is the binary classification result of whether there is a target signal in the anchor box. After preprocessing the feature map S through the 3×3 convolutional layer of the RPN network, the feature extraction result is shown in Equation (10):

[0042] S′ = ReLU(z feat *S + b feat ) (10)

[0043] Among them, ReLU(·) is the rectified linear unit activation function, z feat , b feat are the weights and bias vectors of the feature extraction layer respectively, and * represents the convolution operation;

[0044] After preprocessing, the coordinates of the anchor box defined on S′ are (x′ k , y′ k , σ′ k , θ′ k ), is the binary classification result of each anchor box, indicating the probability that the k-th anchor box contains the target within the (i, j)-th pixel point in the S′ map, as shown in Equation (11):

[0045]

[0046] Among them, is the sigmoid activation function, z cls , b cls are the weights and bias vectors of the classification layer respectively;

[0047] The regression branch is used to predict the bounding box coordinates of the mask corresponding to the signal segment, and gradually approximate the predicted anchor box coordinates to the true bounding box through training to reduce the regression error;

[0048] Based on the regression prediction offset the coordinates of the anchor box are adjusted as shown in Equation (12):

[0049]

[0050] The RPN network performs non-maximum suppression (NMS) on the adjusted anchor boxes, and combines P cls Selects the anchor boxes with higher confidence. After being adjusted by Δ reg N high-quality candidate regions are obtained Each E′ n Represents the coordinates of the candidate region on S′. Map the N high-quality candidate regions To the coordinate system of the feature map S, and obtain the candidate regions The represented coordinates are Where n is the index of the final candidate box, and s is the downsampling ratio of the feature map to the input image;

[0051] Step 3.3: Under the guidance of the candidate regions determined in Step 3.2 Extract N feature image fragments of size W r ×H r From S to form a region matrix The nth candidate region The generated feature fragment As shown in Equation (13):

[0052]

[0053] Where Is the sampling feature value of the pixel point (i, j) in S after the position of the candidate region Is mapped to the feature map S. ω is the weight of bilinear interpolation; The output module of the instance segmentation network will

[0054] Input them into the classification branch, the bounding box regression branch, and the mask prediction branch respectively, as shown in Equation (14):

[0055]

[0056] Where P is the classification probability result, including the classification probabilities of the signal segment, the noise segment, and the confusion segment. Δ is the bounding box regression result, including the horizontal and vertical coordinates of the upper left corner of the bounding box, the length and width of the box. M is the binary mask prediction result indicating whether the pixel point at this position belongs to the classification, Is the combined function of the data dimensionality reduction Flatten and the fully connected layer, Is the function of the convolutional layer. Z and B are the output weights and biases of the corresponding branches respectively;

[0057] Step 4: The network uses the loss function and the predicted Intersection over Union (IoU) as the discrimination criteria, and selects the optimizer strategy of SGD combined with LARS for training. Record the data of the three branches of the above network and determine the training metrics: loss and IoU. Update the network parameters or complete the training according to the training situation, which specifically includes the following sub-steps:

[0058] Step 4.1: The training process of the instance segmentation network involves classification, bounding box regression, and mask prediction tasks. The multi-task learning loss is shown in Equation (15):

[0059]

[0060] Use cross-entropy loss to determine the classification loss As shown in Equation (16):

[0061]

[0062] Among them, is the probability of the model output for each category;

[0063] Use SLL (Smooth L1 Loss) to determine the bounding box regression loss As shown in Equation (17):

[0064]

[0065] Among them, are the coordinates corresponding to the ideal bounding box and the coordinates corresponding to the predicted bounding box respectively;

[0066] Determine the mask loss through binary cross-entropy As shown in Equation (18):

[0067]

[0068] Among them, are the indication results of the real mask and the predicted mask at the same pixel point respectively;

[0069] Step 4.2: Initialize the network weights to ensure that the variances of the inputs of each layer are consistent;

[0070] The training optimizer adopts the SGD (Stochastic Gradient Descent) combined with the LARS (Layer-wise Adaptive Rate Scaling) optimizer strategy to relieve the learning pressure of large-scale deep tasks, integrating the stability of SGD and the advantages of the adaptive learning rate of LARS, enabling the network to still converge quickly when using large-batch deep training. The iteration of the SGD optimizer strategy is shown in Equation (19):

[0071]

[0072] where θ t is the value of the parameter vector at the t-th iteration, is the loss function with respect to the gradient of the parameters, λθ t is the weight decay term used to prevent overfitting, and η is the global learning rate;

[0073] Since the gradient differences of different layers of the deep model are large, and in large-batch training, a fixed learning rate will affect the gradient stability. The layer-wise learning rate dynamic adjustment mechanism of LARS is adopted to improve the training accuracy, enhance the model performance, and improve the convergence speed, as shown in Equation (20):

[0074]

[0075] where is the adaptive learning rate of the parameters of the i-th layer, and ε is a very small constant to prevent the learning rate from dividing by zero error;

[0076] Finally, the optimizer update of SGD combined with LARS is shown in (21):

[0077]

[0078] During the training process, the initial learning rate is set to χ0 in the initial state, and the total number of epochs is set to μ;

[0079] Step 4.3: After each epoch of training is completed, use the validation set to evaluate the performance of the model, determine the IoU and the loss value. The IoU is shown in Equation (22):

[0080]

[0081] where u is the true mask label, is the predicted mask classification;

[0082] When the following situations are encountered during the training process, the training will end:

[0083] 1. When the number of training rounds reaches the preset value μ, stop the training and check the parameters. Fine-tune the network parameters or increase the number of training rounds as needed, and then retrain.

[0084] 2. When the loss value or IoU value of the validation set has not decreased or the fluctuation has not exceeded the critical value for μ′ consecutive rounds, terminate the training in advance to avoid overfitting. Use the test set to test the output model and define it as the optimal network.

[0085] It further includes Step Five: Deploy the optimal instance segmentation network obtained in Step Four to the satellite communication signal analysis device. Generate the STFT image of the received signal through Step Two, perform mask prediction using Step Three, extract the pixel-level segmentation mask of the signal instance in the time-frequency image, and apply the optimal instance segmentation network trained in Step Four to estimate and track the carrier center frequency situation of the satellite communication signal. This can enhance the ability of the instance segmentation network to extract important image features, selectively ignore unimportant features, introduce an adaptive learning rate during the training process, and achieve high-precision and fast carrier center frequency estimation in the satellite communication environment with rapid Doppler frequency shift and strong noise, meeting the core requirements of the satellite communication network for high-precision and low-latency signal tracking.

[0086] Beneficial effects:

[0087] 1. For a carrier center frequency estimation method based on adaptive learning rate fine-tuning according to the present invention, an instance segmentation network is used to perform pixel-level mask segmentation on the signal STFT image. The carrier center frequency is estimated based on the center point of the frequency domain distribution area of the segmentation mask, that is, the midpoint of the signal segment mask in the time-frequency diagram, reducing the dependence on artificial feature extraction and complex time-frequency domain joint analysis.

[0088] 2. For a carrier center frequency estimation method based on adaptive learning rate fine-tuning according to the present invention, through the optimizer strategy of SGD combined with LARS, the learning rate is automatically adjusted according to the parameters of each layer and the norm of the gradient, improving the transfer efficiency of target features. The model convergence speed is improved through the adaptive learning rate, while maintaining the signal segmentation performance, and it can quickly complete the training according to different communication scenarios to adapt to the environment.

[0089] 3. For a carrier center frequency estimation method based on adaptive learning rate fine-tuning according to the present invention, by freezing some high-level parameters, the problems of complex training of deep network parameters and slow gradient flow are alleviated, enhancing the model's ability to segment the signal contour, and being able to improve the estimation accuracy of the signal carrier center frequency under low signal-to-noise ratio conditions.

[0090] 4. A carrier center frequency estimation method based on adaptive learning rate fine-tuning according to the present invention can deploy the optimal instance segmentation network to edge devices through a lightweight model, effectively reducing the memory overhead of carrier center frequency estimation while ensuring a millisecond-level response speed, and adapting to the deployment requirements of high refresh rate and high-precision signal monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] Figure 1 It is a flowchart of a carrier center frequency estimation method based on adaptive learning rate fine-tuning according to the present invention;

[0092] Figure 2 It is a schematic diagram of the network structure in the embodiment;

[0093] Figure 3 It is a flowchart of network training in the embodiment;

[0094] Figure 4 It is a curve of the change of IoU in each training round of the network training process in the embodiment;

[0095] Figure 5 It is the test result of the root mean square relative error (RMSRE) of carrier center frequency estimation in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0096] In order to better illustrate the purpose and advantages of the present invention, the content of the invention will be further described below with reference to the drawings and examples.

[0097] Example 1:

[0098] In this embodiment, a random modulation signal is generated in the QPSK modulation mode, and the signal is sampled at a sampling rate of 10 GHz at the receiving end and subjected to STFT transformation to construct a network training set, a validation set and a test set, where the carrier center frequency f c is from 1 GHz to 4 GHz, the signal-to-noise ratio SNR range of Gaussian channel transmission is from -20 dB to 0 dB, and the symbol rate 1 / T range is from 1 GBaud to 2 GBaud. Taking the above configuration as an example, a carrier center frequency estimation method based on adaptive learning rate fine-tuning according to the present invention is applied for carrier center frequency estimation, as Figure 1 shown, including the following steps:

[0099] Step 1: Establish a satellite communication system model and generate signal data samples;

[0100] Wherein the transmitting end transmits a modulated random signal, and the channel is affected by noise interference;

[0101] At the receiving end, the signal is sampled to obtain a received signal model of length 10,000, as shown in Equation (1):

[0102]

[0103] where r k = exp(j2πm k / 4), m k ∈ {0, 1, 2, 3} is the symbol sequence of the QPSK modulation signal, is additive white Gaussian noise with a mean of 0 and a variance of is the root raised cosine shaping filter function, as shown in Equation (2):

[0104]

[0105] where T s is the sampling period, T is the symbol period, and α is the filter roll-off factor;

[0106] The signal-to-noise ratio SNR of the signal is as shown in Equation (3):

[0107]

[0108] Step 2: Perform STFT using a 256-point Hann window function. The selected window function characterizes the signal in the time-frequency domain, as shown in Equation (4):

[0109]

[0110] where is the sliding window function, w is the time index, and h is the frequency index;

[0111] The model uses the STFT transform to generate the time-frequency amplitude spectrum of the signal and converts the time-frequency amplitude spectrum to a logarithmic scale for processing. The STFT image transform is as shown in Equation (5):

[0112]

[0113] where is the amplitude value corresponding to each pixel, determined by the time index w and the frequency index h;

[0114] Finally, an STFT image is formed. Its information includes the time axis, i.e., the width of the image matrix is 512 pixels, and the image frequency axis, i.e., the height of the image matrix is 384 pixels, as shown in Equation (6):

[0115]

[0116] Perform COCO annotation on the images to form a network training set, validation set, and test set. Finally, a training set and validation set containing 4,000 images and a test set of 1,000 images are formed, with the coverage range including the signal-to-noise ratio SNR and carrier center frequency f proposed in this embodiment. c within the range of;

[0117] Step 3. According to Figure 2 the network structure shown, in this embodiment, in the instance segmentation backbone network, freezing operations are performed on the high-level layers of ResNet50 and FPN, and fine-tuning operations are performed on the low-level layers. The generated feature fusion image passes through RPN to obtain candidate regions. After RoI Align processing, classification probability results P, bounding box regression results Δ, and mask prediction results M are obtained for updating the network parameters used for training, specifically including the following sub-steps:

[0118] Step 3.1: As Figure 2 shown in the structure of the backbone network in, in this embodiment, the STFT image is used as the input of the backbone network, that is, F0 = I. In the residual module stage i of ResNet50, the convolutional layer functions of each layer are combined and defined as Conv i , and its output size is W i ×H i , where W i , H i are the width and height of the feature image in the stage i layer, and the specific sizes are shown in Table 1:

[0119] Table 1 Output image sizes of each residual module of ResNet50

[0120] Residual module <![CDATA[W i > <![CDATA[H i > Stage 1 128 pixels 96 pixels Stage 2 64 pixels 48 pixels Stage 3 32 pixels 24 pixels Stage 4 16 pixels 12 pixels

[0121] The convolutional output of the residual module stage i in the ResNet50 network is shown in Equation (7):

[0122]

[0123] where, is the output function of the max pooling layer, b i is the bias of the convolutional layer and the parameter λ i is the learnable parameter weight, which changes adaptively during the training process;

[0124] After being processed by four residual modules, ResNet50 inputs the feature map to the P5 layer of the FPN network as the input image of the FPN network, that is, The FPN network propagates the image features from the bottom up and uses the upsampling function to uniformly adjust the dimensions of the feature maps of each layer of the FPN to Wj-1 ×H j-1 , obtain P j The feature map vector of the layer fusion is shown in Equation (8):

[0125]

[0126] Among them, The learnable weight parameter ξ i is used to balance the influence of the low-level features and the high-level features on the fused image;

[0127] FPN fuses the features of layers P2 to P5 to obtain the final output feature map shown in Equation (9)

[0128]

[0129] Among them, Adjust the dimension of the feature map of layer P5 in FPN to 128 pixels × 96 pixels;

[0130] During the training process, since the freezing operation is adopted for the high-level part of the backbone network, the parameters λ3, λ4, ξ3, and ξ4 do not change with the training after being preset;

[0131] Step 3.2: As shown in the structure of the regional candidate network in Figure 2 , the RPN network sets 1000 anchor boxes at each pixel point (i, j) on the feature map, and the coordinates of each anchor box defined on the feature map S are E k =(x k , y k , σ k , θ k );

[0132] Among them, x k , y k are the abscissa and ordinate of the upper left corner point of the anchor box respectively, σ k , θ k are the width and height of the anchor box respectively, and k is the index of the anchor box;

[0133] The output of the classification branch is the binary classification result of whether there is a target signal in the anchor box. The feature map S is preprocessed through the 3×3 convolutional layer of the RPN network to obtain the feature extraction result As shown in Equation (10):

[0134] S′ = ReLU(z feat *S + b feat ) (10)

[0135] Among them, ReLU(·) is the rectified linear unit activation function, z feat , b featThey are the weights of the feature extraction layer and the bias vector respectively, and * represents the convolution operation;

[0136] After preprocessing, the coordinates of the anchor box defined on S′ are (x′ k , y′ k , σ′ k , θ′ k ), is the binary classification result of each anchor box, indicating the probability that the k-th anchor box contains the target within the (i, j)-th pixel in the S′ map, as shown in Equation (11):

[0137]

[0138] where is the sigmoid activation function, z cls , b cls are the weights of the classification layer and the bias vector respectively;

[0139] The regression branch is used to predict the bounding box coordinates of the mask corresponding to the signal segment, and through training, the predicted anchor box coordinates are gradually approximated to the true bounding box to reduce the regression error;

[0140] Based on the offset predicted by regression, the coordinates of the anchor box are adjusted as shown in Equation (12):

[0141]

[0142] The RPN network performs non-maximum suppression (NMS) on the adjusted anchor boxes, combines P cls selects the anchor boxes with higher confidence, and after being adjusted by Δ reg , 1000 high-quality candidate regions are obtained Each E′ n represents the coordinates of the candidate region on S′, and the 1000 high-quality candidate regions are mapped to the coordinate system of the feature map S, and the coordinates represented by the candidate region are n is the index of the final candidate box, s is the downsampling ratio of the feature map to the input image, and s = 8;

[0143] Step 3.3: Guided by the candidate regions determined in Step 3.2, 1000 feature image patches with a size of 14 pixels × 14 pixels are extracted from S to form a region matrix The feature patch generated by the n-th candidate region is as shown in Equation (13):

[0144]

[0145] Among them, is the sampling feature value of the pixel point (i, j) in S after the position mapping of the candidate region to the feature map S, and ω is the weight of bilinear interpolation; The output module of the instance segmentation network will

[0146] be respectively input into the classification branch, the bounding box regression branch, and the mask prediction branch, as shown in Equation (14):

[0147]

[0148] Among them, P is the classification probability result, including the classification probabilities of the signal, noise segment, and confusion segment, Δ is the bounding box regression result, including the horizontal and vertical coordinates of the upper left corner of the bounding box, the length and width of the box, and M is the binary mask prediction result indicating whether the pixel point at this position belongs to the classification. is the combined function of the data reduction Flatten and the fully connected layer, is the convolutional layer function, and Z and B are the output weights and biases of the corresponding branches respectively;

[0149] Step Four: According to the training process as Figure 3 shown, the network uses the loss function and the predicted Intersection over Union (IoU) as the discrimination criterion, selects the optimizer strategy of SGD combined with LARS for training, records the data of the three branches of the above network and determines the training metrics: loss and IoU, and then updates the network parameters or completes the training according to the training situation, specifically including the following sub-steps:

[0150] Step 4.1: The training process of the instance segmentation network involves classification, bounding box regression, and mask prediction tasks. The multi-task learning loss is as shown in Equation (15):

[0151]

[0152] The cross-entropy loss is used to determine the classification loss as shown in Equation (16):

[0153]

[0154] Among them, is the probability of each category output by the model;

[0155] The SLL (Smooth L1 Loss) is used to determine the bounding box regression loss as shown in Equation (17):

[0156]

[0157] wherein, are the coordinates corresponding to the ideal bounding box and the coordinates corresponding to the predicted bounding box respectively;

[0158] The mask loss is determined by binary cross - entropy As shown in Equation (18):

[0159]

[0160] wherein, are the indication results of the ground - truth mask and the predicted mask at the same pixel point respectively;

[0161] Step 4.2: As Figure 3 shown in the optimizer strategy, in this embodiment, the network weights are initialized to ensure the same variance of the inputs of each layer;

[0162] The training optimizer adopts the SGD combined with the LARS optimizer strategy to relieve the learning pressure of large - scale deep tasks, and integrates the stability of SGD and the advantage of the adaptive learning rate of LARS, so that the network can still converge quickly when using large - batch deep training. The iteration of the SGD optimizer strategy is shown in Equation (19):

[0163]

[0164] where θ t is the value of the parameter vector at the t - th iteration, is the gradient of the loss function with respect to the parameter, λθ t is the weight decay term used to prevent over - fitting, and η is the global learning rate;

[0165] Since the gradient differences of different layers of the deep model are large, and in large - batch training, a fixed learning rate will affect the gradient stability. The layer - by - layer learning rate dynamic adjustment mechanism of LARS is adopted to improve the training accuracy, improve the model performance, and improve the convergence speed, as shown in Equation (20):

[0166]

[0167] wherein, is the adaptive learning rate of the parameters of the i - th layer, and ε is a very small constant to prevent the learning rate division by zero error;

[0168] Finally, the optimizer update formula of SGD combined with LARS is as shown in (21):

[0169]

[0170] During the training process, the initial learning rate is set to χ0 in the initial state, and the total number of rounds is set to μ;

[0171] Step 4.3: After each round of training is completed, use the validation set to evaluate the performance of the model, determine the IoU and the loss value. The IoU is shown in Equation (22):

[0172]

[0173] where u is the true mask label, is the predicted mask classification;

[0174] When the following situations occur during the training process, the training will end:

[0175] 1. When the number of training rounds reaches 25 rounds, stop training and check the parameters. Fine-tune the network parameters or increase the number of training rounds as needed, and retrain;

[0176] 2. When the loss value of the validation set has not decreased or the fluctuation has not exceeded the critical value of 1×10 -3 for four consecutive rounds, terminate the training in advance to avoid overfitting. Use the test set to test the output model and define it as the optimal network;

[0177] 3. When the IoU value of the validation set has not decreased or the fluctuation has not exceeded the critical value of 5×10 -4 for four consecutive rounds, terminate the training in advance to avoid overfitting. Use the test set to test the output model and define it as the optimal network;

[0178] Figure 4 is the convergence situation of the network during the training process. Under the same dataset and hardware conditions, after 25 rounds of training, using the SGD combined with the LARS optimizer strategy reduces the training convergence rounds by 42.9% compared to only using the SGD optimizer, and at the same time the IoU reaches 0.958, indicating that the method of the present invention can effectively improve the transfer efficiency of target features and accelerate the network training cycle.

[0179] It also includes Step Five: Deploy the optimal instance segmentation network obtained in Step Four to the satellite-ground communication receiving terminal developed in Python. Generate the STFT image of the received signal through Step Two, perform mask prediction using Step Three, extract the pixel-level segmentation mask of the signal instance in the time-frequency image, and apply the optimal instance segmentation network obtained in Step Four training to track and analyze the carrier center frequency situation of the satellite communication signal.

[0180] According to the operation information shown in Table 2, the complete estimation process of the carrier center frequency takes 3 ms and occupies 35 MiB of memory resources, proving that the method of the present invention can meet the requirements of high timeliness for carrier center frequency estimation in satellite communication;

[0181] Table 2 Software operation information of satellite-ground communication receiving terminal

[0182] Backbone ResNet50 FPN Input_shape [512,384] Cuda 0 (i.e., no graphics card resources required) Time_consuming 2.96ms Peak memory 35.2MiB

[0183] The comparison results of the carrier center frequency estimation RMSRE between the method of the present invention and other methods are as follows Figure 5 shown. The method of the present invention can enhance the ability of the instance segmentation network to extract image features. When SNR = -20dB, the RMSRE of the carrier center frequency estimation is reduced to 2.15%, which is 95.3% lower than the RMSRE of the traditional FFT method and 98.8% lower than the traditional MLE method. The method of the present invention can achieve high-precision carrier center frequency estimation under the condition of strong noise in satellite communication.

[0184] The above specific description further details the purpose, technical solution and beneficial effects of the invention. It should be understood that the above is only a specific embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A carrier center frequency estimation method based on adaptive learning rate fine-tuning, characterized in that: Based on the instance segmentation network Mask R-CNN, combined with the symmetry features of STFT image signal segments, fine-tune the underlying parts of the instance segmentation backbone network ResNet50 and the feature extraction network (Feature Pyramid Networks, FPN), freeze the high-level features, and adopt an adaptive learning rate optimizer strategy during training to ensure stable updates for each layer. Finally, a lightweight instance segmentation network is formed to improve the training efficiency of the deep network and the key feature extraction ability, and can achieve fast and accurate carrier center frequency estimation in a low signal-to-noise ratio environment.

2. The carrier center frequency estimation method based on adaptive learning rate fine-tuning according to claim 1, wherein: It includes the following steps: Step 1: Establish a satellite communication system model and generate signal data samples; Step 2: Characterize the signal in the time-frequency domain using the Short Time Fourier Transformer (STFT) and the selected symmetric window function for characterization; The model uses STFT transformation to generate the signal time-frequency amplitude spectrum and converts the signal time-frequency amplitude spectrum to the logarithmic scale for processing; Finally form the STFT image Its information includes the time axis, i.e., the width W of the image matrix, and the image frequency axis, i.e., the height H of the image matrix, with the unit of pixel; Annotate the images in the COCO (Common Objects in Context) manner to form a network training set, validation set, and test set; Step 3: In the instance segmentation backbone network, freeze the high-level parts of ResNet50 and FPN, and fine-tune the low-level parts. The generated feature fusion image passes through the Region Proposed Network (RPN) to obtain candidate regions. After RoI Align processing, the classification probability result P, the bounding box regression result Δ, and the mask prediction result M are obtained for updating the network parameters used in training; Step 4: The network uses the loss function and the predicted Intersection over Union (IoU) as the discrimination criterion, and selects the optimizer strategy of SGD combined with LARS for training. Record the data of the three branches of the above network and determine the training metrics: loss and IoU, and update the network parameters or complete the training according to the training situation.

3. The carrier center frequency estimation method based on adaptive learning rate fine-tuning according to claim 2, wherein: In Step 1, the transmitting end sends the modulated random signal, and the channel is affected by noise interference; At the receiving end, sample the signal to obtain a received signal model of length K, as shown in Equation (1): where r k is the symbol sequence of the modulation signal, f c is the carrier center frequency, is the Gaussian additive noise with a mean of 0 and a variance of is the root raised cosine shaping filter function, as shown in Equation (2): Among them, T s is the sampling period, T is the symbol period, and α is the filter roll-off coefficient; The signal-to-noise ratio SNR of the signal is as shown in Equation (3):

4. The carrier center frequency estimation method based on adaptive learning rate fine-tuning according to claim 2, wherein: In step two, the short-time Fourier transform (STFT) and the selected symmetric window function are used to characterize the signal in the time-frequency domain as shown in Equation (4): wherein, is a sliding window function, ζ is the window length, w is the time index, and h is the frequency index.

5. The carrier center frequency estimation method based on adaptive learning rate fine-tuning according to claim 2, wherein: The STFT image transformation in Step 2 is as shown in Equation (5): Among them, is the amplitude value corresponding to each pixel, which is determined by the time index w and the frequency index h.

6. The carrier center frequency estimation method based on adaptive learning rate fine-tuning according to claim 2, characterized in that: Step 2 finally forms the STFT image As shown in Equation (6):

7. The carrier center frequency estimation method based on adaptive learning rate fine-tuning according to claim 2, wherein: Step 3 specifically includes the following sub-steps: Step 3.1: Use the STFT image as the input to the backbone network, i.e., F0 = I. In the residual module stage i of ResNet50, the convolutional layer functions are combined and defined as Conv i , and its output size is W i ×H i , where W i and H i are the width and height of the feature image in the stage i layer. The convolutional output of the residual module stage i in the ResNet50 network is shown in Equation (7): Among them, is the output function of the max pooling layer, b i is the bias of the convolutional layer and λ i is the learnable parameter weight, which adapts and changes during the training process; After being processed by four residual modules, ResNet50 feeds the feature map into the P5 layer of the FPN network as the input image of the FPN network, that is the FPN network propagates the image features from the bottom up and, through the upsampling function uniformly adjusts the dimensions of the feature maps of each layer of the FPN to W j-1 ×H j-1 , obtaining the fused feature map vector of the P j layer as shown in Equation (8): Among them, the learnable weight parameter ξ i is used to balance the influence degrees of low-level features and high-level features on the fused image; The output layer of the FPN fuses the features of the P2 to P5 layers to obtain the final output feature map shown in Equation (9). Among them, Adjust the dimension of the feature map of the P5 layer in the FPN to During the training process, since the high-level parts of the backbone network are frozen, the parameters λ3, λ4, ξ3, and ξ4 do not change with training after being preset; Step 3.2: The RPN network sets K anchor boxes at each pixel point (i, j) on the feature map, and the coordinates of each anchor box defined on the feature map S are Among them, x k , y k are the abscissa and ordinate of the upper left corner point of the anchor box respectively, and σ k , are the width and height of the anchor box respectively, and k is the index of the anchor box; The output of the classification branch is a binary classification result indicating whether the target signal is included in the anchor box. The feature map S is preprocessed through the 3×3 convolutional layer of the RPN network to obtain the feature extraction result As shown in Equation (10): S′ = ReLU(z feat *S + b feat ) (10) where ReLU(·) is the rectified linear unit activation function, and z feat , b feat are the weights of the feature extraction layer and the bias vector respectively, and * represents the convolution operation; After preprocessing, the coordinates of the anchor boxes defined on S′ are which is the binary classification result for each anchor box, representing the probability that the k-th anchor box contains an object within the (i, j)-th pixel in the S′ image, as shown in Equation (11): Among them, is the sigmoid activation function, and z cls , b cls are the classification layer weights and the bias vector, respectively; The regression branch is used to predict the bounding box coordinates of the mask corresponding to the signal segment, and gradually approximate the predicted anchor box coordinates to the real bounding box through training to reduce the regression error; Offset based on regression prediction Adjust the coordinates of the anchor box as shown in Equation (12): The RPN network performs non-maximum suppression (NMS) on the adjusted anchor boxes and combines P cls Selects the anchor boxes with higher confidence, and after being adjusted by Δ reg N high-quality candidate regions are obtained Each E′ n Represents the coordinates of the candidate region on S′. Map the N high-quality candidate regions To the coordinate system of the feature map S, and the coordinates represented by the obtained candidate regions Are Where n is the index of the final candidate box, and s is the downsampling ratio of the feature map to the input image; Step 3.3: Under the guidance of the candidate region determined in Step 3.2 , extract N feature image segments of size W r ×H r from S to form a region matrix The nth candidate region The generated feature segment As shown in Equation (13): Among them, after the position of the candidate region is mapped to the feature map S, the sampling feature value of the pixel point (i, j) in S within it, and ω is the weight of bilinear interpolation; The output module of the instance segmentation network sends to the classification branch, the bounding box regression branch, and the mask prediction branch respectively, as shown in Equation (14): Among them, P is the classification probability result, including the classification probabilities of the signal segment, the noise segment, and the confusion segment. Δ is the bounding box regression result, including the horizontal and vertical coordinates of the upper left corner of the bounding box, the length and width of the box. M is the binary mask prediction result indicating whether the pixel at this position belongs to the classification. It is the combined function of the data dimensionality reduction Flatten and the fully connected layer. It is the function of the convolutional layer. Z and B are the output weights and biases of the corresponding branches respectively.

8. The carrier center frequency estimation method based on adaptive learning rate fine-tuning according to claim 2, wherein: Step 4 specifically includes the following sub-steps: Step 4.1: The training process of the instance segmentation network involves classification, bounding box regression, and mask prediction tasks. The multi-task learning loss is as shown in Equation (15): The cross-entropy loss is used to determine the classification loss As shown in Equation (16): Among them, is the probability of the model output for each category; Use SLL (Smooth L1 Loss) to determine the bounding box regression loss As shown in Equation (17): wherein, are the coordinates corresponding to the ideal bounding box and the coordinates corresponding to the predicted bounding box, respectively; Determine the mask loss by binary cross-entropy As shown in Equation (18): wherein, are respectively the indication results of the real mask and the predicted mask at the same pixel points; Step 4.2: Initialize the network weights to ensure that the variances of the inputs of each layer are consistent; The training optimizer adopts the SGD (Stochastic Gradient Descent) combined with the LARS (Layer-wise Adaptive Rate Scaling) optimizer strategy to relieve the learning pressure of large-scale deep tasks, integrating the stability of SGD and the advantages of the adaptive learning rate of LARS, enabling the network to still converge quickly when using large-batch deep training. The iteration of the SGD optimizer strategy is shown in Equation (19): where θ t is the value of the parameter vector at the t-th iteration, is the loss function gradient of the parameters, λθ t is the weight decay term used to prevent overfitting, and η is the global learning rate; Since the gradient differences of different layers of the deep model are relatively large, and in large-batch training, a fixed learning rate will affect the gradient stability. The layer-wise dynamic adjustment mechanism of the learning rate of LARS is adopted to improve the training accuracy, enhance the model performance, and increase the convergence speed, as shown in Equation (20): in, is the adaptive learning rate of the i-th layer parameters, and ε is a very small constant to prevent the learning rate from dividing by 0; Finally, the optimizer update of SGD combined with LARS is obtained as shown in (21): During the training process, the initial learning rate is set to χ0 in the initial state, and the total number of epochs is set to μ; Step 4.3: After each epoch of training is completed, use the validation set to evaluate the performance of the model, determine the IoU and the loss value. The IoU is shown in Equation (22): where u is the true masked label, is the predicted masked classification.

9. The carrier center frequency estimation method based on adaptive learning rate fine-tuning according to claim 2, wherein: The training process in Step 4 will end in the following situations:

1. When the number of training epochs reaches the preset value μ, stop training and check the parameters. Fine-tune the network parameters or increase the number of training epochs as needed, and retrain; 2. When the loss value or the IoU value of the validation set has not decreased or the fluctuation has not exceeded the critical value for μ′ consecutive epochs, terminate the training in advance to avoid overfitting. Use the test set to test the output model and define it as the optimal network.

10. The carrier center frequency estimation method based on adaptive learning rate fine-tuning according to claim 2, wherein: Step Five: Deploy the optimal instance segmentation network obtained in Step Four to the satellite communication signal analysis device. Generate the STFT image of the received signal through Step Two, perform mask prediction using Step Three, extract the pixel-level segmentation mask of the signal instance in the time-frequency image, and apply the optimal instance segmentation network trained in Step Four to estimate and track the carrier center frequency situation of the satellite communication signal. It can enhance the ability of the instance segmentation network to extract important image features and selectively ignore unimportant features. An adaptive learning rate is introduced during the training process to achieve high-precision and fast carrier center frequency estimation in the satellite communication environment with fast-changing Doppler frequency shift and strong noise, meeting the core requirements of the satellite communication network for high-precision and low-latency signal tracking.