Signal bandwidth estimation method based on jump connection enhanced instance segmentation

By introducing a jump connection instance segmentation network in ResNet101 and FPN, the problem of insufficient accuracy and range of signal bandwidth estimation in mobile communication is solved, and high-precision signal bandwidth estimation in large bandwidth and low signal-to-noise ratio environments is realized.

CN120282195APending Publication Date: 2025-07-08BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510417496.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing signal bandwidth estimation technology has insufficient estimation accuracy and estimation range in the environment of large bandwidth and low signal-to-noise ratio of mobile communications, and there is a problem of feature degradation in deep network training.

Method used

An instance segmentation network based on jump connection enhancement is adopted. By introducing jump connections in ResNet101 and FPN, the feature fusion capability is optimized, and combined with STFT transformation and instance segmentation network training, accurate estimation of signal bandwidth is achieved.

Benefits of technology

It improves the accuracy and range of signal bandwidth estimation, reduces data dependence, enhances the generalization performance of the network and signal characteristic resolution, and meets the high accuracy and low latency requirements of mobile communications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120282195A_ABST
    Figure CN120282195A_ABST
Patent Text Reader

Abstract

The invention discloses a signal bandwidth estimation method based on jump connection enhancement instance segmentation, and belongs to the technical field of signal bandwidth estimation in mobile communication. The method effectively improves the multi-scale feature fusion capability of the network through the step-by-step introduction of the jump connection in the backbone network and the feature extraction network of the instance segmentation network, finally obtains the optimal instance segmentation network of an enhanced structure through training, improves the image feature resolution, and improves the image quality. The convergence stability and the feature fusion capability of the training process are improved, and accurate bandwidth estimation of large-bandwidth signals in a low signal-to-noise ratio environment can be realized. According to the method, data dependence can be effectively reduced, and generalization performance is improved; the large-scale training efficiency is improved, and the segmentation performance and accuracy are enhanced; the capability of distinguishing the fuzzy boundary in the signal time-frequency image is enhanced, the signal bandwidth estimation error is reduced, and the signal bandwidth estimation range is expanded; operation memory resource consumption is reduced, the real-time requirement of a communication system is met, and resources are saved for other key functions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a signal bandwidth estimation method based on skip connection enhanced instance segmentation, belonging to the technical field of signal bandwidth estimation in mobile communication. Background Art

[0002] In scenarios such as ultra-dense networking, dynamic spectrum sharing, and multi-user communication, accurate estimation of signal bandwidth has become a prerequisite technology for realizing intelligent spectrum management, improving network energy efficiency, and pre-suppressing interference. The signal bandwidth estimation technology is one of the key foundations for the stable operation of communication systems, which can accurately determine the spectrum occupancy of signals, realize dynamic allocation of spectrum resources, and improve the utilization efficiency of intelligent network spectra.

[0003] Currently, the main signal bandwidth estimation methods include the following two categories:

[0004] Traditional signal bandwidth estimation methods include analyzing signal data in the time domain, frequency domain, and joint domain. The Auto-Correlative Spectrum Function (ACF) method estimates the bandwidth based on the signal frequency domain power spectral density and signal time domain correlation, and can accurately estimate the narrowband signal bandwidth, but the estimation accuracy of large bandwidth signals is extremely low. The Continuous Wavelet Transformer (CWT) method processes signals by combining multi-dimensional domains, has strong multi-resolution analysis capabilities, and is suitable for complex signal processing. However, both the ACF method and the CWT method still have limitations in the bandwidth estimation range, and have low accuracy in low signal-to-noise ratio environments, and cannot meet the bandwidth estimation requirements of mobile communication systems.

[0005] Bandwidth estimation based on deep learning mainly trains and tests signal images to estimate signal bandwidth. Traditional convolutional neural network fusion with adversarial neural network, extreme learning machine and other enhanced network bandwidth estimation methods can adapt to the signal characteristics of real-time transformation in mobile communication and show strong robustness. However, the problem of blurred image features caused by large bandwidth and strong noise in mobile communication seriously affects the estimation accuracy. Methods in the visual fields such as semantic segmentation and instance segmentation, such as UNet, have not been widely applied to the field of signal analysis. Their original design is for general object segmentation tasks, and there are certain limitations in the field of bandwidth estimation technology with high-precision requirements. At the same time, the feature degradation and convergence instability phenomena in large-scale training and deep network structures pose higher requirements for bandwidth estimation methods.

[0006] Therefore, in the context of the large bandwidth and low signal-to-noise ratio characteristics of mobile communication, there are obvious deficiencies in the estimation accuracy and estimation range of existing bandwidth estimation techniques. It is necessary to provide an advanced method to improve the estimation accuracy and estimation range of bandwidth estimation techniques in the field of mobile communication, while alleviating the problem of feature degradation caused by deep networks and large-scale training. Summary of the Invention

[0007] Aiming at the deficiencies of the existing technology and the spectral characteristics of ultra-large bandwidth in mobile communication and the transmission characteristics in a strong noise environment, the main purpose of the present invention is to provide a signal bandwidth estimation method based on skip connection enhanced instance segmentation. By gradually introducing skip connections in the backbone network ResNet101 and the feature extraction network (Feature Pyramid Networks, FPN) of the instance segmentation network Mask R-CNN, the multi-scale feature fusion ability of the network is effectively improved. Finally, the optimal instance segmentation network with an enhanced structure is obtained through training, the resolution of image features is improved, the convergence stability and feature fusion ability of the training process are enhanced, and accurate bandwidth estimation of large bandwidth signals in a low signal-to-noise ratio environment can be achieved.

[0008] The object of the present invention is achieved by the following technical solutions:

[0009] A signal bandwidth estimation method based on skip connection enhanced instance segmentation disclosed by the present invention includes the following steps:

[0010] Step 1: Establish a mobile communication system model and generate signal data samples;

[0011] Among them, the transmitting end sends a modulated random signal, and the channel is affected by noise interference;

[0012] At the receiving end, the signal is sampled to obtain a received signal model of length K, as shown in Equation (1):

[0013]

[0014] Among them, r k is the symbol sequence of the modulated signal, f c is the center frequency of the signal carrier, is the Gaussian additive noise with a mean of 0 and a variance of is the root raised cosine shaping filter function, as shown in Equation (2):

[0015]

[0016] Among them, T s is the sampling period, T is the symbol period, and α is the filter roll-off coefficient;

[0017] The signal-to-noise ratio SNR of the signal is as shown in Equation (3):

[0018]

[0019] Step 2: Characterize the signal in the time-frequency domain using the Short Time Fourier Transformer (STFT) and the selected symmetric window function, as shown in Equation (4): as follows:

[0020]

[0021] where is the sliding window function, λ is the window length, w is the time index, and h is the frequency index;

[0022] The model uses the STFT transform to generate the time-frequency amplitude spectrum of the signal, and converts the time-frequency amplitude spectrum of the signal to the logarithmic scale for processing. The STFT image transform is shown in Equation (5):

[0023]

[0024] where is the amplitude value corresponding to each pixel, determined by the time index w and the frequency index h;

[0025] Finally, the STFT image is formed. Its information includes the time axis, i.e., the width W of the image matrix, and the image frequency axis, i.e., the height H of the image matrix, in pixels, as shown in Equation (6):

[0026]

[0027] Label the image in the COCO (Common Objects in Context) manner to form the network training set, validation set, and test set;

[0028] Step 3: Gradually introduce skip connections in the instance segmentation backbone network ResNet101 and FPN. The generated feature fusion image passes through the Region Proposed Network (RPN) to obtain candidate regions. After RoIAlign processing, the classification probability result P, the bounding box regression result Δ, and the mask prediction result M are obtained for updating the network parameters used for training, specifically including the following sub-steps:

[0029] Step 3.1: Use the STFT image as the input to the backbone network, i.e., F0 = I. In the residual module stage i of ResNet101, the convolutional layer functions are combined and defined as Conv i , and its output size is W i ×H i , where Wi , H i are the width and height of the feature image in the stage i layer. The convolutional output of the residual module stage i in the ResNet101 network is shown in Equation (7):

[0030]

[0031] where, is the output function of the max pooling layer, b i is the bias of the convolutional layer and

[0032] After being processed by four residual modules, ResNet101 inputs the feature map shown in Equation (8) into the P5 layer of the FPN network;

[0033]

[0034] where, S5 is the input of the P5 layer of the FPN network, β i is the skip feature transfer function of the residual module,

[0035] The skip connection connects stage 1 to stage 3 in ResNet101 with P2 to P4 in the FPN network, and through the dimension matching convolution function and the upsampling function respectively adjust the dimensions of the feature maps of each layer of ResNet101 and FPN to W j-1 ×H j-1 , and obtain the fused feature map vector of the P j layer, as shown in Equation (9):

[0036]

[0037] where, The learnable weight parameters γ and ξ are used to balance the influence of low-level features and high-level features on the fused image;

[0038] FPN fuses the features of the P2 to P5 layers and performs an upsampling operation on the feature map of the P5 layer to obtain the final output feature map as shown in Equation (10):

[0039]

[0040] where, adjusts the dimension of the feature map of the P5 layer in FPN to

[0041] Step 3.2: The RPN network sets K anchor boxes at each pixel point (i, j) on the feature map. The coordinates of each anchor box defined on the feature map S are E k =(x k ,y k ,σ k ,θ k );

[0042] Among them, x k ,y k are the abscissa and ordinate of the upper left corner point of the anchor box respectively, σ k ,θ k are the width and height of the anchor box respectively, and k is the index of the anchor box;

[0043] The output of the classification branch is the binary classification result of whether there is a target signal in the anchor box. The feature map S is preprocessed through the 3×3 convolutional layer of the RPN network to obtain the feature extraction result As shown in Equation (11):

[0044] S′ = ReLU(z feat *S + b feat ) (11)

[0045] Among them, ReLU(·) is the rectified linear unit activation function, z feat , b feat are the weights of the feature extraction layer and the bias vector respectively, and * represents the convolution operation;

[0046] After preprocessing, the coordinates of the anchor box defined on S′ are (x′ k ,y′ k ,σ k ′,θ k ′), is the binary classification result of each anchor box, indicating the probability that the k-th anchor box in the (i, j)-th pixel point of the S′ map contains the target, as shown in Equation (12):

[0047]

[0048] Among them, is the sigmoid activation function, z cls , b cls are the weights of the classification layer and the bias vector respectively;

[0049] The regression branch is used to predict the bounding box coordinates of the mask corresponding to the signal segment, and gradually approximate the predicted anchor box coordinates to the true bounding box through training to reduce the regression error;

[0050] Based on the offset predicted by regression, the coordinates of the anchor box are adjusted as shown in Equation (13):

[0051]

[0052] The RPN network performs non-maximum suppression (NMS) on the adjusted anchor boxes, and combines with P cls Select the anchor boxes with higher confidence, and after Δ reg Adjustment, N high-quality candidate regions are obtained Each E n ′ represents the coordinates of the candidate region on S′. Map the N high-quality candidate regions to the coordinate system of the feature map S to obtain the candidate regions The coordinates represented are n is the index of the final candidate box, and s is the downsampling ratio of the feature map to the input image;

[0053] Step 3.3: Under the guidance of the candidate regions determined in Step 3.2 Extract N feature image segments of size W r ×H r from S to form a region matrix The nth candidate region The generated feature segment As shown in Equation (14):

[0054]

[0055] Among them, is the sampling feature value of the pixel point (i, j) in S after the position of the candidate region is mapped to the feature map S, and ω is the weight of bilinear interpolation;

[0056] The output module of the instance segmentation network will be respectively input into the classification branch, the bounding box regression branch, and the mask prediction branch, as shown in Equation (15):

[0057]

[0058] Among them, P is the classification probability result, including the classification probabilities of the signal, noise segments, and confusion segments, Δ is the bounding box regression result, including the horizontal and vertical coordinates of the upper left corner of the bounding box, the length and width of the box, M is the binary mask prediction result indicating whether the pixel point at this position belongs to the classification, is the combined function of data dimensionality reduction Flatten and the fully connected layer, is the function of the convolutional layer, and Z and B are the output weights and biases of the corresponding branches respectively;

[0059] ​Step 4: The network uses the loss function and the predicted Intersection over Union (IoU) as the discriminant criteria, selects the SGD optimizer and the learning rate decay strategy for training, records the data of the three branches of the network, and determines the training metrics: loss and IoU. Subsequently, the network parameters are updated or the training is completed according to the training situation, which specifically includes the following sub-steps:

[0060] Step 4.1: The training process of the instance segmentation network involves classification, bounding box regression, and mask prediction tasks. The multi-task learning loss is shown in Equation (16):

[0061]

[0062] The cross-entropy loss is used to determine the classification loss As shown in Equation (17):

[0063]

[0064] Among them, is the probability of each category output by the model;

[0065] The Smooth L1 Loss (SLL) is used to determine the bounding box regression loss As shown in Equation (18):

[0066]

[0067] Among them, are the coordinates corresponding to the ideal bounding box and the coordinates corresponding to the predicted bounding box, respectively;

[0068] The binary cross-entropy is used to determine the mask loss As shown in Equation (19):

[0069]

[0070] Among them, are the indication results of the true mask and the predicted mask at the same pixel point, respectively;

[0071] Step 4.2: Initialize the network weights to ensure that the variances of the inputs of each layer are consistent;

[0072] The SGD optimizer strategy is adopted for the training optimizer, as shown in Equation (20):

[0073]

[0074] Among them, θ t is the value of the parameter vector at the t-th iteration, is the loss function Gradient of the parameter, λθ t is the weight decay term used to prevent overfitting, and η is the global learning rate;

[0075] During the training process, the initial learning rate is set to χ0 in the initial state, and the total number of epochs is set to μ;

[0076] The learning rate decay strategy is enabled during the training process. After every μ0 epochs, the learning rate is multiplied by ε to accelerate convergence and improve the model stability;

[0077] Step 4.3: After each epoch of training is completed, use the validation set to evaluate the performance of the model, determine the IoU and the loss value. The IoU is shown in Equation (21):

[0078]

[0079] where u is the true mask label, is the predicted mask classification;

[0080] When the following situations are encountered during the training process, the training will end:

[0081] 1. When the number of training epochs reaches the preset value μ, stop training and check the parameters. Fine-tune the network parameters or increase the number of training epochs as needed, and then retrain;

[0082] 2. When the loss value or the IoU value of the validation set has not decreased or the fluctuation has not exceeded the critical value for μ′ consecutive epochs, terminate the training in advance to avoid overfitting. Use the test set to test the output model and define it as the optimal network;

[0083] It further includes Step Five: Deploy the optimal instance segmentation network obtained in Step Four to the mobile communication signal analysis device. Generate the STFT image of the received signal through Step Two, perform mask prediction using Step Three, extract the pixel-level segmentation mask of the signal instance in the time-frequency image, and calculate the signal bandwidth according to the vertical height of the mask in the frequency domain direction. This can enhance the ability of the instance segmentation network to extract image features, perform real-time processing on the continuous time-frequency image sequence, apply the optimal instance segmentation network trained in Step Four to estimate and analyze the bandwidth situation of the mobile communication signal, and achieve high-precision signal bandwidth estimation under the conditions of large bandwidth and strong noise in mobile communication, meeting the core requirements of mobile communication networks for high-precision and low-latency spectrum management.

[0084] Beneficial effects:

[0085] 1. A signal bandwidth estimation method based on skip connection enhanced instance segmentation of the present invention uses an instance segmentation network to perform pixel-level mask segmentation on the signal STFT image, and calculates the signal bandwidth based on the vertical height of the mask in the frequency domain, which can effectively reduce data dependence and improve the network generalization performance.

[0086] 2. A signal bandwidth estimation method based on skip connection enhanced instance segmentation of the present invention introduces skip connection, optimizes the gradient transfer path of the backbone network, improves the ability to obtain underlying image features through backpropagation, improves large-scale training efficiency, and enhances the network mask segmentation performance and accuracy.

[0087] 3. A signal bandwidth estimation method based on skip connection enhanced instance segmentation of the present invention fully fuses the underlying detail features obtained by the ResNet101 residual module with the multi-scale features of each layer of the FPN network through skip connection, and fully fuses with the multi-scale features of each layer of the FPN network, enhancing the model's ability to distinguish blurred boundaries in the signal time-frequency image, reducing the signal bandwidth estimation error, and expanding the signal bandwidth estimation range.

[0088] 4. A signal bandwidth estimation method based on skip connection enhanced instance segmentation of the present invention can deploy the optimal instance segmentation network to edge devices through model lightweighting, achieve fast bandwidth estimation, reduce the consumption of system operating memory resources, meet the real-time requirements of the communication system, and save resources for the system to implement other key functions. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] Figure 1 is a flowchart of a signal bandwidth estimation method based on skip connection enhanced instance segmentation of the present invention;

[0090] Figure 2 is a schematic diagram of the network structure in the embodiment;

[0091] Figure 3 is a flowchart of network training in the embodiment;

[0092] Figure 4 is a curve of the change of IoU in each training round of the network training process in the embodiment;

[0093] Figure 5 is the test result of the root mean square relative error (RMSRE) of signal bandwidth estimation in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0094] To better illustrate the purpose and advantages of the present invention, the following further describes the content of the invention with reference to the drawings and examples.

[0095] Example 1:

[0096] In this example, a random modulation signal is generated in the 8PSK modulation mode, and the signal is sampled at a sampling rate of 10 GHz at the receiving end and subjected to STFT transformation to construct a network training set, a validation set, and a test set, where the signal carrier center frequency fc For a 3 GHz Gaussian channel with a transmission signal-to-noise ratio (SNR) range of -10 dB to 10 dB and a symbol rate 1 / T range of 0.5 GBaud to 2.3 GBaud, taking the above settings as an example, a signal bandwidth estimation method based on skip connection enhanced instance segmentation of the present invention is applied for bandwidth estimation, as Figure 1 shown, which includes the following steps:

[0097] Step 1: Establish a mobile communication system model and generate signal data samples;

[0098] Among them, the transmitting end sends a random signal modulated by 8PSK, and the channel is affected by noise interference;

[0099] At the receiving end, the signal is sampled to obtain a received signal model with a length of 10,000, as shown in Equation (1):

[0100]

[0101] where r k = exp(j2πm k / 8), m k ∈{0, 1, 2, …, 7} is the symbol sequence of the 8PSK modulated signal, is Gaussian additive noise with a mean of 0 and a variance of is the root-raised cosine shaping filter function, as shown in Equation (2):

[0102]

[0103] where T s is the sampling period, T is the symbol period, and α is the filter roll-off coefficient;

[0104] The signal-to-noise ratio SNR of the signal is as shown in Equation (3):

[0105]

[0106] Step 2: Perform STFT using a 256-point Hanning window function. The selected window function characterizes the signal in the time-frequency domain, as shown in Equation (4):

[0107]

[0108] where, is the sliding window function, w is the time index, and h is the frequency index;

[0109] The model uses the STFT transform to generate the time-frequency amplitude spectrum of the signal, and converts the time-frequency amplitude spectrum of the signal to the logarithmic scale for processing. The STFT image transform is as shown in Equation (5):

[0110]

[0111] Among them, is the amplitude value corresponding to each pixel, which is determined by the time index w and the frequency index h;

[0112] Finally, an STFT image is formed Its information includes the time axis, that is, the width of the image matrix is 512 pixels, and the image frequency axis, that is, the height of the image matrix is 384 pixels, as shown in Equation (6):

[0113]

[0114] The image is COCO-annotated to form a network training set, a validation set, and a test set. Finally, a training set and a validation set containing 12,000 images and a test set of 2,000 images are formed, and the coverage range includes the SNR and symbol rate 1 / T ranges mentioned in this embodiment;

[0115] Step Three: According to the network structure as Figure 2 shown, skip connections are introduced step by step in the instance segmentation backbone network ResNet101 and FPN. The generated feature fusion image passes through the Region Proposed Network (RPN) to obtain candidate regions. After RoI Align processing, the classification probability result P, the bounding box regression result Δ, and the mask prediction result M are obtained for updating the network parameters used for training, which specifically includes the following sub-steps:

[0116] Step 3.1: As shown in the structure of the backbone network in Figure 2 , the STFT image is used as the input of the backbone network, that is, F0 = I. In the residual module stage i of ResNet101, the convolutional layer functions are combined and defined as Conv i , and its output size is W i ×H i , where W i , H i are the width and height of the feature image in the stage i layer, and the specific sizes are shown in Table 1:

[0117] Table 1 Output image sizes of each residual module of ResNet101

[0118] Residual module <![CDATA[W i > <![CDATA[H i > Stage 1 128 pixels 96 pixels Stage 2 64 pixels 48 pixels Stage 3 32 pixels 24 pixels Stage 4 16 pixels 12 pixels

[0119] The convolutional output of the residual module stage i in the ResNet101 network is as shown in Equation (7):

[0120]

[0121] Among them, is the output function of the max pooling layer, b i is the bias of the convolutional layer and

[0122] After being processed by four residual modules, ResNet101 inputs the feature map shown in Equation (8) into the P5 layer of the FPN network;

[0123]

[0124] where S5 is the input of the P5 layer of the FPN network, β i is the skip feature transfer function of the residual module,

[0125] The skip connection connects stage 1 to stage 3 in ResNet101 with P2 to P4 of the FPN network, and through the dimension matching convolution function and the upsampling function u(·), the dimensions of the feature maps of each layer of ResNet101 and FPN are uniformly adjusted to W j-1 ×H j-1 , to obtain the fused feature map vector of the P j layer, as shown in Equation (9):

[0126]

[0127] where, The learnable weight parameters γ and ξ are used to balance the influence of low-level features and high-level features on the fused image;

[0128] FPN fuses the features of P2 to P5 layers and performs an upsampling operation on the feature map of the P5 layer to obtain the final output feature map as shown in Equation (10):

[0129]

[0130] where, The dimension of the feature map of the P5 layer in FPN is adjusted to 128 pixels × 96 pixels;

[0131] Step 3.2: As Figure 2 In the structure of the regional candidate network, the RPN network sets 10 anchor boxes at each pixel point (i, j) on the feature map, and the coordinates of each anchor box defined on the feature map S are E k =(x k ,y k ,σ k ,θ k );

[0132] where, xk , y k are the abscissa and ordinate of the upper left corner point of the anchor box, respectively, and σ k , θ k are the width and height of the anchor box respectively, and k is the index of the anchor box;

[0133] The output of the classification branch is the binary classification result of whether there is a target signal in the anchor box. After preprocessing the feature map S through the 3×3 convolutional layer of the RPN network, the feature extraction result is obtained As shown in Equation (11):

[0134] S′ = ReLU(z feat * S + b feat ) (11)

[0135] where ReLU(·) is the rectified linear activation function, z feat , b feat are the weights of the feature extraction layer and the bias vector respectively, and * represents the convolution operation;

[0136] After preprocessing, the coordinates of the anchor box defined on S′ are (x′ k , y′ k , σ k ′, θ k ′), is the binary classification result for each anchor box, representing the probability that the k-th anchor box contains a target within the (i, j)-th pixel point in the S′ map, as shown in Equation (12):

[0137]

[0138] where is the sigmoid activation function, z cls , b cls are the weights of the classification layer and the bias vector respectively;

[0139] The regression branch is used to predict the bounding box coordinates of the mask corresponding to the signal segment, and through training, the predicted anchor box coordinates are gradually approximated to the true bounding box to reduce the regression error;

[0140] Based on the offset predicted by regression, the coordinates of the anchor box are adjusted as shown in Equation (13):

[0141]

[0142] The RPN network performs non-maximum suppression (NMS) on the adjusted anchor boxes, and combines P cls to select the anchor boxes with higher confidence, and through Δ regAfter adjustment, 2000 high-quality candidate regions are obtained. Each E n ′ represents the coordinates of the candidate region on S′. Map the N high-quality candidate regions to the coordinate system of the feature map S to obtain the candidate regions The coordinates represented are where n is the index of the final candidate box, s is the downsampling ratio of the feature map to the input image and s = 8;

[0143] Step 3.3: Under the guidance of the candidate regions determined in Step 3.2 extract 2000 feature image fragments with a size of 7 pixels × 7 pixels from S to form a region matrix The nth candidate region The generated feature fragment is shown in Equation (14):

[0144]

[0145] where is the sampling feature value of the pixel point (i, j) in S after mapping the position of the candidate region to the feature map S, and ω is the weight of bilinear interpolation;

[0146] The output module of the instance segmentation network inputs into the classification branch, the bounding box regression branch, and the mask prediction branch respectively, as shown in Equation (15):

[0147]

[0148] where P is the classification probability result, including the classification probabilities of the signal, noise segments, and confusion segments, Δ is the bounding box regression result, including the horizontal and vertical coordinates of the upper left corner of the bounding box, the length and width of the box, and M is the binary mask prediction result indicating whether the pixel point at this position belongs to the classification. is the combined function of data dimensionality reduction Flatten and the fully connected layer, is the convolutional layer function, and Z and B are the output weights and biases of the corresponding branches respectively;

[0149] Step Four: According to the training structure as Figure 3 shown, the network is trained using the loss function and IoU as the discriminant criteria, adopting the SGD optimizer and the learning rate decay strategy. Record the data of the three branches of the above network and determine the training metrics: loss and IoU. Then update the network parameters or complete the training according to the training situation, specifically including the following sub-steps:

[0150] Step 4.1: AsFigure 3 As shown in the multi-task loss function, the constructed network training process involves classification, bounding box regression, and mask prediction tasks. The multi-task learning loss is shown in Equation (16):

[0151]

[0152] The cross-entropy loss is used to determine the classification loss As shown in Equation (17):

[0153]

[0154] where is the probability of each class output by the model;

[0155] The SLL (Smooth L1 Loss) is used to determine the bounding box regression loss As shown in Equation (18):

[0156]

[0157] where are the coordinates corresponding to the ideal bounding box and the coordinates corresponding to the predicted bounding box, respectively;

[0158] The binary cross-entropy is used to determine the mask loss As shown in Equation (19):

[0159]

[0160] where are the indication results of the true mask and the predicted mask at the same pixel point, respectively;

[0161] Step 4.2: As Figure 3 shown in the optimizer strategy, the network weights are initialized to ensure that the variances of the inputs of each layer are consistent; the training optimizer adopts the SGD optimizer strategy, as shown in Equation (20):

[0162]

[0163] where θ t is the value of the parameter vector at the t-th iteration, is the loss function is the gradient of the parameter, λθ t is the weight decay term used to prevent overfitting, and η is the global learning rate;

[0164] During the training process, the initial learning rate is set to 0.02 in the initial state, and the total number of training rounds is set to 25;

[0165] During the training process, a learning rate decay strategy is enabled. At the 8th, 15th, and 22nd training epochs respectively, the learning rate is multiplied by 0.1 to accelerate convergence and improve model stability;

[0166] Step 4.3: After each training epoch is completed, use the validation set to evaluate the performance of the model, determine the IoU and the loss value. The IoU is shown in Equation (21):

[0167]

[0168] where u is the true mask label, is the predicted mask classification;

[0169] When the following situations occur during the training process, the training will end:

[0170] 1. When the number of training epochs reaches 25, stop training and check the parameters. Modify the network parameters or increase the number of training epochs as needed, and retrain;

[0171] 2. When the loss value of the validation set has not decreased or the fluctuation has not exceeded the critical value of 5×10 -4 for four consecutive epochs, terminate the training in advance to avoid overfitting. Use the test set to test the output model and define it as the optimal network;

[0172] 3. When the IoU value of the validation set has not decreased or the fluctuation has not exceeded the critical value of 2×10 -4 for four consecutive epochs, terminate the training in advance to avoid overfitting. Use the test set to test the output model and define it as the optimal network;

[0173] Figure 4 is the convergence situation of the network during the training process. Under the same dataset and hardware conditions, after 25 epochs of training, the number of iterations required for the network with skip connections to converge is reduced by 28.5% compared to the Mask R-CNN network, and the IoU is improved to 0.991, indicating that the method of the present invention can optimize the gradient transfer path of the backbone network and improve the mask segmentation accuracy at the same time;

[0174] Step Five: Deploy the optimal instance segmentation network obtained in Step Four to the spectrum monitoring terminal device developed in Python. Generate the STFT image of the received signal through Step Two, perform mask prediction using Step Three, extract the pixel-level segmentation mask of the signal instance in the time-frequency image, and evaluate the signal bandwidth according to the vertical height of the mask in the frequency domain direction. The complete process takes 18 ms and occupies 74.4 MiB of memory resources. Under the condition of ensuring accuracy, the method of the present invention can save resource overhead and achieve efficient bandwidth estimation;

[0175] The comparison result of the bandwidth estimation RMSRE between the present invention and other methods is as Figure 5As shown, the training strategy can enhance the ability of the instance segmentation network to extract image features. When SNR = -10dB, the RMSRE of signal bandwidth estimation is reduced to 5.88%. Compared with the traditional ACF method, the RMSRE is reduced by 98.6%, and compared with the CWT method, it is reduced by 98.3%. Under the conditions of the embodiments, the method of the present invention can measure the signal bandwidth of 0.675GHz to 3.105GHz with an error of RMSRE ≤ 5.88%. Compared with the range of 0.675GHz to 2GHz of the traditional method, the improvement reaches 1.83 times, indicating that the method of the present invention can achieve more accurate signal bandwidth estimation under the conditions of large bandwidth and strong noise in mobile communication.

[0176] The above specific description further details the purpose, technical solution and beneficial effects of the invention. It should be understood that the above is only a specific embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A signal bandwidth estimation method based on skip connection enhanced instance segmentation, characterized in that: By gradually introducing skip connections in the backbone network ResNet101 and the feature extraction network (Feature Pyramid Networks, FPN) of the instance segmentation network Mask R-CNN, the multi-scale feature fusion ability of the network is effectively improved. Finally, the optimal instance segmentation network with an enhanced structure is obtained through training, the resolution of image features is improved, the convergence stability and feature fusion ability of the training process are enhanced, and accurate bandwidth estimation of large-bandwidth signals can be achieved in a low signal-to-noise ratio environment.

2. The signal bandwidth estimation method based on skip connection enhanced instance segmentation according to claim 1, wherein: Specifically, it includes the following steps: Step 1: Establish a mobile communication system model and generate signal data samples; Step 2: Characterize the signal in the time-frequency domain by using the Short Time Fourier Transformer (STFT) and the selected symmetric window function for characterization; The model uses the STFT transform to generate the signal time-frequency amplitude spectrum and converts the signal time-frequency amplitude spectrum to the logarithmic scale for processing; Finally form the STFT image Its information includes the time axis, i.e., the width W of the image matrix, and the image frequency axis, i.e., the height H of the image matrix, with the unit being pixels; The images are labeled in the COCO (Common Objects in Context) manner to form a network training set, a validation set, and a test set; Step 3: Gradually introduce skip connections in the instance segmentation backbone network ResNet101 and FPN. The generated feature fusion image passes through the Region Proposed Network (RPN) to obtain candidate regions. After RoI Align processing, the classification probability result P, the bounding box regression result Δ, and the mask prediction result M are obtained for updating the network parameters used for training; Step 4: The network uses the loss function and the predicted Intersection over Union (IoU) as the discrimination criteria, selects the SGD optimizer and the learning rate decay strategy for training, records the three-branch data of the above network and determines the training metrics: loss and IoU, and then updates the network parameters or completes the training according to the training situation; Step 5: Deploy the optimal instance segmentation network obtained in Step 4 to the mobile communication signal analysis device. Generate the STFT image of the received signal through Step 2, use Step 3 for mask prediction, extract the pixel-level segmentation mask of the signal instance in the time-frequency image, and calculate the signal bandwidth according to the vertical height of the mask in the frequency domain direction, which can enhance the ability of the instance segmentation network to extract image features, perform real-time processing on the continuous time-frequency image sequence, apply the optimal instance segmentation network trained in Step 4 to estimate and analyze the bandwidth situation of mobile communication signals, and achieve high-precision signal bandwidth estimation under the conditions of large bandwidth and strong noise in mobile communication, meeting the core requirements of mobile communication networks for high-precision and low-latency spectrum management.

3. The signal bandwidth estimation method based on jump connection enhanced instance segmentation according to claim 2, characterized in that: In Step 1, the transmitting end sends a modulated random signal, and the channel is affected by noise interference; At the receiving end, the signal is sampled to obtain a received signal model of length K, as shown in Equation (1): where r k is the symbol sequence of the modulation signal, f c is the center frequency of the signal carrier, is Gaussian additive noise with a mean of 0 and a variance of is the root raised cosine shaping filter function as shown in Equation (2): Among them, T s is the sampling period, T is the symbol period, and α is the filter roll-off coefficient; The signal-to-noise ratio SNR of the signal is as shown in Equation (3):

4. The method for estimating signal bandwidth based on skip connection enhanced instance segmentation according to claim 2, wherein: The signal is characterized in the time-frequency domain by using STFT and the selected symmetric window function, as shown in Equation (4): as shown in Equation (4): Among them, is a sliding window function, λ is the window length, w is the time index, and h is the frequency index.

5. The signal bandwidth estimation method based on skip connection enhanced instance segmentation according to claim 2, wherein: The STFT image transform is as shown in Equation (5): Among them, is the amplitude value corresponding to each pixel, which is determined by the time index w and the frequency index h.

6. The method for estimating signal bandwidth based on jump connection enhanced instance segmentation according to claim 2, wherein: STFT image As shown in Equation (6):

7. The method for estimating signal bandwidth based on jump connection enhanced instance segmentation according to claim 2, wherein: Step 3 specifically includes the following sub-steps: Step 3.1: Use the STFT image as the input to the backbone network, i.e., F0 = I. In the residual module stage i of ResNet101, the convolutional layer functions are combined and defined as Conv i , and its output size is W i ×H i , where W i , H i are the width and height of the feature image in the stage i layer. The convolutional output of the residual module stage i in the ResNet101 network is shown in Equation (7): Among them, is the output function of the max pooling layer, b i is the bias of the convolutional layer and After being processed by four residual modules, ResNet101 inputs the feature map shown in Equation (8) into the P5 layer of the FPN network; Among them, S5 is the input of the P5 layer of the FPN network, β i is the skip feature transfer function of the residual module, The skip connection connects stages 1 to 3 in ResNet101 to P2 to P4 in the FPN network, and through the dimension matching convolution function and the upsampling function respectively unify the dimensions of the feature maps of each layer of ResNet101 and FPN to W j-1 ×H j-1 , obtaining the fused feature map vector of the P j layer, as shown in Equation (9): Among them, The learnable weight parameters γ and ξ are used to balance the influence degrees of low-level features and high-level features on the fused image; The FPN fuses the features of the P2 to P5 layers and performs an upsampling operation on the feature map of the P5 layer to obtain the final output feature map As shown in Equation (10): Among them, Adjust the dimension of the feature map of the P5 layer in the FPN to Step 3.2: The RPN network sets K anchor boxes at each pixel point (i, j) on the feature map, and the coordinates of each anchor box defined on the feature map S are where x k , y k are the abscissa and ordinate of the upper-left corner point of the anchor box respectively, and σ k , are the width and height of the anchor box respectively, and k is the index of the anchor box; The output of the classification branch is the binary classification result of whether the target signal is included in the anchor box. The feature map S is preprocessed through the 3×3 convolutional layer of the RPN network to obtain the feature extraction result As shown in Equation (11): S′ = ReLU(z feat *S + b feat ) (11) where ReLU(·) is the rectified linear activation function, and z feat , b feat are the weights of the feature extraction layer and the bias vector respectively, and * represents the convolution operation; After preprocessing, the coordinates of the anchor boxes defined on S′ are is the binary classification result for each anchor box, indicating the probability that the k-th anchor box contains the target within the (i, j)-th pixel in the S′ map, as shown in Equation (12): Among them, is the sigmoid activation function, and z cls , b cls are the weights of the classification layer and the bias vector, respectively; The regression branch is used to predict the bounding box coordinates of the mask corresponding to the signal segment, and through training, the predicted anchor box coordinates are gradually approximated to the true bounding box to reduce the regression error; Offset based on regression prediction Adjust the coordinates of the anchor box as shown in Equation (13): The RPN network performs non-maximum suppression (NMS) on the adjusted anchor boxes and combines P cls Selects the anchor boxes with higher confidence levels. After being adjusted by Δ reg N high-quality candidate regions are obtained Each E n ′ represents the coordinates of the candidate region on S′. Map the N high-quality candidate regions to the coordinate system of the feature map S to obtain the candidate regions The coordinates represented are where n is the index of the final candidate box and s is the downsampling ratio of the feature map to the input image; Step 3.3: Under the guidance of the candidate region determined in Step 3.2 extract N feature image segments of size W r ×H r from S to form a region matrix The nth candidate region The generated feature segment As shown in Equation (14): Among them, after the position of the candidate region is mapped to the feature map S, the sampling feature value of the pixel point (i, j) in S within it, and ω is the weight of bilinear interpolation; The output module of the instance segmentation network will be input into the classification branch, the bounding box regression branch, and the mask prediction branch respectively, as shown in Equation (15): Among them, P is the classification probability result, including the classification probabilities of the signal, noise segment, and confusion segment, Δ is the bounding box regression result, including the horizontal and vertical coordinates of the upper left corner of the bounding box, the length and width of the box, and M is the binary mask prediction result indicating whether the pixel at this position belongs to the classification. It is the common function of the data dimensionality reduction Flatten and the fully connected layer. It is the function of the convolutional layer. Z and B are the output weights and biases of the corresponding branches respectively.

8. The method for estimating signal bandwidth based on skip connection enhanced instance segmentation according to claim 2, characterized in that: Step 4 specifically includes the following sub-steps: Step 4.1: The training process of the instance segmentation network involves classification, bounding box regression, and mask prediction tasks. The multi-task learning loss is shown in Equation (16): Use cross-entropy loss to determine the classification loss As shown in Equation (17): wherein, is the probability of each category output by the model; Use SLL (Smooth L1 Loss) to determine the bounding box regression loss As shown in Equation (18): wherein, are the coordinates corresponding to the ideal bounding box and the coordinates corresponding to the predicted bounding box, respectively; Determine the mask loss through binary cross-entropy As shown in Equation (19): wherein, are respectively the indication results of the real mask and the predicted mask at the same pixel points; Step 4.2: Initialize the network weights to ensure that the variances of the inputs of each layer are consistent; The training optimizer adopts the SGD optimizer strategy, as shown in Equation (20): where θ t is the value of the parameter vector at the t-th iteration, is the gradient of the loss function with respect to the parameters, λθ t is the weight decay term used to prevent overfitting, and η is the global learning rate; During the training process, the initial learning rate is set to χ0 in the initial state, and the total number of epochs is set to μ; The learning rate decay strategy is enabled during the training process. After every μ0 epochs, the learning rate is multiplied by ε to accelerate convergence and improve the model stability; Step 4.3: After each epoch of training is completed, use the validation set to evaluate the performance of the model, determine the IoU and the loss value. The IoU is shown in Equation (21): where u is the true masked label, is the predicted masked classification.

9. The signal bandwidth estimation method based on jump connection enhanced instance segmentation according to claim 2, wherein: When the following situations occur during the training process of Step 4, the training will end:

1. When the number of training epochs reaches the preset value μ, stop training and check the parameters. Fine-tune the network parameters or increase the number of training epochs as needed, and retrain; 2. When the loss value or IoU value of the validation set has not decreased or the fluctuation has not exceeded the critical value for μ′ consecutive epochs, terminate the training in advance to avoid overfitting. Use the test set to test the output model and define it as the optimal network.