A short wave signal detection method and system

CN122594788BActive Publication Date: 2026-09-11BEIJING HAIGE SHENZHOU COMM TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611079511.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-09-11
Estimated Expiration
2046-07-21

AI Technical Summary

Technical Problem

现有各类方法的特征提取模块均将两个物理意义完全不同的维度视为同质空间进行混合处理,无法针对各维度的物理特性进行差异化特征提取,限制了特征表达的判别力

Benefits of technology

(1)时频双维度解耦特征提取。以非对称深度可分离卷积替代对称卷积骨干,将时间维度(码元通断节奏)与频率维度(载波带宽与频率漂移)解耦建模,一维时间核捕获时序规律,一维频率核捕获谱形特征,避免对称卷积核混合处理两个物理维度导致特征判别力受限。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122594788B_ABST
    Figure CN122594788B_ABST
Patent Text Reader

Abstract

The application discloses a short-wave signal detection method and system, and belongs to the technical field of short-wave radio monitoring; the method comprises the following steps: collecting with Beidou second pulses as sampling trigger references, and establishing a Beidou time anchor point and frame coordinate bidirectional mapping; inputting after pretreatment into an asymmetric multi-scale feature encoder, and outputting a multi-scale feature map; performing three-branch center line attribute prediction synthesis initial bounding box after layer-by-layer attention enhancement in the channel-frequency-time three-dimensional mode; sequentially executing three-level semantic post-processing of frequency domain sidelobe suppression, symbol level merging and transmission segment level aggregation, and restoring the fragment detection box to a complete communication transmission segment; converting the start and end time into a microsecond-level absolute time stamp according to the bidirectional mapping, and outputting a structured result. The application realizes full-process automatic detection from radio frequency collection to absolute time stamp and transmission end fingerprint feature, and is suitable for a low signal-to-noise ratio detection scene of a short-wave narrowband keying signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of shortwave radio signal monitoring and intelligent signal processing technology, specifically to a shortwave signal detection method and system, applicable to low signal-to-noise ratio intelligent detection of shortwave narrowband keyed communication signals, including but not limited to Morse signals, amplitude shift keying signals (ASK / OOK), and frequency shift keying signals (FSK). Background Technology

[0002] Shortwave narrowband keyed communication signals are still widely used in emergency rescue, maritime communications, and military communications due to their simple equipment configuration, strong long-distance transmission capability, and strong anti-interference characteristics. In radio spectrum monitoring and signal reconnaissance scenarios, rapid detection and parameter extraction of such signals are crucial for signal alarm, content interception, and target tracing.

[0003] Traditional shortwave signal detection methods mainly rely on energy detection, matched filtering, and cyclostationary analysis. However, in low signal-to-noise ratio (SNR≤8dB) and complex electromagnetic environments, the false alarm rate and false negative rate are often insufficient for practical applications. In recent years, researchers have begun to combine time-frequency analysis with deep target detection networks, transforming the shortwave signal detection problem into a target detection problem on a time-frequency map. Existing technical approaches mainly fall into three categories: The first is the transfer application of YOLO series and its variants. Represented by YOLOv5, YOLOv8, and YOLOX, these methods perform bounding box regression on the time-frequency map using pre-set anchor boxes or without anchor boxes. The backbone network often employs symmetric convolutional structures such as CSPDarkNet and EfficientNet. The second category is methods based on Transformer attention architectures. These methods capture the global feature dependencies of the signal on the time-frequency map through a self-attention mechanism. After embedding the time-frequency map into blocks, they homogeneously encode the time and frequency dimensions with attention to achieve signal localization and classification. The third category is methods based on centerline modeling. Centerline Networks (CLNs), for example, use ResNet-18 as the feature extraction backbone network. They predict the position and extent of the signal centerline on the time-frequency map through a branching structure and obtain the signal bounding box from the centerline heatmap using connected component decoding. However, through systematic research, the inventors of this invention have discovered that the aforementioned existing deep learning-based shortwave signal detection methods generally suffer from the following unrecognized and unresolved technical shortcomings: First, there is a mismatch between the feature extraction backbone network and the two-dimensional physical characteristics of the time-frequency image. YOLO variants use symmetric convolutional backbones such as CSPDarkNet, CLN uses a ResNet-18 symmetric convolutional backbone, and the Transformer attention method, after block embedding, also treats the time and frequency dimensions as homogeneous spatial position sequences. However, the time-frequency image has a fundamentally different physical structure from natural images—the frequency axis reflects the carrier frequency distribution and bandwidth characteristics of the signal, while the time axis reflects the on / off timing arrangement of the signal's symbols. For example, with a Morse signal, the time axis represents a rhythmic sequence of dots, dashes, and intervals; with an FSK signal, the time axis represents a time-slot sequence of frequency jumps. Existing methods' feature extraction modules treat these two physically completely different dimensions as homogeneous spaces and process them together, failing to extract differentiated features based on the physical characteristics of each dimension, thus limiting the discriminative power of feature representation. In situations with low SNR or frequency jitter, it is difficult to utilize temporal context information to enhance detection robustness.

[0004] Second, feature extraction lacks multi-scale context awareness, making it difficult to adapt to signal variations with different code rates and bandwidths. Existing models achieve multi-scale fusion solely through deconvolution upsampling and adding features to the corresponding stage, lacking an explicit multi-scale context aggregation mechanism. In actual communication, signals such as Morse and FSK exhibit significant differences in symbol duration on their time-frequency maps due to large code rate variations (from tens to hundreds of codes per minute) caused by channel environment (fading, multipath). Simultaneously, different transmitters also exhibit differences in bandwidth and frequency drift range. Current technologies cannot simultaneously detect short-duration point pulses and long-duration scratch pulses at the same feature location, resulting in poor consistency in detecting signals at different scales.

[0005] Third, the detection head lacks deformation adaptive capability and cannot cope with the non-rigid signal morphology caused by frequency jitter. Existing CLN models use standard 3×3 and 1×1 convolutional layers for three-branch attribute prediction, with the sampling position of the convolutional kernel remaining fixed. However, hand-crafted Morse signals exhibit non-rigid frequency jitter in the time-frequency plot due to factors such as key contact jitter and channel Doppler effects—the signal bright bars are not strictly horizontal straight lines, but fluctuate vertically or even bend. The rigid sampling grid of standard convolution cannot adaptively fit this deformation, leading to centerline prediction bias and bounding box regression errors.

[0006] Fourth, the lack of an explicit attention mechanism tailored to the characteristics of time-frequency maps limits their feature representation capabilities. Existing models do not introduce adaptive attention mechanisms, treating all feature channels, frequency positions, and time positions equally. They fail to adaptively weight these key dimensions, resulting in background noise and irrelevant interference features being treated the same, suppressing the response intensity of useful signals. For example, the fully convolutional structure of the CLN model has no attention module, and the attention of the YOLO variant only applies to the channel dimension and does not extend to the frequency and time dimensions specific to time-frequency maps.

[0007] Fifth, the lack of time synchronization capability with real monitoring systems prevents the use of detection results for time correlation analysis. Existing models only output the relative time range of the signal on the time-frequency graph, without establishing a mapping relationship with an absolute time reference. In practical applications such as radio monitoring and signal intelligence (SIGINT), time correlation analysis of multi-station, multi-channel detection signals is required to determine the signal source and track transmitter activity patterns. Detection results lacking absolute timestamps cannot meet these engineering requirements, severely limiting the practical deployment of the algorithm.

[0008] Sixth, the detection results remain at the "pixel-level" signal localization level, failing to achieve "communication semantic-level" signal reconstruction. The YOLO variant outputs isolated bounding boxes, the Transformer method outputs foreground region markers, and CLN's connected component decoding outputs only independent time-frequency segments. None of these methods utilize the inherent protocol hierarchy of shortwave keying (SSK) communication signals to perform semantic-level aggregation of the detection results. Taking the Morse signal as an example, its hierarchy is as follows: symbols form characters, characters form words, and complete word sequences form communication transmission segments, with each level separated by standard intervals (character interval = 3 times symbol duration, word interval = 7 times symbol duration). FSK and other digital modulation signals also have hierarchical relationships such as frame structure and time slot guard intervals. Existing methods do not utilize this prior knowledge, and the fragmented detection boxes output cannot be reconstructed into complete communication transmission segments.

[0009] Seventh, the detection results lack information output at the signal feature level and do not possess transmitter fingerprinting capabilities. Existing methods output information limited to bounding box coordinates and signal classification labels, failing to extract individual transmitter features from the detected signals. For manually keyed signals such as Morse, individual differences exist in the dot-stroke duration ratio and symbol interval consistency among different transmitters, forming a handwriting-like operational fingerprint that can be used for cross-communication association and transmitter identification. For machine-transmitted signals such as FSK and OOK, differences in hardware parameters such as crystal oscillator stability and power amplifier characteristics also constitute a transmitter hardware fingerprint, which can be used for equipment tracking and model identification. Existing methods rely entirely on manual measurement for these transmitter fingerprint features, failing to meet the needs of automated signal association and tracing in large-scale monitoring.

[0010] Therefore, there is an urgent need for a signal detection method and system that can overcome the shortcomings of the above-mentioned technologies for shortwave narrowband keying communication signals. Summary of the Invention

[0011] Based on the above analysis, the present invention aims to disclose a shortwave signal detection method and system; to solve the above problems, and to achieve fully automated processing from raw radio frequency acquisition to the output of symbol-level structured detection information with absolute timestamp and transmitter fingerprint features in complex electromagnetic environments with low signal-to-noise ratio.

[0012] This invention discloses a shortwave signal detection method, comprising: S1. BeiDou Time Synchronization Acquisition and Time-Frequency Transformation: The shortwave receiver uses the rising edge of the BeiDou second pulse signal as the sampling trigger reference to align the starting sampling point of each channel's sampling sequence with the absolute integer second of BeiDou; the acquired shortwave time domain signal is transformed into time-frequency data frames to obtain time-frequency data frames, and a bidirectional mapping table is established between the absolute time anchor point of BeiDou and the relative frame coordinates of the time-frequency data frames; after preprocessing, it is input into an asymmetric multi-scale feature encoder, which learns the on / off rules of symbol by decoupling the asymmetric depthwise separable convolution and the time-series axial self-attention through the decoupling of the frequency and time dimensions, and outputs a multi-scale feature map; S2. Apply attention weights layer by layer along the channel, frequency and time dimensions to the multi-scale feature map to obtain the attention-enhanced feature map; S3. Perform three-branch centerline attribute prediction on the attention-enhanced feature map. The first branch inserts a time-axis self-attention layer along the time axis. Based on the unified time reference, calculate the correlation weight between features of each time step along the time dimension. Apply a causal mask so that the current time step only interacts with the historical time step to capture the symbol on / off time-series features. Output the centerline heatmap. The second branch outputs the position offset map. The third branch outputs the height map. Synthesize the initial bounding box. S4. Communication semantic level post-processing: The initial bounding box is processed sequentially as follows: frequency domain sidelobe suppression, symbol-level merging with a character-level interval threshold that is adaptive to the signal symbol duration to obtain a character-level transmission segment, and transmission segment-level aggregation with a transmission segment interval threshold that is adaptive to the signal symbol duration to obtain a complete communication transmission segment. S5. BeiDou Absolute Timestamp Synthesis and Structured Output of Signal Features: Based on the bidirectional mapping table, the start and end times of the transmission segment are converted into absolute timestamps; the output includes a structured detection result containing absolute start timestamp, absolute end timestamp, center frequency, upper and lower frequency bounds, confidence score, estimated symbol rate, symbol duration statistics, signal-to-noise ratio estimate, transmission segment type, and signal feature vector.

[0013] Another aspect of the present invention discloses a shortwave signal detection system, comprising: a BeiDou time synchronization module, a shortwave radio frequency direct acquisition receiver, a signal processing unit, and a data storage unit; wherein, The BeiDou timing module, including the BeiDou antenna and timing module, is used to receive BeiDou satellite signals and output clock information and BeiDou second pulse signals; The shortwave radio frequency direct sampling receiver includes a shortwave reconnaissance antenna and a radio frequency direct sampling receiver, which is connected to the Beidou time synchronization module via a clock line. It is used to receive the clock information as a sampling reference and the Beidou second pulse signal as a sampling trigger reference, and to collect broadband I / Q data in the target frequency band so that the starting sampling point of each channel sampling sequence is aligned with the absolute whole second. The signal processing unit is connected to the shortwave radio frequency direct acquisition receiver via a high-speed data bus interface. It is used to receive the broadband I / Q data, execute the shortwave signal detection method described above, and output the signal detection result with an absolute timestamp. The data storage unit is connected to the signal processing unit via a high-speed data bus interface and is used to store the signal detection results and the corresponding time-frequency data.

[0014] Compared with the prior art, the present invention has the following technical advantages: (1) Time-frequency dual-dimensional decoupled feature extraction. Asymmetric depth separable convolution is used to replace the symmetric convolution backbone, decoupling the time dimension (symbol on / off rhythm) and the frequency dimension (carrier bandwidth and frequency drift) into a model. A one-dimensional time kernel captures the temporal pattern, and a one-dimensional frequency kernel captures the spectral features, avoiding the limitation of feature discrimination caused by the mixed processing of the two physical dimensions by the symmetric convolution kernel.

[0015] (2) Full-link BeiDou time synchronization. From sampling trigger to result output, the entire link uses the BeiDou second pulse as a unified time reference, establishes a two-way mapping between BeiDou time anchor point and frame coordinate, and accurately converts the start and end time of the detection result into microsecond-level absolute timestamps, which meets the needs of multi-station TDOA positioning and signal tracing, and fills the gap of existing pure laboratory algorithms that do not have absolute timestamp capabilities.

[0016] (3) Three-dimensional attention adaptive enhancement. Attention weights are applied layer by layer along the three directions of channel-frequency-time. Channel attention filters important features, frequency attention focuses on carrier bandwidth and frequency drift neighborhood, and time attention focuses on symbol dot-dash-interval timing neighborhood. The three attentions work together to enhance signal features and suppress background noise in a series residual manner.

[0017] (4) Causal-temporal axial self-attention dual modeling. Timing-axial self-attention is embedded in both the backbone network and the first branch of the detector head. A causal mask is applied so that the current interaction only occurs with historical time steps, capturing the irreversible on / off rhythm of the signal symbols; parameters are shared across frequency rows to learn frequency-independent general symbol patterns. Unlike DTF-AT's bidirectional self-attention along both time and frequency axes, this invention applies causal self-attention only along the time axis, stemming from the physical characteristics of each frequency row carrying independently transmitted signals.

[0018] (5) Deformation-adaptive detection and multi-scale perception. The deformable convolution in the first layer of the detection head applies anisotropic offset constraints—a small offset in the time direction adapts to symbol jitter, and a larger offset in the frequency direction adapts to frequency drift, so that the sampling position adapts to the non-rigid signal shape; the multi-scale dilated convolution in the middle layer works in conjunction with the hollow spatial pyramid pooling to simultaneously perceive short-time pulses and long-time signals.

[0019] (6) Communication semantic level signal restoration. Utilizing the protocol-level prior of the keying signal, after three levels of post-processing—frequency domain sidelobe suppression, symbol-level merging, and transmission segment-level aggregation—fragmented frame-level detection boxes are sequentially restored into character-level transmission segments and complete communication transmission segments, achieving a leap from pixel localization to semantic restoration; the interval threshold is adaptively determined based on the estimated symbol duration.

[0020] (7) Automatic extraction of transmitter fingerprint features. The signal-to-noise ratio is extracted from the frequency attention map, the frequency drift variance is extracted from the deformable convolution offset sequence, and the symbol rate, dot-to-dash ratio and consistency score are calculated from the symbol duration statistics to form the transmitter fingerprint feature vector, which supports the identification of the sender of manual keying signals and the tracking of machine-transmitted signals.

[0021] (8) Engineering practicality. It realizes the closed loop of the entire process from radio frequency acquisition to structured intelligence output, and can be directly deployed at shortwave monitoring stations without the need for additional timing equipment and manual feature measurement. Attached Figure Description

[0022] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. Figure 1 This is a flowchart of the shortwave signal detection method in an embodiment of the present invention; Figure 2 This is a schematic diagram of the feature extraction network structure in an embodiment of the present invention; Figure 3 This is a structural diagram of an asymmetric depth-separable convolutional block in an embodiment of the present invention; Figure 4 This is a schematic diagram of the single-segment signal detection result in an embodiment of the present invention; Figure 5This is a schematic diagram of the multi-segment signal detection results in an embodiment of the present invention; Figure 6 This is a schematic diagram of the detection results in a multi-channel scenario in an embodiment of the present invention; Figure 7 This is a schematic diagram showing the components and connections of the shortwave signal detection system in an embodiment of the present invention. Detailed Implementation

[0023] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and, together with the embodiments of the present invention, serve to illustrate the principles of the present invention.

[0024] Example 1 One embodiment of the present invention discloses a shortwave signal detection method, such as... Figure 1 As shown, it includes the following steps: S1. The shortwave receiver uses the BeiDou second pulse signal as the sampling trigger reference to align the starting sampling point of each channel sampling sequence with the absolute integer second of BeiDou; it performs time-frequency transformation on the acquired shortwave time domain signal to obtain time-frequency data frames, and establishes a bidirectional mapping table between the BeiDou absolute time anchor point and the frame coordinates; after preprocessing, it performs multi-scale feature extraction on the time-frequency data frames and outputs multi-scale feature maps. S2. Apply attention weights layer by layer along the channel, frequency and time dimensions to the multi-scale feature map to obtain the attention-enhanced feature map; S3. Perform centerline attribute prediction on the attention-enhanced feature map, output the centerline heatmap, position offset map and height map, and synthesize the initial bounding box; S4. Perform communication semantic level post-processing on the initial bounding box, sequentially performing frequency domain sidelobe suppression, symbol level merging and transmission segment level aggregation to restore the fragmented detection box into a complete communication transmission segment. S5. Convert the start and end times of the transmission segment into absolute timestamps according to the bidirectional mapping table; output a structured detection result containing absolute timestamps, center frequency, symbol parameters, signal-to-noise ratio estimates, and signal feature vectors.

[0025] Specifically, S1 includes: S1-1: The BeiDou timing module receives the BeiDou satellite timing signal and outputs a reference clock to the shortwave receiver as the sampling clock reference. Simultaneously, it outputs a BeiDou second pulse signal as the sampling trigger reference. The shortwave receiver, within the target frequency band, uses the rising edge of the BeiDou second pulse signal as the sampling trigger time to acquire broadband I / Q data, aligning the starting sampling point of each channel's sampling sequence with the absolute integer second. The broadband I / Q data is then subjected to parallel multiplexing to obtain a narrowband time domain signal. The shortwave receiver acquires broadband I / Q data within the target frequency band (1.5~30MHz) at a sampling rate of 40MHz. The BeiDou timing module receives BeiDou timing signals through the BeiDou antenna and outputs a 10MHz sampling clock to the shortwave receiver as a sampling reference, while simultaneously outputting a BeiDou second pulse signal as a sampling trigger reference. The shortwave receiver uses the rising edge of the BeiDou second pulse signal as the sampling trigger reference to align the starting sampling point of each channel's sampling sequence with the absolute integer second, thus obtaining a shortwave time-domain signal time-aligned by the BeiDou second pulse signal. The broadband data is then subjected to parallel multiplexing to obtain a narrowband time-domain signal with a sampling rate of 9600Hz.

[0026] S1-2. Perform sliding window framing, short-time Fourier transform, logarithmic compression, global percentile normalization, and time stamping preprocessing on the narrowband time domain signal to obtain a time-stamped single-channel time-frequency map; using the BeiDou second pulse as a reference, record the frame index and intra-frame sampling point offset of the rising edge of the second pulse in each time-frequency data frame to establish a bidirectional mapping table between the BeiDou absolute time anchor point and the frame coordinates. Preprocessing includes: 1) Sliding window framing: The narrowband time domain signal is processed by sliding window framing, with 131,456 sampling points per frame (approximately 13.69 seconds, covering 2 to 3 complete transmission cycles), and adjacent frames overlap by 25% (stepping 98,592 sampling points, approximately 10.27 seconds).

[0027] 2) Short-time Fourier transform; add a Hamming window to each frame of data ( After calculating a 512-point FFT, adjacent frames overlap by 384 points (75% overlap rate, frame shift of 128 points), resulting in a two-dimensional time-frequency energy matrix with a size of 256 (height = 512 / 2 = 256 non-redundant frequency bins) × 1024 (time frame number) pixels; frequency coverage range 0~4800Hz, frequency resolution 18.75Hz / bin, and time resolution approximately 0.0133 seconds / frame (128 / 9600).

[0028] 3) Logarithmic compression; for each spectral amplitude value application Transforms the compressed dynamic range; compresses the dynamic range by approximately four orders of magnitude while preserving an accurate mapping of zero values.

[0029] 4) Global percentile normalization; select the spectrograms of the first min(M,20) windows as the statistical sampling set, where M is the total number of windows in the current data, and calculate the 50th percentile after sorting all pixel values ​​( ) and the 99.5th percentile ( This is used as a global normalization parameter; for the spectrograms of all windows, the values ​​are first cropped to... Interval, then calculate Z-score standardization: , This is the set of all pixel values ​​within the statistical sampling interval; resulting in a standardized single-channel time-frequency map data frame with a mean of 0 and a standard deviation of 1.

[0030] 5) Establishment of a bidirectional mapping table: Based on the BeiDou second pulse, record the frame index and intra-frame sampling point offset of the rising edge of the second pulse in each time-frequency data frame, and establish a mapping entry consisting of BeiDou whole second time code, second pulse phase count value, frame index, and intra-frame sampling offset to form a bidirectional mapping table between the BeiDou absolute time anchor point and the frame coordinates for subsequent retrieval during S5 absolute timestamp synthesis.

[0031] S1-3. Input the single-channel time-frequency map into the asymmetric multi-scale feature encoder. Extract deep feature maps through multi-stage asymmetric depth-separable convolution. Apply temporal axial self-attention to the deep features to learn the on / off rules of symbols. Then, upsample through transposed convolution and fuse with the feature maps of the preceding stages by skip connection. Embed multi-scale hollow spatial pyramid pooling aggregation context in the intermediate layer to output a multi-scale feature map that has both high-level semantics and low-level details.

[0032] like Figure 2 As shown, the asymmetric multi-scale feature encoder includes: an asymmetric feature extraction backbone module, a temporal axial self-attention module, a transposed convolutional decoder module, and a spatial pyramid pooling module (ASPP).

[0033] (a) An asymmetric feature extraction backbone module is connected to a single-channel time-frequency map. Feature extraction and spatial downsampling are performed through multiple stages of asymmetric depth separable convolutional residual units (ADSC units). The ADSC unit decouples and models the physical characteristics of the time-frequency map in two dimensions using one-dimensional depth convolution on the time axis and one-dimensional depth convolution on the frequency axis. Then, it performs channel fusion through pointwise convolution and connects with the input residual. Each stage is downsampled by independent max pooling. The deep feature map is output to the self-attention module on the temporal axis at the end stage. The intermediate feature maps of each stage are output to the transposed convolutional decoder module for skip connection fusion.

[0034] (b) The time-axis self-attention module is connected to the deep feature map and decoupled into frequency row-level time series along the frequency dimension. After layer normalization, multi-head self-attention is applied to model symbol timing dependence. After restoring the feature map size, the time-enhanced feature map is output to the transposed convolutional decoder module. All frequency rows share the same set of self-attention parameters, enabling the model to learn the general symbol on / off rules independent of the carrier frequency. The number of parameters is only 1 / 16 of that of the independent parameter scheme. Unlike the design of DTF-AT which applies bidirectional self-attention along the time and frequency axes, this module applies self-attention only along the time axis, which is due to the physical characteristics of each frequency row carrying independent transmitted signals.

[0035] (c) The transposed convolutional decoder module receives the temporal enhancement feature map, performs multi-stage transposed convolutional upsampling, and merges it with the corresponding stage feature map of the asymmetric feature extraction backbone module through channel splicing. The intermediate layer feature map is output to the hole space pyramid pooling module ASPP, and the final upsampling outputs a multi-scale feature map. (d) The hollow spatial pyramid pooling module is connected to the intermediate layer feature map of the decoder. It captures the multi-scale receptive field context through parallel multi-diffraction convolution branches. After channel splicing and compression, it is fused with the output residual of the decoder. The enhanced feature map is returned to the transposed convolution decoder module for further upsampling.

[0036] The asymmetric feature extraction backbone module includes: (1) Dual-path asymmetric Stem layer.

[0037] The input is a standardized single-channel time-frequency plot (size 1×256×1024). Initial feature projection is first performed through two parallel branches: Branch A: 1×7 convolution (along the time axis) → 7×1 convolution (along the frequency axis), outputting 16 channels; Branch B: 7×1 convolution (along the frequency axis) → 1×7 convolution (along the time axis), outputting 16 channels.

[0038] The outputs of the two branches are concatenated along the channels to obtain an initial feature map with dimensions of 32×256×1024. This design enables the network to extract features differentially from the time and frequency dimensions from the first layer, and the interchange of the convolution order of the two branches further enhances the complementarity of the features.

[0039] (2) Four-stage asymmetric depth separable residual feature extraction.

[0040]

[0041] Among them, the structure of asymmetric depth-separable convolutional blocks is as follows: Figure 3 As shown, this module consists of a first convolutional layer, a second convolutional layer, a third convolutional layer, a residual connection layer, and an activation layer in sequence.

[0042] Both the first and second convolutional layers employ depthwise convolution (DWConv), performing convolution operations independently on each channel in the spatial dimension without mixing information between channels. Specifically, the first convolutional layer uses a 1×7 depthwise convolution to extract symbol on / off timing features along the time axis; the second convolutional layer uses a 7×1 depthwise convolution to extract carrier bandwidth and frequency drift features along the frequency axis. These two layers are decoupled to model the physical characteristics of the time-frequency plot in the time and frequency dimensions, respectively.

[0043] The third convolutional layer uses 1×1 pointwise convolution (PWConv) to fuse information between channels.

[0044] In the first and second convolutional layers, group normalization (GN, number of groups = 32) and Mish activation are performed sequentially after convolution; the third convolutional layer only performs group normalization (GN) and does not apply an activation function to maintain the feature distribution after channel fusion.

[0045] The residual connection layer adds the output of the third convolutional layer to the module input features element by element, and the result is output after nonlinear transformation by the Mish function of the activation layer.

[0046] Independent 2×2 max pooling downsampling layers are set up between each stage to decouple the downsampling operation from feature extraction, so that the features before spatial downsampling retain complete information. The deep feature map (512×16×64) output by Stage4 is input into the subsequent temporal axial self-attention module.

[0047] The time-axis self-attention module includes: The frequency row decoupling submodule decouples the deep feature map into independent time series according to the frequency row. The multi-head self-attention submodule applies layer normalization to the time series and then performs multi-head self-attention processing. Each head projects the channel dimension features into a low-dimensional query matrix, key matrix, and value matrix through three unbiased convolutional kernels, calculates scaled dot product attention, and concatenates the outputs of each head. The channel dimension is then restored by convolutional fusion and connected to the original input residual. The feedforward network submodule processes the output of the multi-head self-attention submodule using a feedforward network. The feedforward network contains two convolutional kernels, activated by GELU (Gaussian Error Linear Unit), and the intermediate dimension is expanded, compressed, and restored. It is then connected to the input residual of the feedforward network to obtain the feedforward enhanced features. The feature recombination submodule reassembles the feedforward enhanced features into a two-dimensional feature map according to the frequency row, and outputs the temporal enhanced feature map to the transposed convolutional decoder module.

[0048] The transposed convolutional decoder module includes: The UpConvBlock3 module receives the temporal enhancement feature map, which is then upsampled by transposed convolution and fused with the corresponding stage feature map of the asymmetric feature extraction backbone module through channel splicing, and output to the UpConvBlock2 module. The UpConvBlock2 module receives the output of the UpConvBlock3 module, continues to upsample and merges it with the corresponding stage feature map of the asymmetric feature extraction backbone module through channel splicing. The main output is sent to the UpConvBlock1 module, and the intermediate layer feature map is output to the hollow space pyramid pooling module. The UpConvBlock1 module receives the output of the UpConvBlock2 module and fuses it with the enhanced feature map returned by the Hollow Space Pyramid Pooling module. It then continues to upsample and merges the feature map with the corresponding stage feature map of the Asymmetric Feature Extraction Backbone Module through channel splicing before outputting it to the output sub-module. The output submodule performs final-stage upsampling on the output of the UpConvBlock1 module to output multi-scale feature maps.

[0049] The void space pyramid pooling module includes: The parallel dilated convolution submodule connects to the intermediate layer feature map of the transposed convolution decoder module and captures multi-scale receptive field context through multiple parallel dilated convolution branches with different dilation rates. The channel splicing submodule splices the outputs of each dilated convolution branch along the channel dimension. The compression fusion submodule performs channel compression on the stitched feature map to obtain context-aggregated features. The residual connection submodule performs residual fusion between the context aggregation features and the intermediate layer feature map of the transposed convolutional decoder module, and outputs the enhanced feature map to the UpConvBlock1 module of the transposed convolutional decoder module.

[0050] A specific implementation plan for multi-scale feature map extraction based on the above network structure. The time-stamped single-channel time-frequency map (size 1×256×1024) is first passed through a dual-path asymmetric STEM layer: branch A undergoes a 1×7 convolution (along the time axis, outputting 16 channels) → a 7×1 convolution (along the frequency axis, outputting 16 channels); branch B undergoes a 7×1 convolution (along the frequency axis, outputting 16 channels) → a 1×7 convolution (along the time axis, outputting 16 channels); the outputs of the two branches are concatenated along the channels to obtain an initial feature map with a size of 32×256×1024.

[0051] The feature extraction process then proceeds through four stages: Stage 1 (input size 256×1024, containing 2 ADSC units, output 64 channels, size 128×512), Stage 2 (containing 2 ADSC units, output 128 channels, size 64×256), Stage 3 (containing 2 ADSC units, output 256 channels, size 32×128), and Stage 4 (containing 2 ADSC units, output 512 channels, size 16×64). Each stage from Stage 1 to Stage 3 is followed by an independent 2×2 max pooling process (stride = 2) for spatial downsampling. Stage 4 does not undergo downsampling; its output deep feature map (512×16×64) is fed into the subsequent temporal axial self-attention module.

[0052] The structure of each ADSC unit is as follows: Input → [1×7DWConv+GN+Mish] → [7×1DWConv+GN+Mish] → [1×1PWConv+GN] → concatenated with the input residual → Mish activation → output. DWConv is a depthwise convolution, with a 1×7 kernel capturing symbol on / off timing features along the time axis and a 7×1 kernel capturing carrier bandwidth and frequency drift features along the frequency axis; PWConv is a 1×1 pointwise convolution for channel fusion; GN is GroupNorm (number of groups = 32); and the activation function is Mish. The deep feature map output from Stage 4 is designated F_deep, with dimensions of 512×16×64.

[0053] Temporal axial self-attention is applied to the F_deep feature map (512×16×64): the feature map is transformed from [B,512,16,64] to [B×16,64,512], that is, each frequency row is treated as an independent 64-frame time series; after applying layer normalization to the sequence, it is input into an 8-head multi-head self-attention module—each head projects the 512-dimensional feature into a 64-dimensional query, key, and value matrix through three 1×1 convolutions (without bias), and the scaled dot product attention is calculated: ; The 64-dimensional outputs from the 8 heads are concatenated to restore 512 dimensions, fused through 1×1 convolutions, and then connected to the residual of the original input. This is followed by a feedforward network (two 1×1 convolutions, GELU activation, intermediate expansion ratio 4: 512→2048→512) and the residual is then connected. Finally, the feature map shape is restored to the [B, 512, 16, 64] output. All frequency rows share the same set of self-attention parameters, enabling the model to learn general symbol on / off rules independent of carrier frequency—for each frequency row, the sequence of 64 time frames shares the same attention parameters, capturing the duration, interval, and rhythmic dependencies of symbols, unrestricted by specific frequency positions.

[0054] Three UpConvBlocks are used for transposed convolutional upsampling decoding. The structure of each UpConvBlock is as follows: 4×4 transposed convolution (stride=2, padding=1, no bias) → GroupNorm → Mish activation; if there are skip connections, the upsampling result is concatenated with the feature map of the corresponding Stage of the asymmetric feature extraction backbone module, and then fused by 3×3 convolution (padding=1, no bias) → GroupNorm → Mish activation. In the specific upsampling path UpConvBlock3: The F_deep feature map (512×16×64) after temporal self-attention enhancement is upsampled to 256×32×128 through transposed convolution, and concatenated with the 256×32×128 of Stage3 along the channel (256+256=512 channels), and then fused into 256 channels through 3×3 convolution; UpConvBlock2: Outputs 128×64×256, which is concatenated with Stage2's 128×64×256 (128+128=256 channels), and then fused into 128 channels by 3×3 convolution. The intermediate layer feature map is output to the dilated spatial pyramid pooling module. The dilated spatial pyramid pooling module applies five parallel branches to the 128×64×256 feature map output by UpConvBlock2: (1) 1×1 standard convolution; (2) 3×3 dilated convolution (dilation rate = 6); (3) 3×3 dilated convolution (dilation rate = 12); (4) 3×3 dilated convolution (dilation rate = 18); (5) global average pooling + 1×1 convolution + bilinear upsampling. Each branch output is 128 channels, which are concatenated along the channels (5×128=640 channels) and then compressed to 128 channels by 1×1 convolution, and fused with the residual output by UpConvBlock2. This module captures multi-scale receptive field context (minimum receptive field 3×3, maximum theoretical receptive field 37×37) from the same feature map through multiple sets of parallel dilated convolutions with different dilation rates, enabling the simultaneous perception of short-term symbol pulses and long-term sustained signals at the same location.

[0055] UpConvBlock1: Receives the output of UpConvBlock2 (128×64×256), concatenates it with the enhanced feature map returned by ASPP (128×64×256) along the channels (128+128=256 channels), fuses it through 3×3 convolution, transposes it for upsampling to 64×128×512, concatenates it with Stage1's 64×128×512 (64+64=128 channels), and fuses it through 3×3 convolution to 64 channels; Output submodule: Perform final-stage transposed convolution upsampling on the output of UpConvBlock1 (64×128×512) to output a multi-scale feature map with a size of 96×256×1024.

[0056] This pyramid-shaped encoder-decoder fusion structure enables the final multi-scale feature map to possess both high-level semantic information (signal category and overall shape) and low-level detailed information (time-frequency edges and fine structure), which helps to uniformly perceive signals with different code rates and bandwidths.

[0057] Specifically, S2 includes: S2-1. Apply attention weights along the channel dimension to the multi-scale feature map to obtain a channel-weighted feature map; Specifically, for the multi-scale feature map Global adaptive average pooling and global max pooling are performed along the spatial dimension (256×1024) to obtain two 96-dimensional vectors. The two vectors are fed into a shared two-layer MLP (96→96 / 16=6→96, with ReLU activation in between), added together, and then activated by Sigmoid to generate a 96-dimensional channel weight vector. Adaptive weighting is applied to the 96 feature channels to output a channel-weighted feature map. ;in, This indicates element-wise multiplication.

[0058] S2-2. Apply attention weights along the frequency dimension to the channel-weighted feature map to obtain the frequency-weighted feature map; After pooling the channel-weighted feature map along the time axis, a frequency attention map is generated by a one-dimensional convolution kernel along the frequency axis, and then applied in the form of residuals to obtain the frequency-weighted feature map. Specifically, the channel-weighted feature map Average pooling and max pooling are performed along the time axis (1024 dimensions) to obtain two pooling results of [96, 256, 1]. These results are then concatenated along the channel dimension to form [192, 256, 1]. A 7×1 convolution kernel is then applied along the frequency axis (padding=3) to output a frequency attention map of [1, 256, 1]. After sigmoid activation, the residual is multiplied by the channel-weighted feature map to obtain the frequency-weighted feature map. The 7×1 nuclei of frequency attention can sense frequency neighborhood information of about 7 frequency bins (about 131 Hz), which is sufficient to cover the bandwidth and frequency drift range of narrowband keying signals.

[0059] S2-3. Apply attention weights along the time dimension to the frequency-weighted feature map to obtain the attention-enhanced feature map; After pooling the frequency-weighted feature map along the frequency axis, a temporal attention map is generated by a one-dimensional convolution kernel along the time axis. The attention-enhanced feature map is then multiplied with the frequency-weighted feature map in residual form to obtain the attention-enhanced feature map. Specifically, frequency-weighted feature maps After pooling along the frequency axis (256 dimensions), it is convolved along the time axis with a 1×7 kernel (padding=3) to produce a temporal attention map of [96,1,1024]. The attention-enhanced feature map is obtained by applying it in the form of residuals. The 1×7 kernel of time attention can perceive temporal neighborhood information for about 7 frames (about 93ms) and can distinguish the duration difference between short-time symbols and long-time symbols.

[0060] Specifically, in S3, the attention-enhanced feature map is input into a three-branch centerline attribute prediction network; the first layer of each branch uses deformable convolution to learn the spatial offset, and the middle layers use multi-scale dilated convolution to expand the receptive field; wherein, The sampling position offset of the deformable convolution is subject to anisotropic constraints; the offset along the time axis is limited to a first preset range, corresponding to the physical range of symbol duration jitter; the offset along the frequency axis is limited to a second preset range, corresponding to the physical range of Doppler frequency shift and frequency drift; the first preset range is smaller than the second preset range; the resolution of the three-branch output feature map is uniformly set to a downsampling factor R=4 relative to the original time-frequency map; The first branch is used for temporal axis centerline heatmap prediction. A temporal axis self-attention layer is inserted along the time axis, expanding the feature map after dilated convolution into a frequency row-level time series. Scaling dot product attention is independently performed on each frequency slice, projected into a query matrix, key matrix, and value matrix by three parallel one-dimensional convolutional layers. A lower triangular causal mask is applied when calculating the attention score, ensuring that the current time step only calculates correlation weights with historical time steps to capture the irreversible temporal rhythm of the signal. After the temporal axis self-attention output, it is predicted point-by-point by a one-dimensional convolutional layer along the time axis. The final layer is activated to output a centerline heatmap, where each pixel value represents the confidence probability that the corresponding time-frequency position belongs to the horizontal centerline of the narrowband signal. The second branch is used for outputting the position offset map; the last layer of the second branch is a convolutional layer, which outputs the position offset map. The third branch is used for height map output; the last layer of the third branch is a convolutional layer, which outputs the height map.

[0061] In a specific implementation, the attention-enhanced feature map (96×256×1024) is input into a three-branch centerline attribute prediction network. The first layer of each branch uses a 3×3 deformable convolution (DeformableConv2d, output 128 channels, stride=1, padding=1). Through a parallel 2×3×3 convolution branch (input 96 channels → output 18 channels), the (x,y) spatial offset of the 9 sampling positions of each 3×3 convolution kernel is learned, so that the effective sampling position of the convolution kernel can adaptively fit non-rigid signal shapes such as frequency drift and frequency jitter. At the same time, Mish activation and BatchNorm normalization are used. Anisotropic constraints are applied to the sampling position offset of the deformable convolution—the offset along the time axis is limited to ±2 pixels, corresponding to the physical range of symbol duration jitter; the offset along the frequency axis is limited to ±5 pixels, corresponding to the physical range of Doppler frequency shift and frequency drift.

[0062] The middle two layers use 3×3 dilated convolutions with dilation rates of 2 and 3 (outputting 128 channels, padding=2 and 3); the dilated convolution with d=2 makes the interval between adjacent sampling points 2 pixels, expanding the effective receptive field to 5×5; d=3 makes the interval 3 pixels, expanding the effective receptive field to 7×7; the progressively expanding receptive field allows the same feature point to refer to an increasingly wider spatial context, which helps to handle signals with different bit rates, different bandwidths and different degrees of distortion.

[0063] in, The first branch is used for predicting the time-series axial centerline heat map; The first branch inserts a temporal axis self-attention layer along the time axis; the feature map after dilated convolution is unfolded along the time dimension, and scaling dot product attention is independently performed on each frequency slice. This is then projected into a query, key, and value matrix through three parallel one-dimensional convolutional layers (kernel size 1). A lower triangular causal mask is applied when calculating the attention score, ensuring that the current time step only calculates correlation weights with historical time steps to capture the dot-stroke rhythm temporal features of the signal. After the temporal axis self-attention output, it undergoes point-by-point prediction through two one-dimensional convolutional layers along the time axis (kernel 3, channels 256→128→1). The final layer is activated by a Sigmoid function to output a centerline heatmap, with dimensions 1×64×256, representing the confidence probability that each spatial location belongs to the horizontal centerline of the narrowband signal. During training, a Gaussian heatmap ( As the supervision target, the Focal Loss loss function is adopted: Calculate for positive samples , Calculate for negative samples ; in, =2, =4; To predict probabilities, This is a real label; By using an exponential weighting method, the optimization focus is concentrated on difficult samples (samples with low prediction confidence) and a small number of positive samples (centerline pixels), which effectively alleviates the problem of extreme positive and negative sample imbalance where centerline pixels account for only 0.1% to 1%.

[0064] The second branch is used for outputting the position offset map; The last layer of the second branch is a 1×1 convolution (outputting 1 channel), which outputs a position offset map with a size of 1×64×256. It is used to correct the vertical (frequency direction) position deviation of the center line caused by feature map downsampling (R=4) and provide continuous positioning with sub-pixel accuracy. SmoothL1 Loss is calculated only at pixel positions where the confidence of the center line heatmap is ≥0.5. .

[0065] The third branch is used for heightmap output; The third branch's final layer is a 1×1 convolution (outputting 2 channels), outputting a height map with dimensions of 2×64×256, which predicts the distances from the center line to the upper and lower boundaries of the signal region, allowing the prediction of asymmetric signal bandwidth shapes (such as the case of unilateral broadening caused by frequency drift); it also uses SmoothL1 Loss and is calculated only at the effective location of the center line.

[0066] The joint training loss function for the three parallel branches is: ; in, For the first branch loss, For the second branch loss, For the third branch loss; To compensate for the centerline continuity loss, the difference in centerline confidence at adjacent positions is calculated frame by frame along the time axis, and the breakage of the predicted centerline on the time axis is penalized. As an L1 sparse regularization term, the L1 norm is calculated for all pixels in the centerline heatmap to suppress noise and false alarms in low-confidence regions. The overall training adopts an end-to-end approach, with all parameters trained using the Adam optimizer (learning rate) via backpropagation. Weight decay Jointly optimize the backbone network and attribute representation module Training is divided into two phases: The first phase (10 epochs) freezes all parameters of the asymmetric feature extraction backbone module and trains only the three-branch detection head. The detection head converges quickly by warming up and cosine annealing learning rate scheduling. The second phase (150 epochs) unfreezes all parameters and performs end-to-end joint optimization, and adopts training techniques such as early stopping (patience=30), mixed precision training, model exponential moving average (Model EMA), and gradient clipping.

[0067] Specifically, S4 includes: S4-1, Thermal properties of the centerline Figure 2 After valueization, connected components are labeled, and the optimal time position is located by accumulating column confidence. The position offset and height mean are combined to map back to the original spectral coordinates to generate a signal bounding box without redundancy. The centerline heatmap is binarized with a threshold of 0.5. An eight-neighbor disjoint-set data structure is used for two passes of the connected component labeling algorithm to extract all connected regions. The following decoding operation is performed on each connected component: 1) Calculate the sum of the centerline confidence scores of all active pixels in a column, and take the column with the largest cumulative sum as the representative centerline time position of the signal. ; 2) In In the column, the row position with the highest confidence level of the center line is taken as the discrete center point. Adding the value of that position to the position offset map yields a continuous center point. ; 3) For all active pixels within the entire connected component, accumulate the position offset and vertical height values, count them separately, and then calculate the average. 、 、 ; 4) Using a downsampling factor of R=4, map the feature map coordinates back to the original spectrogram coordinates: , ; ; ; 5) Crop the coordinate values ​​to and Within range, discard or The degenerate box; where, The smallest column index for the connected components. The index of the largest column in the connected component. This represents the width of the time dimension of the original time-frequency plot. The original time-frequency plot has a frequency dimension height. Based on the connected component decoding scheme, only one bounding box is regressed for each connected component, which automatically avoids redundant predictions and eliminates the need for the non-maximum suppression (NMS) post-processing step required by traditional object detectors.

[0068] Because there is a 25% overlap between sliding windows, the same persistent signal may be detected multiple times, requiring post-processing merging.

[0069] S4-2, Frequency Domain Sidelobe Suppression: For detection box pairs that overlap in time and whose frequency center distance is less than or equal to the first frequency threshold, suppress the detection box with lower confidence. For time-overlapping detection frames, if the distance between the frequency centers of the two frames is less than or equal to a first frequency threshold, the detection frame with lower confidence is suppressed. The first frequency threshold is set according to the STFT frequency resolution and is used to filter out repeated detections of the same signal on different frequency bins due to spectral leakage. In this embodiment, the first frequency threshold is set to 8 frequency bins (approximately 150Hz).

[0070] S4-3, Symbol-level merging: Detection frames with a center frequency difference less than or equal to the second frequency threshold are grouped by frequency. Within each group, they are sorted by time, and adjacent detection frames with a time interval less than or equal to the character-level interval threshold are merged into a character-level transmission segment. The character-level interval threshold is adaptively equal to N1 times the estimated symbol duration of the current frequency group, and N1 is determined according to the standard proportional relationship between character-level interval and symbol duration in the target signal communication protocol. After sorting by time within each group, adjacent detection boxes with a time interval less than or equal to the character-level interval threshold are merged into a single character-level transmission segment. The earliest start time is taken as the start of the merged segment, the latest end time as the end of the merged segment, and the maximum confidence level as the merged score. In this embodiment, the second frequency threshold is set to 3 frequency bins (approximately 56Hz).

[0071] When the target signal is a Morse signal, N1=3, corresponding to the standard character interval (3 times the symbol duration); when the target signal is a digital modulation signal such as FSK, N1 is configured according to its frame structure and time slot protection interval standard.

[0072] S4-4, Transmission segment-level aggregation; merge character-level transmission segments with time intervals less than or equal to the inter-transmission segment interval threshold into a complete communication transmission segment; the inter-transmission segment interval threshold is adaptively equal to N2 times the estimated symbol duration, N2 is greater than N1, and N2 is determined according to the standard proportional relationship of the inter-transmission segment interval in the target signal communication protocol. When the target signal is a Morse signal, N2=7, corresponding to the standard word interval (7 times the symbol duration); when the target signal is a digital modulation signal such as FSK, N2 is configured according to its frame structure and time slot guard interval standard.

[0073] S4-5, Code rate adaptive estimation; cluster analysis is performed on the duration of each detection frame and the interval duration between frames within the current frequency group, and the frames are divided into symbol signal clusters, character-level interval clusters and transmission segment-level interval clusters. The mean of the short-time symbol clusters in the symbol signal clusters is used as the estimated symbol duration for the adaptive configuration of the character-level interval threshold and the transmission segment interval threshold. The mean of the short-time symbol clusters in the symbol signal cluster is used as the estimated symbol duration. This is for adaptive configuration of the S4-3 character-level interval threshold and the S4-4 transmission segment interval threshold.

[0074] Detection frames with a center frequency difference ≤ the second frequency threshold are grouped by frequency; after sorting by time within each group, adjacent detection frames with a time interval ≤ the character-level interval threshold are merged into a character-level transmission segment, and the earliest start time is taken as the start of the merged segment, the latest end time is taken as the end of the merged segment, and the maximum confidence level is taken as the score of the merged segment.

[0075] S5, BeiDou absolute timestamp synthesis and signal feature structured output.

[0076] S5-1, Absolute timestamp synthesis: Based on the start and end frame index of the transmission segment, the adjacent BeiDou absolute time anchor points are retrieved in the bidirectional mapping table, and the start and end times are converted into absolute timestamps with sub-microsecond precision through linear interpolation. For each signal transmission segment processed by S4, based on its start and end frame indices, the adjacent BeiDou absolute time anchor points are retrieved from the bidirectional mapping table established in S1-2. Relative frame coordinates are then converted to absolute timestamps via linear interpolation. Each mapping entry contains (BeiDou integer second timecode, second pulse phase count value, frame index, intra-frame sampling offset), where the second pulse phase counter is the count value (0~9,999,999) of a 10MHz sampling clock within each second pulse period, providing 0.1... The accuracy reaches sub-microsecond levels. After interpolation, the absolute timestamp accuracy reaches the sub-microsecond level.

[0077] S5-2, Signal-to-noise ratio (SNR) estimation: Extract the ratio of the mean attention weights of the signal frequency region to the noise frequency region from the frequency attention map generated in S2 to obtain the SNR estimation. From the frequency attention map generated during the frequency attention enhancement process, the attention weights corresponding to the signal frequency region are extracted, and their mean is calculated as the signal response intensity; the attention weights of the noise frequency region outside the signal frequency region are extracted, and their mean is calculated as the noise response intensity; the ratio of the two is calculated to obtain the signal-to-noise ratio estimate.

[0078] S5-3, Frequency drift variance extraction: The frequency drift variance is obtained by statistically analyzing the variance along the time axis from the sampling position offset sequence learned by the deformable convolution in S3 along the frequency axis. From the sampling position offset sequence learned by deformable convolution along the frequency axis, the offsets belonging to the same signal transmission segment are extracted, and the variance is calculated along the time axis to characterize the transmitter frequency stability.

[0079] S5-4. Calculate the symbol rate, dot-to-stroke ratio, symbol duration consistency score, and on / off ratio from the symbol duration statistics results of S4-5 to form the signal fingerprint feature vector. The symbol duration consistency score is obtained by weighting the short-time symbol duration variation coefficient, the long-time symbol duration variation coefficient, and the character-level interval duration variation coefficient, and is used to determine whether different signal transmission segments come from the same transmission source.

[0080] The fingerprint feature vector construction process includes: (1) Estimate the symbol rate; (2) Dot-to-stroke ratio: The ratio of the long-time symbol mean to the short-time symbol mean in a symbol signal cluster; (3) Symbol duration consistency score: It is calculated by weighted sum of short-time symbol duration variation coefficient (short-time symbol duration standard deviation / short-time symbol duration mean), long-time symbol duration variation coefficient and character-level interval duration variation coefficient, and is used to determine whether different signal transmission segments come from the same transmission source; (4) On / off ratio: The ratio of the total duration of the signal on period to the total duration of the transmission segment.

[0081] The fingerprint features described above are combined with the S5-2 signal-to-noise ratio estimate and the S5-3 frequency drift variance to form a signal feature vector. This vector, along with the transmission segment semantic labels (character-level fragments / word-level / complete communication-level) and the number of original detection boxes participating in the merging, serves as the structured detection result output. Each final detection result output includes: absolute start timestamp, absolute end timestamp, center frequency, frequency upper and lower bounds, confidence score, estimated symbol rate, symbol duration statistics, dot-to-dash ratio, on / off ratio, signal-to-noise ratio estimate, frequency drift variance, symbol duration consistency score, transmission segment type, number of detection boxes participating in the merging, and signal feature vector.

[0082] To verify the effectiveness of the method, this embodiment also performed model detection in single-segment signal, multi-segment signal, and multi-channel scenarios.

[0083] like Figure 4The figure shows a schematic diagram of the model detection results for a single signal segment. This figure illustrates the detection performance of the shortwave signal detection method in a real-world shortwave narrowband keying signal scenario. The time-frequency background is shortwave band data time-aligned with the BeiDou second pulse signal at 15:03 on July 8, 2020. The horizontal axis represents absolute time (accurate to the microsecond level), and the vertical axis represents frequency. The solid-line boxes in the figure represent the model detection results, which highly overlap with the real signal, indicating that the model can accurately detect the target signal in the 0-1500Hz strong background interference band. The time deviation between the detection box and the ground truth box is less than 0.5 seconds, and the frequency boundaries are closely aligned. The structured detection results for this signal are as follows: Absolute start timestamp 2020-07-08 15:03:12.345678, Absolute end timestamp 2020-07-08 15:03:18.912345, Center frequency 1023.5Hz, Frequency upper and lower bounds [986.0, 1061.0]Hz, Confidence score 0.937, Estimated symbol rate 58WPM, Symbol duration statistical point mean 40.2ms (σ=3.1) / stroke mean 121.5ms (σ=5.8), Point-to-stroke ratio 1:3.02, On / off ratio 0.62, Estimated signal-to-noise ratio 4.8dB, Frequency drift variance 2.3 The results show the following: symbol duration consistency score of 0.89; transmission segment type of complete communication level; number of detection boxes involved in merging of 15; and signal feature vector [4.8, 2.3, 58, 3.02, 0.62, 0.89]. These results validate the precise location capability of timing-axis self-attention for signal timing boundaries, the enhancement effect of three-dimensional attention mechanism on frequency selectivity, the overall perception capability of centerline modeling for time-frequency structure, the ability of three-level semantic post-processing to reconstruct complete communication transmission segments, and the effective extraction of transmitter fingerprint features. The absolute timestamp accuracy reaches the microsecond level, meeting the timing requirements for shortwave signal monitoring.

[0084] like Figure 5 The figure shows a schematic diagram of the model detection results for multiple signal segments. This diagram illustrates the detection performance of the shortwave signal detection method in a multi-target scenario. The time-frequency background is shortwave band data time-aligned with the BeiDou second pulse signal at 15:00 on July 8, 2020. The solid lines in the figure represent the model detection results; both the long-time signal on the left and the short-time signal on the right are independently detected by the model in the form of bounding boxes, and the bounding boxes of the two signal segments highly overlap with the corresponding real signals. Both signal segments output complete structured detection results (including absolute timestamps, symbol features, and fingerprint feature vectors), verifying the model's ability to independently locate the temporal boundaries of multiple signal segments, the adaptability of centerline modeling to signal length differences, and the multi-target perception capability of the 3D attention mechanism in complex backgrounds.

[0085] like Figure 6The figure shows a schematic diagram of the model detection results in a multi-channel scenario. This diagram demonstrates the batch detection performance of the shortwave signal detection method across nine independent channels, covering various typical scenarios including no signal, short signal, long signal, and single / multiple signal combinations. The solid lines represent model predictions: channels 1 and 6 represent a pure noise background with no signal, where the model correctly rejects detections without false alarms; channels 2, 4, 5, and 7 represent single-signal scenarios, where the detection boxes highly overlap with the actual signals; channels 3, 8, and 9 represent dense multi-signal scenarios, where the model can detect parallel signals separately with independently separated bounding boxes. The detection results for each channel include a microsecond-level absolute timestamp aligned with the BeiDou second pulse signal and a complete signal feature vector, verifying the model's detection stability, temporal consistency, and generalization ability under complex electromagnetic environments in multi-channel parallel processing scenarios.

[0086] Example 2 Another embodiment of the present invention discloses a shortwave signal detection system, such as Figure 7 As shown, it includes: a BeiDou time synchronization module, a shortwave radio frequency direct acquisition receiver, a signal processing unit, and a data storage unit; among which, The BeiDou timing module, including the BeiDou antenna and timing module, is used to receive BeiDou satellite signals and output clock information and BeiDou second pulse signals; The shortwave radio frequency direct sampling receiver includes a shortwave reconnaissance antenna and a radio frequency direct sampling receiver, which is connected to the Beidou time synchronization module via a clock line. It is used to receive the clock information as a sampling reference and the Beidou second pulse signal as a sampling trigger reference, and to collect broadband I / Q data in the target frequency band so that the starting sampling point of each channel sampling sequence is aligned with the absolute whole second. The signal processing unit is connected to the shortwave radio frequency direct acquisition receiver via a high-speed data bus interface. It is used to receive the broadband I / Q data, execute the shortwave signal detection method as described in Embodiment 1, and output the signal detection result with an absolute timestamp. The data storage unit is connected to the signal processing unit via a high-speed data bus interface and is used to store the signal detection results and the corresponding time-frequency data.

[0087] In this embodiment, the signal processing unit is an industrial control computer or server with an integrated GPU accelerator card, used to carry out the deep learning model inference calculation in the method described in Embodiment 1; the shortwave radio frequency direct sampling receiver adopts a direct radio frequency sampling architecture, with a sampling rate of not less than twice the bandwidth of the target frequency band, to satisfy the Nyquist sampling theorem; the data storage unit adopts a large-capacity solid-state storage array, supporting real-time writing of time-frequency data frames and playback of historical data.

[0088] In this embodiment, the more specific technical details and beneficial effects are the same as in Embodiment 1. Please refer to them for details, and they will not be repeated here.

[0089] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for detecting shortwave signals, characterized in that, include: S1. The shortwave receiver uses the BeiDou second pulse signal as the sampling trigger reference to align the starting sampling point of each channel sampling sequence with the absolute integer second of BeiDou. The collected shortwave time-domain signal is transformed by time-frequency to obtain time-frequency data frames, and a two-way mapping table between the BeiDou absolute time anchor point and the frame coordinates is established. After preprocessing, multi-scale feature extraction is performed on the time-frequency data frames to output multi-scale feature maps; S2. Apply attention weights layer by layer along the channel, frequency and time dimensions to the multi-scale feature map to obtain the attention-enhanced feature map; S3. Perform centerline attribute prediction on the attention-enhanced feature map, output the centerline heatmap, position offset map and height map, and synthesize the initial bounding box; S4. Perform communication semantic level post-processing on the initial bounding box, sequentially performing frequency domain sidelobe suppression, symbol level merging and transmission segment level aggregation to restore the fragmented detection box into a complete communication transmission segment. S5. Convert the start and end times of the transmission segment into absolute timestamps according to the bidirectional mapping table. The output includes structured detection results containing absolute timestamps, center frequency, symbol parameters, signal-to-noise ratio estimates, and signal feature vectors. S3 includes: The attention-enhanced feature map is input into a three-branch centerline attribute prediction network; the first layer of each branch uses deformable convolution to learn the spatial offset, and the middle layers use multi-scale dilated convolution to expand the receptive field; wherein... The sampling position offset of the deformable convolution is subject to anisotropic constraints; the offset along the time axis is limited to a first preset range, corresponding to the physical range of symbol duration jitter; the offset along the frequency axis is limited to a second preset range, corresponding to the physical range of Doppler frequency shift and frequency drift; the first preset range is smaller than the second preset range. The first branch is used for temporal axis centerline heatmap prediction. A temporal axis self-attention layer is inserted along the time axis, expanding the feature map after dilated convolution into a frequency row-level time series. Scaling dot product attention is independently performed on each frequency slice, projected into a query matrix, key matrix, and value matrix by three parallel one-dimensional convolutional layers. A lower triangular causal mask is applied when calculating the attention score, ensuring that the current time step only calculates correlation weights with historical time steps to capture the irreversible temporal rhythm of the signal. After the temporal axis self-attention output, it is predicted point-by-point by a one-dimensional convolutional layer along the time axis. The final layer is activated to output a centerline heatmap, where each pixel value represents the confidence probability that the corresponding time-frequency position belongs to the horizontal centerline of the narrowband signal. The second branch is used for outputting the position offset map; the last layer of the second branch is a convolutional layer, which outputs the position offset map. The third branch is used for height map output; the last layer of the third branch is a convolutional layer, which outputs the height map.

2. The shortwave signal detection method according to claim 1, characterized in that, S1 includes: S1-1: The BeiDou timing module receives the BeiDou satellite timing signal and outputs a reference clock to the shortwave receiver as the sampling clock reference. Simultaneously, it outputs a BeiDou second pulse signal as the sampling trigger reference. The shortwave receiver, within the target frequency band, uses the rising edge of the BeiDou second pulse signal as the sampling trigger time to acquire broadband I / Q data, aligning the starting sampling point of each channel's sampling sequence with the absolute integer second. The broadband I / Q data is then subjected to parallel multiplexing to obtain a narrowband time domain signal. S1-2. Perform sliding window framing, short-time Fourier transform, logarithmic compression, global percentile normalization, and time stamping preprocessing on the narrowband time domain signal to obtain a time-stamped single-channel time-frequency map; using the BeiDou second pulse as a reference, record the frame index and intra-frame sampling point offset of the rising edge of the second pulse in each time-frequency data frame to establish a bidirectional mapping table between the BeiDou absolute time anchor point and the frame coordinates. S1-3. Input the single-channel time-frequency map into the asymmetric multi-scale feature encoder. Extract deep feature maps through multi-stage asymmetric depth-separable convolution. Apply temporal axial self-attention to the deep features to learn the on / off rules of symbols. Then, upsample through transposed convolution and fuse with the feature maps of the preceding stages by skip connection. Embed multi-scale hollow spatial pyramid pooling aggregation context in the intermediate layer to output a multi-scale feature map that has both high-level semantics and low-level details.

3. The shortwave signal detection method according to claim 2, characterized in that, The asymmetric multi-scale feature encoder includes: An asymmetric feature extraction backbone module is connected to a single-channel time-frequency map. Feature extraction and spatial downsampling are performed through multiple stages of asymmetric depthwise separable convolutional residual units. The asymmetric depthwise separable convolutional residual units model the time-frequency map in two dimensions by decoupling one-dimensional deep convolution on the time axis and frequency axis. The model is then fused through pointwise convolution and connected to the input residual. Each stage is downsampled using independent max pooling. The final stage outputs a deep feature map to the temporal axis self-attention module, and the intermediate feature maps from each stage are output to the transposed convolutional decoder module for skip connection fusion. The temporal axis self-attention module is connected to the deep feature map, decouples the deep feature map into a frequency row-level time series according to the frequency dimension, shares the same set of self-attention parameters for all frequency rows to learn the symbol rhythm pattern independent of the carrier frequency, and outputs the temporal enhancement feature map to the transposed convolutional decoder module after restoring the feature map size. The transposed convolutional decoder module receives the temporal enhancement feature map, performs multi-stage transposed convolutional upsampling, and merges it with the corresponding stage feature map of the asymmetric feature extraction backbone module through channel splicing. The intermediate layer feature map is output to the hollow space pyramid pooling module, and the final upsampling outputs a multi-scale feature map. The hollow spatial pyramid pooling module receives the feature map from the intermediate layer of the decoder. It captures the multi-scale receptive field context through parallel multi-diffraction convolutional branches. After channel concatenation and compression, it is fused with the decoder output residual. The enhanced feature map is then returned to the transposed convolutional decoder module for further upsampling.

4. The shortwave signal detection method according to claim 3, characterized in that, The time-series axial self-attention module includes: The frequency row decoupling submodule decouples the deep feature map into independent time series according to the frequency row. The multi-head self-attention submodule applies layer normalization to the time series and then performs multi-head self-attention processing. Each head projects the channel dimension features into a low-dimensional query matrix, key matrix, and value matrix through three unbiased convolutional kernels, calculates scaled dot product attention, and concatenates the outputs of each head. The channel dimension is then restored by convolutional fusion and connected to the original input residual. The feedforward network submodule processes the output of the multi-head self-attention submodule using a feedforward network. The feedforward network contains two convolutional kernels, which are activated by GELU. The intermediate dimension is expanded and then compressed and restored, and connected with the input residual of the feedforward network to obtain the feedforward enhanced features. The feature recombination submodule reassembles the feedforward enhanced features into a two-dimensional feature map according to the frequency row, and outputs the temporal enhanced feature map to the transposed convolutional decoder module.

5. The shortwave signal detection method according to claim 4, characterized in that, The transposed convolutional decoder module includes: The UpConvBlock3 module receives the temporal enhancement feature map, which is then upsampled by transposed convolution and fused with the corresponding stage feature map of the asymmetric feature extraction backbone module through channel splicing, and output to the UpConvBlock2 module. The UpConvBlock2 module receives the output of the UpConvBlock3 module, continues to upsample and merges it with the corresponding stage feature map of the asymmetric feature extraction backbone module through channel splicing. The main output is sent to the UpConvBlock1 module, and the intermediate layer feature map is output to the hollow space pyramid pooling module. The UpConvBlock1 module receives the output of the UpConvBlock2 module and fuses it with the enhanced feature map returned by the Hollow Space Pyramid Pooling module. It then continues to upsample and merges the feature map with the corresponding stage feature map of the Asymmetric Feature Extraction Backbone Module through channel splicing before outputting it to the output sub-module. The output submodule performs final-stage upsampling on the output of the UpConvBlock1 module to output multi-scale feature maps.

6. The shortwave signal detection method according to claim 1, characterized in that, S2 includes: S2-1. Apply attention weights along the channel dimension to the multi-scale feature map to obtain a channel-weighted feature map; The multi-scale feature map is subjected to global adaptive average pooling and global max pooling along the spatial dimension, and channel weight vectors are generated by sharing two layers of MLP and Sigmoid activation. These weight vectors are then multiplied with the original feature map channel by channel to output a channel-weighted feature map. S2-2. Apply attention weights along the frequency dimension to the channel-weighted feature map to obtain the frequency-weighted feature map; After pooling the channel-weighted feature map along the time axis, a frequency attention map is generated by a one-dimensional convolution kernel along the frequency axis. After activation, the frequency attention map is multiplied by the channel-weighted feature map in the form of residuals to output the frequency-weighted feature map. S2-3. Apply attention weights along the time dimension to the frequency-weighted feature map to obtain the attention-enhanced feature map; After pooling the frequency-weighted feature map along the frequency axis, a temporal attention map is generated by a one-dimensional convolution kernel along the time axis. After activation, the attention map is multiplied by the frequency-weighted feature map in the form of residuals to output the attention-enhanced feature map.

7. The shortwave signal detection method according to claim 1, characterized in that, S4 includes: S4-1. After binarizing the centerline heatmap, the connected components are marked. The optimal time position is located by accumulating column confidence. The position offset and height mean are combined to map back to the original spectral coordinates to generate a signal bounding box without redundancy. S4-2. For detection box pairs that overlap in time and whose frequency center distance is less than or equal to the first frequency threshold, suppress the detection box with lower confidence. S4-3. Group the detection frames whose center frequency difference is less than or equal to the second frequency threshold by frequency. After sorting them by time within each group, merge adjacent detection frames whose time interval is less than or equal to the character-level interval threshold into a character-level transmission segment. The character-level interval threshold is adaptively equal to N1 times the estimated symbol duration of the current frequency group. N1 is determined according to the standard ratio between character-level interval and symbol duration in the target signal communication protocol. S4-4. Merge character-level transmission segments with time intervals less than or equal to the inter-segment interval threshold into a complete communication transmission segment; the inter-segment interval threshold is adaptively equal to N2 times the estimated symbol duration, N2 is greater than N1, and N2 is determined according to the standard proportional relationship of the inter-segment interval in the target signal communication protocol. S4-5. Perform cluster analysis on the duration of each detection frame and the interval duration between frames within the current frequency group, and classify them into symbol signal clusters, character-level interval clusters and transmission segment-level interval clusters. Use the mean of the short-time symbol clusters in the symbol signal clusters as the estimated symbol duration for adaptive configuration of the character-level interval threshold and the transmission segment interval threshold.

8. The shortwave signal detection method according to claim 7, characterized in that, S5 includes: S5-1. Based on the start and end frame index of the transmission segment, retrieve the adjacent BeiDou absolute time anchor points in the bidirectional mapping table, and convert the start and end times into absolute timestamps with sub-microsecond precision through linear interpolation. S5-2. Extract the ratio of the mean attention weights of the signal frequency region to the noise frequency region from the frequency attention map generated in S2 to obtain the signal-to-noise ratio estimate. S5-3. Statistically calculate the variance along the time axis from the sampling position offset sequence learned along the frequency axis direction from the deformable convolution in S3, and obtain the frequency drift variance. S5-4. Calculate the symbol rate, dot-to-stroke ratio, symbol duration consistency score, and on / off ratio from the symbol duration statistics results of S4-5 to form the signal fingerprint feature vector. The symbol duration consistency score is obtained by weighting the short-time symbol duration variation coefficient, the long-time symbol duration variation coefficient, and the character-level interval duration variation coefficient, and is used to determine whether different signal transmission segments come from the same transmission source.

9. A shortwave signal detection system, characterized in that, include: The system includes a BeiDou time synchronization module, a shortwave radio frequency direct acquisition receiver, a signal processing unit, and a data storage unit; among which... The BeiDou timing module, including the BeiDou antenna and timing module, is used to receive BeiDou satellite signals and output clock information and BeiDou second pulse signals; The shortwave radio frequency direct sampling receiver includes a shortwave reconnaissance antenna and a radio frequency direct sampling receiver, which is connected to the Beidou time synchronization module via a clock line. It is used to receive the clock information as a sampling reference and the Beidou second pulse signal as a sampling trigger reference, and to collect broadband I / Q data in the target frequency band so that the starting sampling point of each channel sampling sequence is aligned with the absolute whole second. The signal processing unit is connected to the shortwave radio frequency direct acquisition receiver via a high-speed data bus interface, and is used to receive the broadband I / Q data, execute the shortwave signal detection method as described in any one of claims 1-8, and output the signal detection result with an absolute timestamp. The data storage unit is connected to the signal processing unit via a high-speed data bus interface and is used to store the signal detection results and the corresponding time-frequency data.

Citation Information

Patent Citations

  • Portable high-precision short wave monitoring direction-finding system

    CN117375742A

  • Fusion networking method and system based on satellite communication and short-wave communication

    CN118233936A