Single pulse candidate recognition based on deep learning
By adopting a deep learning-based recognition method in the field of single-pulse search, combining CNN and Transformer models, the problem of difficulty in extracting non-periodic pulse signal features is solved, and the recognition effect of high accuracy and low false alarm rate is achieved.
Patent Information
- Application Number
- CN202510320112.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-13
AI Technical Summary
In the field of single pulse search, machine learning is relatively few applications, and non-periodic pulse signal feature extraction is difficult, resulting in low recognition accuracy.
The single-pulse candidate recognition method based on deep learning is adopted, combined with CNN and Transformer models, and the recognition accuracy is improved through feature fusion and optimization architectures (such as CoAtNet, MBConv module, and ECA module).
A 99.59% recognition accuracy, a 0.42% false positive rate and a 99.6% recall rate were achieved, which was significantly better than the existing methods, improving the recognition accuracy of single-pulse candidates and reducing the false positive rate.
Smart Images

Figure CN120145157A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of radio astronomy, specifically to the recognition of single-pulse candidates based on deep learning. Background Art
[0002] The "Sky Eye" FAST radio telescope in China generates millions of pulsar candidates during a single sky survey, with a huge amount of data. Traditional recognition methods are difficult to handle. In the FAST sky survey data, the proportion of interference signals is large, and the pulse signals are extremely few. Moreover, the existing models have a low recognition accuracy for pulsar candidates.
[0003] Pulsar search is divided into periodic search and single-pulse search. Periodic search identifies periodic signals through techniques such as fast Fourier transform; single-pulse search uses filters to screen candidates. Although a large number of interference signal candidates will be generated, it can discover non-periodic pulse signals such as rotating radio transients (RRATs) and fast radio bursts (FRBs). However, due to the difficulty of extracting the features of non-periodic pulse signals, the application of machine learning in the field of single-pulse search is relatively less. Although there have been some attempts to apply machine learning to single-pulse search, there are still many deficiencies and urgent improvements are needed. Therefore, we propose the recognition of single-pulse candidates based on deep learning. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] In view of the deficiencies of the prior art, the present invention provides the recognition of single-pulse candidates based on deep learning, which has the advantages of improving the recognition accuracy of single-pulse candidates, etc., and solves the problems that due to the difficulty of extracting the features of non-periodic pulse signals, the application of machine learning in the field of single-pulse search is relatively less, and although there have been some attempts to apply machine learning to single-pulse search, there are still many deficiencies and urgent improvements are needed.
[0006] (2) Technical Solutions
[0007] To achieve the above purpose of improving the recognition accuracy of single-pulse candidates, the present invention provides the following technical solutions: The recognition of single-pulse candidates based on deep learning includes the following steps:
[0008] S1. Use Sigproc to convert the Fits file in CRAFTS data into a Filterbank format file;
[0009] S2. Use Heimdall to generate a group of Cand files, and the Cand files store information such as the signal-to-noise ratio of the pulse signal, candidate sample index, candidate time offset, number of filters, DM index, DM, earliest sample number, and latest sample number;
[0010] S3. Use the Python package Your to convert the signal information in the Cand file into an H5 format file;
[0011] S4. Perform feature fusion on the two features of frequency-time and DM-time in the H5 file;
[0012] S5. Input the fused features into a model that combines CNN and Transformer for recognition.
[0013] Preferably, the model that combines CNN and Transformer is based on CoAtNet, and feature fusion is performed on the basis of CoAtNet.
[0014] Preferably, the model that combines CNN and Transformer includes an MBConv module. The MBConv module uses a 1×1 convolutional kernel to perform convolution on the input channels, then uses a k×k (3×3 or 5×5) convolutional kernel, passes it into the SE module for compression, excitation, and multiplication operations, and then performs convolution through a 1×1 convolutional kernel, and finally outputs through a Dropout layer.
[0015] Preferably, the SE module in the MBConv module can be replaced by an ECA module. The ECA module uses a one-dimensional convolutional kernel after the global average pooling layer, and the convolutional kernel size k is adaptively changed through the formula:
[0016]
[0017] Adaptive change.
[0018] Preferably, the model that combines CNN and Transformer is trained and recognized using one of the three architectures of C-C-C-T, C-C-T-T, and C-T-T-T.
[0019] Preferably, for the model that combines CNN and Transformer, when training the model, 1:1 positive and negative samples are selected as the training set to avoid the problem of poor performance caused by unbalanced data sets.
[0020] Preferably, the model that combines CNN and Transformer uses accuracy, false positive rate, and recall rate as model performance indicators. The accuracy formula is:
[0021]
[0022] The false positive rate formula is:
[0023]
[0024] The recall rate formula is:
[0025]
[0026] Preferably, after the model combining CNN and Transformer is trained, based on the known energy distribution law of pulsar signals in different frequency bands, the energy data of the identified single-pulse candidates in each frequency band are compared and analyzed. If the energy distribution of the candidates deviates from the known law by more than a preset threshold, it is determined as interference signal and excluded.
[0027] (III) Beneficial Effects
[0028] Compared with the prior art, the present invention provides the identification of single-pulse candidates based on deep learning, having the following beneficial effects:
[0029] 1. For the identification of single-pulse candidates based on deep learning, through the selection and experimental verification of large-scale training sets and test sets, the C-C-T-T architecture is determined as the optimal solution. In the experiment for CRAFTS data, this method achieves an identification accuracy of 99.59%, a false positive rate of 0.42%, and a recall rate of 99.6%, significantly superior to other existing methods. Specifically, first, the Sigproc tool is used to convert the original Fits file into the Filterbank format, then Heimdall is used to generate the Cand file containing key information, and it is converted into the H5 format with the help of the Python package Your for subsequent processing. Finally, these features are input into the optimized model for identification. This method not only improves the identification accuracy of single-pulse candidates, but also effectively reduces the false alarm rate and enhances the recall ability for real pulse signals, providing a more efficient and accurate research tool for the field of radio astronomy.
[0030] 2. For the identification of single-pulse candidates based on deep learning, the model combining CNN and Transformer particularly adopts the MBConv module and the ECA module, reducing the complexity while maintaining high performance, further enhancing the feature extraction ability and the model generalization ability. Through a series of comparative experiments, it is proved the superior performance of the C-C-T-T architecture under different signal-to-noise ratio conditions. Especially when the SNR is bounded by 8, its performance indicators exceed those of the FETCH series models. This shows that the present invention can not only efficiently process a large amount of data, but also accurately identify single-pulse candidates in a complex radio astronomy environment, greatly promoting the development of pulsar-related research. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 Schematic diagram of the method for identifying single-pulse candidates of the present invention;
[0032] Figure 2 Diagram of the MBConv module of the present invention;
[0033] Figure 3 Schematic diagrams of the accuracy rate, false positive rate, and recall rate of datasets of different scales of the present invention. Detailed implementation manners
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the embodiments and drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0035] Please refer to Figures 1 - 3 , the single-pulse candidate recognition based on deep learning includes the following steps:
[0036] S1. Use Sigproc to convert the Fits file in the CRAFTS data into a Filterbank format file;
[0037] S2. Use Heimdall to generate a set of Cand files, and the Cand files store information such as the signal-to-noise ratio, candidate sample index, candidate time offset, number of filters, DM index, DM, earliest sample number, and latest sample number of the pulse signal;
[0038] S3. Use the Python package Your to convert the signal information in the Cand file into an H5 format file;
[0039] S4. Perform feature fusion on the two features of frequency-time and DM-time in the H5 file;
[0040] S5. Input the fused features into a model combining CNN and Transformer for recognition.
[0041] Embodiment 1:
[0042] Currently, convolutional neural networks are more commonly used in single-pulse searches, and the attention mechanism has not been widely applied. CNN has strong generalization ability and is good at extracting local features, but its ability to model global information is relatively weak and its invariance to transformations such as rotation is poor. While Transformer can effectively process global information and adjust corresponding weights according to different inputs, but it lacks the inductive bias characteristics of CNN and cannot effectively extract local features, resulting in poor generalization ability. The present invention can give full play to the respective advantages by combining CNN and Transformer, and effectively improve the model capacity and generalization ability.
[0043] MBConv is an inverted linear bottleneck layer with depthwise separable convolutions, a module introduced in MobileNetV2, as Figure 2 shown. The MBConv module of the present invention adopts a lightweight design, an inverted residual structure, dilated convolutions, and depthwise separable convolutions, and can effectively extract multi-scale feature information, featuring high efficiency and flexibility. MBConv first uses a 1×1 convolutional kernel to convolve one channel of the input, and one convolutional kernel is only responsible for one channel; then it uses a k×k (3×3 or 5×5) convolution, then passes it into the SE module for compression, excitation, and multiplication operations; then it is convolved through a 1×1 convolutional kernel; and finally, it outputs through a Droupout layer.
[0044] In the present invention, the ECA module is an efficient channel attention module for enhancing the performance of convolutional neural networks. This module directly uses a one-dimensional convolution after the global average pooling layer. The ECA module avoids dimensional reduction and effectively captures the interaction relationships between all channels with only a small number of parameters involved. Among them, the convolutional kernel size k adaptively changes through a function, and k is defined as:
[0045]
[0046] The ECA module can effectively maintain performance while reducing its complexity through an adaptive convolutional kernel, and has better feature extraction ability and generalization ability.
[0047] The Transformer mechanism is a method of extracting feature information using multi-head self-attention layers and feed-forward neural network layers. Self-Attention is the core component of the Transformer. Its main purpose is to capture the dependencies of the sequence and achieve global dependency modeling. Among them, the multi-head attention module calculates self-attention with different weights, connects the results, and uses another matrix to adjust the results to the embedding dimension and passes it into the next module. Simply put, Self-Attention aggregates the result obtained by aggregating its query and the corresponding key with the value again and outputs a vector. Its core formula is:
[0048]
[0049] In the formula, Q represents Query, K represents Key, and V represents Value. The Self-Attention mechanism can process the entire sequence, is not limited by the local window size, and can perform parallel calculations, which is particularly suitable for reducing the computational complexity when dealing with long sequence data.
[0050] Example 2:
[0051] The model combining CNN and Transformer of the present invention is based on CoAtNet. The output of the CoAtNet model is two columns of numerical values. CoAtNet compares the numerical values with index 0 and index 1, and takes the index value with the larger numerical value as the prediction result. Therefore, four classification results based on the binary classification confusion matrix are used to calculate the performance metrics of the model. The four classification results are as follows:
[0052] (1) True Positives (TP): The number of signals that can be correctly classified as pulse signals;
[0053] (2) False Positives (FP): The number of pulse signals misclassified as interference signals, which is called the number of false positives;
[0054] (3) True Negatives (TN): The number of signals that can be correctly classified as interference signals;
[0055] (4) False Negatives (FN): The number of interference signals misclassified as pulse signals, which is called the number of false negatives.
[0056] In the experiment, accuracy, false positive rate, and recall rate are used as the performance metrics of the model. Accuracy is the ratio of the number of signals correctly classified as pulse signals and interference signals in the test set to the total number of candidates, representing the ability to correctly classify candidates. Accuracy is defined as:
[0057]
[0058] The false positive rate is the ratio of the number of interference signals misclassified as pulse signals in the test set to the number of true interference signals. The false positive rate is defined as:
[0059]
[0060] The recall rate is the ratio of the number of signals correctly classified as pulse signals in the test set to the number of true pulse signals. The recall rate is defined as:
[0061]
[0062] Example 3:
[0063] In the CRAFTS data files of the FAST sky survey project, 239 Fits files were selected. After being processed through the single-pulse search process, they were converted into 616,256 candidate files in H5 format. A part of these files was selected as our training set and test set. Most of the candidate files are interference signals. When the model is trained on an imbalanced dataset, it will overtrain on the interference signals, causing overfitting problems, and will misclassify a large number of real single-pulse signals as interference signals. To avoid the problem of poor performance of the deep learning classifier caused by the imbalanced dataset, a 1:1 positive and negative sample ratio was selected as the training set.
[0064] When a model has a higher model capacity, it can better fit the details and complexity in the training data. However, when the training data is insufficient, it may cause overfitting problems, resulting in poor generalization ability of the model. Therefore, in this paper, experiments were conducted using training sets of different sizes to understand how the data size of the training set affects the performance of the model. In this paper, three architectures of the model were trained using training set sizes of 2000, 5000, 8000, 12000, 15000, and 18000 respectively, and then the performance of the model was evaluated using the same test set. The test set consists of 5000 real pulse signals and 5000 RFI signals, and their accuracy, false positive rate, and recall rate were calculated. The experimental results are as Figure 3 shown.
[0065] It can be seen from Figure 3 that when the size of the training set exceeds 15000, the accuracy of the model tends to be stable, indicating that the proposed model requires a large amount of training data, at least 15000 samples to achieve optimal performance. When the dataset is small, the accuracy of C-C-C-T is higher, and when the dataset is large, the accuracy of C-C-T-T is higher.
[0066] Pulsar signals and interference signals have different strengths. Strong pulsar signals are more likely to be recognized as real pulsar signals by the model, and strong interference signals are more likely to be misjudged by the model. To examine how the signal-to-noise ratio (SNR) of the signal affects the performance of the model, based on a 1:1 positive and negative sample ratio in the training set, the training set with a sample size of 15000 was trained at a ratio of 1:1 with SNR <= 7 and SNR > 7, SNR <= 8 and SNR > 8, SNR <= 9 and SNR > 9, and SNR <= 10 and SNR > 10 respectively. Tests were conducted on the same test set (5000 real pulse signals and 5000 RFI signals), and their accuracy, false positive rate, and recall rate were calculated.
[0067] As shown in the table:
[0068] Table 1 SNR bounded by 7
[0069]
[0070] Table 2 SNR is bounded by 8
[0071]
[0072] Table 3 SNR is bounded by 9
[0073]
[0074] Table 4 SNR is bounded by 10
[0075]
[0076] As can be seen from the table, the model trained with SNR bounded by 8 has the best performance on the test set. The subsequent comparative experiments will be carried out on the training set with SNR bounded by 8.
[0077] In this embodiment, the proposed method is compared with the existing method. Seven models of FETCH (composed of combinations of DenseNet, Xception, and VGG respectively) are trained using the same training set with a sample size of 15,000, and tested on the same test set (5,000 real pulsar signals and 5,000 RFI signals). The test results are shown in Table 4.
[0078] Table 4 Comparison of test results
[0079]
[0080] As can be seen from the table, the accuracies of C-C-C-T, C-C-T-T, and C-T-T-T are 0.31%, 0.25%, and 0.14% higher than that of Model A with the highest accuracy in FETCH respectively, which proves that the method proposed in the present invention is more effective.
[0081] On this basis, the model is optimized by replacing the SE module in MBConv with the ECA module. The ECA module can effectively maintain performance while reducing its complexity through an adaptive convolution kernel, and has better feature extraction ability and generalization ability. The test results are shown in the following table:[[]]
[0082]
[0083] For the optimized model, the accuracies are increased by 0.03%, 0.02%, and 0.11% respectively compared with the original model, further improving the ability of the model to identify pulsar candidates.
[0084] In summary, for the single-pulse candidate recognition based on deep learning, through the selection and experimental verification of large-scale training sets and test sets, the C-C-T-T architecture was determined as the optimal solution. In the experiments on CRAFTS data, this method achieved an identification accuracy of 99.59%, a false positive rate of 0.42%, and a recall rate of 99.6%, significantly outperforming other existing methods. Specifically, first, the Sigproc tool was used to convert the original Fits file into the Filterbank format. Then, Heimdall was utilized to generate the Cand file containing key information, and it was converted into the H5 format with the help of the Python package Your for subsequent processing. Finally, these features were input into the optimized model for identification. This method not only improves the identification accuracy of single-pulse candidates but also effectively reduces the false alarm rate and enhances the recall ability for real pulse signals, providing a more efficient and accurate research tool for the field of radio astronomy.
[0085] Moreover, for the single-pulse candidate recognition based on deep learning, the MBConv module and the ECA module were specifically adopted in the model that combines CNN and Transformer, reducing the complexity while maintaining high performance, and further enhancing the feature extraction ability and the model generalization ability. Through a series of comparative experiments, the superior performance of the C-C-T-T architecture under different signal-to-noise ratio conditions was demonstrated. Especially under the condition where the SNR is bounded by 8, its performance indicators exceeded those of the FETCH series models. This indicates that the present invention can not only efficiently process large amounts of data but also accurately identify single-pulse candidates in a complex radio astronomy environment, greatly promoting the development of pulsar-related research and solving the problems that due to the difficulty in extracting features of non-periodic pulse signals, the application of machine learning in the field of single-pulse search is relatively less, and although there have been some attempts to apply machine learning to single-pulse search, there are still many deficiencies and urgent improvements needed.
[0086] All relevant modules involved in this system are either hardware system modules or functional modules that combine computer software programs or protocols in the prior art with hardware. The computer software programs or protocols themselves involved in this functional module are all well-known technologies to those skilled in the art, and they are not the improvements of this system; the improvement of this system lies in the interaction relationship or connection relationship between each module, that is, the overall structure of the system is improved to solve the corresponding technical problems to be solved by this system.
[0087] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. Single pulse candidate recognition based on deep learning, characterized by: The following steps are involved: S1. Use Sigproc to convert the Fits file in the CRAFTS data into a Filterbank format file; S2. Generate a set of Cand files using Heimdal l, wherein the Cand files store information about the signal-to-noise ratio of the pulse signal, the candidate sample index, the candidate time offset, the number of filters, the DM index, the DM, the earliest sample number, and the latest sample number; S3. Use the Python package Your to convert the signal information in the Cand file into an H5 format file; S4, feature fusion of the frequency-time and DM-time features in the H5 file; S5. Combine the fused feature input with the CNN and Transformer models for recognition.
2. The single pulse candidate recognition based on deep learning according to claim 1, characterized in that: The model combining CNN and Transformer is based on CoAtNet, and feature fusion is performed on the basis of CoAtNet.
3. The single pulse candidate recognition based on deep learning according to claim 1, characterized in that: The model combining CNN and Transformer includes an MBConv module, which uses a 1×1 convolution kernel to convolve the input channel, then uses k×k (3×3 or 5×5) convolution, passes it into the SE module for compression, excitation and multiplication operations, then convolves it with a 1×1 convolution kernel, and finally outputs it through a Dropout layer.
4. The single pulse candidate recognition based on deep learning according to claim 3, characterized in that: The SE module in the MBConv module can be replaced by an ECA module, which uses a one-dimensional convolution after the global average pooling layer, and the convolution kernel size k is calculated by the formula: Adapt to change.
5. The single pulse candidate recognition based on deep learning according to claim 1, characterized in that: The model combining CNN and Transformer adopts one of the three architectures of CCCT, CCTT and CTTT for training and recognition.
6. The single pulse candidate recognition based on deep learning according to claim 1, characterized in that: The model combining CNN and Transformer selects 1:1 positive and negative samples as training sets when training the model to avoid poor performance caused by an unbalanced data set.
7. The single pulse candidate recognition based on deep learning according to claim 1, characterized in that: The model combining CNN and Transformer uses accuracy, false positive rate and recall rate as model performance indicators. The accuracy formula is: The formula for false positive rate is: The recall formula is:
8. The single pulse candidate recognition based on deep learning according to claim 1, characterized in that: After the training of the combined CNN and Transformer model is completed, based on the known energy distribution patterns of pulsar signals in different frequency bands, the energy data of the identified single pulse candidates in each frequency band are compared and analyzed. If the energy distribution of the candidate deviates from the known pattern by more than a preset threshold, it is judged as an interference signal and excluded.