A coherent dispersion method for pulsar baseband data
Through GPU acceleration and multi-threaded parallel processing of the HASHPIPE architecture, combined with VDIF data frame parsing and lost frame repair technology, the problems of high computing resources, insufficient precision and poor flexibility in pulsar data processing are solved, and efficient and real-time pulsar signal processing is achieved.
Patent Information
- Application Number
- CN202411717283.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-11-27
AI Technical Summary
Existing pulsar data processing technology has problems such as high computing resource requirements, insufficient processing accuracy, poor flexibility and slow processing speed. In particular, there is a clear gap in the independent research and development of pulsar terminals and processing technology.
The HASHPIPE architecture based on GPU acceleration is adopted to achieve coherent dispersion elimination through multi-threaded parallel processing. Combined with VDIF data frame parsing and lost frame repair technology, the processing flow of pulsar baseband data is optimized.
It significantly improves the real-time and accuracy of pulsar signal processing, solves the problem of data frame loss, improves processing efficiency and data integrity, and meets the processing needs of radio telescope observation data.
Smart Images

Figure CN119671841B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of radio telescopes, and in particular relates to a coherent de-dispersion method for pulsar baseband data, which is based on a GPU and a HASHPIPE architecture. Background Art
[0002] Pulsars were one of the four great discoveries of astronomy in the 1960s, their discovery confirming the existence of neutron stars. Pulsars possess extremely strong magnetic fields, emitting radio radiation from their polar caps while rotating rapidly on their axes. When this radiation sweeps past radio telescopes on Earth, it is detected as a periodic pulse signal. By measuring pulsar signals, the arrival times (TOA) and pulse profiles of the pulses can be determined, enabling the study of extreme physical processes in the universe, such as pulsar magnetosphere dynamics, internal structure, cosmic background gravitational waves, and ultra-dense matter. Pulsars, with their stable period and extremely slow rate of change, are comparable in stability to atomic clocks, making them suitable reference sources for time calibration. Furthermore, the radiation signatures of pulsars can be used to probe the average magnetic field distribution of the interstellar medium and the Milky Way. These extreme physical properties make pulsars extremely valuable astrophysical laboratories.
[0003] However, the electromagnetic waves emitted by pulsars are affected by dispersion, scattering, and scintillation effects in the interstellar medium during propagation, resulting in variations in the average pulse width and shape of the pulsar signal received by radio telescopes. Dispersion broadens the pulse profile, scattering further blurs the signal, and scintillation causes fluctuations in pulse intensity. Of these effects, the time delay caused by dispersion is far greater than that caused by scattering. Therefore, to obtain an accurate pulse shape, pulsar observation data must be subjected to de-dispersion processing.
[0004] Currently, there are two main methods for de-dispersion: incoherent de-dispersion and coherent de-dispersion. Incoherent de-dispersion is an earlier approach that uses a multi-channel filter to split the pulsar baseband observation signal into multiple independent channels. Appropriate time delays are applied to the data in each frequency band, thereby aligning and summing the received pulse signals. However, while incoherent de-dispersion eliminates dispersion effects to a certain extent, it cannot completely eliminate dispersion within each small channel, resulting in certain limitations in high-precision observations.
[0005] Coherent dispersion cancellation was first proposed by Hankins and Rickett in 1975. Although initially unrealized due to computer limitations, its theoretical approach is unique. Coherent dispersion cancellation effectively eliminates the effects of dispersion within the passband by converting the time-domain signal into the frequency domain, multiplying it by the interstellar medium's transmission function, and then converting it back to the time domain. Advances in computer and data storage technology, particularly the parallel computing power of graphics processing units (GPUs), have enabled real-time processing of massive amounts of observational data, significantly improving both observational sensitivity and temporal resolution.
[0006] In other words, incoherent dedispersion cannot completely eliminate the residual dispersion within the frequency channel, while coherent dedispersion will not have this problem. Therefore, the pulsar baseband data after sampling and processing by the existing radio telescope terminal is more inclined to undergo coherent dedispersion processing to remove the dispersion effect introduced by the pulsar signal during the propagation of the interstellar medium, thereby improving the time resolution and signal quality of the data.
[0007] In recent years, the in-depth study of pulsars has promoted the development of research such as millimeter-wave pulsar observations and gravitational wave detection, and has put forward higher requirements on the performance of observation equipment.
[0008] Internationally, countries like Australia and the United States have gradually developed a heterogeneous pulsar digital terminal system architecture combining FPGA, CPU, and GPU. In this architecture, the FPGA is used for ADC control and signal preprocessing, the CPU handles data transmission and logic control, and the GPU executes core algorithms such as coherent dispersion cancellation. This architecture offers wide bandwidth and high-speed real-time signal processing capabilities and has been applied in multiple pulsar observation projects.
[0009] For example, the Parkes Telescope in Australia has installed multiple pulsar terminal systems, such as BPSR, CASPER, and UWL. These systems utilize the aforementioned heterogeneous architecture. The CASPER system processes dual-polarization signals using IBOBs, CPUs, and multiple GPU nodes, achieving full-bandwidth coherent de-dispersion from low to high frequencies. Similarly, the GUPPI pulsar terminal at the Green Bank Telescope (GBT) in the United States also provides a wide-bandwidth coherent de-dispersion observation mode, and the VEGAS system was developed based on this.
[0010] In China, pulsar observation equipment has been steadily developing. The Tianma 65-meter radio telescope at the Shanghai Astronomical Observatory, the Nanshan 26-meter telescope at the Xinjiang Astronomical Observatory, the 40-meter telescope at the Yunnan Astronomical Observatory, the 40-meter telescope at the National Time Service Center, and the FAST 500-meter telescope at the National Astronomical Observatory all employ incoherent or coherent dispersion elimination techniques to varying degrees. Since 1999, when the Xinjiang Nanshan Observatory introduced an analog multi-channel filter bank dispersion elimination system from the University of Manchester in the UK, it has continuously improved its pulsar observation equipment. In 2009, it introduced the Australian pulsar digital filter bank PDFB system, an incoherent dispersion elimination system.
[0011] The Tianma 65-meter radio telescope is currently equipped with multiple terminal systems, including the DIBAS (Digital Backend System) digital terminal, which is based on the VEGAS system developed by the National Radio Astronomy Observatory (NRAO). It features a pulsar observation mode similar to GUPPI and supports both incoherent and coherent de-dispersion data processing methods. The DIBAS digital terminal is an unpublished technology. The DIBAS hardware system consists of a digital acquisition unit, a high-performance computing unit, and a storage unit. It utilizes a GPU-based heterogeneous computing model to further optimize and improve coherent de-dispersion technology. Because the terminal was jointly developed by China and the United States, the data obtained from DIBAS after observations are all processed data files according to scientific objectives, not raw baseband data. As we do not yet have the technology to post-process pulsar data, we may be unable to perform subsequent observations and processing of pulsar data if the equipment unexpectedly stops working.
[0012] In other words, because Shanghai Tianma uses imported pulsar equipment, the pulsar data it observes is already processed. The details of this technology have not been made public, and they don't have coherent dispersion elimination technology. Most domestic research institutions use incoherent dispersion elimination.
[0013] In summary, there are indeed some obvious gaps in current pulsar data processing technology at home and abroad, especially in the calculation of massive data, especially in the independent research and development of pulsar terminals and processing technology. The specific shortcomings of pulsar data processing technology include:
[0014] 1. High computing resource requirements: Existing technologies often require a large amount of computing resources for high-precision data processing and dispersion reduction algorithms. Traditional computing methods require high hardware investment and may cause system overload when processing large amounts of data.
[0015] 2. Insufficient processing accuracy: Some existing technologies may not strictly control the accuracy of data during processing, resulting in errors when performing complex processing such as coherent de-dispersion of pulsar signals. This will affect the final data quality and analysis results.
[0016] 3. Poor flexibility: Traditional data processing methods often lack sufficient flexibility and adaptability, making it difficult to adjust to different data characteristics and processing requirements. Modifying algorithms and adapting to new requirements may require a lot of manual adjustments and optimizations.
[0017] GPU accelerated processing: In order to improve processing efficiency, GPU is usually used for accelerated computing in existing technologies, but the processing speed is still not fast enough. Summary of the Invention
[0018] The purpose of the present invention is to provide a coherent dedispersion method for pulsar baseband data, so as to accelerate the coherent dedispersion process by applying multi-threaded parallel processing, and further optimize the GPU acceleration processing process by using HASHPIPE architecture parallel processing to significantly improve processing performance.
[0019] To achieve the above object, the present invention provides a coherent de-dispersion method for pulsar baseband data, comprising:
[0020] S1: Based on the HASHPIPE architecture, a data transmission method is established between the data reading thread net_thread.c, the input buffer inputbuf, the multi-threaded parallel coherent de-dispersion thread, the output buffer outputbuf and the file writing thread output_thread.c;
[0021] S2: In the data reading thread net_thread.c, the data frame of the VDIF data is checked, the time information is read, and the frame header is deleted to obtain the data in the required format and store it in the input buffer inputbuf;
[0022] S3: Execute multi-threaded parallel coherent de-dispersion threads using data from the input buffer inputbuf and write the results to the output buffer outputbuf. During this process, data splitting, FFT, frequency conversion, complex multiplication, frequency conversion, IFFT, de-filling detection power, and Stokes parameter calculation are implemented in sequence through function calls;
[0023] S4: Call the file writing thread output_thread.c to write the processed data in the output buffer outputbuf into the file.
[0024] The step S1 specifically includes: step S11: using the buffer setting programs databuf.h and databuf.c to complete the configuration of the input buffer inputbuf and the output buffer outputbuf, and define the data structure related to the frame number of the VDIF data.
[0025] The size of the input buffer inputbuf is three times the amount of data processed each time, which is nvalid×nfft×nsub, where nsub refers to the number of subbands of VDIF data, nfft refers to the number of FFT data performed simultaneously in GPU parallel processing, and nvalid is the valid data length of a single subband in an FFT;
[0026] The size of the output buffer outputbuf is three times the amount of data written each time Msamp×mchan, where mchan is the total number of channels generated and Msamp is the length of the processed data for the user-set single channel number; and / or
[0027] The number of input buffers inputbuf and output buffers outputbuf is multiple.
[0028] The step S2 specifically includes:
[0029] S21: Use the data reading thread net_thread.c to read the VDIF baseband data and read the frame header information;
[0030] S22: Parse the time in the frame header of the data frame and calculate the corresponding UTC time and MJD time;
[0031] S23: Remove the frame header and store the processed data into the input buffer inputbuf that has been set.
[0032] The step S22 further includes: checking the frame number in the second in the frame header by detecting changes in the second number and the frame number per second in the data frame header, and filling the gap by generating an equal amount of white noise frames when data frames are found to be lost.
[0033] The step S3 specifically includes:
[0034] S31: Execute the main program module gpu_thread.cu to implement the initialization of the multi-threaded parallel coherent de-dispersion thread;
[0035] S32: After the main program module gpu_thread.cu completes the initialization of the multi-threaded parallel coherent de-dispersion thread, the thread status of the hashpipe framework is determined to ensure the correct operation of the thread; next, the data status of the input buffer inputbuf and the output buffer outputbuf is checked to see if they are available; if they are, step S33 is executed;
[0036] S33: Execute the main program module gpu_thread.cu to implement the multi-threaded parallel coherent de-dispersion thread process.
[0037] The step S33 specifically includes:
[0038] S331: For each data processing, the main program module gpu_thread.cu reads a specified amount of data from the input buffer inputbuf and enters the format conversion function separatevdifData, calling the format conversion function separatevdifData to implement data conversion and data arrangement processing of the VDIF data;
[0039] S332: Call the data splitting function unpack_and_padd to split the data into nfft×nsub data blocks with valid length nvalid, where nfft is the number of FFT data to be performed simultaneously in GPU parallel processing, nsub refers to the number of subbands of VDIF data, and nvalid is the valid data length of a single subband in one FFT;
[0040] S333: calling the FFT plan to perform a fast Fourier transform based on multi-threaded parallel processing on the nfft×nsub group of data blocks;
[0041] S334: calling the complex multiplication function PointwiseComplexMultiply to perform a complex multiplication operation on the data block and the chirp function generated in step S31 to obtain a data block after dispersion elimination;
[0042] S335: calling the IFFT plan to perform an inverse Fourier transform based on multi-threaded parallel processing on the data block that ultimately generates the total number of channels mchan, so as to convert the de-dispersed data back into the time domain;
[0043] S336: calling the unpadd detection power function transpose_unpadd_and_detect to remove the overlapping parts of the time domain data and calculate the power sum on the two polarization channels;
[0044] S337: Perform Stokes parameter calculation and quantization, and save to output buffer outputbuf, and terminate the current multi-threaded parallel coherent dispersion thread.
[0045] In step S334, before performing complex multiplication operation, it also includes: deleting negative frequency part and retaining positive frequency part; and / or before step S334 and after step S334, it includes: calling channel spectrum positive and negative frequency conversion function swap_spectrum_halve to flip the positive and negative frequencies of data.
[0046] The multi-threaded parallel processing includes one of GPU parallel processing, FPGA acceleration parallel processing, multi-core CPU parallel processing and special hardware processor parallel processing.
[0047] The file is a filterbank file; and the coherent dispersion method further includes step S5: performing folding drawing in dspsr pulsar software according to the generated filterbank file.
[0048] Compared with the existing domestic pulse star baseband data processing technology, the coherent dispersion method of the pulse star baseband data of the application combines GPU acceleration technology and HASHPIPE architecture, realizes efficient real-time processing of pulse star signals, especially in the application of coherent dispersion algorithm, the parallel processing capability of GPU greatly improves the calculation speed. This improvement significantly reduces the processing delay, so that the system can complete the data analysis more quickly. In addition, the application introduces VDIF data frame analysis and lost frame repair technology, ensures the integrity and continuity of data in the transmission process, solves the problem that the data frame is difficult to recover after being lost in the prior art.
[0049] The prior art often encounters problems such as insufficient processing speed and difficulty in frame loss repair when processing pulse star observation data. The application combines GPU acceleration and multi-threaded processing, not only greatly improves the processing efficiency, but also ensures the integrity of the data through intelligent frame loss detection and repair mechanism. This improvement improves the real-time performance and accuracy of pulse star signal processing, better meets the processing needs of radio telescope observation data. BRIEF DESCRIPTION OF DRAWINGS
[0050] Figure 1 is a schematic diagram of the data frame format of the existing VDIF data.
[0051] Figure 2 is a schematic diagram of the data frame header format of the existing VDIF data.
[0052] Figure 3 is a frequency distribution diagram of eight channels of left-handed or right-handed.
[0053] Figure 4 It is a schematic diagram of the principle of processing one of the two original data types.
[0054] Figure 5 It is an overall structural block diagram of the coherent de-dispersion method of pulsar baseband data of the present invention.
[0055] Figure 6 It is a structural block diagram of the coherent de-dispersion algorithm part of the coherent de-dispersion method for pulsar baseband data of the present invention.
[0056] Figure 7 It is a schematic diagram of the principle of adding and removing overlapping parts of the coherent de-dispersion method of pulsar baseband data of the present invention.
[0057] Figures 8-10 This is a result diagram of three pulsars obtained by the coherent de-dispersion method of pulsar baseband data of the present invention. DETAILED DESCRIPTION
[0058] The present invention will be further described below with reference to specific examples. It should be understood that the following examples are only used to illustrate the present invention and are not intended to limit the scope of the present invention.
[0059] The basic principle of coherent dedispersion of the pulsar baseband data coherent dedispersion method of the present invention is as follows:
[0060] The following will introduce the basic theory of coherent de-dispersion and the derivation process of the core algorithm. The pulsar radiation emitted by the pulsar will pass through the interstellar medium (ISM) before being received by the radio telescope and will be affected by its refraction, i.e. dispersion. The refractive index μ is:
[0061]
[0062] where f p is the plasma frequency, f is the electromagnetic wave frequency, when f<f p When the electromagnetic waves cannot pass through the interstellar medium, this is also the key condition that constrains whether the pulsar radiation signal can be received by the radio telescope on the ground. The plasma frequency f p It can be expressed by the following formula. The other dispersion constant D can be expressed by the following formula:
[0063]
[0064] n e The electron number density is approximately equal to 0.03 cm -3 , e represents charge, m e The values in the dispersion constant formula are all known, where c is the speed of light.
[0065] So according to the above two formulas, the plasma frequency f can be calculated p The values of and dispersion constant D are:
[0066]
[0067] The pulsar radiation signal travels through the interstellar medium at a group velocity of v g =cμ, from the formula we can know that v g <c, when the distance d from the pulsar to the Earth is the time delay t due to the interstellar medium, the following formula can be used to calculate:
[0068]
[0069] By substituting the dispersion constant D and setting the dispersion measure
[0070] The result of the last row in the above formula can be deduced.
[0071] In coherent dedispersion, the dispersion delay of the interstellar medium on the pulsar radiation signal appears as a convolution of the time domain signal, while in the frequency domain, it is equivalent to the product of the pulse radiation signal and the dispersion delay effect caused by the interstellar medium.
[0072] S(v0+v)=S initial (v0+v)×H(v0+v)
[0073] Among them, S is the signal received by the radio telescope, S initial is the original signal of pulse radiation, v0 is the center frequency of the channel, ν is the frequency offset within the channel, and H(ν0+ν) represents the transmission function of the interstellar medium in the frequency domain.
[0074] The delay effect caused by the interstellar medium can be expressed as a phase rotation that depends on the frequency and the propagation path length. In the frequency domain, the interstellar medium only has a delay effect on the frequency domain phase of the signal. The wave number is the wavelength, d is the distance the pulsar radiation signal travels in the interstellar medium, and c = λv.
[0075] By the formula Performing Taylor expansion, we can get:
[0076] So the transmission function of the interstellar medium in the frequency domain is H(ν0+ν):
[0077]
[0078] Taylor expansion of the above formula and taking the first three terms yields:
[0079]
[0080] The first of these three is a constant phase shift, the second term According to the shift theorem, it corresponds to the delay in the time domain, and the third term is the dispersion effect of the in-band signal, which is the part we need to eliminate.
[0081] Therefore, the present invention uses the previously derived dispersion DM to replace the distance d and plasma frequency f of the pulsar radiation from the pulsar to the earth. p , the final transfer function can be obtained, so that the formula can be used to perform coherent de-dispersion processing on the radio telescope received signal S.
[0082] The final transfer function is:
[0083]
[0084] Where DM represents the dispersion, v0 is the center frequency of the channel, ν is the frequency offset within the channel, and k DM represents the dispersion constant.
[0085] The basic principle of VDIF format data decoding based on the coherent de-dispersion method of pulsar baseband data of the present invention is as follows:
[0086] The baseband data generated by the VLBI terminal is in VDIF (VLBI Data Interchange Format) format, and the VLBI data packet consists of numerous data frames.
[0087] The data frame format of the VLBI data packet belongs to the existing technology. The VLBI Data Interchange Format (VDIF) Specification document explains the VDIF data format in detail. Each data frame is as follows Figure 1 As shown, a data frame includes a data frame header and a data array. A data frame includes data of one or more frequency subbands, and the frame header contains information about the data array within the data frame. This section introduces the data frame format to explain the specific effects of the data decoding program designed for this data frame format.
[0088] Among them, the basic format of the data frame header is as follows Figure 2 As shown, the information in the data frame header is the basis for subsequent data processing.
[0089] exist Figure 2In the present application, one word represents 4 bytes, and the frame header consists of 8 words, which is equivalent to 32 bytes, which is the standard 32-byte frame header. The word is only a division method, and multiple words are transmitted at the same time. The VDIF data transmission is carried out according to the structure of the data frame, and the data collected by the data acquisition terminal in real time is transmitted in the form of a data frame and saved to a data file. It should be noted that the digital acquisition and transmission part belongs to the prior art.
[0090] Word 0 (4 bytes):
[0091] Bit 31: Invalid data flag I1 (valid = 0, invalid = 1);
[0092] Bit 30: Legacy mode L1 (0 represents a standard 32-byte frame header; 1 represents a 16-byte frame header); in the present application, the data frame header format used for processing data is a standard 32-byte frame header, so when reading the frame header information, the legacy mode L1 needs to be read to confirm that it is a standard 32-byte frame header;
[0093] Bits 29-0: Seconds from reference epoch 30 .
[0094] Word 1 (4 bytes):
[0095] Bits 31-30: Un-assigned (should be set to all 0) Un-assigned2;
[0096] Bits 29-24: Reference epoch Ref Epoch6 (from 00:00 UTC on January 1, 2000, plus one every six months); the following text is an example of the data frame header information read by the program designed in the present application, where the epoch below is the reference epoch data of bits 29-24, the value is 44, which means that from January 1, 2000, it is added by 1 every six months, which is January 1, 2022, as the reference time, and the second number information is recorded in bits 23-0, starting from this time as a reference to record the second number;
[0097] Bits 23-0: Data frame number within second 24(Starting from zero, an integer number of data frames per second); the observation bandwidth of the telescope is very wide, and the amount of data generated per second is also very large. A single VDIF data frame cannot accommodate all the data. Taking the frame header information read by the program designed by the present invention as an example, the data volume of a single data frame is 8192 bytes, but according to the bandwidth and sampling rate of the telescope, the data volume of 1 second is 8*64000000 bytes, which means that multiple data frames are needed to store the data per second. The data frame number within the second records the amount of data in the current second. If the data frame number within the second is a, it means that the data currently recorded starts from (a-1)*8192 bytes in the current second.
[0098] Word 2 (4 bytes):
[0099] Bits 31-29: VDIF version number V3 (used to support up to seven future VDIF frame header formats);
[0100] Bits 28-24: Logarithm of the number of channels in the data array (log2(#chns)5) (base 2; the number of channels must be a power of 2). Existing telescope observation data is channeled. According to the frame header information below, bits 28-24 have a fraction of 4, which represents the number of channels as 2 raised to the fourth power, or 16 channels. This VDIF file records a starting frequency of 2268 MHz, a bandwidth of 64 MHz per frequency channel, and a total of eight channels, with both left-hand and right-hand rotations, or 2 × 8 = 16 channels.
[0101] Bit 23-0: Data frame length (including header) 24 , the unit is 8 bytes; the data frame number within the second above records the total number of data frames per second, and this part records the number of data stored in the data array in a single data frame.
[0102] Word 3 (4 bytes):
[0103] Bit 31: Data type C1 (0 indicates real data; 1 indicates complex data);
[0104] Bits 30-26: bits / sample-I5 (maximum 32 bits / sample);
[0105] Bits 25-16: Thread ID 10 , 0 to 1023);
[0106] Bits 15-0: Station ID 16 , a 2-character ASCII standard ID or a numeric ID assigned by a global site);
[0107] The number of bits per sample indicates the number of bytes occupied by a single data item in the data array. The thread ID and site ID are not used later, but these two items are included in the data frame header in VDIF, so they are briefly introduced here.
[0108] Word 4-7 (4 bytes): Extended User Data 32 .
[0109] In the present invention, the data type processed by the coherent de-dispersion method for pulsar baseband data is a VDIF multi-channel data type. Therefore, in the present invention, the VDIF data is pre-processed, specifically including:
[0110] Step A1: Read the data frame header information and determine the exact time point of data observation through the three parameters of the data frame header: seconds, legacymode, and epoch. Seconds is the number of seconds since the reference epoch, epoch is the reference epoch, and legacymode is the legacy mode.
[0111] Step A2: Use the three functions UTCconvertToMJD, UTC2JD, and VDIFconvertToMJD to calculate the data frame header information of the VDIF data into UTC time and MJD time;
[0112] UTC and MJD are two astronomical time recording methods. UTC time can be calculated using the existing method described above. There is an existing conversion formula between UTC and MJD time. MJD time can be calculated using this formula and applied when generating pulsar processing files.
[0113] Step A3: Check and supplement the data frame by detecting changes in the second number and the frame number per second in the data frame header.
[0114] Among them, four functions are used: data frame header check function check_vdif_header, data frame header check function only_check_vdif_header, white noise generation function
[0115] generate_gaussian_noise and data frame supplement function generate_and_write_vdif_frame, where check_vdif_header implements the function of checking whether the data frame is lost. If there is a loss, the white noise generation function generate_gaussian_noise and the data frame supplement function generate_and_write_vdif_frame are called to use white noise to fill the lost data frame. White noise can only be used to fill the lost data array, not to fill the data frame header; the data frame header check function only_check_vdif_header is used to check the data frame again after completing the above operations.
[0116] For example, for the data file information of the Tianma telescope observing pulsar B0329+54, seconds = 68133904, legacymode = 0, frame = 0, epoch = 44, framelength = 1028, nhan = 4, version = 0, stationid = 0, treadid = 0, and nbits = 7. Therefore, the observed UTC time and the abbreviated Julian day time are accurately calculated. The frame length can also be used to calculate the number of bytes of the frame header and data array of a single data frame in the VDIF file: 8224 bytes (the framelength read data is 1028, so the size of the entire data frame is 1028 × 8 = 8224 bytes. Excluding the 32-byte frame header, the data array length is 8224-32 = 8192 bytes, where the frame header size is 32 bytes and the data array length is 8192 bytes. Using nhan to output the data, the total number of data channels can be calculated as 2 4 = 16 channels, 16 channels include eight channels of left-hand signals and eight channels of right-hand channels. A single channel has 64,000,000 samples per second, but the VDIF data is arranged as follows: right-hand sample 1 of the first channel, left-hand sample 1 of the first channel, right-hand sample 1 of the second channel, left-hand sample 1 of the second channel, etc. After sample 1 is completed, sample 2 is arranged in this order. The specific arrangement is as follows: Figure 3 and Figure 4 As shown in the figure, this is one of the arrangements. The size of a single data in a channel is 8 bits, that is, a char type integer data.
[0117] The frequency distribution of the eight channels of left-hand or right-hand rotation is as follows Figure 3 As shown, the starting frequency of the first channel is 1968MHz, the bandwidth of a single channel is 64MHz, and the final cutoff frequency of the eight channels is 2480MHz.
[0118] In the VDIF data array, there are two ways to arrange the original data. The system includes processing of two original data types, where the schematic diagram of processing one of the two original data types is as follows: Figure 4 Specifically, before processing, the data is first arranged in the order of left-hand and right-hand signals, and then arranged in sequence by eight channels. Another arrangement is the opposite, first arranging in sequence by eight channels and then arranging in the order of left-hand and right-hand signals.
[0119] To perform coherent de-dispersion of pulsars, the data must first be pre-processed. This pre-processing of the VDIF data involves removing the VDIF frame header and rearranging the data to prepare it for subsequent processing. This data rearrangement involves storing the left-handed and right-handed VDIF data in two separate data buffers, and aligning the VDIF data according to a channel-first, time-second logic.
[0120] Therefore, the two raw data are finally processed to form the following Figure 4 The arrangement format is shown on the right. After this rearrangement, the header information of each data frame is first removed, leaving only the valid data. The left and right rotational data are then stored in two separate data buffers, as shown on the right side of the figure below. In this new arrangement, the first time point of all channels is displayed first, followed by the second time point of all channels. This data structure is more suitable for subsequent processing operations, especially when performing coherent de-dispersion, helping to improve processing efficiency and data access continuity.
[0121] The coherent de-dispersion method of pulsar baseband data of the present invention is implemented by a coherent de-dispersion system based on GPU and HASHPIPE architecture. The coherent de-dispersion system based on GPU and HASHPIPE architecture is mainly divided into two parts, namely the HASHPIPE architecture part and the coherent de-dispersion algorithm part. The HASHPIPE architecture part is mainly as follows: Figure 5 As shown, the coherent de-dispersion algorithm runs through the HASHPIPE architecture, and its specific structure is as follows Figure 6 shown.
[0122] The coherent de-dispersion method for pulsar baseband data of the present invention can realize coherent de-dispersion processing of multi-channel baseband pulsar data (multi-channel data here refers to the number of channels of baseband data and the number of channels in VDIF data), and the final total number of channels generated in the final pulsar filterbank file can be selected according to user needs (the final total number of channels generated here refers to the number of channels obtained by splitting each single channel in the program, which will be explained in detail below). The generated filterbank file can be directly used to complete the final folding drawing in the dspsr software, and due to the channel splitting function of the program, the user can see the details of the pulsar pattern in each frequency channel, and can also compare the gap between pulsar signals under different frequency channel frequency bandwidths in more detail.
[0123] The implementation process of the coherent de-dispersion method for pulsar baseband data of the present invention is described in detail below.
[0124] The coherent de-dispersion method of pulsar baseband data of the present invention comprises:
[0125] Step S1: Based on the HASHPIPE architecture, a data transmission method is established among the data reading thread net_thread.c, the input buffer inputbuf, the multi-threaded parallel coherent de-dispersion thread, the output buffer outputbuf and the file writing thread output_thread.c;
[0126] HASHPIPE is a high-availability shared pipeline engine that accelerates multi-threaded parallel processing and rapidly transfers data between threads through a shared buffer. Based on the HASHPIPE architecture, data is transferred sequentially along the data reading thread net_thread.c, the input buffer inputbuf, the multi-threaded parallel coherent de-dispersion thread, the output buffer outputbuf, and the file writing thread output_thread.c. This HASHPIPE architecture enables cyclical reading and writing between the two shared buffers, ensuring efficient data processing.
[0127] The step S1 specifically includes:
[0128] Step S11: using the buffer setting programs databuf.h and databuf.c, the configuration of the input buffer inputbuf and the output buffer outputbuf is completed, and the data structure related to the frame number of the VDIF data is defined.
[0129] Among them, the size of the input buffer inputbuf is set to three times the amount of data processed each time, nvalid×nfft×nsub, where nsub refers to the number of subbands of VDIF data. In this embodiment, nsub takes 8 channels as an example, nfft refers to the number of FFT data set to be performed simultaneously in GPU parallel processing, and nvalid is the valid data length of a single subband in an fft, that is, the amount of data processed each time is nvalid data amount of the number of subband channels that can be performed nfft times.
[0130] The values of nvalid, nfft, and nsub are set manually.
[0131] The number of input buffers inputbuf is multiple to improve data processing speed. In this embodiment, the number of input buffers inputbuf is 3, and the purpose of buffers of the same size is to achieve parallel processing.
[0132] The input buffer inputbuf works as follows:
[0133] 1) Data reading: The data reading thread net_thread.c is responsible for reading data; after the received data frame header is removed, the remaining data will be stored in the first input buffer.
[0134] 2) Filling the input buffer inputbuf: When the first input buffer inputbuf is full, the data reading thread net_thread.c will store the newly received data in the second input buffer inputbuf; this process is cyclic, and the data is stored in turn among the three input buffers inputbuf, forming a circular queue.
[0135] 3) Input buffer: If the read thread processes data quickly but the GPU is slower, the input buffer can effectively prepare the data. This way, when the GPU is reading data from the first input buffer, if its processing speed is slower, the second input buffer can provide the subsequent data, thus reducing the GPU's waiting time and improving overall processing efficiency.
[0136] The output buffer (outputbuf) is set to three times the amount of data written each time (Msamp x mchan). Here, mchan is the total number of channels generated, which is the product of the number of baseband data subband channels and the number of user-split channels. Msamp is the length of the processed data for the user-set number of channels. The values of Msamp and mchan are also manually set.
[0137] There are multiple output buffers outputbuf, which also play a role in data preparation for subsequent reading of the file writing thread output_thread.c. In this embodiment, there are three output buffers outputbuf.
[0138] Therefore, the final data block size after data processing will change from nvalid × nfft × nsub to Msamp × mchan data size. The same as the input data format, three output buffers of the same size are set to achieve parallel processing.
[0139] Important data parameters in the coherent de-dispersion system are also defined in the buffer setting programs databuf.h and databuf.c, which facilitates subsequent program calls.
[0140] Step S2: VDIF data preprocessing, including: in the data reading thread net_thread.c, checking the data frame of the VDIF data, reading the time information and deleting the frame header, obtaining the data in the required arrangement format and storing it in the input buffer inputbuf.
[0141] Step S2 specifically includes:
[0142] Step S21: Using the data reading thread net_thread.c to read the VDIF baseband data and read the frame header information;
[0143] Step S22: Parse the time in the frame header of the data frame to calculate the corresponding UTC time and MJD time;
[0144] Since pulsar data requires extremely high time accuracy, step S22 also includes checking the second frame number in the frame header by detecting changes in the second number and the frame number per second in the data frame header to ensure frame continuity; and when a data frame is found to be lost, it is filled by generating an equal amount of white noise frames to avoid data loss problems in subsequent pulsar folding and drawing.
[0145] Step S23: Remove the frame header and store the processed data in the input buffer that has been set up. The system will give a prompt each time the processed data is successfully stored. After this step is completed, the data is saved in the input buffer, waiting for further processing.
[0146] Step S3: using the data in the input buffer inputbuf to execute the multi-threaded parallel coherent de-dispersion thread and write the result to the output buffer outputbuf. In this process, data splitting, FFT, frequency conversion, complex multiplication, frequency conversion, IFFT, de-filling detection power and Stokes parameter calculation are sequentially implemented through function calls;
[0147] Step S3 is the core step. The multi-threaded parallel coherent de-dispersion thread is composed of the coherent de-dispersion kernel function gpu_kernel.cu, the function declaration module gpu_thread.h and the main program module gpu_thread.cu.
[0148] The coherent de-dispersion kernel function gpu_kernel.cu defines all the functions required for coherent de-dispersion, that is, Figure 6 The definitions of all functions in the middle column, except for FFT and IFFT (including the format conversion function separatevdifData, the data unpacking function unpack_and_padd, the channel spectrum positive and negative frequency conversion function swap_spectrum_halve, the complex multiplication function PointwiseComplexMultiply, the unpadding detection power function transpose_unpadd_and_detect, and the Stokes parameter calculation and quantization function), are all in the coherent dispersion kernel function gpu_kernel.cu. The function declaration module gpu_thread.h declares the functions defined in the coherent dispersion kernel function gpu_kernel.cu.
[0149] The principle of coherent dedispersion has been explained in detail above. The time delay caused by the pulsar radiation signal passing through the interstellar medium can be expressed as the product of the interstellar medium effect in the frequency domain and the pulsar signal in the frequency domain, as shown below:
[0150] S initial (v+v0)=H -1 (v+v0)S(v+v0)
[0151] The inverse of the transfer function H -1 (v+v0) is called the chirp function.
[0152] Therefore, the overall process of step S3 is as follows Figure 6 The main program module gpu_thread.cu is used to complete the initialization of the multi-threaded parallel coherent de-dispersion thread, the process of the multi-threaded parallel coherent de-dispersion thread (i.e., the main thread), and the calling of all functions.
[0153] Step S3 specifically includes:
[0154] Step S31: executing the main program module gpu_thread.cu to implement the initialization of multi-threaded parallel coherent de-dispersion threads;
[0155] The initialization process includes variable initialization, reading the parameter file para, CPU and GPU memory allocation (memory must be allocated in advance during initialization for each portion of data after function processing), FFT plan generation, and chirp function calculation.
[0156] Since the function calls the cufft library in CUDA, the FFT in this cufft library needs to generate an FFT plan in advance. The FFT plan records the number of data points processed by a single FFT and the form of FFT processing. The setting here is CUFFT_C2C. Subsequent FFT only needs to call this FFT plan and can run directly without setting too many parameters.
[0157] The parameter file para includes N bin (single processing data size), noverlap (filling length), outputfilename (output fillerbank file name), source_name (processing pulsar source name, such as B0329+54), dm_start (dispersion start value), dm_step (chromatic step length), ndm (number of dispersions), nchan (number of user split channels), etc. The para file is provided each time according to the user's processing requirements.
[0158] It can be seen from the formula that the chirp function is related to the frequency bandwidth of the pulsar and the number of single data processing points. The frequency bandwidth of the pulsar baseband data is fixed, and the number of single data processing points is also set as a fixed value in the Para parameter file. Therefore, the chirp function calculation can be completed according to the known parameters in the initialization stage.
[0159] The calculation of the chirp function is achieved through the calculation formula of the transfer function, which is:
[0160]
[0161] Where DM represents the dispersion, k DM represents the dispersion constant, v0 is the center frequency of the subband or channel, the bandwidth is Δv (all in MHz), and v is the frequency offset within the channel or subband, satisfying -Δv / 2<ν<Δν / 2.
[0162] The initialization process of the main program module gpu_thread.cu can be completed after the cache area setting programs databuf.h and databuf.c are used to initialize the input buffer inputbuf and the output buffer outputbuf and the data reading thread net_thread.c for thread initialization, without waiting for the data reading thread net_thread.c to complete the preprocessing of the first data block. This is because the initialization steps of the HASHPIPE architecture program can be implemented in parallel according to the current running status.
[0163] That is to say, both the data reading thread net_thread.c and the main program module gpu_thread.cu have the functions of thread initialization and the main program of the thread. The operation logic in the HASHPIPE architecture is to first complete the cache area setting programs databuf.h and databuf.c and then complete the initialization of the data reading thread net_thread.c and the main program module gpu_thread.cu.
[0164] Step S32: After the main program module gpu_thread.cu completes the initialization of the multi-threaded parallel coherent de-dispersion thread, the thread status of the hashpipe framework is determined to ensure the correct operation of the thread; next, the data status of the input buffer inputbuf and the output buffer outputbuf is checked to ensure the consistency of data processing; if it is determined to be available, step S33 is executed.
[0165] Step S33: executing the main program module gpu_thread.cu to implement the multi-threaded parallel coherent de-dispersion thread process.
[0166] Step S33 specifically includes:
[0167] Step S331: For each data processing, the main program module gpu_thread.cu reads a specified amount of data from the input buffer inputbuf and enters the format conversion function separatevdifData, calling the format conversion function separatevdifData to implement data conversion and data arrangement processing of the VDIF data;
[0168] Step S332: calling the data splitting function unpack_and_padd to split the data into nfft×nsub data blocks with valid length nvalid, where nfft is the number of FFT data to be performed simultaneously in GPU parallel processing, nsub refers to the number of subbands of VDIF data, and nvalid is the valid data length of a single subband in one FFT;
[0169] Due to the characteristics of the discrete convolution theorem, there will be mixed edges in the output time series. To avoid this phenomenon, we use the overlap preservation method.
[0170] The specific operation process of calling the data splitting function unpack_and_padd is as follows Figure 6 shown.
[0171] like Figure 7 As shown, in step S332, in the data splitting function unpack_and_padd, the data is split into nfft×nsub data blocks of length nvalid, where nvalid is the valid data length of a single subband in an FFT. An overlap of length noverlap is added to both ends of each data block. After adding the overlap, the total data length reaches nfft×nsub data points of length nbin, where nbin is the number of FFT calculation points and the size of nbin is nvalid + 2×nverlap.
[0172] It should be noted that the data content of the overlapping part also comes from the input data.
[0173] Step S333: Call the FFT plan to perform GPU parallel processing-based Fast Fourier Transform (FFT) on the nfft×nsub group of data blocks.
[0174] like Figure 6 As shown, the FFT program has been initialized in the initialization step of step S31, so the operation can be directly performed here. In this process, the number of calculation points of FFT is nbin, and Fourier transform is performed on nfft×nsub group data blocks each time.
[0175] After FFT, the data is converted from the time domain to the frequency domain, and the data aliasing is reduced by the overlapping preservation method.
[0176] Step S334: calling the complex multiplication function PointwiseComplexMultiply to perform a complex multiplication operation on the data block and the chirp function generated in step S31, thereby obtaining a data block after dispersion elimination.
[0177] In the complex multiplication function PointwiseComplexMultiply, complex multiplication is performed on the pulsar baseband data in the frequency domain (i.e., data block) and the chirp function generated during initialization, thereby completing the dispersion elimination within the channel.
[0178] In this embodiment, since the VDIF data format is real, a Fourier transform can be performed in a complex format using the imaginary zero-padding method. According to signal system principles, a real signal exhibits a symmetrical spectrum after Fourier transform. Therefore, before performing the complex multiplication operation in step S334, the negative frequency spectrum is deleted and the positive frequency spectrum is retained.
[0179] Since the fft plan calls the cufft library, the positive and negative frequencies of the spectrum in the Fourier transform result will be swapped. Therefore, in the present invention, a channel spectrum positive and negative frequency conversion function swap_spectrum_halve is also designed to restore the shape of the spectrum and avoid subsequent data processing problems caused by library functions. This channel spectrum positive and negative frequency conversion function swap_spectrum_halve needs to be performed before and after the complex multiplication function PointwiseComplexMultiply, that is, after FFT and before IFFT; that is, before step S334 and after step S334, it includes: calling the channel spectrum positive and negative frequency conversion function swap_spectrum_halve to flip the positive and negative frequencies of the data.
[0180] After step S334 (before IFFT), while calling the channel spectrum positive and negative frequency conversion function swap_spectrum_halve to restore the spectrum channels, the user's required channel division is performed to convert the number of frequency channels from the number of subbands nsub of the VDIF data to the final total number of channels mchan. Here, nchan is the number of user-split channels. That is, a single subband of the VDIF data is split into nchan channels, and the final total number of channels mchan and the number of user-split channels nchan satisfy mchan = nsub × nchan.
[0181] Step S335: calling the IFFT plan to perform an inverse Fourier transform (IFFT) based on GPU parallel processing on the data block that ultimately generates the total number of channels mchan, so as to convert the de-dispersed data back to the time domain.
[0182] Similar to the FFT scheme, the IFFT scheme also completes initialization during the initialization step. However, unlike the FFT, the length of each single data processed by the IFFT is 1 / (2 × nchan) of the data length during the Fourier transform. The 2 in the denominator is due to the deletion of half the data, while the 1 / nchan is due to the frequency channel splitting of a single subband during the frequency-domain to time-domain conversion. nchan is the number of channels the user has split. This satisfies the user's need for channel segmentation.
[0183] Step S336: Call the unpadd detection power function transpose_unpadd_and_detect to Figure 7 The overlapping parts of the time domain data are removed and the power sum on the two polarization channels is calculated;
[0184] For data that has been restored to the time domain, completed coherent dedispersion and data channel splitting, it is necessary to process the individual data in the de-pad detection power function transpose_unpadd_and_detect to remove the overlapping parts added at both ends of the previous valid data and calculate the power sum of the data on the two polarization channels in preparation for subsequent quantization operations.
[0185] Step S337: Calculate and quantize the Stokes parameters and save them to the output buffer outputbuf, and then terminate the current multi-threaded parallel coherent de-dispersion thread.
[0186] In this embodiment, the compute_block_sums, compute_channel_statistics, and redigitize kernel functions are called to perform Stokes parameter calculation and data quantization.
[0187] Because the power sums obtained in the previous steps are floating-point numbers, and the final data to be stored is of type unsigned char, the compute_block_sums, compute_channel_statistics, and redigitize kernel functions quantize the data and convert the floating-point numbers to char types. compute_block_sums sums the data, compute_channel_statistics calculates the mean and standard deviation, and redigitize quantizes the data based on the previously calculated mean and standard deviation, converting the floating-point numbers to unsigned char types. The processed data is then saved to the output buffer (outputbuf), effectively writing it to the output block. Finally, the input and output buffers (outputbuf) are updated, terminating the current multi-threaded parallel coherent de-dispersion process. Only after redigitization does the data become usable and can be written to the output buffer (outputbuf).
[0188] Step S4: calling the file writing thread output_thread.c to write the processed data in the output buffer outputbuf into a file, especially into a filterbank file that meets the processing requirements of the dspsr pulsar software.
[0189] In order to ensure the correctness of the file format and the integrity of the data, before calling the file writing thread output_thread.c to perform the write operation, the file header of the filterbank file is initialized and written according to the data information of the parameter file; after the file header is written, the data in the output buffer outputbuf is written to this filterbank file.
[0190] Among them, the data information of the parameter file includes key information such as the name of the observation source, the right ascension and declination of the observation source, the sampling rate, the number of channels, the starting frequency, and the ending frequency.
[0191] Step S5 (optional): Perform folding drawing in the dspsr pulsar software based on the generated filterbank file.
[0192] Among them, dspsr is a mature pulsar software. Using the corresponding commands, you can draw the filterbank file that has undergone coherent dedispersion, so that the data can be presented to the user in a visual way.
[0193] Experimental results:
[0194] After the coherent de-dispersion system for pulsar baseband data based on GPU and HASHPIPE architecture was completed, a large amount of data was used to test the technology. The main pulsar data tested were the following pulsars: B0329+54, B0355+54, J0248+6021
[0195] like Figure 8 Shown are the data for pulsar B0329+54, with a flux of 203 mjy, right ascension 03:32:59.4096, declination 54:34:43.329, DM value 26.7641, and a rotation period frequency of about 1.399;
[0196] like Figure 9 Shown are the data for pulsar B0355+54, with a flux of 23 mjy, right ascension 03:58:53.7238, declination 54:13:13.784, DM value 57.1420, and a rotation period frequency of approximately 6.394;
[0197] like Figure 10 Shown are data for pulsar J0248+6021, with a flux of 14 mjy, right ascension 02:48:18.617, declination 60:21:34.72, DM value 370, and a rotation period frequency of approximately 4.606.
[0198] Figures 8-10The resulting images for three pulsars show that the initial baseband pulsar data consisted of eight subbands of VDIF data. Using the designed coherent dispersion elimination system, each subband was split into 64 subchannels, for a total of 8 x 64 channels. The pulse patterns for each channel can be viewed in the DSPSR folded plot. Because the Tianma telescope currently observes pulsars in the S- and X-bands, the system's previous data acquisition terminal only covers S-band observations. Therefore, the resulting image only shows pulsar signals around 2200 to 2300 MHz. This system covers a wide frequency range, laying the foundation for subsequent ultra-wideband data processing.
[0199] In other embodiments, the GPU can be replaced with other types of processors capable of multithreading, so GPU parallel processing can be replaced with other types of multithreaded parallel processing.
[0200] FPGA-accelerated parallel processing: FPGAs (field-programmable gate arrays) can serve as an alternative to GPUs, particularly in applications requiring high throughput and low latency. Implementing coherent de-dispersion algorithms and data frame parsing algorithms on FPGAs provides real-time data processing capabilities and hardware-level optimization for improved processing speed and efficiency.
[0201] Multi-core CPU parallel processing: Similar processing acceleration can be achieved using CPUs with multiple processing cores. By optimizing multi-core processing and task scheduling, parallel processing capabilities similar to those of GPUs can be achieved. While this approach may be slightly inferior to GPU performance, it is still effective for certain application scenarios.
[0202] Dedicated hardware processor parallel processing: Design and use of custom hardware processors (application-specific integrated circuits) specifically for pulsar data processing. These processors can be optimized for specific processing requirements, thereby improving processing speed and efficiency.
[0203] The coherent de-dispersion method of pulsar baseband data of the present invention can achieve the following beneficial effects:
[0204] 1. Improved processing speed: This invention achieves efficient data processing by utilizing GPU acceleration and the HASHPIPE architecture. The GPU has powerful parallel computing capabilities, which can significantly increase the speed of pulsar data processing. The HASHPIPE architecture can add parallel processing channels to the original GPU, thereby increasing the speed.
[0205] GPU accelerated computing technology uses the parallel computing capabilities of GPUs to accelerate complex coherent de-dispersion algorithms during pulsar data processing, improving data processing efficiency. It has significant performance advantages, especially when processing large-scale observation data.
[0206] The multi-threaded data processing architecture based on HASHPIPE adopts the HASHPIPE framework and ensures the real-time performance of data processing through multi-threaded parallel processing. It is suitable for application scenarios of large-scale data processing, especially the real-time observation and processing needs of radio telescopes.
[0207] 2. Optimizing Computing Resource Utilization: By optimizing data processing algorithms and leveraging modern hardware acceleration technologies, this invention reduces computing resource requirements and improves system processing efficiency. This reduces hardware investment costs while simultaneously improving processing performance. Specifically, for VDIF data format decoding, an efficient decoding algorithm is employed to reduce CPU burden. Coherent dispersion elimination maximizes GPU resource utilization through parallel processing and efficient algorithms.
[0208] 3. Improve processing accuracy: This invention designs precise algorithms and data processing procedures to ensure that the accuracy of data can be controlled when processing pulsar signals, reduce errors, and improve data reliability and analysis accuracy.
[0209] Coherent dispersion-eliminating processing technology converts the pulsar signals received by the radio telescope into the frequency domain and applies a dispersion-eliminating algorithm in the frequency domain to eliminate the pulse profile broadening problem caused by the interstellar medium, thereby improving the time resolution and accuracy of the signal.
[0210] 4. Increased Flexibility and Adaptability: The technical solution of the present invention is highly flexible and can be adjusted and optimized based on different data characteristics and requirements, adapting to various processing scenarios and improving the system's adaptability and scalability. Specifically, the present invention employs dynamic parameter adjustment: using parameter files to input data characteristics and adjust algorithm parameters to process different types of data; and modular design: the system adopts a modular architecture, and individual processing modules can be replaced or adjusted as needed. These steps enhance the system's adaptability and scalability, enabling it to adapt to a variety of processing scenarios.
[0211] 5. The decoding and processing of baseband VDIF format data is more accurate and efficient: For the baseband data collected by radio telescopes, the VDIF (VLBI Data Interchange Format) format is currently mainly used. VDIF is a general term for formats. The data in the VDIF format are generally the same. Different VDIF files have different frame header information, and the data arrangement formats in different VDIF files are also different. The existing VDIF decoding method is only suitable for current VDIF data. The coherent de-dispersion method for pulsar baseband data of the present invention uses baseband data from other digital acquisition terminals of Tianma. The format of the baseband data is VDIF format. A special decoding method is designed, and GPU is used for accelerated decoding, which is faster and more accurate.
[0212] VDIF data parsing and lost frame repair technology determines the integrity of the data by detecting the seconds and frame numbers of VDIF format data frames. When a data frame is lost, it uses Gaussian noise to fill the lost part to ensure data continuity and integrity.
[0213] The above-mentioned technical points jointly ensure the real-time data processing capability and high efficiency of the present invention.
[0214] In summary, compared to existing domestic pulsar baseband data processing technologies, this invention combines GPU acceleration technology with the HASHPIPE architecture to achieve efficient real-time processing of pulsar signals. In particular, the parallel processing capabilities of the GPU significantly increase computational speed when applying the coherent de-dispersion algorithm. This improvement significantly reduces processing latency, enabling the system to complete data analysis more quickly. Furthermore, this invention introduces VDIF data frame parsing and lost frame repair technology, ensuring data integrity and continuity during transmission, resolving the existing difficulty in recovering lost data frames.
[0215] Existing technologies often encounter problems such as slow processing speed and difficulty repairing lost frames when processing large amounts of pulsar observation data. This invention, by combining GPU acceleration with multithreaded processing, not only significantly improves processing efficiency but also ensures data integrity through intelligent lost frame detection and repair mechanisms. This improvement enhances the real-time performance and accuracy of pulsar signal processing, better meeting the processing requirements of radio telescope observation data.
[0216] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of the present invention. Various modifications are possible. Any simple, equivalent changes and modifications made in accordance with the claims and description of the present invention are within the scope of protection of the patent claims. Anything not fully described in this invention is conventional technology.
Claims
1. A coherent de-dispersion method for pulsar baseband data, characterized in that: include: Step S1: Based on the HASHPIPE architecture, a data transmission method is established among the data reading thread net_thread.c, the input buffer inputbuf, the multi-threaded parallel coherent de-dispersion thread, the output buffer outputbuf and the file writing thread output_thread.c; Step S2: In the data reading thread net_thread.c, the data frame of the VDIF data is checked, the time information is read, and the frame header is deleted to obtain the data in the required format and store it in the input buffer inputbuf; Step S3: using the data in the input buffer inputbuf to execute the multi-threaded parallel coherent de-dispersion thread and write the result to the output buffer outputbuf. In this process, data splitting, FFT, frequency conversion, complex multiplication, frequency conversion, IFFT, de-filling detection power and Stokes parameter calculation are sequentially implemented through function calls; Step S4: calling the file writing thread output_thread.c to write the processed data in the output buffer outputbuf into the file; The step S3 specifically includes: Step S31: executing the main program module gpu_thread.cu to implement the initialization of multi-threaded parallel coherent de-dispersion threads; Step S32: After the main program module gpu_thread.cu completes the initialization of the multi-threaded parallel coherent de-dispersion thread, the thread status of the hashpipe framework is determined to ensure the correct operation of the thread; next, the data status of the input buffer inputbuf and the output buffer outputbuf is checked to see if they are available; if so, step S33 is executed; Step S33: executing the main program module gpu_thread.cu to implement the multi-threaded parallel coherent de-dispersion thread process; The step S33 specifically includes: Step S331: For each data processing, the main program module gpu_thread.cu reads a specified amount of data from the input buffer inputbuf and enters the format conversion function separatevdifData, calling the format conversion function separatevdifData to implement data conversion and data arrangement processing of the VDIF data; Step S332: calling the data splitting function unpack_and_padd to split the data into nfft×nsub data blocks with valid length nvalid, where nfft is the number of FFT data to be performed simultaneously in GPU parallel processing, nsub refers to the number of subbands of VDIF data, and nvalid is the valid data length of a single subband in one FFT; Step S333: calling the FFT plan to perform FFT based on multi-thread parallel processing on the nfft×nsub group of data blocks; Step S334: calling the complex multiplication function PointwiseComplexMultiply to perform a complex multiplication operation on the data block and the chirp function generated in step S31 to obtain a data block after dispersion elimination; Step S335: calling the IFFT plan to perform IFFT based on multi-thread parallel processing on the data block that ultimately generates the total number of channels mchan, so as to convert the de-dispersed data back to the time domain; Step S336: calling the unpadd detection power function transpose_unpadd_and_detect to remove the overlapping parts of the time domain data and calculate the power sum on the two polarization channels; Step S337: Calculate and quantize the Stokes parameters and save them to the output buffer outputbuf, and then terminate the current multi-threaded parallel coherent de-dispersion thread; In step S332, in the data splitting function unpack_and_padd, the data is split into nfft×nsub data blocks of length nvalid, where nvalid is the valid data length of a single subband in an FFT, and overlapping parts of length noverlap are added at both ends of each data block; and / or In step S334, before performing the complex multiplication operation, the method further includes: deleting the negative frequency spectrum part and retaining the positive frequency part; and / or Before and after step S334 , the process includes: calling the channel spectrum positive and negative frequency conversion function swap_spectrum_halve to flip the positive and negative frequencies of the data.
2. The coherent de-dispersion method for pulsar baseband data according to claim 1, characterized in that: The step S1 specifically includes: step S11: using the buffer setting programs databuf.h and databuf.c to complete the configuration of the input buffer inputbuf and the output buffer outputbuf, and define the data structure related to the frame number of the VDIF data.
3. The coherent de-dispersion method for pulsar baseband data according to claim 2, characterized in that: The size of the input buffer inputbuf is three times the amount of data processed each time, which is nvalid × nfft × nsub, where nsub refers to the number of subbands of VDIF data, nfft refers to the number of simultaneous FFT data set in GPU parallel processing, and nvalid is the valid data length of a single subband in an FFT; The size of the output buffer outputbuf is three times the amount of data written each time Msamp×mchan, where mchan is the total number of channels generated and Msamp is the length of the processed data for the user-set single channel number; and / or The number of input buffers inputbuf and output buffers outputbuf is multiple.
4. The coherent de-dispersion method for pulsar baseband data according to claim 1, characterized in that: Step S2 specifically includes: Step S21: using the data reading thread net_thread.c to read the VDIF baseband data and read the frame header information; Step S22: Parse the time in the frame header of the data frame to calculate the corresponding UTC time and MJD time; Step S23: remove the frame header and store the processed data into the input buffer inputbuf that has been set.
5. The coherent de-dispersion method for pulsar baseband data according to claim 4, characterized in that: The step S22 further includes: checking the frame number in the second in the frame header by detecting changes in the second number and the frame number per second in the data frame header, and filling the gap by generating an equal amount of white noise frames when data frames are found to be lost.
6. The coherent de-dispersion method for pulsar baseband data according to claim 1, characterized in that: The multi-threaded parallel processing includes one of GPU parallel processing, FPGA accelerated parallel processing, multi-core CPU parallel processing and dedicated hardware processor parallel processing.
7. The coherent de-dispersion method for pulsar baseband data according to claim 1, characterized in that: In step S4, the file is a filterbank file; and the coherent de-dispersion method further includes step S5: performing folding drawing in dspsr pulsar software according to the generated filterbank file.
Citation Information
Patent Citations
Pulsar dispersion elimination device and method
CN111292222A
Pulsar coherent dispersion elimination system and method
CN111368252A