Data processing method and apparatus, storage medium, device, and program product
By fitting pulsar signals with the Lorentz distribution function in radio telescope data, suspected pulsar samples were identified, solving the problem of low efficiency caused by large data volume and achieving efficient and accurate pulsar identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECH SHANGHAI
- Filing Date
- 2022-03-11
- Publication Date
- 2026-07-17
AI Technical Summary
Radio telescope observations generate a massive amount of pulsar data, and manual processing to identify suspected pulsar signals is inefficient.
By acquiring the dispersion value and signal-to-noise ratio data points in the sample to be tested, the peak target data points are determined. The data points are fitted based on the Lorentz distribution function to generate the first fitting function. The outer envelope data points are determined within the fitting width to generate the second fitting function. The similarity and parameters of the fitting functions are judged to determine the suspected pulsar sample.
It improved the search efficiency of pulsar signals, reduced manual processing, enhanced identification accuracy, achieved a recall rate of 95.14%, reduced the number of false positives by an order of magnitude, and reduced noise data.
Smart Images

Figure CN116776032B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to a data processing method, apparatus, storage medium, device, and program product. Background Technology
[0002] Currently, radio telescopes are generally used to observe pulsars. However, radio telescopes generate a huge amount of data, with peak values approaching 40 gigabytes per second. If manual processing is used to identify signals that are suspected to be pulsars, it would require a lot of manpower and would be inefficient. Summary of the Invention
[0003] This application provides a data processing method, apparatus, storage medium, device, and program product that can improve the search efficiency of pulsar signals.
[0004] On one hand, a data processing method is provided, the method comprising: acquiring a test sample, the test sample comprising multiple data points, the data points comprising dispersion values and signal-to-noise ratios; determining a target data point as a peak value among the multiple data points; fitting multiple data points in the test sample within a preset width range of the target data point based on a preset Lorentz distribution function to determine a first fitting function, the first fitting function comprising a fitting width; determining an outer envelope data point based on the multiple data points within the fitting width range of the target data point, wherein the outer envelope data point has the highest signal-to-noise ratio when the dispersion values are the same; fitting the outer envelope data point based on the preset Lorentz distribution function to determine a second fitting function, the second fitting function comprising fitting parameters; and determining the test sample as a suspected pulsar sample when the similarity between the first fitting function and the second fitting function is greater than a preset similarity and the fitting parameters satisfy a preset condition.
[0005] On the other hand, a data processing apparatus is provided, comprising an acquisition module, a first determination module, a first fitting module, a second determination module, a second fitting module, and a third determination module. The acquisition module acquires a test sample, which includes multiple data points, each including a dispersion value and a signal-to-noise ratio (SNR). The first determination module determines a target data point among the multiple data points as a peak value. The first fitting module fits multiple data points within a preset width range of the target data point in the test sample based on a preset Lorentz distribution function to determine a first fitting function, which includes a fitting width. The second determination module determines an outer envelope data point based on the multiple data points within the fitting width range of the target data point, wherein the outer envelope data point has the highest SNR when the dispersion values are the same. The second fitting module fits the outer envelope data point based on a preset Lorentz distribution function to determine a second fitting function, which includes fitting parameters. The third determination module determines the test sample as a suspected pulsar sample when the similarity between the first fitting function and the second fitting function is greater than a preset similarity and the fitting parameters meet preset conditions.
[0006] In another aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program adapted for loading by a processor to perform the steps in the data processing method as described in any of the above embodiments.
[0007] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing a computer program, the processor executing the steps of the data processing method as described in any of the above embodiments by calling the computer program stored in the memory.
[0008] On the other hand, a computer program product is provided, including computer instructions that, when executed by a processor, implement the steps in the data processing method as described in any of the above embodiments.
[0009] The data processing method, data processing apparatus, computer-readable storage medium, computer equipment, and computer program product of this application acquire data points containing dispersion values and signal-to-noise ratios in a sample to be tested, determine the target data points of the peak values, and fit a first fitting function with data points within a preset width range near the target data points. This determines the fitting width, which is closer to the width of the data point distribution corresponding to the pulsar signal than the preset width range. Then, the fitting is performed again based on the outer envelope data points within the fitting width to obtain a second fitting function that is closer to the data point distribution of the pulsar signal. Finally, it is determined whether the similarity between the second and first fitting functions reaches a preset similarity and whether the fitting parameters of the second fitting function meet preset conditions, thereby determining whether the second fitting function conforms to the Lorentz distribution and thus whether the sample to be tested is a suspected pulsar sample. This saves manpower and improves the search efficiency for pulsar samples. Furthermore, the high similarity between the second and first fitting functions—that is, the difference between generating the second fitting function based on the outer envelope data points and generating the first fitting function based on the data points within the fitting width—prevents errors caused by excessively large outer envelope data points and improves the accuracy of the second fitting function. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A schematic diagram of the Lorentz distribution function provided in the embodiments of this application.
[0012] Figure 2 This is a schematic diagram of the structure of the data processing system provided in the embodiments of this application.
[0013] Figure 3 This is a flowchart illustrating the data processing method provided in an embodiment of this application.
[0014] Figure 4 This is a schematic diagram illustrating the principle of the data processing method provided in the embodiments of this application.
[0015] Figure 5 This is a flowchart illustrating the data processing method provided in an embodiment of this application.
[0016] Figure 6 This is a flowchart illustrating the data processing method provided in an embodiment of this application.
[0017] Figure 7This is a schematic diagram illustrating the principle of the data processing method provided in the embodiments of this application.
[0018] Figure 8 This is a schematic diagram illustrating the principle of the data processing method provided in the embodiments of this application.
[0019] Figure 9 This is a flowchart illustrating the data processing method provided in an embodiment of this application.
[0020] Figure 10 This is a flowchart illustrating the data processing method provided in an embodiment of this application.
[0021] Figure 11 This is a flowchart illustrating the data processing method provided in an embodiment of this application.
[0022] Figure 12 This is a flowchart illustrating the data processing method provided in an embodiment of this application.
[0023] Figure 13 This is a flowchart illustrating the data processing method provided in an embodiment of this application.
[0024] Figure 14 This is a schematic diagram illustrating the principle of the data processing method provided in the embodiments of this application.
[0025] Figure 15 This is a schematic diagram of the structure of the data processing apparatus provided in the embodiments of this application.
[0026] Figure 16 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0028] This application provides a data processing method, apparatus, computer device, and storage medium. Specifically, the data processing method of this application can be executed by a computer device, which can be a terminal or a cloud server, etc. The terminal can be a smartphone, tablet, laptop, desktop computer, smart TV, smart speaker, wearable smart device, smart vehicle terminal, etc. The terminal can also include a client, which can be a cloud gaming client, a client applet, a video client, a browser client, or an instant messaging client, etc. The cloud server can be an independent physical cloud server, a cloud server cluster or distributed system composed of multiple physical cloud servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0029] The embodiments of this application can be applied to various scenarios such as artificial intelligence, computer vision, and image recognition.
[0030] First, some of the nouns or terms that appear in the description of the embodiments of this application are explained as follows:
[0031] Pulsars, also known as neutron stars, are a type of neutron star that periodically emits pulse signals. The study of pulsars can provide useful information for research in astronomy and physics, such as general relativity and the interstellar medium. Pulsar search methods can be divided into two categories: periodic searches and single-pulse searches.
[0032] The purpose of monopulse search is to find pulse signals with long periods or isolated pulse signals without obvious periodicity, in order to determine whether they are pulsars.
[0033] The Lorentz distribution, also known as the Cauchy distribution, is a continuous probability distribution named after Augustine-Louis Cauchy and Hendrik Lorentz. Its probability density function is: Where x0 is the location parameter defining the peak position of the distribution, and γ is the scale parameter representing half the width at half the maximum value. The effect of different parameters on the Lorentz distribution is as follows: Figure 1 As shown.
[0034] A cloud server is a server that runs games in the cloud and has functions such as data processing.
[0035] A terminal refers to a type of device that has rich human-computer interaction methods, internet access capabilities, typically runs various operating systems, and possesses strong processing power. Terminals include smartphones, living room TVs, tablets, in-vehicle terminals, handheld game consoles, etc.
[0036] Please refer to Figure 2 , Figure 2 This is a schematic diagram of the structure of a data processing system provided in an embodiment of this application. The data processing system includes a terminal 10 and a cloud server 20, etc.; the terminal 10 and the cloud server 20 are connected via a network, such as a wired or wireless network.
[0037] The terminal 10 can be used to display a graphical user interface. The terminal 10 is used to interact with the user through the graphical user interface, such as downloading and installing a corresponding client, running a corresponding mini-program, or displaying a corresponding graphical user interface by logging into a website. In this embodiment, the terminal 10 can receive and display suspected pulsar samples filtered by the cloud server 20, allowing the user to confirm the suspected pulsar samples and find the accurate pulsar sample.
[0038] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the priority of the embodiments.
[0039] Various embodiments of this application provide a data processing method, which can be executed by a data processing system, such as by a terminal 10 or a cloud server 20, or by both a terminal 10 and a cloud server 20. This application embodiment uses the execution of the data processing method by the cloud server 20 as an example for illustration.
[0040] Please see Figures 3 to 13 , Figure 3 , Figure 5 , Figure 6 , Figures 9 to 13 These are all schematic flowcharts of the data processing methods provided in the embodiments of this application. Figure 4 , Figure 7 and Figure 8 These are schematic diagrams illustrating the principle of the data processing method provided in the embodiments of this application. The data processing method includes:
[0041] Step 011: Obtain the sample to be tested. The sample to be tested includes multiple data points, including dispersion value and signal-to-noise ratio.
[0042] After obtaining pulsar data through radio telescope observations, the pulsar data can be preprocessed to obtain the sample to be tested. For example, by preprocessing the pulsar data using the open-source pulsar software PRESTO, the dispersion measure (DM) and signal-to-noise ratio (SNR) can be obtained, where the dispersion measure is the amount of dispersion of the pulsar, and the SNR is the ratio of the pulsar signal to the noise.
[0043] The sample to be tested includes multiple data points, each of which includes a dispersion value and a signal-to-noise ratio. For example... Figure 4 As shown, a coordinate system is formed with the dispersion value DM as the horizontal axis and the signal-to-noise ratio SNR as the vertical axis, and all data points are located within the coordinate system.
[0044] Step 012: Determine the target data point as the peak value among multiple data points.
[0045] After establishing a coordinate system, to quickly search for data point regions conforming to the Lorentz distribution, we can first search for one or more peaks in the coordinate system. A peak can be understood as the data point with the highest signal-to-noise ratio (SNR) within a specific width range, where the SNR of data points on either side of the peak is lower than the peak's SNR. Here, the specific width range refers to the range defined by the DM value within the coordinate system, such as... Figure 4 The DM value ranges from 50 to 150.
[0046] Optionally, the graph formed by the data points in the coordinate system (e.g., connecting the outermost data points to form an image) can be used to detect whether the graph conforms to a peak graph. If so, the data point with the highest signal-to-noise ratio in the image can be determined as the data point corresponding to the peak.
[0047] Once the data points corresponding to the peak values are determined, all data points corresponding to the peak values can be used as target data points.
[0048] Please see Figure 5 Optionally, step 012 may include:
[0049] Step 0121: Compare the dispersion values of any two data points to determine one or more peak data points with the highest signal-to-noise ratio within a predetermined width range.
[0050] The magnitude relationship between any two data points can be determined by comparing their dispersion values. By selecting data points within a predetermined width range and determining the peak value within that range based on the magnitude relationship between any two data points, the peak value within that range can be quickly determined.
[0051] The predetermined width range is an empirical value, which can be determined based on the broadband range of the DM values of the data points of the detected pulsar samples.
[0052] It is understandable that for data points that conform to the Lorenz distribution, the peak value is generally located at the midpoint of a predetermined width range. Therefore, based on the position of the peak value within the selected predetermined width range, the width range can be redefined. Then, based on the data points within the redefined width range, it can be determined whether the data point corresponding to the peak value is still the peak value. If so, the data point corresponding to the peak value can be determined as the peak data point.
[0053] In this way, one or more peak data points can be identified in the sample to be tested (e.g. Figure 4 S is shown.
[0054] Step 0122: Determine the target data point based on the peak data point.
[0055] Once the peak data point is determined, the target data point can be determined based on the peak data point.
[0056] Optionally, all peak data points can be used as target data points.
[0057] Please see Figure 6 Optionally, step 0122 may include:
[0058] Step 01221: Determine the peak data points with a signal-to-noise ratio greater than a preset value as the target data points.
[0059] It is understandable that signals with excessively low signal-to-noise ratios (SNR) may contain significant errors, and pulsar signals generally have high SNRs. Therefore, by summarizing the SNRs of a large number of real pulsar signals obtained in practice, a preset value (i.e., an empirical value) can be obtained. By comparing the preset value with peak data points, peak data points that are less than or equal to (less than) the preset value can be quickly eliminated, while only peak data points that are greater than or equal to the preset value are retained as target data points (e.g., ...). Figure 4 As shown in the diagram (M), this reduces the amount of data processing.
[0060] Step 013: Based on the preset Lorentz distribution function, fit multiple data points in the test sample within the preset width range of the target data points to determine the first fitting function, which includes the fitting width.
[0061] Please see Figure 7 After determining the target data point, the preset width range can be determined based on empirical values of the range of DM values corresponding to the predetermined pulsar signal. For example, the preset width range can be determined based on the DM value of the target data point (e.g., DM0). The preset width range can be [DM0-d0, DM0+d0]. d0 can be an empirical value, such as d = 50, 70, 80, etc.
[0062] Then, select multiple data points within a preset width range and fit them according to the preset Lorentz distribution function to obtain the first fitting function.
[0063] In this embodiment, the preset Lorentz distribution function can be: The Lorentz distribution function includes a first fitting parameter X0, a second fitting parameter γ, and a third fitting parameter A, where A is a parameter representing the amplitude. It can be understood that different values of A will not change the shape of the Lorentz distribution function.
[0064] The DM value of each data point within a preset width range is taken as x in the Lorentz distribution function, and the signal-to-noise ratio of each data point is taken as x in the Lorentz distribution function. The first fitting function can be obtained by fitting multiple data points, where x0, γ and A of the first fitting function are all determined.
[0065] Please see Figure 8 The fitting width of the first fitting function can be determined according to the first fitting curve S1 corresponding to the first fitting function. For example, taking the DM value of the target data point as the center, the ratio of the first area of the region enclosed by the first fitting curve S1 and the horizontal axis where the DM value is located to the total area of all regions enclosed by the first fitting curve S1 and the horizontal axis where the DM value is located within the fitting width range of the first fitting curve S1 (such as the fitting width range of [DM0-d1, DM0+d1]) is greater than or equal to a preset ratio (such as the preset ratio can be 0.65, 0.95, etc.).
[0066] Step 014: Determine the outer envelope data points based on multiple data points within the fitting width range of the target data points. Under the condition that the dispersion values are the same, the signal-to-noise ratio of the outer envelope data points is the largest.
[0067] After determining the fitting width, all data points within the fitting width range can be obtained, and the outer envelope data points of all data points within the fitting width range can be determined. The outer envelope data points are the data points located on the outermost edge within the fitting width range. That is to say, under the same dispersion value, the signal-to-noise ratio of the outer envelope data points is the largest.
[0068] Step 015: Based on the preset Lorentz distribution function, fit the outer envelope data points to determine the second fitting function, which includes fitting parameters.
[0069] Please combine Figure 7 and Figure 8 After determining all outer envelope data points within the fitting width range, all outer envelope data points are fitted again based on the preset Lorentz distribution function to obtain the second fitting function S2. The second fitting function S2 includes a first fitting parameter and a second fitting parameter, wherein the first fitting parameter is X0 and the second fitting parameter is γ.
[0070] Specifically, the first fitting parameter X0 is the DM value (i.e., DM0) of the target data point, and the second fitting parameter is half the absolute value of the difference between the two dispersion values corresponding to half the signal-to-noise ratio of the target data point in the second fitting function.
[0071] Step 016: If the similarity between the first fitting function and the second fitting function is greater than the preset similarity and the fitting parameters meet the preset conditions, the sample to be tested is determined to be a suspected pulsar sample.
[0072] After obtaining the second fitting function, it can be determined whether the second fitting function conforms to the Lorentz distribution function corresponding to the pulsar signal.
[0073] It can be understood that for the same pulsar signal data points, the difference will not be too large. However, if there are anomalies in the outer envelope data points (such as the outer envelope data points being far away from other data points within the fitting width range (i.e., the signal-to-noise ratio difference is large)), it will lead to a large difference between the second fitting function and the first fitting function generated based on the outer envelope data points, thereby affecting the accuracy of the judgment of the sample under test.
[0074] Therefore, it is necessary to ensure that the similarity between the first fitting function and the second fitting function is greater than (greater than or equal to) the preset similarity, such as 0.6, 0.7, 0.8, etc.
[0075] Furthermore, the Lorentz distribution function of the pulsar signal generation also exhibits certain patterns. Besides ensuring the similarity is greater than a preset similarity, the fitting parameters of the second fitting function must also meet preset conditions. These preset conditions can be determined based on the parameters of the Lorentz distribution function corresponding to the acquired real pulsar signal. The preset conditions include determining whether the fitting parameters are greater than a preset parameter threshold, which is an empirical value.
[0076] Optionally, the fitting parameters can satisfy the preset conditions as follows: the first fitting parameter is greater than (greater than or equal to) the first preset value and the second fitting parameter is less than (less than or equal to) the second preset value. The first preset value can be 5, 6, etc., and the second preset value can be 100, 110, etc.
[0077] Please see Figure 9 Data processing methods also include:
[0078] Step 017: Determine the similarity based on the sum of the differences between the two signal-to-noise ratios corresponding to the same dispersion value in the first and second fitting functions. The larger the sum of the differences, the smaller the similarity.
[0079] When calculating the similarity between the first and second fitting functions, the difference between the two signal-to-noise ratios corresponding to the same dispersion value in the first and second fitting functions can be calculated. Then, the similarity is determined based on the sum of the differences. The larger the sum of the differences, the smaller the similarity.
[0080] Optionally, when calculating similarity, a DM value can be selected at predetermined intervals, and then a preset number of DM values can be selected, distributed on both sides of the DM value of the target data point. After calculating the difference between the two signal-to-noise ratios corresponding to each DM value, the similarity can be determined based on the sum of the differences.
[0081] Optionally, please refer to Figure 8 Alternatively, we can first determine the first fitting curve S1 of the first fitting function in the coordinate system and the second fitting curve S2 of the second fitting function in the coordinate system. Then, we can determine the similarity by the area of the region enclosed by the first fitting curve and the second fitting curve. The larger the area, the smaller the similarity.
[0082] Please see Figure 10 Data processing methods also include:
[0083] Step 018: Based on the preset Lorentz distribution function, fit multiple data points within the fitting width range of the target data point in the test sample to determine the third fitting function.
[0084] After fitting the first fitting function based on the data points within a preset width range, since the preset width range includes not only the data points corresponding to the pulsar signal but also some data points of non-pulsar signals, after determining the first fitting function, a third fitting function is obtained by fitting the data points within the fitting width range determined by the first fitting function. It can be understood that the fitting width is closer to the width of the data point distribution corresponding to the pulsar signal than the preset width range, making the third fitting function closer to the Lorentz distribution function corresponding to the pulsar signal.
[0085] Step 019: Update the fitting width according to the third fitting function.
[0086] Then, the fitting width is re-determined based on the third fitting function obtained from the refit. The method for determining the fitting width of the third fitting function is the same as that for the first fitting function, and will not be repeated here.
[0087] Subsequently, when determining the second fitting function, the second fitting function is determined based on the outer envelope data points within the newly determined fitting width range, thereby improving the accuracy of the second fitting function and thus improving the accuracy of the judgment on whether the sample to be tested is a suspected pulsar sample.
[0088] Please see Figure 11 Step 016 includes:
[0089] Step 0161: If the similarity between the first fitting function and the second fitting function corresponding to any target data point is greater than the preset similarity and the fitting parameters meet the preset conditions, the sample to be tested is determined to be a suspected pulsar sample.
[0090] It is understandable that there can be multiple target data points in the sample to be tested. Each target data point can determine the corresponding first fitting function and second fitting function, thereby determining the similarity and fitting parameters corresponding to each target data point.
[0091] When determining whether a sample to be tested is a suspected pulsar sample, it can be determined that the sample to be tested is a suspected pulsar sample if the similarity corresponding to any target data point is greater than the preset similarity and the fitting parameters meet the preset conditions.
[0092] In one example, the preset similarity is 0.6, the first preset value is 5, and the second preset value is 100.
[0093] In the first fitting curve corresponding to the first fitting function of the current target data point, multiple first curve points are selected at predetermined intervals (e.g., 40), such as (20, 6), (40, 7), (60, 10), (80, 17), and (100, 20); in the second fitting curve corresponding to the second fitting function of the current target data point, multiple first curve points are selected at the same predetermined intervals (i.e., 40), such as (20, 6), (40, 6.5), (60, 8), (80, 16), and (100, 20).
[0094] The similarity is calculated as 1 - |(6-6) + (7-5) + (10-8) + (17-16) + (20-20)| / (6 + 7 + 10 + 17 + 20) = 11 / 12 > 0.6, meaning the similarity is greater than the preset similarity. For example, if the first fitting parameter X0 of the second fitting function is 6 and the second fitting parameter γ is 80, then the first fitting parameter is greater than the first preset value, and the second fitting parameter is less than the second preset value. Thus, the sample to be tested can be determined to be a suspected pulsar sample.
[0095] Please see Figure 12 Step 016 includes:
[0096] Step 0162: If the similarity between the first fitting function and the second fitting function corresponding to the previous target data point is less than the preset similarity, or the fitting parameters do not meet the preset conditions, obtain the similarity and fitting parameters corresponding to the next target data point; wherein, the next target data point is the target data point adjacent to the target data point whose similarity and fitting parameters have been determined.
[0097] Step 0163: Determine whether the similarity is greater than the preset similarity.
[0098] Step 0164: Determine whether the fitted parameters meet the preset conditions.
[0099] Step 0165: If the similarity is greater than the preset similarity and the fitting parameters meet the preset conditions, the sample to be tested is determined to be a suspected pulsar sample.
[0100] Step 0166: If the similarity is less than or equal to the preset similarity or the fitting parameters do not meet the preset conditions, determine whether there is a next target data point;
[0101] If a next target data point exists, proceed to step 0162.
[0102] Step 0167: If there is no next target data point, determine that the sample to be tested is not a suspected pulsar sample.
[0103] For steps 0163 to 0165, please refer to the description of step 016, which will not be repeated here.
[0104] There can be multiple target data points. If the similarity of the current target data point (i.e. the previous target data point) is less than or equal to the preset similarity, or if the fitting parameters do not meet the preset conditions, the similarity and fitting parameters of the next target data point can be obtained.
[0105] The processing order of target data points can be sorted according to the size of the DM value of the target data points. For example, the smaller the DM, the smaller the sorting number, and the processing order can be carried out in ascending order of the sorting number.
[0106] The similarity and fitting parameters of the target data points are judged sequentially according to the order. After all target data points have been judged, if the similarity of all target data points is less than the preset similarity or the fitting parameters do not meet the preset conditions, that is, there is no next target data point, it can be determined that the sample to be tested is not a suspected pulsar sample.
[0107] Alternatively, during the process of judging target data points in sequence, if the similarity of a target data point is greater than the preset similarity and the fitting parameters meet the preset conditions, stop obtaining the similarity and fitting parameters of the next target data point, and directly determine the sample to be tested as a suspected pulsar sample, thereby reducing the amount of computation.
[0108] Optionally, the similarity and fitting parameters of the target data points can be judged in advance, and the judgment result of each target data point can be generated. After establishing a coordinate system, the data points of the test sample, the second fitting curve corresponding to the target data points, and the judgment result of each target data point can be marked in the coordinate system to generate the test sample image. Then the terminal can display the test sample image, so as to facilitate users to quickly judge whether the test sample is a pulsar sample.
[0109] Please refer to it again. Figure 7 , Figure 7 It contains three target data points, and each of the three target data points is labeled with the corresponding second fitting curve of the second fitting function. The judgment result of each second fitting curve is also labeled, which is either a suspected pulsar or a non-suspected pulsar. For example, the judgment result of the first second fitting curve is a suspected pulsar, while the judgment results of the second and third second fitting curves are both non-suspected pulsars.
[0110] Please see Figure 13 Step 0162 includes:
[0111] Step 1621: Determine the deletion width based on the fitting width of the second fitting function and the predetermined width.
[0112] After determining the similarity and fitting parameters of a target data point, in order to prevent related data points from affecting the determination of subsequent target data points, the deletion width can be determined based on the fitting width of the second fitting function and the predetermined width.
[0113] Optionally, the predetermined width can be an empirical value. It can be understood that the fitting width range only corresponds to a part of the second fitting curve of the second fitting function. Therefore, in order to prevent the data points related to the second fitting function from affecting the similarity of the next target data point and the judgment of the fitting parameters, the predetermined width is increased on the basis of the fitting width to ensure the accuracy of the deletion width.
[0114] Step 01622: Delete the data points within the width range of the sample to be tested;
[0115] Once the deletion width is determined, all data points within the deletion width range can be deleted, thereby ensuring the accuracy of the similarity of subsequent target data points and the judgment of fitting parameters.
[0116] Please combine Figure 7 and Figure 14In one example, the target data point is M(DM0, SN0), the fitting width is d1 (i.e., the fitting width range is [DM0-d1, DM0+d1]), and the predetermined width is d2. Then, the deletion width range is [DM0-d1-d2, DM0+d1+d2]. All data points within this deletion width range are deleted, thereby affecting the accuracy of the second fitting function for subsequent target data points.
[0117] Step 01623: Based on the data points in the test sample after deleting data points, determine the similarity and fitting parameters of the next target data point.
[0118] Then, based on the data points in the test sample after deleting data points, the similarity and fitting parameters of the next target data point are determined, thus ensuring the accuracy of the similarity and fitting parameters of the next target data point.
[0119] All of the above technical solutions can be combined in any way to form optional embodiments of this application, and will not be described in detail here.
[0120] This application embodiment acquires data points containing dispersion values and signal-to-noise ratios from the sample to be tested, and determines the target data points for peak values. A first fitting function is obtained by fitting data points within a preset width range near the target data points, thereby determining the fitting width. This fitting width, compared to the preset width range, is closer to the width of the data point distribution corresponding to the pulsar signal. Then, a second fitting function is obtained based on the outer envelope data points within the fitting width, resulting in a data point distribution that is also closer to the pulsar signal. Finally, it is determined whether the similarity between the second and first fitting functions reaches a preset similarity, and whether the fitting parameters of the second fitting function meet preset conditions, thereby determining whether the second fitting function conforms to the Lorentz distribution, and thus determining whether the sample to be tested is a suspected pulsar sample. Furthermore, the second and first fitting functions have a high similarity; that is, the difference between generating the second fitting function based on the outer envelope data points and generating the first fitting function based on the data points within the fitting width is not significant, preventing errors caused by excessively large outer envelope data points and improving the accuracy of the second fitting function.
[0121] Experiments showed that this application achieved a recall rate of 95.14% on a test set of known pulsars, and reduced the number of false positives on a negative test set of 20,000 samples to approximately 3,000. This application reduces the number of samples required for searching by an order of magnitude while maintaining the recall rate for positive samples, and filters out more noisy data, thereby reducing the workload of manual image interpretation and improving the efficiency of pulsar searches.
[0122] To facilitate better implementation of the data processing method of this application embodiment, this application embodiment also provides a data processing apparatus. Please refer to... Figure 15 , Figure 15 This is a schematic diagram of the structure of a data processing apparatus 1000 provided in an embodiment of this application. The data processing apparatus 1000 may include:
[0123] The acquisition module 1011 is used to acquire the sample to be tested, which includes multiple data points, including dispersion value and signal-to-noise ratio.
[0124] The first determining module 1012 is used to determine the target data point as the peak value among multiple data points;
[0125] The first determining module 1012 is specifically used for:
[0126] Compare the dispersion values of any two data points to determine one or more peak data points with the highest signal-to-noise ratio within a predetermined width range;
[0127] Determine the target data point based on the peak data point.
[0128] The first determining module 1012 is specifically used for:
[0129] The peak data points with a signal-to-noise ratio greater than a preset value are identified as the target data points.
[0130] The first fitting module 1013 is used to fit multiple data points in the test sample within a preset width range of the target data points based on a preset Lorentz distribution function, so as to determine the first fitting function, the first fitting function including the fitting width.
[0131] The second determining module 1014 is used to determine the outer envelope data point based on multiple data points within the fitting width range of the target data point, wherein the signal-to-noise ratio of the outer envelope data point is the largest when the dispersion values are the same.
[0132] The second fitting module 1015 is used to fit the outer envelope data points based on a preset Lorentz distribution function to determine the second fitting function, which includes fitting parameters.
[0133] The third determining module 1016 is used to determine the sample to be tested as a suspected pulsar sample when the similarity between the first fitting function and the second fitting function is greater than the preset similarity and the fitting parameters meet the preset conditions.
[0134] The third determining module 1016 is specifically used for:
[0135] If the similarity between the first and second fitting functions corresponding to any target data point is greater than the preset similarity and the fitting parameters meet the preset conditions, the sample to be tested is determined to be a suspected pulsar sample.
[0136] The third determining module 1016 is also specifically used for:
[0137] If the similarity between the first and second fitting functions corresponding to the previous target data point is less than or equal to the preset similarity, or if the fitting parameters meet the preset conditions, the similarity and fitting parameters corresponding to the next target data point are obtained; wherein, the next target data point is the target data point adjacent to the target data point whose similarity and fitting parameters have been determined.
[0138] The third determining module 1016 is also specifically used for:
[0139] The deletion width is determined based on the fitting width of the second fitting function and the predetermined width;
[0140] Delete the data points within the specified width range in the sample to be tested;
[0141] Based on the data points in the test sample after deleting data points, determine the similarity and fitting parameters of the next target data point.
[0142] The data processing device 1000 also includes:
[0143] The fourth determining module 1017 is used to determine the similarity based on the sum of the differences between the two signal-to-noise ratios corresponding to the same dispersion value in the first fitting function and the second fitting function. The larger the sum of the differences, the smaller the similarity.
[0144] The data processing device 1000 also includes:
[0145] The third fitting module 1018 is used to fit multiple data points within the fitting width range of the target data point in the test sample based on a preset Lorentz distribution function, so as to determine the third fitting function.
[0146] Update module 1019 is used to update the fitting width based on the third fitting function.
[0147] Each module in the aforementioned data processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0148] The data processing device can be integrated into a terminal 10 and / or a cloud server 20 that has storage and a processor and thus computing power, or the data processing device can be the terminal 10 and / or the cloud server 20.
[0149] Optionally, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0150] Figure 16 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. The computer device may be... Figure 2 The terminal 10 or cloud server 20 shown. Figure 16 As shown, the computer device 4000 may include: a communication interface 4010, a memory 4020, a processor 4030, and a communication bus 4040. The communication interface 4010, memory 4020, and processor 4030 communicate with each other via the communication bus 4040. The communication interface 4010 is used for data communication between the data processing device 1000 and external devices. The memory 4020 can be used to store software programs and modules, and the processor 4030 runs the software programs and modules stored in the memory 4020, such as the software programs for corresponding operations in the aforementioned method embodiments.
[0151] Optionally, the processor 4030 can call software programs and modules stored in the memory 4020 to perform the following operations: acquire a test sample, the test sample including multiple data points, the data points including dispersion values and signal-to-noise ratios; determine the target data point as the peak among the multiple data points; based on a preset Lorentz distribution function, fit multiple data points in the test sample within a preset width range of the target data point to determine a first fitting function, the first fitting function including a fitting width; determine the outer envelope data points based on the multiple data points within the fitting width range of the target data points, where the signal-to-noise ratio of the outer envelope data points is the largest when the dispersion values are the same; fit the outer envelope data points based on the preset Lorentz distribution function to determine a second fitting function, the second fitting function including fitting parameters; and determine the test sample as a suspected pulsar sample if the similarity between the first fitting function and the second fitting function is greater than a preset similarity and the fitting parameters meet preset conditions.
[0152] This application also provides a computer-readable storage medium for storing a computer program. This computer-readable storage medium can be applied to a computer device, and the computer program causes the computer device to execute the corresponding flow in the data processing method of the embodiments of this application; for brevity, further details are omitted here.
[0153] This application also provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding flow in the data processing method of the embodiments of this application. For simplicity, further details are omitted here.
[0154] This application also provides a computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding flow in the data processing method of the embodiments of this application. For simplicity, further details are omitted here.
[0155] It should be understood that the processor in the embodiments of this application may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0156] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0157] It should be understood that the above-described memory is exemplary and not a limiting description. For example, the memory in the embodiments of this application may also be static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DR RAM), etc. That is to say, the memory in the embodiments of this application is intended to include, but is not limited to, these and any other suitable types of memory.
[0158] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0159] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0160] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0161] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0162] In addition, the functional modules in the embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0163] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer or cloud server 20) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0164] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A data processing method, characterized in that, include: Obtain a test sample, which includes multiple data points, including dispersion values and signal-to-noise ratio; Identify the target data point as the peak value among multiple data points; Based on a preset Lorentz distribution function, multiple data points in the test sample within a preset width range of the target data point are fitted to determine a first fitting function, wherein the first fitting function includes a fitting width. Based on multiple data points within the fitting width range of the target data points, an outer envelope data point is determined, wherein the signal-to-noise ratio of the outer envelope data point is maximized when the dispersion values are the same. Based on a preset Lorentz distribution function, the outer envelope data points are fitted to determine a second fitting function, which includes fitting parameters. and If the similarity between the first fitting function and the second fitting function is greater than a preset similarity, and the fitting parameters meet preset conditions, the sample to be tested is determined to be a suspected pulsar sample.
2. The data processing method as described in claim 1, characterized in that, The determination of the target data point as the peak value among multiple data points includes: Compare the dispersion values of any two data points to determine one or more peak data points with the highest signal-to-noise ratio within a predetermined width range; The target data point is determined based on the peak data point.
3. The data processing method as described in claim 2, characterized in that, Determining the target data point based on the peak data point includes: The peak data points with a signal-to-noise ratio greater than a preset value are identified as the target data points.
4. The data processing method as described in claim 1, characterized in that, Also includes: Based on the preset Lorentz distribution function, multiple data points within the fitting width range of the target data points in the test sample are fitted to determine the third fitting function; The fitting width is updated according to the third fitting function.
5. The data processing method as described in claim 1, characterized in that, Also includes: The similarity is determined by the sum of the differences between the two signal-to-noise ratios corresponding to the same dispersion value in the first fitting function and the second fitting function. The larger the sum of the differences, the smaller the similarity.
6. The data processing method as described in claim 4, characterized in that, The fitting parameters include a first fitting parameter and a second fitting parameter. The first fitting parameter is the dispersion value of the target data point, and the second fitting parameter is half the absolute value of the difference between the two dispersion values corresponding to half the signal-to-noise ratio of the target data point in the second fitting function. The fitting parameters satisfy the preset conditions including: the first fitting parameter is greater than a first preset value and the second fitting parameter is less than a second preset value.
7. The data processing method as described in claim 1, characterized in that, The target data points are multiple, and the data processing method further includes: If the similarity between the first fitting function and the second fitting function corresponding to any target data point is greater than a preset similarity, and the fitting parameters meet the preset conditions, the sample to be tested is determined to be a suspected pulsar sample.
8. The data processing method as described in claim 1, characterized in that, The target data points are multiple, and the data processing method further includes: If the similarity between the first fitting function and the second fitting function corresponding to the previous target data point is less than the preset similarity, or the fitting parameters do not meet the preset conditions, the similarity and fitting parameters corresponding to the next target data point are obtained; wherein, the next target data point is a target data point adjacent to the previous target data point whose similarity and fitting parameters have been determined; The process of repeatedly executing the step of obtaining the similarity and fitting parameters corresponding to the next target data point when the similarity between the first fitting function and the second fitting function corresponding to the previous target data point is less than the preset similarity or the fitting parameters do not meet the preset conditions is repeated until the similarity and fitting parameters corresponding to all target data points are determined, or the similarity between the first fitting function and the second fitting function corresponding to any target data point is greater than the preset similarity and the fitting parameters meet the preset conditions.
9. The data processing method as described in claim 8, characterized in that, The step of obtaining the similarity and fitting parameters corresponding to the next target data point includes: The deletion width is determined based on the fitting width of the second fitting function and the predetermined width; Delete the data points within the deletion width range in the sample to be tested; Based on the data points in the test sample after deleting data points, the similarity and fitting parameters of the next target data point are determined.
10. A data processing apparatus, characterized in that, The device includes: An acquisition module is used to acquire a sample to be tested, the sample to be tested including multiple data points, the data points including dispersion value and signal-to-noise ratio; The first determining module is used to determine the target data point as the peak value among multiple data points; The first fitting module is used to fit multiple data points in the test sample within a preset width range of the target data points based on a preset Lorentz distribution function, so as to determine a first fitting function, wherein the first fitting function includes a fitting width. The second determining module is used to determine the outer envelope data point based on multiple data points within the fitting width range of the target data point, wherein the signal-to-noise ratio of the outer envelope data point is the largest when the dispersion values are the same. The second fitting module is used to fit the outer envelope data points based on a preset Lorentz distribution function to determine a second fitting function, which includes fitting parameters. The third determining module is used to determine the sample to be tested as a suspected pulsar sample when the similarity between the first fitting function and the second fitting function is greater than a preset similarity and the fitting parameters meet preset conditions.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted for loading by a processor to perform the steps of the data processing method as described in any one of claims 1-9.
12. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program, and the processor executing the steps of the data processing method according to any one of claims 1-9 by calling the computer program stored in the memory.
13. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the data processing method according to any one of claims 1-9.