Launching line star identification method and system based on LAMOST spectral data

By applying the deep residual network model Hα-ResNet based on LAMOST spectral data in astronomical data processing, the problem of low accuracy of multi-classification recognition of radial star in traditional methods is solved, and efficient classification of multi-classification and accurate identification of radial star is achieved.

CN120045996AActive Publication Date: 2025-05-27GUANGZHOU UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510101363.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Traditional astronomical data processing methods are difficult to effectively process massive spectral data, resulting in low multi-classification recognition accuracy of radiated stars and unable to achieve efficient classification of multi-radiation stars.

Method used

Using the deep residual network model Hα-ResNet based on LAMOST spectral data, the shortcut residual path is constructed, data preprocessing and feature extraction are performed to realize binary classification and six-category identification of resolution spectral data in LAMOST Dr9.

Benefits of technology

The multi-classification recognition accuracy of transmitting stars is improved, efficient classification of multi-radiation stars is achieved, and the classification and identification accuracy of spectral data is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045996A_ABST
    Figure CN120045996A_ABST
Patent Text Reader

Abstract

The invention discloses an emission line star identification method and system based on LAMOST spectral data, and the method comprises the steps: constructing LAMOST Dr9 medium-resolution spectral data, and carrying out the data preprocessing, and obtaining a spectral data feature matrix; based on the deep residual network, introducing a shortcut residual path, and constructing an H alpha-ResNet model; carrying out dichotomy training on the H alpha-ResNet model, constructing a trained dichotomy H alpha-ResNet model, and carrying out dichotomy identification to obtain launch line star sample data; and performing six-classification training on the H alpha-ResNet model after preprocessing based on the launch line star sample data, constructing the trained six-classification H alpha-ResNet model, and performing six-classification identification to obtain a launch line star identification result. According to the invention, the multi-emission line stars can be classified, and the classification and identification accuracy of the spectral data can be improved. The method and the system for identifying the transmitting line stars based on the LAMOST spectral data can be widely applied to the technical field of astronomical data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of astronomical data processing, and in particular to an emission-line star identification method and system based on LAMOST spectrum data. Background Art

[0002] Modern astronomical observation equipment, such as radio telescopes, optical telescopes, and space telescopes, cannot do without the support of electronic information technology. For example, radio telescopes need to use electronic equipment to receive and process radio signals from the universe, while optical telescopes need to use electronic equipment to convert the received light signals into electrical signals for processing. As astronomy enters the era of big data, the amount of data collected by astronomical observation equipment is huge, and complex processing and analysis are required to obtain valuable scientific results. It is further necessary to use electronic information technology, such as data mining, machine learning and other methods, to efficiently process and analyze massive amounts of astronomical data.

[0003] Faced with massive amounts of spectral data and complex astrophysical problems, traditional data processing methods are restricted by the limitations of manual feature extraction and pattern recognition, and are unable to fully explore the potential of the data and provide accurate classification and identification results. In addition, due to the lack of high-confidence emission-line star samples, the accuracy of multi-classification models is difficult to exceed 75%. As a result, most of the current emission-line star classifications are still at the stage of classifying a single emission-line star, and cannot achieve classification of multiple emission-line stars. Summary of the invention

[0004] In order to solve the above technical problems, the purpose of the present invention is to provide an emission-line star identification method and system based on LAMOST spectral data, which can classify multi-emission-line stars and improve the classification and identification accuracy of spectral data.

[0005] The first technical solution adopted by the present invention is: a method for identifying emission-line stars based on LAMOST spectral data, comprising the following steps:

[0006] Construct LAMOSTDr9 medium-resolution spectral data and perform data preprocessing to obtain the spectral data feature matrix;

[0007] Based on the deep residual network, a shortcut residual path is introduced to build the Hα-ResNet model;

[0008] Based on the spectral data feature matrix, the Hα-ResNet model is trained for binary classification, and the trained binary Hα-ResNet model is constructed to perform binary classification recognition on the LAMOSTDr9 medium-resolution spectral data to obtain emission-line star sample data.

[0009] After preprocessing the emission line star sample data, the Hα-ResNet model is trained for six-class classification, and the trained six-class Hα-ResNet model is constructed and used to perform six-class classification recognition on the LAMOST Dr9 medium-resolution spectral data to obtain the emission line star recognition results.

[0010] Furthermore, the step of constructing the LAMOST Dr9 medium-resolution spectral data and performing data preprocessing to obtain the spectral data feature matrix specifically includes:

[0011] Obtain the observation data through the LAMOST official website database and obtain the emission line star candidate catalog through the target database;

[0012] Cross-match the observation data and the emission line star candidate catalog with the LAMOST Dr9 database respectively to construct the LAMOST Dr9 medium-resolution spectral data;

[0013] Extract the target spectral data segment from the LAMOST Dr9 medium-resolution spectral data to obtain the spectral red-end data;

[0014] Perform data preprocessing on the spectral red-end data to obtain the spectral data feature matrix.

[0015] Furthermore, the step of performing data preprocessing on the spectral red-end data to obtain the spectral data feature matrix specifically includes:

[0016] Perform median filtering on the spectral red-end data to obtain the filtered spectral red-end data;

[0017] Perform normalization on the filtered spectral red-end data to obtain the normalized spectral red-end data;

[0018] Perform data cleaning on the normalized spectral red-end data to obtain the cleaned spectral red-end data;

[0019] Perform feature selection on the cleaned spectral red-end data, assign labels, and store them in binary file format to obtain the spectral data feature matrix.

[0020] Furthermore, the Hα-ResNet model specifically includes a main path residual block, a shortcut residual path, an activation function, a pooling layer, and a linear layer. The output ends of the main path residual block and the shortcut residual path are both connected to the activation function, and the activation function, the pooling layer, and the linear layer are connected in sequence, where:

[0021] The main path residual block includes a first convolutional layer, a first batch normalization layer, a first activation function layer, a second convolutional layer, and a second batch normalization layer;

[0022] The shortcut residual path includes a third convolutional layer and a third batch normalization layer.

[0023] Further, the step of performing binary classification training on the Hα-ResNet model based on the spectral data feature matrix, constructing the trained binary classification Hα-ResNet model, and performing binary classification recognition on the LAMOST Dr9 medium-resolution spectral data to obtain the emission line star sample data specifically includes:

[0024] Input the spectral data feature matrix into the Hα-ResNet model;

[0025] Based on the main path residual block of the Hα-ResNet model, perform feature extraction processing on the spectral data feature matrix to obtain the first spectral residual feature data;

[0026] Based on the shortcut residual path of the Hα-ResNet model, perform feature extraction processing on the first spectral residual feature data to obtain the second spectral residual feature data;

[0027] Combine the first spectral residual feature data and the second spectral residual feature data to obtain the spectral residual feature data;

[0028] Based on the activation function and pooling layer of the Hα-ResNet model, perform activation pooling processing on the spectral residual feature data to obtain the pooled spectral residual feature data;

[0029] Based on the linear layer of the Hα-ResNet model, map the pooled spectral residual feature data to the class space, obtain the class prediction result of the spectral data, and output the trained binary classification Hα-ResNet model;

[0030] Store the LAMOST Dr9 medium-resolution spectral data in an HDF5 file, and perform binary classification recognition on the LAMOST Dr9 medium-resolution spectral data based on the trained binary classification Hα-ResNet model to obtain the emission line star sample data.

[0031] Further, the step of performing six-classification training on the Hα-ResNet model after preprocessing based on the emission line star sample data, constructing the trained six-classification Hα-ResNet model, and performing six-classification recognition on the LAMOST Dr9 medium-resolution spectral data to obtain the emission line star recognition result specifically includes:

[0032] Perform cross-processing on the emission line star sample data and the Simbad database to obtain the emission line star sample data with labels;

[0033] Perform data preprocessing on the emission line star sample data with labels to obtain the preprocessed emission line star sample data;

[0034] Perform six-class training on the Hα-ResNet model based on the preprocessed emission-line star sample data, and construct the trained six-class Hα-ResNet model;

[0035] Perform six-class recognition on the LAMOST Dr9 medium-resolution spectroscopic data based on the trained six-class Hα-ResNet model to obtain the emission-line star recognition result.

[0036] Further, the step of performing data preprocessing on the emission-line star sample data with labels to obtain the preprocessed emission-line star sample data specifically includes:

[0037] Based on the emission-line star sample data with labels, obtain the blue-end data and red-end data of the emission-line star sample data;

[0038] Perform median filtering, normalization, and data cleaning on the blue-end data and red-end data of the emission-line star sample data respectively to obtain the cleaned blue-end data and cleaned red-end data;

[0039] Merge the cleaned blue-end data and cleaned red-end data to obtain the preprocessed emission-line star sample data.

[0040] The second technical solution adopted by the present invention is: An emission-line star recognition system based on LAMOST spectroscopic data, including:

[0041] The first module is used to construct the LAMOST Dr9 medium-resolution spectroscopic data and perform data preprocessing to obtain the spectroscopic data feature matrix;

[0042] The second module is used to construct the Hα-ResNet model based on the deep residual network and introduce a shortcut residual path;

[0043] The third module is used to perform two-class training on the Hα-ResNet model based on the spectroscopic data feature matrix, construct the trained two-class Hα-ResNet model, and perform two-class recognition on the LAMOST Dr9 medium-resolution spectroscopic data to obtain the emission-line star sample data;

[0044] The fourth module is used to perform six-class training on the Hα-ResNet model after preprocessing based on the emission-line star sample data, construct the trained six-class Hα-ResNet model, and perform six-class recognition on the LAMOST Dr9 medium-resolution spectroscopic data to obtain the emission-line star recognition result.

[0045] The beneficial effects of the method and system of the present invention are as follows: By constructing the medium-resolution spectral data of LAMOST Dr9 and performing data preprocessing, the accuracy of the model is significantly improved through the selection of data features. Then, based on the deep residual network, a shortcut residual path is introduced to construct the Hα-ResNet model. And based on the spectral data feature matrix, binary classification training is carried out on the Hα-ResNet model. After constructing the trained binary classification Hα-ResNet model, binary classification recognition is performed on the medium-resolution spectral data of LAMOST Dr9. By inputting the preprocessing results into the Hα-ResNet model for training and testing, a model that can efficiently and accurately identify whether it is an emission-line star is trained. Finally, cross-processing is performed with the Simbad database to obtain the sample data of emission-line stars with labels. After preprocessing the data, six-classification training is carried out on the Hα-ResNet model. After constructing the trained six-classification Hα-ResNet model, six-classification recognition is performed on the medium-resolution spectral data of LAMOST Dr9 to obtain high-quality detailed classification labels of emission-line stars, realizing the classification of multiple emission-line stars. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is the flowchart of the steps of a method for identifying emission-line stars based on LAMOST spectral data according to the present invention;

[0047] Figure 2 is the structural block diagram of a system for identifying emission-line stars based on LAMOST spectral data according to the present invention;

[0048] Figure 3 is the schematic diagram of preprocessing red-end data provided by a specific embodiment of the present invention;

[0049] Figure 4 is the structural schematic diagram of the Hα-ResNet model provided by a specific embodiment of the present invention;

[0050] Figure 5 is the schematic diagram of processing the HDF5 file provided by a specific embodiment of the present invention;

[0051] Figure 6 is the schematic diagram of preprocessing red-end and blue-end data provided by a specific embodiment of the present invention;

[0052] Figure 7 is the schematic diagram of the confusion matrix of the binary classification Hα-ResNet model provided by a specific embodiment of the present invention;

[0053] Figure 8 is the schematic diagram of the confusion matrix of the six-classification Hα-ResNet model provided by a specific embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0055] First of all, it should be noted that in the past, the commonly used method was to use spectral features for classification. According to characteristic parameters such as the position, intensity, width, and redshift of emission lines, emission-line stars can be distinguished from other types of celestial bodies. Traditional statistical methods, feature engineering, and machine learning algorithms (such as support vector machines, decision trees, and neural networks) have all been applied to the identification of emission-line stars. At present, deep learning algorithms have achieved remarkable results in many fields, including image recognition, speech recognition, and natural language processing. In the identification of emission-line stars, deep learning algorithms have also gradually been applied. For example, convolutional neural networks (CNNs) are widely used to process image data. These deep learning algorithms can automatically extract features from raw data and have high accuracy and generalization ability. However, due to the uncertainty of spectral data features, the classification effect of deep learning in the field of multi-spectral classification generally cannot exceed 90%, and the six-class classification with similar spectral features is difficult to exceed 75%. Coupled with the lack of high-confidence samples for the detailed classification of emission-line stars, it will further lead to a significant reduction in the training and test accuracy of the model.

[0056] Based on this, the embodiment of the present invention preprocesses the medium-resolution spectral data in LAMOST Dr9 to obtain a feature matrix of the spectral data, inputs the feature matrix into the Hα-ResNet model for training and testing, and then uses the trained model for the binary classification recognition of emission lines and non-emission lines in the medium-resolution data of LAMOST Dr9; cross-reference with Simbad according to the classification results to obtain the specific categories of emission-line stars. Then train a new Hα-ResNet six-classification model to screen the medium-resolution data of LAMOST Dr9 again to obtain labeled emission-line star candidates. Through the binary classification model, high-confidence emission-line star classification samples can be obtained, and the Hα-ResNet six-classification model can efficiently and accurately identify the categories of emission-line stars.

[0057] Refer to Figure 1 , the present invention provides a method for identifying emission-line stars based on LAMOST spectral data, and the method includes the following steps:

[0058] S100. Construct the medium-resolution spectral data of LAMOST Dr9 and perform data preprocessing to obtain a spectral data feature matrix;

[0059] Specifically, obtain the observation data through the LAMOST official website database and obtain the emission-line star candidate catalog through the target database; cross-match the observation data and the emission-line star candidate catalog with the LAMOST Dr9 database respectively to construct the LAMOST Dr9 medium-resolution spectral data; perform target spectral data segment extraction processing on the LAMOST Dr9 medium-resolution spectral data to obtain the spectral red-end data; perform data preprocessing on the spectral red-end data to obtain the spectral data feature matrix.

[0060] In this embodiment, the observation data can be directly obtained from the LAMOST official website, and the emission-line star candidates can be obtained from high-quality academic papers, and then cross-matched with the LAMOST Dr9 database to obtain the original data for model training. After obtaining the observation data, the spectral red-end data is extracted from the observation data. The present invention embodiment takes the FITS file as an example for illustration. First, clarify the extensions in the FITS file. The first element is the primary extension, which has no data. The second to the last ones are all corresponding spectral files with data. The second file stores the blue-end data, and the third file stores the red-end data, arranged in sequence. Eliminate the bad files with a signal-to-noise ratio lower than 10 and no data. Then open the FITS file to find the corresponding red-end spectrum and read the data. Preprocess the extracted spectrum and store it in a.npy file.

[0061] Among them, it should be noted that the step process of data preprocessing for the spectral red-end data is as follows: perform median filtering processing on the spectral red-end data to obtain the filtered spectral red-end data; perform normalization processing on the filtered spectral red-end data to obtain the normalized spectral red-end data; perform data cleaning processing on the normalized spectral red-end data to obtain the cleaned spectral red-end data; perform feature selection and assign labels to the cleaned spectral red-end data and then store it in the binary file format to obtain the spectral data feature matrix.

[0062] In this embodiment, median filtering is first performed, using the median_filter function from the scipy.ndimage library. Effective suppression of image noise is achieved while maintaining the edge clarity and detail features of the image as much as possible; further normalization is performed, firstly median normalization is performed to avoid local feature loss, and the specific processing process involves scaling or converting the data values ​​in each local area or feature so that it has a uniform distribution or scale while maintaining the relative differences and characteristics between these local areas or features. Afterwards, global normalization is performed by scaling the values ​​of the entire data set (all features) to a uniform range (such as 0 to 1), which is achieved by subtracting the minimum value of the data and then dividing by the difference between the maximum and the minimum value. This normalization method is very important for ensuring the comparability between different features or different samples, especially in machine learning, which can avoid problems caused by differences in numerical ranges, and the goal is to ensure that all spectral data are compared on the same scale; and then data cleaning is performed to remove dirty data such as missing values ​​and outliers from the original data; and then data feature selection is performed, and according to the emission lines at specific positions of the spectrum, specific lengths are intercepted on the left and right to determine the size of the spectral data. Specifically, around 6565.5 angstroms, 200 data points are selected on each side. The processed data is as follows Figure 3 As shown; finally, the processed data is stored in a .npy file, each piece of processed data is given a label, and the label and data are stored in two .npy files respectively.

[0063] In summary, this embodiment can help reduce the noise in the data through median filtering, which helps to improve the training effect and generalization ability of the model; normalization helps to optimize the convergence speed of the algorithm and prevent certain features from dominating the training due to their large numerical range; data cleaning can remove or correct erroneous, missing or abnormal data points, improve the overall quality of the data set, and thus improve the performance of the model; feature selection helps to reduce the complexity of the model, reduce the risk of overfitting, and improve computational efficiency. For the identification of emission-line stars, this step will greatly improve the model accuracy; storing data in .npy files can easily load data into the model training and testing process, avoiding repeated data preprocessing steps, saving time and computing resources, and ensuring the reproducibility of the experiment, because the same data preprocessing steps will always produce the same output, which is crucial for the verification of scientific experiments and the reliability of results; .npy files can be easily shared with others, facilitating team collaboration and verification of results; the format of .npy files is compatible with a variety of programming languages ​​and data analysis tools, which facilitates data analysis and model training using different tools and platforms.

[0064] S200, based on the deep residual network, introduce the shortcut residual path and build the Hα-ResNet model;

[0065] Specifically, as Figure 4 shown, the Hα-ResNet model specifically includes a main path residual block, a shortcut residual path, an activation function, a pooling layer, and a linear layer. The output ends of the main path residual block and the shortcut residual path are both connected to the activation function, and the activation function, the pooling layer, and the linear layer are connected in sequence. Among them, the main path residual block includes a first convolutional layer, a first batch normalization layer, a first activation function layer, a second convolutional layer, and a second batch normalization layer, and the shortcut residual path includes a third convolutional layer and a third batch normalization layer.

[0066] S300. Perform binary classification training on the Hα-ResNet model based on the spectral data feature matrix, construct the trained binary classification Hα-ResNet model, and perform binary classification recognition on the LAMOST Dr9 medium-resolution spectral data to obtain the emission line star sample data;

[0067] Specifically, input the spectral data feature matrix into the Hα-ResNet model; based on the main path residual block of the Hα-ResNet model, perform feature extraction processing on the spectral data feature matrix to obtain the first spectral residual feature data; based on the shortcut residual path of the Hα-ResNet model, perform feature extraction processing on the first spectral residual feature data to obtain the second spectral residual feature data; combine the first spectral residual feature data and the second spectral residual feature data to obtain the spectral residual feature data; based on the activation function and the pooling layer of the Hα-ResNet model, perform activation pooling processing on the spectral residual feature data to obtain the pooled spectral residual feature data; based on the linear layer of the Hα-ResNet model, map the pooled spectral residual feature data to the class space, obtain the class prediction result of the spectral data, and output the trained binary classification Hα-ResNet model; store the LAMOST Dr9 medium-resolution spectral data in an HDF5 file, and perform binary classification recognition on the LAMOST Dr9 medium-resolution spectral data based on the trained binary classification Hα-ResNet model to obtain the emission line star sample data.

[0068] In this embodiment, the Hα-ResNet model defines two main classes: ResidualBlock (residual block) and ResNet_Model. Both of these classes are based on nn.Module of the PyTorch framework. The ResidualBlock class implements a basic residual block structure, which is the basic unit for constructing a deep residual network. The residual block solves the problem of vanishing gradients or exploding gradients in the training of deep neural networks by introducing a "shortcut" connection. This invention includes two 3x3 convolutional layers, each followed by a batch normalization layer. The forward propagation first passes through the first convolutional layer and the batch normalization layer, and then applies the ReLU activation function. Then it passes through the second convolutional layer and the batch normalization layer. Finally, the output of the main path is added to the output of the shortcut connection, and the ReLU activation function is applied again to obtain the final output. The ResNet_Model class defines the Hα-ResNet model based on the residual block, including the type of the residual block (block), the number of residual blocks in each layer (num_block), and the number of output classes (num_class). When initializing, the first convolutional layer (conv1), the batch normalization layer (bn1), the ReLU activation function, and the max pooling layer (maxpool) are defined. layer1 and layer2 are residual block layers constructed by the _make_layer method. The last linear layer (linear) maps the features to the class space. The _make_layer method is used to construct a layer composed of multiple residual blocks. The strides list contains the stride of each layer. The stride of the first residual block is specified by the parameter, and the rest are all 1. Traverse the strides, create a residual block for each stride, and update self.in_planes to match the input channel number of the next residual block. The forward propagation method includes the input x passing through the first convolutional layer, the batch normalization layer, the ReLU activation function, and the max pooling layer, and then passing through layer1 and layer2 in sequence. The average pooling layer is applied to reduce the size of the feature map. The feature map is flattened and passed through the linear layer to obtain the final class prediction. A model with a test accuracy higher than 99.5% is trained and the model is saved. The confusion matrix of the model is as Figure 7 shown.

[0069] S400. After preprocessing the emission-line star sample data, perform six-class training on the Hα-ResNet model, construct the trained six-class Hα-ResNet model, and perform six-class recognition on the LAMOST Dr9 medium-resolution spectral data to obtain the emission-line star recognition result.

[0070] Specifically, the emission-line star sample data is cross-processed with the Simbad database to obtain the emission-line star sample data with labels; the emission-line star sample data with labels is preprocessed to obtain the preprocessed emission-line star sample data; the Hα-ResNet model is trained for six-classification based on the preprocessed emission-line star sample data to construct the trained six-classification Hα-ResNet model; the six-classification recognition of the LAMOST Dr9 medium-resolution spectroscopic data is performed based on the trained six-classification Hα-ResNet model to obtain the emission-line star recognition result.

[0071] In this embodiment, the trained two-classification Hα-ResNet model is applied to the LAMOST Dr9 medium-resolution data already stored in the HDF5 file to screen out the emission-line stars. Specifically, the operations of reading the HDF5 file and preprocessing the data are as Figure 5 shown. Open the HDF5 file, select the red-end data and perform the next data preprocessing operation; perform median filtering, normalization, data cleaning, and data feature selection operations; input the preprocessed data into the model; output the file names of the emission-line stars that meet the conditions according to the model prediction results and save them in a.txt file.

[0072] Furthermore, it should be noted that the process of data cross with the Simbad database is to obtain the RA and DEC of the star catalog by cross the file names of the emission-line stars saved in the.txt file with the LAMOST Dr9 official website; upload the obtained star catalog to the TOPCAT software and cross it with the Simbad database to obtain the emission-line stars with labels, and further preprocess the data of the emission-line stars with labels. First, eliminate the bad files with a signal-to-noise ratio lower than 10 and no data, and then read the blue-end and red-end data in the FITS file; then perform median filtering, normalization, data cleaning, and merging of the blue-end and red-end data on the data. The processed data is as Figure 6 shown.

[0073] In summary, by inputting the preprocessing results into the Hα-ResNet model for training and testing, a model that can efficiently and accurately identify the category of emission-line stars is trained, which makes up for the shortcomings of low recognition accuracy and poor processing efficiency of traditional methods, and greatly improves the accuracy of the model through the selection of data features. HDF5 (Hierarchical Data Format version 5) can store and organize large amounts of data. It supports compression and chunking, and large data sets can be stored and accessed more efficiently. HDF5 files contain metadata about their contents, making the files easier to understand. It allows users to define new data types, allowing them to adapt to various complex data structures and support multidimensional arrays, making them very suitable for storing and processing scientific data. HDF5 files allow multiple data sets to be stored in the same file, which helps to organize and manage related data, and supports parallel processing, which makes it possible to use it in high-performance computing environments. The obtained classification results are crossed with Simbad to obtain label classification, so that high-quality emission-line star subdivision category labels can be obtained.

[0074] Further pre-train the model again, load the processed data into the model, and divide the training set and test set into 8:2 ratio; train the Hα-ResNet model and ensure that the test accuracy of the model is higher than 87%. The model confusion matrix is ​​as follows: Figure 8 As shown; Save the trained model.

[0075] Further, it should be explained that the process of data preprocessing for the emission-line star spectral data with labels is to obtain the blue-end data of the emission-line star spectral data and the red-end data of the emission-line star spectral data based on the emission-line star spectral data with labels; perform median filtering, normalization and data cleaning on the blue-end data of the emission-line star spectral data and the red-end data of the emission-line star spectral data respectively to obtain the cleaned blue-end data and the cleaned red-end data; merge the cleaned blue-end data and the cleaned red-end data to obtain the preprocessed emission-line star spectral data.

[0076] In this embodiment, according to Figure 5 The process shown in the figure obtains the blue-end and red-end data by reading the FITS file; then performs median filtering, normalization, data cleaning, and merging of the blue-end and red-end data on the data; inputs the preprocessed data into the model; and finally outputs the emission line star file name that meets the conditions according to the model prediction results and saves it in a .txt file.

[0077] In summary, in the embodiment of the present invention, a binary classification Hα-ResNet model is first used to obtain emission line star samples with high confidence, so as to solve the problem that the accuracy of the model cannot be improved due to the scarcity of samples. Then, a six-classification Hα-ResNet model is used to provide more emission line star samples with high confidence for the LAMOST database. At the same time, the high performance of the deep learning model is used to solve the problems of low accuracy and poor efficiency of traditional recognition methods.

[0078] Referring to Figure 2 , an emission line star recognition system based on LAMOST spectral data includes:

[0079] The first module 201 is used to construct LAMOST Dr9 medium-resolution spectral data and perform data preprocessing to obtain a spectral data feature matrix;

[0080] The second module 202 is used to construct an Hα-ResNet model based on the deep residual network and introduce a shortcut residual path;

[0081] The third module 203 is used to perform binary classification training on the Hα-ResNet model based on the spectral data feature matrix, construct a trained binary classification Hα-ResNet model and perform binary classification recognition on the LAMOST Dr9 medium-resolution spectral data to obtain emission line star sample data;

[0082] The fourth module 204 is used to perform preprocessing on the emission line star sample data and then perform six-classification training on the Hα-ResNet model, construct a trained six-classification Hα-ResNet model and perform six-classification recognition on the LAMOST Dr9 medium-resolution spectral data to obtain an emission line star recognition result.

[0083] The content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0084] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A method for identifying emission-line stars based on LAMOST spectral data, characterized in that: The following steps are involved: Construct LAMOST Dr9 medium-resolution spectral data and perform data preprocessing to obtain the spectral data feature matrix; Based on the deep residual network, a shortcut residual path is introduced to build the Hα-ResNet model; Based on the spectral data feature matrix, the Hα-ResNet model is trained for binary classification, and the trained binary Hα-ResNet model is constructed to perform binary classification recognition on the LAMOST Dr9 medium-resolution spectral data to obtain emission-line star sample data. After preprocessing the emission-line star sample data, the Hα-ResNet model was trained for six categories. The trained six-category Hα-ResNet model was constructed and six-category recognition was performed on the LAMOST Dr9 medium-resolution spectral data to obtain the emission-line star identification results.

2. According to claim 1, a method for identifying emission-line stars based on LAMOST spectrum data, characterized in that: The step of constructing LAMOST Dr9 medium-resolution spectral data and performing data preprocessing to obtain a spectral data feature matrix specifically includes: Obtain observation data through the LAMOST official website database and obtain the emission-line star candidate catalog through the target database; The observation data and emission-line star candidate catalogs were cross-matched with the LAMOST Dr9 database to construct the LAMOST Dr9 medium-resolution spectral data; The target spectrum data segment is extracted from the LAMOST Dr9 medium-resolution spectrum data to obtain the spectrum red end data; The red end data of the spectrum is preprocessed to obtain the spectral data feature matrix.

3. According to claim 2, a method for identifying emission-line stars based on LAMOST spectrum data, characterized in that: The step of preprocessing the spectral red end data to obtain a spectral data feature matrix specifically includes: Perform median filtering on the red end data of the spectrum to obtain filtered red end data of the spectrum; The filtered red-end spectrum data is normalized to obtain normalized red-end spectrum data; Performing data cleaning on the normalized red-end spectrum data to obtain cleaned red-end spectrum data; The cleaned spectral red-end data is subjected to feature selection and assigned labels, and then stored in a binary file format to obtain a spectral data feature matrix.

4. According to claim 3, a method for identifying emission-line stars based on LAMOST spectrum data, characterized in that: The Hα-ResNet model specifically includes a main path residual block, a shortcut residual path, an activation function, a pooling layer, and a linear layer. The output end of the main path residual block and the output end of the shortcut residual path are both connected to the activation function. The activation function, the pooling layer, and the linear layer are connected in sequence, wherein: The main path residual block includes a first convolution layer, a first batch normalization layer, a first activation function layer, a second convolution layer and a second batch normalization layer; The shortcut residual path includes a third convolutional layer and a third batch normalization layer.

5. According to claim 4, a method for identifying emission-line stars based on LAMOST spectrum data, characterized in that: The step of performing binary classification training on the Hα-ResNet model based on the spectral data feature matrix, constructing the trained binary classification Hα-ResNet model and performing binary classification recognition on the LAMOST Dr9 medium-resolution spectral data to obtain emission-line star sample data specifically includes: Input the spectral data feature matrix into the Hα-ResNet model; Based on the main path residual block of the Hα-ResNet model, feature extraction processing is performed on the spectral data feature matrix to obtain first spectral residual feature data; Based on the shortcut residual path of the Hα-ResNet model, feature extraction processing is performed on the first spectral residual feature data to obtain second spectral residual feature data; Combining the first spectrum residual feature data with the second spectrum residual feature data to obtain spectrum residual feature data; Based on the activation function and pooling layer of the Hα-ResNet model, the spectral residual feature data is activated and pooled to obtain the pooled spectral residual feature data; Based on the linear layer of the Hα-ResNet model, the pooled spectral residual feature data is mapped to the category space to obtain the category prediction result of the spectral data and output the trained binary classification Hα-ResNet model; The medium-resolution spectral data of LAMOST Dr9 are stored in HDF5 files, and binary classification recognition is performed on the medium-resolution spectral data of LAMOST Dr9 based on the trained binary classification Hα-ResNet model to obtain the emission-line star sample data.

6. The emission-line star identification method based on LAMOST spectrum data according to claim 5, characterized in that: The step of performing six-category training on the Hα-ResNet model after preprocessing the emission-line star sample data, constructing the trained six-category Hα-ResNet model, and performing six-category recognition on the LAMOST Dr9 medium-resolution spectral data to obtain the emission-line star recognition result specifically includes: Cross-process the emission-line star sample data with the Simbad database to obtain the emission-line star sample data with labels; Performing data preprocessing on the emission-line star sample data with labels to obtain preprocessed emission-line star sample data; Based on the preprocessed emission-line star sample data, the Hα-ResNet model is trained for six categories, and a trained six-category Hα-ResNet model is constructed; Based on the trained six-category Hα-ResNet model, six-category recognition was performed on the LAMOST Dr9 medium-resolution spectral data to obtain the emission-line star identification results.

7. The emission-line star identification method based on LAMOST spectrum data according to claim 6, characterized in that: The step of performing data preprocessing on the emission-line star sample data with labels to obtain the preprocessed emission-line star sample data specifically includes: Based on the emission-line star sample data with labels, blue-end data of the emission-line star sample data and red-end data of the emission-line star sample data are obtained; The blue-end data of the emission-line star sample data and the red-end data of the emission-line star sample data are respectively subjected to median filtering, normalization and data cleaning to obtain cleaned blue-end data and cleaned red-end data; The cleaned blue-end data and the cleaned red-end data are merged to obtain the preprocessed emission-line star sample data.

8. An emission-line star identification system based on LAMOST spectrum data, characterized in that: Includes the following modules: The first module is used to construct LAMOST Dr9 medium-resolution spectral data and perform data preprocessing to obtain the spectral data feature matrix; The second module is used to introduce a shortcut residual path based on a deep residual network and build a Hα-ResNet model; The third module is used to perform binary classification training on the Hα-ResNet model based on the spectral data feature matrix, build the trained binary classification Hα-ResNet model, and perform binary classification recognition on the LAMOST Dr9 medium-resolution spectral data to obtain emission-line star sample data; The fourth module is used to perform six-category training on the Hα-ResNet model based on the emission-line star sample data after preprocessing, construct the trained six-category Hα-ResNet model, and perform six-category recognition on the LAMOST Dr9 medium-resolution spectral data to obtain the emission-line star recognition results.

Citation Information

Patent Citations

  • LAMOST low signal-to-noise ratio celestial body spectrum classification method based on YOLOv7 network

    CN116883734A

  • Hyperspectral image lithology identification method and system based on object-oriented spatial spectrum enhanced convolutional neural network model

    CN118298313A

  • Spectral mixture process conditioned by spatially-smooth partitioning

    US20050047663A1