Method for Identifying Molecular Cloud Clumps Based on Semi-Supervised Deep Learning
By applying semi-supervised deep learning method in molecular cloud cluster detection, using 3D convolutional neural network to extract features and perform semi-supervised learning, the problems of low efficiency and false targets in the existing technology are solved, and high accuracy and efficient automated certification are achieved.
Patent Information
- Application Number
- CN202311124385.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-01
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-09-01
AI Technical Summary
When detecting molecular cloud clusters in large-scale sky survey data, the prior art is inefficient and requires manual parameter adjustment, resulting in the existence of false targets that adversely affect scientific analysis.
Using a semi-supervised deep learning method, 3D convolutional neural network is used to extract the characteristics of molecular cloud clusters, and through semi-supervised learning training models, their generalization capabilities and data utilization are improved, and automated molecular cloud cluster authentication is achieved.
It realizes high-accuracy molecular cloud cluster certification, reduces manual intervention, improves detection efficiency, and ensures the correctness of the detection results.
Smart Images

Figure CN117315329B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of molecular cloud clump detection, and particularly relates to a method for identifying molecular cloud clumps based on semi-supervised deep learning. Background Art
[0002] Modern astronomy has confirmed that the formation site of stars is inside molecular cloud clumps. Detecting molecular cloud clumps, determining the various parameters of molecular cloud clumps, and statistical analysis of molecular cloud clumps are of great significance for the study of the laws of stellar evolution. With the development of many large-scale CO surveys at home and abroad, it is a huge challenge to detect molecular cloud clumps from a large amount of molecular cloud data. For the candidate molecular cloud clumps detected by traditional molecular cloud clump detection algorithms, it is necessary to combine manual identification to eliminate false targets in the candidates, so as to ensure the reliability of the clump targets used for scientific analysis. In large-scale surveys, the number of candidate molecular cloud clumps often reaches tens of thousands, and large-scale manual identification is unrealistic. Therefore, there is an urgent need for an automated molecular cloud clump authentication algorithm that can replace manual inspection.
[0003] Currently, algorithms widely used in molecular cloud clump detection include GaussClumps, ClumpFind, FellWalker, etc. The GaussClumps algorithm uses an iterative method to continuously perform three-dimensional Gaussian fitting on the data at the maximum peak position and regards the fitting region that meets the conditions as a molecular cloud clump. This method takes a long time. The ClumpFind detection algorithm uses the contour line method to divide the pixels of the contour line containing the extreme point to the extreme point as a molecular cloud clump. This algorithm is widely used. The FellWalker algorithm starts from the point with lower data intensity and searches upward along the direction of the maximum gradient to find local maxima, and divides the points on all paths that converge at the same peak position to the same molecular cloud clump. The comprehensive performance of this algorithm is good.
[0004] In the process of applying these algorithms, the detection parameters of the algorithms can be manually selected by using the prior knowledge of data and observation instruments and combining the morphological characteristics of the target to be searched, so as to achieve better detection results. However, the setting of algorithm parameters is based on sufficient prior knowledge of data and the analysis of detection results, and often requires repeated parameter tuning to meet the detection requirements. Usually, researchers will first design the initial parameters of the algorithm detection according to experience, and then observe whether the detected candidates meet the requirements. If the detection results are not ideal, the parameters will be modified and the detection will be repeated until the best detection effect is achieved. However, for large-scale sky survey data, such an approach is inefficient and cannot meet the needs of practical applications. If the parameter settings are unreasonable, a large number of false molecular cloud clumps introduced by background noise or other interferences will appear in the detection results, and manual inspection is often required to further screen the candidates of molecular cloud clumps. Otherwise, the existence of false clumps will have an adverse impact on subsequent scientific analysis. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a method for identifying molecular cloud clumps based on semi-supervised deep learning. This method uses a 3D convolutional neural network to extract the features of molecular cloud clumps and adopts semi-supervised learning to train the model, aiming to improve its generalization ability and data utilization rate. The present invention solves the problem of insufficient labeled samples in traditional supervised learning and enables the model to better adapt to new, unlabeled samples, achieving high accuracy in the verification of molecular cloud clumps.
[0006] The technical solution adopted by the present invention is as follows:
[0007] A method for identifying molecular cloud clumps based on semi-supervised deep learning, comprising the following steps:
[0008] Step 1: Obtain candidates of molecular cloud clumps;
[0009] Step 2: Train the SS-3D-Clump model based on the candidates of molecular cloud clumps and output the probability value of the molecular cloud clumps;
[0010] Step 3: Determine the probability threshold. When the output probability value exceeds the probability threshold, it is considered that the candidate is a molecular cloud clump; otherwise, the candidate is not a molecular cloud clump.
[0011] In the said Step 1, the ClumpFind algorithm is used to obtain candidates of molecular cloud clumps, including the following steps:
[0012] S1.1: Preprocess the size of the molecular cloud clump data:
[0013] Put the molecular cloud clump candidate into a cube with a volume of 30×30×30 pixels, and then extract the data without molecular cloud clumps from the measured data as background data and fill it into the cube area with a mask value of 0;
[0014] S1.2: Normalization processing of the intensity of the molecular cloud clump candidate:
[0015] After a single molecular cloud clump candidate is normalized, the intensity value of the molecular cloud clump data is linearly mapped to the range of [0,1] through linear transformation. The normalization formula is as follows:
[0016]
[0017] Among them, x represents the intensity of the three-dimensional molecular cloud clump data, x max represents the maximum value of the intensity, and x min represents the minimum value of the intensity.
[0018] The step 2 includes the following steps:
[0019] S2.1: Use the 3D convolutional neural network CNN in the SS-3D-Clump model to extract the features of the molecular cloud clump candidate;
[0020] S2.2: Based on the features extracted in S2.1, use the Constrained-KMeans algorithm to cluster the extracted features to obtain the pseudo-labels of the molecular cloud clumps;
[0021] S2.3: Use the classifier network in the SS-3D-Clump model to classify the molecular clump candidates to generate the model classification labels;
[0022] S2.4: Calculate the loss using the pseudo-labels and classification labels obtained in step S2.2 and step S2.3;
[0023] The binary cross-entropy is used to calculate the loss. The loss calculation formula is as follows:
[0024]
[0025] Among them, N is the total number of samples; represents the pseudo-label of the i-th sample, taking values of 0 or 1; i = 1, 2, 3,..., N; is the probability of the classifier output label.
[0026] S2.5: The SS-3D-Clump model measures the convergence degree of the model by monitoring the difference in the recognition results of the model for the molecular cloud clump candidate dataset in two adjacent rounds during the training process;
[0027] The Normalized Mutual Information (NMI) is used to measure the information shared between two different assignments A and B of the same data, and is defined as:
[0028]
[0029] Among them, MI(A; B) is the mutual information between variables A and B; H(A) represents the entropy of A; H(B) represents the entropy of B; when NMI(A; B) is close to 1, it means that A and B have no difference, indicating that the recognition results of the model are consistent after two rounds of training before and after, indicating that the model converges.
[0030] S2.6: When the model converges or reaches the termination condition, the training is completed.
[0031] In the above 2.2, the principle of the Constrained-KMeans algorithm is as follows:
[0032] Given a sample set X = {x1, x2,..., x m}, where x i represents the i-th sample, and this sample can be described by a vector or a matrix.
[0033] Assume that the set of a small number of labeled samples is where S j is a non-empty sample set belonging to the j-th clustering cluster. k represents the number of clustering clusters, represents the union of S1, S2,..., S k , that is, the set S contains samples of k categories.
[0034] Directly use the set S as the "seed set" to initialize the k clustering centers of the KMeans algorithm, and do not change the cluster membership of the seed samples during the iterative update process of the clustering clusters. In this way, the Constrained-KMeans with constraints is obtained.
[0035] In the above step 3, after the SS-3D-Clump model training converges, for the molecular cloud clump candidates, according to the data preprocessing method in S1.1 in step 1, the preprocessed data is obtained; the preprocessed data is input into the SS-3D-Clump model, the SS-3D-Clump model extracts features based on the input data, the classifier classifies based on the features, and finally outputs the probability of belonging to the molecular cloud clump; the user determines the probability threshold according to the work requirements. When the output probability value exceeds the user's threshold, it is considered that this candidate is a molecular cloud clump; otherwise, this candidate is not a molecular cloud clump.
[0036] A method for identifying molecular cloud clumps based on semi-supervised deep learning according to the present invention has the following technical effects:
[0037] 1: The SS-3D-Clump model has achieved high accuracy in the identification of molecular cloud clumps, and its main advantages are as follows:
[0038] (1) After being trained on datasets constructed in three different density regions, the SS-3D-Clump model showed an accuracy of 0.933, a recall rate of 0.955, a precision rate of 0.945, and an F1 of 0.950 on the corresponding test datasets.
[0039] (2) The SS-3D-Clump model effectively captures the key features of molecular cloud clumps through 3D convolutional neural networks, such as intensity, rotation angle, and background noise.
[0040] (3) The SS-3D-Clump model exhibits strong generalization ability, can adapt to new unlabeled samples, and always maintains high accuracy on molecular cloud clump data in different regions.
[0041] (4) The SS-3D-Clump model can be integrated with existing molecular clump detection algorithms to form a framework for automatic detection and identification of molecular cloud clumps.
[0042] 2: The present invention designs and develops an automated molecular cloud clump identification algorithm that can replace manual inspection. When the accuracy of automated molecular cloud clump identification is relatively high, the detection algorithm for the front-end molecular cloud clumps can complete the detection under the initial parameter settings to obtain candidates for molecular cloud clumps. Since no manual parameter tuning has been carried out, some false clumps are inevitably introduced into the candidates, and then an automated authentication process is used to eliminate the incorrect targets. In this way, the detection efficiency of molecular cloud clumps can be greatly improved, and the correctness of the detection results can be ensured.
[0043] 3: The method of combining semi-supervised clustering algorithms and deep features, as a type of semi-supervised deep learning, can make full use of limited labeled data and a large amount of unlabeled data to improve the effect of molecular cloud clump identification. This method can reduce the burden of manually labeling data while improving the accuracy and scalability of molecular cloud clump identification. The main objective of the present invention is to develop an automated identification method for molecular cloud clump candidates, which uses 3D convolutional neural networks to extract the features of molecular cloud clumps and adopts semi-supervised learning to train the model, aiming to improve its generalization ability and data utilization rate. It solves the problem of insufficient labeled samples in traditional supervised learning and enables the model to better adapt to new, unlabeled samples, achieving high accuracy in the verification of molecular cloud clumps. Moreover, the model can be integrated with any detection algorithm to construct a framework for automatic detection and identification of molecular cloud clumps. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 Flow chart for identifying molecular cloud clumps with the SS-3D-Clump model;
[0045] Figure 2 Schematic diagram of the structure of the 3D CNN and classifier in Step 1;
[0046] Figure 3 Example of a true molecular cloud clump verified manually;
[0047] Figure 4 Example of a false molecular cloud clump verified manually;
[0048] Figure 5 Feature map extracted from the SS-3D-Clump model of positive samples of molecular cloud clumps;
[0049] Figure 6 Principle diagram of the Constrained-KMeans algorithm clustering illustrated by a two-dimensional scatter plot;
[0050] Figure 7 Recognition results and probability distribution of candidate molecular clumps in the high-density region HDR test dataset by the SS-3D-Clump model;
[0051] Figure 8 Performance graph of the SS-3D-Clump model during the training process in the high-density region HDR;
[0052] Figure 9 Recognition results of the SS-3D-Clump model for candidate molecular cloud clumps. Detailed implementation
[0053] Molecular cloud clump verification method based on semi-supervised deep learning (Semi-Supervised Deep Learning for Molecular Clump Verification, SS-3D-Clump).
[0054] Aiming at the characteristics of the molecular cloud clump data being a three-dimensional voxel data type of position-position-velocity (PPV), this method first uses a 3D convolutional neural network to extract the features of candidate molecular cloud clumps, and adopts a semi-supervised learning method to train the model to improve the generalization ability of the model and the utilization rate of data. It solves the problem of insufficient labeled samples in traditional supervised learning and enables the model to better adapt to new, unlabeled samples, achieving a relatively high accuracy in the verification of candidate molecular cloud clumps.
[0055] The technical solution of the present invention first uses the ClumpFind algorithm to obtain candidate molecular cloud clumps, but other molecular cloud clump detection algorithms can also be used. Then, the feature extraction part of SS-3D-Clump is used to extract the depth features of the candidates. The Constrained-KMeans algorithm uses the depth features as the basis and a small number of artificially labeled samples as seeds to generate pseudo-labels for the candidates. In addition, the SS-3D-Clump classifier assigns prediction labels to the candidates according to their depth features. The difference between the two labels is used to optimize the parameters of the SS-3D-Clump model. Generally speaking, SS-3D-Clump iteratively clusters the depth features through the Constrained-KMeans algorithm and uses the subsequent class labels as supervision to update the entire network weights of SS-3D-Clump. Finally, the SS-3D-Clump model is combined with existing molecular cloud clump detection algorithms to form a framework for automatically detecting and identifying molecular cloud clumps.
[0056] The model identification of molecular cloud clumps includes the following steps:
[0057] Step A: Use the ClumpFind algorithm to obtain candidate molecular cloud clumps;
[0058] (1.1): Processing the size of candidate molecular cloud clumps:
[0059] When using the ClumpFind algorithm to detect molecular cloud clumps, it returns a mask to mark the area of the molecular cloud clumps. Each molecular cloud clump is identified by a unique integer, such as 1, 2, etc.). The data of each molecular cloud clump is extracted using the mask information provided by the algorithm. Since the volume sizes of each candidate molecular cloud clump are different, and the input data size of the SS-3D-Clump model is fixed, it is necessary to preprocess the size of the molecular cloud clump data. The candidate molecular cloud clumps are placed in a cube with a volume of 30×30×30 pixels. Then, the data without molecular cloud clumps is extracted from the measured data as background data and filled into the cube area with a mask value of 0.
[0060] (1.2): Normalization processing of the intensity of candidate molecular cloud clumps:
[0061] After the size of the candidate molecular cloud clumps is uniformly processed, there are intensity differences among different molecular cloud clumps, which will have an adverse impact on the stability and accuracy of the SS-3D-Clump model. To eliminate this difference, it is necessary to perform normalization processing on the data. After a single candidate molecular cloud clump is normalized, the intensity value of the molecular cloud clump data is linearly mapped to the range of [0,1]. The normalization formula is as follows:
[0062]
[0063] Among them, x represents the intensity of the three-dimensional molecular cloud clump data, x max represents the maximum value of the intensity, and x min represents the minimum value of the intensity. Normalization makes it easier for the SS-3D-Clump model to learn useful features in the molecular cloud clump data and enables the network to converge quickly.
[0064] Step B: Train the SS-3D-Clump identification model based on the molecular cloud clump candidates. The preprocessed molecular cloud clump candidates are sent into SS-3D-Clump in batches for training. Since this model is trained based on a semi-supervised learning method, when the model training is completed, the identification work is also completed. When the model training converges, the probability value that the molecular cloud clump candidate belongs to the molecular cloud clump is output. According to the output probability value, it can be determined whether the candidate is a real molecular cloud clump. In step B, the SS-3D-Clump identification model is trained based on the molecular cloud clump candidates. It is specifically divided into the following steps:
[0065] (2.1): Use the 3D convolutional neural network (CNN) in the SS-3D-Clump model to extract the features of the molecular clump candidates;
[0066] (2.2): Based on the extracted features, use the Constrained-KMeans algorithm to cluster the extracted features to obtain the pseudo-labels of the molecular cloud clumps; The principle of the Constrained-KMeans algorithm is as follows:
[0067] Given the sample set X = {x1, x2,..., x m} and assuming that a small number of labeled samples are where is the sample belonging to the j-th clustering cluster. Directly use S as the "seed" to initialize the k clustering centers of the KMeans algorithm, and do not change the cluster membership of the seed samples during the iterative update process of the clustering clusters. In this way, the Constrained-KMeans with constraints is obtained;
[0068] (2.3): Use the classifier network in the SS-3D-Clump model to classify the molecular clump candidates and generate the model classification labels;
[0069] (2.4): Calculate the loss using the pseudo-labels and classification labels obtained in steps (2.2) and (2.3), and use binary cross-entropy to calculate the loss. The loss calculation formula is as follows:
[0070]
[0071] where N is the total number of samples, denotes the pseudo-label of the i-th sample, taking values 0 or 1, and is the probability of the classifier output label.
[0072] (2.5): In the way of backpropagation, using the loss obtained in step (2.4), optimize the parameters of the SS-3D-Clump model; so that the SS-3D-Clump model can correctly identify the molecular cloud clump candidates. Since the SS-3D-Clump model is semi-supervised learning and cannot provide real labeled data to monitor the training quality of the model, the SS-3D-Clump model measures the convergence degree of the model by monitoring the difference in the identification results of the molecular cloud clump candidate data set by the model in two adjacent rounds during the training process. Use the Normalized Mutual Information (NMI) to measure the information shared between two different assignments A and B of the same data, defined as:
[0073]
[0074] where MI is the mutual information and H is the entropy. When NMI is close to 1, it means there is no difference between the two, indicating that the identification results of the model in the previous and subsequent two rounds of training are consistent, showing that the model converges.
[0075] (2.6): When the model converges or reaches the termination condition, complete the training;
[0076] Step C: Identification of molecular cloud clump candidates:
[0077] After the SS-3D-Clump model training converges, for the molecular cloud clump candidates, according to the data preprocessing method in step A, obtain the preprocessed data. Input this data into the SS-3D-Clump model. The SS-3D-Clump extracts features based on the input data, the classifier classifies based on the features, and finally outputs the probability of belonging to the molecular cloud clump. The user determines the probability threshold according to the work requirements. When the output probability value exceeds the user's threshold, it is considered that the candidate is a molecular cloud clump; otherwise, the candidate is not a molecular cloud clump.
[0078] Example:
[0079] As Figure 1 shown, the model identification of molecular cloud clumps includes the following steps:
[0080] Step S1: Use the ClumpFind algorithm to detect molecular clump candidates;
[0081] Step S2: Train the SS-3D-Clump recognition model based on the molecular cloud clump candidates. Since this model is trained in a semi-supervised learning manner, when the model training is completed, the recognition work is also completed.
[0082] In step S2, the structures of the 3D CNN and the classifier are as Figure 2 shown, consisting of 5 3D CNNs with 16, 32, 32, 32 filters and 4 fully connected layers. The ReLU activation function is applied between each layer, and Dropout is used during the training process. Training the SS-3D-Clump recognition model based on the molecular cloud clump candidates is specifically divided into the following steps:
[0083] Step S2.1: Use the 3D CNN in the SS-3D-Clump model to extract the features of the molecular clump candidates;
[0084] Step S2.2: Based on the extracted features, use the Constrained-KMeans algorithm to cluster the extracted features to obtain pseudo-labels.
[0085] Manually recognize the molecular cloud clump candidates obtained by the ClumpFind algorithm to obtain the labels of the candidates, and use these labeled candidates as the seed samples. Use the seed sample set to constrain the KMeans algorithm. The principle of the Constrained-KMeans algorithm is as follows:
[0086] Given the sample set X = {x1, x2, …, x m}, assume that a small number of labeled samples are where is the sample belonging to the j-th clustering cluster. Directly use S as the "seed" to initialize the k clustering centers of the KMeans algorithm, and do not change the cluster membership of the seed samples during the iterative update process of the clustering clusters. In this way, the Constrained-KMeans with constraints is obtained;
[0087] Step S2.3: Use the classifier network in the SS-3D-Clump model to classify the molecular clump candidates and generate classification labels;
[0088] Step S2.4: Calculate the loss using the pseudo-labels and classification labels obtained in steps S2.2 and S2.3;
[0089] Step S2.5: Adopt the backpropagation method and use the loss obtained in step S2.4 to optimize the parameters of the SS-3D-Clump model;
[0090] Step S2.6: When the model converges or reaches the termination condition, complete the training;
[0091] Step S3: Identification of molecular cloud clump candidates.
[0092] Semi-supervised learning based on the features of the deep network is carried out as follows:
[0093] Step1: Use the ClumpFind algorithm to detect molecular cloud clump candidates, Figure 3 and Figure 4 respectively show examples of true and false molecular cloud clumps identified manually. Figure 3 and Figure 4 The two rows of subgraphs in show the integrated intensity maps of two clumps respectively. For each row of subgraphs, from left to right are the integrated maps of the molecular cloud clump candidates in the x-y, x-v, and y-v planes.
[0094] Step2: Train the SS-3D-Clump model based on the molecular cloud clump candidates. Since this model is trained in a semi-supervised learning manner, when the model training is completed, the identification work is also completed;
[0095] In Step2, training the SS-3D-Clump identification model based on the molecular cloud clump candidates is specifically divided into the following steps:
[0096] Step2.1: Use the 3D CNN in the SS-3D-Clump model to extract the features of the molecular clump candidates. To qualitatively evaluate the feature extraction performance of the SS-3D-Clump model, we plotted the intermediate feature extraction results of the SS-3D-Clump model, using manually verified positive and negative molecular cloud clump samples as inputs. Figure 5 The intermediate results of the SS-3D-Clump model when the input is a positive molecular cloud clump sample. The left panel shows the integrated map of the molecular clump in the x-y plane. The middle panel consists of four subgraphs (arranged in a 2x2 grid), representing 4 of the 16 results obtained after the Stem operation in the SS-3D-Clump model. The right panel shows 32 feature results obtained after the ResB1 operation;
[0097] Step2.2: Use the Constrained-KMeans algorithm to cluster the extracted features based on the extracted features to obtain pseudo-labels. To illustrate the working principle of the ConstrainedKMeans algorithm, we simulated five clusters with a two-dimensional Gaussian distribution as labeled seed data, as Figure 6 shown, and then randomly generated points within the coverage of these five clusters. Then use the Constrained-KMeans algorithm to cluster the random points;
[0098] Step2.3: Use the classifier network in the SS-3D-Clump model to classify the molecular clump candidates and generate classification labels. Figure 7 Figure 7 shows the identification results and probability distributions of the candidate molecules in the HDR test dataset by the SS-3D-Clump model. The x-axis represents the probability assigned by the SS-3D-Clump model to the molecular cloud clump candidates belonging to the molecular cloud clumps. The blue histogram represents the probability distribution of all candidate molecules in the test dataset. The green histogram represents the probability distribution of the candidate samples correctly identified as negative samples by the SS-3D-Clump model, and the orange represents the probability distribution of the candidate samples correctly identified as positive samples;
[0099] Step2.4: Calculate the loss using the pseudo-labels and classification labels obtained in Step2.2 and Step2.3;
[0100] Step2.5: In a backpropagation manner, use the loss obtained in Step2.4 to optimize the parameters of the SS-3D-Clump model, and use stochastic gradient descent for optimization;
[0101] Step2.6: When the model converges or reaches the termination condition, the training is completed. The convergence condition is that the label change rate of the model in two consecutive times is very low. The mutual information between the two consecutive labels is used to quantitatively measure whether the model converges. Figure 8 Figure 8 shows the performance of the SS-3D-Clump model during the HDR training process in the high-density region. Left figure: The evolution of the clustering ability of the SS-3D-Clump model in different training epochs. Middle figure: The clustering label reallocation rate of each clustering iteration of the SS-3D-Clump model. Right figure: The overall performance of the SS-3D-Clump model in the test dataset;
[0102] Step3: Identification of molecular cloud clump candidates:
[0103] Figure 9 Figure 9 shows the identification results of the SS-3D-Clump model for the candidate molecular cloud clumps. The first two rows show the results of four molecular cloud clumps, and the last row shows the results of two non-molecular cloud clumps. The first row of each subfigure represents the maximum intensity, total flux, and total number of pixels of the respective molecular cloud clump. The second row marks the centroid position of the molecular cloud clump. For example, the centroid position of the upper-left molecular cloud clump is: l = 11.186°, b = -1.171°, v = 38.893 km / s. The third row is the identification result of the SS-3D-Clump model for each candidate. The second number represents the probability that the model believes the candidate belongs to the molecular cloud clump: for example, the probability that the candidate shown in the upper-left corner belongs to the molecular clump is 100%.
Claims
1. A method for identifying molecular cloud clumps based on semi-supervised deep learning, characterized in that It includes the following steps: Step 1: Obtain candidate molecular cloud clumps; Step 2: Train the SS-3D-Clump model based on the candidate molecular cloud clumps, and output the probability value of the molecular cloud clumps; Step 3: Determine the probability threshold. When the output probability value exceeds the probability threshold, it is considered that the candidate is a molecular cloud clump; otherwise, the candidate is not a molecular cloud clump; The said Step 2 includes the following steps: S2.1: Use the 3D convolutional neural network CNN in the SS-3D-Clump model to extract the features of the candidate molecular cloud clumps; S2.2: Based on the features extracted in S2.1, use the Constrained-KMeans algorithm to cluster the extracted features to obtain the pseudo-labels of the molecular cloud clumps; S2.3: Use the classifier network in the SS-3D-Clump model to classify the candidate molecular cloud clumps to generate the model classification labels; S2.4: Calculate the loss using the pseudo-labels and classification labels obtained in Step S2.2 and Step S2.3; The binary cross-entropy is used to calculate the loss, and the loss calculation formula is as follows: where N is the total number of samples; represents the pseudo-label of the i-th sample, taking values 0 or 1; i = 1, 2, 3, …, N; is the probability of the output label of the classifier; S2.5: The SS-3D-Clump model measures the convergence degree of the model by monitoring the difference in the recognition results of the candidate molecular cloud clump dataset by the model in two adjacent rounds during the training process; The Normalized Mutual Information (NMI) is used to measure the information shared between two different assignments A and B of the same data, and is defined as: where MI(A; B) is the mutual information between variables A and B; H(A) represents the entropy of A; H(B) represents the entropy of B; when NMI(A; B) is close to 1, it means that A and B have no difference, indicating that the recognition results of the model after two rounds of training are consistent, indicating that the model converges; S2.6: When the model converges or reaches the termination condition, the training is completed.
2. The method for identifying molecular cloud clumps based on semi-supervised deep learning according to claim 1, wherein: In the said Step 1, the ClumpFind algorithm is used to obtain the candidate molecular cloud clumps, including the following steps: S1.1: Preprocess the size of the molecular cloud clump data: Place the candidate molecular cloud clumps in a cube with a volume of 30×30×30 pixels, and then extract the data without molecular cloud clumps from the measured data as background data and fill it into the cube area with a mask value of 0; S1.2: Normalize the intensity of the candidate molecular cloud clumps: After a single candidate molecular cloud clump is normalized, the intensity value of the molecular cloud clump data is linearly mapped to the range of [0,1], and the normalization formula is as follows: Among them, x represents the intensity of the three-dimensional molecular cloud clump data, and x max represents the maximum value of the intensity, and x min represents the minimum value of the intensity.
3. The method for identifying molecular cloud clumps based on semi-supervised deep learning according to claim 1, wherein: In the said 2.2, the Constrained-KMeans algorithm is as follows: Given a sample set X = {x1, x2, …, x m}, let the set of a small number of labeled samples be where S j is a non-empty sample set belonging to the j-th clustering cluster; k represents the number of clustering clusters, represents the union of S1, S2, …, S k ; Directly use the set S as the "seed set" to initialize the k clustering centers of the KMeans algorithm, and do not change the cluster membership of the seed samples during the iterative update process of the clustering clusters; in this way, the Constrained-KMeans with constraints is obtained.
4. The method for identifying molecular cloud clumps based on semi-supervised deep learning according to claim 2, characterized in that: In step 3, after the SS-3D-Clump model training converges, for the molecular cloud clump candidates, preprocessed data is obtained according to the data preprocessing method in S1.1 of step 1. The preprocessed data is input into the SS-3D-Clump model. The SS-3D-Clump model extracts features based on the input data, and the classifier classifies based on the features, and finally outputs the probability of belonging to a molecular cloud clump. The user determines the probability threshold according to the work requirements. When the output probability value exceeds the user's threshold, the candidate is considered to be a molecular cloud clump; otherwise, the candidate is not a molecular cloud clump.