Marker and method for identifying central precocious girls in vitro
By using the free miRNA markers has-miR-584-5p, hsa-miR-625-3p and has-miR-7-5p in the serum, the inadaptation and diagnosis difficulties of patients in the diagnosis of central precocious puberty in the prior art were solved, and high accuracy and low cost central precocious puberty detection was achieved.
Patent Information
- Application Number
- CN202510519683.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-31
- Publication Date
- 2025-07-11
AI Technical Summary
Existing central precocious puberty diagnosis methods such as GnRH induced trials have problems with patient discomfort and multiple venous blood withdrawals, and cannot effectively distinguish central precocious puberty from healthy individuals, limiting their application in outpatient clinics.
The free miRNA markers has-miR-584-5p, hsa-miR-625-3p and has-miR-7-5p in the serum were used as detection markers. MiRNAs with significant differential expression were screened through second-generation sequencing and data analysis, and in vitro identification model was constructed. The subset of features was screened by recursive feature elimination cross-validation method was used to screen the ROC curve, and AUC values were drawn to achieve accurate central precocious diagnosis.
It provides a peripheral blood detection method, simplifies the material extraction process, improves the accuracy of diagnosis and patient acceptance, and is within an acceptable cost, suitable for ordinary skilled personnel to operate, and can efficiently distinguish central precocious puberty from healthy individuals.
Smart Images

Figure CN120290710A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and more particularly, to a method for detecting biomarkers for identifying central precocious puberty in girls in vitro and a method therefor. Background Art
[0002] Central precocious puberty (Central precocious puberty) is also known as true precocious puberty and complete precocious puberty, and is the result of the premature activation and maturation of the hypothalamic-pituitary-gonadal (HPG) axis. Currently, the incidence of precocious puberty in children shows an increasing trend year by year and the age of onset is gradually getting younger. At present, it has become the second largest endocrine disease in children after obesity. The so-called central precocious puberty refers to girls clinically diagnosed by GnRH stimulation test. The clinical confirmation standard is based on the "Consensus on the Diagnosis and Treatment of Central Precocious Puberty 2015". The greatest harm to children is the premature closure of the epiphysis, which leads to short stature, and also significantly increases the risk of breast cancer, endometrial cancer, obesity, type 2 diabetes and cardiovascular diseases in adulthood.
[0003] At present, neither growth rate, height, body mass index, etc., nor bone age examination, hypothalamic pituitary magnetic resonance examination, uterine and ovarian ultrasound examination, endocrine hormone examination, etc. can effectively diagnose central precocious puberty. The clinically recognized diagnostic standard for central precocious puberty is the gonadotropin-releasing hormone (GnRH) stimulation test. The GnRH stimulation test is to inject gonadorelin intravenously, and venous blood is drawn before injection and at 30 minutes, 60 minutes, and 90 minutes after injection to measure the stimulating concentrations of luteinizing hormone (LH) and follicle-stimulating hormone (FSH). If the LH peak value > 5.0 IU / L and LH / FSH > 0.6 after gonadorelin injection, central precocious puberty can be diagnosed. Some patients will have systemic and local allergic reactions after injecting gonadorelin. Therefore, the GnRH stimulation test is not suitable for outpatient clinics and requires inpatient examination. Moreover, since this examination requires repeated venous blood sampling, patients often have a resistant attitude, so the application of the GnRH stimulation test in the differential diagnosis of central precocious puberty is limited.
[0004] MicroRNA (miRNA) is a class of non-coding short RNAs approximately 19-25 nucleotides in length. It can degrade the target gene mRNA or inhibit its translation by completely or incompletely pairing with the 3'UTR of the target gene mRNA. Past studies have shown that miRNA is involved in a variety of regulatory pathways, including development, viral defense, hematopoiesis, organ formation, cell proliferation and death, and so on. In recent years, a large number of studies have shown a close relationship between the abundance changes of miRNA and the occurrence and development of various diseases. Among them, circulating miRNA can exhibit significantly different expression profiles according to different physiological and pathological states of individuals, so it can be used to distinguish normal and disease states. At present, circulating miRNA has been involved in the auxiliary differential diagnosis of various aspects such as cancer, neurodegenerative diseases, cardiovascular and cerebrovascular diseases, aging, etc., but it has not been developed in central precocious puberty.
[0005] Therefore, it is necessary to develop in vitro detection markers, corresponding detection methods and reagents with clinical application value for central precocious puberty to facilitate the rapid and convenient detection of central precocious puberty population and early clinical intervention. Summary of the Invention
[0006] The object of the present invention is to provide a method for detecting markers for in vitro identification of central precocious puberty in girls, which is patient-friendly and simple in sampling compared with the GnRH stimulation test, and has the advantage of high accuracy compared with other clinical detection methods.
[0007] In the first aspect of the present application, a detection marker for in vitro identification of central precocious puberty in girls is provided. The marker is a free miRNA marker from human serum, and the miRNA marker includes: one or more of has-miR-584-5p, hsa-miR-625-3p, and has-miR-7-5p.
[0008] Further, the miRNA marker is a mature miRNA in serum.
[0009] Further, the expression level of the miRNA marker in blood is a relative expression level, and there are significant statistical differences between girls with central precocious puberty and healthy girls. The miRNA marker can distinguish girls with central precocious puberty from healthy girls, and can also distinguish girls with central precocious puberty from girls with simple premature thelarche.
[0010] Furthermore, a training group was composed of free miRNA samples from the sera of girls with central precocious puberty and normal healthy girls. The samples were subjected to next-generation sequencing and data analysis. Based on the miRNAs with significantly different expressions between the sera of girls with central precocious puberty and normal healthy girls, the recursive feature elimination cross-validation (RFECV) method was used to screen out miRNAs with statistical significance as markers.
[0011] The present invention also provides a method for in vitro identification of detection markers for girls with central precocious puberty, which is characterized by including the following steps:
[0012] (a) A training group was composed of free miRNA samples from the sera of girls with central precocious puberty and normal healthy girls. After the samples were respectively subjected to library preparation, next-generation sequencing and data analysis, by comparing with the positions of miRNAs in the human reference genome, the expression level RPM of each miRNA in the sample was determined.
[0013] (b) Using the expression level RPM of miRNAs as independent variables, random forest was used for feature screening to select the feature subset with the most predictive ability for the performance of the prediction model.
[0014] (c) The obtained feature subsets were permuted and combined, and the ROC curve was plotted and the AUC value was calculated one by one. The combination with the highest AUC was taken as the marker.
[0015] Further, in step (a), to obtain the RPM of the expression level of each miRNA in the sample, the following steps are specifically included:
[0016] (a1) After the sample was subjected to library preparation and next-generation sequencing, the off-machine data was obtained, and the off-machine data was subjected to data quality control and preprocessing by a quality control tool to obtain valid data with low-quality sequences and sequencing adapters removed.
[0017] (a2) The sequences of the valid data were aligned with the human reference genome sequence to obtain the miRNA position information located in the human reference genome sequence. Among them, the miRNA position information was taken from the miRBase database. When the 5' end of a certain sequence was consistent with the 5' end of a certain miRNA, this sequence was recorded as the sequencing sequence of this miRNA.
[0018] (a3) The expression level RPM of each miRNA in the test sample was determined. Among them, RPM (reads per million) is the unit of the expression level. The expression level RPM of a certain miRNA is the millionth percentage of the total amount of the sequencing sequence of this miRNA in the total amount of all sequencing sequences that can be aligned to the human reference genome of the sample.
[0019] Further, in step (b), select the feature subset that has the most predictive ability for the performance of the prediction model, specifically including the following steps:
[0020] (b1) Perform Min-Max normalization on the RPM values of the expression levels of each miRNA. First, calculate the minimum value X min and the maximum value X max of each feature. For each feature j and sample i, use formula S1 to calculate the value X scaled after data scaling. The formula S1 is:
[0021]
[0022] (b2) Label the two types of feature subsets, that is, {"central precocious puberty": 0, "NC": 1}, and train on the serum samples of central precocious puberty and normal healthy samples;
[0023] (b3) Use the RFECV algorithm to gradually converge and retain the feature subset with higher accuracy. Through such cyclic iteration, it finally converges to several miRNAs. The several miRNAs that finally converge need to meet the following characteristics: (1) The miRNA combination shows the highest eigenvalue; (2) The number of miRNAs can be within the acceptable range of the computer when doing permutations and combinations.
[0024] In some embodiments, control the convergence speed by modifying the min_features_to_select parameter to obtain the effects of different numbers of features, select the feature subset with the highest accuracy, and thus determine the number of feature variables.
[0025] Further, in step (c), obtain the highest AUC combination, specifically including the following steps:
[0026] (c1) Perform permutations and combinations on the feature subset obtained in b3, and draw the ROC curve and calculate the AUC value for each combination one by one;
[0027] (c2) Take the miRNA combination with the maximum AUC as the biomarker.
[0028] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0029] (1) Peripheral blood samples are easier to obtain, have strong clinical operability and very little trauma, which is conducive to the subjects undergoing such tests and has a broad application prospect. Moreover, the serum miRNA has good stability and relatively large content, and the difficulty of extraction, library construction and sequencing is relatively low. What is required are all conventional experimental techniques and reagents and drugs that are easily purchased;
[0030] (2) Compared with the existing means for detecting central precocious puberty, the miRNA markers protected by the present invention have higher detection ability for the identification of central precocious puberty, and the experimental cost based on next-generation sequencing is also within an acceptable range;
[0031] (3) By using the miRNA markers of the present invention and the expression level RPM of the miRNA markers in the test sample, and adopting a simple formula for calculation, it can be determined whether the individual of the test sample suffers from central precocious puberty. The data analysis method is not complicated either, so it can be quickly mastered by ordinary technical personnel. Brief Description of the Drawings
[0032] When reading in combination with the following attached Figure 1 drawings, the above and other features of the content of the present application will be more fully described. It can be understood that these drawings only depict several embodiments of the content of the present application, and therefore should not be considered as limiting the scope of the content of the present application. By using the attached drawings, the content of the present application will be more clearly and detailedly described.
[0033] Figure 1 ROC curve graphs for three miRNA combinations of has-miR-584-5p, hsa-miR-625-3p and has-miR-7-5p arranged according to the AUC from high to low, with the highest AUC value of 0.9885; four miRNA combinations of hsa-miR-625-3p, has-miR-874-3p, hsa-miR-181c-3p and hsa-miR-1290 with the second highest AUC value of 0.9722; and three miRNA combinations of has-miR-584-5p, has-miR-652-3p and hsa-miR-181c-3p with an AUC value of 0.9685.
[0034] Figure 2 ROC curve graph for the validation cohort.
[0035] Figure 3 Bar graph showing the effectiveness of miRNA samples in the validation cohort. Detailed Description of the Embodiments
[0036] The following embodiments are described to assist in understanding the present application. The embodiments are not and should not in any way be construed as limiting the scope of protection of the present application.
[0037] For the experimental methods without specific conditions in the following examples, they were carried out according to conventional experimental conditions, such as those described in Molecular Cloning: A Laboratory Manual by Sambrook et al. (New York: Cold Spring Harbor Laboratory Press, 1989), or according to the conditions recommended by the manufacturer. Unless otherwise stated, percentages and parts are calculated by weight. Unless otherwise specified, the materials used in the examples are commercially available products. Example 1: Obtaining the training set samples
[0038] The applicant collected 28 peripheral venous blood samples from girls with central precocious puberty from July 2022 to February 2023. Each sample contained 10 mL of peripheral blood, with an average age of 8.6 years and an age distribution of 8.1 - 9.1 years. At the same time, the applicant collected 23 peripheral venous blood samples from normal healthy girls (i.e., healthy controls without various diseases, the same below). Each sample contained 10 mL of peripheral blood, with an average age of 8.2 years and an age distribution of 7.8 - 8.6. These two groups of samples were used as the training set samples. There was no statistically significant difference in age between these two groups of samples, and the gender was all female, thus meeting the principle of gender and age matching. For each peripheral blood sample, sequencing library preparation and second-generation sequencing were performed to obtain the off-machine data. Example 2: Obtaining the miRNA expression profiles of samples through sequencing libraries
[0039] For each training set sample, the following reagents and steps were used for library preparation and second-generation sequencing:
[0040] (1) After collecting 10 mL of peripheral blood samples with dry collection tubes, they were left to stand at 4°C for more than half an hour. Subsequently, 400 g of free RNA was obtained, and the supernatant was taken after centrifugation at 400 g for 10 minutes at 4°C. Further, the supernatant was taken after centrifugation at 1800 g for 10 minutes at 4°C to obtain the serum sample, which was stored in a -80°C refrigerator.
[0041] (2) 50 - 200 ng of serum-free RNA was extracted from the above serum samples using Qiagen miRNeasy Serum / Plasma Kit (product number: 217184), diluted to a total volume of 4 μL with ultrapure water (free of DNase and RNase, the same below), and placed in a 200 μL thin-walled PCR tube.
[0042] (3) 1 μL of adapter RA3 with a concentration of 10 μM was added to the solution obtained in step (2), mixed well, and reacted at 70°C for 2 minutes, then immediately cooled on ice. The sequence of RA3 is 5’-TGGAATTCTCGGGTGCCAAGG-3’.
[0043] (4) Add 2 μL of HML (Ligation Buffer) (Illumina, catalog number 15013206), 1 μL of RNase Inhibitor (Illumina, catalog number 15003548), and 1 μL of T4 RNA Ligase 2 Deletion Mutant (Epicentre, catalog number LR2D11310K) to the solution obtained in step (3), mix well, and incubate at 28 °C for 1 hour; (5) Add 1 μL of STP (Stop Solution) (Illumina, catalog number 15016304) to the solution obtained in step (4), mix well, and incubate at 28 °C for 15 minutes;
[0044] (6) Take a new PCR tube, add 1.1 μL of adapter RA5 with the base sequence 5’-GUUCAGAGUUCUACAGUCCGACGAUC-3’, the concentration of RA5 is 10 μM, incubate at 70 °C for 2 minutes, and immediately place on ice for cooling after the reaction;
[0045] (7) Add 1.1 μL of 10 mM ATP (Illumina, catalog number 15007432) to the solution obtained in step (6), and then add 1.1 μL of T4 RNA ligase (Illumina, catalog number 1000587) and mix well;
[0046] (8) Take 3 μL from the solution obtained in step (7) and add it to the solution obtained in step (5), mix well, and react at 28 °C for 1 hour;
[0047] (9) Add 1 μL of RNA RT Primer (10 μM) to the solution obtained in step (8), mix well, react at 70 °C for 2 minutes to perform reverse transcription reaction to obtain the first strand of DNA. The sequence of the reverse transcription primer RT Primer is 5’-CCTTGGCACCCGAGAATTCCA-3’, and immediately place on ice for cooling after the reaction;
[0048] (10) Add 2 μL of 5× First Strand Buffer (Thermo, catalog number 1889832), 0.5 μL of dNTP Mix (12.5 mM, Illumina, catalog number 11318102), 1 μL of 100 mM DTT (Thermo, catalog number 1850670), 1 μL of RNase Inhibitor, and 1 μL of SuperScript II Reverse Transcriptase (Thermo, catalog number 2008270) to the solution obtained in step (9), mix well, and incubate at 50 °C for 1 hour;
[0049] (11) Add 25 μL of PML (PCR Mix) (Illumina, catalog number 15022681), 2 μL of Primer1 (10 μM), and 2 μL of Primer2 (10 μM) to the solution obtained in step (10). After mixing, perform a PCR reaction: pre-denature at 98 °C for 30 s, denature at 98 °C for 10 s, anneal at 60 °C for 30 s, extend at 72 °C for 15 s. After performing 18 cycles, extend at 72 °C for 10 min and store at 4 °C. Among them, the sequence of Primer1 is 5’-CAAGCAGAAGACGGCATACGAGATGTCGTGATGTGACTGGAGTTCCTTGGCACCCGAGAATTCCA-3’, the sequence of Primer2 is 5’-AATGATACGGCGACCACCGAGATCTACACGTTCAGAGTTCTACAGTCCGA-3’, and the 8 bases “GTCGTGAT” in Primer1 are the index sequence;
[0050] (12) Perform 6% polyacrylamide gel electrophoresis on the PCR product obtained in step (11) at a voltage of 120 V for 1 h, stain with Gelred dye at a concentration of one ten-thousandth for 5 minutes, then observe and photograph under an ultraviolet lamp. Cut and recover the band between 149 and 169. After detecting the fragment length range using an Agilent 2100 Bioanalyzer and quantifying the concentration using an Invitrogen Qubit, send it to the Illumina Novaseq 6000 sequencing platform for sequencing. The sequencing read length is 75 bp, and the sequencing mode is single-end sequencing to obtain the off-machine data.
[0051] For the off-machine data of the training group samples, perform data analysis using the following steps to obtain the expression level RPM of each miRNA in the samples:
[0052] (1) For the off-machine data of the samples, use FastQC, Cutadpat, and Trimmomatic for data quality control and preprocessing (using default parameters) to obtain valid data with low-quality sequences and sequencing adapters removed;
[0053] (2) Remove the base sequence in RA5 from the 5’ end of the sequences in the valid data, and then use the sequence alignment software Bowtie to align the obtained sequences to the human reference genome sequence (allowing a maximum of 1 base mismatch) to obtain the location information mapped to the human reference genome;
[0054] (3) Compare the positions of the obtained sequences with the miRNA positions in the human reference genome to determine the RPM (reads per million) of each miRNA expression in the sample. Information on the RPM values of 466 miRNAs was obtained. Among them, the miRNA position information was taken from the miRBase database (http: / / www.mirbase.org / ). When the 5'-end of a certain sequence is consistent with the 5'-end of a certain miRNA, this sequence is recorded as the sequencing sequence of this miRNA; the RPM (reads per million) of each miRNA expression is the percentage of the total amount of the sequencing sequence of this miRNA in the total amount of all sequencing sequences that can be aligned to the reference genome of the sample per million. Example 3: Obtaining miRNA Markers for the Identification of Central Precocious Puberty
[0055] Using the MinMaxScaler tool in the scikit-learn library, perform Min-Max normalization on the RPM values of the expression of each miRNA using formula S1. The central precocious puberty group is defined as 0, and the NC group is defined as 1. Use the RFECV algorithm to gradually converge and retain features. Here, control the convergence speed to 9 feature subsets. The code is as follows:
[0056] n_features = 466
[0057] while n_features > 9:
[0058] m = int(n_features / 2)
[0059] if int(n_features / 2) > 100:
[0060] m = int(n_features / 2)
[0061] else:
[0062] m = 9
[0063] rfecv = RFECV(estimator = rf, step = 1, cv = 5, scoring = accuracy, verbose = 0, min_features_to_select = m)
[0064] The accuracy rates during the continuous convergence process are as follows:
[0065] 399 miRNAs: 0.931034482758620793; 253 miRNAs: 0.7931034482758621; 174 miRNAs: 0.7931034482758621; 91 miRNAs: 0.8275862068965517; 57 miRNAs: 0.7931034482758621; 39 miRNAs: 0.8620689655172413; 36 miRNAs: 0.7931034482758621; 26 miRNAs: 0.9255172413793104; 18 miRNAs: 0.9310344827586207; 12 miRNAs: 0.896551724137931; 9 miRNAs: 0.9310344827586207.
[0066] It can be seen that the effect of retaining 9 features and the effect of retaining 18 features can achieve the same accuracy rate. However, the workload of permutation and combination of 9 features is less than that of 18 features and is within the acceptable range of the computer.
[0067] They are has-miR-584-5p, has-miR-625-3p, has-miR-652-3p, has-miR-7-5p, has-miR-874-3p, hsa-miR-181c-3p, hsa-miR-1290, hsa-miR-454-3p and hsa-let-7i-3p respectively. The 9 miRNAs obtained are subjected to permutation and combination, and a total of 511 combinations are obtained. The ROC curves are plotted one by one and the AUC values are calculated. Arranged according to the AUC from high to low, a combination of three miRNAs, has-miR-584-5p, hsa-miR-625-3p and has-miR-7-5p, has the highest AUC value of 0.9885. In addition, a combination of four miRNAs, hsa-miR-625-3p, has-miR-874-3p, hsa-miR-181c-3p and hsa-miR-1290, has the second highest AUC value of 0.9722; a combination of three miRNAs, has-miR-584-5p, has-miR-652-3p and hsa-miR-181c-3p, has an AUC value of 0.9685, as Figure 1 shown. Example 4: Verification of the correctness of miRNA markers for the identification of central precocious puberty
[0068] Using 148 subjects, including 92 girls with central precocious puberty (CPP) and 56 normal (NC) girls as a validation cohort, a total of 511 permutations of the ROC curve for each 9-miRNA panel were plotted, and the ROC curve was plotted one by one and the AUC value was calculated. As Figure 2 shown.
[0069] It was confirmed that the combination of three miRNAs, has-miR-584-5p, hsa-miR-625-3p, and has-miR-7-5p, had the highest AUC value. In addition, the combination of four miRNAs, hsa-miR-625-3p, has-miR-874-3p, hsa-miR-181c-3p, and hsa-miR-1290, had the second-highest AUC value; the combination of three miRNAs, has-miR-584-5p, has-miR-652-3p, and hsa-miR-181c-3p, had a lower AUC value.
[0070] And QPCR was used to quantify these three markers.
[0071] As Figure 3 shown, the figure shows the effectiveness of miRNA samples in the validation cohort. The bar chart shows the K values of miRNA samples (has-miR-584-5p, hsa-miR-625-3p, and has-miR-7-5p) to verify the effectiveness of differentiating CPP and NC.
[0072] The clinical efficacy of CPP / NC (sensitivity: 86.96%, specificity: 100%, accuracy: 91.89%) in the validation cohort was further determined.
[0073] Analysis showed that compared with the GnRH stimulation test used in clinical practice, the assay based on the identified miRNA panel could achieve an accuracy of approximately 91.89% in differentiating CPP and NC.
[0074] Therefore, the above miRNA markers can well distinguish patients with central precocious puberty from healthy people and can be used as an auxiliary diagnostic basis for the identification of central precocious puberty.
[0075] Although the present application has disclosed multiple aspects and embodiments, other aspects and embodiments will be obvious to those skilled in the art. Without departing from the concept of the present application, several modifications and improvements can still be made, which all fall within the protection scope of the present application. The multiple aspects and embodiments disclosed in the present application are only for illustrative purposes and are not intended to limit the present application. The actual protection scope of the present application is subject to the claims.
Claims
1. An in vitro detection marker for identifying central precocious puberty in girls, characterized in that, The biomarker is a free miRNA biomarker derived from human serum, and the miRNA biomarker includes one or more of has-miR-584-5p, hsa-miR-625-3p, and has-miR-7-5p.
2. The in vitro identification detection marker for central precocious puberty in girls according to claim 1, wherein The miRNA biomarker is a mature miRNA in serum.
3. The in vitro detection marker for identifying central precocious puberty girls as described in claim 1, wherein, The expression level of the miRNA biomarker in blood is a relative expression level, and there are significant statistical differences between girls with central precocious puberty and healthy girls. The miRNA biomarker can distinguish girls with central precocious puberty from healthy girls, and can also distinguish girls with central precocious puberty from girls with simple premature thelarche.
4. The detection marker for in vitro identification of central precocious puberty in girls according to claim 1, wherein The free miRNA samples of sera from girls with central precocious puberty and normal healthy girls are used to form a training group. The samples are subjected to next-generation sequencing and data analysis. Based on the miRNAs with significantly different expressions between the serum samples of girls with central precocious puberty and those of normal healthy girls, the recursive feature elimination cross-validation method is used to screen out miRNAs with statistical significance as biomarkers.
5. The method for in vitro identification of detection markers for girls with central precocious puberty according to any one of claims 1-4, characterized in that, It includes the following steps: (a) The free miRNA samples of sera from girls with central precocious puberty and normal healthy girls are used to form a training group. After the samples are respectively subjected to library preparation, next-generation sequencing, and data analysis, by comparing with the positions of miRNAs in the human reference genome, the expression level RPM of each miRNA in the samples is determined. (b) Using the expression level RPM of miRNAs as independent variables, random forest is used for feature screening to select the feature subset that has the most predictive ability for the performance of the prediction model. (c) The obtained feature subsets are arranged and combined, and the ROC curve is plotted and the AUC value is calculated one by one, and the combination with the highest AUC is taken as the biomarker.
6. The method for in vitro identification of detection markers for girls with central precocious puberty according to claim 5, characterized in that, In step (a), to obtain the RPM of the expression level of each miRNA in the sample, it specifically includes the following steps: (a1) After the samples are subjected to library preparation and next-generation sequencing, the off-machine data is obtained. The off-machine data is subjected to data quality control and preprocessing through a quality control tool to obtain valid data with low-quality sequences and sequencing adapters removed. (a2) The sequences of the valid data are aligned with the human reference genome sequence to obtain the miRNA position information located in the human reference genome sequence. Among them, the miRNA position information is taken from the miRBase database. When the 5'-end of a certain sequence is consistent with the 5'-end of a certain miRNA, this sequence is recorded as the sequencing sequence of this miRNA. (a3) Determine the expression level RPM of each miRNA in the test sample. Among them, RPM (reads per million) is the unit of the expression level. The expression level RPM of a certain miRNA is the millionth percentage of the total amount of the sequencing sequence of this miRNA in the total amount of all sequencing sequences that can be aligned to the human reference genome of this sample.
7. The method for in vitro identification of detection markers for girls with central precocious puberty according to claim 5, wherein In step (b), to select the feature subset that has the most predictive ability for the performance of the prediction model, it specifically includes the following steps: (b1)Perform Min-Max normalization on the RPM values of the expression levels of each miRNA. First, calculate the minimum value X of each feature min and the maximum value X max . For each feature j and sample i, use formula S1 to calculate the value X after data scaling scaled . The formula S1 is as follows: ; (b2) Label the two types of feature subsets, i.e., {"central precocious puberty": 0, "NC": 1}, and train on central precocious puberty and normal healthy serum samples; (b3) Use the RFECV algorithm to gradually converge, retain the feature subset with higher accuracy. Through such iterative cycles, finally converge to several miRNAs. The several miRNAs that finally converge need to meet the following characteristics: (1) The miRNA combination exhibits the highest eigenvalue; (2) The number of miRNAs can be within the acceptable range of the computer when doing permutations and combinations.
8. The method for in vitro identification of detection markers for girls with central precocious puberty according to claim 5, characterized in that, Control the convergence speed by modifying the min_features_to_select parameter, obtain the effects of different numbers of features, select the feature subset with the highest accuracy, and thus determine the number of feature variables.
9. The method for in vitro identification of detection markers for girls with central precocious puberty according to claim 5, characterized in that, In step (c), obtain the highest AUC combination, which specifically includes the following steps: (c1) Make permutations and combinations of the feature subsets obtained in b3, draw the ROC curve for each combination and calculate the AUC value; (c2) Take the miRNA combination with the maximum AUC as the biomarker.