Premature ovarian failure diagnosis marker based on metabonomics and artificial intelligence technology and application of premature ovarian failure diagnosis marker

Through the combination of metabolomics and artificial intelligence, high sensitivity and specific auxiliary diagnostic markers of premature ovarian failure are screened out, and an efficient diagnostic model is constructed, which solves the shortcomings in the diagnosis of premature ovarian failure in the existing technology and achieves efficient and accurate diagnostic effects.

CN120161149AInactive Publication Date: 2025-06-17ZHEJIANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510647247.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-17
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN120161149A_ABST
    Figure CN120161149A_ABST
Patent Text Reader

Abstract

The invention discloses a premature ovarian failure diagnostic marker based on metabonomics and artificial intelligence technology and application thereof, and relates to the field of clinical examination and diagnosis. A metabonomics technology and an artificial intelligence data analysis technology are integrated to screen out a specific biomarker suitable for premature ovarian failure (POI) diagnosis, and an efficient intelligent diagnosis model is constructed. The model is high in sensitivity and excellent in specificity, can effectively distinguish an ovarian healthy individual from a POI patient, and ensures the applicability and robustness of the model in different clinical samples. According to the method, only a peripheral blood sample needs to be collected for detection, additional tissue biopsy is not needed, patient wounds and detection cost are reduced, clinical feasibility is improved, the diagnosis process is simple, detection time is short, and early discovery and intervention of POI are facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of clinical laboratory diagnosis, and particularly to ovarian premature failure diagnostic markers based on metabolomics and artificial intelligence technology and their applications. Background Art

[0002] Premature ovarian insufficiency (POI), also known as primary ovarian insufficiency, refers to the decline or loss of ovarian function in women before the age of 40. It is an important disease affecting female reproductive function and endocrine health, with a global prevalence of approximately 3.7%. Epidemiological surveys show that approximately 90% of adult women in China have varying degrees of ovarian health problems. However, the specific pathogenesis of POI remains unclear, and there is currently no effective method to fully restore ovarian function. Therefore, the research on biomarker detection for female ovarian health, early diagnosis of POI, and precise treatment strategies has important clinical value.

[0003] Currently, the routine diagnosis of POI mainly relies on indicators such as the detection of six hormones (such as follicle-stimulating hormone FSH), determination of anti-Müllerian hormone (AMH) levels, assessment of ovarian volume, and antral follicle count (AFC). Although these detection methods can provide a basic diagnosis basis for POI, it is difficult to reflect the biological differences between individuals and the individualized characteristics of disease progression. Therefore, exploring individual differential biomarkers that can distinguish ovarian health status from POI and establishing an efficient early screening and precise diagnosis system are one of the key challenges in the current fields of reproductive medicine and gynecological endocrinology.

[0004] Metabolomics is an important discipline developed after genomics, transcriptomics, and proteomics. Different from other omics studies, metabolomics mainly focuses on the dynamic changes of metabolites in biological systems (cells, tissues, or organisms) after being stimulated by genetic or environmental factors. Metabolites are the final products of gene and protein functional activities and can more directly reflect the physiological state of biological systems. Therefore, metabolomics has important application value in biomarker discovery and precision medicine research and has become one of the important tools for systems biology research.

[0005] Due to the high-throughput characteristics of metabolic data, traditional principal component analysis (PCA) or linear regression methods have certain limitations in processing complex data and may overlook some key biological information. Metabolomics data analysis faces challenges such as high dimensionality, complex non-linear relationships, and difficulty in feature screening. Therefore, it is urgent to introduce more advanced data mining techniques to improve the screening efficiency of biomarkers and the accuracy of diagnostic models. Machine learning-based algorithms (such as XGBoost, random forest, deep learning, etc.) can effectively address these problems and improve the sensitivity and specificity of biomarker screening.

[0006] Therefore, combining metabolomics with artificial intelligence technology and using machine learning models for feature screening and classification modeling may break through the bottleneck of POI biomarker discovery and promote the development of precision medicine. Summary of the Invention

[0007] Based on the deficiencies in the prior art, the present invention provides a set of auxiliary diagnostic markers for premature ovarian failure discovered based on metabolomics and artificial intelligence technology and application methods. This set of markers has high sensitivity and specificity for premature ovarian failure and can be used for the auxiliary diagnosis of premature ovarian failure. The metabolite indicators provided by the present invention further contribute to the etiology analysis and precise treatment of patients.

[0008] The object of the present invention can be achieved by the following technical solutions: (1) Collect serum samples of premature ovarian failure patients and healthy individuals from different subject groups as analysis samples; (2) Perform targeted metabolomics analysis on each analysis sample using mass spectrometry technology, conduct data quality control, statistical analysis, and standardization processing on metabolite information to obtain a two-dimensional matrix, where each column is metabolite data and each row is an analysis sample; (3) Based on the metabolomics detection data of premature ovarian failure patients and the control group, establish and compare initial machine learning prediction models for the detection data through multiple machine learning algorithms to obtain the optimal preferred algorithm and the initial machine learning prediction model; (4) According to the ranking of metabolite contribution degrees provided by the initial machine learning prediction model, screen and analyze metabolites with higher contribution degrees as candidate metabolite features, and conduct comprehensive evaluation and screening such as differential significance and ROC analysis on the candidate metabolite features to determine the auxiliary diagnostic markers for premature ovarian failure; (5) According to the determined auxiliary diagnostic markers for premature ovarian failure, set a preferred feature combination from them to construct an artificial intelligence diagnostic model for premature ovarian failure; (6) Use the validation data set to verify and evaluate the prediction performance of the model, and perform iterative optimization to finally determine the metabolite combination and diagnostic model applicable to the auxiliary diagnosis of premature ovarian failure; (7) By detecting the relative contents of the metabolite combination in the subject and inputting it into a preferred machine learning model for calculation, the auxiliary diagnosis of premature ovarian failure can be achieved. The related metabolic indicators can also be applied to the auxiliary diagnosis and prediction of ovarian hypofunction and primary ovarian insufficiency. Therefore, the present invention claims the metabolite combination and its application.

[0009] In the embodiments of the present invention, the plasma samples of 38 subjects were analyzed, including 17 patients with premature ovarian failure and 21 healthy control groups. We randomly divided the subjects into a training set (12 patients with premature ovarian failure and 18 healthy control groups) and a validation set (5 patients with premature ovarian failure and 3 healthy control groups). The training set was used to screen and analyze serum metabolites related to premature ovarian failure, and the validation set was used for internal validation.

[0010] Specifically, in the embodiments of the present invention, the serum samples of patients with premature ovarian failure and healthy people in different subject groups were collected as analysis samples. A total of 111 preferred targeted metabolomics indicators were covered, including metabolites in the tricarboxylic acid cycle (citric acid, α-ketoglutaric acid, etc.), fatty acid metabolism (acetylcarnitine, palmitic acid, etc.), steroid hormone precursors (pregnenolone, dehydroepiandrosterone), etc. The detection of oxidative stress markers and mitochondrial function metabolites such as NADH and adenosine triphosphate (ATP) was increased. We obtained the normalized metabolite data of the samples using mass spectrometry technology.

[0011] In the embodiments of the present invention, 11 machine learning algorithms were screened to model and compare the performance of the metabolome detection data. The top three preferred algorithms for prediction accuracy were: CatBoost, XGBoost, and RandomForest.

[0012] In the embodiments of the present invention, the CatBoost algorithm was used to construct an initial machine learning prediction model, and the modeling parameters were as follows: the learning rate (learning_rate) was 0.03, the number of base learners (iterations) was 1000, the maximum depth of the tree (max_depth) was 6, and the l2_leaf_reg (L2 regularization parameter) was 3.

[0013] According to the contribution degree ranking of the metabolite indicators provided by the initial machine learning prediction model, and performing diagnostic performance evaluations such as differential significance and ROC analysis on the candidate metabolite features, it was found that acetoacetic acid, arginine, fatty acid C22:0, fatty acid C22:1, hyodeoxycholic acid, histidine, citrulline, ornithine, reduced nicotinamide adenine dinucleotide, and fumaric acid were used as auxiliary diagnostic markers for premature ovarian failure.

[0014] To reduce the detection cost, it is necessary to reduce the number of input metabolite features while ensuring the diagnostic performance of the artificial intelligence prediction model. For this purpose, according to the discovered auxiliary diagnostic markers for premature ovarian failure, an optimal feature combination is selected, and the CatBoost algorithm is used to construct an artificial intelligence diagnostic model based on the optimal features. In the embodiments of the present invention, the optimal metabolite feature combinations are: (1) a combination consisting of 5 metabolites, acetoacetic acid, arginine, fatty acid C22:1, fatty acid C22:0, and hyodeoxycholic acid; (2) a combination consisting of 10 metabolites, acetoacetic acid, arginine, citrulline, fatty acid C22:1, histidine, fatty acid C22:0, hyodeoxycholic acid, ornithine, NADH, and fumaric acid.

[0015] In the preferred embodiment of the present invention, the CatBoost classification models constructed based on the optimal feature combinations (1) and (2) demonstrated excellent prediction performance in the test set. Specifically, for the combination (1), based on 5 metabolites, the model had a prediction accuracy higher than 0.87, a specificity higher than 0.68, a recall rate of 1.0, a positive predictive value greater than 0.83, a negative predictive value of 1, an F1 score higher than 0.91, and a Kappa coefficient greater than 0.80. For the combination (2), based on 10 metabolites, the model had a prediction accuracy higher than 0.87, a specificity higher than 0.75, a recall rate of 1.0, a positive predictive value greater than 0.80, a negative predictive value of 1.0, an F1 score higher than 0.89, and a Kappa coefficient greater than 0.75.

[0016] These data fully demonstrate that the model has a stable prediction effect and high accuracy. The model can quickly diagnose whether it is premature ovarian failure, especially can diagnose early premature ovarian failure, and has the characteristics of accuracy, high sensitivity, and universality, and has clinical application and promotion value.

[0017] The present invention also provides the application of the auxiliary diagnostic markers for premature ovarian failure in the preparation of auxiliary diagnostic products for premature ovarian failure, ovarian hypofunction, and / or primary ovarian insufficiency.

[0018] The present invention also provides the application of substances for detecting the auxiliary diagnostic markers for premature ovarian failure in the preparation of auxiliary diagnostic products for premature ovarian failure, ovarian hypofunction, and / or primary ovarian insufficiency.

[0019] Preferably, the substance is a substance for detecting the content of diagnostic markers in serum.

[0020] Furthermore, the substance is the instruments and / or reagents required for a gas / liquid chromatography-mass spectrometry instrument for detecting the auxiliary diagnostic markers for premature ovarian failure.

[0021] The present invention also provides an auxiliary diagnostic kit for premature ovarian failure, and the auxiliary diagnostic kit for premature ovarian failure contains substances for detecting the auxiliary diagnostic markers for premature ovarian failure.

[0022] Based on the above steps, the present invention also provides a screening method for auxiliary diagnostic markers for early premature ovarian failure. The markers obtained by this method have good sensitivity and specificity for the auxiliary diagnosis of premature ovarian failure, and can especially reflect the differences of each individual, which is of great significance for the precise treatment of premature ovarian failure.

[0023] The present invention also provides a method for constructing an artificial intelligence diagnosis model for premature ovarian failure, comprising the following steps: (1) Setting the preferred metabolite feature combinations: The preferred metabolite feature combinations are two combinations of auxiliary diagnostic markers for premature ovarian failure; The first one consists of 5 metabolites, namely acetoacetic acid, arginine, fatty acid C22:1, fatty acid C22:0 and hyodeoxycholic acid; The second one consists of 10 metabolites, namely acetoacetic acid, arginine, fatty acid C22:1, fatty acid C22:0, hyodeoxycholic acid, histidine, citrulline, ornithine, reduced nicotinamide adenine dinucleotide and fumaric acid; (2) Construction of the artificial intelligence diagnosis model for premature ovarian failure: According to the preferred metabolite feature combinations, use the CatBoost algorithm to construct an artificial intelligence diagnosis model for premature ovarian failure. The modeling parameters are as follows: the learning rate (learning_rate) is 0.03, the number of base learners (iterations) is 1000, the maximum depth of the tree (max_depth) is 6, the L2 regularization parameter is 3. The following indicators are used to evaluate the prediction performance of the model, including accuracy, specificity, recall rate, positive predictive value, negative predictive value, F1 score and Kappa coefficient. Calculate the above indicators through cross-validation to evaluate the stability and generalization ability of the model, and obtain the final artificial intelligence diagnosis model for premature ovarian failure.

[0024] The method for constructing this model is simple, and has high sensitivity and specificity for the detection of premature ovarian failure, providing a strong technical guarantee for the prevention, timely diagnosis and treatment of premature ovarian failure.

[0025] Advantages of the present invention: The present invention integrates metabolomics technology and artificial intelligence data analysis technology to screen specific biomarkers suitable for the diagnosis of premature ovarian insufficiency (POI) and constructs an efficient artificial intelligence diagnosis model. The biomarker screening method is scientific, rigorous, and highly operable. By using high-throughput metabolomics technology and combining machine learning algorithms for feature screening, the accuracy and reliability of diagnostic biomarker screening are improved. The construction of the diagnosis model is simple, and the prediction performance is good. Using an artificial intelligence-based classification algorithm, the model has high sensitivity and excellent specificity, and can effectively distinguish between ovarian healthy individuals and POI patients in the validation set, verifying the applicability and robustness of the model in clinical samples.

[0026] The present invention only needs to collect peripheral blood samples for detection, without additional tissue biopsy, reducing patient trauma and detection costs, improving clinical feasibility. The diagnostic process is simple and the detection time is short, which helps to achieve early detection and intervention of POI.

[0027] Compared with traditional hormone detection methods, it overcomes the limitations of hormone level fluctuations, provides a more stable and objective diagnostic basis, is suitable for the assessment of ovarian function in women of different age groups, and can achieve efficient screening at the early stage of POI, superior to the diagnostic methods based on single hormone indicators. It has high clinical application value and broad promotion potential. This method is applicable to routine physical examinations, assisted reproductive assessments, and ovarian function monitoring, which helps the application of precision medicine in the field of reproductive health, can effectively replace the existing hormone diagnostic methods, the diagnostic process is simple and rapid, which is conducive to the discovery and timely treatment of premature ovarian failure, and has high clinical application and promotion value. Brief Description of the Drawings

[0028] Figure 1 It is a schematic diagram of the principle of the technical solution of the present invention.

[0029] Figure 2 It is the result of constructing an initial machine learning prediction model through various machine learning algorithms and comparing the algorithm performances.

[0030] Figure 3 It is the metabolite feature contribution degree and ranking provided by the initial machine learning prediction model constructed by the CatBoost algorithm.

[0031] Figure 4 For 9 candidate metabolites, the relative contents and differential significance analysis of acetoacetic acid, arginine, fatty acid C22:0, citrulline, homoserine, fatty acid C22:1, hyodeoxycholic acid, histidine, and reduced nicotinamide adenine dinucleotide in samples of healthy people and premature ovarian failure patients. P value: calculated by the unpaired analysis method of t-test using Prism10 software; those with P≤0.05 all indicate significant differences, and those with P>0.05 all indicate no significant differences; scatter points represent parallel samples.

[0032] Figure 5 For 6 candidate metabolites, the relative contents and significant differences of homoarginine, phosphoenolpyruvate, ursodeoxycholic acid, ornithine, homocitrulline and fumaric acid in samples of healthy people and patients with premature ovarian failure were analyzed. P-value: calculated using the unpaired analysis method of t-test in Prism 10 software; those with P≤0.05 all indicate significant differences, and those with P>0.05 all indicate no significant differences; scatter points represent parallel samples.

[0033] Figure 6 For 5 metabolic markers of premature ovarian failure, the ROC curve analysis results of acetoacetic acid, arginine, fatty acid C22:0, fatty acid C22:1 and hyodeoxycholic acid were obtained, the area under the curve (AUC) was calculated, and their sensitivity and specificity were evaluated to judge their diagnostic performance for samples of premature ovarian failure.

[0034] Figure 7 For 5 metabolic markers of premature ovarian failure, the ROC curve analysis results of histidine, reduced nicotinamide adenine dinucleotide, citrulline, ornithine and fumaric acid were obtained, the area under the curve (AUC) was calculated, and their sensitivity and specificity were evaluated to judge their diagnostic performance for samples of premature ovarian failure. Specific implementation manners

[0035] The implementation schemes of the present application will be described in detail below in conjunction with embodiments, but the present application is not limited to these embodiments. The test methods used in the following embodiments are all conventional methods unless otherwise specified; the materials, reagents, etc. used are all reagents and materials that can be obtained from commercial channels unless otherwise specified.

[0036] The schematic diagram of the technical solution principle of the present invention is as Figure 1 shown.

[0037] Example 1 Targeted metabolomics analysis of patients with premature ovarian failure

[0038] 1. Research subjects: Plasma samples from 38 subjects from Peking University Third Hospital were analyzed in this study, including 17 patients with premature ovarian insufficiency (POI) and 21 healthy controls. All subjects were < 40 years old, and factors such as environmental factors, infections, genetic diseases, autoimmune history, and history of radiotherapy, chemotherapy, or pelvic surgery were excluded. The POI group met the diagnostic criteria of the European Society of Human Reproduction and Embryology (ESHRE 2023), that is, the basal FSH (follicle-stimulating hormone) level was > 25 IU / L in two tests at an interval of ≥ 4 weeks, and chromosomal karyotype analysis excluded abnormalities such as Turner syndrome. The control group had regular menstrual cycles (21–35 days), anti-Müllerian hormone (AMH) levels of 1.5–4.0 ng / mL, antral follicle count (AFC) of 5–15, and basal FSH levels < 10 IU / L.

[0039] The subjects were randomly divided into a training set (12 POI patients and 18 healthy controls) and a validation set (5 POI patients and 3 healthy controls). The training set was used to screen and analyze POI-related serum metabolites, and the validation set was used for internal validation.

[0040] 2. Sample collection and quality control system: All subjects had their cubital venous blood collected after 8 hours of fasting, left to stand for 30 minutes (4 °C), then centrifuged at 3000g for 15 minutes (4 °C), aliquoted into 500 μL / tube, and stored in a -80 °C ultra-low temperature freezer. Protease inhibitors and antioxidants were added to the samples to prevent metabolite degradation. The quality control (QC) system included standard mixture standards (S1–S5). To the test samples (50 μL), 200 μL of protein precipitant containing internal standard was added, mixed well, and centrifuged at 13200 rpm for 4 minutes at low temperature. 100 μL of the supernatant was taken for detection.

[0041] 3. Metabolomics analysis using LC-MS technology: The HPLC-MS / MS system was a Dionex Ultimate 3000 liquid chromatograph (Thermo Fisher), the mass spectrometry detection was an API 3200 Q TRAP mass spectrometer (AB Sciex, USA), the chromatographic column was MSLab-AA-C18 (150 mm × 4.6 mm, 5 μm), the mobile phase was phase A (0.1% formic acid aqueous solution), phase B (0.1% formic acid acetonitrile solution), the gradient elution conditions were a column temperature of 50 °C, an injection volume of 5 μL, a flow rate of 1 mL / min, and the mass spectrometry parameters were electrospray ionization (ESI), multiple reaction monitoring (MRM) mode, ion spray voltage (IS) of 5.5 kV, a temperature of 500 °C, and the collision gas was nitrogen.

[0042] 4. Metabolomics detection and data normalization: Targeted quantitative metabolomics detection covered the tricarboxylic acid cycle (citric acid, α-ketoglutaric acid), fatty acid metabolism (acetylcarnitine, palmitic acid), steroid hormone precursors (pregnenolone, dehydroepiandrosterone), etc., and extended to oxidative stress markers and mitochondrial function metabolites (such as NADH, ATP), and a total of 111 metabolite features were detected.

[0043] One quality control (QC) sample was added to every 10 test samples, and metabolites with a relative standard deviation (RSD) > 20% were excluded. The intensity trend of the QC sample was corrected by LOESS regression, and the peak area was converted to the absolute concentration based on the standard curve of the external standard method to ensure the verification of the linear range (R² > 0.99) and the limit of detection (LOD / LOQ). All metabolite data were processed by z-score normalization.

[0044] Example 2 Construction of the initial machine learning model and screening of auxiliary diagnostic markers for premature ovarian failure 1. Construction of the initial machine learning model: Multiple machine learning methods were used to analyze the metabolome data. The metabolomics data was a two-dimensional matrix with each row representing metabolite information and each column representing the content of different metabolic markers of the analysis samples. As Figure 2 shown, candidate machine learning models were constructed based on 13 algorithms such as CatBoost and XGBoost (eXtreme Gradient Boosting). By comparing the classification accuracies of different machine learning algorithms, the most suitable and optimal algorithm was selected as CatBoost, and the initial machine learning model was constructed using the CatBoost algorithm. The modeling parameters were as follows: the learning rate was 0.03, the number of base learners (iterations) was 1000, the maximum depth of the tree (max_depth) was 6, and the l2_leaf_reg (L2 regularization parameter) was 3.

[0045] 2. Screening of auxiliary diagnostic markers for premature ovarian failure based on artificial intelligence: Based on the initial machine learning prediction model of the CatBoost algorithm, the contribution degrees of all 111 metabolite features to the model were analyzed, and they were sorted from large to small according to the contribution degrees. As Figure 3 shown, candidate differential metabolites between premature ovarian failure patients and healthy control populations were thus revealed. We used the unpaired t-test analysis method of Prism10 software to analyze the significance of differences in the top-ranked candidate metabolite features, and the results were as Figure 4 and Figure 5As shown, it was found that there were 10 metabolites showing significant differences (P value ≤ 0.05), including acetoacetic acid, arginine, fatty acid C22:0, fatty acid C22:1, hyodeoxycholic acid, histidine, citrulline, ornithine, reduced nicotinamide adenine dinucleotide, and fumaric acid.

[0046] 3. Diagnostic performance analysis of markers: For the above-mentioned metabolomic markers with significant differences, ROC analysis was performed using Prism 10 software, and the area under the curve (AUC) was calculated. At the same time, their sensitivity and specificity were evaluated to judge their diagnostic performance for premature ovarian failure. The results are as Figure 6 and Figure 7 shown. Acetoacetic acid (AUC = 0.92; sensitivity = 0.94; specificity = 0.81), arginine (AUC = 0.86; sensitivity = 0.82; specificity = 0.71), fatty acid C22:0 (AUC = 0.86; sensitivity = 0.71; specificity = 0.90), fatty acid C22:1 (AUC = 0.87; sensitivity = 0.94; specificity = 0.86), hyodeoxycholic acid (AUC = 0.94; sensitivity = 0.94; specificity = 0.81), histidine (AUC = 0.72; sensitivity = 0.76; specificity = 0.67), citrulline (AUC = 0.86, sensitivity = 0.82; specificity = 0.86), ornithine (AUC = 0.80; sensitivity = 0.82; specificity = 0.71), reduced nicotinamide adenine dinucleotide (AUC = 0.76; sensitivity = 0.82; specificity = 0.62), and fumaric acid (AUC = 0.75; sensitivity = 0.71; specificity = 0.71) all showed good diagnostic performance in the samples and could be used as auxiliary diagnostic markers for premature ovarian failure.

[0047] Example 3 Construction of a diagnostic prediction model for premature ovarian failure based on preferred features 1. Setting the preferred feature combination: To reduce the detection cost, while ensuring the diagnostic performance of the artificial intelligence prediction model, the number of input metabolite features needs to be reduced. Therefore, metabolite features need to be selected from the auxiliary diagnostic markers for premature ovarian failure and set as the preferred feature combination.

[0048] The preferred feature combinations are the following two: (1) Consisting of 5 metabolites, including acetoacetic acid, arginine, fatty acid C22:1, fatty acid C22:0, and hyodeoxycholic acid. (2) Consisting of 10 metabolites, including acetoacetic acid, arginine, fatty acid C22:1, fatty acid C22:0, hyodeoxycholic acid, histidine, citrulline, ornithine, reduced nicotinamide adenine dinucleotide, and fumaric acid.

[0049] 2. Construction of the diagnosis model for premature ovarian failure based on the optimal metabolite feature combination: According to the above optimal metabolite feature combination, use the CatBoost algorithm to construct a machine learning model, and the modeling parameters are as follows: the learning rate is 0.03, the number of base learners (iterations) is 1000, the maximum depth of the tree (max_depth) is 6, and the L2 regularization parameter is 3.

[0050] 3. The CatBoost classification model constructed based on the optimal feature combination demonstrated excellent prediction performance in the test set (validation set). Specifically, for the combination (1), based on 5 metabolite features, the model prediction accuracy was higher than 0.87, the specificity was higher than 0.68, the recall rate was 1.0, the positive predictive value was greater than 0.83, the negative predictive value was 1, the F1 score was higher than 0.91, and the Kappa coefficient was greater than 0.80. For the combination (2), based on 10 metabolite features, the model prediction accuracy was higher than 0.87, the specificity was higher than 0.75, the recall rate was 1.0, the positive predictive value was greater than 0.80, the negative predictive value was 1.0, the F1 score was higher than 0.89, and the Kappa coefficient was greater than 0.75.

Claims

1. A marker for the auxiliary diagnosis of premature ovarian failure, characterized in that: The auxiliary diagnostic marker for premature ovarian failure is any one or a combination of the following: acetoacetate, arginine, fatty acid C22:0, fatty acid C22:1, hyodeoxycholic acid, histidine, citrulline, ornithine, reduced nicotinamide adenine dinucleotide and fumaric acid.

2. Use of the auxiliary diagnostic marker for premature ovarian failure according to claim 1 in the preparation of auxiliary diagnostic products for premature ovarian failure, ovarian dysfunction and / or primary ovarian insufficiency.

3. Use of a substance for detecting the auxiliary diagnostic marker for premature ovarian failure according to claim 1 in the preparation of auxiliary diagnostic products for premature ovarian failure, ovarian dysfunction and / or primary ovarian insufficiency.

4. The use according to claim 3, characterized in that: The substance is a substance used to detect the content of diagnostic markers in serum.

5. The use according to claim 4, characterized in that: The substances are instruments and / or reagents required for a gas / liquid chromatography-mass spectrometer for detecting auxiliary diagnostic markers of premature ovarian failure.

6. A premature ovarian failure auxiliary diagnosis kit, characterized in that: The premature ovarian failure auxiliary diagnosis kit comprises a substance for detecting the premature ovarian failure auxiliary diagnosis marker described in claim 1.

7. A method for constructing an artificial intelligence diagnostic model for premature ovarian failure, characterized in that: The following steps are involved: The preferred metabolite feature combination is set, and the preferred metabolite feature combination is any one of the following two combinations: (1) It is composed of five metabolites, namely acetoacetate, arginine, fatty acid C22:1, fatty acid C22:0 and hyodeoxycholic acid; (2) It is composed of 10 metabolites, namely acetoacetate, arginine, fatty acid C22:1, fatty acid C22:0, hyodeoxycholic acid, histidine, citrulline, ornithine, reduced nicotinamide adenine dinucleotide, and fumaric acid; According to the metabolite preferred feature combination (1) or combination (2), the CatBoost algorithm was used to construct an artificial intelligence diagnosis model for premature ovarian failure. The modeling parameters were as follows: learning rate was 0.03, the number of base learners was 1000, the maximum depth of the tree was 6, and the L2 regularization parameter was 3. The artificial intelligence diagnosis model for premature ovarian failure was obtained.

8. The method for constructing an artificial intelligence diagnostic model for premature ovarian failure according to claim 7, characterized in that: The following indicators were used to evaluate the predictive performance of the model, including accuracy, specificity, recall, positive predictive value, negative predictive value, F1 score and Kappa coefficient. The above indicators were calculated through cross-validation to evaluate the stability and generalization ability of the model.

Citation Information

Patent Citations

  • Ovarian premature senility related gene whole exon amplification and detection method

    CN110564839A

  • Diabetic nephropathy prediction method, system and device based on machine learning

    CN113470816A

  • Compositions and methods for modulating ovarian follicular initiation

    WO2005007687A1