Construction method of central precocious puberty prediction model and storage medium
By constructing a multi-feature-based prediction model for central precocious puberty, using the Lasso regression algorithm to screen predictive factors and train a machine learning model, the adverse reaction problem in the diagnosis of central precocious puberty was solved, achieving efficient and low-cost prediction of precocious puberty, which is suitable for primary hospitals.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-03
AI Technical Summary
Current diagnostic methods for central precocious puberty rely on gonadotropin stimulation tests, which carry risks of adverse reactions such as allergies and are often resisted by children. Furthermore, these specialized stimulation tests are difficult to conduct in primary care hospitals.
A multi-feature-based prediction model for central precocious puberty was constructed. By acquiring patient data of female patients with precocious puberty, Lasso regression algorithm was used to screen independent predictors, a machine learning model was trained, and the model parameters were adjusted through a validation group and an external test group to finally obtain the optimal prediction model. A preset number of independent predictors were then input to predict the probability of precocious puberty.
It can replace some gonadotropin stimulation tests, reduce medical costs and patient discomfort, has high predictive accuracy, and is suitable for primary hospitals.
Smart Images

Figure CN121789995A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical data processing, specifically to a method for constructing and storing a multi-feature-based prediction model for central precocious puberty. Background Technology
[0002] Central precocious puberty is a common endocrine disorder in children, which can cause advanced bone age and premature closure of the epiphyses, leading to impaired height. Premature sexual characteristics may also cause psychological and behavioral abnormalities. Early identification of the risks of this disease is beneficial for timely treatment. Currently, the diagnosis of central precocious puberty relies on the gonadotropin stimulation test. This test requires multiple intravenous blood samples to be drawn within 2 hours after the injection of gonadotropins. This carries the risk of adverse reactions such as allergies, and children often resist it. Furthermore, most primary care hospitals have difficulty performing this specialized stimulation test. Summary of the Invention
[0003] In view of the above problems, this application provides a method for constructing a central precocious puberty prediction model based on multiple features and a storage medium, which solves the problem that the existing diagnosis of central precocious puberty relies on gonadotropin stimulation test, which requires injection of gonadotropin and multiple intravenous blood sampling tests within 2 hours, which poses the risk of adverse reactions such as allergies and is often resisted by children.
[0004] To achieve the above objectives, the inventors provide a method for constructing a central precocious puberty prediction model based on multiple features, comprising the following steps:
[0005] Acquire patient data from several female patients with precocious puberty. The patient data includes several candidate predictive factors, including actual age, disease duration, height standard deviation, body mass index, breast Tanner stage, vulvar Tanner stage, basal luteinizing hormone, basal follicle-stimulating hormone, bone age-to-actual age difference, bone age-to-height standard deviation, uterine volume, and maximum ovarian volume.
[0006] The acquired patient data were randomly divided into a training group, a validation group, and an external testing group.
[0007] The Lasso regression algorithm is used to select a predetermined number of independent predictors from several candidate predictors.
[0008] Based on a preset number of independent predictors after screening, several machine learning models are trained using patient data from the training group.
[0009] The parameters of several machine learning models were adjusted using patient data from the validation group.
[0010] The generalization ability of several machine learning models was validated using patient data from an external test group, and the optimal machine learning model was selected from these models as a prediction model for central precocious puberty.
[0011] In some embodiments, after obtaining patient data from a number of female patients with precocious puberty, the method further includes the following steps:
[0012] Classify and transform skewed continuous variables in patient data.
[0013] In some embodiments, the machine learning model includes logistic regression analysis model, linear support vector machine, random forest, XGBoost and fully connected neural network.
[0014] In some embodiments, the linear support vector machine uses a Gaussian radial basis function kernel, the random forest has 300 decision trees, the XGBoost has a learning rate of 0.1 and 50 trees, and the fully connected neural network contains two hidden layers and a SeLU activation function.
[0015] In some embodiments, the external test group consists of three.
[0016] Another technical solution is also provided: a storage medium storing a computer program, which, when executed by a processor, performs the following steps:
[0017] Acquire patient data from several female patients with precocious puberty. The patient data includes several candidate predictive factors, including actual age, disease duration, height standard deviation, body mass index, breast Tanner stage, vulvar Tanner stage, basal luteinizing hormone, basal follicle-stimulating hormone, bone age-to-actual age difference, bone age-to-height standard deviation, uterine volume, and maximum ovarian volume.
[0018] The acquired patient data were randomly divided into a training group, a validation group, and an external testing group.
[0019] The Lasso regression algorithm is used to select a predetermined number of independent predictors from several candidate predictors.
[0020] Based on a preset number of independent predictors after screening, several machine learning models are trained using patient data from the training group.
[0021] The parameters of several machine learning models were adjusted using patient data from the validation group.
[0022] The generalization ability of several machine learning models was validated using patient data from an external test group, and the optimal machine learning model was selected from these models as a prediction model for central precocious puberty.
[0023] In some embodiments, after obtaining patient data from a number of female patients with precocious puberty, the method further includes the following steps:
[0024] Classify and transform skewed continuous variables in patient data.
[0025] In some embodiments, the machine learning model includes logistic regression analysis model, linear support vector machine, random forest, XGBoost and fully connected neural network.
[0026] In some embodiments, the linear support vector machine uses a Gaussian radial basis function kernel, the random forest has 300 decision trees, the XGBoost has a learning rate of 0.1 and 50 trees, and the fully connected neural network contains two hidden layers and a SeLU activation function.
[0027] In some embodiments, the external test group consists of three.
[0028] Unlike existing technologies, the above-mentioned technical solution acquires patient data from several female patients with precocious puberty. This data includes several candidate predictive factors. The patient data is divided into a training group, a validation group, and an external test group. A predetermined number of independent predictive factors are selected from the candidate predictive factors using the Lasso regression algorithm. Based on these selected independent predictive factors, several machine learning models are trained using patient data from the training group. The parameters of these machine learning models are then adjusted using patient data from the validation group. Finally, the generalization ability of these machine learning models is validated using patient data from the external test group. The optimal machine learning model is then obtained as the central precocious puberty prediction model. By inputting the predetermined number of independent predictive factors into the central precocious puberty prediction model, the probability of a user developing precocious puberty can be predicted, potentially replacing some gonadotropin stimulation tests. Furthermore, the selected independent predictive factors are all routine clinical indicators, requiring no additional testing, thus reducing medical costs and patient discomfort.
[0029] The above description of the invention is merely an overview of the technical solution of this application. In order to enable those skilled in the art to better understand the technical solution of this application and to implement it based on the description and drawings, and to make the above-mentioned objectives and other objectives, features and advantages of this application easier to understand, the following description is provided in conjunction with the specific embodiments and drawings of this application. Attached Figure Description
[0030] The accompanying drawings are only used to illustrate the principles, implementation methods, applications, features, and effects of specific embodiments of this application and other related content, and should not be considered as limitations on this application.
[0031] In the accompanying drawings of the instruction manual:
[0032] Figure 1 This is a flowchart illustrating a method for constructing a central precocious puberty prediction model based on multiple features, as described in a specific implementation.
[0033] Figure 2 This is another flowchart illustrating the construction method of the central precocious puberty prediction model based on multiple features as described in the specific implementation.
[0034] Figure 3 This is a schematic diagram of the structure of the storage medium described in a specific embodiment.
[0035] The reference numerals used in the above figures are explained as follows:
[0036] 310. Storage medium,
[0037] 320. Processor. Detailed Implementation
[0038] To illustrate the possible application scenarios, technical principles, implementable specific solutions, and achievable objectives and effects of this application in detail, the following description, in conjunction with the listed specific embodiments and accompanying drawings, provides a detailed explanation. The embodiments described herein are merely illustrative of the technical solutions of this application and are therefore intended to limit the scope of protection of this application.
[0039] In this document, the term "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The term "embodiment" appearing in various places throughout the specification does not necessarily refer to the same embodiment, nor does it specifically limit its independence or connection with other embodiments. In principle, in this application, as long as there are no technical contradictions or conflicts, the technical features mentioned in each embodiment can be combined in any way to form corresponding implementable technical solutions.
[0040] Unless otherwise defined, the technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the use of related terms herein is merely for the purpose of describing particular embodiments and is not intended to limit this application.
[0041] In the description of this application, the term "and / or" is used to describe the logical relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A exists, B exists, and A and B exist simultaneously. Additionally, the character " / " in this document generally indicates that the preceding and following objects have an "or" logical relationship.
[0042] In this application, terms such as “first” and “second” are used only to distinguish one entity or operation from another, and do not necessarily require or imply any actual quantity, hierarchy or order relationship between these entities or operations.
[0043] Unless otherwise specified, the use of terms such as “comprising,” “including,” “having,” or other similar expressions in this application is intended to cover non-exclusive inclusion, which does not exclude the presence of additional elements in a process, method, or product that includes the stated elements, such that a process, method, or product that includes a list of elements may include not only those defined elements but also other elements not expressly listed, or elements inherent to such a process, method, or product.
[0044] As understood in the Examination Guidelines, in this application, expressions such as "greater than," "less than," and "exceeding" are understood to exclude the stated number; expressions such as "above," "below," and "within" are understood to include the stated number. Furthermore, in the description of the embodiments in this application, "multiple" means two or more (including two), and similar expressions related to "multiple" are also understood in this way, such as "multiple groups" and "multiple times," unless otherwise explicitly specified.
[0045] In the description of the embodiments of this application, the space-related expressions used, such as "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "vertical," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," and "circumferential," indicate the orientation or positional relationship based on the orientation or positional relationship shown in the specific embodiments or drawings. They are only for the purpose of describing the specific embodiments of this application or for the reader's understanding, and do not indicate or imply that the device or component referred to must have a specific position, a specific orientation, or be constructed or operated in a specific orientation. Therefore, they should not be construed as limitations on the embodiments of this application.
[0046] Unless otherwise expressly specified or limited, the terms "installation," "connection," "linking," "fixing," and "setting," as used in the description of the embodiments of this application, should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral setting; it can be a mechanical connection, an electrical connection, or a communication connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be the internal connection of two components or the interaction between two components. For those skilled in the art to which this application pertains, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.
[0047] Please see Figure 1 This embodiment provides a method for constructing a central precocious puberty prediction model based on multiple features, including the following steps:
[0048] Step S110: Obtain patient data for several female patients with precocious puberty. The patient data includes several candidate predictive factors, including actual age, disease duration, height standard deviation, body mass index, breast Tanner stage, vulvar Tanner stage, basal luteinizing hormone, basal follicle-stimulating hormone, bone age-actual age difference, bone age-height standard deviation, uterine volume, and maximum ovarian volume.
[0049] Step S120: Randomly divide the acquired patient data into a training group, a validation group, and an external test group;
[0050] Step S130: Select a preset number of independent predictors from several candidate predictors based on the Lasso regression algorithm;
[0051] Step S140: Based on the preset number of independent predictors after screening, train several machine learning models using patient data in the training group;
[0052] Step S150: Adjust the model parameters of several machine learning models using patient data from the validation group;
[0053] Step S160: Validate the generalization ability of several machine learning models using patient data from an external test group, and select the optimal machine learning model from among the several machine learning models as the prediction model for central precocious puberty.
[0054] By acquiring patient data from several female patients with precocious puberty, including several candidate predictive factors, the patient data was divided into a training group, a validation group, and an external test group. A predetermined number of independent predictive factors were selected from the candidate predictive factors using the Lasso regression algorithm. Based on these selected independent predictive factors, several machine learning models were trained using patient data from the training group. The model parameters of the trained machine learning models were then adjusted using patient data from the validation group. Finally, the generalization ability of the machine learning models was validated using patient data from the external test group. The optimal machine learning model was ultimately obtained as the central precocious puberty prediction model. By inputting the predetermined number of independent predictive factors into the central precocious puberty prediction model, the probability of a user developing precocious puberty can be predicted, potentially replacing some gonadotropin stimulation tests. Furthermore, the selected independent predictive factors are all routine clinical indicators, requiring no additional testing, thus reducing medical costs and patient discomfort.
[0055] Please see Figure 2 In some embodiments, after obtaining patient data from several female patients with precocious puberty, the following steps are further included:
[0056] Step S210: Perform classification transformation on the skewed continuous variables in the patient data.
[0057] By applying classification transformation rules to skewed distribution variables in patient data, Tanner staging is grouped and merged to improve the model's adaptability to heterogeneous clinical data. Specifically, skewed continuous variables such as disease duration and baseline luteinizing hormone (LH) are classified and transformed: disease duration is divided into <0.5 years, 0.5-1 years, and ≥1 year; baseline LH is divided into <0.20 IU / L, 0.20-0.83 IU / L, and ≥0.83 IU / L. Tanner staging (breast and pubic hair) is grouped and merged (breast staging B=2 vs B>2; pubic hair staging P=1 vs P>1).
[0058] In some embodiments, the machine learning model includes a logistic regression analysis model, a linear support vector machine, a random forest, XGBoost, and a fully connected neural network. By training these five machine learning models—logistic regression analysis model, linear support vector machine, random forest, XGBoost, and fully connected neural network—a machine learning model with high prediction accuracy can be obtained as a central precocious convergence prediction model. In other embodiments, two or more machine learning models can be selected from the logistic regression analysis model, linear support vector machine, random forest, XGBoost, and fully connected neural network to obtain the optimal machine learning model as the central precocious convergence prediction model. Alternatively, other machine learning models besides the above five can be selected to train the optimal machine learning model as the central precocious convergence prediction model.
[0059] In some embodiments, the linear support vector machine uses a Gaussian radial basis function kernel, the random forest has 300 decision trees, the XGBoost has a learning rate of 0.1 and 50 trees, and the fully connected neural network contains two hidden layers and a SeLU activation function. By using a Gaussian radial basis function kernel in the SVM, a random forest with 300 decision trees, an XGBoost with a learning rate of 0.1 and 50 trees, and a neural network with two hidden layers (8 / 4 neurons) and a SeLU activation function, parameter tuning of the desired machine learning model can train a machine learning model with higher prediction accuracy.
[0060] In some embodiments, the external test group consists of three groups. By evaluating the generalization ability of several machine learning models trained on the three external test groups, the optimal machine learning model can be selected as the central precocious puberty prediction model.
[0061] In some embodiments, a method for constructing a central precocious puberty prediction model based on multiple features includes the following steps:
[0062] (1) Data collection and preprocessing;
[0063] Study subjects: Data from 2148 female patients with precocious puberty (PP) were collected in a multicenter retrospective study. They were divided into a training group (1048 cases), a validation group (262 cases), and three external test groups (270, 278, and 290 cases, respectively). The study covered 12 candidate predictive factors, including chronological age, disease duration, height standard deviation, body mass index, breast Tanner stage, vulvar Tanner stage, basal luteinizing hormone, basal follicle-stimulating hormone, bone age-chronological age difference, bone age-height standard deviation, uterine volume, and maximum ovarian volume.
[0064] Data transformation: skewed continuous variables (such as disease duration and baseline luteinizing hormone) were classified and transformed (disease duration was divided into <0.5 years, 0.5-1 years, and ≥1 year; baseline luteinizing hormone was divided into <0.20 IU / L, 0.20-0.83 IU / L, and ≥0.83 IU / L); Tanner staging (breast and pubic hair) was merged and grouped (breast staging B=2 vs B>2; pubic hair staging P=1 vs P>1).
[0065] (2) Predictor selection;
[0066] Using the Lasso regression algorithm, based on the 1 standard error criterion of λ=0.00077, eight independent predictors were selected: actual age, disease duration, height standard deviation, baseline luteinizing hormone, bone age-actual age difference, bone age-height standard deviation, uterine volume, and maximum ovarian volume.
[0067] (3) Model building and training;
[0068] Based on the eight selected predictive factors, five machine learning models were constructed: logistic regression, linear support vector machine (SVM), random forest, XGBoost, and fully connected neural network. The models were optimized through the following steps:
[0069] Parameter tuning: SVM uses Gaussian radial basis kernel function, random forest has 300 decision trees, XGBoost has learning rate of 0.1 and number of trees of 50, and neural network has 2 hidden layers (8 / 4 neurons) and SeLU activation function;
[0070] Internal validation: Model parameters were adjusted using validation set data (262 cases), and evaluation metrics included C-statistic (AUC), accuracy, and calibration curve;
[0071] External validation: Three independent test groups (a total of 838 cases) were used to verify the generalization ability of the model, and the optimal model (SVM) was finally selected.
[0072] (4) Model application;
[0073] The optimal SVM model is deployed as an online diagnostic tool. After inputting the values of 8 predictive factors, it outputs the probability of central precocious puberty to assist clinical decision-making.
[0074] For the first time, Lasso regression was used to screen out eight highly specific diagnostic indicators for central precocious puberty (including bone age-to-actual age difference, bone age-specific height standard deviation, and other bone age-related parameters, as well as reproductive organ volume indicators), solving the problems of redundant predictive factors and poor clinical operability in traditional models.
[0075] Multi-center machine learning model construction: Based on 2148 samples from 4 centers, a three-stage validation strategy of "training-validation-external testing" (5 groups of data) was adopted to ensure the stability of the model in different clinical scenarios. The average accuracy of the SVM model in external validation reached 72.1%, and the AUC reached 0.827.
[0076] Clinical translational tools: Develop lightweight online calculators to achieve non-invasive and rapid assessment, reduce the use of gonadotropin stimulation tests (avoiding multiple blood draws), and make them suitable for primary hospitals and large-scale screening.
[0077] Data preprocessing methods: Classification transformation rules were designed for skewed distribution variables (such as baseline luteinizing hormone), and Tanner staging was merged and grouped to improve the model's adaptability to heterogeneous clinical data.
[0078] It has the following advantages:
[0079] High diagnostic efficacy: The SVM model performed excellently in multicenter validation (AUC 0.827-0.850), with an accuracy of 72.1%-78.6%, and can replace some gonadotropin stimulation tests.
[0080] Clinical applicability: All eight predictive factors are routine clinical examination indicators (bone age, ultrasound, and baseline hormones), requiring no additional testing and reducing medical costs.
[0081] Wide applicability: Verified by multi-center data from both northern and southern China, it is suitable for the diagnosis of central precocious puberty in Chinese girls, and is especially suitable for promotion in primary healthcare institutions.
[0082] Please see Figure 3 In another embodiment, a storage medium 310 stores a computer program, which, when executed by a processor 320, performs the following steps:
[0083] Acquire patient data from several female patients with precocious puberty. The patient data includes several candidate predictive factors, including actual age, disease duration, height standard deviation, body mass index, breast Tanner stage, vulvar Tanner stage, basal luteinizing hormone, basal follicle-stimulating hormone, bone age-to-actual age difference, bone age-to-height standard deviation, uterine volume, and maximum ovarian volume.
[0084] The acquired patient data were randomly divided into a training group, a validation group, and an external testing group.
[0085] The Lasso regression algorithm is used to select a predetermined number of independent predictors from several candidate predictors.
[0086] Based on a preset number of independent predictors after screening, several machine learning models are trained using patient data from the training group.
[0087] The parameters of several machine learning models were adjusted using patient data from the validation group.
[0088] The generalization ability of several machine learning models was validated using patient data from an external test group, and the optimal machine learning model was selected from these models as a prediction model for central precocious puberty.
[0089] By acquiring patient data from several female patients with precocious puberty, including several candidate predictive factors, the patient data was divided into a training group, a validation group, and an external test group. A predetermined number of independent predictive factors were selected from the candidate predictive factors using the Lasso regression algorithm. Based on these selected independent predictive factors, several machine learning models were trained using patient data from the training group. The model parameters of the trained machine learning models were then adjusted using patient data from the validation group. Finally, the generalization ability of the machine learning models was validated using patient data from the external test group. The optimal machine learning model was ultimately obtained as the central precocious puberty prediction model. By inputting the predetermined number of independent predictive factors into the central precocious puberty prediction model, the probability of a user developing precocious puberty can be predicted, potentially replacing some gonadotropin stimulation tests. Furthermore, the selected independent predictive factors are all routine clinical indicators, requiring no additional testing, thus reducing medical costs and patient discomfort.
[0090] In some embodiments, after obtaining patient data from a number of female patients with precocious puberty, the method further includes the following steps:
[0091] Classify and transform skewed continuous variables in patient data.
[0092] By applying classification transformation rules to skewed distribution variables in patient data, Tanner staging is grouped and merged to improve the model's adaptability to heterogeneous clinical data. Specifically, skewed continuous variables such as disease duration and baseline luteinizing hormone (LH) are classified and transformed: disease duration is divided into <0.5 years, 0.5-1 years, and ≥1 year; baseline LH is divided into <0.20 IU / L, 0.20-0.83 IU / L, and ≥0.83 IU / L. Tanner staging (breast and pubic hair) is grouped and merged (breast staging B=2 vs B>2; pubic hair staging P=1 vs P>1).
[0093] In some embodiments, the machine learning model includes a logistic regression analysis model, a linear support vector machine, a random forest, XGBoost, and a fully connected neural network. By training these five machine learning models—logistic regression analysis model, linear support vector machine, random forest, XGBoost, and fully connected neural network—a machine learning model with high prediction accuracy can be obtained as a central precocious convergence prediction model. In other embodiments, two or more machine learning models can be selected from the logistic regression analysis model, linear support vector machine, random forest, XGBoost, and fully connected neural network to obtain the optimal machine learning model as the central precocious convergence prediction model. Alternatively, other machine learning models besides the above five can be selected to train the optimal machine learning model as the central precocious convergence prediction model.
[0094] In some embodiments, the linear support vector machine uses a Gaussian radial basis function kernel, the random forest has 300 decision trees, the XGBoost has a learning rate of 0.1 and 50 trees, and the fully connected neural network contains two hidden layers and a SeLU activation function. By using a Gaussian radial basis function kernel in the SVM, a random forest with 300 decision trees, an XGBoost with a learning rate of 0.1 and 50 trees, and a neural network with two hidden layers (8 / 4 neurons) and a SeLU activation function, parameter tuning of the desired machine learning model can train a machine learning model with higher prediction accuracy.
[0095] In some embodiments, the external test group consists of three groups. By evaluating the generalization ability of several machine learning models trained on the three external test groups, the optimal machine learning model can be selected as the central precocious puberty prediction model.
[0096] Finally, it should be noted that although the above embodiments have been described in the text and drawings of this application, this should not limit the scope of patent protection of this application. Any technical solutions that are based on the essential concept of this application and utilize the content described in the text and drawings of this application, resulting in equivalent structural or procedural substitutions or modifications, as well as the direct or indirect application of the technical solutions of the above embodiments to other related technical fields, are all included within the scope of patent protection of this application.
Claims
1. A method for constructing a central precocious puberty prediction model based on multiple features, characterized in that, Includes the following steps: Acquire patient data from several female patients with precocious puberty. The patient data includes several candidate predictive factors, including actual age, disease duration, height standard deviation, body mass index, breast Tanner stage, vulvar Tanner stage, basal luteinizing hormone, basal follicle-stimulating hormone, bone age-to-actual age difference, bone age-to-height standard deviation, uterine volume, and maximum ovarian volume. The acquired patient data were randomly divided into a training group, a validation group, and an external testing group. The Lasso regression algorithm is used to select a predetermined number of independent predictors from several candidate predictors. Based on a preset number of independent predictors after screening, several machine learning models are trained using patient data from the training group. The parameters of several machine learning models were adjusted using patient data from the validation group. The generalization ability of several machine learning models was validated using patient data from an external test group, and the optimal machine learning model was selected from these models as a prediction model for central precocious puberty.
2. The method for constructing a central precocious puberty prediction model based on multiple features according to claim 1, characterized in that, After obtaining patient data from several female patients with precocious puberty, the following steps are also included: Classify and transform skewed continuous variables in patient data.
3. The method for constructing a central precocious puberty prediction model based on multiple features according to claim 1, characterized in that, The machine learning models include logistic regression analysis models, linear support vector machines, random forests, XGBoost, and fully connected neural networks.
4. The method for constructing a central precocious puberty prediction model based on multiple features according to claim 3, characterized in that, The linear support vector machine uses a Gaussian radial basis kernel function, the random forest has 300 decision trees, the XGBoost has a learning rate of 0.1 and 50 trees, and the fully connected neural network contains two hidden layers and a SeLU activation function.
5. The method for constructing a central precocious puberty prediction model based on multiple features according to claim 1, characterized in that, There are three external test groups.
6. A storage medium storing a computer program, characterized in that, The computer program, when executed by the processor, performs the following steps: Acquire patient data from several female patients with precocious puberty. The patient data includes several candidate predictive factors, including actual age, disease duration, height standard deviation, body mass index, breast Tanner stage, vulvar Tanner stage, basal luteinizing hormone, basal follicle-stimulating hormone, bone age-to-actual age difference, bone age-to-height standard deviation, uterine volume, and maximum ovarian volume. The acquired patient data were randomly divided into a training group, a validation group, and an external testing group. The Lasso regression algorithm is used to select a predetermined number of independent predictors from several candidate predictors. Based on a preset number of independent predictors after screening, several machine learning models are trained using patient data from the training group. The parameters of several machine learning models were adjusted using patient data from the validation group. The generalization ability of several machine learning models was validated using patient data from an external test group, and the optimal machine learning model was selected from these models as a prediction model for central precocious puberty.
7. The storage medium according to claim 6, characterized in that, After obtaining patient data from several female patients with precocious puberty, the following steps are also included: Classify and transform skewed continuous variables in patient data.
8. The storage medium according to claim 6, characterized in that, The machine learning models include logistic regression analysis models, linear support vector machines, random forests, XGBoost, and fully connected neural networks.
9. The storage medium according to claim 8, characterized in that, The linear support vector machine uses a Gaussian radial basis kernel function, the random forest has 300 decision trees, the XGBoost has a learning rate of 0.1 and 50 trees, and the fully connected neural network contains two hidden layers and a SeLU activation function.
10. The storage medium according to claim 6, characterized in that, There are three external test groups.