Method and device for classifying gestational diabetes mellitus at early pregnancy stage
By using cfDNA fragment omics data and neural network models, combined with characteristic parameters and clinical information, the accuracy and cost issues of early prediction of gestational diabetes were solved, efficient gestational diabetes classification was achieved, and the risk of complications was reduced.
Patent Information
- Application Number
- CN202410369398.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-28
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies have problems with high cost or poor accuracy in predicting gestational diabetes, especially in the classification of healthy people and people with gestational diabetes. Proteomic analysis technology is immature and the prediction results of clinical groups are inaccurate.
Using cell-free DNA (cfDNA) fragment omics data combined with a neural network model, a model for predicting gestational diabetes in early pregnancy was constructed through feature parameter screening and analysis. Feature parameters such as terminal motifs, MDS data, CGN/NCG, fetal concentration and TSS data were used for classification, and early prediction was performed in combination with clinical information.
Highly accurate prediction of gestational diabetes was achieved in the early stages of pregnancy (11-13 weeks), with an AUC of 0.856-0.964 and an accuracy of 0.8217-0.8821, reducing the risk of pregnancy complications and being cost-effective.
Smart Images

Figure CN120766985A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gestational diabetes, and in particular to a classification method and device for early pregnancy gestational diabetes. Background Art
[0002] Gestational Diabetes Mellitus (GDM) refers to a disorder of glucose metabolism that occurs during pregnancy. GDM is a common pregnancy complication characterized by high blood sugar levels during pregnancy. While these conditions typically resolve spontaneously after the pregnancy, there is a nearly 50% chance of developing type 2 diabetes postpartum. GDM has a significant impact on both maternal and fetal health. For mothers, GDM increases the risk of complications such as cesarean section, shoulder dystocia, birth trauma, and pregnancy-induced hypertension (including eclampsia). Furthermore, GDM may negatively impact the mother's long-term cardiovascular metabolism. Randomized clinical trials have shown that early pregnancy intervention for pregnant women at high risk for GDM can significantly reduce the incidence of GDM, minimize maternal and fetal complications, and improve pregnancy outcomes. Therefore, it is crucial to assess the risk of GDM in the first and second trimesters and implement targeted interventions for high-risk women.
[0003] Cell-free DNA (cfDNA) refers to DNA fragments present in body fluids, including those released by host cells. Studies have shown that analysis of cfDNA can provide biological information about an individual, including genomic mutations, methylation patterns, and chromosomal abnormalities. In medicine, cfDNA research has been widely applied in the diagnosis and monitoring of various diseases, achieving significant progress in oncology and prenatal diagnosis.
[0004] Currently, patents related to early pregnancy prediction of GDM mainly cover the following aspects: clinical data, metagenome, proteome, and metabolome. There are also some recent reports on lipidome. For example:
[0005] Yan-Ting Wu et al. analyzed the medical records of 16,819 women diagnosed with gestational diabetes. Using the K-Nearest Neighbor algorithm, they identified seven features from sociodemographic characteristics, early pregnancy clinical variables, and laboratory parameters. These features were then used for prediction. The results showed that this method achieved a predictive accuracy (expressed as an area under the curve (AUC)) of 0.77.
[0006] Hou, G., et al. from the BGI-Research Institute of Life Sciences conducted a study in which they collected samples from 100 patients with GDM and 100 normal controls during the first and second trimesters. The results showed that significant quantitative differences in seven triglycerides and five diglycerides were found between the first and second trimesters. By analyzing these quantitative differences in triglycerides and diglycerides, the research team developed a prediction model and evaluated its accuracy using the area under the curve (AUC). The model achieved an AUC of 0.88.
[0007] There are also studies using microbial genomes to predict GDM. For example, Jinfeng Wang et al. collected saliva from 44 GDM patients and performed 16s rRNA sequencing. Based on two oral microbes, Lautropia Neisseria and Veillonella, they predicted GDM with an AUC of 0.83.
[0008] In 2013, Juha P Rasanen et al. showed that fibronectin and pregnancy-specific glycoprotein (PSG) can be used to predict GDM, and their model AUC can reach 0.91.
[0009] In 2021, Brian J Koos et al. studied 46 healthy pregnant women and 46 pregnant women with GDM and analyzed the urine metabolome. Their research showed that the AUC of the model for predicting GDM based on 7 urine metabolites could reach 0.97.
[0010] Currently, there are still few studies on the relationship between cfDNA and GDM. In 2019, Xiaosong Yuan et al. investigated the relationship between cfDNA and various pregnancy diseases. They found that the total amount of cfDNA was related to the risk of GDM
[10] .
[0011] However, the above technologies have the following problems: 1) There are problems of high cost and immature technology in the proteomic field: proteomic analysis technology is still in the development stage and faces challenges in technological maturity and standardization; 2) The current prediction results of the clinical group (clinical group refers to the method of prediction based on clinical indicators and characteristics) are poor and cannot accurately reflect the current pregnancy situation. Summary of the Invention
[0012] The present invention aims to provide a method and device for classifying gestational diabetes in early pregnancy, so as to solve the technical problems in the prior art of high cost or poor accuracy in classifying healthy people and people with gestational diabetes.
[0013] To achieve the above objectives, according to one aspect of the present invention, a method for classifying early pregnancy gestational diabetes is provided. The method comprises: obtaining a set of characteristic parameters for a sample to be tested, wherein the set of characteristic parameters is extracted from cfDNA fragment omics data of gestational diabetes, and the set of characteristic parameters includes at least one of the following: terminal motifs, MDS data, CGN / NCG data, fetal concentration, and TSS data; inputting the characteristic parameters into a model for predicting early pregnancy gestational diabetes, performing feature classification using the model, and outputting a classification result for the sample to be tested.
[0014] Furthermore, the model for predicting gestational diabetes in early pregnancy includes: a first classification layer, and a second classification layer connected to the first classification layer, the first classification layer includes at least one of the following: a first neural network layer, a second neural network layer, a third neural network layer, a fourth neural network layer and a fifth neural network layer, and the types of the first classification layer network include a fully connected neural network, a neural network and a convolutional neural network.
[0015] Furthermore, the characteristic parameters are input into the early pregnancy prediction model for gestational diabetes, and the early pregnancy prediction model for gestational diabetes performs feature classification, and outputs the classification results of the sample to be tested, including: using the first neural network layer to process the terminal motif to obtain the first intermediate data, and the first neural network layer is a fully connected neural network; using the second neural network layer to process CGN / NCG to obtain the second intermediate data, and the second neural network layer is a neural network; using the third neural network layer to process the fetal concentration to obtain the third intermediate data, and the third neural network layer is a neural network; using the fourth neural network layer to process the MDS data to obtain the fourth intermediate data, and the fourth neural network layer is a neural network; using the fifth neural network layer to process the TSS data to obtain the fifth intermediate data, and the fifth neural network layer is a convolutional neural network; using the second classification layer to process the intermediate data output by the first classification layer to obtain the classification result, and the intermediate data output by the first classification layer includes the first intermediate data, the second intermediate data, the third intermediate data, the fourth intermediate data and the fifth intermediate data.
[0016] Furthermore, the first neural network layer is a fully connected neural network for processing terminal motifs; the second neural network layer is a neural network for processing CGN / NCG; the third neural network layer is a neural network for processing fetal concentration; the fourth neural network layer is a neural network for processing MDS data; the fifth neural network layer is a convolutional neural network for processing TSS data; and the second classification layer is a fully connected neural network.
[0017] Furthermore, when obtaining MDS data, it includes: scoring the diversity of terminal motifs according to the occurrence probability Pi of a certain terminal motif.
[0018] Further, the CGN / NCG obtaining includes: performing ratio analysis on the fragments starting with CGN and the fragments starting with NCG; and / or, the fetal concentration obtaining includes obtaining the fetal concentration by using a method of estimating the fetal concentration by calculating the proportion of the partial bin.
[0019] Further, the TSS data generating includes: analyzing the coverage of the transcription start site region; preferably, the analyzing the coverage of the transcription start site region includes: defining the transcription start site region: selecting the transcription start site of the target gene, and determining the fixed value of the transcription start site; calculating the coverage: analyzing the aligned sequencing data file by using the depth function, calculating the number of times of the sequencing reads covering in the transcription start site region, to obtain the coverage; normalizing the coverage: normalizing the coverage of the transcription start site region to generate the TSS data.
[0020] Further, the classification method further includes clinical information analysis, and the clinical information includes one or more of height, weight, age, BMI, smoking, and alcohol consumption.
[0021] According to another aspect of the present application, a device for classifying early pregnancy gestational diabetes is provided. The device includes: an obtaining unit configured to obtain a feature parameter set of a sample to be tested, wherein the feature parameter set is extracted from cfDNA fragmentomic data of gestational diabetes, and the feature parameter set includes at least one of the following: end motif, MDS data, CGN / NCG, fetal concentration, and TSS data; and a classification unit configured to input the feature parameters into an early pregnancy gestational diabetes prediction model, perform feature classification by using the early pregnancy gestational diabetes prediction model, and output a classification result of the sample to be tested.
[0022] According to another aspect of the present application, an electronic device is provided. The electronic device includes: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any of the above-mentioned classification methods of early pregnancy gestational diabetes by executing the executable instructions.
[0023] According to another aspect of the present application, a computer readable storage medium is provided. The computer readable storage medium includes a stored computer program, wherein the computer readable storage medium controls the device where the computer readable storage medium is located to execute any of the above-mentioned classification methods of early pregnancy gestational diabetes when the computer program is running.
[0024] By using the feature parameters related to gestational diabetes to construct an early pregnancy gestational diabetes prediction model, the healthy population and the gestational diabetes population can be accurately classified by using the technical solution of the present application.
[0025] Furthermore, the present invention combines clinical information and cell-free DNA (cfDNA) fragment omics data to apply a neural network model to construct a model for predicting gestational diabetes in early pregnancy. Specific experimental data show that accurate prediction of gestational diabetes mellitus (GDM) can be performed in early pregnancy (11-13 weeks), with an AUC of (0.856-0.964, 95% CI) and an accuracy of (0.8217-0.8821, 95% CI). Compared with the existing technology level, this method has a higher predictive ability. By accurately predicting GDM in early pregnancy, medical professionals and pregnant women can take early intervention measures to reduce the risk of pregnancy complications. This technology not only performs well in terms of accuracy, but also focuses on cost-effectiveness, and therefore has the potential for practical application. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0027] Figure 1 A flowchart showing the prediction of early pregnancy GDM by cfDNA fragment omics is shown; and
[0028] Figure 2 The AUC effects of cfDNA fragmentomic data and clinical data on the prediction of early pregnancy GDM are shown. DETAILED DESCRIPTION
[0029] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0030] It should be noted that the terms "first," "second," and the like in the description and claims of the present invention are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.
[0031] To facilitate those skilled in the art to understand the present invention, some terms or nouns involved in the embodiments of the present invention are explained below:
[0032] Name explanation:
[0033] End Motifs: End Motifs, a few nucleotides at the end of DNA, usually refers to 4-6bp nucleotides. In this application, the end motif refers to the 4 bases at the double-end of cfDNA.
[0034] MDS: motif diversity score, terminal sequence diversity score.
[0035] CGN / NCG: The ratio of fragments starting with CGN to those starting with NCG.
[0036] Fetal concentration: The percentage of free DNA from the fetus in the peripheral blood of a pregnant woman.
[0037] TSS data: TSS stands for Transcription Start Site score.
[0038] Collective knowledge: A collection of published medical classification standards, patient clinical records, datasets collected from patient samples, etc.
[0039] In response to the technical problems in the prior art for classifying healthy people and people with gestational diabetes, the present invention has achieved significant results by screening characteristic parameters relative to GDM through cell-free DNA (cfDNA) fragment omics data. These characteristic parameters may include terminal motifs, MDS data, CGN / NCG, fetal concentration, TSS data, length of cfDNA fragments, etc. In addition, preferably, the present invention can also be combined with some basic and easily available clinical data, such as height, weight, age, smoking and drinking conditions. The application of the technical solution of the present invention can find significant differences related to GDM. By screening and analyzing these characteristic parameters, the accuracy of classifying healthy people and people with early gestational diabetes can be improved, which is helpful for early diagnosis and intervention of GDM, thereby improving pregnancy outcomes and preventing the occurrence of related complications. Among them, in a specific embodiment of the present invention, in order to process data and establish an accurate prediction model, the present invention adopts a specific neural network method. This neural network structure can construct a highly accurate GDM prediction model by learning and training data based on the differences between different features.
[0040] According to an exemplary embodiment of the present application, a method for classifying early pregnancy gestational diabetes mellitus is provided. The method comprises: obtaining a set of feature parameters of a sample to be tested, wherein the set of feature parameters is extracted from cfDNA fragmentomic data of gestational diabetes mellitus, and the set of feature parameters comprises at least one of the following: end motif, MDS data, CGN / NCG, fetal concentration, and TSS data; inputting the feature parameters into an early pregnancy gestational diabetes mellitus prediction model, performing feature classification by the early pregnancy gestational diabetes mellitus prediction model, and outputting a classification result of the sample to be tested.
[0041] By using the technical solution of the present application, the early pregnancy gestational diabetes mellitus prediction model is constructed by using feature parameters related to gestational diabetes mellitus, and the healthy population and the gestational diabetes mellitus population can be accurately classified.
[0042] Further, the present application combines clinical information and cfDNA fragmentomic data, and uses a neural network model to construct an early pregnancy gestational diabetes mellitus prediction model. In specific experimental data, it is shown that accurate gestational diabetes mellitus (GDM) prediction can be made in early pregnancy (11-13 weeks), with an AUC of (0.856-0.964, 95% CI) and an accuracy of (0.8217-0.8821, 95% CI). Compared with the prior art, the method has high prediction ability. By accurately predicting GDM in early pregnancy, medical professionals and pregnant women can take early intervention measures to reduce the risk of pregnancy complications. This technology not only performs well in accuracy, but also focuses on cost-effectiveness, and therefore has potential for practical application.
[0043] The beneficial effects of the present application will be further illustrated in conjunction with the examples below.
[0044] Example 1
[0045] The present embodiment provides a method for classifying early pregnancy gestational diabetes mellitus. The method comprises: obtaining a set of feature parameters of a sample to be tested, wherein the set of feature parameters is extracted from cfDNA fragmentomic data of gestational diabetes mellitus, and the set of feature parameters comprises at least one of the following: end motif, MDS data, CGN / NCG, fetal concentration, and TSS data; inputting the feature parameters into an early pregnancy gestational diabetes mellitus prediction model, performing feature classification by the early pregnancy gestational diabetes mellitus prediction model, and outputting a classification result of the sample to be tested.
[0046] In some preferred embodiments, the model for predicting gestational diabetes in early pregnancy includes: a first classification layer, and a second classification layer connected to the first classification layer, the first classification layer includes at least one, two, three or four of the following: a first neural network layer, a second neural network layer, a third neural network layer, a fourth neural network layer and a fifth neural network layer, and the types of the first classification layer network include a fully connected neural network, a neural network and a convolutional neural network.
[0047] In some preferred embodiments, feature parameters are input into a model for predicting gestational diabetes in early pregnancy, and the model for predicting gestational diabetes in early pregnancy performs feature classification, and outputs classification results of the sample to be tested, including: using the first neural network layer to process the terminal motif to obtain first intermediate data, and the first neural network layer is a fully connected neural network; using the second neural network layer to process CGN / NCG to obtain second intermediate data, and the second neural network layer is a neural network; using the third neural network layer to process fetal concentration to obtain third intermediate data, and the third neural network layer is a neural network; using the fourth neural network layer to process MDS data to obtain fourth intermediate data, and the fourth neural network layer is a neural network; using the fifth neural network layer to process TSS data to obtain fifth intermediate data, and the fifth neural network layer is a convolutional neural network; using the second classification layer to process the intermediate data output by the first classification layer to obtain a classification result, and the intermediate data output by the first classification layer includes the first intermediate data, the second intermediate data, the third intermediate data, the fourth intermediate data, and the fifth intermediate data.
[0048] In some preferred embodiments, the first neural network layer is a fully connected neural network for processing terminal motifs; the second neural network layer is a neural network for processing CGN / NCG; the third neural network layer is a neural network for processing fetal concentration; the fourth neural network layer is a neural network for processing MDS data; the fifth neural network layer is a convolutional neural network for processing TSS data; and the second classification layer is a fully connected neural network.
[0049] In some preferred embodiments, when obtaining MDS data, the method includes scoring the diversity of terminal motifs according to the occurrence probability Pi of a terminal motif. For example, in a specific embodiment, the calculation formula 1 is used to calculate and obtain the MDS data:
[0050] In some preferred embodiments, obtaining CGN / NCG includes: performing ratio analysis using fragments starting with CGN and fragments starting with NCG; and / or, obtaining fetal concentration includes obtaining fetal concentration by estimating the fetal concentration by calculating the ratio of partial bins, for example, using the SeqFF method to obtain fetal concentration.
[0051] In some preferred embodiments, generating the TSS data comprises: analyzing coverage of the transcription start site region; preferably, analyzing coverage of the transcription start site region comprises: defining the transcription start site region: selecting a transcription start site of a gene of interest, and determining a fixed value of the transcription start site; calculating coverage: analyzing the aligned sequencing data file using a depth function to calculate the number of times sequencing reads cover the transcription start site region to obtain the coverage, for example, in a specific embodiment, using the SAMtools depth function to analyze the aligned BAM file to calculate the coverage of the transcription start site region; normalizing the coverage: normalizing the coverage of the transcription start site region to generate the TSS data.
[0052] In some preferred embodiments, the early-pregnancy prediction model for gestational diabetes mellitus is pre-constructed, and in constructing the early-pregnancy prediction model for gestational diabetes mellitus, the steps include: obtaining cfDNA fragmentomic data of pregnant women with gestational diabetes mellitus and normal pregnant women; analyzing the cfDNA fragmentomic data to screen characteristic parameters related to gestational diabetes mellitus; inputting the screened characteristic parameters as training set samples into a neural network model machine for training to obtain the early-pregnancy prediction model for gestational diabetes mellitus.
[0053] Further, screening the characteristic parameters related to gestational diabetes mellitus includes: starting from an empty logical model, taking a forward feature selection step to optimize the early-pregnancy prediction model for gestational diabetes mellitus; using likelihood ratio chi-square test to evaluate the adjustment of the characteristic parameters to the early-pregnancy prediction model for gestational diabetes mellitus, wherein the characteristic parameters with the smallest chi-square test p value (for example: p<0.05) are selected into the early-pregnancy prediction model for gestational diabetes mellitus as the characteristic parameters related to gestational diabetes mellitus.
[0054] In some preferred embodiments, the pregnant women with gestational diabetes mellitus and the normal pregnant women are in the 12th week of pregnancy.
[0055] In some preferred embodiments, the classification method further comprises clinical information analysis, and the clinical information includes one or more of height, weight, age, BMI, smoking, and alcohol consumption.
[0056] Embodiment 2
[0057] In this embodiment, by collecting cell-free DNA (cfDNA) in the maternal peripheral blood plasma and analyzing and extracting the characteristics of cfDNA fragmentomics, early risk prediction markers of GDM can be screened. Based on the differences of these markers, a specific neural network is constructed to process the data, thereby establishing a high-accuracy GDM prediction model.
[0058] This embodiment mainly includes the following steps, such as Figure 1 As shown:
[0059] Step 1: Maternal peripheral blood was collected from pregnant women with GDM and healthy women between 11 and 16 weeks of gestation (depending on sample availability) and immediately stored at 4°C. Plasma was then separated within 8 hours. Plasma was immediately stored at -80°C until further processing.
[0060] It is worth noting that the sample used in this embodiment is plasma, but other body fluid samples such as amniotic fluid, serum, urine, etc. can also be used in other embodiments.
[0061] Step 2: cfDNA sequencing.
[0062] cfDNA was extracted from 200 μl of plasma using the MagPure Circulating DNA Mini KF Kit (Magen). This kit uses a specific chemistry and centrifugation step to separate cfDNA from plasma. The extracted cfDNA was used for library preparation using the MGIEasy DNA Library Preparation Kit (MGI). The prepared cfDNA library was quantified using the Qubit dsDNA Kit (Invitrogen) to determine its concentration and quality. The library was circularized using the MGIEasy Circularization Kit (MGI) to generate single-stranded DNA (ssDNA) circles. The purified ssDNA circles were quantified using the Qubit ssDNA Assay Kit (Invitrogen) to assess their concentration and quality. The ssDNA circles were amplified by rolling circle amplification (RCA) to generate DNA nanospheres (DNBs). After RCA and DNB generation, the final products were quantified using the Qubit ssDNA Assay Kit (Invitrogen) and then loaded onto the DNBSEQ platform (MGI) for multiplex sequencing using a paired-end 100-base-pair strategy.
[0063] Step 3: Fragmentomics data analysis of cfDNA.
[0064] Fetal fraction was analyzed using the SeqFF method.
[0065] The end motif and end sequence diversity score (MDS) are determined by the sequence of the last four bases in the sequencing data. The MDS algorithm refers to Formula 1:
[0066]
[0067] Where Pi is the probability of occurrence of a motif.
[0068] Methylation analysis (MA) refers to the method of CGN / NCG (Q. Zhou, et. al. Epigenetic analysis of cell-free DNA by fragmentomic profiling, Proc. Natl. Acad. Sci. U.S.A. 119(44) (2022) e2209852119) which analyzes the ratio of fragments starting with CGN and NCG in cfDNA fragments.
[0069] The calculation method of the Transcription Start Site score (TSS score) is as follows:
[0070] Define the TSS region: select the transcription start site of all mRNA genes and determine the region 500bp upstream and downstream of the TSS;
[0071] Calculate the coverage: use the SAMtools depth function to analyze the aligned BAM file and calculate the coverage of the TSS region (coverage represents how many sequencing reads cover the position in this region);
[0072] Standardize the coverage: considering the influence of sequencing depth and sequencing bias, standardize the coverage of the TSS region, for example, divide a 1kb region into three parts and then calculate the coverage of each part.
[0073] Step 4: Parameter screening.
[0074] According to preliminary experimental results, most of the features of the TSS gene contain more noise and cannot have an effective impact on the final result. Therefore, the invention takes the step of forward feature selection to optimize the model. In the process of feature selection, start with an empty logical model and gradually add features to the model to find the features with the best performance for each step. The feature selection method is based on the method of C Hu et al. (Hu, C., Liu, Y., Lu, Z., Zhao, S., Han, X., & Xiong, J. (2021). Smartphone location spoofing attack in wireless networks. In Security and Privacy in Communication Networks: 17th EAI International Conference, SecureComm 2021, Virtual Event, September 6-9, 2021, Proceedings, Part II 17 (pp. 295-313). Springer International Publishing), as shown in Formula 2:
[0075]
[0076] Where e represents the natural constant and y' represents the probability of GDM positivity.
[0077] In the initial stage, there are no components in the β and f vectors. Starting from an empty model, a feature is gradually added to it, and the selection is made by comparing the performance improvement of the model with the previous model. In this embodiment, the best performing likelihood ratio chi-square test is used to evaluate the performance improvement, where the feature with the smallest chi-square test p value (p < 0.05) will be selected into the model. Then, the best performing feature is selected for the next step. Forward selection continues until there are no qualified features that can further improve the model, that is, the chi-square test p value is less than 0.05. The selected features will be used to construct a feature-specific neural network model (SNN) model.
[0078] Step 5: Build a feature-specific neural network model (SNN).
[0079] The structure of SNN is as follows Figure 1As shown in Figure 3, in the model, different features of cfDNA are divided into multiple groups, representing different aspects. Because each type of feature has different patterns, in this example, a feature-specific neural network (SNN) layer is set to process the data for final classification. In this example, a fully connected neural network is constructed using the dense layer in Keras to process the terminal motifs, as shown in Formula 3:
[0080] L=[Relu(W m […[Relu(W j [Relu(W i ·x i +b i )]+b k )]…]+b n )] (Formula 3)
[0081] L represents the output of the fully connected layer; W i represents the weight matrix; X i represents the input matrix of each layer; b i represents the input bias of the first layer; b k Indicates the bias of the next layer; b n Represents the input bias of the last layer.
[0082] Use W i , W j ,...W m To represent the weight matrix for layers i, j, ..., each layer processes the output of the previous layer and generates the output of the next layer using the ReLU function. This is because terminal motif features are typically uniform and have similar values; fully connected neural networks can achieve the best classification models. For MDS and CGN / NCG, a simple single neuron is used for processing, as shown in Equation 4:
[0083]
[0084] M represents the output result of MDS or CGN / NCG, and b represents the bias.
[0085] Here, x is the input matrix for MDS or CGN / NCG, and w is an n×1 weight matrix. Here, three CGN / NCG features are used as an example. This is because these features are relatively simple and contain less information. Therefore, only one neuron can be output to the classification layer. Next, a convolutional layer is used to process the TSS score, as shown in Equation 5:
[0086]
[0087] R represents the output result of TSS; j represents the x-axis coordinate of the convolution kernel; k represents the y-axis coordinate of the convolution kernel; m-j represents the current convolution x coordinate, and n-k represents the current convolution y coordinate.
[0088] This is based on the observation that there are many intrinsic relationships between different genes in the TSS score matrix. The convolution layer learns and emphasizes the relationship between two genes, thereby helping the final classification task.
[0089] In summary, the design of the feature-specific neural network (SNN) layer enables the present technical solution to process feature vectors containing different feature groups (e.g. 93 dimensions in the present embodiment), making full use of the unique features of each feature group to improve classification performance.
[0090] Step 6: Classification and prediction of GDM.
[0091] After processing the input data using the specific neural network, the output of each subnetwork is connected to form a single input matrix. This connected input matrix is then used for classification and regression tasks. By combining the outputs of different subnetworks, the model can utilize the information learned from each subnetwork and make predictions or classifications. The calculation of classification is shown in equation 6:
[0092] class = softmax(p i ) = softmax([LMMR]·W c +b c ) (Equation 6)
[0093] p i represents the probability of whether it is GDM, LMR represents the output layer combined input layer of L, M, and R described above, W c represents the weight matrix, and b c represents the bias.
[0094] In the present embodiment, a fully connected layer with a Softmax activation function is used, and a cross-entropy loss function is used for classification output. Since BMI has no negative output, the regression output adopts a Relu activation function and weights. The loss function of the regression output is the mean square error, and the area under the ROC curve (AUC) is selected as the main evaluation index of the SNN.
[0095] Example 3
[0096] A total of 299 GDM samples and 299 healthy controls were recruited in the hospital for analysis, as shown in Table 1.
[0097] Table 1
[0098]
[0099] The method in Example 2 was used:
[0100] The dataset was randomly shuffled and split into training and test sets, with 299 samples for GDM and 299 samples for control. The split was done in 75% training and 25% test. The testing procedure included 10-fold cross-validation. Thus, 10 models were trained on the training data and tested on the test data 10 times, resulting in 10 AUCs with 95% confidence intervals. To this end, the inventors generated the final input by the features obtained from the parameter screening method on the early pregnancy samples, resulting in a 88-feature vector, including 31 End Motifs (31 End Motifs refer to 31 types of end motifs, corresponding to a 31-dimensional vector, where each dimension represents whether the sample contains such end motif or the content of such end motif in the cfDNA of the sample, each dimension represents the content of a 4bp fragment), 4 MDS (in this example, according to the fragment length, three subgroups are divided: short: <150bp peak: 160-170bp, long: >250bp, plus all (all fragments), a sample has 4 values), 3 CGN / NCG (including the measurement value of CGN, the measurement value of NCG and the ratio between them CGN / NCG), 1 fetal concentration and 50 TSS scores (in this example, the activities of the transcription start points of 50 genes of interest correspond to 50 features) plus 4 clinical indicators age, BMI, smoking and drinking to model the incidence of GDM, resulting in a model with an accuracy of 88.21% and AUC = 0.91 as shown in Figure 2 Using only clinical information can achieve AUC = 0.73, while all fragmentomics plus clinical data can achieve 0.91. It can be seen that the use of cfDNA fragmentomics data plus clinical information integrated together can significantly improve the prediction effect of the model.
[0101] Example 4
[0102] The present embodiment provides a classification device for early pregnancy gestational diabetes. The classification device comprises: an acquisition unit configured to acquire a feature parameter set of a to-be-tested sample, wherein the feature parameter set is extracted from cfDNA fragmentomics data of gestational diabetes, and the feature parameter set comprises at least one of the following: end motif, MDS data, CGN / NCG, fetal concentration and TSS data; a classification unit configured to input the feature parameters into an early pregnancy prediction model for gestational diabetes, perform feature classification by the early pregnancy prediction model for gestational diabetes, and output a classification result of the to-be-tested sample.
[0103] In some preferred embodiments, the model for predicting gestational diabetes in early pregnancy includes: a first classification layer, and a second classification layer connected to the first classification layer, the first classification layer includes at least one of the following: a first neural network layer, a second neural network layer, a third neural network layer, a fourth neural network layer and a fifth neural network layer, and the type of the first classification layer network includes a fully connected neural network, a neural network and a convolutional neural network.
[0104] In some preferred embodiments, feature parameters are input into a model for predicting gestational diabetes in early pregnancy, and the model for predicting gestational diabetes in early pregnancy performs feature classification, and outputs classification results of the sample to be tested, including: using the first neural network layer to process the terminal motif to obtain first intermediate data, and the first neural network layer is a fully connected neural network; using the second neural network layer to process CGN / NCG to obtain second intermediate data, and the second neural network layer is a neural network; using the third neural network layer to process fetal concentration to obtain third intermediate data, and the third neural network layer is a neural network; using the fourth neural network layer to process MDS data to obtain fourth intermediate data, and the fourth neural network layer is a neural network; using the fifth neural network layer to process TSS data to obtain fifth intermediate data, and the fifth neural network layer is a convolutional neural network; using the second classification layer to process the intermediate data output by the first classification layer to obtain a classification result, and the intermediate data output by the first classification layer includes the first intermediate data, the second intermediate data, the third intermediate data, the fourth intermediate data, and the fifth intermediate data.
[0105] In some preferred embodiments, the first neural network layer is a fully connected neural network for processing terminal motifs; the second neural network layer is a neural network for processing CGN / NCG; the third neural network layer is a neural network for processing fetal concentration; the fourth neural network layer is a neural network for processing MDS data; the fifth neural network layer is a convolutional neural network for processing TSS data; and the second classification layer is a fully connected neural network.
[0106] In some preferred embodiments, when obtaining MDS data, the method includes scoring the diversity of terminal motifs according to the occurrence probability Pi of a terminal motif. For example, in a specific embodiment, the calculation formula 1 is used to calculate and obtain the MDS:
[0107] In some preferred embodiments, obtaining CGN / NCG includes: performing ratio analysis using fragments starting with CGN and fragments starting with NCG; and / or, obtaining fetal concentration includes obtaining fetal concentration by estimating the fetal concentration by calculating the ratio of partial bins, for example, using the SeqFF method to obtain fetal concentration.
[0108] In some preferred embodiments, generating the TSS data comprises: analyzing the coverage of the TSS region; preferably, analyzing the coverage of the TSS region comprises: defining the TSS region: selecting the TSS of the gene of interest, and determining a fixed value of the TSS; calculating the coverage: using a depth function to analyze the aligned sequencing data file, calculating the number of times the sequencing reads cover the TSS region, obtaining the coverage, for example, in a specific embodiment, using the SAMtools depth function to analyze the aligned BAM file, calculating the coverage of the TSS region; normalizing the coverage: normalizing the coverage of the TSS region to generate the TSS data.
[0109] Embodiment 5
[0110] The embodiment provides a method for constructing a model for predicting gestational diabetes in early pregnancy. The method comprises: obtaining cfDNA fragmentomic data of pregnant women with gestational diabetes and normal pregnant women; analyzing the fragmentomic data of the cfDNA to screen characteristic parameters related to gestational diabetes; inputting the screened characteristic parameters as training set samples into a neural network model machine for training to obtain a model for predicting gestational diabetes in early pregnancy.
[0111] In some preferred embodiments, the cfDNA fragmentomic data comprises at least one of the following: end motif, MDS data, CGN / NCG, fetal concentration and TSS data.
[0112] In some preferred embodiments, screening the characteristic parameters related to gestational diabetes comprises: starting from an empty logical model, taking a forward feature selection step to optimize the model for predicting gestational diabetes in early pregnancy; using likelihood ratio chi-square test to evaluate the adjustment of the characteristic parameters to the model for predicting gestational diabetes in early pregnancy, wherein the characteristic parameters with the smallest chi-square test p value (for example: p<0.05) are selected into the model for predicting gestational diabetes in early pregnancy as the characteristic parameters related to gestational diabetes.
[0113] In some preferred embodiments, when obtaining the MDS data, it comprises: scoring the diversity of end motifs according to the occurrence probability Pi of each end motif, for example, in a specific embodiment, the calculation formula 1 is used to calculate the MDS:
[0114] In some preferred embodiments, obtaining the CGN / NCG comprises: using the ratio analysis of fragments starting with CGN and fragments starting with NCG; and / or, obtaining the fetal concentration comprises using the method of estimating fetal concentration by calculating the proportion of partial bins to obtain fetal concentration, for example, using the SeqFF method to obtain fetal concentration.
[0115] In some preferred embodiments, generating the TSS data comprises: analyzing the coverage of the transcription start site region; preferably, analyzing the coverage of the transcription start site region comprises: defining the transcription start site region: selecting the transcription start site of the gene of interest, and determining a fixed value of the transcription start site; calculating the coverage: using a depth function to analyze the aligned sequencing data file, calculating the number of times the sequencing reads cover the transcription start site region to obtain the coverage, for example, in a specific embodiment, using the SAMtools depth function to analyze the aligned BAM file to calculate the coverage of the transcription start site region; normalizing the coverage: normalizing the coverage of the transcription start site region to generate the TSS data.
[0116] In some preferred embodiments, the early pregnancy prediction model for gestational diabetes is pre-constructed, and in constructing the early pregnancy prediction model for gestational diabetes, the steps include: obtaining cfDNA fragmentomic data of pregnant women with gestational diabetes and normal pregnant women; analyzing the fragmentomic data of cfDNA to screen feature parameters related to gestational diabetes; inputting the screened feature parameters as training set samples into a neural network model machine for training to obtain the early pregnancy prediction model for gestational diabetes.
[0117] Further, screening the feature parameters related to gestational diabetes includes: starting from an empty logical model, taking a forward feature selection step to optimize the early pregnancy prediction model for gestational diabetes; using likelihood ratio chi-square test to evaluate the adjustment of the feature parameters to the early pregnancy prediction model for gestational diabetes, wherein the feature parameters with the smallest chi-square test p value (for example: p<0.05) are selected into the early pregnancy prediction model for gestational diabetes as the feature parameters related to gestational diabetes.
[0118] In some preferred embodiments, the pregnant women with gestational diabetes and the normal pregnant women are in the 12th week of pregnancy.
[0119] In some preferred embodiments, the classification method further comprises clinical information analysis, and the clinical information includes one or more of height, weight, age, BMI, smoking, and alcohol consumption.
[0120] Embodiment 6
[0121] This embodiment provides a device for constructing a model for predicting gestational diabetes in early pregnancy. The device includes: an information acquisition module configured to obtain cfDNA fragment omics data and clinical information from pregnant women with gestational diabetes and normal pregnant women; a data analysis and feature screening module configured to analyze the cfDNA fragment omics data and screen for characteristic parameters related to gestational diabetes; and a machine learning module configured to use the characteristic parameters screened in the data analysis and feature screening module as input features of a training set sample to perform machine learning training on a neural network model, thereby obtaining a model for predicting gestational diabetes in early pregnancy.
[0122] In some preferred embodiments, the analysis of the fragment omics data of cfDNA in the data analysis and feature screening module includes one or more of: terminal motifs, MDS data, CGN / NCG, fetal concentration and TSS data.
[0123] Preferably, the data analysis and feature screening module is configured to perform fetal concentration analysis including analysis using the SeqFF method; and / or the terminal sequence diversity score is calculated using Formula 1: Wherein Pi is the probability of occurrence of a certain motif; and / or methylation analysis includes analysis using the ratio of fragments starting with CGN and starting with NCG; and / or transcription start site score includes analysis of the coverage of the transcription start site region; preferably, analysis of the coverage of the transcription start site region includes: defining the transcription start site region: selecting the transcription start site of the target gene, and determining the fixed value of the transcription start site; calculating the coverage: using the SAMtools depth function to analyze the aligned BAM file, and calculate the coverage of the transcription start site region, the coverage of the transcription start site region indicates how many sequencing reads cover the position in the region; standardizing the coverage: standardizing the coverage of the transcription start site region.
[0124] In some preferred embodiments, the feature parameters screened for gestational diabetes in the data analysis and feature screening module are set as follows: starting from an empty logistic model, taking a forward feature selection step to optimize the model; using a likelihood ratio chi-square test to evaluate the improvement of model performance by the feature parameters, wherein the feature parameter with the smallest chi-square test p value (p<0.05) is selected into the model as the feature parameter associated with gestational diabetes.
[0125] In some preferred embodiments, the machine learning training of the neural network model in the machine learning module is set to include: different features of cfDNA are divided into multiple groups, the neural network layer of specific features is set to process the data to obtain the optimal parameters, and the final classification is performed; preferably, a fully connected neural network is constructed using a fully connected layer to process the terminal motif, each layer will process the output of the previous layer and generate the output of the next layer; neurons are used to process the terminal motif and the terminal sequence diversity score (MDS) and output to the classification layer; neurons are used to process methylation analysis and output to the classification layer; and a convolutional layer is used to process the transcription start site score.
[0126] In some preferred embodiments, the device also includes a classification module, and the classification module is configured to classify and predict gestational diabetes using an early pregnancy prediction model for gestational diabetes; preferably, the classification module is configured to process the input data using a neural network layer, and the output of each subnetwork is connected to form a single input matrix, and the input matrix is then used for classification and regression tasks. By combining the outputs of different subnetworks, the early pregnancy prediction model for gestational diabetes utilizes information learned from each subnetwork and performs classification based on collective knowledge.
[0127] In some preferred embodiments, the clinical information includes one or more of age, BMI, smoking and drinking; preferably, the pregnant women with gestational diabetes and the normal pregnant women are in the 12th week of pregnancy.
[0128] Example 7
[0129] This embodiment provides an electronic device comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any of the above-mentioned methods for classifying early pregnancy gestational diabetes by executing the executable instructions.
[0130] This embodiment provides a computer-readable storage medium including a stored computer program, wherein when the computer program is executed, the device containing the computer-readable storage medium is controlled to execute any of the above-mentioned methods for classifying early pregnancy gestational diabetes mellitus.
[0131] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0132] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0133] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, centralized on a single computing device, or distributed across a network of multiple computing devices. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0134] As can be seen from the above description, the above-described embodiments of the present invention achieve the following technical effects: The method has high predictive capabilities. By accurately predicting GDM in early pregnancy, medical professionals and pregnant women can take early intervention measures to reduce the risk of pregnancy complications. This technology not only demonstrates excellent accuracy but is also cost-effective, thus possessing potential for practical application.
[0135] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for classifying gestational diabetes mellitus in early pregnancy, characterized in that: The classification method includes: Obtaining a characteristic parameter set of the sample to be tested, wherein the characteristic parameter set is obtained by extracting cfDNA fragment omics data of gestational diabetes, and the characteristic parameter set includes at least one of the following: terminal motif, MDS data, CGN / NCG, fetal concentration, and TSS data; The characteristic parameters are input into a model for predicting gestational diabetes in early pregnancy, and the model for predicting gestational diabetes in early pregnancy performs characteristic classification, and outputs a classification result of the sample to be tested.
2. The classification method according to claim 1, characterized in that The early pregnancy gestational diabetes prediction model includes: a first classification layer, and a second classification layer connected to the first classification layer, the first classification layer includes at least one of the following: a first neural network layer, a second neural network layer, a third neural network layer, a fourth neural network layer and a fifth neural network layer, and the type of the first classification layer network includes a fully connected neural network, a neural network and a convolutional neural network.
3. The classification method according to claim 2, characterized in that Inputting the characteristic parameters into a model for predicting gestational diabetes in early pregnancy, performing feature classification by the model for predicting gestational diabetes in early pregnancy, and outputting the classification result of the sample to be tested includes: Processing the terminal motif using the first neural network layer to obtain first intermediate data, wherein the first neural network layer is a fully connected neural network; Processing the CGN / NCG using the second neural network layer to obtain second intermediate data, wherein the second neural network layer is a neural network; Processing the fetal concentration using the third neural network layer to obtain third intermediate data, wherein the third neural network layer is a neuronal network; Processing the MDS data using the fourth neural network layer to obtain fourth intermediate data, wherein the fourth neural network layer is a neural network; Processing the TSS data using the fifth neural network layer to obtain fifth intermediate data, wherein the fifth neural network layer is a convolutional neural network; The second classification layer is used to process the intermediate data output by the first classification layer to obtain the classification result, and the intermediate data output by the first classification layer includes the first intermediate data, the second intermediate data, the third intermediate data, the fourth intermediate data and the fifth intermediate data.
4. The classification method according to claim 2, characterized in that: The first neural network layer is a fully connected neural network, Used to process the terminal motif; the second neural network layer is a neural network, used to process the CGN / NCG; the third neural network layer is a neural network, used to process the fetal concentration; the fourth neural network layer is a neural network, used to process the MDS data; the fifth neural network layer is a convolutional neural network, used to process the TSS data; the second classification layer is a fully connected neural network.
5. The classification method according to claim 1, characterized in that: When obtaining the MDS data, the method includes scoring the diversity of the terminal motifs according to the occurrence probability Pi of a certain terminal motif.
6. The classification method according to claim 1, characterized in that The acquisition of CGN / NCG includes: performing ratio analysis using fragments beginning with CGN and fragments beginning with NCG; and / or, The acquisition of the fetal concentration includes obtaining the fetal concentration by estimating the fetal concentration by calculating the proportion of the partial sequencing window.
7. The classification method according to claim 1, characterized in that Generating the TSS data includes: analyzing the coverage of the transcription start site region; Preferably, the analysis of the coverage of the transcription start site region comprises: Define the transcription start site region: select the transcription start site of the target gene and determine the fixed value of the transcription start site; Calculate coverage: Use the depth function to analyze the aligned sequencing data files, calculate the number of times the sequencing reads cover the transcription start site region, and obtain the coverage; Normalizing coverage: Normalizing the coverage of the transcription start site region to generate the TSS data.
8. The classification method according to claim 1, characterized in that: The classification method also includes clinical information analysis, and the clinical information includes one or more of height, weight, age, BMI, smoking and drinking.
9. A classification device for early pregnancy gestational diabetes, characterized in that: The classification device comprises: an acquisition unit configured to acquire a characteristic parameter set of the sample to be tested, wherein the characteristic parameter set is obtained by extracting cfDNA fragment omics data of gestational diabetes, and the characteristic parameter set includes at least one of the following: terminal motif, MDS data, CGN / NCG, fetal concentration, and TSS data; The classification unit is configured to input the characteristic parameters into the early pregnancy prediction model for gestational diabetes mellitus, perform characteristic classification by the early pregnancy prediction model, and output the classification result of the sample to be tested.
10. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the method for classifying early pregnancy gestational diabetes according to any one of claims 1 to 8 by executing the executable instructions.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the method for classifying early pregnancy gestational diabetes according to any one of claims 1 to 8.