Method for predicting time of death based on stacking integrated method to integrate multi-omics data

By integrating multi-omics data through the Stacking ensemble model, the subjectivity and inaccuracy of traditional time of death estimation are resolved, achieving higher prediction accuracy and robustness. In particular, the application of the multi-omics joint model provides a more reliable method for estimating time of death in forensic medicine.

CN115862738BActive Publication Date: 2026-04-24SHANXI MEDICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANXI MEDICAL UNIV
Filing Date
2022-12-31
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Traditional methods for estimating time of death are subjective and empirical. Existing biomarker screening is not stable enough, resulting in inaccurate time of death prediction. Current technologies are difficult to effectively combine multi-omics data for accurate time of death prediction.

Method used

Stacking ensemble models are used to integrate multi-omics data, including metabolomics, protein chips, and infrared spectroscopy. By screening the optimal basic model and performing correlation analysis, single-omics and multi-omics stacking ensemble models are constructed. Combined with machine learning algorithms such as Adaboost, Logistic Regression, and Random Forest, multi-omics stacking ensemble models are built to improve prediction accuracy.

Benefits of technology

It improves the accuracy and generalization ability of time of death prediction. The prediction accuracy of the multi-omics stacking ensemble model reaches 0.93 and the AUC value reaches 0.98, which is significantly better than the single-omics model and has stronger robustness and generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115862738B_ABST
    Figure CN115862738B_ABST
Patent Text Reader

Abstract

The present application relates to the field of forensic science, and specifically relates to a death time prediction method based on a Stacking integrated method and comprehensive multi-omics data, comprising the following steps: collecting rat skeletal muscle samples, using metabolomics, protein chip and infrared spectrum detection technology to extract the expression amount of related biomarkers in the tissue; inputting the expression amount data of the biomarkers into a plurality of basic models respectively to deduce the death time, and screening out a single-omics optimal basic model with the best death time prediction performance; screening out two basic models with the lowest correlation with the single-omics optimal basic model to jointly construct a single-omics Stacking model; and connecting the above single-omics Stacking integrated model in series to construct a multi-omics integrated model. The present application provides a new method and new idea for predicting death time by combining multi-omics multi-molecular markers, and lays a foundation for applying a multi-omics joint machine learning model to death time deduction practice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of forensic medicine, specifically a method for predicting time of death based on a stacking ensemble approach that integrates multi-omics data. Background Technology

[0002] Postmortem time (PMI) refers to the time between the discovery and examination of a body and the occurrence of death. Accurately estimating the PMI is crucial for determining the time of crime and solving cases. Because the PMI is influenced by numerous factors, traditional methods of estimation, such as early postmortem phenomena, are subjective and empirical. Although current research shows that small molecules such as metabolites, RNA, and proteins have potential applications in estimating PMI, screening for more stable biomarkers and developing more accurate estimation methods remain significant bottlenecks for forensic scientists.

[0003] In recent years, with the development of molecular biology techniques, high-throughput analytical methods such as metabolomics, protein chips, and infrared spectroscopy have been used to detect the degradation patterns of biomarkers in the body after death, providing a basis for estimating the time of death. Different omics technologies have yielded a wealth of data types, and how to screen biomarkers related to the time of death from massive amounts of high-dimensional data has become a core aspect of current time-of-death prediction. Furthermore, because the biological processes after death are extremely complex, relying on a single omics detection method only describes the complex molecular biological changes after death in a limited dimension. Utilizing multi-omics technologies can accurately capture changes at different material levels after death, making the estimation of the time of death more accurate and universal.

[0004] With the development of big data and artificial intelligence, machine learning, relying on its strong analytical capabilities and fast computation speed, can provide a more reliable and stable prediction method for time of death estimation. Furthermore, the correct application of stacking ensemble models offers a new approach to multi-group collaboration. Stacking ensemble models integrate multiple base models, and during the integration process, the base models need to meet the construction requirement of being "good but different." Stacking ensemble models combine the strengths of each base model, complementing their weaknesses to achieve better generalization and robustness. Therefore, stacking ensemble models can integrate different omics features and reveal the complex biological changes after death as a whole, thereby improving the accuracy of time of death prediction. Summary of the Invention

[0005] This invention provides a Stacking ensemble model that integrates multiple molecular markers from multiple omics, aiming to provide a prediction model for time of death estimation that is highly accurate, efficient, and has strong generalization ability and robustness. Specifically, it is a time of death prediction method based on the Stacking ensemble method that integrates multi-omics data.

[0006] This invention is achieved through the following technical solution: a method for predicting time of death based on multi-omics data using the Stacking ensemble method, comprising the following steps:

[0007] 1) Collect rat skeletal muscle samples at different time points of death, and use metabolomics, protein chips and infrared spectroscopy to extract the expression levels of relevant biomarkers in the tissues;

[0008] 2) Input the expression data of biomarkers detected by the three omics into multiple basic models to infer the time of death. Select the best basic model of the single omics with the best performance in predicting the time of death for each of the three omics.

[0009] 3) To meet the diverse construction requirements of the Stacking ensemble model, correlation analysis was performed on the six best-performing basic models in each omics, and the two basic models with the lowest correlation to the best basic model of a single omics were selected to jointly construct a single omics Stacking model.

[0010] 4) Concatenate the above single-omics stacking ensemble models to construct a multi-omics ensemble model;

[0011] 5) Repeat step 1) and input the expression data of biomarkers from the three omics tests of the unknown rat skeletal muscle sample into the multi-omics ensemble model for time of death prediction.

[0012] As a further improvement to the technical solution of the present invention, in step 2), the number of the various basic models is eight, namely Adaboost, Logistic Regression, Random Forest, Multilayer Perceptron, Support Vector Machine, Gradient Boosting Tree, Stochastic Gradient Descent, and LightGBM.

[0013] As a further improvement to the technical solution of the present invention, in step 3), the correlation analysis utilizes the Pearson correlation coefficient.

[0014] As a further improvement to the technical solution of the present invention, the multi-omics stacking ensemble model in step 4) is constructed using the pipeline method.

[0015] This invention establishes a multi-omics stacking ensemble model for predicting time of death based on multi-omics and multi-molecular biological expression profiling combined with machine learning. It employs a progressive construction strategy that builds an optimal single-omics model, a single-omics stacking ensemble model, and a multi-omics stacking ensemble model through mutual correlation, effectively improving the robustness and generalization ability of time of death prediction. This invention provides a novel method and approach for predicting time of death using multi-omics and multi-molecular markers, laying the foundation for applying multi-omics joint machine learning models to the practice of time of death prediction. Attached Figure Description

[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 The expression levels of biomarkers detected by the three omics methods are compared with the accuracy and AUC values ​​of eight basic learners. Figure 1 A is a comparison diagram of basic metabolomics models. Figure 1 B is a comparison diagram of basic protein chip models. Figure 1 C is a comparison diagram of basic infrared spectral models.

[0018] Figure 2 Heatmaps related to the stacking ensemble model of metabolomics.

[0019] Figure 3 This is a confusion matrix diagram of a metabolomics stacking ensemble model.

[0020] Figure 4 ROC plot for a metabolomics stacking ensemble model.

[0021] Figure 5 Heatmaps related to the protein chip stacking integration model.

[0022] Figure 6 This is a confusion matrix diagram of the protein chip stacking ensemble model.

[0023] Figure 7 The ROC plot is for the protein chip stacking ensemble model.

[0024] Figure 8 The relevant heatmap for the infrared spectroscopy stacking ensemble model.

[0025] Figure 9 This is a confusion matrix diagram of the infrared spectroscopy stacking ensemble model.

[0026] Figure 10 The ROC plot is for the infrared spectroscopy stacking ensemble model.

[0027] Figure 11 A diagram illustrating the construction and evaluation of a multi-omics stacking ensemble model. Figure 11 A is the ROC plot of the multi-omics stacking ensemble model. Figure 11 B is the confusion matrix diagram of the multi-omics stacking ensemble model.

[0028] Figure 12 This is a flowchart of the death time prediction method based on the Stacking ensemble method that integrates multi-omics data, as described in this invention.

[0029] Figure 13 This is a flowchart of the construction process for a multi-omics stacking ensemble model. Detailed Implementation

[0030] The technical solution of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0031] Example

[0032] 1. Animal experiment grouping:

[0033] The rat samples were healthy adult male Sprague-Dewley rats (10-12 weeks old, weighing 250-300 g), provided by the Experimental Animal Center of Shanxi Medical University, animal license number SCXK[Jin][2009-0001). Male and female rats were housed separately in cages containing standard food and water. All rats were randomly divided into 14 groups, including a control group (0d, n=8) and 13 experimental groups (1, 2, 3, 5, 7, 9, 12, 15, 18, 21, 24, 27, 30d, n=8). All rats were anesthetized by intraperitoneal injection of 3% sodium pentobarbital (0.13 mL / 100g) and then euthanized by cervical dislocation. They were then placed in a climate chamber with a temperature of 16±2℃, humidity of 50±5%, and a 12-hour light / dark cycle. After skeletal muscle samples were harvested from rats at specified time points, they were frozen in liquid nitrogen, wrapped in aluminum foil, and placed in a -80°C freezer for subsequent analysis. The animal carcasses were kept intact before sampling, and to meet the needs of forensic work, no duplicate sampling was performed on any rat. To prevent interference from other drugs, all skeletal muscle samples were not treated with antibacterial agents or any other preservatives. All rat samples were randomly divided into two groups. The first group (98 samples, n=7) was used to build a machine learning model. The dataset was divided into training and validation sets in a 7:3 ratio. The training set was used to optimize the model's hyperparameters. Then, the model was selected based on the performance of the validation set. Finally, the second group (14 samples, n=1) formed the test set to test the model's generalization ability.

[0034] 2. Metabolomics workflow for detecting biomarkers:

[0035] Rat skeletal muscle tissue was slowly thawed on ice, and 200 mg ± 5 mg of tissue was weighed and placed in an EP tube containing 800 μL of cold acetonitrile (-4℃). Two zirconia beads were added to the EP tube, and then the tube was placed in an MM400 ball mill and vortexed 5 times at 30 rpm for 30 s. The EP tube was then vortexed for 20 seconds and placed on ice for 10 min. The solution was centrifuged at 13000 rpm and 4℃ for 30 min in a SIGMA2-16PK centrifuge, and 400 μL of the supernatant was extracted and freeze-dried for 180 min in a centrifugal concentrate freeze dryer. Finally, 200 μL of acetonitrile-water (4 / 1) mixture was added to the freeze-dried precipitate to dissolve it, vortexed for 1 min, centrifuged at 13000 rpm for 30 min at 4℃, filtered through a 0.22 μm PVDF membrane, and added to a sample vial for UPLC-HRMS analysis.

[0036] Ultra-high performance liquid chromatography (UHPLC) was performed via a heated electrospray ionization source connected to a mass spectrometer, with a mass scan range of m / z 80–1200 Da. The chromatographic column used was an ACQUITY UPLC. TMThe HSST3 column was used at 45℃ with an injection volume of 5 μL. Mobile phase A was 0.1% formic acid aqueous solution, and mobile phase B was 0.1% formic acid acetonitrile solution, with a flow rate of 0.3 mL / min. Gradient elution program: ESI was used to acquire positive and negative ions, with positive and negative spray voltages of 3.0 kV and 2.7 kV, respectively. The capillary and heater were set at 320℃ and 300℃, respectively, with a gas flow rate of 11 L / min. The nebulization pressure was 40 psi, the sheath gas temperature was 325℃, the sheath gas flow rate was 11 L / min, and the capillary voltage was 4000 V. Mass spectrometry data acquisition used Centroid mode bar graphs at a acquisition rate of 1.4 spectrum / s, with full scan / dd-MS2 mode. The raw data were imported into CompoundDiscoverer 3.0 software for preprocessing, including peak identification and alignment. The obtained data were identified based on the mass spectra and the standard database mzcloud. Small molecule metabolites were identified based on retention time and mass spectra. Metabolites mainly include amino acids, fatty acids, amides, aldehydes, nucleotides, ketones, acetyl derivatives, oil-based derivatives, and other substances.

[0037] 3. Protein chip workflow for detecting biomarkers:

[0038] 200 mg of muscle tissue was placed in a centrifuge tube, and pure water at a ratio of 1:3.5 (w / v) and the corresponding proportion of protease inhibitor PMSF were added. Two grinding beads were added to the tube, and the mixture was homogenized three times in a ball mill for 30 seconds each time. The mixture was incubated on ice for 1 hour, centrifuged at 12000×g for 15 minutes, and 500 μL of the supernatant was transferred to a new microcentrifuge tube. Gel preparation and chip sample preparation were performed according to the Agilent Protein 230 kit instructions. 4 μL of sample and 2 μL of denaturant were added to a 500 μL LEP tube, and the mixture was thoroughly mixed. The sample solution and ladder were heated in a 95°C water bath for 5 minutes, then rapidly cooled. 84 μL of deionized water was added for further dilution, and 6 μL of the diluted solution was loaded into the corresponding wells of the protein chip. Before chip operation, the metal probes of the bioanalyst were cleaned with deionized water, and a hardware short-circuit test was performed. Protein expression profile data were obtained using Agilent 2100 Expert software. Chromatographic peaks were calibrated and identified using internal standards (lower and upper markers), and peak positions were corrected and adjusted. Protein peaks with fluorescence intensities below 10 FU were removed to eliminate interference from impurities. Peaks with a migration time error of ±0.2 s were labeled as peaks of the same protein.

[0039] 4. Infrared spectroscopy workflow for detecting biomarkers:

[0040] Pre-cool the cryostat to -20°C, sterilize the scalpel with alcohol, install the scalpel, adjust the sectioning angle, and adjust the section thickness to 8μm. Remove the tissue frozen at -80°C and place it on ice. Cut a 1cm × 1cm × 1cm section, place it on a sample holder, fix it with OCT embedding medium, and freeze and flatten it on the freezing stage. Place the cooled sample on the sample head, smooth the sample surface, and after continuous sectioning to reveal complete tissue, take 5 specimens and place them on a CaF2 slide. Dry the CaF2 slide with a hairdryer for later examination. Fill the MCT detector with liquid N2 and cool for 10 minutes. Set the infrared microscope to visible light mode and select transmission mode. Place the CaF2 slide on the microscope stage and adjust the focus. Open the built-in OPUS software, load the measurement parameters, and set the scanning range to 4000cm. -1 -900cm -1 4cm resolution -1 The scan was repeated 32 times. Images of the selected area were acquired in visible light mode under a microscope, then the background spectrum of the blank area on the CaF2 slide was acquired in infrared mode. Three selected points on the sample were then scanned sequentially. To reduce error, the three repeatedly acquired spectra for each sample were automatically averaged into one spectrum. UnscramblerX10.4 software was used for preprocessing the data, including multiplicative scatter correction (MSC), baseline correction, and second derivative. The substances detected by infrared spectroscopy mainly included lipids (1800 cm⁻¹). -1 -1700cm -1 ), Amide I (1700cm) -1 -1500cm -1 ), Amide II and III (1350cm) -1 -1200cm -1 ) as well as nucleic acids and carbohydrates (1200cm -1 -900cm -1 )wait.

[0041] 5. Construct a single-axis optimal basic model

[0042] Expression data of biomarkers from three omics detection methods were input into eight basic learners for time-of-death inference. The eight basic models are Adaboost, Logistic Regression, Random Forest (RF), Multilayer Perceptron Classifier (MLPC), Support Vector Machine (SVM), Gradient Boosted Decision Tree (GBDT), Stochastic Gradient Descent (SGD), and Lightgbm (Light Gradient Boosting Machine).

[0043] (1) Optimal basic model of metabolomics

[0044] from Figure 1 As shown in Figure A, the accuracies of the eight basic models are 0.63, 0.66, 0.66, 0.60, 0.66, 0.70, 0.60, and 0.50, respectively, and the AUC values ​​are 0.92, 0.90, 0.92, 0.93, 0.92, 0.93, 0.72, and 0.90, respectively. Considering the overall predictive performance of each model, the gradient boosting tree model, with an accuracy of 0.72 and an AUC value of 0.93, is considered the optimal basic model for metabolomics-based time of death estimation.

[0045] (2) Optimal Basic Model of Protein Chip

[0046] like Figure 1 As shown in B, unlike metabolomics, logistic regression became the optimal basic model for protein chips because the logistic regression model had a prediction accuracy of 0.80 and an AUC of 0.98, making it the best predictive performance among the eight models.

[0047] (3) Optimal basic model of infrared spectroscopy

[0048] like Figure 1 As shown in Figure C, the multilayer perceptron model has the highest prediction accuracy, with a prediction accuracy of 0.80 and an AUC of 0.95. In summary, the optimal basic model under the three omics frameworks only achieves a prediction accuracy of around 0.80, which is insufficient to meet the practical needs of forensic work. Therefore, it is necessary to construct more advanced models to improve prediction accuracy.

[0049] 6. Construct a single-array stacking ensemble model:

[0050] (1) Metabolomics Stacking Integration Model

[0051] From the relevant heat map ( Figure 2 As can be seen, logistic regression, multilayer perceptron, and gradient boosting tree (GBP) are the least correlated in metabolomics, with correlations of 0.53 and 0.44, respectively. Therefore, logistic regression, multilayer perceptron, and GBP are used as sub-models in the metabolomics stacking ensemble model. The confusion matrix of the metabolomics stacking ensemble model (…) Figure 3 ) and ROC plot ( Figure 4 The results showed that the prediction accuracy of 0.73 and the AUC of 0.95 were significantly higher than the optimal basic model for metabolomics.

[0052] (2) Protein chip stacking ensemble model

[0053] like Figure 5 As shown, the Adaboost model and gradient boosting tree have the weakest correlation with the optimal base model logistic regression for protein chips. Therefore, Logistic regression, Adaboost, and gradient boosting tree are used as sub-models in the protein chip stacking ensemble model. The results of the protein chip stacking ensemble model show that the prediction accuracy and AUC value are 0.83 (…). Figure 6 ) and 0.98 ( Figure 7 Although the AUC is the same as the optimal basic model for protein chips, the prediction accuracy is improved.

[0054] (3) Infrared Spectroscopy Stacking Integrated Model

[0055] Such as correlation heatmap ( Figure 8 The results show that Gradient Boosting Tree (0.11) and LightGBM (0.33) have the weakest correlation with the optimal base model for infrared spectroscopy, the Multilayer Perceptron. Therefore, the Multilayer Perceptron, Gradient Boosting Tree, and LightGBM are considered as sub-models of the Stacking ensemble model for infrared spectroscopy. Figure 9 and Figure 10 It can be seen that the prediction accuracy of the infrared spectroscopy stacking ensemble model (0.83) is better than that of the optimal basic model, while the AUC value is the same as that of the optimal basic model.

[0056] (4) Multi-omics Stacking Integration Model

[0057] Even after constructing a single-omics stacking ensemble model, the prediction accuracy of time of death is still less than 90%. Therefore, based on the above-mentioned construction of a single-omics stacking ensemble model, a multi-omics stacking ensemble model was further established.

[0058] Based on the above, we constructed the optimal basic models of metabolomics, protein chips, and infrared spectroscopy, and established a stacking ensemble model of metabolomics, protein chips, and infrared spectroscopy using correlation analysis. On this basis, we used the pipeline function in Python to connect the stacking models of the three omics in series, and constructed a multi-omics stacking ensemble model with the three stacking ensemble models of metabolomics, protein chips, and infrared spectroscopy as sub-models and stochastic gradient descent (SCD) as the meta-model.

[0059] (5) Evaluation of single-omics and multi-omics stacking ensemble models

[0060] To evaluate the generalization ability and robustness of the single-omics optimal base model, the single-omics stacking ensemble model, and the multi-omics stacking ensemble model, we used test sets to test the estimation performance of each model. One sample from each time point (1, 2, 3, 5, 7, 9, 12, 15, 18, 21, 24, 27, 30 days, n=1) from the control group (0dn=1) and the experimental group were used as the test set. These test set samples were then placed into the three previously constructed single-omics optimal base models, three single-omics stacking ensemble models, and the multi-omics stacking ensemble model. After fitting calculations, the predicted values ​​of each sample in the test set were compared with the actual values ​​to calculate the accuracy and AUC. The generalization ability and robustness of each model were evaluated based on the accuracy and AUC values. The results (Table 1) show that the multi-omics ensemble model had the best predictive performance. Only one sample in the test set was mispredicted, with a prediction accuracy of 0.93 and the highest AUC of 0.98. Figure 11 Compared to models built from single omics, the multi-omics stacking ensemble model not only combines biomarkers detected from three omics levels but also integrates multiple different machine learning models. The results show that the multi-omics stacking ensemble model has better generalization ability and robustness, laying the foundation for multi-omics combined with multi-molecular biomarkers and joint machine learning to infer time of death. Table 1: Comparison of predictive performance of the optimal single-omics base model, the single-omics ensemble model, and the multi-omics stacking ensemble model.

[0061]

[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting time of death based on stacking ensemble methods integrating multi-omics data, characterized in that: Includes the following steps: 1) Collect rat skeletal muscle samples at different time points of death, and use metabolomics, protein chips and infrared spectroscopy to determine the expression levels of relevant biomarkers in the tissues; The specific methods for detecting biomarkers in the metabolomics workflow are as follows: After slowly thawing rat skeletal muscle tissue on ice, 200 mg ± 5 mg of tissue was weighed and placed in an EP tube containing 800 μL of acetonitrile at -4°C. Two zirconia beads were added to the EP tube, and then the tube was placed in an MM400 ball mill and oscillated at 30 times / s × 30 s for 5 times. The EP tube was then vortexed for 20 seconds and placed on ice for 10 min. The solution was centrifuged at 13000 rpm at 4°C for 30 min in a SIGMA2-16PK centrifuge, and 400 μL of the supernatant was extracted and freeze-dried for 180 min in a centrifugal concentrate freeze dryer. Finally, 200 μL of the acetonitrile-water mixture was added to the freeze-dried precipitate to dissolve it, oscillated for 1 min, centrifuged at 13000 rpm at 4°C for 30 min, filtered through a 0.22 μm PVDF membrane, and added to a sample vial for UPLC-HRMS analysis. Ultra-high performance liquid chromatography (UHPLC) was performed by connecting a heated electrospray ionization source to a mass spectrometer, with a mass scan range of m / z 80-1200 Da. The chromatographic column used was an ACQUITYUPLC™ HSST3 column, with a column temperature of 45℃ and an injection volume of 5 μL. Mobile phase A was 0.1% formic acid aqueous solution, and mobile phase B was 0.1% formic acid acetonitrile solution, with a flow rate of 0.3 mL / min. ESI was used to acquire positive and negative ions, with positive and negative spray voltages of 3.0 kV and 2.7 kV, respectively; the capillary and heater were set at 320 °C and 300 °C, respectively, and the gas flow rate was 11 L / min; the nebulization pressure was 40 psi, the sheath gas temperature was 325 °C, the sheath gas flow rate was 11 L / min, and the capillary voltage was 4000 V. Mass spectrometry data acquisition used Centroid mode bar graphs at a acquisition rate of 1.4 spectrum / s, with full scan and dd-MS2 scan modes. The raw data were imported into CompoundDiscoverer 3.0 software for chromatographic peak identification and peak alignment data preprocessing. The obtained data were identified based on the mass spectra and the standard database mzcloud, and small molecule metabolites were identified based on retention time and mass spectra. The specific method for detecting biomarkers using protein chip workflow is as follows: 200 mg of muscle tissue was placed in a centrifuge tube, and pure water and protease inhibitor PMSF were added at a 1:3.5 mass-to-volume ratio. Two grinding beads were added to the tube, and the mixture was homogenized three times in a ball mill for 30 seconds each time. The mixture was incubated on ice for 1 hour, centrifuged at 12000×g for 15 minutes, and 500 μL of the supernatant was transferred to a new microcentrifuge tube. Gel preparation and chip sample preparation were performed according to the instructions of the Agilent Protein 230 kit. 4 μL of sample and 2 μL of denaturant were added to a 500 μL LEP tube and mixed thoroughly. The mixed sample solution and ladder were heated in a 95°C water bath for 5 minutes, then rapidly cooled. 84 μL of deionized water was added for dilution, and 6 μL of the diluted solution was loaded into the corresponding wells of the protein chip. Before chip operation, the metal probes of the bioanalyst were cleaned with deionized water, and a hardware short-circuit test was performed. Protein expression profile data of the samples were obtained using Agilent 2100 Expert software. According to the internal standard lower... The marker and upper marker are used to calibrate and identify chromatographic peaks and correct and adjust their positions; to eliminate interference from impurities, protein peaks with fluorescence intensity below 10 FU are removed; peaks with a migration time error of ±0.2 s are marked as peaks of the same protein. The specific method for detecting biomarkers using infrared spectroscopy workflow is as follows: Pre-cool the cryostat to -20°C, sterilize the scalpel with alcohol, install the scalpel, adjust the sectioning angle, and adjust the section thickness to 8μm. Remove the tissue frozen at -80°C and place it on ice. Cut a 1cm × 1cm × 1cm section, place it on a sample holder, fix it with OCT embedding medium, and freeze and flatten it on the freezing stage. Place the cooled sample on the sample head, smooth the sample surface, and after continuous sectioning to reveal complete tissue, take 5 specimens and place them on a CaF2 slide. Dry the CaF2 slide with a hairdryer until ready for examination. Fill the MCT detector with liquid N2 and cool for 10 minutes. Set the infrared microscope to visible light mode and select transmission mode. Place the CaF2 slide on the microscope stage and adjust the focus. Open the machine's built-in OPUS software, load the measurement parameters, and set the scanning range to 4000cm. -1 -900cm -1 4cm resolution -1 The scan was repeated 32 times. Images of the selected area were acquired in visible light mode of the microscope, and then the background spectrum of the blank area of ​​the CaF2 slide was acquired in infrared mode. Then, the three selected points of the sample were scanned one by one. To reduce errors, the three spectra of each sample were automatically averaged into one spectrum. The data were preprocessed using UnscramblerX10.4 software for multivariate scattering correction, baseline correction and second derivative. 2) The expression data of biomarkers detected by the three omics were input into multiple basic models for time of death inference. The best basic model of the single omics with the best performance in predicting time of death was selected for each of the three omics. The number of multiple basic models was eight, namely Adaboost, Logistic Regression, Random Forest, Multilayer Perceptron, Support Vector Machine, Gradient Boosting Tree, Stochastic Gradient Descent, and LightGBM. 3) To meet the diverse construction requirements of Stacking ensemble models, correlation analysis was performed on the six best-performing basic models in each omics, and the two basic models with the lowest correlation to the best basic model of a single omics were selected to jointly construct a single omics Stacking ensemble model with the best basic model. 4) Concatenate the above single-omics stacking ensemble models to construct a multi-omics stacking ensemble model; 5) Repeat step 1) Input the expression data of biomarkers from the three omics tests of the unknown rat skeletal muscle sample into the multi-omics stacking ensemble model to predict the time of death.

2. The death time prediction method based on the Stacking ensemble method integrating multi-omics data as described in claim 1, characterized in that, In step 3), the correlation analysis uses the Pearson correlation coefficient.

3. The death time prediction method based on the Stacking ensemble method integrating multi-omics data as described in claim 1, characterized in that, In step 4), the multi-omics stacking ensemble model is constructed using the pipeline approach.

Citation Information

Patent Citations

  • Establishment method of severe spinal cord injury prognosis prediction model

    CN112992346A

  • KR1024524330000B1