Time sequence modeling and scoring method for dynamic weight and cross-species data and gene combination

By using a time-series modeling method with dynamic weights and cross-species data, the limitations of data integration and detection methods in the stroke process are addressed, enabling efficient integration and dynamic monitoring of multi-source data, which is suitable for stroke detection in primary healthcare institutions.

CN121171345APending Publication Date: 2025-12-19HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511301129.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing technologies lack cross-organizational integration, dynamic staging, and cross-species migration capabilities in the stroke process. Furthermore, the detection methods have high barriers to entry and insufficient dynamic monitoring capabilities, making them difficult to promote in primary healthcare institutions.

Method used

Employing a dynamic weighting and cross-species data time-series modeling approach, this study integrates multi-source datasets, constructs cross-organizational causal associations and dynamic networks, screens time-specific genes, and trains multi-class time-series models. It then builds cross-species homologous gene mapping and dynamic weight calculation to achieve the integration and scoring of multi-time-point, multi-organization, and multi-species data.

Benefits of technology

It achieves efficient integration of multi-source omics data, improves the systematicness and integrity of the model, enhances the sensitivity and prediction accuracy of pathological evolution trends, reduces detection costs, is suitable for application on grassroots platforms, and has cross-platform comparison capabilities and wide applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121171345A_ABST
    Figure CN121171345A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic weight and cross-species data time sequence modeling and scoring method and a gene combination, and belongs to the technical field of biological information processing and data modeling. According to the technical scheme, on the basis of multi-source group data, through causal network construction, time sequence interpolation and multi-stage model training, feature genes closely related to the stroke process are screened out, the dynamic weight function is constructed, and a scoring system used for staging analysis is formed; according to the method, cross-species gene mapping and dynamic data fusion can be carried out in combination with mouse models and human data, scoring results can correspond to multiple key stages of stroke, and the method has the advantages of being high in expression stability, clear in dynamic trend, good in model interpretation and the like, meanwhile has good universality and is suitable for popularization and application. The method can be expanded and applied to other complex immune process modeling tasks, and has high adaptability, interpretability and popularization potential.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bio-information processing and data modeling, in particular to dynamic weight and cross-species data time series modeling, scoring method and gene combination. BACKGROUND

[0002] In the research process of nervous system related events, the process involved by the immune system has high time sequence and tissue specificity, especially in acute brain injury or ischemia related models, the immune response shows complex dynamic characteristics. With the development of omics technology, researchers have been able to obtain large-scale, multi-time point, multi-tissue source immune related data through RNA-seq, single cell transcriptome (scRNA-seq), spatial transcriptome and qPCR, etc. to provide a rich data basis for studying complex biological processes.

[0003] However, there are still many problems to be solved in the current analysis practice of modeling and quantitative scoring of biological system immune processes, especially in the integration analysis of multi-tissue data related to the time dimension.

[0004] First of all, most public researches often only select a specific single tissue (such as brain tissue, peripheral blood, etc.) as the analysis object, and fail to integrate the time sequence data of key immune cell storage tissues such as skull bone marrow, so it is difficult to fully depict the dynamic migration path of immune cells among multiple tissues and their functional changes. This "local analysis" at the tissue level limits the understanding of the overall behavior of the biological system. More importantly, although research shows that skull bone marrow is an important source of myeloid immune cells in acute pathological states, its signaling pathways and differentiation trajectories are of great reference value for time series modeling, but due to its special location, difficulty in sample collection, and dependence on tissue sampling for technical means, it causes high collection cost and inconvenient clinical operation problems.

[0005] Therefore, these data have high biological accuracy, but lack generalizability, especially difficult to apply in grassroots or dynamic monitoring scenarios.

[0006] In addition, in clinical practice, high-end imaging equipment such as MRI has high reliability in judging whether a stroke has occurred and assessing the degree of stroke, and is often used as a structural assessment standard. However, such equipment has high procurement cost and high use threshold, and the judgment method still relies heavily on the experience interpretation of professional imaging physicians. More importantly, when observing pathological changes in the process of stroke, the conventional imaging method has limitations such as radiation accumulation and long operation period, making it difficult to meet the needs of high frequency, non-invasive and flexible continuous monitoring. In grassroots medical institutions, the popularization rate of such equipment is less than 30%, making it difficult to popularize related methods.

[0007] Some existing molecular detection products attempt to identify stroke through 1-2 blood markers, but their detection targets are usually limited to "whether stroke", rather than dynamic staging or trend modeling; and they rely on laboratory environment operation, and the detection samples are often invasive body fluids such as cerebrospinal fluid, which not only has operation difficulty, but also is easily affected by non-stroke inflammation, and lacks specificity. The detection cost is also high, usually about hundreds of yuan per time.

[0008] Therefore, the current mainstream detection and evaluation technology has the following common technical defects:

[0009] 1) The judgment target is single, and traditional imaging and some biomarker products can only judge whether stroke has occurred, and it is difficult to dynamically stage the disease course or identify the stage trend;

[0010] 2) The test index is narrow, the imaging method relies on structural information, and the biomarker method usually uses a single or two gene indicators, which cannot fully reflect the overall pattern of dynamic expression of the immune system;

[0011] 3) The judgment method relies on subjective or static experience, the imaging result needs to be interpreted by artificial, the molecular method lacks systematic modeling, and only relies on threshold judgment, which is easily affected by noise;

[0012] 4) The dynamic monitoring capability is insufficient, the traditional method is difficult to sample multiple times at low cost, the imaging has radiation exposure, and the molecular product is mostly one-time judgment;

[0013] 5) The popularization threshold is high, the imaging equipment is expensive and the popularization rate is low in primary level; the biomarker product often needs specific laboratory conditions and manual operation;

[0014] 6) The applicable range is limited, some schemes need to collect special body fluids such as cerebrospinal fluid, which has the risk of invasion, and may be interfered by other systemic inflammatory reactions.

[0015] In summary, there is currently a lack of a data processing method that can be used for stroke process modeling and analysis, which can simultaneously have the characteristics of cross-organizational integration, multi-gene feature combination, dynamic staging capability, cross-species migration capability, and easy actual collection and primary adaptation. Therefore, the present application constructs a dynamic weight and cross-species data time series modeling, scoring method and gene combination technical scheme from the perspective of biological data modeling, which can integrate and establish a scoring model for multi-time point, multi-organ, and multi-species data, and is particularly suitable for exploring the feature changes, stage identification, and expression trend quantification of dynamic immune processes. SUMMARY

[0016] The application aims to provide a dynamic weight and cross-species data time series modeling, scoring method and gene combination, which solves the pain points of existing stroke immune staging technology from four aspects of mechanism, data, algorithm and application, has biological explanation, clinical operability and cross-species universality, and is expected to be a bridge for stroke precise diagnosis and treatment from "empirical medicine" to "systematic medicine", and promote the intelligent upgrading of related detection methods, scoring systems and diagnosis and treatment schemes.

[0017] To achieve the above-mentioned purpose, the application provides the following technical scheme: a dynamic weight and cross-species data time series modeling method, comprising the following steps:

[0018] Step A: data acquisition and preprocessing

[0019] A1, induction and construction of multi-source data set, in this step, the required original data source, data format and experimental design are comprehensively combed and archived to ensure the consistency and traceability of cross-species data. Specifically, it includes identification, batch information extraction and preliminary quality evaluation of data in different platforms (such as GEO, ArrayExpress, qPCR laboratory data, etc.), so as to facilitate subsequent conversion and standardization processing, which can include RNA-seq, single cell transcriptome (scRNA-seq) and real-time fluorescent quantitative PCR (qPCR) original data of different species

[0020] A2, standardization of scRNA-seq / RNA-seq and qPCR data, so as to ensure that different data types are integrated and analyzed under the same dimension;

[0021] A3, using brain transcriptome data in multi-source data to interpolate peripheral blood time series expression profile, arranging brain transcriptome data with high space-time resolution according to time points, and performing mathematical interpolation on gene expression values between limited observation points of peripheral blood, that is, interpolating and supplementing the missing time series in the data set, and the interpolation and supplementing method includes one or more combinations of linear interpolation, polynomial interpolation and cubic spline interpolation, wherein the linear interpolation is suitable for the case of smooth change of gene expression in short time interval, and the cubic spline interpolation constructs a C2 continuous smooth curve between all time points, and is more suitable for mutation points. According to the time expression characteristics of different genes, the above methods can be flexibly combined to ensure that the completed time series curve can reflect the real biological change and reduce the interpolation error.

[0022] Step B: cross-tissue causal correlation and dynamic network construction, this step aims to establish a space-time causal network revealing the gene expression between brain and peripheral blood. Through causal inference and co-expression analysis, a dynamic gene interaction network is constructed, which lays a foundation for discovering key regulatory genes and potential biomarkers.

[0023] B1, using a Bayesian network algorithm based on structural equation model (SEM), input the homogenization expression matrix and timestamp information of the multi-source dataset, learn the conditional dependency between gene nodes; filter high-confidence edges through Bootstrapping and BIC criteria to construct a brain-blood causal directed graph, that is, use the preprocessed multi-source dataset in the previous step to construct a Bayesian causal network to represent the brain-peripheral blood gene causal correlation;

[0024] B2, using peripheral blood gene expression as the dependent variable and brain expression at the corresponding time point as the independent variable, establishing a random forest regression model; according to the variable importance score (Mean Decrease in Impurity), filter out the gene subset that contributes most to blood expression prediction, integrate these genes into the dynamic network model, select brain-peripheral blood cross-tissue significant covariation genes based on random forest regression, and include them in the dynamic gene interaction network;

[0025] Step C: time-specific gene screening and initial gene screening, that is, within different staging time windows (such as 1h, 6h, 24h, 72h, 7d, etc.), through the combination of differential expression and statistical model, the gene candidate set that significantly changes over time and has biological significance is preliminarily locked,

[0026] C1, based on two types of machine learning models for feature selection to improve robustness and noise resistance, take the intersection or joint score of the gene importance ranking returned by each model to ensure that the screening results have both sensitivity and specificity, and specifically use random forest and XGBoost algorithms,

[0027] Screening candidate genes that meet the following conditions:

[0028] 1) differential expression fold (|log2FC|>1) and p<0.05 at the corresponding time point, this screening standard ensures that the screened genes have significant changes biologically;

[0029] 2) based on time series mixed effect model (LMM), extract genes that significantly change over time after stroke (p<0.05, |avg_log2FC|>0.3);

[0030] 3) significantly related to brain pathology and cell function state (such as infarct volume, inflammation pathway, neural repair, etc.), that is, through Spearman correlation or multiple regression analysis, ensure the consistency of biological characteristics and pathological performance;

[0031] 4) In non-stroke inflammation models (such as LPS induction), the expression amount has no significant change, specifically, the screening gene is verified in the LPS-induced inflammation mouse model or in vitro cell model, and the absolute value of log2FC is required to be <0.3 and p>0.1 in the model to exclude general inflammatory response genes and improve the accuracy of stroke-specific marker.

[0032] Step D: Multi-class time series model training and secondary reverse screening

[0033] D1, build an XGBoost multi-classification model, divide the samples into different period classes, for example, define four class labels (0-3);

[0034] D2, based on 5-fold stratified cross-validation, take AUC value as evaluation index, optimize model hyperparameters through grid search, use StratifiedKFold of scikit-learn to ensure that the proportion of each class sample in each fold is consistent; carry out grid search on the hyperparameter range such as learning rate, maximum depth, and sub-sample proportion, and take the average AUC value as the optimal model selection standard;

[0035] D3: Perform AUC value curve analysis on the candidate genes screened in step C, calculate the AUC value of each gene in the classification of different stages, and sort them from high to low according to AUC, and eliminate genes with low AUC. Finally, determine a number of AUC highest feature genes as input objects for subsequent dynamic weight calculation, for example, the top 20 genes in the model internal feature importance can be plotted one by one to draw single-factor ROC curves and calculate AUC, and genes with AUC>0.75 are selected as high-quality candidates, and genes with AUC<0.6 are eliminated to improve the signal-to-noise ratio of subsequent SHAP value calculation.

[0036] Step E: Cross-species homologous gene mapping

[0037] E1. Transform cross-species genes through a predefined dictionary. For common species such as humans, mice, and rats, construct a cross-species gene correspondence dictionary (such as based on Ensembl ID and NCBI HomoloGene), map the gene symbols of each species to a unified homologous gene name using the dictionary, and assign weights according to the quality of the homologous relationship (such as one-to-one and one-to-many mapping) to ensure the accuracy and comparability of gene function annotation in subsequent analysis.

[0038] E2. Load another species gene data while synchronously loading the pre-trained species model in steps A-D. The pre-trained model has learned relevant gene expression patterns (such as which genes are more important for stage prediction) from large-scale data, and these patterns can be passed to the cross-species gene data model (i.e. another species mentioned above) through model parameters, reducing the dependence of low-data-volume species on their own data volume.

[0039] E3. Process the cross-species data by the data standardization method in step A, unify the feature distribution, make the feature space of the cross-species data consistent with the space when the initial data model is trained, and ensure that the rules learned by the large data volume species model (such as "high expression of a certain gene corresponds to the hyperacute stage") can be directly migrated to the low data volume species,

[0040] E4. Construct the XGBoost multi-classification model of cross-species gene data according to the method in step D; step F: dynamic weight calculation and time sequence fusion

[0041] F1. Extract the SHAP value of each gene from the trained XGBoost model and perform cross-species SHAP value fusion; the SHAP value of the large data model contains the gene importance learned from the large data volume (such as S100a8 plays a significant role in mouse stroke staging), and after fusion, the feature weight of the low data model will take into account the prior knowledge of the large data volume model, especially when the model data is limited, which can improve the model's ability to identify key genes,

[0042] F2, determine the static weight w i , SHAP, provide the initial weight for the subsequent dynamic weight formula;

[0043] F3, calculate the dynamic weight of each gene according to the following formula:

[0044] w i (t)=w i SHAP·e -λit +α·Adapt i (t)

[0045] Where λi, β are adjustable parameters, and t represents the time unit of occurrence. Through grid search and exponential decay optimization of λi and β on the training set, the weight changes dynamically with time to reflect the contribution degree change of the gene in different stages.

[0046] Preferably, the multi-source data set at least includes GSE225948, GSE233815 of the mouse model, the human serum RNA-seq project CRA016580, and GSE122709.

[0047] Preferably, the spatial transcriptome data of GSE233815 is used to calculate the D1 / D3 / D7 peripheral blood gene expression profile by cubic spline interpolation method, that is, the blood expression is reconstructed by cubic spline interpolation between each block, and the expression dynamics of different regions and time are compared.

[0048] Preferably, in step A2, the scRNA-seq / RNA-seq data is first subjected to log1p conversion to eliminate the influence of extreme values, and then subjected to Z-score standardization to make the expression of each gene in different samples comparable. For qPCR data, calculate the ΔCt value (target gene Ct minus internal reference gene Ct), and then perform Z-score standardization to eliminate experimental batch effects and ensure that different data types are integrated and analyzed on the same scale, which is beneficial to the convergence of machine learning models.

[0049] Preferably, in step F3, β is the adjustable learning rate, which is used to adapt to individual differences and batch effects of each sample in real time, with an initial value of 0.01, and exponentially decays with the number of iterations to prevent early loss of information due to excessive decay of weights.

[0050] Preferably, in step F3, λi is the decay coefficient, which is optimized by grid search and ranges from 0.05 to 3.

[0051] Preferably, in step F, the SHAP value fusion method across species is to take the weighted mean, which not only retains the reliable biological mechanism knowledge in the mouse model, but also combines the clinical reality of human data, so that the evaluation of gene importance is more robust and closer to the actual human scenario.

[0052] A dynamic weight and cross-species data scoring method, the scoring method comprising the following steps,

[0053] G1. Calculate the staging score: Score(t) = ∑(w i (t) × Δct i ),

[0054] G2. Map the staging score to the [0, 1] interval by the Softmax function,

[0055]

[0056] where score k is the staging score of the kth stage, and n is the total number of stages. According to the actual data scoring requirements, the specific stages and the probability of belonging to the stage can be determined.

[0057] A gene combination of a dynamic weight and cross-species data scoring method, the gene combination under qPCR detection means at least includes S100A8, CD14, ARG2, and GAPDH as an internal reference gene, i.e. calculating the relative amount by the ΔCt method, taking GAPDH as the internal reference, achieving high sensitivity and reproducibility.

[0058] A combination of dynamic weight and scoring method of cross-species data, the gene combination of human scRNA-seq / RNA-seq at least includes S100A8, CD14, ARG2, S100A9, IFITM1, ATF3, GM2A; the gene combination of mouse scRNA-seq / RNA-seq at least includes S100a8, Cd14, Arg2, Chil3, Gda, S100a9, Ifitm1, the above-mentioned genes are selected as multi-dimensional features, covering immune inflammation, cell repair and metabolic pathways, and the staging prediction ability is verified by single cell and population level expression data.

[0059] Compared with the prior art, the beneficial effects of the present application are: the present application is a dynamic weight and cross-species data time series modeling and scoring method and gene combination, and the beneficial effects and advantages are embodied in the following aspects:

[0060] I. Efficient integration of multi-source omics data, improving the breadth and depth of model input

[0061] The present application fully integrates RNA-seq, single-cell transcriptome (scRNA-seq), spatial transcriptome and qPCR and other multiple data types, covers multiple key immune tissues such as brain tissue, skull bone marrow and peripheral blood, establishes a unified data processing and standardization process, so that data from different species, different platforms and different tissue sources can be integrated into a unified modeling framework. Compared with the existing analysis method based on single tissue or single time point data, the present application significantly improves the systematicness and integrity of the input data, and provides a more rich and reliable feature basis for subsequent modeling.

[0062] II. Dynamic time series modeling mechanism, overcoming the defect of weak time-varying feature capture ability of static scoring method

[0063] Existing methods generally regard gene expression features as static attributes, and only construct classification models based on single time point data, which cannot reflect the dynamic evolution characteristics in the pathological process. Especially in the study of stroke and other highly time-dependent processes, this method has obvious limitations. The present application innovatively introduces a dynamic weight mechanism, defines the importance weight of gene expression as a function of time, and combines static SHAP value, adaptive gradient adjustment term and time decay factor to model and generate gene weight with time series characteristics at each time point. This mechanism significantly improves the sensitivity of the model to the pathological evolution trend, and avoids the problem that static models perform inconsistently in different time windows.

[0064] III. Introducing cross-species modeling mechanism, improving model generalization ability and verification efficiency

[0065] Traditional human samples have significant limitations in ethics and accessibility, and many mechanism researches rely more on animal models such as mice for simulation and exploration. However, there are great differences between mouse data and human data in expression mode, naming standard and other aspects, which leads to difficult model migration. The present application introduces a cross-species gene mapping dictionary, and through a transfer learning method fused with SHAP values, the feature weights extracted in the mouse model are fused with the human sample model in the modeling stage, and a cross-species shared scoring system is constructed, which effectively improves the stability and prediction accuracy of the model under the condition of insufficient samples. At the same time, this technical scheme can also be extended to other low-data species, providing ideas and data support for in-depth research.

[0066] Four, screening of feature gene combination with strong biological specificity and significant time trend, enhancing model interpretability and biological consistency

[0067] The present application proposes a complete feature gene screening process, covering multiple dimensions such as differential expression analysis, linear mixed effect model, functional pathway mapping, and non-specific inflammation filtering. On this basis, further introduce multi-classification model and SHAP analysis to evaluate the importance of genes, and then combine with ROC curve AUC value for secondary reverse screening, so as to obtain a few core feature genes with strong stability, high biological specificity and prominent time dependence. The finally recommended qPCR gene combination (S100A8, CD14, ARG2) and RNA-seq combination (including S100a8, Cd14, Arg2, Chil3, etc.) are verified to have good expression stability and phase discrimination ability in multiple public databases, which is convenient for subsequent scientific research or development and transformation.

[0068] Five, constructing scoring function and classification threshold, with high resolution stroke timing staging ability

[0069] By constructing a scoring function weighted by dynamic gene weight, the present application divides the four stages of "non-stroke", "stroke super acute phase", "acute phase" and "subacute phase" in the model output layer, and calculates the corresponding probability. This mechanism can convert continuous expression profile into explicit classification label, which is suitable for batch staging processing of high-throughput samples. At the same time, the staging model and the actual pathological evolution process show high consistency in multiple animal experiments, and has strong interpretability.

[0070] Six, highly adapt to qPCR and other basic platforms, reduce detection threshold and improve promotion ability

[0071] Considering that current high-throughput omics technologies have limited popularization in primary medical systems, the application provides a detection scheme based on a portable qPCR device, which can use peripheral blood for sample collection and complete the detection task with a fixed primer group. Compared with imaging, which relies on expensive equipment, and some molecular schemes that require invasive samples such as cerebrospinal fluid, the method has low cost, flexible detection, short cycle, and supports dynamic multiple monitoring, and is especially suitable for resource-limited areas and epidemiological survey projects with limited baseline conditions.

[0072] Seven, provide a unified scoring system with cross-platform and cross-center quantitative comparison capability

[0073] Since the standardized expression matrix and the unified scoring formula are used for model construction, the scoring system proposed by the application does not depend on specific sequencing platforms and can be used for horizontal comparison between different laboratories and different sample groups, which helps to construct a standardized staging scoring system. In the process of multi-center research, cross-project integration and platform migration, the scoring system can significantly improve the consistency of data and the reusability of results, and provide a theoretical basis for subsequent data standardization.

[0074] Eight, the method has wide application range and good expansibility

[0075] Although the application is designed and verified based on the research scene of stroke immune timing process, the dynamic modeling mechanism, weight adjustment method, cross-species migration method and scoring construction process used in the application have universality. For other complex biological problems involving inter-tissue interaction, immune evolution process or pathological staging, such as sepsis, trauma response, autoimmune diseases, etc., the scheme can also be extended or directly reused as an analysis framework, which has strong universal adaptability.

[0076] In summary, by introducing dynamic modeling, cross-species fusion and scoring system construction and other technical means, the application solves the outstanding problems in data integration, model adaptability, interpretability and scene adaptability in the prior art, has significant advantages in both theoretical innovation and engineering implementation, and has good scientific research value and popularization prospect. BRIEF DESCRIPTION OF DRAWINGS

[0077] Figure 1 Part of the code process for screening specific genes in the application;

[0078] Figure 2 q-PCR peripheral blood specific gene expression of the application;

[0079] Figure 3 q-PCR peripheral blood prediction result of the application;

[0080] Figure 4 Prediction result of clinical retrospective data of the application. Detailed Implementation

[0081] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0082] This application proposes a time-series modeling method based on dynamic weights and cross-species data, a scoring method, and a final set of highly expressed gene combinations. Specifically, for the time-series modeling method based on dynamic weights and cross-species data, this application takes stroke staging as an example, and the core multi-source datasets selected are shown in Table 1.

[0083] Table 1

[0084]

[0085] For scRNA-seq / RNA-seq data, log1p transformation was first used to eliminate the influence of extreme values, and then Z-score standardization was used to make the expression levels of each gene comparable in different samples. For qPCR data, ΔCt value (the difference between the Ct of the target gene and the Ct of the internal reference gene) was calculated, and then Z-score standardization was performed to eliminate the experimental batch effect and ensure that different data types were integrated and analyzed under the same scale.

[0086] Then, we can proceed to step A3, which involves interpolating brain transcriptome data from multi-source datasets to extrapolate the temporal expression profile of peripheral blood. This involves sorting the high spatiotemporal resolution brain transcriptome data by time points and performing mathematical interpolation on gene expression values ​​between a limited number of observation points in peripheral blood. Specifically, cubic spline interpolation is used to interpolate brain transcriptome data to extrapolate the temporal expression profile of peripheral blood, which can solve the problem of incomplete time points in the current samples.

[0087] Then, step B, namely the process of cross-organizational causal association and dynamic network construction, and step C, phase-specific gene screening and primary gene screening, are carried out to screen candidate genes that meet the following conditions: 1) the differential expression fold (|log2FC|>1) at the corresponding time point and p<0.05. This screening criterion ensures that the screened genes have significant biological changes.

[0088] 2) Based on the time series mixed effects model (LMM), genes that changed significantly over time after stroke were extracted (p<0.05, |avg_log2FC|>0.3);

[0089] 3) significantly related to brain pathological conditions and cellular functional status (such as infarction volume, inflammatory pathways, nerve repair, etc.), that is, through Spearman correlation or multiple regression analysis, to ensure the consistency of biological characteristics and pathological manifestations;

[0090] 4) In non-stroke inflammatory models (such as LPS induction), the expression level does not change significantly, specifically, the screening gene is verified in the LPS-induced inflammatory mouse model or in vitro cell model, and the absolute value of log2FC is required to be <0.3 and p>0.1 in the model, to exclude general inflammatory response genes and improve the accuracy of stroke-specific markers,

[0091] Specifically, part of the code process of the screening process is as shown in Figure 1 According to the above method process, 243 period-related genes are specifically screened out, and it can be found that the number of this gene combination is large, which is not conducive to inspection and application, and the verification standard of ROC AUC curve analysis also finds that its accuracy is low.

[0092] Therefore, we further perform a secondary screening process, that is, step D of the multi-class time series model training and secondary reverse screening. Specifically, an XGBoost multi-classification model is constructed, and samples are divided into different period classes, for example, four types of labels are defined, specifically, four periods of non-stroke, stroke super acute phase (0-1 day), acute phase (1-3 days), and subacute phase (3-7 days) are divided.

[0093] The method in step D2 is implemented, that is, based on 5-fold stratified cross-validation, taking AUC as the evaluation index, then AUC curve analysis is performed on the candidate genes screened in step C, the AUC value of each gene in the classification is calculated, and the genes with low AUC are removed, and finally a number of AUC highest feature genes are determined as the input object of subsequent dynamic weight calculation.

[0094] The SHAP value of each gene is extracted from the trained XGBoost model to determine the static weight w i , the SHAP value of each gene is extracted from the trained XGBoost model to determine the static weight w i (t)=w i , SHAP·e -λit +α·Adapt i (t), where β is the adjustable learning rate, the initial value is 0.01, which is exponentially decayed with the number of iterations, and λi is the decay coefficient, which is optimized by grid search, and its range is 0.05-3.

[0095] At the same time, the scoring formula and the guidance of the staging range are proposed, and the specific score is calculated by the following formula, that is, the staging score is Score(t)=∑(wi (t) x Δct i ), respectively, to obtain the scores of each partition, and then the original scores of the model are mapped to the interval [0, 1] through a Softmax function, and the formula is:

[0096] And the model established by the present application is verified internally, and the specific data is shown in Table 2.

[0097] Table 2 Internal verification of the model

[0098]

[0099] It can be seen that the overall AUC is greater than 0.9, indicating that the model performs well in distinguishing different categories and has high accuracy.

[0100] The specific experimental process of the verification is as follows:

[0101] Mouse p-MCAO model: C57BL / 6 mice (n=40), divided into sham operation group, 24h, 3d, 4d, 7d groups.

[0102] Detection method: peripheral blood qPCR detection + brain STT staining.

[0103] The specific q-PCR peripheral blood specific gene expression is as shown in Table 2, randomly selected data is brought into the scoring formula, and the specific data after calculation is as follows, and the prediction result is as shown in Table 3. Figure 2 Figure 3

[0104] ①Select the q-pcr data of peripheral blood of sham operation mice:

[0105] Mus GAPDH: 21.96; Mus S100A8: 24.16; Mus CD14: 34.18; Mus Arg2: 34.26

[0106] ②Select the q-pcr data of peripheral blood of mice on the first day of stroke:

[0107] Mus GAPDH: 21.74; Mus S100A8: 23.24; Mus CD14: 31.54; Mus Arg2: 32.99

[0108] ③Select the q-pcr data of peripheral blood of mice on the third day of stroke:

[0109] Mus GAPDH: 21.72; Mus S100A8: 20.19; Mus CD14: 30.10; Mus Arg2: 29.11

[0110] ​​IV. Selecting the q-pcr data of the peripheral blood of the rats on the 4th day after stroke:

[0111] Mus GAPDH: 21.66; Mus S100A8: 17.71; Mus CD14: 26.7; Mus Arg2: 28.51.

[0112] It is calculated that the overall AUC value of the model prediction result is 0.946, that is, the reliability of the technical solution proposed in the application is higher.

[0113] The advantage of the dynamic weight calculation adopted in the application is that the weight of the recovery period gene can be dynamically reduced according to the disease course stage, for example, the weight of Arg2 changes in the chronic period, and the weight distribution is ensured to conform to the biological mechanism, such as the high weight of Cd14 in the hyperacute period.

[0114] Verification is performed by using the public data GSE122709 (D1, D7), and the specific data process is as follows:

[0115] I. Selecting the non-stroke clinical RNA-seq data to be processed:

[0116] S100A8: 23403; CD14: 3109; ARG2: 18; ATF3: 115; GM2A: 758; S100A9: 26098; IFITM1: 3183

[0117] II. Selecting the clinical RNA-seq data to be processed on the 1st day after stroke:

[0118] S100A8: 42856; CD14: 13039; ARG2: 24; ATF3: 213; GM2A: 1472; S100A9: 28839; IFITM1: 2058

[0119] III. Selecting the clinical RNA-seq data to be processed on the 7th day after stroke:

[0120] S100A8: 43395; CD14: 9395; ARG2: 46; ATF3: 219; GM2A: 1231; S100A9: 55089; IFITM1: 4482

[0121] and the prediction result is as shown in Figure 4 At the same time, the AUC value of the model prediction result is 0.912, which has high reliability.

[0122] For specific data conversion, the scRNA-seq / RNA-seq data is subjected to log1p conversion and then Z-score standardization. For example, the original counts value is shown in Table 3,

[0123] Table 3 Real data snippet of raw counts values

[0124]

[0125] Preprocessing: log1p transformation (mitigate high dispersion):

[0126] log1p(S100a8) = log(1 + 133.77) ~ 4.89,

[0127] log1p(Cd14) = log(1 + 18.75) ~ 3.08, and the rest transformed to [0.58, 1.76, 1.56, 5.06, 2.10].

[0128] Normalization (Z-score): Adjusted features based on training set distribution:

[0129] [1.92, 1.65, -0.32, 0.41, -0.18, 2.05, 0.29].

[0130] And according to this result, the XGBoost model output is specifically:

[0131] Non-stroke (0): -2.1, Hyperacute phase (1): 4.2, Acute phase (2): 1.8, Subacute phase (3): -0.5.

[0132] Softmax probability calculation:

[0133] P(0) ~ 0.12 / 75.23 ~ 0.002, P(1) ~ 66.68 / 75.23 ~ 0.89

[0134] P(2) ~ 6.05 / 75.23 ~ 0.08, P(3) ~ 0.61 / 75.23 ~ 0.01

[0135] Result: Hyperacute phase probability 0.89, consistent with the sample true label (Brain_MCAO is a hyperacute phase model).

[0136] Similarly, specifically, after calculating the ΔCt value for the qPCR data, for example, the original Ct value is shown in Table 4 below,

[0137] Table 4, original Ct values

[0138]

[0139] The calculation process is as follows:

[0140] ΔCt calculation (target gene Ct - internal control gene Ct):

[0141] ACtS100A8 = 23.5670 - 21.7206 = 1.8464

[0142] ACtCD14 = 31.3952 - 21.7206 = 9.6746

[0143] ACtArg2 = 32.6291 - 21.7206 = 10.9085

[0144] Standardization (Z-score): Adjust the scale based on the mean (μ) and standard deviation (σ) of the training set:

[0145] Normalized features: [1.44, 0.89, 1.07],

[0146] Calculate gene weights based on SHAP values (feature importance) and time decay:

[0147] w i (t) = w i , SHAP·e -λit + α·Adapt i (t), where λ i is the decay coefficient

[0148] According to the input normalized features [1.44, 0.89, 1.07] and dynamic weights, the XGBoost model outputs are as follows:

[0149] Non-stroke (0): -2.3, Hyperacute phase (1): 3.9, Acute phase (2): 1.5, Subacute phase (3): -0.8, calculated by Softmax probability,

[0150] · The specific process is: p(0) = e^{-2.3} ≈ 0.10, p(1) = e^{3.9} ≈ 49.40, p(2) = e^{1.5} ≈ 4.48, p(3) = e^{-0.8} ≈ 0.45

[0151] · Sum: 0.10 + 49.40 + 4.48 + 0.45 ≈ 54.43

[0152] · Probability calculation P(0)&=0.10 / 54.43≈0.002,P(1)=49.40 / 54.43≈0.91,P(2)&=4.48 / 54.43≈0.08,P(3)=0.45 / 54.43≈0.008

[0153] Result: The probability of hyperacute phase is the highest (0.91), consistent with the true label of the sample.

[0154] It is concluded that P(1) is the highest probability of hyperacute phase (0.91), consistent with the true label of the sample.

[0155] In combination with the above modeling and screening process, a cross-species data fusion process is performed to improve the estimation probability of low data volume species data, i.e., step E: cross-species homologous gene mapping. Taking human gene data as an example, and using a predefined dictionary to map the genes of humans and mice, the specific use of python code mapping is as follows:

[0156] python

[0157] # Human-mouse homologous gene mapping table

[0158] HUMAN_TO_MOUSE_GENE_MAP = {

[0159] 'S100A8': 'S100a8', 'CD14': 'Cd14', 'ARG2': 'Arg2',

[0160] 'ATF3': 'Atf3', 'GM2A': 'Gm2a', 'S100A9': 'S100a9', 'IFITM1': 'Ifitm1'}

[0161] def cross_species_standardization(human_data, mouse_scaler):

[0162] # Process human data with mouse standardization parameters (cross-species standardization)

[0163] human_data_mapped =

[0164] human_data.rename(columns=HUMAN_TO_MOUSE_GENE_MAP)

[0165] return mouse_scaler.transform(human_data_mapped)

[0166] After the data processing and model building process described above, the training process actually loads a pre-trained mouse model, i.e., the human model has a "knowledge starting point". The python code is as follows:

[0167] python

[0168] def load_human_data(data_format):

[0169] X, y, scaler = load_rnaseq_human_data(data_format)

[0170] # Load pre-trained mouse model

[0171] mouse_model_path = "mouse_model.json"

[0172] mouse_model = xgb.XGBClassifier()

[0173] mouse_model.load_model(mouse_model_path) # Load mouse model parameters

[0174] return X, y, scaler, mouse_model

[0175] By the above process, the risk of having less data for the human model can be reduced, and the dependence on the amount of data can be reduced. It should be noted that in order to reduce the distribution difference between mouse and human gene expression (such as numerical range, mean and variance), the code uses the mouse's standardization parameters to process human data, that is: python def cross_species_standardization(human_data, mouse_scaler):

[0176] # Use mouse's standardizer (scaler) to process human data

[0177] human_data_mapped =

[0178] human_data.rename(columns=HUMAN_TO_MOUSE_GENE_MAP) # Gene name mapping

[0179] return mouse_scaler.transform(human_data_mapped) # Standardize human data using mouse's mean / variance

[0180] The difference in processing human gene data is that the SHAP values of each gene are extracted from the trained XGBoost model, and the SHAP values are fused across species. Specifically, the SHAP value fusion method is to take the weighted mean, which uses the code to dynamically calculate the gene weight based on the fused SHAP values through the DynamicWeightCalculator class, while retaining a certain adaptive adjustment space, that is:

[0181] Dynamic weight calculation (combined with SHAP value, time decay and adaptive term)

[0182] weights = [self.weight_calculator.calculate_weight(i, time) for i in range(len(scaled_values[0]))]

[0183] This method of taking the average can make the technical ideas and data provided by the technical solution of the present application continue to be applied in future long-term research. This is because in the process of cross-species data modeling, especially in the early stages of research, the number of human samples is often insufficient. At this time, the present application proposes to use the SHAP value extracted from the mouse model as prior knowledge for supplementation, avoiding overestimation or underestimation of the role of certain genes in the human model due to insufficient training samples. With the development of related fields, human data (or other species data) gradually becomes sufficient, and this mean fusion method can gradually weaken the role of mouse data, while human data naturally dominates in weight distribution. At the same time, this method does not rely on complex parameter adjustment in the implementation process and is easy to operate.

[0184] According to the specific example calculation of the above method, the original data is shown in Table 5,

[0185] Table 5 Human hyperacute sample data

[0186]

[0187] Preprocessing:

[0188] log1p conversion: log1p (S100A8) = log (1 + 42856) ≈ 10.67,

[0189] log1p (CD14) = log (1 + 13039) ≈ 9.47,

[0190] The rest after conversion is [3.22, 5.36, 7.20, 10.27, 7.63].

[0191] Standardization (Z-score): Adjusted features based on human training set: [1.82, 1.56, -0.41, 0.32, -0.18, 1.69, 0.25]

[0192] According to the results, the XGBoost model outputs are as follows: non-stroke (0): -2.3, hyperacute phase (1): 3.8, acute phase (2): 1.5, subacute phase (3): -0.7

[0193] Softmax probability calculation:

[0194] P(0) ~ 0.10 / 50.78 ~ 0.002, P(l) ~ 44.70 / 50.78 ~ 0.88

[0195] P(2) ~ 4.48 / 50.78 ~ 0.088, P(3) ~ 0.50 / 50.78 ~ 0.01

[0196] Result: The probability of hyperacute phase is 0.88, which is consistent with the true label of the sample (D1 is a human hyperacute phase sample), that is, the modeling is feasible and the scoring method is reliable.

[0197] The "dynamic weight and cross-species data time series modeling, scoring method and gene combination" provided by the application aims at the core technical problems of multi-source omics data analysis, time series staging modeling of stroke process, feature gene screening and scoring method construction, and provides an innovative and systematic analysis solution. The application not only breaks through the static nature and low universality of the traditional gene screening and staging model, but also introduces advanced algorithm means such as cross-species mapping, dynamic weight adjustment and multi-dimensional data fusion, which embodies significant technical advantages in the aspects of model interpretability, stability, flexibility, adaptability and the like, and has wide scientific research and practical application value. The technical scheme of the application is started from the perspective of biological data modeling, mainly provides a model for scientific research and data analysis, and does not involve any clinical diagnosis or disease treatment. The feature screening process, dynamic model structure and scoring mechanism proposed can be regarded as a biological information analysis framework. The framework is constructed based on public data and faces scientific research scenarios, emphasizes algorithm, process and reusability design, and what needs to be emphasized is that the dynamic modeling framework constructed by the application is not limited to the stroke scenario, and the dynamic weight calculation, cross-species data integration and staging scoring mechanism proposed are general algorithm logic, which can be translated into other disease models involving time evolution and immune regulation, such as sepsis onset cycle analysis, tissue damage repair dynamic tracking, or chronic inflammatory disease process modeling

[0198] Although embodiments of the application have been shown and described, it is to be understood that the application is not limited to the details of the foregoing embodiment, and that various changes in form and details can be made without departing from the spirit and scope of the application. The scope of the application is defined by the appended claims and their equivalents.

Claims

1. A method for time-series modeling of dynamic weights and cross-species data, characterized in that, The method comprises the following steps: Step A: data acquisition and preprocessing A1, inductive construction of multi-source data set; A2, standardization of scRNA-seq / RNA-seq and qPCR data; A3, using the brain transcriptome data in the multi-source data to interpolate and calculate the peripheral blood time series expression profile, and the method for interpolating and supplementing the missing time series in the data set comprises one or more ways of combination of linear interpolation, polynomial interpolation and cubic spline interpolation; Step B: cross-organizational causal correlation and dynamic network construction B1, using the preprocessed multi-source data set in the previous step to construct a Bayesian causal network to represent the brain-peripheral blood gene causal correlation; B2, screening brain-peripheral blood cross-organizational significant covariant genes based on random forest regression, and incorporating them into a dynamic gene interaction network; Step C: time-specific gene screening and initial gene screening, that is, in different staging time windows (such as 6h, 24h, 72h, 7d, etc.), the method combining differential expression and statistical model is used to preliminarily lock the gene candidate set which significantly changes with time and has biological significance; C1, using random forest and XGBoost algorithm to screen candidate genes meeting the following conditions: 1) differential expression fold (|log2FC|>1) and p<0.05 at the corresponding time point; 2) extracting genes significantly changed with time after stroke based on time series mixed effect model (LMM) (p<0.05, |avg_log2FC|>0.3); 3) significantly related to brain pathological conditions and cell function state (such as infarction volume, inflammation pathway, nerve repair, etc.); 4) no significant change in expression in non-stroke inflammation models (such as LPS induction). Step D: multi-class time series model training and secondary reverse screening D1, constructing an XGBoost multi-classification model to classify samples into different stages of indicators; D2, based on 5-fold stratified cross-validation, taking ROCAUC as the evaluation index, and optimizing the model hyperparameters through grid search; D3: performing ROC AUC curve analysis on the candidate genes screened in step D2, calculating the AUC values of each gene in the staging classification, and removing the genes with low AUC values to finally determine a number of AUC highest feature genes as input objects for subsequent dynamic weight calculation. Step E: cross-species homologous gene mapping E1. Transforming cross-species genes through a predefined dictionary; E2. Loading another species gene data synchronously loads the pre-trained species model in steps A-D; E3. Processing cross-species data according to the data standardization method in step A; E4. Constructing an XGBoost multi-classification model of cross-species gene data according to the method in step D; Step F: dynamic weight calculation and time series fusion F1. Extracting the SHAP values of each gene from the trained XGBoost model and performing cross-species SHAP value fusion; F2. Determine static weights w i , SHAP; F3. Calculate the dynamic weight of each gene according to the following formula: w i (t) = w i , SHAP · e -λit + a · Adapt i (t) Where λi and β are adjustable parameters.

2. The method of claim 1, wherein: The multi-source data set at least includes GSE225948, CRA016580, GSE233815, GSE122709.

3. The method of claim 1, wherein: The spatial transcriptome data of GSE233815 is calculated by three times of spline interpolation method to obtain D1 / D3 / D7 peripheral blood gene expression profile.

4. The method of claim 1, wherein: In step A2, the scRNA-seq / RNA-seq data is subjected to log1p conversion and then Z-score standardization, and the ΔCt value of the qPCR data is calculated and then Z-score standardization is performed.

5. The method of claim 1, wherein: In step F3 β is an adjustable learning rate for real-time adaptation of individual differences and batch effects for each sample, with an initial value of 0.01 and exponential decay with the number of iterations.

6. The method of claim 1, wherein: In step F3, λi is an attenuation coefficient, which is optimized by grid search, and its range is 0.05-3.

7. The method of claim 1, wherein: In step F, the cross-species SHAP value fusion method is to take the weighted mean.

8. A dynamic weight and cross-species data scoring method, characterized in that: The scoring method comprises the following steps, G1. Staging score calculation: Score(t) =∑(w i (t) x Δct i ), G2. mapping the staging score to the [0, 1] interval by a Softmax function, where score k is the staging score for the kth stage, and n is the total number of stages.

9. A combination of dynamic weighting and scoring method of cross-species data, characterized in that, The gene combination of qPCR at least includes S100A8, CD14, ARG2, and GAPDH as an internal reference gene.

10. A combination of dynamic weighting and scoring method of cross-species data, characterized in that, The gene combination of human scRNA-seq / RNA-seq at least includes S100A8, CD14, ARG2, S100A9, IFITM1, ATF3, GM2A; The gene combination of mouse scRNA-seq / RNA-seq at least includes S100a8, Cd14, Arg2, Chil3, Gda, S100a9, Ifitm1.