Method for determining path data influencing progression of malignant blood particle deficiency type sepsis

By filtering and analyzing the initial gene sequence data, a significant enriched pathway in granulocytic sepsis, a malignant hematologic disease, was identified. This solved the problems of high data noise and lack of specificity in traditional techniques, and achieved highly accurate and reproducible pathway data identification.

CN121725877APending Publication Date: 2026-03-24SAILI CHUANGXIN MEDICAL TECHNOLOGY (SHANGHAI) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Traditional techniques often suffer from vaguely defined study cohort populations and a lack of specific data for patients with malignant hematological diseases, resulting in noisy results, an inability to effectively integrate multidimensional information, a failure to deeply couple clinical information with molecular data, and an inability to effectively determine the biological progression pathways of granulocytic sepsis, a malignant hematological disease.

Method used

By filtering the initial gene sequence data, clean gene sequence data is obtained. Gene quantitative analysis is then performed to obtain gene expression data. Next, differential gene expression analysis is performed, and finally, functional enrichment analysis is conducted to identify significantly enriched pathway data.

Benefits of technology

It significantly improves the signal-to-noise ratio and data reliability, sensitively and robustly identifies real differentially expressed genes related to phenotypes, forms an objective quantitative judgment of key pathways, improves the accuracy and reproducibility of pathway judgment, and provides highly reliable technical support for disease mechanism analysis and biomarker/target screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121725877A_ABST
    Figure CN121725877A_ABST
Patent Text Reader

Abstract

The invention relates to a method for determining path data influencing the progress of malignant blood particle deficiency type sepsis. The method comprises the following steps: filtering initial gene sequence data to obtain clean gene sequence data; performing gene quantitative analysis on the clean gene sequence data to obtain gene expression quantity data; performing gene differential expression analysis on the gene expression quantity data to obtain differential expression gene data; and performing function enrichment analysis on the differential expression gene data to obtain significant enrichment pathway data. By adopting the method, the path data can be determined to have novelty and specificity.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biomedical technology, in particular to a method and device for determining pathway data affecting the progression of granulocytopenia sepsis in malignant hematopathy and a computer device. BACKGROUND

[0002] Sepsis is the leading cause of death due to severe infection, which not only seriously threatens human health, but also brings a huge economic burden to the medical and health sectors. Granulocytopenia sepsis in malignant hematopathy is a special type of sepsis, and the incidence of sepsis is higher in patients with neutropenia (granulocytopenia), the progression is faster, and the mortality is higher. Clinical studies have shown that more than 80% of hematological tumor patients will develop granulocytopenia-related infections after chemotherapy, with a mortality rate of 10%-30%. The risk of sepsis in patients with malignant hematopathy is 15 times that of the general population, and 16% of patients with granulocytopenia will develop sepsis after blood stream infection, with a mortality rate of up to 55%. Therefore, it is crucial to elucidate the mechanism of the occurrence and progression of granulocytopenia sepsis in malignant hematopathy, and to provide early warning and intervention, in order to reduce the mortality rate of infection in patients with malignant hematopathy and granulocytopenia, and to improve the cure rate and survival rate.

[0003] In the traditional technology, the definition of the research cohort is ambiguous, and all sepsis populations are widely included, which has strong heterogeneity. There is a lack of cohort data specifically for patients with malignant hematopathy, resulting in a large amount of noise in the results. Relying on public databases, there is a lack of targeted clinical phenotype data, making it difficult to discover unique biological progression pathways that occur in the presence of infection in the granulocytopenia state after chemotherapy in malignant hematopathy. The data analysis method is single, and cannot effectively integrate multi-dimensional information, resulting in a high false positive rate. Clinical information and molecular data cannot be deeply coupled and analyzed. As a result, the final pathway data found is mostly known, lacking novelty and specificity. SUMMARY

[0004] Therefore, it is necessary to provide a method and device for determining pathway data affecting the progression of granulocytopenia sepsis in malignant hematopathy, which has novelty and specificity.

[0005] In a first aspect, the present application provides a method for determining pathway data affecting the progression of granulocytopenia sepsis in malignant hematopathy, comprising: filtering the initial gene sequence data to obtain clean gene sequence data; performing gene quantitative analysis on the clean gene sequence data to obtain gene expression data; performing gene differential expression analysis on the gene expression data to obtain differential expression gene data; performing functional enrichment analysis on the differential expression gene data to obtain significantly enriched pathway data.

[0006] Secondly, this application also provides a pathway data determination device for influencing the progression of the hematologic malignancy granulocytosis, comprising: The data filtering module is used to filter the initial gene sequence data to obtain clean gene sequence data; The quantitative analysis module is used to perform quantitative gene analysis on the clean gene sequence data to obtain gene expression level data; The differential expression analysis module is used to perform differential gene expression analysis on the gene expression level data to obtain differentially expressed gene data. The enrichment analysis module is used to perform functional enrichment analysis on the differentially expressed gene data to obtain significantly enriched pathway data.

[0007] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any step of a method for determining pathway data affecting the progression of the malignant hematologic disease granulocytic sepsis.

[0008] The aforementioned method, apparatus, and computer equipment for identifying pathway data influencing the progression of agranulocytic sepsis, a malignant hematological disease, significantly improves the signal-to-noise ratio and data reliability by performing quality filtering on initial gene sequence data to remove adapter contamination, low-quality, and unusable reads. Subsequently, gene quantification is performed on the clean data to standardize dimensions and reduce systematic bias caused by sequencing depth and inter-sample differences, thereby obtaining comparable expression levels. Based on this, differential expression analysis can sensitively and robustly identify truly differentially expressed genes related to the phenotype, reducing false positives caused by random fluctuations. Finally, through functional enrichment analysis of differentially expressed genes, scattered gene-level signals are summarized into interpretable biological pathway-level evidence, forming an objective quantitative judgment of key pathways. The processing achieves layered constraints and stepwise convergence, ensuring the novelty and specificity of pathway data while improving the accuracy and reproducibility of pathway identification. It also facilitates cross-batch and cross-platform migration and application, providing highly reliable technical support for disease mechanism analysis and biomarker / target screening. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a diagram illustrating the application environment of a method for determining pathway data that influences the progression of agranulocytic sepsis, a malignant hematological disease, in one embodiment. Figure 2 This is a flowchart illustrating a method for determining pathway data that influences the progression of the hematologic malignancy granulocytic sepsis in one embodiment; Figure 3 This is a general technical roadmap for a method to determine pathway data affecting the progression of granulocytic sepsis, a malignant hematological disease, in one embodiment. Figure 4 This is a structural block diagram of a pathway data determination device that affects the progression of the hematologic malignancy agranulocytic sepsis in one embodiment; Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0011] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0012] This application provides a method for determining pathway data affecting the progression of granulocytic sepsis, a malignant hematological disease, which can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on other network servers. Server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0013] In one exemplary embodiment, such as Figure 2 As shown, a method for determining pathway data that influences the progression of granulocytic sepsis, a malignant hematological disease, is provided, and this method is applied to... Figure 1 Taking the server in the example, the explanation includes the following steps 202 to 208. Wherein:

[0014] Step 202: Filter the initial gene sequence data to obtain clean gene sequence data.

[0015] Step 204: Perform gene quantification analysis on the clean gene sequence data to obtain gene expression level data.

[0016] Step 206: Perform differential gene expression analysis on the gene expression data to obtain differentially expressed gene data.

[0017] Step 208: Perform functional enrichment analysis on the differentially expressed gene data to obtain significantly enriched pathway data.

[0018] The initial gene sequence data consists of raw reads output by the sequencer without any quality control or filtering.

[0019] Clean gene sequence data refers to clean reads that are obtained from initial gene sequence data after quality control processes such as adapter removal and low-quality / N ratio filtering, and can be used for downstream analysis.

[0020] Among them, gene quantification analysis is a process that uses clean gene sequence data as input, and calculates the expression abundance of each gene in each sample by comparison or quasi-alignment and combined with length and sequencing depth correction.

[0021] Among them, gene expression data are comparative expression matrices (such as counts, TPM / CPM or their transformed values) organized by "gene × sample" after quantitative analysis.

[0022] Gene differential expression analysis is a process that uses a design matrix to statistically test the gene expression levels of two or more groups of samples and controls multiple comparisons to identify genes with significantly different expression levels.

[0023] Among them, the differentially expressed gene data is a list of differentially expressed genes and a set of related statistical indicators obtained by screening according to preset thresholds (such as FDR and |log2FC|).

[0024] Functional enrichment analysis is a process that uses differentially expressed genes as input and detectable genes as background to assess the enrichment of gene sets in biological pathways against a pathway / functional database.

[0025] Among them, the significantly enriched pathway data are a list of pathways that have reached the significance standard after multiple tests and corrections, along with their corresponding overlapping genes, statistical significance, and related annotation information.

[0026] Specifically, the initial gene sequence data is first quality-assessed, and then adapter removal, length trimming, and low-quality base shearing are performed accordingly. Subsequently, reads containing adapter residues, reads with average quality values ​​below a set lower limit, reads with N base ratios exceeding the threshold, and reads matching contamination references (such as rRNA / microbial vectors) are removed sequentially according to thresholds. Finally, clean gene sequence data that has passed quality control is output, and statistics before and after filtering (number of reads, Q30 percentage, GC distribution, and sequence length distribution) are generated.

[0027] Clean gene sequence data are aligned or quasi-aligned to a reference genome / transcriptome for quantification, yielding a raw count of sample × transcript / gene. After summarizing at the gene level according to transcript-to-gene mapping, length correction and sample depth normalization (e.g., RPK / TPM or CPM) are performed, along with appropriate numerical transformations (e.g., log2(x+1) or variance stabilization) and low-expression filtering. This results in a comparable gene expression matrix with accompanying quality control metrics at the sample and gene levels (sequencing depth, alignment rate, zero-value ratio, etc.).

[0028] A design matrix was constructed based on the gene expression matrix, and known confounding factors (batch, sex, cell ratio, etc.) were corrected for or incorporated into the model. A suitable statistical framework was used to perform hypothesis testing on inter-group expression differences, outputting the effect size (fold change) and significance index (P-value / adjusted P-value) for each gene. The differentially expressed gene set, i.e., the differentially expressed gene data, was obtained by analyzing a preset threshold (e.g., FDR < 0.05 and |log2FC| > 1). Simultaneously, stability checks were performed on the results (volcano plot, MA plot, principal component / sample cluster consistency) to verify the reliability of the analysis.

[0029] Using all successfully quantified and quality-controlled genes from differentially expressed gene data as the background set, functional enrichment analysis (commonly hypergeometric / enrichment tests or rank methods) is performed on predefined pathway / functional databases (such as KEGG / GO / Reactome) to obtain the raw significance of each pathway. False discovery rate correction is applied to multiple comparisons, and a list of significantly enriched pathways is obtained by combining additional rules such as minimum number of overlapping genes and consistency of effect direction. The significantly enriched pathway data are then obtained by combining the name of each pathway, a list of overlapping genes, corrected significance indicators, and a visual summary (bar chart / bubble chart / pathway icon annotation).

[0030] The aforementioned method for identifying pathway data influencing the progression of agranulocytic sepsis, a malignant hematological disease, significantly improves signal-to-noise ratio and data reliability by quality filtering of initial gene sequence data to remove adapter contamination, low-quality, and unusable reads. Subsequently, gene quantification is performed on the clean data to standardize dimensions and reduce systematic bias caused by sequencing depth and inter-sample differences, thereby obtaining comparable expression levels. Based on this, differential expression analysis can sensitively and robustly identify truly differentially expressed genes related to phenotypes, reducing false positives caused by random fluctuations. Finally, functional enrichment analysis of differentially expressed genes summarizes scattered gene-level signals into interpretable biological pathway-level evidence, forming an objective quantitative judgment of key pathways. This process achieves layered constraints and gradual convergence, ensuring the novelty and specificity of pathway data while improving the accuracy and reproducibility of pathway identification. It also facilitates cross-batch and cross-platform migration and application, providing highly reliable technical support for disease mechanism analysis and biomarker / target screening.

[0031] In an exemplary embodiment, differential gene expression analysis is performed on gene expression level data to obtain differentially expressed gene data, including steps 302 to 306. Wherein:

[0032] Step 302: Perform expression morphology transformation on the gene expression data to obtain the gene expression matrix.

[0033] Step 304: Perform statistical significance analysis on the gene expression matrix to obtain the analytical expression matrix.

[0034] Step 306: Perform biological significance analysis on the expression matrix to obtain differentially expressed gene data.

[0035] Among them, expression morphology conversion is the process of converting the original gene expression data into an expression form that can be directly modeled after normalization between samples and numerical transformation (such as log or variance stabilization), low expression filtering and format unification.

[0036] The gene expression matrix is ​​a two-dimensional data table organized in the form of "gene × sample" and has been standardized and transformed as necessary.

[0037] Among them, statistical significance analysis is a process of performing hypothesis testing on inter-group expression differences based on the design matrix and performing multiple comparison correction to obtain the effect size and significance index of each gene.

[0038] Among them, the expression level analysis matrix is ​​an extended data table used for comprehensive judgment, which is formed by adding or replacing statistics (such as log2FC, P value, FDR, dispersion, etc.) on the basis of the gene expression level matrix.

[0039] Among them, biological significance analysis is a process of screening statistical results for biological significance to determine candidate genes by combining rules such as effect size threshold, expression robustness, directional consistency and known pathway / marker association.

[0040] Specifically, gene expression data are constructed into a matrix according to "gene × sample" and the sample order and metadata are standardized. Then, standardization and numerical transformation are performed (such as normalization between samples based on TMM / RLE and variance stabilization using log2(x+1) or VST). At the same time, low expression and extreme outlier entries are removed according to a preset threshold, and the gene expression matrix for statistical modeling is output.

[0041] A design matrix is ​​constructed based on the gene expression matrix, and groups and necessary covariates (such as batch, age, sex, cell ratio, etc.) are included in the model. The effect size and dispersion are estimated by fitting an appropriate statistical framework (such as negative binomial generalized linear model or linear model-empirical Bayesian method) to the genes. The original p-value is calculated and multiple test correction (FDR) is performed. The log2FC, p-value, FDR and other statistics of each gene are compiled together with the original expression to form an analytical expression matrix.

[0042] Based on biological criteria, target genes were screened from the expression matrix. This included setting effect size thresholds (e.g., |log2FC|≥1) while controlling for statistical significance, minimum expression and replication support requirements, directional consistency, and known marker / pathway relevance. The selected genes were functionally annotated and integrated to ultimately generate differentially expressed gene data (including a gene list and corresponding effect sizes, FDRs, and key annotations).

[0043] In this embodiment, expression morphology transformations such as standardization and variance stabilization of gene expression data are performed to reduce technical biases caused by sequencing depth, batch size, and outliers, and to improve data comparability and distribution modelability. Then, statistical significance analysis is used to quantify the effect size and confidence level of each gene and multiple comparison control is implemented to suppress false positives and false negatives from the source. Finally, biological significance analysis is combined with rules such as effect threshold, expression robustness, and pathway / marker association to perform secondary screening, ensuring that the output differentially expressed genes have both statistical reliability and biological interpretability, thereby significantly improving the accuracy, reproducibility, and practical value of downstream pathway analysis and target discovery.

[0044] In an exemplary embodiment, gene expression level data undergoes expression morphology transformation to obtain a gene expression level matrix, including steps 402 to 408. Wherein:

[0045] Step 402: Perform sample layer component calibration on the gene expression data to obtain component-calibrated expression data.

[0046] Step 404: Perform graph regularization pseudo-count analysis on the component calibration expression data to obtain continuous expression data.

[0047] Step 406: Based on the continuous expression data and the effective transcript length of the gene, the matrix expression values ​​are length recalibrated to obtain recalibrated expression data.

[0048] Step 408: Perform variance stabilization transformation on the recalibrated expression data to obtain the gene expression matrix.

[0049] Among them, sample layer composition calibration is a process that robustly normalizes the overall expression level and composition bias at the sample scale to eliminate differences in sequencing depth and total amount between samples.

[0050] Among them, the component-calibrated expression data are expression data obtained after sample layer component calibration, which have eliminated the differences in the total sample size and can be compared across samples.

[0051] Among them, graph regularization pseudo-counting analysis is based on injecting controlled pseudo-counts into low-abundance and zero-inflation entries into the gene relationship graph and applying graph smoothing to alleviate sparsity.

[0052] Among them, continuous expression data are expression data that are smoother near zero, numerically continuous, and easier to model after graph regularization pseudo-counting analysis.

[0053] Among them, the effective transcript length is the equivalent transcript length used for expression correction after taking into account factors such as the distribution of comparable regions and fragment lengths.

[0054] The matrix expression value is the quantitative value corresponding to each cell in the expression matrix organized by "gene × sample".

[0055] Among them, length recalibration is a process that normalizes the expression level gene by gene based on the effective transcript length to eliminate the influence of length differences.

[0056] Among them, relabeled expression data are expression levels that are comparable among genes after length relabeling.

[0057] Among them, variance stabilization transformation is a process that weakens mean-variance coupling by using monotonic transformations (such as logarithmic or VST) to make the variance approximately constant at different expression levels.

[0058] Specifically, after adjusting the original gene expression data to a "gene × sample" format, scale factors and compositional bias indicators are calculated at the sample level. Robust methods (such as the median ratio method in DESeq, RLE, or TMM trimmed mean method) are preferentially used to estimate the scale factor for each sample relative to a pseudo-reference sample (constructed based on the geometric mean of each gene across all samples). During the estimation process, extreme genes (low expression, extreme fold change, and genes suspected of technical noise) are symmetrically trimmed. Threshold checks are performed on key quality control features of the samples (total read count, alignment, proportion of zero values, rRNA / mitochondrial ratio), and if necessary, the scale factor is gently shrunk or abnormal samples are marked to prevent overcorrection. To avoid global drift, the geometric mean of all scale factors is constrained to 1, and Winsorization can be optionally used to truncate the scale factor within a preset interval (e.g., 5%–95th percentile) to suppress extreme values ​​in single samples. Finally, the expression values ​​of all genes in the sample are linearly scaled according to the scaling factor of each sample. Only one sample layer calibration is performed without changing the relative order between genes. At the same time, the mapping relationship between the original measure (count, TPM / CPM) and the calibrated measure and the calibration log (method used, pruning ratio, outlier sample markers and final scaling factor) are preserved, and the component calibration expression data are output.

[0059] A priori relational graph with genes as nodes is constructed using component-calibrated expression data (prioritizing the fusion of co-expression information from pathway / interaction databases and the same batch of data). Edges are weighted according to evidence strength and relevance, and degree normalization and weak edge truncation are performed to obtain a sparse, connected, and robust gene graph. Adaptive pseudo-counting injection is performed on low-abundance and zero-inflation units in the matrix. Based on the median expression and variability of each gene in its first / second-order neighborhood, a minimal and controlled positive offset is added to the unit according to preset upper and lower limits (no injection or minimal injection for high-confidence zero values), and the pseudo-count magnitude and coverage ratio are recorded simultaneously for auditing. Graph smoothing regularization is performed on this gene graph to propagate local structural information and alleviate measurement sparsity (random walk diffusion with stopping probability, Laplace-constrained least squares, or total variational denoising can be used). The smoothing amplitude is controlled by a single intensity hyperparameter, and cross-validation / robustness criteria are used to prevent over-smoothing and signal leakage. To avoid amplifying extreme values, upper bound truncation and monotonicity constraints are applied to the smoothed units (without changing the direction of significant differences between samples). At the same time, the strategy of minimal modification to high-expression units is retained to maintain the true fold change. Finally, continuous expression data with continuous numerical values, low zero expansion, and better modelability in the neighborhood is output, along with accompanying processing metadata (graph construction source and parameters, pseudo-count injection threshold and ratio, smoothing strength and convergence index).

[0060] Based on continuous expression data and the effective transcript length of each gene, all transcripts of the same gene are integrated with reference annotations, and the effective length is calculated (preferably using "effective length" that considers the distribution of comparable regions and fragment lengths; for multiple isoforms, representative values ​​at the gene level are obtained by aggregation according to comparable weights or equal weights; for entries with missing or conflicting annotations, the most recent version of the annotation is used for backfilling or imputation based on the length quantiles of similar genes, and a minimum length limit is set for very short transcripts to avoid numerical instability). At the gene level, the expression value of each sample is normalized gene-by-gene length according to the corresponding effective length (recalibrated by the expression / length ratio, maintaining non-negativity and monotonicity, without changing the relative order of genes within the sample), and a scale constraint is applied to the total recalibration within the sample to avoid global drift (e.g., fixing the geometric mean to 1 or normalizing the total to a sample-specific scale). At the same time, abnormal amplification caused by excessively short lengths is gently truncated or contracted to suppress extreme values. To ensure traceability, the recalibrated expression value corresponding to each gene / sample is output along with the effective length, interpolation and truncation markers used, and the monotonic mapping from continuous expression data to recalibrated expression data and parameter logs are retained to form recalibrated expression data.

[0061] Based on the mean-variance relationship estimated for all genes and samples, an appropriate family of transformations is selected from the recalibrated expression data (preferably using variance stabilization transformations or logarithmic transformations supported by empirical Bayesian dispersion estimation, and minimum positive offset for zero and extremely low values ​​to avoid numerical divergence). Then, a monotonic and scale-comparable transformation is applied to the expression values ​​of each sample at once, with the transformation strength selected using cross-validation or information criteria to reduce mean-driven heteroscedasticity. To suppress outlier amplification, mild truncation and robust standardization are applied to the transformation results, and positional scaling correction is performed on the sample dimension to eliminate residual sample-level systematic bias. Optional row centering is also performed on the gene dimension to facilitate subsequent linear or generalized linear modeling. During the transformation process, the monotonicity of the order between genes and groups and the consistency of effect direction are maintained, avoiding over-smoothing of high-confidence differential signals. Key metadata (transformation type and parameters, offset, truncation threshold, standardization coefficients, and quality diagnostic indicators) is recorded. Finally, a gene expression matrix with approximately constant variance is output at different expression levels.

[0062] In this embodiment, systematic bias introduced by differences in sequencing depth and sample composition was eliminated and cross-sample comparability was improved by calibrating the sample layer components. Subsequently, graph regularization pseudo-counting analysis was used to alleviate zero expansion and sparsity problems while preserving local co-expression structures, thereby improving the signal-to-noise ratio in the low expression range. On this basis, the effective transcript length of the gene was recalibrated for the matrix expression values, which unified the counting bias caused by the length difference of different genes and ensured the consistency of dimensions among genes. Finally, variance stabilization transformation was performed to weaken mean-variance coupling, reduce the influence of heteroscedasticity and extreme values, and make the data more consistent with the statistical assumptions of linear / generalized linear models, significantly improving the sensitivity and specificity of difference testing and improving the robustness and reproducibility of the results.

[0063] In an exemplary embodiment, gene quantitative analysis is performed on clean gene sequence data to obtain gene expression level data, including steps 502 to 506. Wherein:

[0064] Step 502: Perform gene length correction on the clean gene sequence data to obtain gene length corrected data.

[0065] Step 504: Perform total sequencing depth correction on the clean gene sequence data to obtain sample depth correction data.

[0066] Step 506: Divide the gene length correction data by the sample depth correction data to obtain gene expression level data.

[0067] Gene length correction involves normalizing the original counts gene-by-gene based on the comparable length of each gene (or its effective transcript) to eliminate counting bias caused by length differences.

[0068] Among them, gene length-corrected data is the expression metric obtained after gene length correction, which has achieved length comparability between genes.

[0069] Among them, total sequencing depth correction: expression values ​​are normalized between samples based on the effective total sequencing volume (sample scale factor) of each sample to eliminate the influence of sequencing depth differences.

[0070] Among them, the sample depth correction data is the scale factor or its corresponding scale obtained by correcting the total sequencing depth of the samples, which is used to perform proportional conversion and comparison of the expression values ​​of each sample.

[0071] Specifically, clean gene sequence data are quantified against a reference genome / transcriptome to obtain raw gene layer counts, and the effective transcript length of each gene is aggregated (considering the distribution of comparable regions and fragment lengths, isotype weights, and minimum length limits). Gene-by-gene normalization is performed on the corresponding gene counts within each sample based on this effective length, and gene length-corrected data that has been corrected for length differences is output.

[0072] Based on the quantitative results corresponding to clean gene sequence data, the total number of comparable effective sequences for each sample is calculated as the sample scale factor (which can be a standardized scale of the sum of the number of effective reads after quality control and the length correction value, or an equivalent robust estimate). The scale factor of abnormal samples is gently truncated / shrunken to avoid the influence of extreme values, thus forming sample depth correction data.

[0073] Normalize the gene length-corrected data by dividing it by the corresponding sample scale factor of the sample depth-corrected data (i.e., by scaling it up using the sample scale factor), and obtain comparable gene expression data across samples and genes (e.g., expression measures equivalent to TPM / CPM) while keeping the non-negativity and effect direction unchanged.

[0074] In this embodiment, gene length correction is first applied to clean gene sequence data to eliminate counting bias caused by different gene transcript lengths, making genes within the same sample comparable. Subsequently, the same batch of data is corrected according to the total sequencing depth of the samples to offset the systematic bias introduced by differences in sequencing volume and library complexity between samples. On this basis, the length correction results are proportionally converted using a sample scaling factor (i.e., gene length-corrected data is divided by sample depth-corrected data) to form an expression metric that can be directly compared across samples and genes (a dimensionally consistent expression level equivalent to TPM / CPM). This significantly improves the accuracy and stability of expression estimation, reduces the risk of false positives / false negatives, and provides a data foundation with higher signal-to-noise ratio and reproducibility for subsequent differential analysis and pathway enrichment.

[0075] In an exemplary embodiment, before performing gene length correction on the clean gene sequence data to obtain gene length-corrected data, the method further includes steps 602 to 606. Wherein:

[0076] Step 602: The multiple alignment reads in the clean gene sequence data are probabilistically redistributed to obtain the multiple alignment weight parameters.

[0077] Step 604: Perform prior estimation of sample layer fragment length on the clean gene sequence data to obtain prior parameters for fragment length.

[0078] Step 606: Based on the clean gene sequence data and the GC content of the reference transcript, perform GC bias calibration on the read weights of the initial gene sequence data to obtain the GC weight parameters.

[0079] Step 608: Perform probability allocation learning on the multi-position reads in the clean gene sequence data to obtain the multi-position allocation parameters.

[0080] Step 610: Perform library strand specificity determination on the clean gene sequence data to obtain strand specificity parameters.

[0081] Among them, a multiple alignment read is a sequencing read that can be aligned to two or more gene / transcription locations and cannot be uniquely located.

[0082] Among them, probability redistribution is the process of distributing the contribution of multiple alignment reads among candidate positions according to probability proportions based on alignment quality, positional features and prior information.

[0083] Among them, the multiple alignment weight parameter is a normalized probability weight vector that is formed at the candidate position of each multiple alignment read and sums to 1, which is used for weighted counting.

[0084] Among them, the prior estimation of fragment length in the sample layer is the process of inferring the distribution of fragment length in the library by aligning reads / read pairs with high confidence within a single sample to obtain its prior.

[0085] Among them, the prior parameter of fragment length is a parameterized or discretized representation of the fragment length distribution (such as mean, variance, quantile or bucketed probability table), which is used for effective length and accessibility correction.

[0086] Among them, the reference transcript GC content is the proportion of GC in the reference transcript or the window covered by the read segment, and is used as a covariate for bias modeling.

[0087] The read weight is a scalar coefficient assigned to each read, used to adjust its contribution to the expression estimation based on quality, bias, and location uncertainty.

[0088] Among them, GC bias calibration is a process that models the systematic bias of read abundance with GC content and corrects it accordingly to eliminate quantitative bias caused by GC.

[0089] Among them, the GC weight parameter is a set of read segment-level correction factors obtained from GC bias calibration and normalized / truncated.

[0090] Among them, multi-position reads are reads that can be compatible with multiple transcripts or multiple genes within the quasi-alignment / equivalence class framework.

[0091] Among them, probability assignment learning is a process of iteratively learning the posterior assignment probability of multi-localization reads on each candidate target by combining compatibility and abundance methods such as EM / variational methods.

[0092] Among them, the multi-location allocation parameter is the posterior allocation probability of the multi-location reads to each candidate transcript / gene, which is used to summarize and count proportionally.

[0093] Among them, the library chain specificity determination is the process of judging whether the library retains chain information and its direction pattern (such as FR / RF / non-chain specificity) based on the consistency between the reading direction and the reference direction.

[0094] Among them, the chain specificity parameter is the setting of the pattern and direction of the chain specificity of the library and its confidence level, which is used to constrain subsequent counting and quantification.

[0095] Specifically, within each sample of clean gene sequence data, reads aligned to multiple gene / transcript positions are identified (using alignment markers and MAPQ ≤ threshold as multiple alignment criteria). All candidate locations and their alignment features (alignment score, edit distance, soft splicing ratio, splice consistency, k-mer uniqueness of the site, neighborhood repetition / low complexity markers, effective length and reachability of candidate transcripts) are collected for each read. These features are scaled and robustly standardized. A predefined interpretable scoring model (such as weighted linear combination or hierarchical scoring based on sequential logic) is used to calculate the relative confidence of each candidate location. The candidate confidence vectors for the same read are normalized to form probability weights (ensuring the sum of all candidates is 1). To prevent extreme cases from dominating, minimum and maximum weight truncation, equalization rules for tied scores, and distance penalties are set. The candidate set size, final weight entropy, and reason for elimination for each read are recorded. The output is the multiple alignment weight parameters (the weight sparse matrix of read × candidate location and the threshold log) used for subsequent quantitative proportional inclusion.

[0096] For clean gene sequence data, paired end samples are identified, and "properly paired and highly confident" read pairs (properly paired, same chromosome, consistent orientation, MAPQ > threshold, no abnormal insertions or deletions) are selected, and the insert length is calculated. For single-end samples, the insert length is approximated based on the high-confidence alignment position and the library fragment model. Outliers in the original length distribution are removed (e.g., mild pruning at the 1% and 99th percentiles). Piecewise histogram and kernel density are used for joint fitting to balance the tail and main peak descriptions, and the discretized prior (probability table of equal-width or equal-frequency buckets), continuous approximation (mean, standard deviation, skewness, kurtosis), and effective interval (minimum / maximum achievable length) are output. For cross-sample consistency and traceability, goodness-of-fit indices (e.g., KS statistic) and tail skewness alerts are provided, and fragment length prior parameters are generated for effective length calculation and subsequent probability allocation.

[0097] Based on the alignment anchor positions of clean gene sequence data and the GC content of reference transcripts, the corresponding GC features are calculated for each read (or fragment). (The GC of the read sequence itself can be used, or the GC of the reference window covered by the read can be used as a proxy, with the window length matching the read length / insertion length.) The ratio of observed abundance to expected abundance in each GC stratification is statistically analyzed, and a smooth bias curve (such as LOESS / spline) is fitted. The reciprocal of the bias curve is used as the read-level correction factor. Quantile truncation and in-sample renormalization are applied to factors that are too large or too small to ensure that the overall scale does not drift. To prevent coupling between GC and covariates such as length / location, stratified correction is introduced (first stratify by length buckets or transcript regions, and then perform GC correction). The GC weight parameters of each read are output (or the average weight is given at the equivalence class level), along with the fitting residuals, truncation ratio, and outlier alerts.

[0098] Multi-location reads compatible with multiple transcripts in clean gene sequence data are aggregated into "equivalence classes" (the same set of candidate transcripts). The relative abundance of each transcript is initialized (using unique location reads or uniform priors). The EM step is iteratively executed, considering the aforementioned fragment length prior and GC / multiple alignment weights. In the E step, the posterior allocation probability is calculated based on the current abundance, effective length, and read compatibility with each candidate (ensuring the sum of probabilities of each read among candidates is 1, multiplied by the multiple alignment and GC correction weights). In the M step, reads within the equivalence class are weighted according to the allocation probability, and the abundance and effective length correction terms for each transcript are updated. Convergence is determined by the log-likelihood gain or relative parameter change being less than a threshold. A maximum number of iterations and a lower bound are set to avoid oscillations. Finally, stable multi-location allocation parameters (posterior probabilities of reads / equivalence classes to candidate transcripts and convergence diagnostics) are output, providing a probability allocation basis for the gene / transcript layer summary count.

[0099] For each sample of clean gene sequence data, select high-confidence reference sites (or junctions with known splicing directions) that are unidirectional, transcribed, and sufficiently far from neighboring genes. Statistically calculate the proportion of reads consistent with the reference sense and antisense strands. Combine this with theoretical directional patterns from the library preparation protocol (e.g., FR-firststrand, RF-secondstrand, unstranded) for hypothesis testing and consistency assessment (setting minimum coverage thresholds and uncertainty intervals). If a pattern consistently dominates in multiple independent subsets (random sampling or stratified by chromosome) and passes significance testing, the sample is identified as a strand-specific pattern and its direction is recorded. If statistical evidence is insufficient or the two directions are approximately equal, it is regressed to "non-strand-specific." Finally, output strand-specific parameters (pattern type, direction, confidence level, judgment threshold, and anomaly alerts). These parameters are then used in subsequent quantification to filter reads with inconsistent directions or adjust the included directions, avoiding expression bias caused by strand information mismatch.

[0100] In this embodiment, by re-allocating the probabilities of multiple alignment reads, their information can be reasonably distributed according to credibility, avoiding biased inclusion of individual targets. Secondly, establishing a priori fragment length at the sample level effectively compensates for differences in insertion length and accessibility, providing a stable benchmark for subsequent calculations of effective length and compatibility. Thirdly, GC bias calibration of read weights based on the GC content of reference transcripts can significantly weaken the systematic influence of amplification and sequencing chemistry on extreme GC fragments, restoring the true abundance gradient. Subsequently, probabilistic allocation learning is performed on multi-location reads to form a convergent posterior allocation among candidate transcripts / genes, improving the quantitative resolution of complex gene families and homologous sequences. Finally, library chain specificity determination is completed, and the direction of read inclusion is constrained accordingly, which can effectively reduce reverse counting and false signals, thereby improving the overall accuracy, robustness, and cross-sample comparability of expression level estimation, providing a higher signal-to-noise ratio and reproducible data foundation for differential analysis and pathway enrichment.

[0101] In an exemplary embodiment, functional enrichment analysis is performed on differentially expressed gene data to obtain significantly enriched pathway data, including steps 702 to 704. Wherein:

[0102] Step 702: Set the hypergeometric distribution test parameters corresponding to the functional enrichment analysis.

[0103] Step 704: Perform hypergeometric distribution test on the differentially expressed gene data according to the hypergeometric distribution test parameters to obtain significantly enriched pathway data.

[0104] Among them, the hypergeometric distribution test parameters are four discrete counting elements used for pathway enrichment calculation: the total number of background genes, the number of genes covered by the pathway, the number of differentially expressed genes, and the number of genes at the intersection of the two.

[0105] Among them, the hypergeometric distribution test is a statistical test that calculates the probability of "at least so many" intersection genes observed in a pathway under fixed population and sample sizes, based on the above parameters, to assess whether the pathway is significantly enriched.

[0106] Specifically, using "all genes that passed quality control and were successfully quantified" in this batch of data as a unified background set, gene identifier standardization was completed (Ensembl / Entrez / gene symbol standardization, removing aliases and duplicate mappings and recording unmapped items). Then, pathway / functional databases and their versions (such as KEGG 2025-xx, GO-BP 2025-xx, Reactome 2025-xx) were selected and locked. The number of genes covered by each pathway in the background set was pre-calculated, and the differentially expressed gene sets were stratified by direction (all / upregulated / downregulated views were available) to determine the size of the differentially expressed gene sets and the list of genes intersecting with each pathway. Subsequently, configure the test lateralization (right tail, detection enrichment), minimum overlap threshold (e.g., ≥3), reporting cap (thresholds for maximum and minimum pathway size to avoid maximal / minimum set bias), multiple comparison processing method (e.g., Benjamini–Hochberg controlled FDR), significance threshold (e.g., FDR < 0.05), and sorting rules (prioritizing FDR, followed by original significance and overlap), and solidify all hypergeometric distribution test parameters and database timestamps into the audit log.

[0107] For each pathway, four prepared hypergeometric distribution test parameters (total background genes, number of genes covered by the pathway, number of genes in the differential set, and number of genes in the actual intersection) are matched, and right-tailed cumulative probability calculation is performed to obtain the original significance. To ensure numerical stability, precise calculation or high-precision libraries are enabled for extreme counting scenarios, and pathways with unavailable / missing mappings are skipped with the reasons recorded. The original significance of all pathways is simultaneously sent to the multiple comparison correction process to generate FDR values, and then filtered according to preset thresholds and minimum overlap rules to obtain significantly enriched pathway data. For each selected pathway, the significantly enriched pathway data is output as complete fields (pathway name and ID, database and version, background size, pathway size, differential set size, list of genes in the intersection, original significance, FDR, orientation view, ranking), with quality diagnostics (such as enrichment score distribution, λ estimation, QQ / histogram summary) and visualization-ready data (the matrix required for bar / bubble charts). When upregulation / downregulation stratification is enabled, two sets of results and a merged summary are provided to support subsequent biological interpretation and validation design.

[0108] In this embodiment, by pre-setting the hypergeometric distribution test parameters (total number of background genes, number of pathway-covered genes, number of differentially expressed genes, and number of overlapping genes) for functional enrichment analysis, systematic bias caused by inconsistencies in background selection and counting methods can be eliminated at the source, ensuring the traceability and reproducibility of results. Based on this, the hypergeometric distribution test is performed on the differentially expressed gene data according to the set parameters, which can objectively measure the intensity of pathway enrichment with right-tailed probability, robustly distinguish between real functional pathway signals and random overlap, and thus output statistically sound, cross-sample, and cross-database comparable significantly enriched pathway data, improving the efficiency and credibility of subsequent biological interpretation and validation.

[0109] In an exemplary embodiment, a hypergeometric distribution test is performed on the differentially expressed gene data according to the hypergeometric distribution test parameters to obtain significantly enriched pathway data, including steps 802 to 804. Wherein:

[0110] Step 802: Based on the hypergeometric distribution test parameters, calculate the cumulative probability of pathway genes corresponding to the differentially expressed gene data to obtain the pathway enrichment random probability.

[0111] Step 804: Correct the false positive rate of the random probability of pathway enrichment to obtain significantly enriched pathway data.

[0112] Among them, pathway genes are a set of genes that belong to a predefined set of biological pathways / functions and are annotated and included in the pathway entries.

[0113] The cumulative probability calculation is the process of calculating the discrete right-tail probability of the occurrence of "actual overlap number and above" given the background size, pathway size and difference set size.

[0114] Among them, the random probability of pathway enrichment is a primitive significance index calculated from the cumulative probability and used to quantify the possibility that a certain pathway will reach its current enrichment level due to random overlap.

[0115] Among them, false positive rate correction is a statistical process that adjusts the original significance of simultaneous testing of multiple pathways by multiple comparisons (such as controlling FDR) to limit the overall false positive rate.

[0116] Specifically, after locking the background set and pathway library version, identifier unification and mapping are performed for each pathway (duplicate removal, elimination of unmapped genes and recording the reasons). Four hypergeometric distribution test parameters are calculated and cached (total number of background genes, number of genes covering the pathway in the background, number of genes in the differential gene set, and number of genes overlapping between the differential set and the pathway). The test direction is confirmed to be right-tailed, and a minimum overlap threshold and upper and lower bounds for pathway size are set to eliminate extreme sets. For each pathway, the cumulative probability of "overlap ≥ observation value" is summed stepwise using the discrete exact method. In extremely low probability scenarios, a high-precision numerical mode and logarithmic domain accumulation are used to avoid underflow. An auditable skip strategy is executed when boundary conditions occur (such as pathway coverage of 0, overlap of 0, pathway too large / too small). After calculation, the original random probability of the pathway is output, along with the pathway ID and version, four counts, overlapping gene list, threshold, and any skip / correction instructions, to obtain the pathway enrichment random probability.

[0117] The pathway enrichment random probabilities of all pathways are treated as a single family for multiple comparison correction. The Benjamini–Hochberg workflow is preferred, and the inputs are sorted from smallest to largest by their original probabilities before input, with the sorting index recorded for backfilling. For scenarios requiring stratified testing (e.g., up / down adjustment of two views or multiple databases side-by-side), the false detection rate is independently controlled within each stratum. If necessary, a more conservative procedure (e.g., Benjamini–Yekutieli) is provided to accommodate more correlated pathway sets. After correction, corrected significance and ranking are generated for each pathway, filtered according to a preset threshold (e.g., FDR < 0.05) and minimum overlap threshold. The reason codes for excluded pathways (exceeding the threshold, insufficient overlap, out-of-bounds size, etc.) and quality diagnostic summaries (p-value histogram, QQ overview, π0 / λ estimation, number of significant pathways) are retained. The final output is significant enriched pathway data, which includes pathway name and ID, database version, background and pathway size, differential set size, overlapping gene list, original probability, corrected significance, and sorting fields.

[0118] In this embodiment, by first calculating the cumulative probability of overlap between differentially expressed genes and pathway genes under a unified background and pathway library, the observed enrichment intensity can be converted into a comparable random occurrence probability, thereby quantifying the evidence strength of each pathway with objective discrete statistics. Subsequently, false positive rate correction is applied to the original random probability of all pathways, which significantly suppresses the false alarm inflation caused by multiple testing and takes into account the differences in the size and overlap of different pathways. Finally, significant enrichment pathway results with sufficient statistical basis, consistent caliber across datasets and databases, and easy priority ranking and biological interpretation are obtained, improving the reliability and reproducibility of downstream validation and decision-making.

[0119] In an exemplary embodiment, based on the hypergeometric distribution test parameters, the cumulative probability of pathway genes corresponding to differentially expressed gene data is calculated to obtain the random probability of pathway enrichment, including steps 902 to 908. Wherein:

[0120] Step 902: Based on the hypergeometric distribution test parameters, perform detectability analysis on the differentially expressed gene data to obtain the effectively expressed gene data.

[0121] Step 904: Based on the effective expressed gene data, perform network centrality weighting on the pathway genes to obtain the pathway weighted set.

[0122] Step 906: Perform equivalent count discretization analysis on the intersection of the effectively expressed gene data and the pathway weighted set to obtain the equivalent count index.

[0123] Step 908: Calculate the right-tailed cumulative probability of the hypergeometric distribution on the equivalent counting index to obtain the random probability of pathway enrichment.

[0124] Among them, detectability analysis is the process of screening out candidate genes that are insufficiently detected or unstable based on a unified background standard and detection threshold (such as minimum expression, sample coverage, mapping uniqueness, etc.).

[0125] Among them, the effectively expressed gene data is the set of genes that have been retained through detectability analysis, meet the requirements of detection sufficiency and stability, and are used for subsequent calculations.

[0126] Among them, network centrality weighting is a process that assigns weights to genes within a pathway based on centrality measures such as degree, betweenness, or eigenvectors on the pathway relationship graph in order to highlight key signals.

[0127] Among them, the pathway-weighted set is the set after pairing each gene in the pathway with its normalized centrality weight, which is used for subsequent intersection and counting with differentially expressed genes.

[0128] Among them, the equivalent count discretization analysis is a step that transforms the weighted continuous measure into a non-negative integer count according to a uniform scaling and rounding rule to meet the prerequisite of the discretization test.

[0129] The equivalent count index is a set of parameters for the four counts (background scale, path scale, difference scale, and intersection equivalent count) formed after discretization.

[0130] Among them, the calculation of the right tail cumulative probability of the hypergeometric distribution is a calculation to evaluate the probability of "intersection not less than the observed value" under the condition of four-item count, so as to quantify the significance of pathway enrichment.

[0131] Specifically, starting with differentially expressed gene data, a unified gene identifier (one-to-one mapping between Ensembl / Entrez / official symbols, eliminating one-to-many and unmapped items and recording the reason code) is established. Candidate genes are compared with "all genes in this batch of data that have passed quality control and been successfully quantified" to establish a detectability control. Subsequently, detectability analysis is performed item by item based on hypergeometric distribution test parameters, simultaneously excluding blacklisted genes (such as high-duplication / low-complexity / pseudogenes) and abnormal annotation entries. Lists of genes that pass and fail the analysis are output separately, and screening statistics (retention rate, trigger frequency of each threshold, comparison of the sufficiency of detection of differences between groups) and audit logs (thresholds, database version, timestamps) are generated, ultimately forming valid expressed gene data.

[0132] Given a fixed pathway library and version, we constructed a pathway subgraph by using effectively expressed gene data as nodes and integrating multi-source evidence (pathway edges, protein interactions, co-expression edges, and regulatory edges). We then standardized the dimensions and weighted the edge weights with confidence levels (source weighting, evidence level discounting, and low-confidence edge truncation). Next, we calculated the centrality indices (degree / betweenness / eigenvector / proximity, etc.) of each gene within the pathway and synthesized a single centrality score using robust methods (quantile merging or weighted geometric mean). We applied Winsor normalization and normalized the sum of all weights within the pathway to ensure all weights were non-negative and summed to 1. Simultaneously, we conducted sensitivity analysis (recalculating centrality by removing one edge / one node) and robustness assessment (Spearman consistency, resampling variance). The output is a weighted pathway set (genes and their weights for each pathway) and its accompanying metadata (graph source, weighting parameters, truncation ratio, and robustness indices).

[0133] For each pathway, the intersection of the effective expressed gene data and the pathway weighted set is taken. The weighted scores of the genes in the intersection are read, and a uniform scaling scale S is selected (e.g., ensuring that the expected sum of the weights within the pathway × S is within a computable range). The real-valued weights are converted into non-negative integer equivalent counts according to predetermined rules (fixed-scale rounding with controlled random rounding to correct for rounding bias, setting a lower limit of 1 and an upper limit of U to avoid extreme inflation). Under the same scaling scale, the three cardinal values ​​(background size, pathway coverage, and difference set size) are simultaneously discretized to ensure consistency with the discrete counting assumptions of the hypergeometric test. The discretization error is evaluated (relative error to the original weight sum, KL divergence approximation, and intersection information loss rate). If the error exceeds the threshold, the scaling scale S or the upper limit U is adaptively adjusted and the calculation is recalculated. Finally, the equivalent count indicators (the discretized background size, pathway size, difference size, and intersection equivalent count for each pathway) and error reports, a list of pruned genes, and parameter logs are output to ensure numerical stability and consistency with statistical assumptions.

[0134] Using equivalent count metrics, discrete right-tailed cumulative probability summation is performed on each pathway within a locked background set and pathway library version (calculating the probability of "observed equivalent intersection number or higher"). High-precision and logarithmic domain accumulation are enabled to prevent underflow in extremely low-probability or large-scale parameter scenarios, and an auditable strategy is implemented for boundary cases (pathway coverage of 0 or intersection of 0 directly returns 1; pathways exceeding the size limit are skipped and the reason is recorded). Input consistency is also checked (compatibility of the four counts, intersection not exceeding the pathway and difference size). The original random probability and accompanying metadata are written to the results table (pathway ID / name, database and version, four discrete counts, list of intersection genes, minimum overlap threshold, calculation path, and any warning information) to obtain the pathway enrichment random probability.

[0135] In this embodiment, detectability analysis of differentially expressed gene data is performed based on hypergeometric distribution test parameters to unify background criteria and eliminate under-detected or unstable entries, reducing systematic bias caused by sequencing depth and mapping uncertainty. Then, network centrality weighting is applied to pathway genes to explicitly model the contributions of information propagation and regulatory hubs within the pathway, improving the biological relevance and sensitivity of enrichment signals. Subsequently, the "effectively expressed gene data ∩ pathway weighted set" is discretized using equivalent counting, ensuring that the weighted continuous measure strictly aligns with the discrete counting premise of the hypergeometric test, suppressing statistical distortion caused by weight and scale mismatch from the source. Finally, right-tailed cumulative probability calculation is performed based on the equivalent count to obtain a dimensionally uniform, comparable, and numerically stable random probability of pathway enrichment, thereby significantly improving the accuracy, repeatability, and cross-dataset transferability of pathway determination while controlling false alarms.

[0136] In a specific embodiment, an application example of a method for determining pathway data affecting the progression of granulocytic sepsis, a malignant hematological disease, is as follows, and the corresponding overall technical roadmap is shown in the figure. Figure 3 As shown.

[0137] 1. Data Analysis Process Raw gene sequence data (raw reads) were obtained using high-throughput sequencing technology. This raw data is similar to a noisy "raw manuscript" containing billions of short sentences (reads), which cannot be directly used to interpret biological meaning. Therefore, a series of standardized computational analysis procedures were performed to extract reliable, topic-relevant biological conclusions from this massive amount of data.

[0138] Data quality control and preprocessing: ensuring the cleanliness and reliability of the "original manuscript". Objective: The raw sequencing data contains some low-quality, interfering, or uninterpretable sequences. The aim is to filter out these unreliable data like a rigorous "editor," ensuring that subsequent analyses are based on high-quality data.

[0139] Tools: SOAPnuke software Key parameter settings: lowQual=20: Sets a quality threshold for bases. Bases with a quality value (Qphred) ≤ 20 are considered "low-quality bases", with a sequencing error rate as high as 1% (meaning that 1 out of 100 such bases may be wrong).

[0140] nRate=0.005: Sets the tolerance for unknown bases "N". N represents a base that the sequencer cannot determine whether it is A / T / C / G. This parameter specifies that if the proportion of N in a sequence exceeds 0.5%, the entire sequence is considered unreliable.

[0141] qualRate=0.5: Sets the overall quality requirement for the entire sequence. If the number of low-quality bases (Qphred≤20) in a sequence exceeds 50% of the total sequence length, the entire sequence is discarded.

[0142] Specific filtering steps in data processing: (1) Remove sequences with adapters: Background: During whole transcriptome sequencing library preparation, known "connector" sequences need to be ligated to both ends of the RNA fragment to be sequenced, serving as the starting point for sequencing. Sometimes, these adapter sequences themselves are also sequenced, thus becoming contaminated with the target sequence.

[0143] Procedure: This step uses sequence alignment to accurately identify and remove "contaminated" data that still contains adapter sequences.

[0144] (2) Remove sequences with a high N ratio: Background: When the sequencing signal is weak or there is interference, some bases may not be recognized, which is represented by "N".

[0145] Operation: Remove sequences containing more than 0.5% of the total length of "N" bases to ensure the certainty of the information.

[0146] (3) Remove low-quality sequences: Background: A sequence consists of multiple bases, each with a quality value. If most bases in a sequence have low quality, the overall reliability of the sequence is extremely low.

[0147] Operation: Calculate the percentage of low-quality bases in each sequence and discard those sequences that account for more than 50%.

[0148] Output: After the above three rigorous screening processes, high-quality Clean reads of no less than 5GB in size were obtained. These clean reads are like a "clean document" that has been proofread and cleaned of typos and irrelevant content, laying a solid and reliable foundation for subsequent in-depth analysis.

[0149] 2. Quantitative gene analysis Objective: The higher the expression level of a gene, the more important its role in life activities is usually. The goal is to accurately calculate how many "copies" (i.e., expression levels) of each gene are produced in a sample.

[0150] Tools: Salmon software (version 1.10.0) Detailed explanation of the method and its principles: (1) Establishing a “library”: First, a known reference genome or transcriptome is needed as a “gene library”.

[0151] (2) Gene classification: Subsequently, the "clean reads" obtained in the previous step are quickly and accurately compared with the "gene library". This process determines which transcriptome each short read comes from.

[0152] (3) Counting and standardization: Raw count: The number of sequences aligned to each transcript. Genes with higher expression levels have more aligned sequences.

[0153] Standardization: Because genes vary in length, and the total number of sequences (total sequence count) may differ between samples, direct comparisons can lead to inconsistent results. This is similar to comparing the importance of a word in two books; you can't simply look at the frequency of occurrence because the total word count of the books differs. Therefore, standardization metrics such as TPM are used to standardize the data.

[0154] Gene quantification analysis was performed using Salmon software (v1.10.0), with TPM as its core output metric. TPM calculation followed its standard definition, meaning the software automatically performed dual corrections on the raw reads for gene length and total sequencing depth. The specific formula was: TPM = (Gene read count / Gene kilobase length) / (Total RPK of all genes) * 1,000,000. In this analysis, except for explicitly specified input / output files, all internal parameters affecting the quantification model, including but not limited to sequence alignment settings, fragment model learning, and bias correction, were maintained at the software's factory default settings to ensure the standardization and reproducibility of the analysis process.

[0155] Output: A table containing all genes and their expression levels (TPM and raw counts) in different samples.

[0156] 3. Differentially expressed gene analysis Objective: To compare gene expression matrices under different experimental conditions (e.g., disease group vs. healthy control group) to identify which genes showed significant changes in expression levels. These genes are likely key molecules leading to phenotypic differences.

[0157] Software used: edgeR Statistical and biological significance of the screening criteria: By employing two dimensions of criteria, we can jointly identify the true "differentially expressed genes": (1) Statistical significance: Indicator: Adjusted p-value (adj. p-value) < 0.05.

[0158] Explanation: The p-value is used to measure whether the difference is caused by random factors. Because tens of thousands of genes are tested simultaneously, the problem of "false positives" can occur (i.e., there is no difference in fact, but the test incorrectly identifies a difference). The adjusted p-value corrects the original p-value using statistical methods, controlling the overall error rate to within 5%. An adjective p-value < 0.05 means that there is more than a 95% confidence that the selected genes have genuinely different expression levels, rather than being random.

[0159] (2) Biological significance: Metric: |log2FoldChange|>1.

[0160] Explanation: FoldChange is the ratio of expression levels between two groups, while log2FoldChange is its base-2 logarithm. |log2FoldChange|>1 means that the fold change in expression level exceeds 2 (2^1 = 2). This ensures that the gene of interest has a sufficiently large change in expression to have potential biological significance, rather than a trivial change.

[0161] Output: A list of differentially expressed genes. These genes are those whose expression levels show both statistical and biological significance across different experimental groups, and are the core target for subsequent functional analysis.

[0162] 4. Functional enrichment analysis – revealing the “functional themes” involved by key genes Objective: The function of a single gene is limited, and life activities are usually accomplished by multiple genes working together to form so-called "pathways" or "functional modules." This study aims to answer the question: "Does this group of differentially expressed genes significantly cluster in a specific biological pathway?" thereby explaining the underlying mechanisms of phenotypic differences from a systemic and macroscopic perspective.

[0163] Methods: KEGG pathway enrichment analysis Tools and principles: Use the `phyper` function in R to perform a hypergeometric distribution test.

[0164] (1) Purpose of analysis Based on a hypergeometric distribution model, this method employs statistical overcharacterization analysis on gene sets composed of differentially expressed genes to identify whether they exhibit significant enrichment across different biological pathways. The aim is to provide potential functional explanations for differentially expressed gene sets from a systems biology perspective, thereby revealing the core biological processes in which they may synergistically participate.

[0165] (2) Analysis Principles and Implementation The hypergeometric distribution test is applicable to problems involving sampling without replacement from a finite population. In this scenario, its model parameters are defined as follows:

[0166] N: represents the total number of background genes, i.e., the set of all genes included in the quantitative analysis.

[0167] K: Represents the number of background genes belonging to a specific pathway or functional category (such as the KEGG pathway).

[0168] n: represents the size of the target gene set screened through differential expression analysis (i.e., the total number of differentially expressed genes).

[0169] k: represents the number of genes in the target gene set that simultaneously belong to this specific pathway.

[0170] The hypergeometric test assesses enrichment significance by calculating the cumulative probability that at least k pathway genes are observed in the target gene set under the assumption of random sampling. This probability value (P-value) is defined by the following formula:

[0171] P-value = 1 - Σ_{i=0}^{k-1} [ C(K, i) * C(NK, ni) / C(N, n) ] Where C(a, b) represents the binomial coefficient, that is, the number of combinations of choosing b elements from a elements.

[0172] The p-value quantifies the random probability of the target gene set being enriched in a specific pathway. To control for the false positive rate caused by multiple hypothesis testing, the calculated raw p-value was subsequently corrected for the false discovery rate, and an FDR < 0.05 was used as the threshold for statistical significance. Finally, all pathways that passed this significance threshold were identified as significantly enriched under the experimental conditions.

[0173] 5. Results Output Through the complete and rigorous computational biology analysis process described above, starting from the most basic sequencing data, and after data cleaning, gene quantification, differential screening, and functional annotation, biological pathways that are significant under all strict statistical criteria were finally identified.

[0174] The neutrophil extracellular trapping network (NETs) and NOD-like receptor signaling pathway were identified as the most significantly associated key pathways. This result provides crucial computational evidence and a clear direction for further research into elucidating the molecular mechanisms of this invention.

[0175] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0176] Based on the same inventive concept, this application also provides an apparatus for determining pathway data affecting the progression of granulocytic sepsis, a malignant hematological disease, in order to implement the aforementioned method for determining pathway data affecting the progression of granulocytic sepsis. For example... Figure 4 As shown, it includes: a data filtering module, a quantitative analysis module, a difference analysis module, and an enrichment analysis module. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more pathway data determination methods for the progression of granulocytic sepsis that affect the progression of malignant hematological diseases provided below can be found in the limitations of a pathway data determination method for the progression of granulocytic sepsis that affect the progression of malignant hematological diseases described above, and will not be repeated here.

[0177] The modules in the aforementioned pathway data determination device for the progression of agranulocytic sepsis, a malignant hematological disease, can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0178] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores server data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a method for determining pathway data affecting the progression of agranulocytic sepsis, a malignant hematological disease.

[0179] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0180] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0181] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0182] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the steps in the above method embodiments.

[0183] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0184] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0185] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0186] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for determining pathway data affecting the progression of granulocytic sepsis, a malignant hematological disease, characterized in that, The method includes: The initial gene sequence data was filtered to obtain clean gene sequence data; Gene quantitative analysis was performed on the clean gene sequence data to obtain gene expression level data; Gene differential expression analysis was performed on the gene expression level data to obtain differentially expressed gene data; Functional enrichment analysis was performed on the differentially expressed gene data to obtain significantly enriched pathway data.

2. The method according to claim 1, characterized in that, The step of performing differential gene expression analysis on the gene expression level data to obtain differentially expressed gene data includes: The gene expression data were subjected to expression morphology transformation to obtain a gene expression matrix; Statistical significance analysis was performed on the gene expression matrix to obtain the analytical expression matrix; Biological significance analysis was performed on the analyzed expression matrix to obtain the differentially expressed gene data.

3. The method according to claim 2, characterized in that, The process of performing expression morphology transformation on the gene expression data to obtain a gene expression matrix includes: The gene expression data were subjected to sample layer component calibration to obtain component-calibrated expression data; Graph regularized pseudo-count analysis was performed on the component calibration expression data to obtain continuous expression data; Based on the continuous expression data and the effective transcript length of the gene, the matrix expression values ​​are length recalibrated to obtain recalibrated expression data; The recalibrated expression data are subjected to variance stabilization transformation to obtain the gene expression matrix.

4. The method according to claim 1, characterized in that, The step of performing gene quantitative analysis on the clean gene sequence data to obtain gene expression level data includes: Gene length correction is performed on the clean gene sequence data to obtain gene length corrected data; The clean gene sequence data were subjected to total sequencing depth correction to obtain sample depth corrected data. The gene expression level data is obtained by dividing the gene length correction data by the sample depth correction data.

5. The method according to claim 4, characterized in that, Before performing gene length correction on the clean gene sequence data to obtain gene length-corrected data, the method further includes: The multiple alignment reads in the clean gene sequence data are probabilistically redistributed to obtain multiple alignment weight parameters. Prior estimation of fragment length is performed on the clean gene sequence data to obtain prior parameters for fragment length; Based on the clean gene sequence data and the GC content of the reference transcript, the read weights of the initial gene sequence data are calibrated for GC bias to obtain GC weight parameters. Probabilistic allocation learning is performed on the multi-positional reads in the clean gene sequence data to obtain multi-positional allocation parameters; The clean gene sequence data were subjected to library strand specificity determination to obtain strand specificity parameters.

6. The method according to claim 1, characterized in that, The functional enrichment analysis of the differentially expressed gene data yields significantly enriched pathway data, including: Set the hypergeometric distribution test parameters corresponding to the functional enrichment analysis; Based on the hypergeometric distribution test parameters, the differentially expressed gene data are subjected to a hypergeometric distribution test to obtain the significantly enriched pathway data.

7. The method according to claim 6, characterized in that, The step of performing a hypergeometric distribution test on the differentially expressed gene data based on the hypergeometric distribution test parameters to obtain the significantly enriched pathway data includes: Based on the hypergeometric distribution test parameters, the cumulative probability of pathway genes corresponding to the differentially expressed gene data is calculated to obtain the random probability of pathway enrichment. The false positive rate is corrected for the random probability of the enriched pathways to obtain the data of the significantly enriched pathways.

8. The method according to claim 7, characterized in that, The step of calculating the cumulative probability of pathway genes corresponding to the differentially expressed gene data based on the hypergeometric distribution test parameters to obtain the pathway enrichment random probability includes: Based on the hypergeometric distribution test parameters, detectability analysis is performed on the differentially expressed gene data to obtain effectively expressed gene data. Based on the effectively expressed gene data, the pathway genes are weighted by network centrality to obtain a pathway weighted set. An equivalent count discretization analysis is performed on the intersection of the effectively expressed gene data and the pathway weighted set to obtain an equivalent count index; The equivalent counting index is subjected to a hypergeometric distribution right-tail cumulative probability calculation to obtain the pathway enrichment random probability.

9. A device for determining pathway data affecting the progression of granulocytic sepsis, a malignant hematological disease, characterized in that, The device includes: The data filtering module is used to filter the initial gene sequence data to obtain clean gene sequence data; The quantitative analysis module is used to perform quantitative gene analysis on the clean gene sequence data to obtain gene expression level data; The differential expression analysis module is used to perform differential gene expression analysis on the gene expression level data to obtain differentially expressed gene data. The enrichment analysis module is used to perform functional enrichment analysis on the differentially expressed gene data to obtain significantly enriched pathway data.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.