Prediction method for expression regulation mechanism of disease marker

By using a four-quadrant neural network model and an adaptive quadrant partitioning optimization equation set, the problem of insufficient accuracy in predicting the expression regulation mechanism of disease biomarkers was solved. This enabled the effective capture of the multidimensional features and dynamic characteristics of gene expression, thereby improving the accuracy and adaptability of predictions.

CN121034409APending Publication Date: 2025-11-28QINGDAO RAISECARE BIOTECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511101856.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing technologies lack accuracy in predicting the regulatory mechanisms of disease biomarker expression, and fail to fully consider the multidimensional characteristics, dynamic properties, and complex regulatory relationships of gene expression.

Method used

A four-quadrant neural network model was adopted, which combined gene expression data and transcriptional regulation network. Feature analysis and prediction were performed by optimizing the equation system through adaptive quadrant partitioning, including contribution value calculation, statistical testing, feature integration and comprehensive output, to construct an expression regulation prediction model.

Benefits of technology

It improves the predictive accuracy and interpretability of disease biomarker expression regulation mechanisms, enhances the adaptability and robustness of the model, and can better capture the dynamic changes in gene expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034409A_ABST
    Figure CN121034409A_ABST
Patent Text Reader

Abstract

A disease marker expression regulatory mechanism prediction method comprises the following steps: acquiring gene expression data, performing quality control and standardization processing on the gene expression data to obtain standardized gene expression data, identifying differentially expressed genes as candidate markers, and predicting a disease marker expression regulatory mechanism. A transcriptional regulation and control network of the candidate markers is constructed, transcription factor data are generated, a multi-layer neural network model is established through a four-quadrant neural network model, and the four-quadrant neural network model comprises a contribution value calculation layer, a statistical test layer, a quadrant division layer, a feature integration layer and a comprehensive output module. The four-quadrant neural network model adopts an adaptive quadrant division optimization equation set to carry out feature analysis and prediction on the candidate markers; compared with the prior art, the method is remarkably improved in the aspects of feature integration, dynamic modeling, self-adaptive optimization and the like, and important theoretical basis and technical support are provided for disease diagnosis, prognosis evaluation and treatment strategy formulation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of disease biomarker technology, and more specifically, relates to a method for predicting the expression regulation mechanism of disease biomarkers. Background Technology

[0002] Biomarkers play an increasingly important role in clinical medicine as crucial evidence for disease diagnosis, prognostic assessment, and treatment strategy development. In recent years, with the rapid development of gene sequencing and spatial omics technologies, researchers have been able to explore the expression and regulatory mechanisms of disease-related biomarkers in a more comprehensive and in-depth manner. However, screening out clinically valuable biomarkers from massive amounts of gene expression data remains a challenging problem.

[0003] Currently, the main methods used for biomarker discovery and prediction are as follows: 1) Differential expression analysis-based methods: These methods compare gene expression differences between disease samples and healthy control samples to screen for candidate biomarker genes with significant changes in expression levels. However, this method relies solely on a single statistical test and easily overlooks the complex regulatory relationships and dynamic characteristics between genes. 2) Machine learning-based methods: These methods construct various predictive models, such as regression analysis, support vector machines, and artificial neural networks, to predict the expression levels of candidate biomarkers. These methods can effectively utilize multidimensional feature information, but the complexity and interpretability of the models are poor, making it difficult to fully leverage the biological knowledge contained in gene expression data. 3) Bionetwork analysis-based methods: These methods construct gene regulatory networks, protein interaction networks, etc., to analyze the function and mechanism of action of candidate biomarkers from a systems biology perspective. However, existing bionetwork analysis methods are usually limited to static topological structure analysis and cannot fully capture dynamic expression regulation processes.

[0004] In summary, existing technologies suffer from insufficient accuracy in predicting the regulatory mechanisms of disease biomarker expression. Summary of the Invention

[0005] In view of this, the present invention provides a method for predicting the expression regulation mechanism of disease biomarkers, comprising the following steps: acquiring gene expression data, performing quality control and standardization processing on the gene expression data to obtain standardized gene expression data, identifying differentially expressed genes as candidate biomarkers, constructing a transcriptional regulation network of the candidate biomarkers and generating transcription factor data, establishing a multi-layer neural network model using a four-quadrant neural network model, wherein the four-quadrant neural network model includes a contribution value calculation layer, a statistical test layer, a quadrant partitioning layer, a feature integration layer, and a comprehensive output module, wherein the four-quadrant neural network model uses an adaptive quadrant partitioning optimization equation set to perform feature analysis and prediction on the candidate biomarkers, performing functional enrichment analysis on the candidate biomarkers, and constructing an expression regulation prediction model for the candidate biomarkers.

[0006] The gene expression data includes transcriptome sequencing data of patient samples and healthy control samples, with a sample size of no less than 50 samples in each sample. The quality control and standardization process includes removing sample data with a sequencing depth of less than 10 million reads and a quality score of less than 30. The differentially expressed genes are genes with an expression fold change of more than 2 and a statistical significance test probability value of less than 0.01.

[0007] The transcriptional regulatory network identifies transcription factors that have a direct regulatory relationship with the candidate biomarkers based on chromatin immunoprecipitation sequencing data and transcription factor binding site data.

[0008] The number of nodes in the four-quadrant neural network model is as follows: the number of nodes in the contribution value calculation layer is 1.5 times the input feature dimension; the number of nodes in the statistical test layer is the same as the number of nodes in the contribution value calculation layer; the number of nodes in the quadrant division layer is 4; the number of nodes in the feature integration layer is 0.5 times the input feature dimension; and the number of nodes in the comprehensive output module is 1.

[0009] The adaptive quadrant partitioning optimization equation set includes a contribution equation, a significance equation, a weight allocation equation, and an integrated output equation. The inputs to the contribution equation include the standardized gene expression data, the weight matrix, the tissue-specific coefficient matrix, the time-series expression parameter matrix, and the dimension transformation coefficient matrix.

[0010] The inputs to the significance equation include the standardized gene expression data, the healthy control sample data, the significance test level parameter, the homogeneity of variance coefficient, the normal distribution parameter of the data, the degree of freedom adjustment coefficient, and the confidence interval range parameter. The inputs to the weight allocation equation include the contribution value vector, the statistical significance test probability value, the smoothing factor coefficient, the weight decay parameter, the iteration update coefficient, the quadrant boundary adjustment parameter, and the feature importance coefficient.

[0011] The inputs to the integrated output equation include feature quadrant assignment labels, weight coefficients, nonlinear mapping parameters, activation function temperature coefficients, gradient scaling factors, learning rate adjustment parameters, and bias term correction coefficients.

[0012] Specifically, the four-quadrant neural network model is evaluated using a cross-validation method. Models with a prediction accuracy higher than 85% and a Matthew correlation coefficient greater than 0.7 are selected. The four-quadrant neural network model is evaluated using independent validation set data, and the sample size of the independent validation set data is not less than 30% of the gene expression data.

[0013] Specifically, the quadrant partitioning layer assigns features whose contribution value is greater than a contribution value threshold and whose statistical significance test probability value is less than a significance threshold to the first quadrant; the quadrant partitioning layer assigns features whose contribution value is greater than a contribution value threshold and whose statistical significance test probability value is greater than a significance threshold to the second quadrant; the quadrant partitioning layer assigns features whose contribution value is less than a contribution value threshold and whose statistical significance test probability value is greater than a significance threshold to the third quadrant; and the quadrant partitioning layer assigns features whose contribution value is less than a contribution value threshold and whose statistical significance test probability value is less than a significance threshold to the fourth quadrant.

[0014] The contribution value calculation layer receives the standardized gene expression data and the transcription factor data, calculates the contribution value of each input feature through the weight matrix, transforms the contribution value using a non-linear activation function, and outputs the contribution value vector. The statistical testing layer receives the contribution value vector and performs statistical significance analysis in conjunction with the standardized gene expression data, calculates the statistical significance test probability value, and outputs the statistical significance test vector.

[0015] The feature integration layer receives the feature quadrant attribution label, assigns weight coefficients to features in different quadrants and performs weighted integration, and outputs the integrated feature vector. The weight coefficients are calculated by the weight allocation equation based on the contribution value vector, the statistical significance test probability value, the smoothing factor coefficient, the weight decay parameter, the iterative update coefficient, the quadrant boundary adjustment parameter, and the feature importance coefficient.

[0016] The integrated output module receives the integrated feature vector and performs feature mapping through the fully connected layer, calculates the probability using the nonlinear activation function, and generates the predicted expression level of the candidate marker based on the nonlinear mapping parameters, the activation function temperature coefficient, the gradient scaling factor, the learning rate adjustment parameter, and the bias term correction coefficient.

[0017] The advantages of this invention are:

[0018] First, the four-quadrant neural network model fully considers the need for multi-dimensional feature integration in the biomarker discovery process. By dividing the feature space into four quadrants, it achieves hierarchical processing and optimized integration of different types of features. This innovative design not only improves the interpretability of the model but also enhances the reliability of the prediction results.

[0019] Secondly, the method of this invention introduces time integral and derivative terms, which can better capture the dynamic changes in gene expression. Compared with traditional methods, this invention has significant advantages in handling temporal and dynamic characteristics, providing a solid theoretical foundation for the dynamic prediction of biomarkers.

[0020] Furthermore, this invention achieves dynamic adjustment of feature weights through an adaptive quadrant partitioning optimization equation set. This adaptive mechanism improves the model's adaptability, enabling it to better cope with complex and ever-changing biological contexts.

[0021] Furthermore, the methodological innovations of this invention are also reflected in: the introduction of multiple feature matrices such as tissue-specific coefficients and temporal expression parameters, enabling multi-level characterization of gene expression regulation; the design of complex weight allocation equations, considering the smoothness, sparsity, and time-varying nature of feature selection; and the integration of output equations to achieve precise control and optimization of prediction results. These innovations further enhance the applicability and robustness of the method of this invention.

[0022] In summary, this invention, through its innovative four-quadrant neural network design and comprehensive feature integration strategy, achieves accurate prediction of the regulatory mechanisms of disease biomarker expression, thus solving the problem of insufficient accuracy in predicting the regulatory mechanisms of disease biomarker expression in existing technologies. Attached Figure Description

[0023] Figure 1 A flowchart of the method provided by the present invention;

[0024] Figure 2 This is a gene expression distribution map from Example 2;

[0025] Figure 3 This is a differential expression volcano plot from Example 2;

[0026] Figure 4 This is a heatmap of transcription factor regulation in Example 2;

[0027] Figure 5 This is a performance evaluation graph of the model in Example 2. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0029] like Figure 1 As shown, the present invention includes the following operational steps:

[0030] S10. Obtain gene expression data related to the target disease, wherein the gene expression data includes transcriptome sequencing data of patient samples and healthy control samples, and the sample size is not less than 50 for each sample.

[0031] S20. Perform quality control and standardization on the gene expression data, remove sample data with sequencing depth less than 10 million reads and quality score less than 30, and perform standardization to obtain standardized gene expression data.

[0032] S30. Differential expression analysis was used to identify differentially expressed genes related to the target disease, and genes with an expression fold change greater than 2 and a statistical significance test probability value less than 0.01 were selected as candidate biomarkers.

[0033] S40. Construct the transcriptional regulatory network of the candidate biomarkers, and based on chromatin immunoprecipitation sequencing data and transcription factor binding site data, identify the transcription factors that have a direct regulatory relationship with the candidate biomarkers and generate transcription factor data.

[0034] S50. A multi-layer neural network model is established using a four-quadrant neural network model. The number of nodes in the four-quadrant neural network model is as follows: the number of nodes in the contribution value calculation layer is 1.5 times the input feature dimension; the number of nodes in the statistical test layer is the same as the number of nodes in the contribution value calculation layer; the number of nodes in the quadrant partitioning layer is 4; the number of nodes in the feature integration layer is 0.5 times the input feature dimension; and the number of nodes in the comprehensive output module is 1. The four-quadrant neural network model uses an adaptive quadrant partitioning optimization equation set to perform feature analysis and prediction on the candidate markers.

[0035] S60. Evaluate the prediction accuracy of the four-quadrant neural network model based on the cross-validation method, and select the model with a prediction accuracy higher than 85% and a Matthew correlation coefficient greater than 0.7.

[0036] S70. Evaluate the generalization ability of the four-quadrant neural network model using independent validation set data, wherein the sample size of the independent validation set data is not less than 30% of the gene expression data;

[0037] S80. Perform functional enrichment analysis on the candidate biomarkers to determine the biological pathways and molecular functions involved by the candidate biomarkers;

[0038] S90. Construct an expression regulation prediction model for the candidate biomarkers. The expression regulation prediction model integrates the transcription factor binding site data, the chromatin immunoprecipitation sequencing data, and the epigenetic modification data to predict the expression level of the candidate biomarkers under different pathological conditions.

[0039] The four-quadrant neural network model includes a contribution value calculation layer, a statistical testing layer, a quadrant partitioning layer, a feature integration layer, and a comprehensive output module. The contribution value calculation layer receives the standardized gene expression data and the transcription factor data, calculates the contribution value of each input feature using a weight matrix, transforms the contribution value using a non-linear activation function, and outputs a contribution value vector. The statistical testing layer receives the contribution value vector and performs statistical significance analysis in conjunction with the standardized gene expression data, calculates the statistical significance test probability value, and outputs a statistical significance test vector. The quadrant partitioning layer simultaneously receives the contribution value vector and the statistical significance test vector, partitions all features into quadrants based on a preset threshold, and assigns features with contribution values ​​greater than the contribution value threshold and statistical significance test probability values ​​less than the significance threshold to the first quadrant. The quadrant partitioning layer assigns features whose contribution value is greater than a contribution value threshold and whose statistical significance test probability value is greater than a significance threshold to the second quadrant; features whose contribution value is less than a contribution value threshold and whose statistical significance test probability value is greater than a significance threshold to the third quadrant; and features whose contribution value is less than a contribution value threshold and whose statistical significance test probability value is less than a significance threshold to the fourth quadrant, and outputs the feature quadrant assignment labels. The feature integration layer receives the feature quadrant assignment labels, assigns the weight coefficients to the features in different quadrants and performs weighted integration, outputting an integrated feature vector. The comprehensive output module receives the integrated feature vector and performs feature mapping through a fully connected layer, uses the nonlinear activation function to perform probability calculation, and generates the predicted expression level value of the candidate biomarker.

[0040] The adaptive quadrant partitioning optimization equation set includes contribution equation, significance equation, weight allocation equation, and integrated output equation;

[0041] The contribution equation is used to calculate the contribution value. The inputs include the standardized gene expression data, the weight matrix, the tissue-specific coefficient matrix obtained from the tissue expression profile database, the time-series expression parameter matrix obtained from the time-series expression data, and the dimension transformation coefficient matrix obtained from principal component analysis. The output is the contribution value vector.

[0042] The significance equation is used to calculate the statistical significance test probability value. The inputs include the standardized gene expression data, the healthy control sample data, the significance test level parameter obtained from the significance level setting, the homogeneity of variance coefficient obtained from the analysis of variance, the normality distribution parameter of the data obtained from the normality test, the degree of freedom adjustment coefficient obtained from the sample size, and the confidence interval range parameter obtained from the confidence interval analysis. The output is the statistical significance test vector.

[0043] The weight allocation equation is used to determine the weight coefficients. The inputs include the contribution value vector, the statistical significance test probability value, the smoothing factor coefficient obtained from data smoothing, the weight decay parameter obtained from weight optimization, the iterative update coefficient obtained from iterative calculation, the quadrant boundary adjustment parameter obtained from boundary optimization, and the feature importance coefficient obtained from feature selection. The output is the weight coefficients.

[0044] The integrated output equation is used to generate the predicted expression level value. The inputs include the feature quadrant attribution label, the weight coefficient, the nonlinear mapping parameter obtained from nonlinear optimization, the activation function temperature coefficient obtained from temperature adjustment, the gradient scaling factor obtained from gradient optimization, the learning rate adjustment parameter obtained from learning rate optimization, and the bias term correction coefficient obtained from bias optimization. The output is the predicted expression level value of the candidate marker.

[0045] The specific implementation methods of the above steps are described in detail below.

[0046] The specific implementation of step S10 is as follows: First, obtain gene expression data of the target disease from relevant gene expression databases or clinical samples. The obtained gene expression data includes transcriptome sequencing data of patient samples and healthy control samples, with no fewer than 50 samples in each group. This sample size ensures the statistical significance of subsequent analyses.

[0047] Secondly, the acquired gene expression data underwent quality control and standardization. This included the following steps: 1) Removing sample data with a sequencing depth of less than 10 million reads. Insufficient sequencing depth leads to inaccurate expression level estimation. 2) Removing sample data with a quality score below 30. Sample data below Q30 are of poor quality and may introduce significant noise. 3) Standardizing the sample data that passed quality control. The standardization method used was z-score standardization, which centers and normalizes the data by subtracting the sample mean and dividing by the sample standard deviation. A shift correction parameter in the range of 0.01 to 0.1 was also introduced to avoid computational problems caused by zero values.

[0048] Through the above quality control and standardization processes, high-quality standardized gene expression data can be obtained, laying the foundation for subsequent differential expression analysis. The purpose of this step is to ensure that the gene expression data used in the analysis is reliable and comparable.

[0049] The specific implementation of step S20 is as follows: Differential expression analysis is used to identify differentially expressed genes related to the target disease. Specifically, this includes the following steps: 1) Calculating the fold change in expression of each gene between disease samples and healthy control samples. The formula for calculating the fold change (FC) is: Where Ed and E c α represents the gene expression values ​​in the disease sample and the control sample, respectively, with α being the expression correction coefficient in the range of 0 to 0.2. 2) Perform a statistical significance test on the fold change in expression. Calculate the p-value using methods such as the t-test or ANOVA, with the formula: Where t is the test statistic, and β is the probability correction parameter in the range of 0.001 to 0.01. 3) Genes with an expression fold change greater than 2 and a statistical significance test p-value less than 0.01 are selected as candidate biomarkers. This screening criterion can effectively select differentially expressed genes that are highly associated with the disease.

[0050] The purpose of this step is to initially screen differentially expressed genes that may serve as disease biomarkers from a large amount of gene expression data, laying the foundation for subsequent construction of transcriptional regulatory networks and establishment of predictive models.

[0051] The specific implementation of step S30 is as follows: Based on the screened candidate biomarkers, a transcriptional regulatory network is constructed. The construction of the transcriptional regulatory network mainly includes the following steps: 1) Collecting chromatin immunoprecipitation sequencing (ChIP-Seq) data to identify transcription factors that have a direct regulatory relationship with the candidate biomarkers. ChIP-Seq technology can detect the sites where transcription factors actually bind to chromatin. 2) Integrating transcription factor binding site data to determine which transcription factors directly regulate the expression of candidate biomarkers. These direct regulatory relationships constitute the edges of the transcriptional regulatory network. 3) Integrating the identified transcription factors and their regulatory relationships with candidate biomarkers into transcriptional regulatory network data. Nodes in the transcriptional regulatory network represent candidate biomarkers and their regulating transcription factors, and edges represent regulatory relationships.

[0052] The purpose of this step is to gain a deeper understanding of the expression regulation mechanisms of candidate biomarkers. Transcriptional regulatory networks can visually demonstrate which transcription factors are involved in the expression regulation of candidate biomarkers, providing a foundation for the subsequent development of expression regulation prediction models.

[0053] The specific implementation of step S40 is as follows: A multi-layer neural network model is established using a four-quadrant neural network model to perform feature analysis and expression level prediction for candidate biomarkers. The specific structure of the four-quadrant neural network model is as follows: 1) Contribution value calculation layer: Receives standardized gene expression data and transcription factor data, calculates the contribution value of each input feature using a weight matrix, and transforms it using a non-linear activation function to output a contribution value vector. The number of nodes in the contribution value calculation layer is 1.5 times the dimension of the input features. 2) Statistical test layer: Receives the contribution value vector and performs statistical significance analysis in conjunction with standardized gene expression data, calculates the statistical significance test p-value, and outputs a statistical significance test vector. The number of nodes in the statistical test layer is the same as that in the contribution value calculation layer. 3) Quadrant partitioning layer: Receives both the contribution value vector and the statistical significance test vector, and partitions all features into quadrants according to preset contribution value thresholds and significance thresholds. This layer contains 4 nodes, corresponding to the 4 quadrants. 4) Feature integration layer: Receives the feature quadrant assignment labels, assigns corresponding weight coefficients to features in different quadrants, performs weighted integration, and outputs an integrated feature vector. The number of nodes in the feature integration layer is 0.5 times the dimension of the input features. 5) Comprehensive output module: Receives the integrated feature vector, performs feature mapping through a fully connected layer, and calculates the predicted expression level of candidate markers using a non-linear activation function. This module contains 1 node.

[0054] The four-quadrant neural network model employs an adaptive quadrant partitioning optimization equation set to perform feature analysis and prediction of candidate biomarkers. The optimization equation set includes contribution equations, significance equations, weight allocation equations, and integrated output equations. These equations fully consider the biological characteristics of gene expression, such as tissue specificity and temporal variation trends, enabling the model to better capture the expression regulatory mechanisms of candidate biomarkers.

[0055] The purpose of this step is to establish a neural network model that integrates feature analysis and expression level prediction, which can effectively utilize gene expression data and transcriptional regulatory network information to predict the expression levels of candidate biomarkers under different pathological conditions.

[0056] The specific implementation of step S50 is as follows: The predictive performance of the four-quadrant neural network model is evaluated using cross-validation. Specifically, this includes: 1) Randomly dividing the original gene expression data into a training set and a validation set. The training set is used for model training, and the validation set is used for performance evaluation. 2) Using 5-fold cross-validation, the training-validation process is repeated, and the prediction accuracy and Matthew correlation coefficient of the model on the validation set are calculated. 3) The model with a prediction accuracy higher than 85% and a Matthew correlation coefficient greater than 0.7 is selected as the final model. Such performance metrics ensure that the model has high prediction accuracy and generalization ability.

[0057] In addition, independent validation set data is needed to further evaluate the generalization ability of the selected model. The sample size of the independent validation set should be no less than 30% of the original gene expression data. By validating the model's predictive performance on independent samples, its performance in real-world applications can be more realistically evaluated.

[0058] The purpose of this step is to use rigorous cross-validation and independent validation methods to screen out four-quadrant neural network models with excellent predictive performance, providing a reliable foundation for subsequent functional analysis and expression regulation prediction.

[0059] The specific implementation of step S60 is as follows: Functional enrichment analysis is performed on the screened candidate biomarkers to determine the biological pathways and molecular functions they participate in. Specifically, this includes: 1) Querying the biological pathways and molecular functions involved by the candidate biomarkers in relevant bioinformatics databases (such as KEGG, GO, etc.) based on their gene IDs. 2) Calculating the significance p-value for each pathway or function using statistical methods such as hypergeometric tests, and screening out biological processes highly correlated with the candidate biomarkers. 3) Integrating the significantly enriched pathway and functional information into intuitive visualizations, such as bubble charts or bar charts, to help understand the mechanisms of action of candidate biomarkers in disease development and progression.

[0060] The purpose of this step is to conduct an in-depth analysis of candidate biomarkers from the perspective of gene function, understand their specific roles in disease-related biological processes, and provide more biological evidence for subsequent expression regulation prediction.

[0061] The specific implementation of step S70 is as follows: Constructing a predictive model for the expression regulation of candidate biomarkers. This model integrates transcription factor binding site data, chromatin immunoprecipitation sequencing data, and epigenetic modification data to predict the expression levels of candidate biomarkers under different pathological conditions. Specifically, it includes: 1) Extracting information on transcription factors that directly interact with candidate biomarkers from the transcriptional regulatory network. 2) Collecting chromatin site data that binds to these transcription factors to identify transcription factor binding sites that affect candidate biomarker expression. 3) Integrating epigenetic modification data, such as histone methylation / acetylation, to analyze how these epigenetic modifications regulate candidate biomarker expression. 4) Constructing a predictive model that comprehensively considers transcription factor regulation and epigenetic modification, capable of predicting the expression levels of candidate biomarkers under different physiological and pathological conditions.

[0062] The purpose of this step is to establish a more complete expression regulation prediction model that considers not only the direct regulation of transcription factors but also the influence of epigenetic regulation, making the prediction results more reliable and targeted, and providing a theoretical basis for subsequent clinical applications.

[0063] The calculation process involved in this invention is described in detail below.

[0064] 1. The standardized processing procedure is expressed as follows: In the formula, X norm σ is the standardized gene expression value; X is the original gene expression value; μ is the sample mean; σ is the sample standard deviation; δ is the offset correction parameter, ranging from 0.01 to 0.1.

[0065] 2. The differential expression analysis process is expressed as follows: In the formula, FC represents the multiple change; E d E represents the gene expression value in the disease sample. c α represents the gene expression value in the control sample; α is the expression level correction coefficient, ranging from 0 to 0.2; P is the statistical significance test probability value; t is the test statistic; β is the probability correction parameter, ranging from 0.001 to 0.01.

[0066] 3. The contribution equation is expressed as follows: In the formula, C is the feature contribution vector; W is the weight matrix; X norm γ represents the standardized gene expression value; T is the tissue-specific coefficient matrix; S is the time-series expression parameter matrix; D is the dimension transformation coefficient matrix; t is the time variable; t0 is the start time point; t1 is the end time point; γ is the contribution value correction vector. The first partial derivative of gene expression value with respect to time represents the rate of change in expression level; The second partial derivative of gene expression value with respect to time represents the acceleration of the change in expression level.

[0067] in:

[0068] 4. The significance equation is expressed as follows: In the formula, P s F is the probability value for statistical significance testing. obs The F-statistic is the number of observed values; F crit η is the critical F-value; V is the homogeneity of variance coefficient; N is the normality parameter of the data; L is the significance test level parameter; A is the degree of freedom adjustment coefficient; I is the confidence interval range parameter; η is the significance correction parameter, ranging from 0.001 to 0.01.

[0069] 5. The weighting equation is expressed as follows: In the formula, W q P represents the weighting coefficients; C represents the contribution value vector; P represents the weighting coefficients. s λ is the probability value for statistical significance testing; R is the smoothing factor coefficient; λ is the weight decay parameter; k is the number of iterations; B is the quadrant boundary adjustment parameter; U is the feature importance coefficient; F is the feature selection coefficient; θ is the weight correction parameter, ranging from 0.01 to 0.1.

[0070] 6. The integrated output equation is expressed as follows: In the formula, O represents the predicted expression level; σ represents the nonlinear activation function; Q represents the feature quadrant assignment label; W q T represents the weighting coefficients; M represents the nonlinear mapping parameters; T represents the weighting coefficients. a L is the temperature coefficient of the activation function; G is the gradient scaling factor; L is the temperature coefficient of the activation function. r b is the learning rate adjustment parameter; ω is the bias term correction coefficient; ω is the output correction parameter. ΔW is the partial derivative of the predicted value with respect to the weights, representing the weight sensitivity; q This refers to the weight update amount; Let be the mixed partial derivative of the predicted value with respect to weights and time, representing the time evolution of weight sensitivity.

[0071] The principles and significance of constructing the above equations are explained below:

[0072] 1. The standardization process equations take into account the central tendency and dispersion of the data distribution, achieve data comparability through mean centering and variance normalization, and introduce offset correction parameters to avoid computational problems caused by zero values;

[0073] 2. The differential expression analysis equation consists of two parts: fold change calculation and significance test. A correction coefficient is used to account for the effects of technical noise and biological variation, and the integral form reflects the essence of hypothesis testing.

[0074] 3. The contribution equation adopts the product form of multiple feature matrices, which reflects the comprehensive effect of different biological features on gene expression regulation. Matrix operations improve computational efficiency; among them: (1) time integral term The cumulative changes in gene expression over the entire time interval were calculated, reflecting the dynamic regulatory process; (2) Second derivative term It describes the acceleration of gene expression changes and reflects the changing trend of regulatory intensity; (3) The introduction of these terms enables the equation to capture the dynamic characteristics of gene expression and the changing patterns of regulatory modes;

[0075] 4. The significance equation is based on the principle of analysis of variance, takes into account the influence of multiple statistical parameters, and uses the ratio form to reflect the relationship between the observed value and the critical value;

[0076] 5. The weight allocation equation uses a decay form to control overfitting, taking into account the influence of feature importance and boundary adjustment, and the iterative update mechanism improves the model's adaptability;

[0077] 6. The integrated output equation uses nonlinear activation to fit complex patterns, taking into account the effects of temperature coefficient and learning rate. Gradient scaling and bias correction improve the accuracy of prediction; among which: (1) partial derivative terms (1) Describes the sensitivity of the predicted value to changes in weights, used to optimize weight updates; (2) Integral term of mixed partial derivatives The accumulation of weight sensitivity changes over time reflects the model’s dynamic adaptability; (3) The introduction of these terms enhances the model’s ability to respond to parameter changes and capture time series features.

[0078] The parameters were obtained as follows: gene expression data were obtained through RNA sequencing experiments using Illumina NovaSeq as the sequencing platform, with a sequencing depth of at least 10M reads; tissue specificity coefficients were obtained from the GTEx database, and time-series expression parameters were determined through time-series experiments; dimensionality transformation coefficients were calculated using principal component analysis, with a variance explanation threshold of 85%; the weight matrix was optimized using the backpropagation algorithm, with an initial learning rate of 0.01; and statistical parameters were calculated using bootstrap sampling, with 1000 resampling attempts. The result is obtained by taking the derivative after spline interpolation, with at least 10 interpolation nodes. It is obtained by numerically differentiating the first derivative using the central difference method; Obtained through automatic differentiation technology, using a reverse mode; The calculation is approximated using the finite difference method, with a time step of 0.01.

[0079] Specifically: Gene expression data were obtained through RNA sequencing experiments. The specific steps included: firstly, total RNA was extracted using TRIzol reagent and its quality was verified, requiring an RNA integrity value (RIN) greater than 8; and secondly, OD was detected using NanoDrop. 260 / OD 280 The ratio was between 1.8 and 2.0. Then, sequencing libraries were constructed using the NEBNext Ultra II RNALibrary Prep Kit. Library fragments of 350-450 bp were selected using a DNA fragment selection kit, and concentration quantification was performed using Qubit. Next, high-throughput sequencing was performed using the Illumina NovaSeq platform in PE150 mode. The raw data volume of each sample was no less than 6G. The obtained raw data was quality assessed using FastQC software, and then adapter removal and quality filtering were performed using Trimmomatic software. The Q30 ratio after filtering was required to be no less than 90%. Finally, the Clean Data was aligned to the reference genome to calculate gene expression values.

[0080] Tissue specificity coefficients were obtained based on tissue expression profile data from the GTEX database. First, RNA sequencing data from all normal tissues in the database were downloaded, and the raw data were standardized. Then, the mean expression value of each gene in each tissue was calculated, and the tissue specificity score was calculated. Among them TS i x is the tissue-specific score for gene i. j x represents the expression value of the gene in tissue j. max The maximum expression value of the gene in all tissues is given by n, where n is the number of tissues. Finally, a specificity coefficient matrix is ​​constructed based on the tissue specificity score, and the matrix elements are obtained by normalizing the scores.

[0081] Temporal expression parameters were obtained through time-series experiments. The experimental setup included six time points: 0h, 6h, 12h, 24h, 48h, and 72h. Three biological replicates were set up for each time point. Sample processing and RNA extraction methods were the same as those used for gene expression data acquisition. After obtaining the expression data, temporal variation parameters were calculated. Where TP is a time-varying parameter, x i Let be the expression value at time point i, and m be the number of time points; at the same time, the trend coefficient, periodicity coefficient and volatility coefficient of the time series expression are calculated, and finally the time series parameter matrix is ​​constructed.

[0082] The dimensionality transformation coefficients were obtained through principal component analysis. First, principal component analysis was performed on the standardized gene expression matrix to calculate the eigenvalues ​​and eigenvectors of the covariance matrix. Then, principal components with a cumulative contribution rate of 85% were selected to construct the transformation matrix D = VΛ. 1 / 2 , where V is the eigenvector matrix and Λ is the eigenvalue diagonal matrix; finally, the transformation coefficients of each gene in the principal component space are calculated through feature projection, thus forming the dimension transformation coefficient matrix.

[0083] The weight matrix optimization employs the backpropagation algorithm of deep learning. The specific steps are as follows: First, the weights are randomly initialized using a normal distribution N(0, 0.01) to generate initial values; then, forward propagation is performed to calculate the loss function, and backpropagation is performed to calculate the gradient. Finally, the weights are updated. Where L is the loss function and η is the learning rate, with an initial value of 0.01. The Adam optimizer is used to dynamically adjust the learning rate, and iterative training is performed until the loss function converges.

[0084] Statistical parameters were obtained using the bootstrap resampling method, with 1000 resampling cycles. For each resampling, the corresponding statistics were calculated, and an empirical distribution was constructed. The homogeneity of variance coefficient V was obtained through the Levene test, the normality parameter N was obtained through the Shapiro-Wilk test, and the degrees of freedom adjustment coefficient was calculated based on the sample size. The confidence interval parameter I was determined using the critical value of the t-distribution; the significance level of all statistical tests was set to 0.05.

[0085] The derivative calculation includes the numerical calculation of the first and second derivatives. The first derivative is calculated by first interpolating the discrete time-point representation data using cubic spline interpolation, setting no fewer than 10 nodes, and then using the central difference formula. The step size h is set to 0.01; the second derivative is based on the first derivative result, and the central difference method is also used. Perform the calculation.

[0086] Partial derivative calculation is divided into two parts: weight sensitivity calculation and mixed partial derivative calculation. The weight sensitivity calculation uses automatic differentiation technology, which tracks gradient propagation and accumulates gradients by constructing a computational graph. Where l is the number of network layers, y i This is the output of the i-th layer; the mixed partial derivatives are calculated using a time discretization method with a step size of 0.01, through the difference formula. The calculations are performed, and finally, the trapezoidal rule is used for numerical integration.

[0087] In the field of predicting the regulatory mechanisms of disease biomarker expression, traditional methods mainly include the following types:

[0088] Differential expression analysis methods mainly include statistical methods such as t-test, analysis of variance (ANOVA), and rank-sum test. These methods identify differentially expressed genes by calculating the expression differences between the disease group and the control group and performing significance tests. However, these methods often ignore the temporal and dynamic characteristics of gene expression regulation and are easily affected by outliers and batch effects. In addition, these methods cannot consider the interaction relationships between genes, resulting in the obtained biomarkers lacking biological connections.

[0089] Machine learning methods include algorithms such as support vector machines (SVM), random forests (RF), and artificial neural networks (ANN). These methods predict gene expression patterns by building complex nonlinear models. However, traditional machine learning methods often treat genes as independent features, failing to fully utilize the structural information of gene regulatory networks. At the same time, they are prone to overfitting when dealing with high-dimensional features, and the interpretability of the models is poor, making it difficult to reveal the underlying biological mechanisms.

[0090] Network analysis methods mainly include co-expression network analysis, regulatory network reconstruction, and pathway enrichment analysis. These methods study expression regulation mechanisms by constructing networks of connections between genes. However, traditional network analysis methods often rely on static network structures, which cannot reflect the dynamic characteristics of gene expression regulation. Furthermore, they have low computational efficiency when dealing with large-scale networks, and the network construction process is easily affected by noisy data.

[0091] Regression-based modeling methods include linear regression, generalized linear models, and time series analysis. These methods attempt to establish mathematical relationships between gene expression and various influencing factors. However, traditional regression methods often assume simple linear relationships between variables, making it difficult to describe complex biological regulatory processes. They also perform poorly in handling nonlinear relationships and interactions, and are sensitive to outliers.

[0092] Ensemble learning methods, such as Boosting and Bagging, improve prediction performance by combining multiple base models. However, traditional ensemble methods lack biological guidance in feature selection and model integration, making it difficult to ensure that the selected features have clear biological significance. Furthermore, the model training process is computationally complex and requires a large amount of computing resources.

[0093] Knowledge graph methods assist in the discovery and verification of biomarkers by integrating existing biological knowledge bases. However, traditional knowledge graph methods often rely on existing annotation information, have limited predictive ability for newly discovered regulatory relationships and unknown mechanisms, and the lag in updating the knowledge base affects the timeliness of prediction results.

[0094] In summary, this invention provides a method for predicting the expression regulation mechanism of disease biomarkers. The method first acquires gene expression data, obtains standardized gene expression data through quality control and standardization, then identifies candidate biomarkers using differential expression analysis, constructs a transcriptional regulatory network based on chromatin immunoprecipitation sequencing data and transcription factor binding site data, and finally uses a four-quadrant neural network model for feature analysis and prediction, and determines the biological significance of the biomarkers through functional enrichment analysis.

[0095] The innovation and technical effects of this invention are mainly reflected in the following aspects:

[0096] The innovative design of the four-quadrant neural network model fully considers the multi-dimensional feature integration requirements in the biomarker discovery process. By dividing the feature space into four quadrants, it achieves hierarchical processing and optimized integration of different types of features. The contribution value calculation layer first performs weighted calculations on the input features, taking into account the importance of the features; the statistical testing layer ensures the reliability of feature selection through rigorous statistical analysis; the quadrant partitioning layer classifies features based on two dimensions: contribution value and statistical significance, making feature selection and integration more biologically meaningful; the feature integration layer achieves effective combination of different types of features through differentiated weight allocation; and the final comprehensive output module uses a non-linear activation function to complete the final prediction.

[0097] The four-quadrant partitioning design principle originates from the "significance-effect size" analysis framework commonly used in biomarker screening, but this invention innovatively extends this framework. The first quadrant (high contribution value, high significance) contains the most valuable core features, which are not only statistically significant but also have strong biological effects. The second quadrant (high contribution value, low significance) collects features that may be limited by sample size but have potential biological significance. The third quadrant (low contribution value, low significance) contains potential background noise features. The fourth quadrant (low contribution value, high significance) contains statistically significant features but may have weaker biological significance. This partitioning method allows the model to more comprehensively evaluate and utilize various feature information.

[0098] Compared to traditional methods, this invention offers the following advantages: First, traditional methods often rely solely on single statistical tests or machine learning models for feature selection, easily overlooking complex relationships between features and their biological context. This invention, however, achieves multi-dimensional feature evaluation and selection through four-quadrant partitioning. Second, traditional methods are often weak in handling temporal and dynamic features, while this invention introduces time integral and derivative terms, better capturing the dynamic changes in gene expression. Third, traditional methods typically use fixed feature weights, while this invention achieves dynamic weight adjustment through adaptive quadrant partitioning and optimization equations, improving the model's adaptability.

[0099] The methodological innovations of this invention are also reflected in the following aspects: by introducing multiple feature matrices such as tissue-specific coefficients and temporal expression parameters, multi-level characterization of gene expression regulation is achieved; by constructing complex weight allocation equations, the smoothness, sparsity, and time-varying nature of feature selection are considered; and by integrating the design of the output equation, precise control and optimization of prediction results are achieved. Especially when processing large-scale gene expression data, this invention effectively reduces the dimensionality of the data and improves computational efficiency by introducing a dimension transformation coefficient matrix.

[0100] The design of the four-quadrant neural network also considers the inherent characteristics of biological systems: the contribution value calculation layer reflects the hierarchical and complex nature of gene expression regulation; the statistical testing layer ensures the reliability and repeatability of feature selection; the quadrant partitioning layer enables hierarchical management and optimal combination of features; and the feature integration layer ensures the effective fusion of information from different sources. This hierarchical design not only improves the interpretability of the model but also enhances the reliability of the prediction results.

[0101] In practical applications, the method of this invention exhibits significant advantages: the model's prediction accuracy is generally higher than 85%, and the Matthew correlation coefficient is greater than 0.7, indicating strong predictive ability; validation on independent validation sets demonstrates good generalization performance; and it shows stable performance across different disease types and datasets. Furthermore, this method not only predicts biomarker expression levels but also reveals the biological processes involved by biomarkers through functional enrichment analysis, providing a new perspective for disease mechanism research.

[0102] In summary, this invention, through its innovative four-quadrant neural network design and comprehensive feature integration strategy, achieves accurate prediction of the regulatory mechanisms of disease biomarker expression, providing important theoretical basis and technical support for disease diagnosis, prognostic assessment, and treatment strategy formulation. The innovativeness and practicality of this method make it a promising candidate for application in the field of biomarker research.

[0103] The following is a specific embodiment 1 of the present invention. The specific implementation of each step in this embodiment 1 is described in detail below: Step S10 is implemented as follows: First, obtain gene expression data of the target disease from relevant gene expression databases or clinical samples. The obtained gene expression data includes transcriptome sequencing data of patient samples and healthy control samples, and the number of samples in each group is not less than 50. Such a sample size ensures the statistical significance of subsequent analyses.

[0104] Secondly, the acquired gene expression data underwent quality control and standardization. The standardization method employed was z-score standardization, which centers and normalizes the data by subtracting the sample mean μ and dividing by the sample standard deviation σ.

[0105] The specific formula is as follows:

[0106]

[0107] Among them, X norm X represents the standardized gene expression value, X represents the original gene expression value, and δ represents the offset correction parameter in the range of 0.01 to 0.1, used to avoid calculation problems caused by zero values.

[0108] In addition, sample data with sequencing depth below 10 million reads or quality scores below 30 need to be discarded. Data with excessively low sequencing depth or poor quality may introduce more noise, affecting subsequent analysis results.

[0109] Through the above quality control and standardization processes, high-quality standardized gene expression data can be obtained, laying the foundation for subsequent differential expression analysis. The purpose of this step is to ensure that the gene expression data used in the analysis is reliable and comparable.

[0110] The specific implementation of step S20 is as follows: Differential expression analysis is used to identify differentially expressed genes related to the target disease. First, the fold change (FC) of each gene expression between disease samples and healthy control samples is calculated, using the formula:

[0111]

[0112] Among them, E d and E c α represents the gene expression values ​​in the disease sample and the control sample, respectively, and α is the expression correction coefficient in the range of 0 to 0.2.

[0113] Then, a statistical significance test was performed on the change in expression fold. The p-value was calculated using methods such as t-test or ANOVA, with the formula:

[0114]

[0115] Where t is the test statistic, β is the probability correction parameter in the range of 0.001 to 0.01, and μ and σ are the sample mean and standard deviation, respectively.

[0116] Finally, genes with an expression fold change greater than 2 and a statistical significance p-value less than 0.01 were selected as candidate biomarkers. This screening criterion can effectively identify differentially expressed genes that are highly associated with the disease.

[0117] The purpose of this step is to initially screen differentially expressed genes that may serve as disease biomarkers from a large amount of gene expression data, laying the foundation for subsequent construction of transcriptional regulatory networks and establishment of predictive models.

[0118] The specific implementation of step S30 is as follows: Based on the screened candidate biomarkers, a transcriptional regulatory network is constructed. First, chromatin immunoprecipitation sequencing (ChIP-Seq) data is collected to identify transcription factors that have a direct regulatory relationship with the candidate biomarkers. ChIP-Seq technology can detect the sites where transcription factors actually bind to chromatin.

[0119] Then, transcription factor binding site data were integrated to identify which transcription factors directly regulate the expression of candidate biomarkers. These direct regulatory relationships constitute the edges of the transcriptional regulatory network. Finally, the identified transcription factors and their regulatory relationships with candidate biomarkers were integrated into transcriptional regulatory network data.

[0120] The purpose of this step is to gain a deeper understanding of the expression regulation mechanisms of candidate biomarkers. Transcriptional regulatory networks can visually demonstrate which transcription factors are involved in the expression regulation of candidate biomarkers, providing a foundation for the subsequent development of expression regulation prediction models.

[0121] The specific implementation of step S40 is as follows: A multi-layer neural network model is established using a four-quadrant neural network model to perform feature analysis and expression level prediction on candidate markers. The specific structure of the four-quadrant neural network model is as follows:

[0122] 1. Contribution value calculation layer: Receives standardized gene expression data X norm The contribution value of each input feature is calculated using the weight matrix W, along with transcription factor data, and then transformed using a non-linear activation function to output a contribution value vector C. The number of nodes in the contribution value calculation layer is 1.5 times the dimension of the input features. The contribution equation is expressed as follows:

[0123]

[0124] Where T is the tissue-specific coefficient matrix, S is the time-series expression parameter matrix, D is the dimension transformation coefficient matrix, and γ is the contribution value correction vector.

[0125] 2. Statistical Testing Layer: This layer receives the contribution value vector and performs statistical significance analysis using standardized gene expression data, calculating the p-value (P0). s Output the statistical significance test vector. The number of nodes in the statistical test layer is the same as that in the contribution value calculation layer. The significance equation is expressed as follows:

[0126]

[0127] Among them, F obs and F crit Here, represents the observed value and the critical F-value, respectively; V is the homogeneity of variance coefficient; N is the normality parameter of the data; L is the significance level parameter; A is the degree of freedom adjustment coefficient; I is the confidence interval range parameter; and η is the significance correction parameter.

[0128] 3. Quadrant Division Layer: This layer receives both contribution value vectors and statistical significance test vectors, and divides all features into quadrants based on preset contribution value thresholds and significance thresholds. This layer contains 4 nodes, each corresponding to one of the 4 quadrants.

[0129] 4. Feature Integration Layer: Receives feature quadrant assignment labels and assigns corresponding weight coefficients W to features in different quadrants. q The features are then weighted and integrated to output an integrated feature vector. The number of nodes in the feature integration layer is 0.5 times the dimension of the input features. The weight allocation equation is expressed as follows:

[0130]

[0131] Where R is the smoothing factor coefficient, λ is the weight decay parameter, k is the number of iterations, B is the quadrant boundary adjustment parameter, U is the feature importance coefficient, F is the feature selection coefficient, and θ is the weight correction parameter.

[0132] 5. Integrated Output Module: This module receives the integrated feature vector, performs feature mapping through a fully connected layer, and calculates the predicted expression level O of the candidate biomarker using a non-linear activation function. This module contains one node. The integrated output equation is expressed as follows:

[0133]

[0134] Where Q is the feature quadrant assignment label, M is the nonlinear mapping parameter, and T... a Let L be the temperature coefficient of the activation function, G be the gradient scaling factor, and L be the temperature coefficient of the activation function. r ω is the learning rate adjustment parameter, b is the bias term correction coefficient, and ω is the output correction parameter.

[0135] The four-quadrant neural network model employs an adaptive quadrant partitioning optimization equation set to perform feature analysis and prediction of candidate biomarkers. This model can fully utilize the biological characteristics of gene expression, such as tissue specificity and temporal variation trends, to predict the expression levels of candidate biomarkers under different pathological states.

[0136] The purpose of this step is to establish a neural network model that integrates feature analysis and expression level prediction, providing a reliable foundation for subsequent functional analysis and expression regulation prediction.

[0137] The specific implementation of step S50 is as follows: The predictive performance of the four-quadrant neural network model is evaluated using cross-validation. First, the original gene expression data is randomly divided into a training set and a validation set. The training set is used for model training, and the validation set is used for performance evaluation. Then, a 5-fold cross-validation method is used to repeatedly perform the training-validation process, calculating the model's prediction accuracy and Matthew correlation coefficient on the validation set.

[0138] Models with a prediction accuracy higher than 85% and a Matthew correlation coefficient greater than 0.7 were selected as the final models. Such performance metrics ensure that the models have high prediction accuracy and generalization ability.

[0139] In addition, independent validation set data is needed to further evaluate the generalization ability of the selected model. The sample size of the independent validation set should be no less than 30% of the original gene expression data. By validating the model's predictive performance on independent samples, its performance in real-world applications can be more realistically evaluated.

[0140] The purpose of this step is to use rigorous cross-validation and independent validation methods to screen out four-quadrant neural network models with excellent predictive performance, providing a reliable foundation for subsequent functional analysis and expression regulation prediction.

[0141] The specific implementation of step S60 is as follows: Functional enrichment analysis is performed on the screened candidate biomarkers to determine the biological pathways and molecular functions they participate in. First, based on the gene ID of the candidate biomarkers, their biological pathways and molecular functions are queried in relevant bioinformatics databases (such as KEGG, GO, etc.).

[0142] Then, statistical methods such as hypergeometric tests were used to calculate the significance p-value for each pathway or function, screening out biological processes highly correlated with candidate biomarkers. Finally, the significantly enriched pathway and functional information was integrated into intuitive visualizations, such as bubble charts or bar charts, to help understand the mechanisms of action of candidate biomarkers in disease development and progression.

[0143] The purpose of this step is to conduct an in-depth analysis of candidate biomarkers from the perspective of gene function, understand their specific roles in disease-related biological processes, and provide more biological evidence for subsequent expression regulation prediction.

[0144] The specific implementation of step S70 is as follows: Constructing a predictive model for the expression regulation of candidate biomarkers. This model integrates transcription factor binding site data, chromatin immunoprecipitation sequencing data, and epigenetic modification data to predict the expression levels of candidate biomarkers under different pathological states.

[0145] First, information on transcription factors that directly interact with candidate biomarkers is extracted from the transcriptional regulatory network. Then, chromatin site data on the binding sites of these transcription factors are collected to identify transcription factor binding sites that affect the expression of candidate biomarkers.

[0146] Next, epigenetic modification data, such as histone methylation / acetylation, are integrated to analyze how these epigenetic modifications regulate the expression of candidate biomarkers. Finally, a predictive model that comprehensively considers transcription factor regulation and epigenetic modifications is constructed, capable of predicting the expression levels of candidate biomarkers under different physiological and pathological conditions.

[0147] The purpose of this step is to establish a more complete expression regulation prediction model that considers not only the direct regulation of transcription factors but also the influence of epigenetic regulation, making the prediction results more reliable and targeted, and providing a theoretical basis for subsequent clinical applications.

[0148] The following is a specific application scenario of the present invention, Example 2. This Example 2 takes chronic lymphocytic leukemia (CLL) as the research object and aims to use the disease biomarker expression regulation mechanism prediction method proposed in this invention to screen out biomarkers with diagnostic and prognostic value from the gene expression data of CLL patients.

[0149] 1. Gene expression data acquisition and preprocessing

[0150] First, transcriptome sequencing data from CLL patients and healthy controls were obtained from publicly available gene expression databases. This experiment collected samples from 80 CLL patients and 60 healthy controls. The following preprocessing steps were performed on these raw sequencing data: 1) Quality control: Samples with a sequencing depth of less than 10 million records or a quality score of less than 30 were removed to ensure data reliability. 2) Standardization: z-score standardization was used, i.e., subtracting the sample mean and dividing by the standard deviation, while introducing a shift correction parameter of 0.05 to obtain the standardized gene expression data matrix X. norm .

[0151] like Figure 2 The figure shows the gene expression distribution of CLL samples and healthy control samples. A probability density function is used to represent the differences in expression patterns between the two groups.

[0152] After the above preprocessing, standardized gene expression data from 70 CLL patient samples and 55 healthy control samples were finally obtained, laying the foundation for subsequent differential expression analysis.

[0153] 2. Screening for differentially expressed genes

[0154] Differential expression analysis was used to screen differentially expressed genes associated with CLL from standardized gene expression data. The specific steps are as follows: 1) Calculate the fold change (FC) of each gene expression between CLL samples and healthy control samples, using the formula:

[0155]

[0156] Among them, E CLL and E control 1) Gene expression values ​​in CLL samples and control samples, respectively, with 0.1 as the expression level correction coefficient. 2) A t-test was used to analyze the statistical significance of the fold change in expression, and the p-value was calculated using the following formula:

[0157]

[0158] Where t is the test statistic, μ and σ are the sample mean and standard deviation, respectively, and 0.005 is the probability correction parameter. 3) Genes with an expression fold change greater than 2 and a p-value less than 0.01 were selected as candidate biomarkers.

[0159] Following the above screening process, a total of 143 differentially expressed genes were selected as candidate biomarkers for CLL. These genes will serve as input features for subsequent four-quadrant neural network models.

[0160] like Figure 3 The figure shows the statistical significance of changes in gene expression.

[0161] 3. Construction of transcriptional regulatory networks

[0162] To gain a deeper understanding of the expression regulation mechanisms of candidate biomarkers, chromatin immunoprecipitation sequencing (ChIP-Seq) data related to CLL were collected, and transcription factors that have a direct regulatory relationship with candidate biomarkers were identified.

[0163] By integrating ChIP-Seq data, 26 transcription factors were identified as being involved in the expression regulation of candidate biomarkers. Table 1 lists these transcription factors and their regulatory relationships with candidate biomarkers.

[0164] Table 1. Transcriptional regulatory network of candidate biomarkers

[0165] transcription factors Regulatory genes STAT3 IGHM, CD79A, TCL1A MYC IGHM, TCL1A, LPL, BIRC3 NF-κB CD79A, BIRC3, BCL2 TP53 ATM, BIRC3, BCL2 RUNX1 CD79A, LPL, MCL1 EGR1 LPL, BIRC3, BCL2 NFATC1 CD79A, LPL, MCL1 … …

[0166] Based on the above transcriptional regulatory relationships, a transcriptional regulatory network for CLL candidate biomarkers was constructed, providing a foundation for subsequent expression regulation prediction.

[0167] like Figure 4 The diagram illustrates the strength of the regulatory relationship between transcription factors and target genes. The horizontal axis represents target genes, and the vertical axis represents transcription factors; the color intensity indicates the regulatory strength. The YlOrRd color system was used to display changes in regulatory strength.

[0168] 4. Establishment and Training of a Four-Quadrant Neural Network Model

[0169] Next, the four-quadrant neural network model proposed in this invention is used to perform feature analysis and expression level prediction on CLL candidate markers. The specific structure of the model is as follows:

[0170] 1) Contribution value calculation layer: Receives standardized gene expression data X norm Based on transcription factor data, the contribution value C of each input feature is calculated using a weight matrix W. This layer contains 215 nodes. The contribution equation is expressed as follows:

[0171]

[0172] Where T is the tissue-specific coefficient matrix, S is the time-series expression parameter matrix, D is the dimension transformation coefficient matrix, and γ is the contribution value correction vector.

[0173] 2) Statistical test layer: Receives the contribution value vector and performs statistical significance analysis based on standardized gene expression data, calculating the p-value (P0). s Output the statistical significance test vector. The number of nodes in this layer is the same as that in the contribution value calculation layer. The significance equation is expressed as follows:

[0174]

[0175] 3) Quadrant Division Layer: This layer receives both the contribution value vector and the statistical significance test vector, dividing all features into four quadrants. This layer contains four nodes, each corresponding to one of the four quadrants.

[0176] 4) Feature Integration Layer: Receives feature quadrant assignment labels and assigns corresponding weight coefficients W to features in different quadrants. q The layers are then weighted and integrated to output an integrated feature vector. This layer contains 72 nodes. The weight allocation equation is expressed as follows:

[0177]

[0178] 5) Integrated Output Module: This module receives the integrated feature vector, performs feature mapping through a fully connected layer, and calculates the predicted expression level O of the CLL candidate marker using a non-linear activation function. This module contains one node. The integrated output equation is expressed as follows:

[0179]

[0180] During model training, a 5-fold cross-validation method was used to evaluate the predictive performance of the four-quadrant neural network model on the CLL dataset. The results show that the model achieved a prediction accuracy of 88% and a Matthew correlation coefficient of 0.75 on the validation set.

[0181] To further validate the model's generalization ability, 25 independent CLL patient samples and 18 healthy control samples were collected as an independent validation set to test the model. On the independent validation set, the model's prediction accuracy was 86%, and the Matthew correlation coefficient was 0.72, indicating that the model has strong adaptability and reliability.

[0182] like Figure 5 The figure shows the performance changes of the model during training.

[0183] 5. Functional enrichment analysis of candidate biomarkers

[0184] Through the above steps, 143 candidate biomarkers related to CLL were identified. To better understand the mechanisms of action of these candidate biomarkers in the occurrence and development of CLL, functional enrichment analysis was performed on them.

[0185] First, the gene IDs of these candidate biomarkers were input into the KEGG and GO databases to query the biological pathways and molecular functions they participate in. Then, the significance p-value for each pathway or function was calculated using the hypergeometric test, and the results of significant enrichment were visualized.

[0186] Table 2 summarizes the 10 most significantly enriched biological pathways and molecular functions of CLL candidate biomarkers. It can be seen that these candidate biomarkers are mainly involved in biological processes closely related to tumorigenesis and development, such as cell cycle regulation, apoptosis, and cell signal transduction.

[0187] Table 2. Results of functional enrichment analysis of CLL candidate biomarkers

[0188] Biological pathways / molecular functions p-value Cell cycle 1.2 x 10 -8 <!-- 16 -->]]> Apoptosis 3.5 x 10 -7 ]]> PI3K-Akt signal path 7.2 x 10 -6 ]]> MAPK signaling path 1.8 x 10 -5 ]]> B-cell receptor signaling pathway 2.4 x 10 -5 ]]> apoptosis regulation 4.1 x 10 -5 ]]> Regulation of protein kinase activity 6.3 x 10 -5 ]]> Cell proliferation regulation 8.7 x 10 -5 ]]> Cytokine activity 1.1 x 10 -4 ]]> Cell signal transduction 1.5 x 10 -4 ]]>

[0189] The above analysis results indicate that the CLL candidate biomarkers screened in this invention are closely related to the occurrence and development of the disease, and are of great significance for a deeper understanding of the molecular pathogenesis and prognosis of CLL.

[0190] 6. Construction of Expression Regulation Prediction Model

[0191] Finally, a predictive model for the expression regulation of CLL candidate biomarkers was constructed that comprehensively considers transcription factor regulation and epigenetic modification.

[0192] First, information on 26 transcription factors that directly interact with candidate biomarkers was extracted from the aforementioned transcriptional regulatory network. Then, by integrating relevant ChIP-Seq data, the regulatory sites where these transcription factors actually bind to candidate biomarker genes were identified.

[0193] In addition, epigenetic modification data related to CLL, such as histone methylation / acetylation, were collected to analyze how these epigenetic modifications regulate the expression of candidate biomarkers.

[0194] Finally, a predictive model that comprehensively considers transcription factor regulation and epigenetic modification was constructed, capable of predicting the expression levels of candidate biomarkers under different physiological and pathological conditions. The key formula of this model is as follows:

[0195]

[0196] Where, x i For the i-th feature input (including transcription factor binding site information and epigenetic modification data), w i Here, w represents the corresponding weight parameters, T represents the observation time range, and b represents the bias term.

[0197] This model can not only predict the expression levels of CLL candidate biomarkers under different pathological states, but also analyze the key regulatory factors leading to changes in expression. This provides an important basis for further understanding the pathogenesis of CLL and discovering new diagnostic and therapeutic targets.

[0198] 7. Comparison of traditional methods vs. the method of this invention

[0199] To better demonstrate the innovation and advantages of the method of this invention, it is compared with traditional biomarker discovery methods, and the results are shown in Table 3.

[0200] Table 3. Comparison of traditional methods vs. the method of this invention

[0201]

[0202] As can be seen from the table above, compared with traditional methods, the four-quadrant neural network model proposed in this invention has significant improvements in feature integration, dynamic modeling, adaptive optimization, and result interpretation. Especially in terms of prediction accuracy, the method of this invention achieves a high prediction accuracy of 86% on the independent validation set, far exceeding that of traditional methods.

[0203] The key to these advancements lies in the fact that the method of this invention fully considers the multidimensional features and dynamic changes inherent in gene expression data, and effectively utilizes these complex factors through innovative algorithm design. Simultaneously, the adaptive optimization strategy improves the model's adaptability, enabling it to better cope with various biological complexities that may arise in practical applications.

[0204] It should be noted that the variables involved in this invention are explained in detail in Table 4 below.

[0205] Table 4. Variable Explanation Table

[0206]

[0207] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

[0208] Technical Principle: The innovative design of the four-quadrant neural network model in this invention originates from the "significance-effect size" analysis framework commonly used in biomarker screening. This framework posits that an ideal biomarker should simultaneously possess statistical significance and biological meaning, meaning it should not only exhibit significant differential expression between disease samples and healthy control samples but also play a crucial regulatory role in the disease's development and progression.

[0209] Based on this idea, this invention divides the feature space into four quadrants. The first quadrant contains features that have both high contribution values ​​and high significance; these features are the most valuable core biomarkers. The second quadrant collects features that may be limited by sample size but have potential biological significance. The third quadrant contains potential background noise features. The fourth quadrant collects statistically significant features but with weak biological significance.

[0210] This multi-dimensional feature segmentation and combination strategy enables a more comprehensive evaluation and utilization of various feature information. Compared to traditional methods that rely solely on single statistical tests or machine learning models, the method of this invention can better uncover the complex regulatory relationships and dynamic changes contained in gene expression data.

[0211] Specifically, the core of this invention lies in the design of a four-quadrant neural network model. The contribution value calculation layer uses a weight matrix to weight the input features, considering their importance; the statistical testing layer ensures the reliability of feature selection through rigorous statistical analysis. The quadrant partitioning layer classifies features based on two dimensions: contribution value and statistical significance, making feature selection and integration more biologically meaningful. The feature integration layer achieves effective combination of different types of features through differentiated weight allocation. Finally, the comprehensive output module uses a non-linear activation function to complete the final prediction.

[0212] This hierarchical design not only improves the model's interpretability but also enhances the reliability of the prediction results. Simultaneously, this invention introduces time integral and derivative terms, which better capture the dynamic changes in gene expression. Furthermore, the adaptive quadrant partitioning optimization equations enable dynamic adjustment of feature weights, further improving the model's adaptability.

Claims

1. A method for predicting the expression regulation mechanism of disease biomarkers, characterized in that, The method includes the following steps: acquiring gene expression data; performing quality control and standardization on the gene expression data to obtain standardized gene expression data; identifying differentially expressed genes as candidate biomarkers; constructing a transcriptional regulatory network for the candidate biomarkers and generating transcription factor data; establishing a multi-layer neural network model using a four-quadrant neural network model, wherein the four-quadrant neural network model includes a contribution value calculation layer, a statistical test layer, a quadrant partitioning layer, a feature integration layer, and a comprehensive output module; the four-quadrant neural network model uses an adaptive quadrant partitioning optimization equation set to perform feature analysis and prediction on the candidate biomarkers; performing functional enrichment analysis on the candidate biomarkers; and constructing an expression regulation prediction model for the candidate biomarkers.

2. The method for predicting the expression regulation mechanism of disease biomarkers according to claim 1, characterized in that, The gene expression data includes transcriptome sequencing data of patient samples and healthy control samples, with a sample size of no less than 50 samples in each sample. The quality control and standardization process includes removing sample data with a sequencing depth of less than 10 million reads and a quality score of less than 30. The differentially expressed genes are genes with an expression fold change of more than 2 and a statistical significance test probability value of less than 0.

01.

3. The method for predicting the expression regulation mechanism of disease biomarkers according to claim 1, characterized in that, The transcriptional regulatory network identifies transcription factors that have a direct regulatory relationship with the candidate biomarkers based on chromatin immunoprecipitation sequencing data and transcription factor binding site data.

4. The method for predicting the expression regulation mechanism of disease biomarkers according to claim 1, characterized in that, The number of nodes in the four-quadrant neural network model is as follows: the number of nodes in the contribution value calculation layer is 1.5 times the input feature dimension; the number of nodes in the statistical test layer is the same as the number of nodes in the contribution value calculation layer; the number of nodes in the quadrant division layer is 4; the number of nodes in the feature integration layer is 0.5 times the input feature dimension; and the number of nodes in the comprehensive output module is 1.

5. The method for predicting the expression regulation mechanism of disease biomarkers according to claim 1, characterized in that, The adaptive quadrant partitioning optimization equation set includes a contribution equation, a significance equation, a weight allocation equation, and an integrated output equation. The inputs to the contribution equation include the standardized gene expression data, the weight matrix, the tissue-specific coefficient matrix, the time-series expression parameter matrix, and the dimension transformation coefficient matrix.

6. The method for predicting the expression regulation mechanism of disease biomarkers according to claim 5, characterized in that, The inputs to the significance equation include the standardized gene expression data, the healthy control sample data, the significance test level parameter, the homogeneity of variance coefficient, the normal distribution parameter of the data, the degree of freedom adjustment coefficient, and the confidence interval range parameter. The inputs to the weight allocation equation include the contribution value vector, the statistical significance test probability value, the smoothing factor coefficient, the weight decay parameter, the iteration update coefficient, the quadrant boundary adjustment parameter, and the feature importance coefficient.

7. The method for predicting the expression regulation mechanism of disease biomarkers according to claim 5, characterized in that, The inputs to the integrated output equation include feature quadrant assignment labels, weight coefficients, nonlinear mapping parameters, activation function temperature coefficients, gradient scaling factors, learning rate adjustment parameters, and bias term correction coefficients.

8. The method for predicting the expression regulation mechanism of disease biomarkers according to claim 1, characterized in that, The four-quadrant neural network model was evaluated using cross-validation. Models with a prediction accuracy higher than 85% and a Matthew correlation coefficient greater than 0.7 were selected. The four-quadrant neural network model was evaluated using independent validation set data, and the sample size of the independent validation set data was not less than 30% of the gene expression data.

9. The method for predicting the expression regulation mechanism of disease biomarkers according to claim 1, characterized in that, The quadrant partitioning layer assigns features whose contribution value is greater than a contribution value threshold and whose statistical significance test probability value is less than a significance threshold to the first quadrant; the quadrant partitioning layer assigns features whose contribution value is greater than a contribution value threshold and whose statistical significance test probability value is greater than a significance threshold to the second quadrant; the quadrant partitioning layer assigns features whose contribution value is less than a contribution value threshold and whose statistical significance test probability value is greater than a significance threshold to the third quadrant; and the quadrant partitioning layer assigns features whose contribution value is less than a contribution value threshold and whose statistical significance test probability value is less than a significance threshold to the fourth quadrant.

10. The method for predicting the expression regulation mechanism of disease biomarkers according to claim 1, characterized in that, The contribution value calculation layer receives the standardized gene expression data and the transcription factor data, calculates the contribution value of each input feature through the weight matrix, transforms the contribution value using a non-linear activation function, and outputs the contribution value vector. The statistical testing layer receives the contribution value vector and performs statistical significance analysis in conjunction with the standardized gene expression data, calculates the statistical significance test probability value, and outputs the statistical significance test vector.

Citation Information

Cited By

  • Neuronal apoptosis risk assessment system based on oxidative stress signal axis

    CN121366734A