Gene expression data difference analysis method and system

The outliers were eliminated by the interquartile distance method and combined with Wilcoxon test and Benjamini-Hochberg correction, the non-normal distribution and outliers of the gene expression data were solved, which improved the accuracy and biological significance of the analysis and reduced the false positive rate.

CN120279990APending Publication Date: 2025-07-08GUANGDONG NO 2 PROVINCIAL PEOPLES HOSPITAL
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510340731.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art processes gene expression data, especially non-normal distributions and data containing outliers, the accuracy and reliability of the analysis results are insufficient, and traditional methods are difficult to provide intuitive, accurate and biologically significant analysis results.

Method used

The outliers were eliminated by interquartile distance method, and multiple hypothesis correction was performed in combination with Wilcoxon test and Benjamini-Hochberg method, and log2FoldChange and importance scores were calculated to provide intuitive gene expression differences analysis.

Benefits of technology

Effectively reduce the interference of outliers on the analysis results, improve the accuracy and biological significance of differential analysis, reduce the false positive rate, and provide robust statistical support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279990A_ABST
    Figure CN120279990A_ABST
Patent Text Reader

Abstract

The invention discloses a gene expression data difference analysis method and system. The method comprises the following steps: acquiring biological expression data and corresponding sample grouping information; preprocessing the gene expression data and dividing the gene expression data into an experimental group and a control group; carrying out independent intra-group abnormal value identification and elimination on the preprocessed expression data; performing independent difference analysis on the expression data before and after the abnormal values are removed to obtain two groups of verification results; and performing multiple correction on the two groups of verification results to obtain a final verification result and output statistical information. The noise problem in gene expression data can be effectively solved, and the accuracy and biological significance of difference analysis are improved. The method has remarkable advantages especially in analysis of gene expression data of small samples and non-normal distribution, and is extremely suitable for biological data capable of being longitudinally compared. The final analysis result not only has high statistical reliability, but also can provide solid data support for subsequent biological research and experimental design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of biostatistics, and particularly relates to a method and system for differential analysis of gene expression data. Background Art

[0002] With the increasing growth of high-throughput omics data, the challenges in bioinformatics and computational biology for extracting effective information and drawing reliable conclusions have become increasingly acute. The non-normal distribution characteristics of gene expression data (usually manifested as discreteness and skewness) have been elaborated in professional literature. Especially when faced with gene expression data (such as FPKM, RPKM, TPM), such data often has large discreteness and skewness, and traditional normal distribution assumptions may not be applicable. When processing these data, traditional statistical methods often assume that the data conforms to a normal distribution or is insensitive to outliers, which may greatly reduce the accuracy and reliability of the analysis results when faced with non-normal distribution or data containing outliers.

[0003] Widely used differential analysis R tools such as "limma" and "DESeq2" mainly perform differential analysis based on raw count data (such as Count values in RNA-seq). Although they can provide significant differences in gene expression, they often ignore the impact of gene expression levels on the calculation of fold changes. Count values cannot be compared longitudinally, that is, it is impossible to judge the high or low gene expression longitudinally according to the data size, and it is impossible to provide a suitable threshold for screening. In particular, researchers often need to rely on standardized expression values such as FPKM, TPM, or RPKM to further judge whether a gene has biological significance, but this process depends on certain sequencing background knowledge, increasing the complexity of the operation and the requirements for data understanding. Moreover, once the selection is inappropriate, it may lead to inaccurate experimental results, increasing the time and resource consumption of the experiment.

[0004] In addition, traditional methods usually fail to intuitively display the longitudinal comparison of gene expression levels in the preliminary analysis, and for genes with low expression or extreme expression, the evaluation of fold changes may be misled, resulting in wrong decisions in subsequent biological analysis. Therefore, how to combine expression levels and fold changes to provide intuitive, accurate, and biologically meaningful analysis results while ensuring statistical power has become a major problem in current bioinformatics research. Summary of the Invention

[0005] To solve the above technical problems, the present invention proposes a method and system for differential analysis of gene expression data to solve the problems existing in the above prior art.

[0006] To achieve the above object, the present invention provides a method for differential analysis of gene expression data, including:

[0007] Obtain gene expression data and corresponding sample grouping information; preprocess the gene expression data, where the sample grouping includes an experimental group and a control group;

[0008] Identify and remove outliers from the preprocessed gene expression data according to the interquartile range principle; perform checks on the gene expression data before and after outlier removal respectively to obtain two sets of check results;

[0009] Perform multiple corrections on the two sets of check results to obtain the final check result and output statistical information.

[0010] Optionally, the process of preprocessing the gene expression data includes: converting the format of the gene expression data, transposing and cleaning the gene expression data in matrix format, with each column representing a sample and each row representing a gene.

[0011] Optionally, the process of removing outliers from the preprocessed gene expression data includes:

[0012] Calculate the quartiles and interquartile range of the corresponding expression values for each gene, independently for the experimental group and the control group; obtain the outlier range based on the quartiles and interquartile range, and remove outliers from the preprocessed gene expression data based on the outlier range.

[0013] Optionally, use the Wilcoxon test to analyze the between-group differences of the gene expression data before and after outlier removal to obtain the test results, where the test results include the P-value, log-fold change value, and importance score.

[0014] Optionally, use the BH method to correct the P-values in the check results for multiple hypothesis testing.

[0015] The present invention also provides a gene expression data differential analysis system, including:

[0016] A data input module for receiving gene expression data and corresponding sample grouping information and preprocessing the gene expression data;

[0017] An outlier detection and removal module for removing outliers from the preprocessed gene expression data;

[0018] A statistical analysis module for using the Wilcoxon test to analyze the between-group differences of the original data and the data after removing outliers, and correcting the P-values in the check results for multiple hypothesis testing;

[0019] A result integration and output module for outputting the Wilcoxon test results, the corrected P-values, and the importance scores, respectively before and after removal.

[0020] Optionally, the data input module converts the format of the gene expression data, transposes and cleans the gene expression data in matrix format, with each column representing a sample and each row representing a gene.

[0021] Optionally, the outlier detection and removal module calculates the quartiles and interquartile range of the corresponding expression values for each gene; obtains the outlier range based on the quartiles and interquartile range, and removes outliers from the preprocessed gene expression data based on the outlier range.

[0022] Optionally, the statistical analysis module uses the Wilcoxon test to perform between-group difference analysis on the gene expression data before and after outlier removal to obtain the test results; uses the BH method to correct the P-values in the test results for multiple hypothesis testing; wherein, the test results include P-values, log fold change values, and significance scores.

[0023] Compared with the prior art, the present invention has the following advantages and technical effects:

[0024] (1) Reducing the influence of outliers on the analysis results

[0025] By using the interquartile range (IQR) method to remove outliers, it is possible to effectively reduce the interference of outliers in the data on the statistical analysis results. Since outliers often significantly affect the calculation of the mean and standard deviation, and thus the accuracy of the P-value, after removing outliers, the analysis results are more stable and can more truly reflect the between-group differences.

[0026] (2) Improving the accuracy of difference analysis

[0027] In the case where the data does not satisfy the normal distribution, the Wilcoxon test, as a non-parametric test method, can more accurately evaluate the differences between the experimental group and the control group, avoiding the bias that may occur in the traditional t-test for non-normal data. Through this method, it is possible to more precisely determine which genes actually have potential biological significance in the case of no significant differences.

[0028] (3) Eliminating the false positive problem caused by multiple testing

[0029] In multiple hypothesis testing, the traditional P-values may show false positives, that is, wrongly judging that some genes have differences. By applying the Benjamini-Hochberg (BH) method to correct the P-values and controlling the false discovery rate (FDR), it is possible to effectively reduce false alarms in multiple testing and ensure the robustness and reliability of the analysis results in multiple hypothesis tests.

[0030] (4) Enhancing the biological significance of the analysis results

[0031] By calculating the log2FoldChange and the importance score, a more comprehensive understanding of the gene expression changes in the experimental group and the control group can be obtained. The log2FoldChange provides the specific fold change of the expression level, while the importance score comprehensively considers the mean and the change amplitude of gene expression, helping researchers identify genes that are biologically meaningful.

[0032] (5) Optimize result interpretation and biological discovery

[0033] In the final output results, the gene expression difference information before and after removing outliers is provided, and a comprehensive evaluation of the P-value, log2FoldChange, and importance score is obtained based on statistical analysis. These results not only help researchers determine whether the gene expression difference is statistically significant but also provide a more reliable basis for further biological research.

[0034] By comprehensively using methods such as non-parametric test (Wilcoxon test), outlier removal, log2FoldChange, and multiple hypothesis testing correction, the present invention can effectively solve the noise problem in gene expression data and improve the accuracy and biological significance of differential analysis. Especially in the analysis of gene expression data with small samples and non-normal distributions, it has significant advantages. The final analysis results not only have high statistical reliability but also can provide solid data support for subsequent biological research and experimental design. Brief Description of the Drawings

[0035] The drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0036] Figure 1 It is a schematic diagram of the method of the embodiment of the present invention. Detailed Description of the Embodiments

[0037] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0038] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0039] Example 1

[0040] As Figure 1As shown in the figure, in this embodiment, a method for differential analysis of gene expression data is provided, including:

[0041] (1) Data input and preprocessing

[0042] The input data includes two parts: gene expression data (expression_data) and sample grouping information (group_data). expression_data is the expression matrix of genes, where rows represent genes and columns represent samples; group_data describes the group information of each sample, usually divided into "experimental group" and "control group". First, convert expression_data into matrix format and transpose it to ensure that the columns in the data are samples and the rows represent different genes.

[0043] Transpose the gene expression data to ensure that each column represents a sample and each row represents a gene. This facilitates subsequent statistical analysis and variable processing.

[0044] (2) Initialization of result data frame

[0045] Define two result data frames: result_before is used to store the analysis results before removing outliers, including information such as the average expression levels, P-values, log2FoldChange, etc. of the experimental group and the control group for each gene. result_after is used to store the analysis results after removing outliers, including P-values, log2FoldChange, and importance scores, etc.

[0046] (3) Outlier detection and removal

[0047] For each gene, first calculate the quartiles (Q1 and Q3) and interquartile range (IQR) of its expression values. Based on the outlier range defined by the IQR, remove the sample data with abnormal expression levels. This can avoid the interference of outliers on statistical analysis and ensure the stability of the analysis results.

[0048] Use the IQR method to detect outliers and remove outliers based on statistical principles to ensure that the analysis results are not affected by individual extreme values. This method is applicable to small sample data sets and can effectively reduce the bias caused by outliers.

[0049] (4) Wilcoxon test

[0050] Perform inter-group difference analysis using the Wilcoxon test on the data without removing outliers. This test method is a non-parametric test and is applicable to gene expression data with non-normal distribution. The test results include P-values, log2-fold change values (log2FoldChange), and importance scores (Important_Score).

[0051] That is:

[0052]

[0053] (5) Statistical analysis after removing outliers

[0054] For the data after removing outliers, repeat the Wilcoxon test and calculate the P-value, log2FoldChange, and importance score. This step can more accurately reflect the gene expression differences after removing outliers.

[0055] The Wilcoxon test is used instead of the traditional t-test, which is suitable for non-normally distributed data and enhances the adaptability and robustness of this method in gene expression difference analysis.

[0056] By calculating log2FoldChange, measure the gene expression changes between the experimental group and the control group, and reflect the contribution of each gene in the analysis through the importance score, ensuring that the analysis results are not only statistically significant but also biologically meaningful.

[0057] (6) Multiple hypothesis testing correction

[0058] Perform multiple hypothesis testing correction on all P-values using the Benjamini-Hochberg (BH) method to reduce the false positive rate. The corrected P-value (i.e., padj) will be added to the result data frame.

[0059] When performing multiple hypothesis tests, use the Benjamini-Hochberg method to correct the P-values, reduce false positive results, and improve the reliability of the analysis.

[0060] (7) Result output

[0061] The final result returns two data frames: result_before: containing all statistical information before removing outliers. result_after: containing all statistical information after removing outliers.

[0062] The new differential analysis strategy based on the rank sum test (Mann-Whitney U test) proposed in this embodiment fully considers the roles of the median and quartiles. These statistics are more suitable for dealing with non-normal data with outliers than the mean and standard deviation. By removing outliers and using robust statistics to estimate the central tendency of the data, combining gene expression levels and fold changes, it avoids blindly pursuing significant differences while ignoring the importance of the expression levels themselves. The method provided in this embodiment not only improves the accuracy of the analysis but also provides more intuitive and interpretable results for users, thus helping users better identify biologically meaningful change characteristics.

[0063] Example 2

[0064] In this embodiment, a gene expression data differential analysis system is provided, including:

[0065] A data input module, which is used to receive gene expression data and corresponding sample grouping information, and preprocess the gene expression data;

[0066] An outlier detection and elimination module, which is used to eliminate outliers from the preprocessed gene expression data;

[0067] A statistical analysis module, which is used to perform inter-group difference analysis on the original data and the data after removing outliers by using the Wilcoxon test, and perform multiple hypothesis test correction on the P value in the test result;

[0068] A result integration and output module, which is used to output the Wilcoxon test result, the corrected P value and the importance score, respectively before and after elimination.

[0069] Furthermore, the data input module converts the format of the gene expression data, transposes and cleans the gene expression data in matrix format, with each column representing a sample and each row representing a gene.

[0070] Furthermore, the outlier detection and elimination module calculates the quartiles and interquartile ranges of the corresponding expression values for each gene; obtains the outlier range based on the quartiles and interquartile ranges, and eliminates outliers from the preprocessed gene expression data based on the outlier range.

[0071] Furthermore, the statistical analysis module performs inter-group difference analysis on the gene expression data before and after outlier elimination by using the Wilcoxon test to obtain the test result; performs multiple hypothesis test correction on the P value in the test result by using the BH method; among them, the test result includes the P value, the log2 fold change value and the importance score.

[0072] Ignoring the effects of outliers and gene expression levels on fold change can affect the accuracy of differential quantification analysis and, in most cases, obscure the biological significance. To avoid these potential errors in applications, the method and system provided in this embodiment offer robust statistical methods for longitudinal biological data analysis, including FPKM, RPKM, and TPM. The package addresses non-normal data distributions and outlier problems through a rank sum test based on the median and interquartile range, ensuring accurate estimation of central tendency even in the case of non-normal data. Outliers are defined as values exceeding 1.5 times the interquartile range (IQR), and the analysis re-evaluates significance after removing these values, thereby improving robustness. The package also employs a weighted algorithm to balance gene expression levels and fold change, avoiding biologically irrelevant results that may arise from focusing solely on fold change. For example, the degree of change of gene A relative to gene B requires consideration of both fold change and expression level. Therefore, a unified value such as TPM is needed for meaningful comparison of expression levels. In summary, the present invention provides a powerful and flexible solution for more accurate and biologically meaningful analysis, facilitating the progress of biological research.

[0073] The above is only the preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for differential analysis of gene expression data, characterized in that, It includes the following steps: Obtain gene expression data and corresponding sample grouping information; preprocess the gene expression data, wherein the sample grouping includes an experimental group and a control group; Identify and remove outliers from the preprocessed gene expression data according to the interquartile range principle; perform verification on the gene expression data before and after outlier removal respectively to obtain two sets of verification results; Perform multiple corrections on the two sets of verification results to obtain the final verification result and output statistical information.

2. The method for analyzing gene expression data differences according to claim 1, wherein The process of preprocessing the gene expression data includes: converting the format of the gene expression data, transposing and cleaning the gene expression data in matrix format, with each column representing a sample and each row representing a gene.

3. The method for analyzing gene expression data differences according to claim 1, wherein The process of removing outliers from the preprocessed gene expression data includes: Calculate the quartiles and interquartile range of the corresponding expression values for each gene, independently for the experimental group and the control group; obtain the outlier range based on the quartiles and interquartile range, and remove outliers from the preprocessed gene expression data based on the outlier range.

4. The method for analyzing gene expression data differences according to claim 3, wherein Use the Wilcoxon test to analyze the between-group differences of the gene expression data before and after outlier removal to obtain the test results, wherein the test results include P-value, log fold change value, and importance score.

5. The method for analyzing gene expression data differences according to claim 4, wherein Use the BH method to perform multiple hypothesis testing correction on the P-values in the verification results.

6. A gene expression data differential analysis system, characterized in that, It includes: A data input module, used to receive gene expression data and corresponding sample grouping information, and preprocess the gene expression data; An outlier detection and removal module, used to remove outliers from the preprocessed gene expression data; A statistical analysis module, used to use the Wilcoxon test to analyze the between-group differences of the original data and the data after removing outliers, and perform multiple hypothesis testing correction on the P-values in the verification results; A result integration and output module, used to output the Wilcoxon test results, the corrected P-values, and the importance scores, respectively before and after removal.

7. The gene expression data difference analysis system according to claim 6, wherein The data input module converts the format of the gene expression data, transposes and cleans the gene expression data in matrix format, with each column representing a sample and each row representing a gene.

8. The gene expression data difference analysis system according to claim 6, wherein The outlier detection and removal module calculates the quartiles and interquartile range of the corresponding expression values for each gene; obtains the outlier range based on the quartiles and interquartile range, and removes outliers from the preprocessed gene expression data based on the outlier range.

9. The gene expression data difference analysis system according to claim 6, wherein The statistical analysis module uses the Wilcoxon test to perform an inter-group difference analysis on the gene expression data before and after outlier removal, and obtains the test results; The BH method is used to correct the P-values in the test results for multiple hypothesis testing; among them, the test results include P-values, log-fold change values, and importance scores.

Citation Information

Patent Citations

  • Differential analysis method and system based on single cell samples of mixed experimental group and control group

    CN114864003A

  • Data analysis method in gene prediction process

    CN116469461A

  • Proteomics data analysis method and device, electronic equipment and storage medium

    CN116884478A

  • Gene differentiation expression analysis method and system

    CN117976046A

  • Typing auxiliary model of adult acute B lymphocytic leukemia based on ensemble learning and construction method thereof

    CN118486370A