Maize stress resistance gene electronic expression spectrum network analysis method and system
By decomposing gene expression levels into batch baselines and biological response components, and combining this with screening key genes based on stress resistance signal intensity, the problem of batch effect misjudgment in maize stress resistance gene expression profile analysis was solved, achieving accurate identification of key genes and preservation of robust characteristics.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING FENGJIE YIJIA AGRICULTURAL TECHNOLOGY CO LTD
- Filing Date
- 2025-07-30
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies struggle to effectively distinguish and model batch effects from real biological signals in maize stress resistance gene expression profiling, leading to false positive and false negative results and failing to accurately identify key gene expression features associated with maize stress resistance traits.
By decomposing gene expression levels into batch baseline components and biological response components, extracting stable response components, and combining them with stress resistance signal intensity for screening, batch effect interference can be eliminated, thus achieving accurate identification of key genes.
It effectively distinguishes and models batch effects from real biological signals in data, avoids false positive and false negative results, and accurately identifies key gene expression features that exhibit robust consistency in multi-source heterogeneous data.
Smart Images

Figure CN121034408B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioinformatics, and in particular to a method and system for analyzing the electronic expression profile of maize stress resistance genes. Background Technology
[0002] In the field of agricultural biotechnology, research institutions are dedicated to elucidating the gene expression regulatory network of maize under adverse conditions in order to breed superior varieties that can withstand environmental stresses such as drought, salinity, and low temperature. The research process involves placing maize in a stress environment, collecting tissue samples (roots, stems, leaves, etc.) at multiple key time points, and using high-throughput sequencing technology to obtain gene expression data, forming electronic expression profiles. However, the raw data is high-dimensional and contains background noise, necessitating computational methods to extract features related to stress resistance traits, such as significant changes in gene expression levels or co-expression patterns.
[0003] In practical research, single experimental data have limitations. Therefore, data integration and analysis strategies are often employed to combine data from different batches within the laboratory and multiple sets of data from public databases to enhance sample size and statistical power. However, this approach introduces the "batch effect," a systematic bias caused by non-biological factors, such as differences in experimental platforms, sequencing technologies, sample processing procedures, or maize genotypes. These biases can cause traditional feature extraction methods to fail, potentially misinterpreting batch effects as key biological characteristics and producing false positive results.
[0004] To address this issue, the conventional approach is to first perform batch effect correction before feature extraction. However, existing correction algorithms often smooth out or remove genuine but weak biological signals when eliminating batch effects, leading to false negative results. For example, the expression characteristics of some key stress-resistance genes with low amplitude but consistent response patterns under various stress conditions, or signaling genes with transient and rapid expression changes in the early stages of stress, are easily filtered out. Therefore, the urgent technical problem to be solved is to design a key feature extraction method that can effectively distinguish and model batch effects from genuine biological signals in the data within a single analytical framework, avoiding false positive and false negative results, and thus accurately identifying the expression characteristics of robust and consistent key genes closely related to stress resistance traits in maize.
[0005] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention
[0006] In view of the shortcomings of the prior art, this application provides a method and system for electronic expression profiling of maize stress resistance genes. It has the advantages of effectively distinguishing and modeling batch effects and real biological signals in the data, avoiding false positive results caused by misjudging batch effects as key biological characteristics, and retaining real but weak biological signals to avoid false positive and false negative results. This allows for the accurate identification of key gene expression features that are closely related to maize stress resistance traits and show robust consistency in multi-source heterogeneous data.
[0007] Firstly, a method for analyzing the electronic expression profile network of maize stress resistance genes, the method comprising:
[0008] S1: Obtain electronic expression profile data of multi-source heterogeneous maize stress resistance genes, and perform normalization processing on the electronic expression profile data;
[0009] S2: The expression levels of each gene in the normalized electronic expression profile data are decomposed into variant components to obtain the batch baseline components and biological response components;
[0010] S3: Extract stable response components from the biological response components that exhibit consistent responses across different experimental batches;
[0011] S4: Obtain the stress resistance signal intensity of each gene in the electronic expression profile data, and screen preliminary key genes based on the stress resistance signal intensity and the stable response components;
[0012] S5: Based on the batch baseline components, further screen the preliminary key genes to exclude genes affected by batch effects and obtain the key genes.
[0013] This application proposes a network analysis method for electronic expression profiles of maize stress resistance genes. By decomposing gene expression levels into batch baseline components and biological response components, and extracting stable response components from the biological response components, preliminary screening is performed based on the stress resistance signal intensity. Further screening is then conducted based on the batch baseline components. This method effectively distinguishes and models batch effects and real biological signals in the data. It avoids false positive results caused by misjudging batch effects as key biological characteristics, while retaining real but weak biological signals to avoid false negative results. This method has the advantage of accurately identifying key gene expression characteristics that are closely related to maize stress resistance traits and exhibit robust consistency in multi-source heterogeneous data.
[0014] Furthermore, step S1 includes:
[0015] S11: Obtain the biological stress condition description attached to each of the electronic expression spectrum data samples, and parse the description to extract the stress type, stress intensity and stress duration;
[0016] S12: The extracted stress type, stress intensity, and stress duration are converted into a preset stress condition standardization system to obtain a standardized stress category;
[0017] S13: Group the electronic expression spectrum data according to the stress category, and normalize the electronic expression spectrum data within each group.
[0018] This application proposes a method for analyzing the electronic expression profile of maize stress-resistance genes. By standardizing and grouping stress conditions, this method improves the accuracy and reliability of data processing.
[0019] Furthermore, step S13 includes:
[0020] S131: Group the electronic expression spectrum data samples according to the standardized stress categories;
[0021] S132: For each group after grouping, select a set of genes whose expression fluctuation is lower than a preset threshold in all samples within that group as a reference gene set;
[0022] S133: Calculate the normalization factor for each sample based on the reference gene set;
[0023] S134: Apply the normalization factor to all gene expression data of the corresponding sample.
[0024] This application proposes a method for electronic expression profiling network analysis of maize stress resistance genes, which uses a reference gene set for intragroup normalization, further improving the accuracy of data normalization processing.
[0025] Furthermore, step S2 includes:
[0026] S21: For the expression level of each gene in the normalized electronic expression profile data, within each experimental batch, determine the control group sample and the treatment group sample;
[0027] S22: Calculate the mean expression level of the gene in the control group samples as the batch baseline component of the gene;
[0028] S23: Calculate the difference between the expression level of the gene in the treatment group sample and the batch baseline component, and use it as the biological response component of the gene.
[0029] This application proposes a method for analyzing the electronic expression profile of maize stress resistance genes. By providing a specific method for decomposing variant components, it can effectively separate batch baseline components and biological response components.
[0030] Furthermore, step S3 includes:
[0031] S31: For each gene, obtain its biological response components in different experimental batches;
[0032] S32: For each biological response component, determine its expression change direction to obtain the set of expression change directions of the gene in each experimental batch;
[0033] S33: Calculate the proportion of consistent directions in the set of expression change directions;
[0034] S34: When the proportion of the consistent direction reaches a preset threshold, it is determined that the gene has cross-batch directional consistency, thereby extracting the stable response component with consistent response.
[0035] Furthermore, step S32 includes:
[0036] S321: For the value of each of the biological response components, set a positive threshold and a negative threshold;
[0037] S322: When the value of the biological response component is greater than the positive threshold, the expression change direction of the biological response component is determined to be upregulation;
[0038] S323: When the value of the biological response component is less than the negative threshold, it is determined that the expression change direction of the biological response component is down-regulation;
[0039] S324: When the value of the biological response component is between the negative threshold and the positive threshold, it is determined that the expression change direction of the biological response component is unchanged;
[0040] S325: The directions of upregulation, downregulation, or no change are taken as the expression change directions of the gene in the experimental batch, and the sets of expression change directions of the gene in each experimental batch are obtained.
[0041] Furthermore, step S33 includes:
[0042] S331: For each gene, obtain the set of expression change directions in each experimental batch;
[0043] S332: Assign a weight to each experimental batch in the set of expression change directions, the weight being proportional to the data quality within the batch;
[0044] S333: Based on the set of expression change directions and the weights, determine the expression change direction with the largest total weight as the consistent direction;
[0045] S334: The ratio of the sum of weights of the consistent direction to the sum of weights of all experimental batches is taken as the proportion of the consistent direction.
[0046] Furthermore, step S4 includes:
[0047] S41: Obtain the stress resistance signal intensity of each gene in the electronic expression profile data;
[0048] S42: Calculate the combined fraction of the proportion of the anti-reverse signal strength and the consistent direction of the stable response component;
[0049] S43: Based on the combined scores, screen preliminary key genes.
[0050] Furthermore, step S5 includes:
[0051] S51: For each of the aforementioned preliminary key genes, obtain the corresponding batch baseline components and biological response components;
[0052] S52: Calculate the relative contribution of the batch baseline component of the preliminary key gene to the biological response component;
[0053] S53: Based on the relative contribution, exclude genes in the preliminary key genes that are mainly contributed by the batch baseline components to obtain the key genes.
[0054] Secondly, a maize stress resistance gene electronic expression profiling network analysis system is provided for implementing the method described in any of the above claims, the system comprising:
[0055] Acquisition module: Acquires electronic expression profile data of multi-source heterogeneous maize stress resistance genes and performs normalization processing on the electronic expression profile data;
[0056] Decomposition module: Decomposes the expression level of each gene in the normalized electronic expression profile data into variant components to obtain batch baseline components and biological response components;
[0057] Extraction module: Extracts stable response components from the biological response components that exhibit consistent responses across different experimental batches;
[0058] First screening module: Obtain the stress resistance signal intensity of each gene in the electronic expression profile data, and screen preliminary key genes based on the stress resistance signal intensity and the stable response components;
[0059] The second screening module further screens the preliminary key genes based on the batch baseline components, excluding genes affected by batch effects, and obtains the key genes.
[0060] Beneficial Effects: The electronic expression profiling network analysis method and system for maize stress resistance genes proposed in this application decomposes gene expression levels into batch baseline components and biological response components, extracts stable response components from the biological response components, performs preliminary screening based on stress resistance signal intensity, and further screens based on batch baseline components. This effectively distinguishes and models batch effects and real biological signals in the data, thus avoiding false positive results caused by misjudging batch effects as key biological characteristics, while retaining real but weak biological signals to avoid false negative results. This method has the advantage of accurately identifying key gene expression characteristics that are closely related to maize stress resistance traits and exhibit robust consistency in multi-source heterogeneous data. Attached Figure Description
[0061] Figure 1 This is a flowchart of a method for analyzing the electronic expression profile of maize stress resistance genes proposed in this application.
[0062] Figure 2 This is a structural diagram of the electronic expression profiling network analysis system for maize stress resistance genes proposed in this application.
[0063] Figure 3 This is an architecture diagram of a maize stress resistance gene electronic expression profiling network analysis system proposed in this application.
[0064] Labeling explanation: 201, Acquisition module; 202, Decomposition module; 203, Extraction module; 204, First filtering module; 205, Second filtering module. Detailed Implementation
[0065] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. The components of the embodiments of this application described and marked in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0066] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0067] Please refer to Figure 1 A method for analyzing the electronic expression profile network of maize stress resistance genes, the method comprising:
[0068] S1: Obtain electronic expression profile data of multi-source heterogeneous maize stress resistance genes and standardize the electronic expression profile data;
[0069] S2: The expression levels of each gene in the normalized electronic expression profile data are decomposed into variant components to obtain the batch baseline components and biological response components;
[0070] S3: Extract stable response components from biological response components that exhibit consistent responses across different experimental batches;
[0071] S4: Obtain the stress resistance signal intensity of each gene in the electronic expression profile data, and screen preliminary key genes based on the stress resistance signal intensity and stable response components;
[0072] S5: Based on the batch baseline components, further screen the preliminary key genes to exclude genes affected by batch effects and obtain the key genes.
[0073] Among them, obtaining electronic expression profile data of maize stress resistance genes from multiple sources and different platforms refers to collecting gene expression data from different experiments, platforms, and treatment conditions. These data may have systematic differences, and the main purpose is to gather a wider range of data resources and increase the sample size and coverage of the analysis. Standardizing the electronic expression profile data refers to preprocessing the obtained raw expression data to reduce non-biological differences between different samples or batches.
[0074] Decomposing the variation components of gene expression means decomposing the total variation of gene expression into baseline components caused by batch effects and response components caused by biological treatments.
[0075] Extracting stable response components that exhibit consistent responses across different experimental batches refers to identifying those components from the decomposed biological response components that show similar patterns of change across multiple experimental batches.
[0076] Obtaining the stress resistance signal intensity of each gene refers to assessing the degree of correlation between the gene and the stress resistance trait in maize, mainly to introduce stress resistance-related information of the gene itself.
[0077] Screening preliminary key genes based on stress resistance signal strength and stable response components refers to combining the stress resistance correlation information of genes with their stable biological responses in different batches to preliminarily identify potential key genes.
[0078] Further screening of key genes based on batch baseline components refers to using the decomposed batch baseline components to verify and optimize the preliminary screening results.
[0079] As a preferred embodiment, the solution of this application is specifically implemented as follows:
[0080] First, we acquired maize electronic expression profile data from different laboratories, sequencing platforms, and stress treatments. Each sample was then grouped according to its stress category, and within each group, a method based on a reference gene set was used for normalization.
[0081] Then, for each standardized gene, within each experimental batch, control and treatment samples were identified. The mean gene expression level of the control samples was calculated as the batch baseline component of that gene in that batch, and the difference between the expression level of the treatment samples and the batch baseline component was calculated as the biological response component.
[0082] Next, for each gene, its biological response components in different experimental batches are obtained, the expression change direction of each biological response component (upregulation, downregulation or no change) is determined, and a weight is assigned to each batch. Based on the weighted set of expression change directions, the proportion of consistent directions is calculated. When the proportion reaches a preset threshold, the gene is determined to have cross-batch directional consistency, and its stable response components are extracted.
[0083] Subsequently, the stress resistance signal intensity of each gene was obtained, and the combination score of the proportion of stress resistance signal intensity in the same direction as the stable response component was calculated. Based on the combination score, preliminary key genes were screened.
[0084] Finally, for each preliminary key gene, the relative contribution of its batch baseline component and biological response component is calculated, and genes mainly contributed by the batch baseline component are excluded to obtain the final key genes.
[0085] Through the above approach, this application solves the problem of effectively distinguishing and modeling batch effects and real biological signals in multi-source heterogeneous maize stress resistance gene electronic expression profile data within a single analytical framework. It avoids false positive results caused by misjudging batch effects as key biological characteristics, while preserving real but weak biological signals to avoid false negative results. This approach accurately identifies key gene expression features that are closely related to maize stress resistance traits and show robust consistency in multi-source heterogeneous data, thus improving the accuracy and reliability of key gene identification.
[0086] Furthermore, step S1 includes:
[0087] S11: Obtain the biological stress condition description attached to each electronic expression profile data sample, and parse the description to extract the stress type, stress intensity, and stress duration;
[0088] S12: The extracted stress type, stress intensity, and stress duration are converted into a preset stress condition standardization system to obtain standardized stress categories;
[0089] S13: Group the electronic expression spectrum data according to the stress category, and normalize the electronic expression spectrum data within each group.
[0090] Among them, biological stress condition description refers to textual information related to electronic expression profile data samples that records the external environment or treatment conditions experienced by the sample.
[0091] Parsing description refers to the process of identifying and extracting structured biological information from descriptions of biological stress conditions using text analysis techniques, such as keyword matching, regular expressions, or natural language processing methods.
[0092] Stress type refers to the specific category of environmental pressure applied to a biological sample, including but not limited to drought, salinity, high temperature, low temperature, and pest infestation.
[0093] Stress intensity refers to the degree or level of applied stress, such as mild, moderate, severe, or specific physicochemical parameter values.
[0094] The duration of stress refers to the length of time a biological sample is exposed to a specific stress condition, such as 1 hour, 12 hours, 24 hours, or several days.
[0095] A predefined stress condition standardization system refers to a set of predefined classification and naming rules used to uniformly describe stress conditions from different sources. It can be constructed using hierarchical classification structures, label systems, or ontology models. A standardized stress category refers to a unified representation obtained by mapping the original stress type, intensity, and duration information to the standardization system. For example, both "mild drought treatment for 24 hours" and "24-hour weak drought" are standardized into the category "drought-mild-24h".
[0096] Grouping electron expression spectrum data according to stress category means grouping electron expression spectrum data samples with the same standardized stress category into the same logical set.
[0097] Normalization of electronic expression profile data within each group after grouping refers to independently applying a normalization algorithm within each group to eliminate non-biological differences between samples within the group.
[0098] For example, in practice, the descriptive information of each sample can be read from the electronic expression profile data file or its associated metadata. This descriptive information may include text such as "drought treatment, mild, 12 hours" or "salt stress, 200 mM NaCl, 24 h".
[0099] Next, these texts are parsed using predefined rules or machine learning models to extract structured information such as "stress type: drought," "stress intensity: mild," and "stress duration: 12 hours." This extracted information is then matched or mapped to a standardized list of stress conditions, for example, using "drought-mild-12 hours" as the standard category. Based on these standard categories, all samples labeled "drought-mild-12 hours" are grouped into one group, all samples labeled "salt stress-200mM-24h" are grouped into another, and so on.
[0100] Finally, for each group, a normalization method is applied independently. For example, the total expression level of each sample within the group can be calculated and the gene expression data of each sample can be adjusted by a scaling factor, or a more complex intragroup normalization algorithm can be used to eliminate technical differences between samples within the group.
[0101] Furthermore, step S13 includes:
[0102] S131: Group the electronic expression spectrum data samples according to the standardized stress category;
[0103] S132: For each group after grouping, select a set of genes whose expression fluctuation is lower than a preset threshold in all samples within that group as a reference gene set;
[0104] S133: Calculate the normalization factor for each sample based on the reference gene set;
[0105] S134: Apply the normalization factor to all gene expression data of the corresponding sample.
[0106] Specifically, after grouping the electronic expression profile data samples according to standardized stress categories, for each group, a set of genes whose expression fluctuations are below a preset threshold across all samples in that group needs to be selected as a reference gene set. Here, expression fluctuations below the preset threshold mean that the gene exhibits a relatively stable expression level across different samples under the same stress category.
[0107] A preset threshold is a numerical standard used to define the stability of an expression, and its setting can be based on experience or data characteristic analysis.
[0108] A reference gene set refers to a collection of genes whose expression is relatively stable under specific conditions. Changes in the expression levels of these genes tend to reflect systematic differences caused by non-biological factors. Furthermore, based on the selected reference gene set, a normalization factor is calculated for each sample.
[0109] The normalization factor is a proportion or offset used to correct for systematic biases between samples, and its calculation depends on the overall expression level of a reference gene set in the sample. Finally, the calculated normalization factor is applied to all gene expression data of the corresponding sample to correct the overall expression level of that sample. In a specific embodiment, calculating the normalization factor for each sample based on the reference gene set can be specifically determined by calculating the median expression level of the reference gene set in each sample. Applying the normalization factor to all gene expression data of the corresponding sample can be specifically achieved by dividing the original expression level of all genes in the sample by the calculated normalization factor.
[0110] Through the above technical solution, this application can more accurately correct the systematic bias between electronic expression profile data samples, especially within the same stress category, thereby providing a "cleaner" and more accurate data input for subsequent variable component decomposition of each gene expression level, which helps to improve the accuracy of batch baseline component and biological response component separation, and thus improve the reliability of subsequent key gene screening.
[0111] Furthermore, step S2 includes:
[0112] S21: For the expression level of each gene in the standardized electronic expression profile data, within each experimental batch, determine the control group sample and the treatment group sample;
[0113] S22: Calculate the mean gene expression level of the control group sample as the batch baseline component of the gene;
[0114] S23: Calculate the difference between the gene expression levels of the treatment group samples and the batch baseline components, and use it as the biological response component of the gene.
[0115] This protocol details the specific methods for decomposing the variant components by first identifying control and treatment samples within each experimental batch. This distinction leverages the inherent control design structure of biological experiments, providing a foundation for subsequent component decomposition.
[0116] Next, the mean expression level of each gene in the control group samples was calculated, and this mean was defined as the batch baseline component of that gene in that experimental batch. This step separates the basal expression level within the batch, i.e., the portion mainly affected by the batch effect, from the expression level in the treatment group.
[0117] Subsequently, the difference between the gene expression levels in the treatment group samples and the baseline composition of that batch was calculated, and this difference was defined as the biological response component of that gene in the treatment group samples. This method of difference calculation directly quantifies the expression changes caused by the treatment relative to the control state of that batch.
[0118] By independently calculating the mean of the control group and the difference of the treatment group within each experimental batch, this scheme can more directly and accurately separate the baseline expression level (affected by batch effect) from the actual expression changes caused by stress treatment.
[0119] This method fully utilizes experimental design information, avoids conflating batch effects with biological responses, and thus provides a purer biological response signal for subsequent steps, improving the accuracy of key gene identification. Compared to simply performing general variable component decomposition, this approach introduces control-treatment comparisons within experimental batches, making the decomposition results more biologically meaningful and effectively distinguishing between systematic variations caused by batches and true variations caused by biological treatments.
[0120] For example, an electronic expression profile data matrix containing multiple experimental batches can be obtained, where each column represents a sample and each row represents a gene. Each sample is labeled with a batch identifier and a treatment status identifier (control or treatment). For a specific gene in the data matrix, all samples are first iterated through, and then grouped according to the sample's batch identifier.
[0121] For the first experimental batch, the control group and treatment group samples belonging to that batch were identified. The average expression level of the gene in all control group samples of that batch was calculated, and this average was recorded as the batch baseline component of the gene in that batch.
[0122] Then, for each treatment group sample in the batch, the difference between the gene expression level in that sample and the previously calculated baseline component for that batch is calculated, and this difference is recorded as the biological response component of that gene in that sample. This process is repeated for all genes in all experimental batches, ultimately yielding the batch baseline component and biological response component for each gene in each sample.
[0123] Furthermore, step S3 includes:
[0124] S31: For each gene, obtain its biological response components in different experimental batches;
[0125] S32: For each biological response component, determine its expression change direction to obtain the set of gene expression change directions in each experimental batch;
[0126] S33: Calculate the proportion of consistent directions in the set of expression change directions;
[0127] S34: When the proportion of consistent directions reaches a preset threshold, the gene is determined to have cross-batch directional consistency, thereby extracting stable response components with consistent responses.
[0128] Determining the direction of expression change refers to determining whether gene expression in this experimental batch is upregulated, downregulated, or unchanged based on the values of biological response components.
[0129] The percentage of consistent expression direction in the set of expression change directions refers to the degree to which the expression change direction of a gene remains consistent across different experimental batches.
[0130] The present application provides an effective mechanism for extracting stable response components from biological response components by introducing the judgment of the direction of gene expression changes and the calculation of the consistency ratio across batches.
[0131] First, for each gene, its biological response components, obtained through variant component decomposition in different experimental batches, are acquired. These components theoretically reflect the gene's true response to biological treatment and form the basis for subsequent analysis. Next, for each biological response component, instead of directly comparing its numerical values, the direction of its expression change (e.g., upregulation, downregulation, or no change) is determined, and these directions are aggregated to form a set of expression change directions for that gene across different experimental batches. This direction-based determination, compared to a direct numerical-based determination, is more robust to potential residual numerical differences between batches and can more effectively capture consistent response patterns across batches, even if the response amplitude varies across batches.
[0132] Then, by calculating the proportion of consistent directions in the set of expression change directions, the stability or consistency of the gene response across different batches was quantified. A higher proportion indicates a more stable response pattern for the gene, and is more likely to be a genuine biological signal.
[0133] Finally, by setting a preset threshold, when the proportion of consistent directions reaches this threshold, the gene is determined to have cross-batch directional consistency, and its corresponding biological response component is extracted as a stable response component with consistent response. This screening method based on the proportion of directional consistency can effectively identify gene responses that are robust in multi-source heterogeneous data from biological response components, avoiding misjudging batch effect residues or random noise as stable signals.
[0134] This approach, combined with previous steps for decomposing variant components, further verifies and screens the stability of these components by using directional consistency after decomposing them into biological response components. This allows for more accurate identification of real biological signals in multi-source heterogeneous data analysis, laying the foundation for subsequent screening of key genes.
[0135] Furthermore, step S32 includes:
[0136] S321: For the value of each biological response component, set a positive threshold and a negative threshold;
[0137] S322: When the value of a biological response component is greater than the positive threshold, the expression change direction of the biological response component is determined to be upregulation;
[0138] S323: When the value of a biological response component is less than the negative threshold, the direction of the change in the expression of the biological response component is determined to be downregulation;
[0139] S324: When the value of a biological response component is between the negative threshold and the positive threshold, the direction of change in the expression of the biological response component is determined to be no change;
[0140] S325: The direction of upregulation, downregulation, or no change is taken as the direction of gene expression change in this experimental batch, and the set of gene expression change directions in each experimental batch is obtained.
[0141] Among them, the positive threshold is the lower limit value used to define the significant upward adjustment of the value of the biological response component; the negative threshold is the upper limit value used to define the significant downward adjustment of the value of the biological response component, which is usually the opposite of the positive threshold or set according to the symmetry of the data distribution.
[0142] The numerical value of biological response components refers to the quantitative representation of the change in gene expression relative to the batch baseline under specific experimental conditions, obtained through the decomposition of variant components.
[0143] This application provides a more accurate and robust method for determining the direction of expression changes of biological response components by setting positive and negative thresholds for the values of biological response components and quantifying the direction of expression changes based on the relationship between the values and the thresholds.
[0144] Specifically, for each biological response component, a positive threshold and a negative threshold are set. This setting step provides a quantitative standard for determining the direction of expression changes. By setting positive and negative thresholds, instead of simply using zero as the boundary, signals with significant expression changes (greater than the positive threshold or less than the negative threshold) and small fluctuations close to zero (between the negative threshold and the negative threshold) can be distinguished, avoiding misjudging small fluctuations as upregulation or downregulation.
[0145] Based on the set positive threshold, the standard for "upregulation" is clearly defined. Only when the value of the biological response component is significantly higher than zero and exceeds the threshold is it considered to truly express an upregulation signal.
[0146] Similarly, based on the set negative threshold, the criteria for "downregulation" are clearly defined. Only when the value of the biological response component is significantly lower than zero and less than the threshold is it considered a true downregulation signal.
[0147] When the value of a biological response component is between the negative and positive thresholds, the direction of change in the expression of the biological response component is determined to be unchanged. This step defines the state of "no change," meaning that although the value of the biological response component may not be zero, its fluctuation amplitude is insufficient to cross the set threshold, and it is considered noise or an insignificant biological response, thus avoiding the over-interpretation of small fluctuations.
[0148] Finally, the directions of upregulation, downregulation, or no change are used as the expression change directions of the gene in that experimental batch, and these are aggregated to obtain a set of expression change directions for the gene in each experimental batch. This step transforms the quantitative judgment result into discrete expression change directions (upregulation, downregulation, no change) and collects this directional information for the gene in different experimental batches, forming a set. This set is the basis for subsequent calculation of the proportion of consistent directions. Through this standardized directional representation, cross-batch directional consistency judgment becomes possible, and the accuracy and robustness of the judgment are improved.
[0149] This method, combined with steps such as obtaining the biological response components of genes in different experimental batches and calculating the proportion of consistent directions in the set of expression change directions, can more accurately identify genes that exhibit stable response patterns in multi-source heterogeneous data, thereby effectively extracting stable response components.
[0150] Furthermore, step S33 includes:
[0151] S331: For each gene, obtain the set of expression change directions in each experimental batch;
[0152] S332: Assign a weight to each experimental batch in the set representing the direction of change, with the weight being proportional to the data quality within the batch;
[0153] S333: Based on the set of expression change directions and their weights, determine the expression change direction with the largest total weight as the consistent direction;
[0154] S334: The ratio of the sum of weights in the consistent direction to the sum of weights in all experimental batches is taken as the proportion of the consistent direction.
[0155] The method proposed in this application first obtains a set of expression change directions for each gene across various experimental batches, which serves as the foundational data for subsequent consistency calculations. Next, a weight is assigned to each experimental batch within the expression change direction set, and this weight is proportional to the data quality within the batch. Since the reliability of data varies across different experimental batches, assigning a larger weight to batches with higher data quality allows subsequent consistency calculations to more fully consider the contribution of high-quality data and reduce the interference of low-quality data on the results.
[0156] Then, based on the set of expression change directions and their assigned weights, the expression change direction with the largest total weight is determined as the consensus direction. This means that when determining the overall consensus trend of genes, it is no longer simply a matter of which direction has the most batches, but rather which direction corresponds to the largest total batch weight, thus ensuring that batches with higher data quality have a greater influence in determining the consensus direction.
[0157] Finally, the ratio of the sum of weights for consistent directions to the sum of weights for all experimental batches is used as the proportion of consistent directions. This calculated proportion is a weighted consistency index that comprehensively reflects the robustness of gene expression change direction across different experimental batches, and the weights reflect the differences in data quality between different batches. This weighted proportion measures the cross-batch consistency of gene expression changes more accurately than a simple frequency proportion, providing a more reliable quantitative basis for subsequent screening of truly stable genes related to stress resistance.
[0158] Furthermore, step S4 includes:
[0159] S41: Obtain the stress resistance signal intensity of each gene in the electronic expression profile data;
[0160] S42: Calculate the combined fraction of the proportion of the strength of the resistance signal and the consistent direction of the stable response component;
[0161] S43: Screen for preliminary key genes based on combination scores.
[0162] Among them, stress resistance signal intensity refers to a quantitative indicator that reflects the degree of change in gene expression under stress conditions; the combination score refers to an evaluation value that comprehensively reflects the proportion of gene stress resistance signal intensity and stable response components in the same direction.
[0163] In one embodiment, the stress resistance signal intensity of each gene in the electronic expression profile data can be specifically calculated as the log2 value of the ratio of the median expression level of the treatment group to the median expression level of the control group. The combination score of the stress resistance signal intensity and the proportion of the stable response component aligned in the same direction can be specifically calculated as the product of the absolute value of the stress resistance signal intensity and the proportion of the stable response component aligned in the same direction. A product threshold can be specifically set to screen preliminary key genes based on the combination score; genes with a product greater than this threshold are selected as preliminary key genes.
[0164] By calculating the combined score of the proportion of stress resistance signal intensity and stable response components in the same direction, and screening preliminary key genes based on this combined score, this scheme can comprehensively consider the strength of gene response and robustness in multi-source heterogeneous data, improve the accuracy of preliminary key gene screening, and reduce the risk of misidentifying genes affected by batch effects or random factors as key genes.
[0165] Furthermore, step S5 includes:
[0166] S51: For each preliminary key gene, obtain its corresponding batch baseline components and biological response components;
[0167] S52: Calculate the relative contribution of the batch baseline components of the preliminary key genes to the biological response components;
[0168] S53: Based on relative contribution, exclude genes in the preliminary key genes that are mainly contributed by the batch baseline components to obtain the key genes.
[0169] The present application provides a precise screening method based on the relative contributions of batch baseline components and biological response components for further screening of preliminary key genes.
[0170] First, for each gene identified as a preliminary key gene in the initial screening, the batch baseline component and biological response component obtained from the variant component decomposition were acquired. This is the basis for subsequent quantitative analysis, and both components need to be obtained simultaneously because they represent the background level of gene expression caused by non-biological factors and the true response caused by biological treatment, respectively.
[0171] Next, the relative contributions of the batch baseline components and biological response components of the preliminary key genes are calculated. This step is the core of this protocol. By calculating the relative proportions of these two components, it quantitatively assesses whether the gene expression changes are primarily driven by batch effects or primarily by biological responses. This quantitative assessment provides a clear basis for subsequent precise screening, distinguishing which genes' "preliminary criticality" reflects a true biological response and which may merely be background differences caused by batch effects.
[0172] Finally, based on the calculated relative contributions, preliminary key genes that are mainly contributed by batch baseline components are excluded, thus obtaining the final key genes. This means setting a threshold or judgment criterion: if the contribution of batch baseline components to the expression level change of a preliminary key gene is higher than that of biological response components, then the "keyness" of that gene is considered to be more of an illusion caused by batch effects and is excluded.
[0173] This secondary screening based on relative contribution effectively eliminates false-positive genes severely affected by batch effects, ensuring that the final selected key genes are those with truly robust biological responses. Combined with step S2 above, which decomposes gene expression levels into variant components to obtain batch baseline and biological response components, this approach allows subsequent screening to move beyond relying solely on the absolute values of single components. Instead, it enables judgment based on the relative relationships between components, demonstrating a stronger ability to distinguish between batch effects and biological responses, thus improving the accuracy and reliability of key gene identification.
[0174] In one embodiment, for each preliminary key gene, its corresponding batch baseline component value and biological response component value can be read from a data structure storing the results of variable component decomposition. Then, the ratio of the absolute value of the biological response component to the sum of the absolute values of the batch baseline components can be calculated as the relative contribution. For example, if the biological response component is R and the batch baseline component is B, the relative contribution can be calculated as |R| / (|R| + |B|). Next, a threshold for the relative contribution can be set, for example, 0.4. For each preliminary key gene, if its calculated relative contribution is less than 0.4, it is determined that the gene is mainly contributed by the batch baseline component and is excluded from the preliminary key gene list. The genes that are ultimately retained are the key genes.
[0175] Please refer to Figure 2 , Figure 3 A maize stress resistance gene electronic expression profiling network analysis system, used to implement any of the above methods, the system comprising:
[0176] Module 201: Acquires electronic expression profile data of multi-source heterogeneous maize stress resistance genes and performs normalization processing on the electronic expression profile data;
[0177] Decomposition Module 202: Decomposes the expression level of each gene in the normalized electronic expression profile data into variant components to obtain the batch baseline components and biological response components;
[0178] Extraction module 203: Extracts stable response components that exhibit consistent responses across different experimental batches from biological response components;
[0179] First screening module 204: acquires the stress resistance signal intensity of each gene in the electronic expression profile data, and screens preliminary key genes based on the stress resistance signal intensity and stable response components;
[0180] Second screening module 205: Based on the batch baseline components, further screen the preliminary key genes, exclude genes affected by batch effects, and obtain the key genes.
[0181] The acquisition module 201 is a component responsible for receiving, reading and initially processing electronic expression spectrum data from different sources and in different formats. It can be implemented using a data interface, a file reader, a data parser and a data preprocessor.
[0182] Among them, the decomposition module 202 refers to the component that performs mathematical transformation on gene expression data to decompose its total variation into components from different sources. It can be implemented using statistical models, machine learning algorithms or signal processing techniques.
[0183] The extraction module 203 refers to the component that identifies and separates features that exhibit consistency or stability under different conditions from the decomposed biological response components. It can be implemented using pattern recognition algorithms, cluster analysis, or consistency scoring mechanisms.
[0184] The first screening module 204 refers to a component that performs preliminary filtering of the gene set according to multiple preset criteria. It can be implemented using threshold-based filtering, sorting algorithms, or multi-index scoring systems.
[0185] The second screening module 205 refers to a component that uses specific component information to re-verify and refine the preliminary screening results. It can be implemented using ranking based on component contribution, statistical testing, or machine learning classifiers.
[0186] This system provides an automated, efficient, and accurate execution platform by breaking down complex data analysis methods into a series of functional modules and defining the data flow and processing logic between these modules.
[0187] Specifically, the acquisition module 201 is first responsible for aggregating massive amounts of raw electronic expression spectrum data from different databases and different experimental batches, and performing necessary format conversions and preliminary quality control to ensure the uniformity and usability of the input data.
[0188] Subsequently, this normalized data is sent to the decomposition module 202, which performs a fine-grained decomposition of the expression levels of each gene, separating the total expression changes into systematic bias caused by the experimental batch (batch baseline component) and the true response caused by biological treatment (biological response component). This decomposition is the core of the system, achieving the separation of batch effects from biological signals within a single framework, avoiding signal loss that may occur with traditional methods of calibration before analysis.
[0189] Next, the extraction module 203 works on the decomposed biological response components, spanning different experimental batches, to identify and extract gene expression features that show stable and consistent responses under multiple conditions. This helps to pinpoint genes that are truly related to maize stress resistance, rather than accidental or batch-specific responses.
[0190] Then, the first screening module 204 combines the stress resistance signal intensity information of the gene itself (e.g., indicators obtained in advance by other methods or calculated from data) with the stable response component information obtained by the extraction module to perform preliminary gene screening and quickly narrow down the range of candidate genes.
[0191] Finally, the second screening module 205 uses the batch baseline components obtained by the decomposition module to perform a second verification of the preliminary screening results, excluding those genes that show a certain response but whose expression changes are mainly affected by the batch effect, thereby obtaining the final set of key genes.
[0192] Through the sequential processing and information transmission of the above modules, the system can automatically complete the entire complex process from raw heterogeneous data to key gene identification, overcoming the challenge of manual processing or simple tools being unable to cope with large-scale, high-dimensional data, and ensuring the accuracy and reliability of the analysis results.
[0193] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any such actual relationship or order between these entities or operations.
[0194] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for analyzing the electronic expression profile network of maize stress resistance genes, characterized in that, The method includes: S1: Obtain electronic expression profile data of multi-source heterogeneous maize stress resistance genes, and perform normalization processing on the electronic expression profile data; S2: The expression levels of each gene in the normalized electronic expression profile data are decomposed into variant components to obtain the batch baseline components and biological response components; S3: For each gene, obtain its biological response components in different experimental batches; for each biological response component, determine its expression change direction to obtain the set of expression change directions of the gene in each experimental batch; calculate the proportion of consistent directions in the set of expression change directions; when the proportion of consistent directions reaches a preset threshold, determine that the gene has cross-batch directional consistency, thereby extracting the stable response components with consistent responses. S4: Obtain the stress resistance signal intensity of each gene in the electronic expression profile data, and screen preliminary key genes based on the stress resistance signal intensity and the stable response components; S5: Based on the batch baseline components, further screen the preliminary key genes to exclude genes affected by batch effects and obtain the key genes.
2. The method for analyzing the electronic expression profile of maize stress resistance genes according to claim 1, characterized in that, Step S1 includes: S11: Obtain the biological stress condition description attached to each of the electronic expression spectrum data samples, and parse the description to extract the stress type, stress intensity and stress duration; S12: The extracted stress type, stress intensity, and stress duration are converted into a preset stress condition standardization system to obtain a standardized stress category; S13: Group the electronic expression spectrum data according to the stress category, and normalize the electronic expression spectrum data within each group.
3. The method for analyzing the electronic expression profile of maize stress resistance genes according to claim 2, characterized in that, Step S13 includes: S131: Group the electronic expression spectrum data samples according to the standardized stress categories; S132: For each group after grouping, select a set of genes whose expression fluctuation is lower than a preset threshold in all samples within that group as a reference gene set; S133: Calculate the normalization factor for each sample based on the reference gene set; S134: Apply the normalization factor to all gene expression data of the corresponding sample.
4. The method for analyzing the electronic expression profile network of maize stress resistance genes according to claim 1, characterized in that, Step S2 includes: S21: For the expression level of each gene in the normalized electronic expression profile data, within each experimental batch, determine the control group sample and the treatment group sample; S22: Calculate the mean expression level of the gene in the control group samples as the batch baseline component of the gene; S23: Calculate the difference between the expression level of the gene in the treatment group sample and the batch baseline component, and use it as the biological response component of the gene.
5. The method for analyzing the electronic expression profile network of maize stress resistance genes according to claim 1, characterized in that, Step S3 includes: S321: For the value of each of the biological response components, set a positive threshold and a negative threshold; S322: When the value of the biological response component is greater than the positive threshold, the expression change direction of the biological response component is determined to be upregulation; S323: When the value of the biological response component is less than the negative threshold, it is determined that the expression change direction of the biological response component is down-regulation; S324: When the value of the biological response component is between the negative threshold and the positive threshold, it is determined that the direction of the expression change of the biological response component is unchanged; S325: The directions of upregulation, downregulation, or no change are taken as the expression change directions of the gene in the experimental batch, and the sets of expression change directions of the gene in each experimental batch are obtained.
6. The method for analyzing the electronic expression profile network of maize stress resistance genes according to claim 5, characterized in that, Step S3 also includes: S331: For each gene, obtain the set of expression change directions in each experimental batch; S332: Assign a weight to each experimental batch in the set of expression change directions, the weight being proportional to the data quality within the batch; S333: Based on the set of expression change directions and the weights, determine the expression change direction with the largest total weight as the consistent direction; S334: The ratio of the sum of weights of the consistent direction to the sum of weights of all experimental batches is taken as the proportion of the consistent direction.
7. The method for analyzing the electronic expression profile of maize stress resistance genes according to claim 6, characterized in that, Step S4 includes: S41: Obtain the stress resistance signal intensity of each gene in the electronic expression profile data; S42: Calculate the combined fraction of the proportion of the anti-reverse signal strength and the consistent direction of the stable response component; S43: Based on the combined scores, screen preliminary key genes.
8. The method for analyzing the electronic expression profile network of maize stress resistance genes according to claim 1, characterized in that, Step S5 includes: S51: For each of the aforementioned preliminary key genes, obtain the corresponding batch baseline components and biological response components; S52: Calculate the relative contribution of the batch baseline component of the preliminary key gene to the biological response component; S53: Based on the relative contribution, exclude genes in the preliminary key genes that are mainly contributed by the batch baseline components to obtain the key genes.
9. A maize stress resistance gene electronic expression profiling network analysis system, characterized in that, The system for implementing the method according to any one of claims 1-8 comprises: Acquisition module: Acquires electronic expression profile data of multi-source heterogeneous maize stress resistance genes and performs normalization processing on the electronic expression profile data; Decomposition module: Decomposes the expression level of each gene in the normalized electronic expression profile data into variant components to obtain batch baseline components and biological response components; Extraction module: For each gene, obtain its biological response components in different experimental batches; for each biological response component, determine its expression change direction to obtain the set of expression change directions of the gene in each experimental batch; calculate the proportion of consistent directions in the set of expression change directions; when the proportion of consistent directions reaches a preset threshold, determine that the gene has cross-batch directional consistency, thereby extracting the stable response components with consistent responses. First screening module: Obtain the stress resistance signal intensity of each gene in the electronic expression profile data, and screen preliminary key genes based on the stress resistance signal intensity and the stable response components; The second screening module further screens the preliminary key genes based on the batch baseline components, excluding genes affected by batch effects, and obtains the key genes.
Citation Information
Patent Citations
Gene detection data cleaning method based on artificial intelligence
CN119046622A
Quantitative methods, systems and apparatuses for gene expression analysis
WO1999058720A1