A method and system for processing genetic sequencing data

CN119601083BActive Publication Date: 2026-08-11DALIAN YIKANG BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-16
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而一种基因测序数据处理方法及系统存在着对基因突变方向预测不准确的问题以及对性状影响变化的预测不准确的问题

Benefits of technology

[0068] Gene quality improvement module: Differential expression analysis is performed based on the crop trait change storage database to obtain crop differential expression data; crop gene anomaly detection is performed based on the crop differential expression data to obtain crop gene anomaly data; crop gene quality improvement is performed based on the crop gene anomaly data to obtain crop gene quality improvement data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119601083B_ABST
    Figure CN119601083B_ABST
Patent Text Reader

Abstract

This invention relates to the field of gene sequencing technology, and more particularly to a gene sequencing data processing method and system. The method includes the following steps: acquiring a crop sample dataset, performing cell disruption and gene sequencing to obtain crop cell gene sequence data; performing variation analysis on the gene sequence data, dividing it into gene mutation sequences and gene sequence deletion data, further performing gene expression disorder analysis to predict disease resistance attenuation and trait changes; storing the analysis results in a gene storage database to achieve the preservation and retrieval of gene expression trait change data; performing differential expression analysis on the crop trait change storage database to detect crop gene abnormalities; and finally obtaining improved crop gene data through crop gene quality improvement methods. This invention optimizes crop gene sequencing, making gene mutation prediction more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene sequencing technology, and in particular to a gene sequencing data processing method and system. Background Technology

[0002] Crop genomics research has become one of the core areas for improving crop yield, quality, disease resistance, and stress tolerance. The rapid development of crop gene sequencing technology, especially high-throughput genome sequencing technology, has greatly promoted research on the correlation between crop genomes and phenotypic characteristics. Deep sequencing of crop genomes allows for a comprehensive understanding of crop genetic information, gene function, variation patterns, and their relationships with traits such as environmental adaptability and disease resistance. Therefore, efficient processing, analysis, and interpretation of gene sequencing data are crucial aspects of crop genomics research. Crop genome data is typically very large and contains various complex information, such as gene mutations, structural variations, and transcriptional regulation. Furthermore, crop genomes often exhibit high gene repetition, complex gene families, non-standard genome structures, and interspecies diversity. However, current gene sequencing data processing methods and systems suffer from inaccurate predictions of gene mutation directions and changes in trait influences. Summary of the Invention

[0003] Therefore, it is necessary to provide a gene sequencing data processing method and system to solve at least one of the above-mentioned technical problems.

[0004] To achieve the above objectives, a gene sequencing data processing method and system includes the following steps:

[0005] Step S1: Obtain crop sample dataset; grind crop sample cells according to the crop sample dataset to obtain crop sample cell fragmentation data; perform cell gene sequencing on the crop sample cell fragmentation data to obtain crop cell gene sequence data;

[0006] Step S2: Perform gene sequence variation analysis on crop gene sequence data to obtain crop gene sequence variation data; classify gene sequence variations based on crop gene sequence variation data to obtain gene mutation sequence data and gene sequence deletion data; perform gene expression disorder analysis based on gene mutation sequence data and gene sequence deletion data to obtain gene expression disorder data.

[0007] Step S3: Based on the gene expression disorder data, predict the decline of crop disease resistance to obtain crop disease resistance decline data; perform trait change analysis on the crop disease resistance decline data and gene expression disorder data to obtain gene expression trait change data; use a preset gene storage database to store the gene expression trait change data to obtain a crop trait change storage database.

[0008] Step S4: Perform differential expression analysis based on the crop trait change storage database to obtain crop differential expression data; perform crop gene anomaly detection based on the crop differential expression data to obtain crop gene anomaly data; perform crop gene quality improvement based on the crop gene anomaly data to obtain crop gene quality improvement data.

[0009] This invention acquires crop sample datasets and disrupts cells, then combines this with cell gene sequencing technology to comprehensively obtain crop gene sequence data, providing accurate foundational data for subsequent gene variation analysis and functional studies. Cell gene sequencing, through efficient full-sequence analysis of the crop genome, reveals gene diversity and variation, laying the foundation for further gene function analysis and trait research. Variation analysis of gene sequence data can identify mutations and deletions in the genome; these variations directly affect crop traits such as resistance, stress tolerance, and growth. Classifying gene sequence variations and distinguishing mutation and deletion types helps determine which gene variations are associated with target traits. Furthermore, gene expression disorder analysis based on this variation data can further screen affected genes and predict their abnormalities in biological processes, helping to identify key pathogenic or harmful gene variations. Analyzing gene expression disorder data can predict crop disease resistance decline, providing a basis for crop health assessment and identifying which genes or variations lead to weakened crop resistance under specific environmental or pathogenic conditions, thus providing data support for disease resistance enhancement. Simultaneously, combining gene expression aberration data with trait change analysis helps to understand the association between genes and crop trait changes, identify key trait genes, and promote precision breeding. Storing this gene expression trait change data in a gene storage database not only enables efficient data management but also provides a systematic data reference for future crop trait improvement. The establishment of this database facilitates continuous updates and provides a foundation for genomics research on different crops, supporting the implementation of gene-phenotype association analysis under different environments. Further differential expression analysis can identify gene expression differences under different conditions at the whole-genome level, helping to understand the response and regulatory mechanisms of crop genes under different environments, treatments, or mutations. Detecting crop gene anomalies based on differential expression data can reveal abnormalities in gene expression, helping to discover potential mutations or variant regions, especially key genes affecting important crop traits (such as yield, stress resistance, and nutrient composition). Combining the results of gene anomaly detection, improvements in crop gene quality can be achieved through precise gene editing or breeding techniques to optimize key genes, thereby improving the overall quality and adaptability of crops and enhancing agricultural production efficiency. Therefore, this invention is an optimization of a traditional gene sequencing data processing method, which solves the problems of inaccurate prediction of gene mutation direction and inaccurate prediction of gene influence on trait changes in the traditional gene sequencing data processing method, and improves the accuracy of predicting gene mutation direction and gene influence on trait changes in crops.

[0010] Preferably, step S1 includes the following steps:

[0011] Step S11: Obtain crop sample dataset;

[0012] Step S12: Grind crop sample cells according to the crop sample dataset to obtain crop sample cell fragmentation data;

[0013] Step S13: Based on the data of cell disruption in the crop sample, centrifuge and filter impurities to obtain the sample cell impurity filtration data;

[0014] Step S14: Perform cell gene sequencing on the sample cell impurity filtering data to obtain crop cell gene sequence data.

[0015] This invention, by acquiring crop sample datasets and disrupting cells, effectively extracts intracellular genetic material from crop samples, providing the first step of raw materials for subsequent genomics analysis. Cell disruption not only ensures the effective release of DNA or RNA from the cell nucleus but also improves the efficiency and accuracy of subsequent analysis. The disrupted cell samples provide high-quality raw material for genomic analysis, ensuring the reliability of subsequent sequencing results. After cell disruption, centrifugation is used for impurity filtration, removing unwanted impurities such as cell debris, proteins, lipids, and other interfering substances. Impurity filtration not only improves sample purity but also reduces the impact of interfering substances on genomic sequencing data, thereby ensuring higher quality and accuracy of the final gene sequence data. This step helps improve the signal-to-noise ratio of the data, making subsequent genomic analysis more precise. Gene sequencing of samples after cell impurity filtration is a crucial step in obtaining crop genomic information. Gene sequencing can capture the genetic information of crops across the entire genome, providing rich data support for subsequent gene function analysis, variation detection, and phenotypic association studies. Sequencing technology can identify the presence and expression of genes by reading gene sequences in samples, thereby revealing the genetic background, variation characteristics, and pathogenic or dominant genes of crops.

[0016] Preferably, step S14 includes the following steps:

[0017] Step S141: Extract nucleic acid from the sample cells by filtering out impurities in the sample cells to obtain a crop sample cell nucleic acid dataset;

[0018] Step S142: Perform DNA digestion based on the crop sample cell nucleic acid dataset to obtain crop cell DNA fragment data;

[0019] Step S143: Amplify the DNA fragment data of crop cells to obtain amplified data of crop cell DNA fragments;

[0020] Step S144: Gene sequencing is performed based on the amplified DNA fragment data of crop cells to obtain crop cell gene sequence data.

[0021] This invention extracts cellular nucleic acids from sample cells by filtering out impurities, ensuring the extraction of high-purity nucleic acid components and providing high-quality raw materials for subsequent genomic analysis. This step is crucial for the success of the entire genomics research, as pure nucleic acid samples avoid interference from impurities during sequencing, improving the accuracy of downstream analysis. High-quality cellular nucleic acid data provides a stable and reliable foundation for subsequent DNA digestion, amplification, and gene sequencing. After extracting cellular nucleic acids, DNA digestion breaks down large DNA molecules into smaller fragments. This process is a common technique in genomic analysis, designed to facilitate subsequent genome sequencing and fragment assembly. Digestion not only ensures the generation of DNA fragments suitable for sequencing but also helps researchers perform targeted analysis of specific genes or regions, identifying structural variations, mutations, and other features of the target fragments. This process provides important data for constructing genomic datasets and gene annotation. Amplification of the DNA fragments after digestion allows for replication based on the existing DNA fragments, obtaining a sufficient number of fragments for subsequent analysis. The effect of DNA fragment amplification is to ensure sufficient abundance of the desired gene sequence in experiments for more accurate gene sequencing. This step is particularly important for research requiring high-throughput sequencing, as it ensures the quality and reliability of sequencing data and avoids sequencing bias or data loss due to insufficient sample size. Finally, gene sequencing based on the amplified DNA fragments provides comprehensive information about the crop genome, offering core data for subsequent gene variation analysis, gene function studies, and phenotypic association studies. Gene sequencing is fundamental to revealing the genetic background, gene structure, and function of crops, efficiently identifying important features such as mutations, variations, duplications, and deletions in crop genes, thus providing precise evidence for crop genetic improvement, resistance development, and precision breeding.

[0022] Preferably, step S2 includes the following steps:

[0023] Step S21: Perform gene sequence variation analysis on crop gene sequence data to obtain crop gene sequence variation data;

[0024] Step S22: Based on the crop gene sequence variation data, divide the gene sequence variation to obtain gene mutation sequence data and gene sequence missing data;

[0025] Step S23: Predict genome stability decay based on gene mutation sequence data to obtain genome stability decay data;

[0026] Step S24: Perform gene deletion anomaly analysis based on gene sequence deletion data to obtain crop gene deletion anomaly data;

[0027] Step S25: Perform gene expression disorder analysis on the genomic stability decay data and crop gene deletion abnormality data to obtain gene expression disorder data.

[0028] This invention analyzes gene sequence variation in crop gene sequence data to gain a deeper understanding of various variation types in the crop genome, including mutations, insertions, and deletions, thereby revealing the genetic changes affecting crop traits. This step provides a foundation for identifying key gene variations, offering crucial information for subsequent trait improvement, disease resistance enhancement, and breeding strategies. Variation analysis provides strong support for studying crop genetic diversity, adaptability, and performance under different environmental conditions. The classification of gene sequence variations further refines the variation types, distinguishing the specific characteristics of gene mutations and gene deletions. By clarifying these variation types, scientists can better understand which variations affect crop traits, especially those leading to the deletion of undesirable genes or the mutation of important genes. Identifying mutation and deletion regions is crucial for precision breeding and crop genome repair, facilitating targeted improvements in crop disease resistance, stress tolerance, and production performance. Gene mutation sequence data can be further used to predict genomic stability decay, revealing the potential impact of mutation accumulation and genomic instability on crops. Genome stability is central to crop genetic quality. Predicting genome stability decay can provide scientific guidance for crop genetic improvement, healthy breeding, and disease prevention, especially by predicting high-risk regions with excessive mutation frequencies, thus providing early warnings for improvement programs. When analyzing gene sequence deletion data, gene deletion anomaly analysis can help identify missing gene regions and assess their potential impact on crops. Gene deletions lead to the loss of specific functions, affecting crop growth, development, or disease resistance. Understanding these functional defects is crucial for precise crop performance improvement. By identifying deletion anomalies, scientists can further analyze the role of missing genes, providing data support for improvement programs. Finally, combining genome stability decay data with crop gene deletion anomaly data for gene expression disorder analysis can help identify gene expression disturbances caused by gene mutations or deletions, thereby understanding how these gene variations affect the overall gene regulatory network of crops. This analysis provides a new perspective on understanding the molecular mechanisms of crop trait changes, providing theoretical basis and technical support for crop genetic improvement, trait optimization, and gene function restoration.

[0029] Preferably, step S23 includes the following steps:

[0030] Step S231: Perform gene mutation accumulation detection based on gene mutation sequence data to obtain gene mutation accumulation data;

[0031] Step S232: Perform chromosome structural break detection on the accumulated gene mutation data to obtain chromosome structural break data;

[0032] Step S233: Based on the accumulated gene mutation data, predict the probability of chromosome rearrangement to obtain chromosome rearrangement probability data;

[0033] Step S234: Calculate the gene duplication probability based on the chromosome rearrangement probability data to obtain the gene duplication probability data;

[0034] Step S235: Perform genome stability decay prediction on chromosome structural breakage data and gene duplication probability data to obtain genome stability decay data.

[0035] This invention utilizes gene mutation sequence data for gene mutation accumulation detection, enabling scientists to gain a deeper understanding of mutation accumulation in crop genomes and identify regions with high mutation frequencies. This accumulation of mutations is often associated with changes in crop traits such as adaptability, disease resistance, and production performance. Timely detection of mutation accumulation helps identify potential genomic instability, assess the impact of mutations on long-term crop adaptability, and provide a basis for further breeding strategies. Building upon gene mutation accumulation detection, chromosome structural break detection can reveal changes in chromosome structure caused by mutation accumulation or external environmental factors. Chromosomal structural breaks can trigger genomic instability, thereby affecting normal gene expression and regulation, leading to disordered or adverse changes in crop traits. Detecting chromosome structural breaks allows for effective monitoring of crop genome integrity, providing necessary information for preventing and repairing genomic damage. Furthermore, predicting chromosome rearrangement probabilities based on gene mutation accumulation data can reveal rearrangement events that can occur during chromosome evolution. These rearrangements lead to gene duplication, deletion, or rearrangement, thus affecting gene function and expression. Understanding the probability of chromosome rearrangements helps predict future changes in crop genomes, assisting researchers in more precise genome design and optimization during breeding. In the context of chromosomal rearrangements, calculating gene duplication probabilities can help analyze which genes are duplicated, leading to overexpression or dysfunction. Gene duplication not only causes abnormal gene expression levels but also affects crop traits such as metabolism and resistance. By calculating gene duplication probabilities, potential "hotspot" regions can be identified, providing targets for subsequent gene modification and mutation repair. Finally, predicting genome stability decay using chromosomal breakage data and gene duplication probability data can comprehensively assess genome stability and provide early warnings for crop genetic improvement. Genome stability is directly related to crop genetic diversity and breeding potential. By predicting genome stability decay, researchers can take corresponding measures to ensure crop genome stability, avoid trait degradation due to genome instability, and safeguard long-term crop productivity and adaptability.

[0036] Preferably, step S24 includes the following steps:

[0037] Step S241: Based on the gene sequence deletion data, predict the loss of protein function to obtain protein function loss data;

[0038] Step S242: Collect protease gene fragment deletion data from the gene sequence deletion data to obtain protease gene fragment deletion data;

[0039] Step S243: Based on the data of missing gene fragments of white enzyme, perform protease inactivation detection to obtain data on protease inactivation in crops;

[0040] Step S244: Perform crop resistance weakening tests on crop protease inactivation data and protein function loss data to obtain crop resistance weakening data;

[0041] Step S245: Based on the data on crop resistance reduction and crop protease inactivation, predict the decline in crop environmental adaptability and obtain the data on the decline in crop environmental adaptability.

[0042] Step S246: Perform gene deletion anomaly analysis based on crop environmental adaptability decline data and crop resistance weakening data to obtain crop gene deletion anomaly data.

[0043] This invention uses gene sequence deletion data to predict protein loss of function, accurately identifying protein loss due to gene deletion. This process provides core data for studying the impact of gene deletion on crop physiological and biochemical functions, helping scientists predict which protein loss in crops will negatively affect traits such as disease resistance and stress tolerance. Protein loss of function prediction provides important clues for crop gene functional annotation and genetic improvement, thus avoiding the introduction of mutations that lead to loss of function when improving crops. Based on gene sequence deletion, by collecting protease gene fragment deletion data, specific fragment deletions affecting protease function can be further precisely identified. These protease gene fragment deletions are closely related to crop growth, development, and disease resistance. By collecting these deleted fragments, researchers can understand which protease deletions affect crop growth and resistance, providing targeted basis for subsequent functional repair and gene improvement. Further protease inactivation detection based on protease gene fragment deletion data can help identify protease loss of function in crops caused by gene deletion or mutation. This detection can clearly identify which protease inactivation affects crop metabolism, disease resistance, and physiological functions, especially since proteases often play an important role in crop defense mechanisms. Detecting protease inactivation allows researchers to provide early warnings of crop resistance and adjust breeding strategies accordingly. Testing crop resistance attenuation using protease inactivation and protein loss data helps assess the decline in crop disease resistance under environmental stress. This test reveals the specific impact of gene deletions or protein loss on crop resistance, predicting whether crop resistance will significantly weaken with these factors, thus providing a basis for resistance-enhancing breeding strategies. Further prediction of environmental adaptability decline based on resistance attenuation and protease inactivation data can reveal changes in crop adaptability under varying environments. Environmental adaptability is a key factor for long-term crop survival and high yield. Predicting adaptability decline allows for the early identification of potential environmental stress risks, helping to develop precise breeding strategies to enhance crop environmental adaptability, particularly against adverse conditions such as drought, high temperatures, and salinity. Finally, gene deletion anomaly analysis of environmental adaptability decline and resistance attenuation data can help identify key gene deletions associated with environmental adaptability decline and resistance weakening. This analysis provides essential information for the precise repair of crop genomes, enabling researchers to further identify and repair gene deletions that lead to poor crop adaptability or weak disease resistance, thus providing theoretical support and technical means for efficient breeding and genome improvement.

[0044] Preferably, step S25 includes the following steps:

[0045] Step S251: Based on the genomic stability decay data, predict the mutation frequency growth to obtain mutation frequency growth data;

[0046] Step S252: Perform gene structure variation analysis based on mutation frequency growth data to obtain gene structure variation data;

[0047] Step S253: Based on gene structure variation data and mutation frequency growth data, predict transcription factor abnormalities to obtain gene transcription factor abnormality data;

[0048] Step S254: Perform genome integrity decay detection on crop gene deletion abnormality data and gene structure variation data to obtain genome integrity decay data;

[0049] Step S255: Detect gene regulatory network disruption based on genome integrity decay data and gene transcription factor abnormality data to obtain gene regulatory network disruption data;

[0050] Step S256: Perform gene expression disorder analysis on gene regulatory network disruption data and genome integrity decay data to obtain gene expression disorder data.

[0051] This invention, by predicting mutation frequency growth based on genomic stability decay data, can effectively identify trends and rates of mutation accumulation in the genome, thereby assessing long-term genome stability. Increased mutation frequency is often closely related to changes in crop traits such as genetic diversity, adaptability, and stress resistance. By accurately predicting mutation frequency growth, researchers can provide early warnings of future genetic trends, offering data support for improving crop stability and breeding efficiency, and providing a basis for early intervention to prevent negative effects of mutation accumulation (such as decreased disease resistance and weakened adaptability). Analyzing gene structural variation based on mutation frequency growth data can reveal large-scale structural changes in the genome, such as gene rearrangements, insertions, and deletions. Gene structural variation has a significant impact on crop phenotypes, developmental processes, and resistance. By analyzing these variations, researchers can gain a deeper understanding of gene structural characteristics and how their changes lead to crop trait expression, providing fundamental data for gene function research, crop improvement, and precision breeding. Furthermore, combining gene structural variation data and mutation frequency growth data to predict transcription factor abnormalities can help researchers identify which transcription factors exhibit functional abnormalities due to gene structural variations. Transcription factors play a crucial role in gene expression regulation. Abnormal transcription factor function leads to disordered gene expression, thereby affecting crop growth, development, and stress resistance. By predicting transcription factor abnormalities, researchers can effectively identify gene regulatory dysregulations that cause crop loss of function or poor traits, providing a basis for the repair and optimization of crop gene regulatory networks. Detecting genome integrity decay by analyzing crop gene deletion abnormalities and gene structural variation data can assess whether there are significant deletions or damages in the crop genome. These issues directly affect crop growth, reproduction, and stress resistance. Genome integrity is the foundation of healthy crop growth; any structural damage leads to decreased crop productivity. Analyzing this data allows for the timely identification of regions in the genome that cause loss of function, thus supporting the improvement of crop genome stability. Furthermore, combining genome integrity decay data with gene transcription factor abnormality data for gene regulatory network disruption detection can reveal potential dysregulated regions within the gene regulatory network. The gene regulatory network is the core of gene expression and metabolic regulation; any disruption leads to abnormal crop development or decreased stress resistance. By detecting disruptions to this network, researchers can pinpoint key regions in the genome that influence crop performance, providing a theoretical basis for improving crop gene expression regulatory networks. Finally, by analyzing gene expression abnormalities using data on gene regulatory network disruption and genome integrity decay, a comprehensive assessment of gene expression problems caused by structural damage or regulatory dysregulation can be achieved. Dysregulation of gene expression not only affects crop development and adaptability but also weakens important traits such as disease resistance and stress tolerance.

[0052] Preferably, step S3 includes the following steps:

[0053] Step S31: Based on the gene expression disorder data, predict the decline of crop disease resistance to obtain crop disease resistance decline data;

[0054] Step S32: Calculate the crop growth rate decline based on the gene expression disorder data to obtain the crop growth rate decline data;

[0055] Step S33: Analyze trait changes based on crop growth rate decline data and crop disease resistance decline data to obtain gene expression trait change data;

[0056] Step S34: Store the gene expression trait change data in a pre-set gene storage database to obtain a crop trait change storage database.

[0057] This invention predicts crop disease resistance decline based on gene expression disorder data, accurately assessing the degree of reduction in crop disease resistance caused by gene expression disturbances. This step, by quantifying the impact of gene expression disorders on the crop immune system, provides crucial data support for improving crop disease resistance, helping researchers identify and repair key genes leading to resistance decline. Through this prediction, researchers can identify the risk of reduced disease resistance in advance, allowing for appropriate interventions to ensure crop survival and reproductive capacity in future environments. Furthermore, calculating crop growth rate decline based on gene expression disorder data helps understand how abnormal gene expression affects crop growth and development. Crop growth rate is a core indicator of crop productivity, and its decline is related to factors such as gene expression disorders and limited nutrient absorption. This calculation effectively reveals which gene expression disorders affect crop growth rate, providing a theoretical basis for improving crop growth performance and growth rate. Analyzing trait changes based on crop growth rate decline and disease resistance decline data allows for a comprehensive assessment of trait changes caused by gene expression disorders. This analysis reveals comprehensive trait changes in crops, including resistance, adaptability, and yield, providing strong guidance for crop genetic improvement. Through a comprehensive analysis of trait changes, researchers can identify which gene variations or expression abnormalities affect the overall performance of crops, thus providing a scientific basis for precision breeding and performance enhancement. Finally, storing gene expression trait change data in a gene storage database ensures efficient archiving and convenient retrieval of this data. The establishment of the database not only facilitates the long-term storage and management of existing data but also provides a convenient infrastructure for future data sharing and genotype-phenotype association analysis. Through centralized storage of trait change data, researchers can quickly access, integrate, and analyze data from different crop species or under different environmental conditions.

[0058] Preferably, step S4 includes the following steps:

[0059] Step S41: Perform differential expression analysis based on the crop trait change storage database to obtain crop differential expression data;

[0060] Step S42: Detect crop gene anomalies based on differential expression data to obtain crop gene anomaly data;

[0061] Step S43: Annotate the crop change storage database based on differential expression data and crop gene abnormality data to obtain crop change annotation data;

[0062] Step S44: Improve crop genetic quality based on crop change annotation data to obtain crop genetic quality improvement data.

[0063] This invention utilizes differential expression analysis of a crop trait variation storage database to identify gene expression differences among different samples, environmental conditions, or crop varieties. This analysis not only helps researchers discover expression changes in crops at specific growth stages or under specific stresses but also reveals which genes play key roles in different phenotypic expressions. Detecting crop gene anomalies based on differential expression data can effectively screen and identify gene abnormalities leading to decreased crop performance. This process helps identify gene variations or functional deficiencies associated with crop disease susceptibility, low yield, and weakened resistance, providing crucial information for subsequent gene repair, gene function analysis, and crop improvement. Annotating the crop variation storage database with differential expression data and crop gene anomaly data allows for a more systematic association of these data with known gene functions and traits. This annotation process not only helps researchers understand how gene anomalies affect crop traits but also provides more background information for better understanding and utilization of the data. Crop gene quality improvement based on crop variation annotation data allows for targeted crop gene modification based on the results of gene anomaly and differential expression analysis. The core objective of this step is to enhance important traits in crops, such as resistance, adaptability, growth rate, and final yield, by identifying key gene regulatory networks.

[0064] The present invention also provides a gene sequencing data processing system for performing the gene sequencing data processing method described above, the gene sequencing data processing system comprising:

[0065] Gene sequencing module: acquires crop sample dataset; grinds crop sample cells based on crop sample dataset to obtain crop sample cell fragmentation data; performs cell gene sequencing on crop sample cell fragmentation data to obtain crop cell gene sequence data;

[0066] Gene expression disorder analysis module: Performs gene sequence variation analysis on crop gene sequence data to obtain crop gene sequence variation data; classifies gene sequence variations based on crop gene sequence variation data to obtain gene mutation sequence data and gene sequence deletion data; performs gene expression disorder analysis based on gene mutation sequence data and gene sequence deletion data to obtain gene expression disorder data.

[0067] Database storage module: Based on gene expression disorder data, predict the decline of crop disease resistance to obtain crop disease resistance decline data; perform trait change analysis on crop disease resistance decline data and gene expression disorder data to obtain gene expression trait change data; use a preset gene storage database to store the gene expression trait change data to obtain a crop trait change storage database.

[0068] Gene quality improvement module: Differential expression analysis is performed based on the crop trait change storage database to obtain crop differential expression data; crop gene anomaly detection is performed based on the crop differential expression data to obtain crop gene anomaly data; crop gene quality improvement is performed based on the crop gene anomaly data to obtain crop gene quality improvement data.

[0069] This invention offers several advantages. By acquiring crop sample datasets and disrupting cells, combined with cell gene sequencing technology, it can comprehensively obtain crop gene sequence data, providing accurate foundational data for subsequent gene variation analysis and functional studies. Cell gene sequencing, through efficient whole-sequence analysis of the crop genome, reveals gene diversity and variation, laying the foundation for further gene function analysis and trait research. Variation analysis of gene sequence data can identify mutations and deletions in the genome, which directly affect crop traits such as resistance, stress tolerance, and growth and development. Classifying gene sequence variations and distinguishing mutation and deletion types helps determine which gene variations are associated with target traits. Furthermore, gene expression disorder analysis based on this variation data can further screen affected genes and predict their abnormalities in biological processes, helping to identify key pathogenic or harmful gene variations. Analyzing gene expression disorder data to predict crop disease resistance decline not only provides a basis for crop health assessment but also identifies which genes or variations lead to weakened crop resistance under specific environmental conditions or pathogen infection, thus providing data support for directions to enhance disease resistance. Simultaneously, combining gene expression disorder data with trait change analysis helps understand the association between genes and crop trait changes, identify key trait genes, and promote precision breeding. Storing this gene expression trait change data in a gene storage database not only enables efficient data management but also provides a systematic data reference for future crop trait improvement. The establishment of this database facilitates continuous updates and provides a foundation for genomics research on different crops, supporting the implementation of gene-phenotype association analysis under different environments. Further differential expression analysis can identify gene expression differences under different conditions at the whole-genome level, helping to understand the response and regulatory mechanisms of crop genes under different environments, treatments, or mutations. Detecting crop gene anomalies based on differential expression data can reveal abnormalities in gene expression, helping to identify potential mutations or variant regions, especially key genes affecting important crop traits (such as yield, stress resistance, and nutrient composition). Combining the results of gene anomaly detection, improvements in crop gene quality can be achieved through precise gene editing or breeding techniques to optimize key genes, thereby improving overall crop quality and adaptability, and enhancing agricultural production efficiency. Therefore, this invention optimizes a traditional gene sequencing data processing method, addressing the problems of inaccurate prediction of gene mutation direction and gene influence on trait changes. It improves the accuracy of predicting crop gene mutation direction and gene influence on trait changes. Attached Figure Description

[0070] Figure 1This is a flowchart illustrating the steps of a gene sequencing data processing method.

[0071] Figure 2 for Figure 1 A detailed flowchart illustrating the implementation steps of step S2.

[0072] Figure 3 for Figure 1 A detailed flowchart illustrating the implementation steps of step S3.

[0073] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0074] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0075] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.

[0076] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0077] To achieve the above objectives, please refer to Figures 1 to 3 A gene sequencing data processing method includes the following steps:

[0078] Step S1: Obtain crop sample dataset; grind crop sample cells according to the crop sample dataset to obtain crop sample cell fragmentation data; perform cell gene sequencing on the crop sample cell fragmentation data to obtain crop cell gene sequence data;

[0079] Step S2: Perform gene sequence variation analysis on crop gene sequence data to obtain crop gene sequence variation data; classify gene sequence variations based on crop gene sequence variation data to obtain gene mutation sequence data and gene sequence deletion data; perform gene expression disorder analysis based on gene mutation sequence data and gene sequence deletion data to obtain gene expression disorder data.

[0080] Step S3: Based on the gene expression disorder data, predict the decline of crop disease resistance to obtain crop disease resistance decline data; perform trait change analysis on the crop disease resistance decline data and gene expression disorder data to obtain gene expression trait change data; use a preset gene storage database to store the gene expression trait change data to obtain a crop trait change storage database.

[0081] Step S4: Perform differential expression analysis based on the crop trait change storage database to obtain crop differential expression data; perform crop gene anomaly detection based on the crop differential expression data to obtain crop gene anomaly data; perform crop gene quality improvement based on the crop gene anomaly data to obtain crop gene quality improvement data.

[0082] In this embodiment of the invention, reference Figure 1 The above is a schematic diagram of the steps of a gene sequencing data processing method according to the present invention. In this example, the gene sequencing data processing method includes the following steps:

[0083] Step S1: Obtain crop sample dataset; grind crop sample cells according to the crop sample dataset to obtain crop sample cell fragmentation data; perform cell gene sequencing on the crop sample cell fragmentation data to obtain crop cell gene sequence data;

[0084] In this embodiment of the invention, a crop sample dataset is collected, including crop samples from different growth stages and varieties. The collected samples undergo rigorous screening to ensure representativeness and freedom from contamination. Subsequently, cell disruption is performed using mechanical or chemical methods, typically liquid nitrogen grinding or an ultrasonic cell disruptor, to break down the cell membranes and release intracellular nucleic acids, proteins, and cellular components. The disrupted samples are then centrifuged to remove unwanted impurities, yielding cell disruption data. Next, high-throughput gene sequencing technologies, such as the Illumina or PacBio platform, are used to sequence the genes in the disrupted cell samples, obtaining cellular gene sequence data from the crops. Gene sequencing involves constructing libraries from the extracted DNA and determining gene sequences through DNA synthesis reactions, thereby obtaining high-quality gene sequence data.

[0085] Step S2: Perform gene sequence variation analysis on crop gene sequence data to obtain crop gene sequence variation data; classify gene sequence variations based on crop gene sequence variation data to obtain gene mutation sequence data and gene sequence deletion data; perform gene expression disorder analysis based on gene mutation sequence data and gene sequence deletion data to obtain gene expression disorder data.

[0086] In this embodiment of the invention, after acquiring gene sequence data, alignment tools such as BWA (Burrows-WheelerAligner) or Bowtie are used to align the original sequence with a reference genome to identify gene sequence variations. For each gene location, GATK (Genome Analysis Toolkit) is used to perform SNP (single nucleotide polymorphism) and InDel (insertion / deletion) detection, thereby obtaining gene sequence variation data for crops. Based on this variation information, gene sequence variations are divided into two categories: gene mutation sequences and gene sequence deletion data. Gene mutation sequences are further confirmed to determine whether they affect functional genes by comparing the detected variations. For deletion data, de novo assembly methods (such as SPAdes or Trinity) are used to reassemble misaligned gene fragments to identify gene deletion regions. Finally, combining mutation sequences and deletion data, differential analysis tools such as DESeq2 are used to analyze changes in gene expression to obtain gene expression disorder data.

[0087] Step S3: Based on the gene expression disorder data, predict the decline of crop disease resistance to obtain crop disease resistance decline data; perform trait change analysis on the crop disease resistance decline data and gene expression disorder data to obtain gene expression trait change data; use a preset gene storage database to store the gene expression trait change data to obtain a crop trait change storage database.

[0088] In this embodiment of the invention, based on gene expression disorder data, data analysis platforms such as "edgeR" or "DESeq2" in R language are used to predict disease resistance decline by combining gene expression levels, mutation and deletion information. By comparing known gene regions related to resistance in the crop genome, bioinformatics methods such as functional annotation and pathway enrichment analysis are used to predict which gene mutations lead to reduced disease resistance. For example, if genes involved in insect resistance, disease resistance, etc., mutate or have decreased expression levels, it will lead to resistance decline. Next, based on the disease resistance decline data and gene expression disorder data, crop trait change analysis is performed. Through multi-dimensional statistical analysis (such as principal component analysis (PCA) or cluster analysis), data on crop trait changes under resistance decline and gene expression disorder are obtained. Finally, these data are stored in a gene storage database, such as MySQL or MongoDB, for unified management and storage, constructing a crop trait change storage database.

[0089] Step S4: Perform differential expression analysis based on the crop trait change storage database to obtain crop differential expression data; perform crop gene anomaly detection based on the crop differential expression data to obtain crop gene anomaly data; perform crop gene quality improvement based on the crop gene anomaly data to obtain crop gene quality improvement data.

[0090] In this embodiment of the invention, a data mining technique (such as differential expression analysis and regression analysis) is used to perform differential expression analysis on the stored trait change data by querying a crop trait change storage database. Differential expression tools such as "limma" are used to compare gene expression differences between different crop samples to identify genes that show significant differences under different traits. Based on these differential expression data, gene anomaly detection methods, such as SNP- and InDel-based genome scanning or gene function annotation, are used to identify whether genes have abnormalities or mutations. This process helps to detect pathogenic gene variations that lead to differences in crop performance. Finally, combining the differential expression data and gene anomaly data, systems biology tools (such as Cytoscape) are used to annotate gene-trait relationships to determine which gene mutations or deletions lead to changes in crop traits. Finally, using gene improvement algorithms, such as gene editing technology (CRISPR-Cas9) or traditional breeding methods, suitable candidate genes are selected for genetic improvement of crops using the obtained crop gene quality improvement data to optimize crop disease resistance, yield, and quality.

[0091] Preferably, step S1 includes the following steps:

[0092] Step S11: Obtain crop sample dataset;

[0093] Step S12: Grind crop sample cells according to the crop sample dataset to obtain crop sample cell fragmentation data;

[0094] Step S13: Based on the data of cell disruption in the crop sample, centrifuge and filter impurities to obtain the sample cell impurity filtration data;

[0095] Step S14: Perform cell gene sequencing on the sample cell impurity filtering data to obtain crop cell gene sequence data.

[0096] In this embodiment of the invention, during the acquisition of crop sample datasets, it is necessary to collect crop samples from different regions, different growth stages, and different environmental conditions to ensure sample diversity and representativeness. Appropriate sampling tools, such as surgical scissors and sampling forceps, are used to extract parts of the crops, such as leaves, stems, or roots, from the farmland. The collected samples are then numbered and grouped according to set conditions to ensure that each group of samples originates from crops grown under the same environmental and growth conditions. Based on the acquired crop sample dataset, cell disruption techniques are used to extract cellular components from the samples. Liquid nitrogen grinding is typically used to disrupt the cells of the collected crop samples. Liquid nitrogen grinding freezes the crop samples at extremely low temperatures, making the cell walls brittle, thereby breaking the cell membrane and releasing intracellular components such as nucleic acids and proteins. Ultrasonic disruption uses the mechanical vibration of high-frequency sound waves to disrupt the cell walls, releasing the cell contents into a solution. Regardless of the method used, the disrupted samples should be centrifuged to remove larger particles and obtain cell disruption data. During cell disruption, temperature and time must be controlled to avoid degradation or denaturation of nucleic acids, proteins, etc., in the samples. After cell lysis, the sample contains various impurities, such as cell membrane fragments, incompletely ruptured cells, and organelles, which interfere with subsequent gene sequencing. Therefore, centrifugation and impurity filtration are necessary. By setting appropriate centrifugation conditions, such as a speed of 3000 rpm and a centrifugation time of 5 minutes, larger cell fragments and impurities precipitate at the bottom of the tube, while smaller cell contents (including nucleic acids) remain in the supernatant. An Eppendorf centrifuge is used with 15 mL tubes to ensure the sample remains uncontaminated during the operation. After centrifugation, the supernatant is aspirated to obtain the cell impurity filtration data. The centrifuged and filtered sample then proceeds to the cell gene sequencing stage, from which high-quality DNA is extracted. Common DNA extraction kits (such as Qiagen DNeasy or Thermo Fisher GeneJET) can be used to extract DNA from the cells. During extraction, the kit's instruction manual must be strictly followed to ensure DNA integrity and purity. After extraction, the DNA is quantified and its quality is assessed using NanoDrop to measure its concentration and purity. Finally, high-throughput gene sequencing technology is used to sequence the extracted DNA. Commonly used sequencing platforms include the Illumina NovaSeq or HiSeq series, which offer extremely high sequencing depth and accuracy, suitable for high-precision crop genome analysis. DNA sequencing yields crop cell gene sequence data.

[0097] Preferably, step S14 includes the following steps:

[0098] Step S141: Extract nucleic acid from the sample cells by filtering out impurities in the sample cells to obtain a crop sample cell nucleic acid dataset;

[0099] Step S142: Perform DNA digestion based on the crop sample cell nucleic acid dataset to obtain crop cell DNA fragment data;

[0100] Step S143: Amplify the DNA fragment data of crop cells to obtain amplified data of crop cell DNA fragments;

[0101] Step S144: Gene sequencing is performed based on the amplified DNA fragment data of crop cells to obtain crop cell gene sequence data.

[0102] In this embodiment of the invention, during the extraction of nucleic acid from filtered cell samples, nucleic acid is extracted from the filtered samples to ensure the acquisition of pure DNA or RNA samples. This process uses standardized cell nucleic acid extraction methods; commonly used extraction kits such as Qiagen DNeasy or Thermo Fisher GeneJET can effectively extract high-quality DNA. Depending on specific needs, an appropriate extraction method is selected (e.g., for samples with thick plant cell walls, CTAB or PVP buffers are used to improve nucleic acid extraction efficiency). During extraction, the temperature and time of the sample need to be controlled to prevent DNA degradation. The extracted nucleic acid is measured for concentration and purity using Nanodrop to ensure the sample is free of impurities and confirm that the DNA quality meets the requirements for gene sequencing. After obtaining the crop sample cell nucleic acid dataset, DNA digestion is performed. Appropriate restriction endonucleases (e.g., EcoRI, BamHI, HindIII, etc.) are selected to digest the DNA according to the characteristics of the crop genome. These enzymes can accurately recognize specific DNA sequences and cut DNA fragments with the same restriction sequences. During DNA digestion, the digestion reaction system must include appropriate buffers and reaction temperatures to ensure maximum enzyme activity. Common reaction conditions include incubation at 37°C for 30 minutes to 1 hour. After enzyme digestion, fragment size is separated by agarose gel electrophoresis to further verify the digestion effect and DNA fragment integrity. Electrophoresis detection confirms the integrity and size range of the DNA fragments, ultimately obtaining crop cell DNA fragment data. The digested crop cell DNA fragment data is then amplified. Polymerase chain reaction (PCR) technology is used to amplify the target DNA fragment. The PCR reaction system includes high-fidelity DNA polymerase, primers (forward and reverse primers), dNTPs, buffer, and sample DNA. Specific primers are designed based on the characteristics of the DNA fragment to be amplified, and PCR conditions are optimized to ensure the best amplification efficiency. The amplification program typically includes a pre-denaturation step (95°C for 5 minutes), denaturation (95°C for 30 seconds), annealing (generally 55-65°C for 30 seconds depending on the primer melting temperature), and extension (72°C for 1 minute), repeated 30-40 times before end extension (72°C for 10 minutes). After amplification, the size and concentration of the PCR products were detected using agarose gel electrophoresis. If the amplified DNA fragments met the requirements, crop cell DNA fragment amplification data were obtained. After DNA fragment amplification, the gene sequencing stage commenced. High-throughput gene sequencing platforms such as Illumina NovaSeq or PacBio were used to sequence the amplified DNA fragments. Before sequencing, the PCR products needed to be purified to remove unreacted primers, dNTPs, and other impurities to ensure sequencing quality.The sequencing process employs appropriate library construction methods, ligating DNA fragments into aptamers for library amplification and purification. After library construction, it is loaded onto a gene sequencing platform for high-throughput sequencing, typically using short-read (Illumina) or long-read (PacBio) technologies to obtain complete gene sequence data. During sequencing, sufficient sequencing depth is ensured for each DNA fragment to guarantee genome coverage and accuracy. After quality control (e.g., removal of low-quality reads, adapter contamination), the sequencing data yields crop cell gene sequence data.

[0103] Preferably, step S1 includes the following steps:

[0104] Step S21: Perform gene sequence variation analysis on crop gene sequence data to obtain crop gene sequence variation data;

[0105] Step S22: Based on the crop gene sequence variation data, divide the gene sequence variation to obtain gene mutation sequence data and gene sequence missing data;

[0106] Step S23: Predict genome stability decay based on gene mutation sequence data to obtain genome stability decay data;

[0107] Step S24: Perform gene deletion anomaly analysis based on gene sequence deletion data to obtain crop gene deletion anomaly data;

[0108] Step S25: Perform gene expression disorder analysis on the genomic stability decay data and crop gene deletion abnormality data to obtain gene expression disorder data.

[0109] As an example of the present invention, reference is made to Figure 2 As shown, in this example, step S2 includes:

[0110] Step S21: Perform gene sequence variation analysis on crop gene sequence data to obtain crop gene sequence variation data;

[0111] In this embodiment of the invention, the genome sequence is aligned and labeled using standard gene alignment tools such as BWA or Bowtie2 to ensure accurate alignment of the sequencing data. Gene sequence variation analysis employs software tools such as GATK (Genome Analysis Toolkit) or Samtools to detect variations in the aligned data, identifying single nucleotide polymorphisms (SNPs) and small insertion / deletion (InDels) variations. During variation detection, appropriate quality thresholds and filtering criteria are set, such as minimum variation quality values ​​and minimum supported read counts, to ensure the reliability and accuracy of variation detection. Through statistical analysis of the variation information, the final crop gene sequence variation data is obtained.

[0112] Step S22: Based on the crop gene sequence variation data, divide the gene sequence variation to obtain gene mutation sequence data and gene sequence missing data;

[0113] In this embodiment of the invention, gene sequence variation information is classified. Existing databases such as dbSNP are used to annotate SNP variations, identifying common variations and potentially pathogenic mutations. For each mutation site, it is further classified according to its gene region, into non-coding region variations and coding region variations. For coding region mutations, they are further distinguished into synonymous mutations, missense mutations, and nonsense mutations. In addition, based on the missing parts of the gene sequence, tools such as BEDTools or VCFtools are used to analyze the variation data, marking the missing regions in the gene sequence and determining whether these regions affect the coding of functional genes. Finally, gene mutation sequence data and gene sequence missing data are obtained through the above classification and analysis.

[0114] Step S23: Predict genome stability decay based on gene mutation sequence data to obtain genome stability decay data;

[0115] In this embodiment of the invention, specialized genome stability analysis tools, such as StabilitySeq or Genomic Instability Scanner (GIS), are used. These tools can combine known genome stability indicators, such as mutation accumulation rate, gene rearrangement frequency, and chromosome number, to perform in-depth analysis of mutation sequence data. During the prediction process, the impact of mutations on gene function is assessed based on the location and type of the variant, thereby analyzing whether the mutation leads to structural instability of the genome. By comparing normal and mutated genome data and combining bioinformatics algorithms to calculate the trend of genome stability decay, genome stability decay data is finally obtained.

[0116] Step S24: Perform gene deletion anomaly analysis based on gene sequence deletion data to obtain crop gene deletion anomaly data;

[0117] In this embodiment of the invention, VCF format data is used to compare missing gene regions with a reference genome to identify whether the missing regions affect the presence of specific functional genes. For each missing region of the crop genome, functional analysis is further performed using gene annotation databases (such as Ensembl, NCBI, etc.) to determine whether the missing genes are related to important physiological characteristics of crops (such as disease resistance, stress resistance, nutrient accumulation, etc.). Based on this, a genome browser (such as IGV) is used to visualize the missing regions to check for large-scale chromosomal deletions or the loss of important genes. Through these analyses, abnormal data on crop gene deletions are finally obtained.

[0118] Step S25: Perform gene expression disorder analysis on the genomic stability decay data and crop gene deletion abnormality data to obtain gene expression disorder data.

[0119] In this embodiment of the invention, genomic stability decay data are combined with gene deletion anomaly data to assess the potential impact of mutations and deletions on gene expression using bioinformatics analysis methods. Transcriptome data (RNA-Seq) are used to compare gene expression patterns in different treatment groups, and statistical analysis methods such as DESeq2 or EdgeR are applied to identify genes whose expression levels significantly change under genomic stability decay or gene deletion anomalies. For genes with anomalous expression, further gene function enrichment analysis is performed using the GO (Gene Ontology) and KEGG (Kyoto Encyclopedia of Genes and Genomes) databases to analyze the functions of these genes in plant resistance, growth, and development. Finally, data on gene expression anomalies are obtained through these analyses.

[0120] Preferably, step S23 includes the following steps:

[0121] Step S231: Perform gene mutation accumulation detection based on gene mutation sequence data to obtain gene mutation accumulation data;

[0122] Step S232: Perform chromosome structural break detection on the accumulated gene mutation data to obtain chromosome structural break data;

[0123] Step S233: Based on the accumulated gene mutation data, predict the probability of chromosome rearrangement to obtain chromosome rearrangement probability data;

[0124] Step S234: Calculate the gene duplication probability based on the chromosome rearrangement probability data to obtain the gene duplication probability data;

[0125] Step S235: Perform genome stability decay prediction on chromosome structural breakage data and gene duplication probability data to obtain genome stability decay data.

[0126] In this embodiment of the invention, gene sequence data is aligned with a reference genome using genome alignment tools such as BWA or Bowtie2. Mutation sites in the samples are identified through alignment, including single nucleotide mutations (SNPs), insertions, and deletions (InDels). Next, the mutation frequency and type of each gene region are statistically analyzed, with particular attention paid to high-frequency mutation regions. For the accumulation of mutation frequencies, software such as VCFtools is used to perform statistical analysis on the VCF-formatted mutation data, calculating the cumulative mutation density to identify the increasing trend of mutation frequency over time or under different conditions in different samples. This analysis yields gene mutation accumulation data, which, combined with chromosome mapping information, is analyzed across the entire genome. Software such as BreakDancer or Delly is used to detect chromosomal structural rearrangements in mutation accumulation regions, identifying chromosomal breaks or rearrangements in the genome. Specifically, these tools can detect chromosomal deletions, inversions, translocations, and other structural variations in the genome based on mutation accumulation. By comparing the locations of genomic breaks and mutations, it is further determined which mutations lead to genomic structural instability, thus obtaining chromosomal structural break data. Statistical models and algorithms are applied to predict the probability of chromosomal rearrangements by comparing gene mutation accumulation data with known chromosomal rearrangement patterns. Tools such as MosaicPlot or CIRCOS are used to visualize chromosomal structures to observe whether regions with high mutation frequencies in the genome have a high risk of rearrangement. Based on existing genomic databases and mutation data, chromosomal rearrangement events in the genome are predicted using model algorithms, and machine learning algorithms or probabilistic analysis methods are used to further optimize the prediction accuracy. This step obtains chromosomal rearrangement probability data, thereby identifying chromosomal regions prone to structural rearrangement during mutation accumulation. Using previously generated chromosomal rearrangement data, genes involved in rearrangement regions are screened, with particular attention paid to genes located in rearrangement hotspots. These genes undergo copy number changes due to chromosomal rearrangements. Next, Copy Number Variation (CNV) analysis tools such as GISTIC or Control-FREEC are used to analyze the duplication of each gene in different individuals using genomic data and calculate the duplication probability of each gene. By analyzing the frequency of gene duplication in different samples, it is determined which genes have a high duplication probability in chromosomal rearrangement regions. The results of the first two steps are combined to assess the relationship between gene mutation accumulation, chromosome breakage, and gene duplication. A genome stability decay prediction model was established by utilizing a specialized stability decay prediction algorithm, combining the correlation between chromosome breakpoints, gene repetition frequencies, and mutation locations. Bioinformatics tools, such as StabilitySeq or Instability-SCAN, were used for data integration and prediction.These tools calculate the mutation frequency, structural rearrangement, and duplication probability of different regions of the genome to comprehensively deduce the trend of genome stability decay. Finally, based on these data, statistical analysis is performed to predict the probability of genome stability decay under different environmental or selective pressures, resulting in the final genome stability decay data.

[0127] Preferably, step S24 includes the following steps:

[0128] Step S241: Based on the gene sequence deletion data, predict the loss of protein function to obtain protein function loss data;

[0129] Step S242: Collect protease gene fragment deletion data from the gene sequence deletion data to obtain protease gene fragment deletion data;

[0130] Step S243: Based on the data of missing gene fragments of white enzyme, perform protease inactivation detection to obtain data on protease inactivation in crops;

[0131] Step S244: Perform crop resistance weakening tests on crop protease inactivation data and protein function loss data to obtain crop resistance weakening data;

[0132] Step S245: Based on the data on crop resistance reduction and crop protease inactivation, predict the decline in crop environmental adaptability and obtain the data on the decline in crop environmental adaptability.

[0133] Step S246: Perform gene deletion anomaly analysis based on crop environmental adaptability decline data and crop resistance weakening data to obtain crop gene deletion anomaly data.

[0134] In this embodiment of the invention, gene deletion data is used to identify missing regions in genes, with particular attention paid to deletion sites that affect protein function. Protein structure prediction software, such as SWISS-MODEL or Phyre2, is used in conjunction with reference genome and functional gene information to analyze the impact of deleted sequences on protein structure and function. By comparing the relationship between deleted sequences in the genome and known functional regions, it is identified which gene deletions lead to protein function loss. Then, by using reference databases such as InterPro or Pfam, the impact of gene deletions on protein functional domains is further predicted, thereby obtaining predicted protein function loss data. Genes containing protease functions are screened, and which protease gene regions are missing are identified based on gene deletion data. By comparing known protease gene databases (such as MEROPS) with specific crop genome information, it is determined which protease genes or gene fragments have been deleted. Genomic analysis tools, such as BEDTools or GATK, are used to extract deletion data for specific genes or gene fragments. This deletion data is mapped to specific protease gene regions to generate protease gene fragment deletion data. Using gene deletion data and the known functional structure of proteases, combined with predicted protease structural changes, it is analyzed whether the deleted fragments lead to loss of protease activity. Molecular dynamics simulation tools such as GROMACS or AutoDock were used to simulate the functional residues of proteases and assess the impact of missing fragments on their active sites. By comparing with functional validation data, such as enzyme activity assays or protease inhibitor experiments, it was determined which proteases lost activity due to gene fragment deletion. Combining experimental data and computational results, data on crop protease inactivation were finally obtained. The impact of these changes on crop disease resistance, pest resistance, and other abilities was analyzed by combining protease inactivation data and protein function loss data. Databases of plant immune response-related genes, such as PlantCyc or PANTHER, were used to identify genes related to immune responses and their functions, and the relationship between protease inactivation and immune response capacity was assessed. By comparing genotype and phenotype data, combined with biostatistical methods such as ANOVA or regression analysis, the impact of protein function loss and protease inactivation on crop resistance reduction was assessed. Experimental data, such as disease resistance experiments or pest experiments, were used to validate the predicted results, ultimately obtaining data on crop resistance reduction. Correlation analysis was performed between resistance reduction data and genotype-phenotype data related to environmental adaptation. Using environmental adaptability models (such as environmental adaptability indices or crop growth models) and resistance weakening data, we can analyze the crop's adaptability under different environmental conditions.By combining crop genomic information and ecological environmental factors, simulation algorithms such as LUMO (Land Use and Modeling Optimization) or GIS (Geographic Information System) tools are applied to predict the trend of crop adaptability decline under adverse environmental conditions. By incorporating climate change models, the impact of weakened resistance on environmental adaptability is predicted, thus obtaining data on crop environmental adaptability decline. Combining environmental adaptability decline data with resistance weakening data, the deletion of which genes is related to both resistance weakening and environmental adaptability decline is analyzed. Key genes related to environmental adaptability and resistance are identified using gene function annotation databases such as Gene Ontology (GO) and KEGG. Combining genomic data, bioinformatics methods, such as gene enrichment analysis and pathway analysis, are employed to identify which gene deletions lead to weakened crop adaptability. Gene mutation detection tools such as GATK or Samtools are used to further verify the accuracy of the deletion data, ultimately identifying crop gene deletion anomalies associated with gene deletion abnormalities.

[0135] Preferably, step S25 includes the following steps:

[0136] Step S251: Based on the genomic stability decay data, predict the mutation frequency growth to obtain mutation frequency growth data;

[0137] Step S252: Perform gene structure variation analysis based on mutation frequency growth data to obtain gene structure variation data;

[0138] Step S253: Based on gene structure variation data and mutation frequency growth data, predict transcription factor abnormalities to obtain gene transcription factor abnormality data;

[0139] Step S254: Perform genome integrity decay detection on crop gene deletion abnormality data and gene structure variation data to obtain genome integrity decay data;

[0140] Step S255: Detect gene regulatory network disruption based on genome integrity decay data and gene transcription factor abnormality data to obtain gene regulatory network disruption data;

[0141] Step S256: Perform gene expression disorder analysis on gene regulatory network disruption data and genome integrity decay data to obtain gene expression disorder data.

[0142] In this embodiment of the invention, genomic stability decay data is analyzed to identify which gene regions have undergone more frequent mutations. Frequency information of mutation events is extracted using genomic variation detection tools (such as GATK, Samtools, etc.), and the cumulative frequency of mutations under different time or environmental conditions is calculated. Statistical analysis methods (such as linear regression, time series analysis) are used to predict the growth trend of mutation frequency. Predicted mutation frequency growth data is obtained by assessing the type of mutation site (such as single nucleotide polymorphisms (SNPs), insertions / deletions, etc.) and the frequency changes of mutation occurrence. The mutation frequency growth data is correlated with changes in gene structure. Genomic annotation information is used to determine which genes or gene regions are affected during the mutation frequency growth process. Structural variation detection tools (such as Manta, Delly, etc.) are used for in-depth analysis of mutation data to identify large-scale structural variations in the genome (such as gene rearrangements, inversions, copy number variations, etc.). Combined with the mutation frequency growth data, it is analyzed whether these structural variations are related to changes in gene expression or loss of function. Visualization tools (such as IGV) are used to present the specific locations of gene structural variations, ultimately obtaining gene structural variation data. Transcription factor genes affected during the mutation frequency growth process are identified. Using transcription factor databases (such as TFDB and JASPAR) combined with gene structural variation data, we analyze whether mutated regions cover transcription factor binding sites or regulatory regions. Gene expression prediction tools (such as CAGE and RNA-Seq analysis) are used to analyze whether these variations lead to abnormal expression or functional changes of transcription factors. By comparing transcription factor action mechanisms in existing literature and databases, we predict the impact of increased mutation frequency on normal transcription factor function. Integrating gene deletion aberration data with gene structural variation data, we analyze which gene deletions and structural variations lead to impaired overall genome integrity. Genome integrity detection tools (such as CheckM and QUAST) are used to assess the quality of genome data, identifying which regions' deletions or structural variations cause genome integrity decay. Combined with existing genome annotation information, we further analyze the potential impact of deleted and mutated regions on gene expression. Quantitative analysis methods (such as copy number variation (CNV) analysis and genome alignment) are used to measure the degree of decay, ultimately obtaining genome integrity decay data. Combining this data with gene transcription factor aberration data identifies which genes lose normal regulation during integrity decay. Gene regulatory network analysis tools (such as Cytoscape and GeneMANIA) were used to construct gene regulatory networks, and genes with impaired regulatory functions were identified based on abnormal transcription factor data. Combined with gene expression data (such as RNA-Seq data), the relationships between impaired transcription factors and their downstream target genes were analyzed to assess their impact on the entire gene regulatory network. By comparing normal and abnormal gene regulatory networks, data on gene regulatory network disruption were obtained.By combining data on disruption of gene regulatory networks, this study assesses whether gene dysregulation leads to abnormal gene expression levels. Gene expression analysis tools (such as DESeq2 and EdgeR) are used to analyze gene expression data and identify which genes are abnormally expressed in the context of genome integrity degradation and disruption of gene regulatory networks. Heatmaps and principal component analysis (PCA) are employed to visually visualize changes in gene expression patterns, further identifying key genes with abnormal expression levels. Experimental validation (such as qPCR and Western blot) confirms the changes in gene expression, ultimately yielding data on gene expression dysregulation.

[0143] Preferably, step S3 includes the following steps:

[0144] Step S31: Based on the gene expression disorder data, predict the decline of crop disease resistance to obtain crop disease resistance decline data;

[0145] Step S32: Calculate the crop growth rate decline based on the gene expression disorder data to obtain the crop growth rate decline data;

[0146] Step S33: Analyze trait changes based on crop growth rate decline data and crop disease resistance decline data to obtain gene expression trait change data;

[0147] Step S34: Store the gene expression trait change data in a pre-set gene storage database to obtain a crop trait change storage database.

[0148] As an example of the present invention, reference is made to Figure 3 As shown, step S3 in this example includes:

[0149] Step S31: Based on the gene expression disorder data, predict the decline of crop disease resistance to obtain crop disease resistance decline data;

[0150] In this embodiment of the invention, in-depth analysis of gene expression disorder data is performed to screen out key genes affecting crop disease resistance. Bioinformatics analysis tools (such as Gene Ontology and KEGG pathway) are used to functionally annotate these genes to determine whether they are closely related to disease resistance-related biological processes or pathways (such as immune responses and pathogen recognition). Combined with crop disease resistance phenotypic data (such as disease resistance test results), statistical models (such as multiple regression analysis and decision trees) are used to establish the relationship between gene expression and disease resistance. This model predicts the potential impact of different gene mutations or expression disorders on crop disease resistance based on changes in gene expression levels, ultimately yielding data on crop disease resistance attenuation.

[0151] Step S32: Calculate the crop growth rate decline based on the gene expression disorder data to obtain the crop growth rate decline data;

[0152] In this embodiment of the invention, gene expression dysregulation data are analyzed to identify which gene dysregulation is associated with changes in growth rate. By consulting existing literature and databases (such as the Plant Genome Database, PANTHER, etc.), genes with dysregulation are compared with known genes affecting growth to construct a list of growth-related genes. Then, combined with crop growth data (such as growth curves, root length, leaf area, etc.), mathematical models (such as linear regression models, logistic regression models) are used to calculate the specific impact of gene expression dysregulation on growth rate. By matching growth rate data at different time points with gene expression data, growth rate decay data are obtained.

[0153] Step S33: Analyze trait changes based on crop growth rate decline data and crop disease resistance decline data to obtain gene expression trait change data;

[0154] In this embodiment of the invention, growth rate attenuation data and disease resistance attenuation data are integrated to determine whether these two attenuation factors are correlated and their combined impact on trait changes. Statistical analysis methods (such as ANOVA and principal component analysis) are used to assess the correlation between gene expression, disease resistance, and growth rate, identifying key genes or genetic variations affecting crop trait changes. By comparing trait data (such as plant height, yield, and disease resistance phenotypic data) with attenuation data, the patterns of crop trait changes under different growth and disease resistance attenuation backgrounds are further determined. Finally, through comprehensive analysis, gene expression trait change data are obtained.

[0155] Step S34: Store the gene expression trait change data in a pre-set gene storage database to obtain a crop trait change storage database.

[0156] In this embodiment of the invention, all generated trait change data, including mutations and expression abnormalities of relevant genes, as well as corresponding phenotypic changes, are organized. Based on data storage requirements, a suitable database format (such as SQL, NoSQL, etc.) is selected to organize and store the data. A dedicated database of crop gene trait changes is constructed using a database management system (such as MySQL, MongoDB, etc.). All data is categorized in the database according to key characteristics such as crop variety, gene, trait type, and phenotype, ensuring convenient data retrieval and analysis. Data is accessed and extracted through data query interfaces in the database (such as SQL queries, APIs, etc.) to provide efficient data support for subsequent genotype-phenotype association analysis, crop breeding research, and genetic improvement. Finally, a complete database for storing crop trait changes is obtained.

[0157] Preferably, step S4 includes the following steps:

[0158] Step S41: Perform differential expression analysis based on the crop trait change storage database to obtain crop differential expression data;

[0159] Step S42: Detect crop gene anomalies based on differential expression data to obtain crop gene anomaly data;

[0160] Step S43: Annotate the crop change storage database based on differential expression data and crop gene abnormality data to obtain crop change annotation data;

[0161] Step S44: Improve crop genetic quality based on crop change annotation data to obtain crop genetic quality improvement data.

[0162] In this embodiment of the invention, all trait change data, including gene expression levels, phenotypic changes, and experimental results under different environmental conditions, are extracted from the database. Differential expression analysis tools (such as DESeq2, edgeR, etc.) are used to perform statistical tests on gene expression in the data, identifying genes with significant differences in expression under different conditions. During the analysis, trait change data and gene expression data are combined, with a focus on screening genes closely related to specific trait changes (such as disease resistance, growth rate, etc.). The expression level of each gene is standardized to further determine which genes show significant differences in expression across different treatment groups or time points. False positive rates are controlled using multiple validation corrections (such as the Benjamini-Hochberg method), ultimately obtaining differential expression data for crops. Genes exhibiting abnormal expression are screened from the differential expression data, especially those with significant expression changes under multiple experimental conditions. Gene anomaly detection methods (such as SNP calling, Indel detection) are used to analyze gene mutations or deletions, checking for functional or structural variations in these genes. Further application of genomic data analysis tools (such as GATK and Samtools) to align genome sequences identifies abnormalities such as mutations, copy number variations (CNVs), or deletions compared to normal genotypes. The results of abnormal gene detection are compared with known gene function data to identify genes affecting crop phenotypes or traits. Ultimately, crop gene anomaly data is obtained, and this data is combined with differential expression data to provide detailed functional annotation for each gene. Using gene function annotation tools (such as BLAST and InterProScan), genes associated with known functions are identified by comparing genome or protein sequences, particularly those related to crop traits such as resistance, nutrient accumulation, and disease resistance. Annotation information (such as gene name, function, and relationship to traits) is appended to each relevant gene in the database, forming detailed crop gene annotation data. By establishing gene function annotation mapping rules, changes in different traits are correlated with gene expression and mutation status, ultimately yielding complete crop variation annotation data. Genes significantly associated with target traits (such as disease resistance, stress resistance, and yield) are screened from variation annotation data. These key genes are then targeted for editing using modern gene editing technologies (such as CRISPR-Cas9 and TALENs). Gene editing techniques are used to introduce gene variations with advantageous traits into the crop genome or to repair gene function loss caused by mutations. In practice, suitable crop varieties are selected to ensure the stable delivery and expression of the edited genes. The edited genes are then transferred to offspring through plant transformation technology, and the gene improvement effect is ultimately verified through field trials.If the improved variety exhibits the expected traits (such as improved disease resistance and increased yield), the genetically modified variety can be further promoted. Ultimately, data on crop genetic quality improvement can be obtained.

[0163] The present invention also provides a gene sequencing data processing system for performing the gene sequencing data processing method described above, the gene sequencing data processing system comprising:

[0164] Gene sequencing module: acquires crop sample dataset; grinds crop sample cells based on crop sample dataset to obtain crop sample cell fragmentation data; performs cell gene sequencing on crop sample cell fragmentation data to obtain crop cell gene sequence data;

[0165] Gene expression disorder analysis module: Performs gene sequence variation analysis on crop gene sequence data to obtain crop gene sequence variation data; classifies gene sequence variations based on crop gene sequence variation data to obtain gene mutation sequence data and gene sequence deletion data; performs gene expression disorder analysis based on gene mutation sequence data and gene sequence deletion data to obtain gene expression disorder data.

[0166] Database storage module: Based on gene expression disorder data, predict the decline of crop disease resistance to obtain crop disease resistance decline data; perform trait change analysis on crop disease resistance decline data and gene expression disorder data to obtain gene expression trait change data; use a preset gene storage database to store the gene expression trait change data to obtain a crop trait change storage database.

[0167] Gene quality improvement module: Differential expression analysis is performed based on the crop trait change storage database to obtain crop differential expression data; crop gene anomaly detection is performed based on the crop differential expression data to obtain crop gene anomaly data; crop gene quality improvement is performed based on the crop gene anomaly data to obtain crop gene quality improvement data.

[0168] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A method for processing gene sequencing data, characterized in that, Includes the following steps: Step S1: Obtain crop sample dataset; grind crop sample cells according to the crop sample dataset to obtain crop sample cell fragmentation data; Cell gene sequencing was performed on the cell fragmentation data of crop samples to obtain crop cell gene sequence data; Step S2: Perform gene sequence variation analysis on crop gene sequence data to obtain crop gene sequence variation data; classify gene sequence variations based on crop gene sequence variation data to obtain gene mutation sequence data and gene sequence deletion data; perform gene expression disorder analysis based on gene mutation sequence data and gene sequence deletion data to obtain gene expression disorder data. Step S2 includes the following steps: Step S21: Perform gene sequence variation analysis on crop gene sequence data to obtain crop gene sequence variation data; Step S22: Based on the crop gene sequence variation data, divide the gene sequence variation to obtain gene mutation sequence data and gene sequence missing data; Step S23: Predict genome stability decay based on gene mutation sequence data to obtain genome stability decay data. Step S23 includes the following steps: Step S231: Perform gene mutation accumulation detection based on gene mutation sequence data to obtain gene mutation accumulation data; Step S232: Perform chromosome structural break detection on the accumulated gene mutation data to obtain chromosome structural break data; Step S233: Based on the accumulated gene mutation data, predict the probability of chromosome rearrangement to obtain chromosome rearrangement probability data; Step S234: Calculate the gene duplication probability based on the chromosome rearrangement probability data to obtain the gene duplication probability data; Step S235: Perform genome stability decay prediction on chromosome structural breakage data and gene duplication probability data to obtain genome stability decay data; Step S24: Perform gene deletion anomaly analysis based on gene sequence deletion data to obtain crop gene deletion anomaly data. Step S24 includes the following steps: Step S241: Based on the gene sequence deletion data, predict the loss of protein function to obtain protein function loss data; Step S242: Collect protease gene fragment deletion data from the gene sequence deletion data to obtain protease gene fragment deletion data; Step S243: Based on the protease gene fragment deletion data, perform protease inactivation detection to obtain crop protease inactivation data; Step S244: Perform crop resistance weakening tests on crop protease inactivation data and protein function loss data to obtain crop resistance weakening data; Step S245: Based on the data on crop resistance reduction and crop protease inactivation, predict the decline in crop environmental adaptability and obtain the data on the decline in crop environmental adaptability. Step S246: Perform gene deletion anomaly analysis based on crop environmental adaptability decline data and crop resistance weakening data to obtain crop gene deletion anomaly data; Step S25: Perform gene expression disorder analysis on the genomic stability decay data and crop gene deletion abnormality data to obtain gene expression disorder data. Step S25 includes the following steps: Step S251: Based on the genomic stability decay data, predict the mutation frequency growth to obtain mutation frequency growth data; Step S252: Perform gene structure variation analysis based on mutation frequency growth data to obtain gene structure variation data; Step S253: Based on gene structure variation data and mutation frequency growth data, predict transcription factor abnormalities to obtain gene transcription factor abnormality data; Step S254: Perform genome integrity decay detection on crop gene deletion abnormality data and gene structure variation data to obtain genome integrity decay data; Step S255: Detect gene regulatory network disruption based on genome integrity decay data and gene transcription factor abnormality data to obtain gene regulatory network disruption data; Step S256: Perform gene expression disorder analysis on gene regulatory network disruption data and genome integrity decay data to obtain gene expression disorder data; Step S3: Predict crop disease resistance attenuation based on gene expression disorder data to obtain crop disease resistance attenuation data; specifically: conduct in-depth analysis of gene expression disorder data to screen out key genes affecting crop disease resistance; use bioinformatics analysis tools to perform functional annotation of these key genes to determine whether they are closely related to disease resistance-related biological processes or pathways; combine crop disease resistance phenotypic data to establish the relationship between gene expression and disease resistance through statistical models; predict the potential impact of different gene mutations or expression disorders on crop disease resistance based on changes in gene expression levels to obtain crop disease resistance attenuation data; perform phenotypic change analysis on crop disease resistance attenuation data and gene expression disorder data to obtain gene expression phenotypic change data; use a pre-set gene storage database to store the gene expression phenotypic change data to obtain a crop phenotypic change storage database; Step S4: Perform differential expression analysis based on the crop trait change storage database to obtain crop differential expression data; perform crop gene anomaly detection based on the crop differential expression data to obtain crop gene anomaly data; perform crop gene quality improvement based on the crop gene anomaly data to obtain crop gene quality improvement data.

2. The gene sequencing data processing method according to claim 1, characterized in that, Step S1 includes the following steps: Step S11: Obtain crop sample dataset; Step S12: Grind crop sample cells according to the crop sample dataset to obtain crop sample cell fragmentation data; Step S13: Based on the cell disruption data of the crop sample, centrifuge and filter impurities to obtain sample cell impurity filtration data; Step S14: Perform cell gene sequencing on the sample cell impurity filtering data to obtain crop cell gene sequence data.

3. The gene sequencing data processing method according to claim 2, characterized in that, Step S14 includes the following steps: Step S141: Extract nucleic acid from the sample cells by filtering out impurities in the sample cells to obtain a crop sample cell nucleic acid dataset; Step S142: Perform DNA digestion based on the crop sample cell nucleic acid dataset to obtain crop cell DNA fragment data; Step S143: Amplify the DNA fragment data of crop cells to obtain amplified data of crop cell DNA fragments; Step S144: Gene sequencing is performed based on the amplified DNA fragment data of crop cells to obtain crop cell gene sequence data.

4. The gene sequencing data processing method according to claim 1, characterized in that, Step S3 includes the following steps: Step S31: Based on the gene expression disorder data, predict the decline of crop disease resistance to obtain crop disease resistance decline data; Step S32: Calculate the crop growth rate decline based on the gene expression disorder data to obtain the crop growth rate decline data; Step S33: Analyze trait changes based on crop growth rate decline data and crop disease resistance decline data to obtain gene expression trait change data; Step S34: Store the gene expression trait change data in a pre-set gene storage database to obtain a crop trait change storage database.

5. The gene sequencing data processing method according to claim 1, characterized in that, Step S4 includes the following steps: Step S41: Perform differential expression analysis based on the crop trait change storage database to obtain crop differential expression data; Step S42: Detect crop gene anomalies based on differential expression data to obtain crop gene anomaly data; Step S43: Annotate the crop change storage database based on differential expression data and crop gene abnormality data to obtain crop change annotation data; Step S44: Improve crop genetic quality based on crop change annotation data to obtain crop genetic quality improvement data.

6. A gene sequencing data processing system, characterized in that, The gene sequencing data processing system is used to perform the gene sequencing data processing method as described in claim 1, and includes: Gene sequencing module: acquires crop sample dataset; grinds crop sample cells based on crop sample dataset to obtain crop sample cell fragmentation data; performs cell gene sequencing on crop sample cell fragmentation data to obtain crop cell gene sequence data; Gene expression disorder analysis module: Performs gene sequence variation analysis on crop gene sequence data to obtain crop gene sequence variation data; classifies gene sequence variations based on crop gene sequence variation data to obtain gene mutation sequence data and gene sequence deletion data; performs gene expression disorder analysis based on gene mutation sequence data and gene sequence deletion data to obtain gene expression disorder data. Database storage module: Based on gene expression disorder data, predict the decline of crop disease resistance to obtain crop disease resistance decline data; perform trait change analysis on crop disease resistance decline data and gene expression disorder data to obtain gene expression trait change data; use a preset gene storage database to store the gene expression trait change data to obtain a crop trait change storage database. Gene quality improvement module: Differential expression analysis is performed based on the crop trait change storage database to obtain crop differential expression data; crop gene anomaly detection is performed based on the crop differential expression data to obtain crop gene anomaly data; crop gene quality improvement is performed based on the crop gene anomaly data to obtain crop gene quality improvement data.

Citation Information

Patent Citations

  • CRISPR-CAS component systems, methods and compositions for sequence manipulation

    CN110982844A

  • Quality evaluation method and screening method of nucleic acid sequencing data

    CN114420214A