Gene detection system

By providing a gene detection system that includes sample collection, gene extraction, sequencing, data analysis and report generation modules, the problem that the existing technology cannot perform gene sequencing and analysis quickly and accurately is solved, efficient and accurate gene detection and data analysis are achieved, and the correlation analysis of the GWAS model is improved.

CN120072045APending Publication Date: 2025-05-30CHERRY VALLEY BREEDING TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510179318.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology cannot quickly and accurately deep sequencing gene samples, obtain high-quality gene sequence information, and convert this information into the original data set, and cannot integrate game theory into the GWAS model of whole genome association analysis to improve the model, and cannot build a genetic-deficient gene association map.

Method used

Provide a gene detection system, including a sample collection module, a gene extraction module, a sequencing module, a data analysis module and a report generation module. The system uses standardized acquisition methods and efficient gene extraction technology to perform deep sequencing using high-throughput sequencing technology, integrates bioinformatics algorithms for data analysis, and integrates game theory into the GWAS model.

Benefits of technology

It realizes rapid and accurate deep sequencing of gene samples, obtains high-quality gene sequence information, and converts this information into the original data set, improves the GWAS model of genome-wide association analysis, and can further explore the correlation between genetic traits and defective genes, and improves the accuracy and reliability of the association model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072045A_ABST
    Figure CN120072045A_ABST
Patent Text Reader

Abstract

The invention discloses a gene detection system which comprises a sample collection module, a gene extraction module, a sequencing module, a data analysis module and a report generation module. The gene extraction module is used for extracting high-quality DNA (Deoxyribonucleic Acid) or RNA (Ribonucleic Acid) from a sample, and the data analysis module is used for processing, analyzing and explaining original data obtained by sequencing by integrating various bioinformatics algorithms and tools. The game theory is integrated into the whole genome association analysis GWAS model to improve the whole genome association analysis GWAS model, the game theory is combined to enable the SNP sites to be more and more detailed, the relevance between deeper hereditary characters and defective genes can be mined, the false positive probability of the result obtained by the association model is smaller, and the accuracy of the result obtained by the association model is improved. A genetic-defect gene association map is constructed according to a whole genome association analysis GWAS model fused with the game theory thought, a defect gene risk prediction model is trained, and the most possible defect gene type is predicted according to genetic characteristics, so that the accuracy is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gene detection, and particularly to a gene detection system. Background Art

[0002] A gene detection system is a technical platform that identifies important biological characteristics such as genetic traits, susceptibility to defective genes, and drug responses by analyzing an individual's genomic information. With the rapid development of genomics, molecular biology, and bioinformatics, gene detection technology has been increasingly widely applied in fields such as medicine, health management, agriculture, and forensic science. The basis of gene detection systems is genomics, which studies the structure, function, variation, and interaction of the genetic information of organisms. The breakthroughs in genomics, especially the completion of the Human Genome Project, have laid the foundation for gene detection technology.

[0003] Currently, the Chinese invention patent with the application number CN202410730267.X discloses a drug-resistant gene detection method and system, which for the first time realizes the detection of multiple drug-resistant genes in a single detection. After screening, the primer sequences and primer dosages are determined. The multiplex amplification primers are tested for reducing the amplification time, and finally the multiplex amplification program is determined. The collected data is analyzed using fragment analysis software. It solves the problem that although the traditional bacterial drug resistance detection method can provide phenotypic information of bacterial drug resistance, it requires long-term cultivation and identification and cannot detect drug-resistant genes. It breaks through the limitations of the existing technology and can detect drug-resistant genes without long-term cultivation and identification, achieving efficient and rapid detection. The existing technology cannot perform deep sequencing of gene samples quickly and accurately, cannot obtain high-quality gene sequence information, and convert this information into a raw data set, cannot integrate game theory into the genome-wide association analysis GWAS model to improve the GWAS model, and cannot construct a genetic-defective gene association map based on the GWAS model integrating the idea of game theory. Summary of the Invention

[0004] The technical problem solved by the present invention is that the existing technology cannot perform deep sequencing of gene samples quickly and accurately, cannot obtain high-quality gene sequence information, and convert this information into a raw data set, cannot integrate game theory into the genome-wide association analysis GWAS model to improve the GWAS model, and cannot construct a genetic-defective gene association map based on the GWAS model integrating the idea of game theory.

[0005] To solve the above technical problem, the present invention provides the following technical solution: A gene detection system, comprising a sample collection module, a gene extraction module, a sequencing module, a data analysis module, and a report generation module: The sample collection module is used to collect several types of biological samples, including blood and cell tissues. The cell tissues include feather root tissues and liver tissues. The biological samples are collected and stored through standardized collection methods and storage conditions. The gene extraction module is used to extract high-quality DNA or RNA from the original sample dataset using gene extraction techniques. The sequencing module is used to perform deep sequencing on the extracted gene samples using high-throughput sequencing technology to obtain gene sequence information, integrate the gene sequence information into the original sample dataset, and process, analyze, and interpret the original sample dataset. The data analysis module is used to integrate several bioinformatics algorithms and tools to identify defective genes and genetic information of gene expression differences, and evaluate the association between the genetic information and specific traits or defective genes. The report generation module is used to visualize the gene detection report.

[0006] Preferably, the sample collection module includes: The standardized collection methods include: When collecting blood, use a special blood collection tube to ensure that contamination or decomposition is avoided during the blood collection process. When collecting cell tissues, obtain them through biopsy or surgery, and perform cryopreservation and rapid fixation on the extracted cell tissues. The standardized storage conditions include: Perform cold chain storage on the collected biological samples, including quickly refrigerating or freezing the blood and cell tissue samples. Equip a special transport box and transport the biological samples in combination with temperature control equipment. Mark each collected biological sample with a unique identifier.

[0007] Preferably, the gene extraction module includes: The gene extraction techniques include magnetic bead method, column method, enzymatic hydrolysis method, manual extraction, and automated extraction. The steps of extracting DNA or RNA include: Step S1: Destroy the cell membrane and nuclear membrane through chemical reagents or physical methods to release DNA or RNA. The physical methods include ultrasonic waves and cryogenic grinding. Step S2: Use protease during the extraction process to remove protein impurities and lipid impurities in the cells. Step S3: Further separate and purify DNA or RNA from the lysate through centrifugation, column purification, and magnetic bead separation methods. Step S4: Quantitatively analyze the extracted DNA or RNA by spectrophotometer and fluorescence quantitative detection methods, and calculate the concentration and purity of the DNA or RNA; Judge the quality of the extracted DNA or RNA. The judgment logic includes: Detect the purity of the extracted DNA or RNA. When the purity is higher than the preset purity threshold, the DNA or RNA has high purity; Detect the structure of the extracted DNA or RNA. When there is no lysis phenomenon in the structure, the DNA or RNA has good integrity; Detect the concentration of the extracted DNA or RNA. When the concentration is within the preset concentration threshold range, the DNA or RNA has a moderate concentration; When the DNA or RNA has high purity, good integrity and moderate concentration, it is judged that the quality of the extracted DNA or RNA is high.

[0008] Preferably, the sequencing module includes: After fragmenting DNA or RNA molecules by physical or chemical methods, add sequencing adapters to ligate the DNA or RNA fragments to the adapters to form a library that can be recognized by the sequencing platform. Sequence the DNA or RNA fragments. The sequencing process includes polymerization reaction, fluorescence signal reading and cycle sequencing. Convert the signals output by the sequencer into digital signals, and generate the original gene sequence data according to the time sequence for the digital signals; Collect the original gene sequence data. The collection process includes: During the sequencing process, divide the DNA fragments into several subsequences. Each subsequence represents the sequence of a fragment. Statistically calculate the number and overlap degree of the subsequences read during the sequencing process. Splice, align and assemble the obtained subsequence information to generate a complete genomic sequence or transcriptomic sequence. Splice the overlapping subsequences into longer sequences through algorithms to generate continuous gene sequences or complete genomes. Align the spliced sequences with the known genomes retrieved from big data to identify specific gene regions and variant site elements in the genome.

[0009] Preferably, the De Bruijn graph assembly algorithm is used to find the overlaps between subsequences, and the subsequences are combined into a continuous sequence according to the overlaps. BLAST is used to align the assembled sequence with a known reference genome to identify key regions in the genome, obtaining regions consistent with the preset reference genome and regions different from the preset reference genome. The regions different from the preset reference genome are the variant regions. Structure variation (SV) tools are used to identify and analyze the variant regions and key regions in the genome. The variant regions and key regions include coding regions, intron and exon regions, variant sites, regulatory regions, transcription factor binding sites, and RNA genes. Ensembl is used to annotate the detected variants and gene regions.

[0010] Preferably, the obtained gene sequence information is converted into an original dataset. The logic for generating the original dataset includes: The original dataset is directly stored in a FASTQ file, which includes the original reads of the sequencing, base quality scores, and sequence data. The data after aligning the original sequencing data with the reference genome is stored in a BAM file. The variant information, which includes SNPs and INDELs, is stored in a VCF file.

[0011] Preferably, the data analysis module includes: The FASTQ file is extracted. The FASTQ file includes the original DNA or RNA sequence data of the sample. FastQC is used for data quality control, and the control process includes removing low-quality sequences, trimming adapter sequences and low-quality bases. Based on the aligned data, GATK is used to call SNPs to identify possible variant sites in the genome. HTSeq is used for quantitative analysis of gene expression levels to obtain the expression levels of each gene. The DESeq2 differential expression analysis tool is used to compare the gene expression differences between different groups. The Gene Ontology database is used for functional enrichment analysis of differentially expressed genes to identify their possible biological processes, cellular components, molecular functions, and the pathways involved.

[0012] Preferably, game theory is incorporated into the genome-wide association study (GWAS) model to improve the GWAS model. The improvement logic includes: When performing single-locus association analysis between genetic traits and defective genes using the above genome-wide association study (GWAS) model, the genetic traits and defective genes are regarded as two parties in a game. When choosing game strategies, both parties need to consider themselves and the other party, and finally select the game strategy that can maximize their own "benefits". When the genetic traits and defective genes are in a game, each party is implicated and restricted by the other party while trying to maximize its own "benefits". The respective maximized quantities obtained from the game depend on the actual gene expression levels. When the interaction effects between the actual gene expression levels match their own expressions, the game results between the genetic traits and defective genes are obtained. The game results are represented by constructing a trait-defective gene matrix model between the actual gene expression levels, and its mathematical expression is: ; Wherein, represents the change in the gene expression level of the genetic trait relative to on the individual, () represents the average gene expression level of the genetic trait on individual a, represents the total amount of genes of the genetic trait on the defective gene, represents the gene expression level of individual 1 in the independent state, represents the gene expression level of individual 2 in the independent state, and represent the interaction effects generated when an individual is affected by another type of individual, is the maximum likelihood estimation gene quantity of individual a, represents the maximum likelihood estimation dynamic range when an individual is affected by another type of individual; Collect the historical interactions between genetic traits and defective genes, classify the historical interactions using k-means clustering to obtain the first interaction type, the second interaction type,..., and the nth interaction type, construct a regression model of the interaction type regarding all predicted defective gene situations, and screen out the defective gene situations that have a significant interaction relationship with the interaction type. The significant judgment criterion is the defective gene situations that exceed the preset significant threshold; The data of expression interaction effects in the trait-defective gene matrix model are added, and the trait-defective gene matrix corresponding to the genetic trait is spliced ​​to obtain a comprehensive change matrix of the gene expression level of the genetic trait for the defective gene, and the comprehensive change matrix is ​​used as an estimated value for the single-site association analysis. According to the estimated value, the GLM or MLM results calculated by Tassel in the calculation process of the whole genome association analysis GWAS model are FDR corrected to obtain SNP sites, and the SNP sites include potentially significant sites.

[0013] Preferably, a genetic-defective gene association map is constructed according to a genome-wide association analysis GWAS model integrating game theory ideas, wherein the genetic-defective gene association map simultaneously displays the defective gene risk of genetic traits and the correlation between the defective gene phenotype and the genetic genotype, and the nodes of the genetic-defective gene association map are SNP and genetic variation nodes, and the edges are the strength of association between the genetic traits and the defective genes and the influence of the genetic traits on the defective genes output by the genome-wide association analysis GWAS model integrating game theory ideas; The genetic-defective gene association map is used as a training set to train a support vector machine SVM to obtain a defective gene risk prediction model, and the defective gene risk prediction model is used to predict the most likely defective gene type based on genetic characteristics.

[0014] Preferably, the report generating module comprises: ANNOVAR is used to associate SNPs with genes, transcripts, and protein structures. Based on the output of the data analysis module, a standardized genetic testing report is automatically generated, covering information on defective genes, defective gene associations, and drug responses.

[0015] The beneficial effects of the present invention are as follows: rapid and accurate deep sequencing of gene samples to obtain high-quality gene sequence information, and converting this information into original data sets, integrating game theory into the whole genome association analysis GWAS model to improve the whole genome association analysis GWAS model, combining game theory to make SNP sites more numerous and detailed, being able to dig out deeper associations between genetic traits and defective genes, and making the probability of false positive results obtained by the association model smaller, constructing a genetic-defective gene association map based on the whole genome association analysis GWAS model integrating game theory ideas, training a defective gene risk prediction model to predict the most likely defective gene type based on genetic characteristics, and greatly improving the accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A basic flow chart of a gene detection system provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0017] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments.

[0018] Referring to Figure 1 , for an embodiment of the present invention, a gene detection system is provided, including a sample collection module, a gene extraction module, a sequencing module, a data analysis module, and a report generation module: The sample collection module is used to collect several types of biological sample types, and the biological sample types include blood and cell tissue. The biological samples are collected and stored through standardized collection methods and storage conditions to ensure the quality and integrity of the samples; The gene extraction module is used to extract high-quality DNA or RNA from the original sample dataset using efficient gene extraction techniques; The sequencing module is used to perform deep sequencing on the extracted gene samples using high-throughput sequencing technology to obtain gene sequence information, synthesize the gene sequence information into an original sample dataset, process, analyze, and interpret the original sample dataset, and provide raw data for subsequent data analysis and interpretation through accurate determination of the gene sequence; The data analysis module is used to integrate several bioinformatics algorithms and tools to identify defective genes and genetic information of gene expression differences, and evaluate the association between the genetic information and specific traits or defective genes; The report generation module is used to visualize the gene detection report.

[0019] The sample collection module includes: The standardized collection methods include: When collecting blood, a dedicated blood collection tube is used to ensure that contamination or decomposition is avoided during the blood collection process; When collecting cell tissue, it is obtained through biopsy or surgery, and the extracted cell tissue is cryopreserved and rapidly fixed to ensure the quality of the tissue sample after acquisition; The standardized storage conditions include: The collected biological samples are stored in a cold chain to ensure the integrity of genetic materials such as DNA and RNA. The cold chain storage includes rapidly refrigerating or freezing blood and cell tissue samples to avoid DNA degradation; A dedicated transport box is equipped to transport the biological samples in combination with temperature control equipment to ensure that the samples always maintain appropriate temperature and environmental conditions during transportation; Each collected biological sample is labeled with a unique identifier to facilitate tracking and management during sample transportation, storage, and analysis, and prevent sample confusion or loss.

[0020] The sample collection module ensures the quality and integrity of various biological samples through standardized collection methods, precise storage and transportation conditions, and strict quality control. The efficient operation of this module is the basis for the accurate analysis and accurate reporting of the gene detection system, providing a reliable data source for subsequent gene extraction, sequencing, and data analysis. At the same time, the system reduces human operation errors and the difficulty of sample management through an automated process, improving the efficiency and accuracy of the overall process.

[0021] The gene extraction module includes: Gene extraction techniques include magnetic bead method, column method, enzymatic digestion method, manual extraction, and automated extraction; The steps for extracting DNA or RNA include: Step S1: Destroy the cell membrane and nuclear membrane through chemical reagents or physical methods to release DNA or RNA. The physical methods include ultrasonic waves and cryogenic grinding; Step S2: Use protease during the extraction process to remove intracellular protein impurities and lipid impurities to ensure the purity of DNA or RNA; Step S3: Further separate and purify DNA or RNA from the lysate through centrifugation, column purification, and magnetic bead separation methods, and use appropriate solutions and buffers to ensure that the separated nucleic acids are not contaminated; Step S4: Quantitatively analyze the extracted DNA or RNA through spectrophotometry and fluorescence quantitative detection methods, calculate the concentration and purity of DNA or RNA. The commonly used standard is to evaluate the purity of DNA or RNA through the A260 / A280 ratio. Ideally, this ratio should be between 1.8 and 2.0; Judge the quality of the extracted DNA or RNA. The judgment logic includes: Detect the purity of the extracted DNA or RNA. When the purity is higher than the manually preset purity threshold, the DNA or RNA has high purity; Detect the structure of the extracted DNA or RNA. When there is no lysis phenomenon in the structure, the DNA or RNA has good integrity; Detect the concentration of the extracted DNA or RNA. When the concentration is within the manually preset concentration threshold range, the DNA or RNA has a moderate concentration; When the DNA or RNA has high purity, good integrity, and moderate concentration, it is judged that the quality of the extracted DNA or RNA is high.

[0022] The gene extraction module is a key link in the entire gene detection process. Through efficient and precise gene extraction technology, high-quality DNA or RNA is extracted from various biological samples, laying the foundation for subsequent gene sequencing, gene expression analysis, and data interpretation. The standardized process ensures the efficiency, consistency, and accuracy of the extraction process, thereby improving the working efficiency and detection accuracy of the entire system.

[0023] The sequencing module includes: After fragmenting DNA or RNA molecules by physical or chemical methods, by adding sequencing adapters, the DNA or RNA fragments are ligated to the adapters to form a library that can be recognized by the sequencing platform. The DNA or RNA fragments are sequenced. The sequencing process includes polymerization reaction, fluorescence signal reading, and cycle sequencing. The signals output from sequencing are converted into digital signals by the sequencer, and the digital signals are used to generate raw gene sequence data according to the time sequence. Collect the raw gene sequence data. The collection process includes: During the sequencing process, DNA fragments are divided into several subsequences. Each subsequence represents the sequence of a fragment. The number and overlap degree of the subsequences read during sequencing are statistically calculated. The higher the sequencing depth, the more subsequences are generated, thereby improving the coverage and accuracy. The overlapping regions enable adjacent subsequences to verify each other, improving the accuracy of splicing. The obtained subsequence information is spliced, aligned, and assembled to generate a complete genomic sequence or transcriptomic sequence. The overlapping subsequences are spliced into longer sequences by algorithms to generate continuous gene sequences or complete genomes. The spliced sequences are aligned with the known genomes retrieved from big data to identify specific gene regions and variant site elements in the genome.

[0024] Use the De Bruijn graph splicing algorithm to find the overlapping parts between subsequences, combine the subsequences into continuous sequences according to the overlapping parts, use BLAST to align the spliced sequences with the known reference genomes, identify the key regions in the genome, obtain the regions consistent with the preset reference genome and the regions different from the preset reference genome. The regions different from the preset reference genome are the variant regions. Use the structural variation SV tool to identify and analyze the variant regions and key regions in the genome. The variant regions and key regions include coding regions, intron and exon regions, variant sites, regulatory regions, transcription factor binding sites, and RNA genes. Use Ensembl to annotate the detected variants and gene regions.

[0025] This process includes the segmentation of DNA fragments, the generation and splicing of subsequences, alignment and variant identification, and finally the detailed analysis of genomes and transcriptomes. By applying modern algorithms and tools, researchers can precisely reconstruct the complete genomic sequence and identify gene regions and variant sites associated with specific phenotypes or defective genes.

[0026] Converting the obtained gene sequence information into an original dataset, the logic for generating the original dataset includes: Directly storing the original dataset into a FASTQ file, which includes the original reads of sequencing, base quality scores, and sequence data; Storing the data after aligning the original sequencing data with the reference genome into a BAM file, which is commonly used for downstream variant detection and gene annotation; Storing the variant information into a VCF file, and the variant information includes SNPs and INDELs, for further genetic analysis, defective gene research, or personalized medical applications.

[0027] Deep sequencing can cover the entire genome to obtain sequence information of the whole genome or target regions, with extremely high detection sensitivity. The sequencing module uses high-throughput sequencing technology to quickly and accurately perform deep sequencing on gene samples, obtain high-quality gene sequence information, and convert this information into an original dataset.

[0028] The data analysis module includes: Extracting the FASTQ file, which includes the original sequence data of DNA or RNA of the sample, and using FastQC for data quality control. The control process includes removing low-quality sequences, trimming adapter sequences and low-quality bases; Based on the aligned data, using GATK to call SNPs to identify possible variant sites in the genome, using HTSeq for gene expression quantification to obtain the expression levels of each gene respectively, comparing the gene expression differences between different groups through the DESeq2 differential expression analysis tool, and using the Gene Ontology database to perform functional enrichment analysis on the differentially expressed genes to identify their possible biological processes, cellular components, molecular functions, and the pathways involved.

[0029] Analyzing the interaction between genes and defective genes through a game theory model to identify which defective genes have important effects on specific defective genes, analyzing the propagation pattern of genes in the population through game theory to reveal how natural selection affects defective genes related to defective genes, and combining the adaptive model and correlation analysis of game theory can predict the defective gene risks that may be caused by different genetic traits in the future evolution process.

[0030] Integrate game theory into the genome-wide association study (GWAS) model to improve the GWAS model. The improvement logic includes: When the GWAS model performs single-locus association analysis between genetic traits and defective genes, consider the genetic traits and defective genes as the two sides of the game. When choosing game strategies, both sides of the game need to take into account themselves and the other side, and finally choose the game strategy that can maximize their own "benefit". When the genetic traits and defective genes are in the game, each side is implicated and restricted by the other side while trying to maximize its own "benefit". The respective maximized quantities obtained from the game depend on the actual gene expression levels. When the interaction effect between the actual gene expression levels is consistent with their own expressions, the game result between the genetic traits and defective genes is obtained. Represent the game result through a trait-defective gene matrix model constructed between the actual gene expression levels. Its mathematical expression is: ; Among them, represents the change in the gene expression level of the genetic trait relative to on an individual, () represents the average gene expression level of the genetic trait on individual a, represents the total gene amount of the genetic trait on the defective gene, represents the gene expression level of individual 1 in the independent state, represents the gene expression level of individual 2 in the independent state, and represent the interaction effects generated when an individual is affected by another type of individual, is the maximum likelihood estimation gene amount of individual a, represents the maximum likelihood estimation dynamic range when an individual is affected by another type of individual; Collect the historical interactions between genetic traits and defective genes, classify the historical interactions using k-means clustering to obtain the first interaction type, the second interaction type,..., and the nth interaction type. Construct a regression model of the interaction type regarding all predicted defective gene situations, and screen out the defective gene situations that have a significant interaction relationship with the interaction type. The significant judgment criterion is the defective gene situations that exceed the preset significant threshold; Add the data expressing interaction effects in the trait-defect gene matrix model, splice the trait-defect gene matrix corresponding to the genetic trait, obtain a comprehensive change matrix of the gene expression level of the genetic trait with respect to the defect gene, use the comprehensive change matrix as the estimated value for single-site association analysis, and perform FDR correction on the GLM or MLM results calculated by tassel in the process of calculating the genome-wide association analysis GWAS model according to the estimated value to obtain SNP sites. The SNP sites include potentially significant sites. Combining the game theory can make the SNP sites more numerous and detailed, can dig deeper into the correlation between genetic traits and defect genes, and make the false positive probability of the results obtained by the association model smaller.

[0031] Construct a genetic-defect gene association map according to the genome-wide association analysis GWAS model integrating the game theory idea. The genetic-defect gene association map simultaneously shows the defect gene risk of the genetic trait and the correlation between the defect gene phenotype and the genetic genotype. The nodes of the genetic-defect gene association map are SNP and genetic variation nodes, and the edges are the association strength between the genetic trait and the defect gene output by the genome-wide association analysis GWAS model integrating the game theory idea and the influence of the genetic trait on the defect gene. Use the genetic-defect gene association map as a training set to train a support vector machine SVM to obtain a defect gene risk prediction model. The defect gene risk prediction model is used to predict the most likely defect gene type according to genetic characteristics. For example: If a certain specific defect gene shows a strong selection advantage and is significantly associated with a high risk of a certain defect gene, then this defect gene may be closely related to this genetic variation; If under specific environmental conditions, a certain genotype shows "adaptive advantage", then this genotype may be related to the high incidence or drug resistance of certain specific defect genes.

[0032] The data analysis module can comprehensively analyze the original high-throughput sequencing data by integrating a variety of bioinformatics tools and algorithms. The analysis results not only help identify defect genes and gene expression differences, but also can deeply explore the relationship between these genetic information and specific traits or defect genes.

[0033] The report generation module includes: Use ANNOVAR to associate SNPs with genes, transcripts, and protein structures to predict their potential functional impacts, and automatically generate a standardized gene detection report covering information on defect genes, defect gene associations, and drug responses according to the output of the data analysis module.

[0034] The present invention performs deep sequencing on gene samples quickly and accurately, obtains high-quality gene sequence information, and converts this information into an original data set. Game theory is incorporated into the genome-wide association analysis (GWAS) model to improve the GWAS model. Combining game theory makes the SNP sites more numerous and detailed, enabling the discovery of deeper associations between genetic traits and defective genes, and reducing the false positive probability of the results obtained by the association model. A genetic-defective gene association map is constructed based on the GWAS model integrating game theory ideas, and a defective gene risk prediction model is trained to predict the most likely defective gene types according to genetic characteristics, greatly improving the accuracy.

[0035] Through high-throughput sequencing and advanced bioinformatics algorithms, the genetic background of breeding ducks can be deeply explored, helping to identify gene variations related to production traits, disease resistance traits, etc. This helps to screen high-quality breeding ducks, promote the transmission of excellent genetic characteristics, extract high-quality DNA or RNA from various biological sample types, ensuring the accuracy and reliability of the data. Using efficient gene extraction and sequencing technologies, accurate gene sequence data can be obtained, avoiding analysis errors caused by sample quality problems. Through in-depth data analysis, the system compares gene expression differences between different populations, enabling the identification of the genetic responses of breeding ducks under different environments and feeding conditions, assisting in more precisely selecting genotypes with excellent traits in breeding work. Through the report generation module, detailed and visual gene detection reports can be automatically generated. These reports not only cover the associations between gene defects and genetic traits but also further analyze information such as drug reactions, providing a basis for accurate breeding decisions. During the breeding process of breeding ducks, the relationship between genotypes and phenotypes can be understood more clearly, thus breeding high-quality breeding ducks that are more adaptable to the environment. By incorporating game theory into the GWAS model, the complex interactions between genes and traits can be predicted more accurately. This model helps to understand the dynamic relationship between genetic traits and defective genes and can more accurately identify the key gene regions affecting the performance of breeding ducks. By constructing a genetic-defective gene association map, the risk prediction of defective genes in breeding ducks is more accurate. This map helps breeding experts quickly identify gene variations related to specific traits when selecting breeding ducks and implement a more efficient breeding strategy for breeding ducks; Overall, this gene detection system can provide more scientific and accurate support during the breeding process of breeding ducks, improving the breeding efficiency and the quality of breeding ducks, ultimately contributing to increased production efficiency, reduced disease occurrence, and improved health levels of the population.

[0036] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. Among them, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM for short), electrically erasable programmable read-only memory (EEPROM for short), erasable programmable read-only memory (EPROM for short), programmable read-only memory (PROM for short), read-only memory (ROM for short), magnetic memory, flash memory, magnetic disk, or optical disk. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the process Figure 1 one process or multiple processes and / or blocks Figure 1 the functions specified in one block or multiple blocks.

[0037] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A gene detection system, characterized in that: Including sample collection module, gene extraction module, sequencing module, data analysis module and report generation module: The sample collection module is used to collect several types of biological samples, including blood and cell tissues, and the cell tissues include feather root tissues and liver tissues, and collect and store biological samples through standardized collection methods and storage conditions; The gene extraction module is used to extract high-quality DNA or RNA from the original sample data set using gene extraction technology; The sequencing module is used to perform deep sequencing on the extracted gene samples using high-throughput sequencing technology to obtain gene sequence information, aggregate the gene sequence information into an original sample data set, and process, analyze and interpret the original sample data set; The data analysis module is used to integrate several bioinformatics algorithms and tools to identify genetic information of defective genes and gene expression differences, and to evaluate the association between the genetic information and specific traits or defective genes; The report generation module is used to visualize the gene detection report.

2. The gene detection system according to claim 1, characterized in that: The sample collection module comprises: Standardized collection methods include: When collecting blood, use dedicated blood collection tubes to ensure that the blood is not contaminated or decomposed during the collection process; When collecting cell tissue, it is obtained through biopsy or surgery, and the extracted cell tissue is frozen, stored and quickly fixed; Standardized storage conditions include: Perform cold chain storage on the collected biological samples, which includes rapid refrigeration or freezing of blood and cell tissue samples; A special transport box is provided to transport the biological sample in combination with temperature control equipment; Each collected biological sample is labeled with a unique identifier.

3. The gene detection system according to claim 2, characterized in that: The gene extraction module comprises: Gene extraction techniques include magnetic bead method, column method, enzymatic method, manual extraction and automated extraction; The steps for extracting DNA or RNA include: Step S1: destroying the cell membrane and the cell nuclear membrane by chemical reagents or physical methods to release DNA or RNA, wherein the physical methods include ultrasound and cryo-grinding; Step S2: using protease to remove protein impurities and lipid impurities in cells during the extraction process; Step S3: further separating and purifying the DNA or RNA from the lysate by centrifugation, column purification and magnetic bead separation; Step S4: quantitatively analyzing the extracted DNA or RNA by a spectrophotometer and a fluorescence quantitative detection method to calculate the concentration and purity of the DNA or RNA; Determine the quality of the extracted DNA or RNA. The judgment logic includes: Detecting the purity of the extracted DNA or RNA, when the purity is higher than a preset purity threshold, the purity of the DNA or RNA is high; Detecting the structure of the extracted DNA or RNA, when the structure has no cleavage phenomenon, the integrity of the DNA or RNA is good; Detecting the concentration of the extracted DNA or RNA, when the concentration is within a preset concentration threshold range, the DNA or RNA concentration is moderate; When the DNA or RNA has high purity, good integrity and moderate concentration, the extracted DNA or RNA is judged to be of high quality.

4. The gene detection system according to claim 3, characterized in that: The sequencing module comprises: After DNA or RNA molecules are fragmented by physical or chemical methods, a sequencing adapter is added to connect the DNA or RNA fragments to the adapter to form a library that can be recognized by the sequencing platform, and the DNA or RNA fragments are sequenced. The sequencing process includes polymerization reaction, fluorescence signal reading and cycle sequencing. The sequencing output signal is converted into a digital signal by a sequencer, and the digital signal is used to generate original gene sequence data in time sequence; The original gene sequence data is collected, and the collection process includes: During the sequencing process, the DNA fragments are divided into several subsequences, each of which represents the sequence of a fragment. The number of subsequences read and the degree of overlap during the sequencing process are statistically calculated, and the obtained subsequence information is spliced, aligned, and assembled to generate a complete genome sequence or transcriptome sequence. The overlapping subsequences are spliced ​​into longer sequences through algorithms to generate a continuous gene sequence or a complete genome. The spliced ​​sequence is compared with the known genome obtained by big data retrieval to identify specific gene regions and variant site elements in the genome.

5. The gene detection system according to claim 4, characterized in that: The De Bruijn graph splicing algorithm is used to find the overlapping parts between subsequences, and the subsequences are combined into a continuous sequence according to the overlapping parts. The spliced ​​sequences are compared with the known reference genome using BLAST to identify the key regions in the genome, and the regions consistent with the preset reference genome and the regions different from the preset reference genome are obtained. The regions different from the preset reference genome are the variable regions. The structural variation SV tool is used to identify and analyze the variable regions and key regions in the genome. The variable regions and key regions include coding regions, introns and exons, variant sites, regulatory regions, transcription factor binding sites and RNA genes. Ensembl is used to annotate the detected variants and gene regions.

6. The gene detection system according to claim 5, characterized in that: The acquired gene sequence information is converted into an original data set. The logic for generating the original data set includes: The raw data set is directly stored in a FASTQ file, wherein the FASTQ file includes raw sequencing reads, base quality scores, and sequence data; The data after comparing the original sequencing data with the reference genome is stored in a BAM file; The variation information is stored in a VCF file, wherein the variation information includes SNP and INDEL.

7. The gene detection system according to claim 6, characterized in that: The data analysis module includes: Extracting a FASTQ file, wherein the FASTQ file includes the original sequence data of DNA or RNA of the sample, and performing data quality control using FastQC, wherein the control process includes removing low-quality sequences, trimming off adapter sequences and low-quality bases; Based on the aligned data, GATK was used to call SNPs and identify possible variant sites in the genome. HTSeq was used to quantify gene expression and obtain the expression levels of each gene. The DESeq2 differential expression analysis tool was used to compare the gene expression differences between different groups. The Gene Ontology database was used to perform functional enrichment analysis on differentially expressed genes to identify their possible biological processes, cellular components, molecular functions, and pathways involved.

8. The gene detection system according to claim 7, characterized in that: Game theory is integrated into the genome-wide association analysis GWAS model to improve the genome-wide association analysis GWAS model. The improvement logic includes: When the genome-wide association analysis GWAS model performs single-point association analysis between genetic traits and defective genes, the genetic traits and defective genes are regarded as the two parties of the game, and the respective maximization amounts obtained by the game depend on the actual gene expression amounts. When the interaction effect between the actual gene expression amounts is consistent with the self-expression, the game result between the genetic traits and the defective genes is obtained, and the game result is represented by constructing a trait-defective gene matrix model between the actual gene expression amounts, and its mathematical expression is: ; in, express Individual genetic traits Relative to Changes in gene expression levels, () indicates genetic traits The average gene expression in individual a, Indicates genetic traits exist The total amount of genes on the defective gene, Indicates the gene expression level of an individual in an independent state. Indicates the gene expression level of two individuals in an independent state. and It represents the interaction effect when an individual is influenced by another type of individual. is the maximum likelihood estimated gene amount of individual a, Represents the maximum likelihood estimate dynamic range when an individual is affected by another type of individual; Collect historical interactions between genetic traits and defective genes, classify the historical interactions using k-means clustering to obtain a first interaction type, a second interaction type, ... and an nth interaction type, construct a regression model of the interaction type with respect to all predicted defective gene situations, and screen out defective gene situations that have a significant interaction relationship with the interaction type, wherein the significant judgment standard is the defective gene situation that exceeds a preset significant threshold; The data of expression interaction effects in the trait-defective gene matrix model are added, and the trait-defective gene matrix corresponding to the genetic trait is spliced ​​to obtain a comprehensive change matrix of the gene expression level of the genetic trait for the defective gene, and the comprehensive change matrix is ​​used as an estimated value for the single-site association analysis. According to the estimated value, the GLM or MLM results calculated by Tassel in the calculation process of the whole genome association analysis GWAS model are FDR corrected to obtain SNP sites, and the SNP sites include potentially significant sites.

9. The gene detection system according to claim 8, characterized in that: A genetic-defective gene association map is constructed according to a genome-wide association analysis GWAS model integrating game theory ideas, wherein the genetic-defective gene association map simultaneously displays the defective gene risk of genetic traits and the correlation between the defective gene phenotype and the genetic genotype, wherein the nodes of the genetic-defective gene association map are SNP and genetic variation nodes, and the edges are the strength of association between the genetic traits and the defective genes and the influence of the genetic traits on the defective genes output by the genome-wide association analysis GWAS model integrating game theory ideas; The genetic-defective gene association map is used as a training set to train a support vector machine SVM to obtain a defective gene risk prediction model, and the defective gene risk prediction model is used to predict the most likely defective gene type based on genetic characteristics.

10. The gene detection system according to claim 9, characterized in that: The report generation module comprises: ANNOVAR is used to associate SNPs with genes, transcripts, and protein structures. Based on the output of the data analysis module, a standardized genetic testing report is automatically generated, covering information on defective genes, defective gene associations, and drug responses.

Citation Information

Patent Citations

  • Drug resistance gene detection method and system

    CN118600052A

  • Evaluating system based on gene testing

    CN110246581A

  • Extraction and detection integrated gene detection system

    CN114005488A

  • Machine learning driven gene discovery and gene editing in plants

    US20220301658A1

Cited By

  • Method for evaluating sequence quality after magnetic bead-nucleic acid separation

    CN120260681A