Data analysis method, system and storage medium for genetic disease gene detection
Through automated screening and variation analysis technology, combined with information entropy calculation and Bayesian inference analysis, and combined with patient phenotypic information, standardized reports are generated, which solves the problems of low efficiency and insufficient accuracy of variation screening in genetic disease gene testing, and realizes efficient and accurate diagnosis and report generation.
Patent Information
- Application Number
- CN202510637839.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-05-19
AI Technical Summary
In existing genetic disease genetic testing technologies, mutation screening efficiency is low, accuracy is insufficient, computing resource utilization is not optimized, report generation is cumbersome and lacks personalization, making it difficult to meet the needs of efficient and accurate clinical applications.
Automated screening and variation analysis technology is used, combined with information entropy calculation and Bayesian inference analysis, combined with patient phenotypic information, to generate standardized reports and optimize computing resource allocation.
It achieves efficient screening of pathogenic variants, improves diagnostic accuracy and work efficiency, reduces human errors, ensures accuracy and consistency of reports, and optimizes computing resource utilization.
Smart Images

Figure CN120164524B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical data analysis technology, and in particular to a data analysis method, system and storage medium for genetic disease gene detection. Background Art
[0002] Genetic testing for genetic diseases analyzes a patient's genomic information to identify potential pathogenic variants, providing a scientific basis for early diagnosis, treatment, and genetic counseling. However, existing genetic testing technologies face several shortcomings, resulting in certain limitations in their clinical applications.
[0003] First, in existing technologies, the variant screening process often relies on manual intervention. Manual methods are inefficient for screening large-scale data and are easily affected by the operator's experience, with a certain risk of misjudgment or omission. Especially with the continuous increase in the amount of genomic data, the time cost and error risk brought about by manual intervention are constantly increasing, making it difficult to meet the needs of efficiency and accuracy. The screening of genetic variants, pathogenicity analysis, and matching with patient phenotypes are usually highly complex and tedious tasks. Existing methods often rely on manual judgment, which greatly affects the accuracy of diagnosis and the efficiency of clinical application.
[0004] Secondly, traditional genetic variation analysis often relies on single-sample data, neglecting the integration of multiple information sources. In the clinical diagnosis of genetic diseases, relying solely on genomic data analysis fails to fully consider a patient's clinical phenotype and family history. This results in a lack of personalized accuracy in variant screening, potentially missing variants that are crucial to a patient's clinical characteristics and compromising a comprehensive diagnosis.
[0005] Furthermore, the current process for generating genetic testing reports is relatively traditional and cumbersome. While genetic testing technology continues to advance, report generation still relies on the physician's judgment to confirm the final results. Reports are often manually completed by physicians, which not only introduces time delays but is also susceptible to information processing errors or inconsistent formats that can affect the final diagnosis. While existing technologies also offer automated report generation, most systems lack standardized templates and flexible customization options, making it difficult to automatically generate appropriate diagnostic reports based on the patient's specific circumstances. This significantly reduces the accuracy and practicality of the reports.
[0006] Furthermore, existing computing resource management is inadequate. In genetic testing for inherited diseases, processing and analyzing large amounts of genetic data consumes significant computing resources. Existing technologies often neglect the efficient allocation of computing resources, resulting in wasted resources. This is particularly true in diverse application scenarios, particularly in clinical settings, where the need for efficient, accurate, and resource-optimized analysis is extremely pressing.
[0007] Therefore, the present invention proposes a data analysis method, system and storage medium for genetic disease gene detection to address the deficiencies of the existing technology. Summary of the Invention
[0008] In response to the shortcomings of the existing technology, the present invention provides a data analysis method, system and storage medium for genetic disease gene detection, which solves the problems of low efficiency, insufficient accuracy and non-optimal computing resource utilization in gene variation screening, pathogenicity analysis and report generation in the existing technology.
[0009] To achieve the above objectives, the present invention is implemented through the following technical solutions: A data analysis method for genetic disease gene detection, comprising the following steps:
[0010] Collect patient samples and perform high-throughput sequencing to obtain raw sequencing data;
[0011] performing quality control and alignment processing on the raw sequencing data to generate variation detection data, and calculating variation information based on the variation detection data;
[0012] Automatically screen suspicious mutation sites and calculate and analyze the amount of mutation information based on information entropy, and screen key mutations according to the preset entropy threshold;
[0013] Bayesian inference analysis was performed on the key variants screened, their pathogenicity probabilities were calculated, and pathogenicity was determined based on the ACMG variant classification criteria;
[0014] Combined with the patient's phenotypic information, the matching degree between the key variants and the patient's phenotypic information is calculated, and the variants that match the patient's phenotypic characteristics are screened out;
[0015] Submit the screened variants to doctors for review and generate the final screening results;
[0016] Generate a standardized genetic testing report based on the final screening results.
[0017] The present invention also provides a data analysis system for genetic disease gene detection, comprising:
[0018] A sample collection module is used to collect patient samples and perform high-throughput sequencing to obtain raw sequencing data;
[0019] a data processing module, configured to perform quality control and alignment processing on the raw sequencing data, generate variation detection data, and calculate variation information based on the variation detection data;
[0020] The mutation screening module is used to automatically screen suspicious mutation sites and calculate and analyze the amount of mutation information based on information entropy, and screen key mutations according to the preset entropy threshold;
[0021] The Bayesian inference analysis module is used to perform Bayesian inference analysis on the selected key variants, calculate their pathogenicity probability, and make pathogenicity determinations based on the ACMG variant classification criteria;
[0022] The phenotypic matching module is used to calculate the matching degree between key variants and patient phenotypic information and screen out variants that match the patient's phenotypic characteristics;
[0023] The doctor review module is used to submit the screened variants to doctors for review and generate the final screening results;
[0024] The report generation module is used to generate a standardized genetic testing report based on the final screening results.
[0025] The present invention also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the method described above is implemented.
[0026] The present invention provides a data analysis method, system, and storage medium for genetic disease gene detection. It has the following beneficial effects:
[0027] 1. This invention utilizes automated screening and variant analysis techniques, using information entropy calculation and Bayesian inference analysis to achieve precise variant screening. This achieves the technical effect of efficiently screening for pathogenic variants. Compared to existing variant screening methods that require significant manual intervention, this method solves the cumbersome and error-prone screening process. The automated process not only improves efficiency but also reduces the incidence of human error.
[0028] 2. This invention combines a patient's clinical data with genetic variation information through a phenotypic matching module to accurately assess the match between variation and phenotype. This allows for the screening of variants that are highly correlated with the patient's clinical characteristics, achieving the technical effect of personalized diagnosis. Compared with existing technologies, traditional methods fail to fully incorporate the patient's phenotypic information, resulting in lower accuracy in variant screening and affecting the reliability of clinical diagnosis.
[0029] 3. This invention automatically generates standardized genetic test reports through a report generation module, ensuring a uniform report format, accurate content, and the ability to meet clinical needs. Compared to the manual report generation process in the prior art, this invention reduces human error and ensures report accuracy and consistency. This standardized output significantly improves the communication efficiency of genetic test results and facilitates rapid decision-making by clinicians.
[0030] 4. This invention uses optimal control theory to optimize the allocation of computing resources, ensuring that the information value of screened variants can be maximized even with limited computing resources. This achieves efficient use of computing resources. Compared to existing technologies that severely waste computing resources, this invention significantly reduces computing costs through resource optimization and improves overall system performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a flow chart of the method of the present invention;
[0032] Figure 2 This is a system architecture diagram of the present invention;
[0033] Figure 3 Schematic diagram of the computer device structure of the present invention.
[0034] Among them, 10. Computer equipment; 11. Processor; 12. Memory; 13. Storage medium. DETAILED DESCRIPTION
[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0036] See also Figure 1 The present invention provides a data analysis method, system, and storage medium for genetic disease gene detection, including the following steps:
[0037] S1. Collect patient samples and perform high-throughput sequencing to obtain raw sequencing data;
[0038] S2. performing quality control and alignment processing on the raw sequencing data to generate variation detection data, and calculating variation information based on the variation detection data;
[0039] S3: Automatically screen suspicious mutation sites and calculate and analyze the amount of mutation information based on information entropy, and screen key mutations according to the preset entropy threshold;
[0040] S4. Perform Bayesian inference analysis on the selected key variants, calculate their pathogenicity probability, and make pathogenicity determinations based on the ACMG variant classification criteria;
[0041] S5. Calculate the matching degree between key variants and patient phenotypic information based on the patient's phenotypic information, and select variants that match the patient's phenotypic characteristics;
[0042] S6. Submit the screened variants to the doctor for review and generate the final screening results;
[0043] S7. Generate a standardized genetic testing report based on the final screening results.
[0044] Step S1 involves collecting patient samples and performing high-throughput sequencing to obtain raw sequencing data. High-throughput sequencing technology enables rapid and accurate acquisition of large-scale genomic data, providing sufficient sample information for variant detection. The core goal of this step is to ensure data quality and integrity, providing a reliable basis for subsequent quality control, variant screening, and pathogenicity inference.
[0045] Typically, the patient's sample can be blood, saliva, or other human specimens. DNA from the sample is first extracted using standardized protocols. High-throughput sequencing technology is then used to sequence the extracted genomic DNA, generating raw sequencing data. This data contains information about various regions of the genome, including potential mutation sites, and provides detailed data support for subsequent genetic analysis.
[0046] In this embodiment, the specific operations for collecting patient samples are as follows:
[0047] Sample Collection: First, a DNA sample is collected from the patient, typically using blood or saliva. Standardized sampling methods are used to ensure the quality and representativeness of each sample. In some cases, the most appropriate sample type may be selected based on the patient's clinical background.
[0048] DNA Extraction: After cell lysis of the collected sample, genomic DNA is extracted. Typically, commercial DNA extraction kits are used to ensure DNA integrity and purity. In some embodiments, other types of extraction methods, such as phenol / chloroform extraction or silica gel column extraction, can also be used to meet different needs.
[0049] High-throughput sequencing: After the extracted DNA undergoes adapter ligation and PCR amplification, it is sequenced on a high-throughput sequencing platform. Sequencing platforms include, but are not limited to, Illumina and PacBio. Generally, these platforms use sequencing-by-synthesis technology, capable of reading thousands of DNA fragments simultaneously. This method enables the generation of massive amounts of data covering the entire genome.
[0050] Data Acquisition Process: During high-throughput sequencing, DNA samples are cut into shorter fragments, which are then decoded by high-precision sequencing equipment. Each DNA fragment is assigned a unique identifier, and the sequencing equipment combines the information from all fragments based on these identifiers to form a complete genomic data set. The raw data generated by this process is typically stored in FASTQ format and includes sequencing quality values and sequencing site information.
[0051] Quality Control and Data Correction: After acquiring raw data, the system performs preliminary quality control. This primarily involves using alignment algorithms to assess the quality of sequencing results and remove low-quality reads. Specifically, quality control includes removing low-quality reads, sequencing errors, and contaminant data. To ensure data accuracy, the system performs multiple corrections on the sequencing data during subsequent analysis.
[0052] Acquisition of raw sequencing data: The raw data obtained through high-throughput sequencing technology contains comparison data with the reference genome. The raw data includes the sequence information of all measured base pairs. The sequence of each data fragment can be expressed as follows:
[0053] ;
[0054] in, To indicate the A sequencing fragment is a DNA sequence extracted from a patient sample; To represent each base in the sequenced fragment, is the first base of the fragment, is the second base, and so on until , Indicates the total length of the sequenced fragment.
[0055] Alternatively, when selecting a sequencing platform, different sequencing depths can be chosen based on sample characteristics and desired coverage. For example, in clinical applications, if high accuracy is required for variant detection, a higher-depth sequencing solution, typically 30x coverage, can be selected. In certain specific cases, lower coverage can also be selected to optimize costs.
[0056] In one possible implementation, to improve sequencing accuracy, secondary sequencing or high-precision sequencing technology (such as PacBioSMRT sequencing technology) can be performed to ensure accurate measurement of the genome in difficult-to-resolve regions.
[0057] For the quality control of sequencing data, the following formula was used to evaluate the quality of sequencing:
[0058] ;
[0059] in, The sequencing quality score (Phred quality score) is used to indicate the reliability of the sequencing results. The higher the quality score, the higher the sequencing accuracy. is the probability of sequencing error, that is, the probability of a certain base being wrong. The value of is between the interval [0,1]. =0.001 means there is approximately one error per kilobase.
[0060] During the data alignment and correction stage, alignment algorithms (such as BWA, Bowtie, etc.) are used to align the sequencing data with the reference genome, and the optimal alignment position is determined by the shortest path algorithm.
[0061] Through the above steps, we can obtain high-quality raw sequencing data, providing accurate and reliable data support for subsequent genetic variation detection. High-throughput sequencing technology not only enables large-scale and high-precision variation detection, but also captures a variety of genetic variations, including SNPs, indels, and structural variants. Through this step, the system can provide an accurate raw data foundation for subsequent genetic data analysis, providing strong technical support for the diagnosis and treatment of genetic diseases.
[0062] In general, this step ensures high-quality genomic data collection through standardized operating procedures. At the same time, through the application of high-throughput sequencing technology, it ensures comprehensive coverage and efficient detection of complex genetic variations, laying a solid foundation for subsequent variation screening and pathogenicity analysis.
[0063] In step S2, quality control and alignment are performed to generate variant detection data, and the amount of relevant information about the variant is further calculated. The purpose of this process is to ensure the accuracy of subsequent variant screening and provide reliable data support for pathogenicity analysis.
[0064] Typically, raw sequencing data may contain quality issues such as sequencing errors, low-quality fragments, and duplicates. Therefore, quality control is essential to remove invalid or low-quality data. This quality-controlled data is then compared to a reference genome to identify potential genetic variants (such as SNPs and indels). Subsequently, calculation of the variant information content helps assess the potential significance of each variant, providing more accurate data support for subsequent analysis.
[0065] In this embodiment, the main tasks of step S2 include: performing quality control on the original sequencing data, comparing it with the reference genome, generating variant detection data, and screening out meaningful variants through information calculation methods.
[0066] First, the raw sequencing data undergoes quality control. In some embodiments, quality control primarily involves removing low-quality reads and removing sequencing errors. Specifically, for each sequenced fragment, the system evaluates its reliability by calculating its quality score (Phred score), using the formula:
[0067] ;
[0068] in, is the quality score; is the probability of sequencing errors for this base. Generally speaking, the quality score (i.e. the error rate is less than ) is considered high-quality data. In this way, low-quality sequencing data will be eliminated to ensure the accuracy of subsequent analysis.
[0069] After removing low-quality data, the next step is to align the quality-controlled data with a reference genome. Common alignment tools (such as BWA, Bowtie, and GATK) are typically used for genome alignment. The goal of alignment is to identify potential genetic variants, such as single nucleotide variants (SNPs) and insertion / deletion variants (Indels), by aligning the raw sequencing data with the reference genome.
[0070] The comparison process is generally carried out through the following formula:
[0071] ;
[0072] in, This score represents the alignment score between the entire sequenced fragment and the reference genome. This score can be used to evaluate the quality of the alignment. The higher the score, the better the match between the sequenced data and the reference genome. To indicate the length of the sequencing fragment, that is, the number of base pairs that need to be aligned; To indicate the bases of sequenced fragments; Represents all bases of the sequenced fragment; To represent the reference genome bases at positions.
[0073] After the alignment is complete, the system will output variant detection data. This data includes information about all variant sites identified in the alignment, such as variant type (SNP, Indel, etc.), variant location, and variant frequency. These variant sites will serve as the basis for subsequent analysis.
[0074] Once the variant detection data is obtained, the next step is to calculate the information content of the variant. The calculation of information content is based on the information entropy theory and is usually calculated using the following formula:
[0075] ;
[0076] in, Information entropy, also known as information quantity, measures the random variable uncertainty or complexity of information; Indicates the number of events or mutation outcomes, usually referring to the number of all possible mutation types; represents a random variable No. possible outcomes (in genomic data, this could be a specific type of variant, or a variant occurring at a specific position); Indicates the The probability of an event (or type of variation) occurring; Representing an event Probability of occurrence The binary logarithm of . A variant with greater information content represents greater unpredictability and is generally considered a more meaningful variant. By calculating the entropy value of all variants, the system can assign an information content score to each variant.
[0077] As an option, during the quality control phase, sequencing data can be subjected to more detailed deduplication processing to remove duplicate sequencing fragments to further improve data quality. In addition, the data coverage can be improved by increasing the sequencing depth of the sample to ensure that every variant site can be accurately detected.
[0078] When choosing an alignment algorithm, you can use different alignment tools based on specific needs. For example, when processing low-complexity genomes, you might use a simpler alignment tool to reduce computing time; when processing high-complexity genomes, you might choose a more precise alignment algorithm to ensure the accuracy of the alignment results.
[0079] By quality-controlling and aligning raw sequencing data, the system can filter out low-quality data, ensuring the accuracy of subsequent analysis. Furthermore, the generation of variant detection data and the calculation of information content provide an effective screening mechanism, making variant analysis results more reliable. The variant information obtained in this step provides a key basis for subsequent pathogenicity analysis, further improving the efficiency and accuracy of genetic variant detection.
[0080] In step S3, suspicious variant sites are automatically screened and the information content of the variants is calculated and analyzed based on the entropy values of these variant sites. By screening key variants based on a preset entropy threshold, the system can more accurately identify potential pathogenic variants. This step is a key part of the entire genetic testing process, directly affecting subsequent pathogenicity analysis and clinical decision support.
[0081] Typically, the variant data screening process uses preset thresholds to filter out low-information or irrelevant variants. By calculating information entropy, the system can assess the "uncertainty" of each variant, i.e., its potential clinical importance. In some embodiments, the system uses entropy as a measure of variant importance, and the threshold setting can be adjusted based on the specific application scenario to achieve higher accuracy and reliability.
[0082] In this embodiment, the main goal of step S3 is to screen out key variants with clinical significance. These variants may be closely related to the occurrence of genetic diseases and therefore require further analysis and confirmation.
[0083] First, the system will screen the variant detection data generated in the previous step. During this process, the system will automatically detect variant sites with high information content. Specifically, the information content of the variant is evaluated by calculating the information entropy of each variant site. The information entropy calculation formula is:
[0084] ;
[0085] in, Information entropy, also known as information quantity, measures the random variable uncertainty or complexity of information; Indicates the number of events or mutation outcomes, usually referring to the number of all possible mutation types; represents a random variable No. possible outcomes (in genomic data, this could be a specific type of variant, or a variant occurring at a specific position); Indicates the The probability of an event (or type of variation) occurring; Representing an event Probability of occurrence The binary logarithm of .
[0086] Specifically, the information entropy is calculated according to the following steps:
[0087] Determine the variant type: For each variant site, the system first identifies its variant type, for example, whether it is a single nucleotide variant (SNP), insertion / deletion (Indel), or other types of genetic variants.
[0088] Calculate the probability of occurrence of mutation: Calculate the probability of occurrence of each mutation type based on the mutation frequency in the sample. This step is analyzed through statistical data to obtain the probability value of each mutation type. .
[0089] Calculate information entropy: Apply the information entropy formula to calculate the entropy value of each variant. A larger information entropy value indicates greater unpredictability of the variant site and may have higher biological significance.
[0090] After calculating the entropy value of each variant, the system will filter out key variants based on the preset entropy threshold. The threshold value is optimized based on clinical needs, the type of gene variation, and the research needs of genetic diseases. For example, an entropy threshold can be set. Variants with entropy values above this threshold are considered key variants. These key variants require further analysis and confirmation and may be pathogenic variants.
[0091] ;
[0092] in, The entropy threshold is preset, and only mutations that meet the conditions will be screened out.
[0093] Alternatively, the system can adaptively adjust the entropy threshold based on the clinical context of the variant. For example, in certain genetic diseases, certain variant types may have a higher pathogenicity. In this case, the entropy threshold can be adjusted based on historical data of that type to ensure that the selected variants are more in line with clinical needs.
[0094] Furthermore, the granularity of entropy calculations can be adjusted. In some embodiments, the entropy calculation method can be adjusted based on the region of variation (e.g., different regions of the genome). For example, variations in certain functional regions may require a higher entropy threshold to avoid filtering out irrelevant variants.
[0095] Through this step, the system automatically selects variants with high information content based on a preset entropy threshold. These variants are often clinically significant and may be closely associated with genetic diseases. The calculation of information entropy provides a quantitative assessment of each variant site, reducing human interference and making the variant screening process more objective and accurate.
[0096] Through this automatic screening and entropy analysis mechanism, the system can greatly improve the efficiency of variant screening, ensuring that subsequent pathogenicity analysis can focus on the most likely relevant variants, thereby providing strong support for clinical decision-making.
[0097] Verification of information entropy screening performance:
[0098] Experimental design:
[0099] Samples: 100 clinical samples (50 cystic fibrosis, 50 DMD), 100 healthy controls;
[0100] Methods: Traditional manual screening was performed by three people manually analyzing according to the ACMG guidelines; the method of the present invention set the entropy threshold to 0.8 to automatically screen key variants.
[0101] The test results are shown in Table 1:
[0102] Table 1:
[0103] index Traditional methods Method of the present invention Single sample processing time 3.5 hours 0.4 hours Pathogenic variant recall rate 84% 96% False screening rate of healthy samples 12% 4%
[0104] The present invention uses dynamic threshold screening based on information entropy to increase efficiency by 8.7 times, improve recall rate by 12%, and reduce false positive rate by 67%.
[0105] In step S4, the system performs Bayesian inference analysis on the key variants screened out, calculates their pathogenicity probability, and determines their pathogenicity according to the ACMG variant classification criteria.
[0106] Generally, Bayesian inference analysis combines known clinical information with genetic variant data to calculate the posterior probability of pathogenicity for each variant based on prior probabilities. This process can improve the accuracy of pathogenicity judgments by considering various types of evidence (such as functional experimental data, family history, and literature reports). Each variant is then classified according to the ACMG variant classification criteria to determine its pathogenicity.
[0107] In this embodiment, the goal of step S4 is to calculate the pathogenicity probability of the variant based on Bayesian inference and determine the pathogenicity classification of the variant according to the ACMG criteria. Through this process, the system can provide a clear pathogenicity conclusion for each variant, providing a scientific basis for clinical decision-making.
[0108] In this example, we first perform Bayesian inference analysis on the selected key variants. The core idea of Bayesian inference is to calculate the posterior probability of the variant based on known prior information and new observation data. Specifically, the Bayesian theorem formula is as follows:
[0109] ;
[0110] in, Represents a given mutation Under these conditions, the variant is pathogenic. The posterior probability of It indicates the probability of observing a variant under the pathogenicity hypothesis, usually provided by experimental data or literature reports; is the prior probability of pathogenicity, that is, the probability that the variant itself is a pathogenic variant; is the marginal probability of a variant, that is, the overall probability of observing that variant.
[0111] Bayesian inference can be used to comprehensively consider multiple lines of evidence (such as family history, phenotypic findings, and functional experiments) to calculate the posterior probability that a variant is pathogenic. The higher the posterior probability, the more likely the variant is to be considered pathogenic.
[0112] Next, the system classifies the variant into several types according to the ACMG criteria based on the calculated pathogenicity probability. The ACMG variant classification criteria divides variants into the following categories:
[0113] Pathogenic (P): Based on sufficient evidence, the variant is likely to cause disease;
[0114] Likely Pathogenic (LP): The variant has some evidence of pathogenicity, but the evidence is insufficient to determine whether it is pathogenic.
[0115] Uncertain Significance (VUS): The pathogenicity of the variant cannot be determined and more evidence is needed;
[0116] Benign (B): The variant is unlikely to cause disease and is usually a common variant;
[0117] Likely Benign (LB): There is insufficient evidence to support the pathogenicity of the variant, and the variant is more likely to be benign.
[0118] The ACMG criteria classify variants by combining different types of evidence, such as functional experiments, family studies, and literature reports. The specific classification is generally based on the following:
[0119] PVS1: Loss-of-function (LOF) variants are known to be pathogenic;
[0120] PS1: has the same amino acid change as a known pathogenic variant;
[0121] PS2: de novo mutation verified by both parents;
[0122] PS3: In vitro and in vivo experiments clearly show that the mutation leads to impaired gene function;
[0123] PS4: The frequency of the variant in the diseased population is significantly higher than that in the control population;
[0124] PM1: mutations are located in known hotspots or functional regions;
[0125] PM2: The mutation has a very low frequency in the normal population;
[0126] PM3: In recessive genetic diseases, a pathogenic or suspected pathogenic variant is detected at the trans position of the variant;
[0127] PM4: changes in protein length caused by in-frame insertions / deletions or loss of stop codons in non-repeat regions;
[0128] PM5: The amino acid change position is the same as that of the known pathogenic variant, but the mutation is different;
[0129] PM6: a novel variant without parental verification;
[0130] PP1: Evidence of co-segregation in the family;
[0131] The mutation and the disease cosegregate in the family;
[0132] PP2: A new missense variant is found in a gene where the missense variant is the cause of a disease and the proportion of benign variants in this gene is very small;
[0133] PP3: Multiple statistical methods predict that the variant will have a deleterious effect on the gene or gene product, including conservation prediction, evolutionary prediction, and splice site impact;
[0134] PP4: The phenotype or family history of the variant carrier is highly consistent with a single gene genetic disease.
[0135] BA1: When the frequency in the population is greater than 5%, it is considered a benign variant.
[0136] BS1: allele frequency is greater than disease incidence;
[0137] BS2: For early, fully penetrant disease, the variant is found in healthy adults;
[0138] BS3: variants confirmed to have no effect on protein function and splicing in in vitro and in vivo experiments;
[0139] BP1: A missense variant is found in a gene in which the cause of a disease is known to be a truncated variant.
[0140] BP3: Deletions / insertions within the repetitive region of unknown function without causing changes in the gene coding frame;
[0141] BP4: Multiple statistical methods predict that the variant will have no effect on the gene or gene product, including conservation prediction, evolutionary prediction, and splice site impact;
[0142] BP7: Synonymous variant and predicted not to affect splicing;
[0143] These types of evidence and classification criteria will be used as a decision basis in this example to help determine the pathogenicity of each variant.
[0144] Alternatively, more clinical and experimental data can be incorporated into the Bayesian inference analysis to further improve the accuracy of the pathogenicity probability. For example, more representative prior information can be obtained by combining the patient's clinical phenotypic characteristics and genotype data, or by incorporating more family history data.
[0145] In addition, in the application of ACMG standards, the system can adjust the weight or threshold according to the specific pathological type (such as single-gene genetic disease, multi-gene genetic disease), so as to more accurately adapt to the classification requirements of specific diseases.
[0146] By combining Bayesian inference with ACMG criteria, the system synthesizes various lines of evidence to provide a pathogenicity assessment for a variant. Bayesian inference analysis not only leverages experimental data, clinical data, family history, and other information, but also dynamically updates pathogenicity probabilities based on new evidence, thereby improving the accuracy of pathogenicity assessments. Combined with ACMG variant classification criteria, variant pathogenicity is subdivided into multiple levels, ensuring scientific and standardized pathogenicity classification.
[0147] This step provides a more reliable basis for the clinical diagnosis of gene mutations, helping doctors make more accurate diagnostic decisions and thus providing effective personalized treatment plans for patients with genetic diseases.
[0148] Step S5 is to evaluate these variants based on the patient's clinical phenotypic characteristics and screen out variants that are highly correlated with the patient's disease manifestations.
[0149] Typically, a patient's phenotypic information, including symptoms, medical history, and family history, is an integral part of the diagnostic process. In this embodiment, by combining phenotypic information with variant data, the system can calculate the degree of match between the variant and the patient's phenotype, further improving the accuracy of variant screening. Only when a variant closely matches the patient's phenotypic characteristics will it be considered a potential pathogenic variant and enter the subsequent verification and diagnosis stages.
[0150] In this embodiment, the goal of step S5 is to combine the patient's clinical phenotypic data and the screened variant information, calculate the matching degree between them, and thus screen out variants that match the patient's phenotypic characteristics. These variants are more likely to be the causative factors of the disease.
[0151] First, the system needs to obtain the patient's clinical phenotype information. This information can be obtained through various channels, such as the patient's medical records, family history, and clinical examination data. Phenotype information typically includes but is not limited to symptoms, signs, age of disease onset, and whether there are similar patients in the family.
[0152] The system then evaluates the relevance of each variant by calculating the degree of match between it and the patient's phenotype. To calculate the match, the system uses the following matching function:
[0153] ;
[0154] in, Indicates the matching score between variant and phenotypic information; Indicates the number of mutation information that needs to be matched; It is The weight of matching a variant feature with a phenotypic feature; It is a variation feature and phenotypic characteristics A matching function is typically used to evaluate the association of a variant with a specific symptom or disease; this function can be evaluated using Boolean matching, distance metrics, or more complex machine learning models.
[0155] For each mutation, the matching function Evaluate whether the variant is associated with a specific phenotype of the patient, such as a specific symptom or family history. Specifically, if the variant is functionally thought to cause the occurrence of certain symptoms, then it will have a higher match.
[0156] Weight This reflects the strength of the association between the phenotypic trait and the variant. Generally, symptoms or pathological conditions with greater clinical relevance are given higher weights. For example, if a variant is associated with a severe genetic symptom, the weight assigned to that variant is higher.
[0157] After calculating the matching degree of each variant After that, the system will match the To screen out variants that match the patient's phenotypic characteristics. Only when the match between the variant and the patient's phenotype is greater than or equal to the threshold, the variant will be considered a potential pathogenic variant highly associated with the disease and enter the subsequent analysis stage.
[0158] ;
[0159] As an option, the matching function Adaptive optimization can be performed based on clinical needs. In some cases, a variant may only be associated with some of a patient's symptoms, rather than all phenotypic features. In this case, the matching function can be weighted based on the importance of different symptoms. For example, a pathogenic variant may be given a higher weight if it matches a lethal symptom.
[0160] In another possible implementation, phenotypic matching could be performed using deep learning or machine learning models for more complex analysis. For example, a model could be trained to automatically identify which phenotypic features are strongly associated with specific variants. Such a model could then be used to optimize match calculations across a large amount of case data.
[0161] By calculating the match between key variants and patient phenotypic information, the system can effectively screen variants that match the patient's clinical characteristics. This process significantly improves the relevance of variant screening and avoids misidentifying unrelated variants as pathogenic. The system selects variants most likely to be associated with the patient's disease based on a match threshold, ensuring that subsequent pathogenicity analysis is more targeted and accurate.
[0162] This screening method that combines clinical phenotypic information not only improves the accuracy of variant screening, but also enhances the clinical applicability of the system, enabling more personalized diagnosis and treatment recommendations based on the characteristics of individual patients.
[0163] HJB resource optimization performance verification:
[0164] Experimental design:
[0165] Hardware: 16-core CPU / 64GB memory server;
[0166] Load scenarios: low load (10 samples), high load (100 samples), extreme load (500 samples).
[0167] The test results are shown in Table 2:
[0168] Table 2:
[0169]
[0170] The HJB resource dynamic allocation strategy increases high-load processing speed by 86%, and maintains stable operation under extreme loads.
[0171] In step S6, based on the patient's clinical phenotype information in step S5, the system screens for variants that closely match the patient's phenotype. This screening process yields a set of key variants potentially associated with the patient's disease. This variant information is then submitted to the physician for final review in step S6, generating the final screening results to inform clinical decision-making.
[0172] Typically, physician review is a further confirmation of automated screening results. Although the system utilizes a series of advanced algorithms to screen and analyze variants, physician review is crucial due to the complexity of the clinical environment and individual differences. Physicians can assess the final screening results based on a variety of information, including the pathogenicity classification of the variant, relevant evidence, and the patient's specific condition. The resulting screening results are used for further clinical diagnosis, genetic counseling, or treatment planning.
[0173] In this embodiment, the goal of step S6 is to submit the variant information screened by the system to doctors for review and generate the final screening results based on the doctors' feedback. This process requires comprehensive consideration of the accuracy of the variant information, clinical relevance, and individual patient needs.
[0174] In this embodiment, the doctor review process mainly includes the following steps:
[0175] Data display and review: The system displays the selected variant information to the doctor. The displayed information includes detailed information of each variant, pathogenicity classification (such as pathogenic, likely pathogenic, unclassified, etc.), relevant evidence (such as functional experimental data, family history, etc.), and the match between the variant and the patient's phenotype. Doctors can view the analysis results of the variant through the interface, including but not limited to:
[0176] Type of variant (SNP, Indel, etc.);
[0177] The site information of the mutation;
[0178] Probability of pathogenicity and ACMG classification;
[0179] The match score between the variant and the patient's phenotypic information.
[0180] In some embodiments, doctors can also use tools provided by the system to mark, edit or annotate variant information to ensure the accuracy and completeness of the data.
[0181] Review process: The doctor will make a final confirmation of the pathogenicity of the variant based on the data displayed by the system and the patient's specific medical history, family history and clinical manifestations. Specifically, the doctor may:
[0182] Evaluate the clinical evidence for the variant and confirm whether it is consistent with clinical experience;
[0183] The pathogenicity of the variant was determined according to the ACMG classification criteria;
[0184] Consider the patient's specific phenotypic information to determine whether the variant is consistent with the patient's clinical characteristics.
[0185] In some cases, your doctor may request further laboratory testing or additional evidence to confirm the pathogenicity of the variant.
[0186] Feedback and result generation: The physician's final review of the variant is fed back to the system, and a final screening report is generated. This report combines the variant's pathogenicity classification, relevant evidence, and clinical relevance to support subsequent clinical decision-making. The final screening results are output in a standardized format for further analysis and reference by the physician. The report may include the following:
[0187] The pathogenicity classification of the variant (e.g., pathogenic, likely pathogenic, benign, etc.);
[0188] Evidence and support for pathogenicity determination;
[0189] the relevance and match between the variant and the patient's phenotype;
[0190] Variants or samples that require further investigation.
[0191] Optionally, physicians can conduct a more detailed, personalized assessment of the selected variants during review. For example, they can adjust the pathogenicity of the variant based on the patient's age, gender, family history, and other information. Furthermore, in some embodiments, the system can provide physicians with personalized screening preferences, allowing them to adjust match thresholds or weights to better meet the diagnostic needs of specific patients.
[0192] In another possible implementation, the physician review module could integrate clinical databases or literature resources to automatically provide physicians with the latest variant-related literature and research results. This could help physicians make more accurate judgments based on more comprehensive information.
[0193] Through the physician review in step S6, the system effectively combines automated screening results with the physician's professional judgment to ensure that the final screening results meet the patient's clinical needs. Physician review not only improves the accuracy of screening results but also provides personalized decision support based on the patient's individual situation and clinical characteristics.
[0194] The resulting screening report provides standardized, detailed, and accurate variant information, providing strong support for genetic disease diagnosis, treatment plan development, and genetic counseling. This process significantly improves the efficiency and accuracy of genetic testing by combining automation and artificial intelligence with physicians' professional judgment.
[0195] Full process clinical verification:
[0196] Experimental design:
[0197] Sample: 500 patients with clinically suspected genetic diseases (covering 50 diseases).
[0198] Gold standard: Sanger sequencing + clinical confirmation results.
[0199] Process: Fully automatic analysis (steps S1-S6), with doctors reviewing system results.
[0200] The test results are shown in Table 3:
[0201] Table 3:
[0202] index result Diagnostic accuracy 98%(490 / 500) Average report generation time 5 hours (traditional process: 24 hours) Doctor review and modification rate 5% (adjust VUS annotations only)
[0203] The efficiency of the entire process has increased by 4.8 times, and the diagnostic accuracy is consistent with the gold standard.
[0204] In step S7, the key variants screened in step S6 are reviewed by a physician, generating final screening results. These results summarize the variant's pathogenicity assessment, its compatibility with the patient's phenotype, and relevant clinical evidence. Based on these final screening results, the goal of step S7 is to generate a standardized genetic testing report. This report serves as an important communication tool between physicians and patients, providing a basis for further genetic disease diagnosis, treatment selection, and genetic counseling.
[0205] Generally, the generation of genetic testing reports must adhere to certain formatting requirements to ensure the completeness, standardization, and clinical operability of the report content. Reports must not only display detailed information about the variants but also provide personalized genetic analysis and recommendations based on the patient's specific circumstances. Therefore, in this embodiment, the system can significantly improve the efficiency of report generation by automatically generating standardized reports while ensuring the accuracy and consistency of the report content.
[0206] In this embodiment, the key task of step S7 is to generate a standardized genetic testing report based on the final screening results. This report will list each screened variant in detail and provide information such as the pathogenicity classification, relevant evidence, and the degree of association with the patient's phenotype for each variant.
[0207] In this embodiment, the process of generating a standardized genetic testing report includes the following steps:
[0208] Report Template Design: The report is first formatted using a pre-set template to ensure it complies with the reporting requirements of the healthcare industry. The report template includes basic patient information, test objectives, method description, variant detection results, and more. The report design must be tailored to the needs of clinicians and comply with relevant legal and medical standards.
[0209] Data integration and generation: Based on the variant information screened in step S6 above, the system automatically fills this data into the report template. Specifically, the report includes the following parts:
[0210] Patient's basic information: such as name, age, gender, family history, etc.
[0211] Purpose and methods of testing: Briefly describe the purpose of the genetic testing and the testing methods used (e.g., high-throughput sequencing).
[0212] Variant information: Lists all filtered variant information, including the type, location, pathogenicity classification, matching degree with the patient's phenotype, relevant evidence, etc.
[0213] Pathogenicity classification: Each variant is classified based on the ACMG criteria. The system will determine the pathogenicity of each variant (e.g., pathogenic, likely pathogenic, benign, etc.) based on Bayesian inference and physician review.
[0214] Clinical Recommendations and Genetic Counseling: Based on the screening results, the system will provide clinical recommendations. For each pathogenic variant, the report will provide the potential clinical impact to assist physicians in decision-making. For patients with genetic diseases, the report may also include genetic counseling recommendations, such as genetic risk assessment and screening of family members.
[0215] Report Formatting and Output: Once the report is complete, the system automatically formats and outputs it. The report format follows a standardized medical report template, allowing physicians to use it directly in clinical practice. In some embodiments, the system also supports outputting the report in multiple formats, such as PDF or printable versions, to facilitate communication between physicians and patients and for archiving.
[0216] Automated Proofreading and Validation: To ensure report accuracy, the system automatically proofreads the data in the report. For example, the system verifies that the variant information is complete, consistent with the screening results, and that the pathogenicity classification meets legal standards. This process helps avoid errors caused by human oversight.
[0217] As an option, genetic test reports can also provide visual charts and images to help physicians more intuitively understand variant results. For example, a chart can be used to display the pathogenicity classification of each variant and its corresponding clinical impact, or a heat map can be used to show the match between the variant and the patient's phenotype. These visualization tools help physicians quickly identify key variants and improve clinical decision-making efficiency.
[0218] In another possible implementation, the report generation process could combine the patient's medical history and other test data to automatically generate personalized treatment recommendations. For example, the system might infer possible treatment options based on factors such as the patient's age, gender, and family medical history, and provide relevant recommendations in the report.
[0219] By automatically generating a standardized genetic test report in step S7, the system can significantly improve the efficiency of outputting genetic test results and ensure the accuracy, completeness, and standardization of the report content. This not only reduces the workload of manual operations but also ensures that clinicians can quickly and accurately obtain the information they need, thereby making scientifically sound diagnostic decisions.
[0220] The final report will provide detailed mutation information, pathogenicity assessment and matching with the patient's phenotype, providing patients with personalized genetic counseling and treatment recommendations, helping doctors make more accurate clinical judgments and promoting the development of personalized medicine.
[0221] See also Figure 2 The present invention also provides a data analysis system for genetic disease gene detection, comprising:
[0222] A sample collection module is used to collect patient samples and perform high-throughput sequencing to obtain raw sequencing data;
[0223] The Sample Collection Module collects biological samples from patients and performs high-throughput sequencing to obtain raw genomic data. Depending on the testing requirements, samples collected may include blood, saliva, skin tissue, and other tissues. After collection, high-throughput sequencing technology is used to read the genomic sequence and generate raw sequencing results containing large-scale genetic data. This raw data provides foundational information for subsequent variant detection, genotyping analysis, and disease diagnosis.
[0224] a data processing module, configured to perform quality control and alignment processing on the raw sequencing data, generate variation detection data, and calculate variation information based on the variation detection data;
[0225] The Data Processing Module is responsible for quality control and alignment of raw sequencing data. Quality control cleans the data and removes low-quality or erroneous reads generated during the sequencing process. Alignment compares the raw data from high-throughput sequencing to a reference genome to identify genetic variants (such as single nucleotide variants and insertions / deletions). This module not only generates variant detection data but also evaluates the characteristics of each variant based on this data, providing detailed variant information for subsequent analysis.
[0226] The mutation screening module is used to automatically screen suspicious mutation sites and calculate and analyze the amount of mutation information based on information entropy, and screen key mutations according to the preset entropy threshold;
[0227] Based on the variant detection data generated by the data processing module, the variant screening module uses automated algorithms to screen for potentially significant variants. This module utilizes techniques such as information entropy analysis to assess the unpredictability and potential significance of each variant, automatically identifying key variants that may be associated with the patient's disease. Using preset thresholds, this module ensures that the selected variants are of high clinical interest and can be further verified and analyzed.
[0228] The Bayesian inference analysis module is used to perform Bayesian inference analysis on the selected key variants, calculate their pathogenicity probability, and make pathogenicity determinations based on the ACMG variant classification criteria;
[0229] The Bayesian Inference Analysis module performs detailed statistical analysis on the identified key variants, calculating their pathogenicity probabilities. By integrating the patient's clinical background, family history, and genetic variation databases, the module infers whether each variant is pathogenic and categorizes it as "pathogenic," "likely pathogenic," or "benign." This module leverages Bayesian inference methods to provide powerful support for clinical decision-making, specifically by taking into account both prior information and emerging evidence when determining the probability of a variant's pathogenicity.
[0230] The phenotypic matching module is used to calculate the matching degree between key variants and patient phenotypic information and screen out variants that match the patient's phenotypic characteristics;
[0231] The Phenotype Matching module further assesses the relevance of variants to the patient's disease by matching the patient's phenotypic information (such as symptoms, clinical signs, and family history) with the selected variants. Based on the phenotypic characteristics, this module calculates the degree of match between each variant and the patient, screening for key variants that align with the patient's clinical presentation. For example, certain variants may be strongly associated with specific genetic disease symptoms or disease types, and the system prioritizes these variants for screening, providing physicians with more personalized diagnostic information.
[0232] The doctor review module is used to submit the screened variants to doctors for review and generate the final screening results;
[0233] The physician review module submits the results of phenotypic matching and variant screening to a physician for final review. During this review, the physician will further confirm the pathogenicity of the variant and conduct clinical verification based on the patient's specific medical history, clinical manifestations, and family history. Physicians can adjust the system's automatic classification results based on their expertise and known genetic disease data to ensure that the selected variants match the patient's actual disease manifestations. Finally, the physician-confirmed screening results will generate a report for clinical decision-making.
[0234] A report generation module is used to generate a standardized genetic testing report based on the final screening results;
[0235] The report generation module automatically generates a standardized genetic testing report based on the screening results reviewed by the physician. The report details the type of each variant screened, its pathogenicity classification, its correlation with the patient's phenotype, and the evidence supporting the pathogenicity of the variant. The report is formatted according to medical standards to ensure that the information is clear and accurate and can provide clinicians with practical genetic information. In some embodiments, the report may include personalized clinical recommendations, such as genetic risk assessment and family member screening, to help physicians develop appropriate treatment plans for patients.
[0236] See also Figure 3 The present invention also provides a computer device, comprising: a processor 11 and a memory 12, wherein the memory 12 stores a computer program executable by the processor, and when the computer program is executed by the processor, the above method is performed.
[0237] The present invention further provides a storage medium, wherein the storage medium 13 stores a computer program, and the computer program executes the above method when executed by the processor 11.
[0238] The storage medium 13 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0239] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A data analysis method for genetic disease gene detection, characterized in that: The following steps are involved: Collect patient samples and perform high-throughput sequencing to obtain raw sequencing data; performing quality control and alignment processing on the raw sequencing data to generate variation detection data, and calculating variation information based on the variation detection data; The steps for calculating the variation information based on the variation detection data include: Annotate genomic regions for variant detection data and identify functional variant sites; Calculate the variation frequency, base change type and position in the reference genome for each variant site; Calculate the information entropy value of each mutation site based on the mutation frequency and base substitution type functional impact score; The information content of the mutation site is evaluated based on the calculated information entropy value; The calculation of information volume is based on information entropy theory and uses the following formula: Where H(X) represents information entropy; n represents the number of events or mutation results; x i represents the i-th possible outcome of the random variable x; p(x i ) represents the probability of the occurrence of the i-th event or mutation type; log2p(x i ) represents event x i The probability of occurrence p(x i )'s binary logarithm; Automatically screen suspicious mutation sites and calculate and analyze the amount of mutation information based on information entropy, and screen key mutations according to the preset entropy threshold; The steps for calculating and analyzing the amount of variation information based on information entropy include: Calculate the information entropy value of each gene variant, which is used to measure the uncertainty of the variant in multiple pathogenicity classifications; According to the preset information entropy threshold, only variants with information entropy values higher than the threshold are retained, and low-information variants are eliminated; Bayesian inference analysis was performed on the key variants screened, their pathogenicity probabilities were calculated, and pathogenicity was determined based on the ACMG variant classification criteria; Combined with the patient's phenotypic information, the matching degree between the key variants and the patient's phenotypic information is calculated, and the variants that match the patient's phenotypic characteristics are screened out; The steps for performing Bayesian inference analysis on the selected key variants include: Each key variant was assigned a prior probability based on a database of known pathogenic variants; Calculate the likelihood of the variant under different evidence conditions, including at least clinical manifestations, functional effects, and family genetic information; Calculate the posterior probability by combining the prior probability with the likelihood of evidence using Bayes' theorem; Based on the calculated posterior probability, the pathogenicity category of the variant is determined and classified; Bayes' theorem formula is as follows: Where P(D|V) represents the posterior probability that a variant is pathogenic D given a variant V; P(V|D) represents the probability of observing a variant under the hypothesis of pathogenicity, which is provided by experimental data or literature reports; P(D) is the prior probability of pathogenicity, that is, the probability that the variant itself is pathogenic; P(V) is the marginal probability of a variant, that is, the overall probability of observing the variant; During the process of screening variants that match the patient's phenotypic characteristics, computing resource allocation optimization is performed based on optimal control theory. The computing resource allocation optimization includes: Maximize the information increment of the selected variant set under the condition of limited computing resources; Solve the optimal control strategy through the Hamilton-Jacobi-Bellman equation to minimize the computational cost and maximize the value of the screened variation information; Submit the screened variants to doctors for review and generate the final screening results; Generate a standardized genetic testing report based on the final screening results.
2. The data analysis method for genetic disease gene detection according to claim 1, characterized in that: The step of obtaining raw sequencing data comprises: Extract genomic DNA from patient samples; Sequence genomic DNA using a high-throughput sequencing platform to obtain raw sequencing files; Perform preliminary quality control on the original sequencing files, remove low-quality data and generate valid data files; Valid data files are converted into standardized formats and saved in a data storage system for subsequent analysis.
3. The data analysis method for genetic disease gene detection according to claim 1, characterized in that: The step of determining pathogenicity based on the ACMG variant classification criteria includes: According to the ACMG variant classification criteria, each variant is assigned a relevant evidence type, which includes at least genetic evidence, functional evidence, and family information; The pathogenicity evidence score of the variant is calculated based on the weight of each evidence type; Based on the scores of each evidence item, the pathogenicity of the variant was determined according to the ACMG criteria and divided into pathogenic, possibly pathogenic, clinically unknown, possibly benign, and benign categories; Based on the judgment results, the variant is classified and its pathogenicity conclusion is reported.
4. The data analysis method for genetic disease gene detection according to claim 1, characterized in that: The steps of submitting the selected variants to doctors for review are as follows: Generate an electronic report of the screened variant information; Submit the electronic report to the physician review system, and the physician will view and analyze the variation information through the review interface; Doctors confirm or adjust the classification of variant information based on clinical experience and patient specific conditions; The doctor's review results are fed back to the system and form the final variant classification results.
5. A data analysis system for genetic disease gene detection, applied to a data analysis method for genetic disease gene detection according to any one of claims 1 to 4, characterized in that: include: A sample collection module is used to collect patient samples and perform high-throughput sequencing to obtain raw sequencing data; a data processing module, configured to perform quality control and alignment processing on the raw sequencing data, generate variation detection data, and calculate variation information based on the variation detection data; The mutation screening module is used to automatically screen suspicious mutation sites and calculate and analyze the amount of mutation information based on information entropy, and screen key mutations according to the preset entropy threshold; The Bayesian inference analysis module is used to perform Bayesian inference analysis on the selected key variants, calculate their pathogenicity probability, and make pathogenicity determinations based on the ACMG variant classification criteria; The phenotypic matching module is used to calculate the matching degree between key variants and patient phenotypic information and screen out variants that match the patient's phenotypic characteristics; The doctor review module is used to submit the screened variants to doctors for review and generate the final screening results; The report generation module is used to generate a standardized genetic testing report based on the final screening results.
6. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data analysis method for genetic disease gene detection according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Data analysis method for genetic disease gene testing, system thereof and storage medium
CN109686439A
Microsatellite unstable site screening and analysis model construction method and device
CN110797078A
Evidence optimization method and device based on Bayesian model, computer equipment and storage medium
CN119719955A
Distributed management system and method for cloud container cluster
CN119728592A