Data analysis method and system for genetic disease gene detection and storage medium

Through automated screening and variation analysis technology, combined with Bayesian inference and patient clinical data, efficient and accurate screening of genetic testing is achieved, solving the problems of low efficiency and insufficient accuracy in the existing technology, and reducing costs through computing resource optimization and improving system performance.

CN120164524AActive Publication Date: 2025-06-17JINAN AIXIN ZHUOER MEDICAL LAB CO LTD

Patent Information

Application Number
CN202510637839.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-17
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The existing genetic testing technology is inefficient and inaccurate in variant screening, pathogenicity analysis and report generation, and the utilization of computing resources is not optimized, making it difficult to meet the needs of efficient and accurate.

Method used

Automatic screening and variation analysis technology is adopted to achieve accurate screening of variation through information entropy calculation and Bayesian inference analysis, personalized diagnosis is carried out by combining patient clinical data, and the accuracy and efficiency of reports are improved through standardized report generation module. At the same time, optimal control theory is used to optimize the allocation of computing resources.

Benefits of technology

It has achieved efficient screening of pathogenic variants, improved the accuracy of diagnosis and reporting accuracy, reduced calculation costs, and improved the overall performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164524A_ABST
    Figure CN120164524A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biomedical data analysis, and discloses a data analysis method and system for genetic disease gene detection and a storage medium, and the method comprises the following steps: collecting a patient sample, and carrying out high-throughput sequencing to obtain original data; performing quality control and comparison processing on the data to generate variation detection data, and calculating variation information amount; suspicious variation sites are automatically screened, variation information amount is analyzed based on information entropy, and key variation is screened according to a preset threshold; performing Bayesian inference analysis on the key variation, calculating pathogenicity probability, and performing pathogenicity judgment based on an ACMG standard; calculating the matching degree of the key variation and the phenotype in combination with the phenotype information of the patient, and screening the variation conforming to phenotype characteristics; submitting the screening result to a doctor for auditing, and generating a final screening result; and generating a standardized gene detection report according to a final result. According to the method, pathogenic variation is efficiently and accurately recognized, and the automation level and clinical application value of genetic disease gene detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of biomedical data analysis, and particularly to a data analysis method, system and storage medium for genetic disease gene detection. Background Art

[0002] Genetic disease gene detection is to identify potential pathogenic variants by analyzing the genomic information of patients, so as to provide a scientific basis for the early diagnosis, treatment and genetic counseling of genetic diseases. However, existing gene detection technologies have multiple deficiencies, resulting in certain limitations in clinical applications.

[0003] First of all, in the prior art, the variant screening process often relies on manual intervention. The manual method has low screening efficiency for large-scale data, and is easily affected by the experience of operators, with a certain risk of misjudgment or missed judgment. Especially with the continuous increase in the amount of genomic data, the time cost and error risk brought by manual intervention are both increasing, making it difficult to meet the requirements of high efficiency and accuracy. The screening of gene variants, pathogenicity analysis, and matching with the patient's phenotype are usually highly complex and cumbersome tasks. Existing methods often rely on manual judgment, greatly affecting the accuracy of diagnosis and the efficiency of clinical applications.

[0004] Secondly, traditional gene variant analysis mostly relies on single detection data, ignoring the combination of multi-faceted information. Especially in the clinical diagnosis of genetic diseases, simply relying on the analysis of genomic data is difficult to comprehensively consider information such as the patient's clinical phenotype and family history. This makes the variant screening process lack personalized accuracy, and some variants crucial to the patient's clinical characteristics will be missed, affecting the comprehensive diagnosis of the disease.

[0005] Furthermore, the current gene detection report generation process is relatively traditional and cumbersome. Although the technology of gene detection is constantly advancing, the report generation still relies on the judgment of doctors to confirm the final results. The report content usually depends on manual filling by doctors, which not only causes time delay, but also easily affects the final diagnosis result due to information processing errors or inconsistent formats. Although there is also automated report generation in the prior art, most systems lack standardized templates and flexible personalized customization options, making it difficult to automatically generate a suitable diagnostic report according to the specific situation of the patient, resulting in a great reduction in the accuracy and practicality of the report.

[0006] In addition, there are also deficiencies in existing computing resource management. In genetic disease gene detection, processing and analyzing a large amount of gene data requires consuming a large amount of computing resources. Existing technologies often neglect the efficient allocation of computing resources, resulting in waste of computing resources. Especially in the case of limited computing resources, it is difficult to maximize the efficiency of data analysis. This is particularly prominent in diverse application scenarios, especially in the clinical environment, where the demand for high efficiency, accuracy, and resource optimization is extremely urgent.

[0007] Therefore, the present invention proposes a data analysis method, system, and storage medium for genetic disease gene detection to solve the deficiencies of the existing technology. Summary of the Invention

[0008] Aiming at the deficiencies of the existing technology, the present invention provides a data analysis method, system, and storage medium for genetic disease gene detection, which solves the problems of low efficiency, insufficient accuracy, and unoptimized utilization of computing resources in gene mutation screening, pathogenicity analysis, and report generation in the existing technology.

[0009] To achieve the above objectives, the present invention is realized through the following technical solutions: A data analysis method for genetic disease gene detection includes the following steps: Collect patient samples and perform high-throughput sequencing to obtain raw sequencing data; Perform quality control and alignment processing on the raw sequencing data to generate variant detection data, and calculate variant information content based on the variant detection data; Automatically screen suspicious variant sites and calculate and analyze variant information content based on information entropy, and screen key variants according to a preset entropy threshold; Perform Bayesian inference analysis on the screened key variants, calculate their pathogenicity probabilities, and perform pathogenicity determination based on the ACMG variant classification criteria; Calculate the matching degree between the key variants and the patient phenotype information in combination with the patient phenotype information, and screen out the variants that conform to the patient phenotype characteristics; Submit the screened variants to a doctor for review and generate a final screening result; Generate a standardized gene detection report based on the final screening result.

[0010] The present invention also provides a data analysis system for genetic disease gene detection, including: A sample collection module for collecting patient samples and performing high-throughput sequencing to obtain raw sequencing data; A data processing module for performing quality control and alignment processing on the raw sequencing data to generate variant detection data, and calculating variant information content based on the variant detection data; A mutation screening module for automatically screening suspicious mutation sites, calculating and analyzing the mutation information content based on information entropy, and screening key mutations according to a preset entropy value threshold; A Bayesian inference analysis module for performing Bayesian inference analysis on the screened key mutations, calculating their pathogenic probabilities, and making pathogenicity judgments based on the ACMG mutation classification criteria; A phenotype matching module for calculating the matching degree between the key mutations and the patient phenotype information by combining the patient phenotype information, and screening out mutations that conform to the patient phenotype characteristics; A doctor review module for submitting the screened mutations to a doctor for review and generating a final screening result; A report generation module for generating a standardized gene detection report based on the final screening result.

[0011] The present invention also provides a storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method as described above is implemented.

[0012] The present invention provides a data analysis method, system and storage medium for genetic disease gene detection. It has the following beneficial effects: 1. The present invention adopts automated screening and mutation analysis technologies, and realizes precise screening of mutations through information entropy calculation and Bayesian inference analysis. It achieves the technical effect of efficiently screening pathogenic mutations. Compared with the mutation screening methods with more manual interventions in the prior art, it solves the problems of cumbersome screening process and easy errors. The automated process not only improves work efficiency but also reduces the incidence of human errors.

[0013] 2. The present invention combines the clinical data of patients with gene mutation information through the phenotype matching module to accurately evaluate the matching degree between mutations and phenotypes. In this way, mutations highly relevant to the clinical characteristics of patients can be screened out, achieving the technical effect of personalized diagnosis. Compared with the prior art, traditional methods fail to fully combine the phenotype information of patients, resulting in lower accuracy of mutation screening and affecting the reliability of clinical diagnosis.

[0014] 3. The present invention automatically generates a standardized gene detection report through the report generation module, ensuring unified report format, accurate content, and meeting clinical requirements. Compared with the process of manually generating reports in the prior art, the present invention reduces human operation errors, ensuring the accuracy and consistency of the reports. This standardized output significantly improves the communication efficiency of gene detection results and facilitates rapid decision-making by clinical doctors.

[0015] 4. The present invention uses optimal control theory to optimize the allocation of computing resources, ensuring that the information value of screening mutations can be maximized under the condition of limited computing resources. It achieves the effect of efficiently using computing resources. Compared with the serious waste of computing resources in the prior art, the present invention significantly reduces the computing cost through resource optimization and improves the overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a flowchart of the method of the present invention; Figure 2 is a system architecture diagram of the present invention; Figure 3 is a schematic structural diagram of a computer device of the present invention.

[0017] Among them, 10 is a computer device; 11 is a processor; 12 is a memory; 13 is a storage medium. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] Please refer to Figure 1 , the embodiments of the present invention provide a data analysis method, system and storage medium for genetic disease gene detection, including the following steps: S1. Collect patient samples and perform high-throughput sequencing to obtain raw sequencing data; S2. Perform quality control and alignment processing on the raw sequencing data to generate variant detection data, and calculate variant information content based on the variant detection data; S3. Automatically screen suspicious variant sites, calculate and analyze variant information content based on information entropy, and screen key variants according to a preset entropy threshold; S4. Perform Bayesian inference analysis on the screened key variants, calculate their pathogenic probabilities, and perform pathogenicity determination based on the ACMG variant classification standard; S5. Calculate the matching degree between the key variants and the patient phenotype information in combination with the patient phenotype information, and screen out the variants that conform to the patient phenotype characteristics; S6. Submit the screened variants to a doctor for review and generate a final screening result; S7. Generate a standardized gene detection report based on the final screening result.

[0020] For step S1, a patient sample is collected and high-throughput sequencing is performed to obtain raw sequencing data. Through high-throughput sequencing technology, large-scale genomic data can be obtained quickly and accurately, providing sufficient sample information for variant detection. The core objective of this step is to ensure the quality and integrity of the data, so as to provide a reliable basis for subsequent quality control, variant screening, and pathogenicity inference.

[0021] Generally, the patient's sample can be blood, saliva, or other human samples. Through standardized operating methods, DNA in the sample is first extracted. Then, high-throughput sequencing technology is used to sequence the extracted genomic DNA, thereby generating raw sequencing data. These data will contain information about various regions of the genome, covering possible variant sites, and can provide detailed data support for subsequent gene analysis.

[0022] In this embodiment, the specific operation of collecting patient samples is as follows: Sample collection: First, a DNA sample is collected from the patient. Usually, blood or saliva is used as the sample source. Through standardized sampling methods, the quality and representativeness of each sample are ensured. In some embodiments, the most suitable sample type may be selected in combination with the patient's clinical background.

[0023] DNA extraction: After the collected sample is lysed, genomic DNA is extracted. Usually, a commercial DNA extraction kit is used for the operation to ensure the integrity and purity of the DNA. In some embodiments, other types of extraction methods, such as phenol / chloroform extraction or silica column method, can also be used to meet different requirements.

[0024] High-throughput sequencing: The extracted DNA is processed by adapter ligation, PCR amplification, etc., and then enters the high-throughput sequencing platform for sequencing. The sequencing platforms include but are not limited to Illumina, PacBio, etc. Generally speaking, the sequencing platform uses the technology of sequencing by synthesis, which can read thousands of DNA fragments at one time. In this way, a vast amount of data covering the entire genome can be obtained.

[0025] Data acquisition process: During high-throughput sequencing, the DNA sample is cut into shorter fragments and then decoded by a high-precision sequencing device. Each DNA fragment is assigned a unique identifier, and the sequencing device combines the information of all fragments according to these identifiers to form a complete genomic data set. The raw data generated by this process is usually stored in FASTQ format and contains sequencing quality values and sequencing site information.

[0026] Quality control and data correction: After obtaining the raw data, the system will perform preliminary quality control. The quality of the sequencing results is mainly evaluated through a comparison algorithm to remove low-quality sequencing fragments. Specifically, quality control includes removing low-quality reads, removing sequencing errors, and removing contaminated data, etc. To ensure the accuracy of the data, the system will perform multiple corrections on the sequencing data during subsequent analysis processes.

[0027] Obtaining raw sequencing data: The raw data obtained through high-throughput sequencing technology contains comparison data with the reference genome. The raw data includes the sequence information of all measured base pairs, and the sequence of each data fragment can be expressed in the following form: ; where represents the th sequencing fragment, which is a DNA sequence extracted from a patient sample; represents each base in the sequencing fragment, is the first base of the fragment, is the second base, and so on until , represents the total length of the sequencing fragment.

[0028] As an option, in the selection of the sequencing platform, different sequencing depths can be selected according to the characteristics of the sample and the required coverage. For example, in clinical applications, if a higher accuracy of variant detection is required, a higher-depth sequencing scheme can be selected, usually with a coverage of 30×. In certain specific cases, a lower coverage can also be selected to optimize costs.

[0029] In a possible implementation, to improve the accuracy of sequencing, second-generation sequencing or high-precision sequencing technologies (such as PacBio SMRT sequencing technology) can also be performed to ensure accurate determination of the genomes in difficult-to-resolve regions.

[0030] For the quality control of sequencing data, the following formula is used to evaluate the quality of sequencing: ; where represents the sequencing quality score (Phred quality score), which is used to represent the reliability of the sequencing results. The higher the quality score, the higher the accuracy of the sequencing; represents the probability of sequencing error, that is, the probability that a certain base measured is incorrect, The value of is in the interval [0,1],

[0031] In the data comparison and correction stage, alignment algorithms (such as BWA, Bowtie, etc.) are used to align the sequencing data with the reference genome, and the best alignment position is determined through the shortest path algorithm.

[0032] Through the above steps, we can obtain high-quality raw sequencing data, providing accurate and reliable data support for subsequent gene variant detection. High-throughput sequencing technology can not only detect variants on a large scale and with high precision, but also capture various types of genetic variations including SNPs, Indels, structural variations, etc. Through the implementation of this step, the system can provide an accurate raw data basis for subsequent gene data analysis, providing strong technical support for the diagnosis and treatment of genetic diseases.

[0033] Generally speaking, this step ensures the collection of high-quality genomic data through a standardized operation process. At the same time, through the application of high-throughput sequencing technology, it ensures the comprehensive coverage and efficient detection of complex gene variations, laying a reliable foundation for subsequent variant screening and pathogenicity analysis.

[0034] For step S2, through quality control and alignment processing, variant detection data is generated, and the relevant information content of the variants is further calculated. The purpose of this process is to ensure the accuracy of subsequent variant screening and provide reliable data support for pathogenicity analysis.

[0035] Generally, the original sequencing data may have certain quality problems, such as sequencing errors, low-quality fragments, repetitive fragments, etc. Therefore, it is necessary to perform quality control on these data to remove invalid or low-quality data. The quality-controlled data will be aligned with the reference genome to identify possible gene variants (such as SNPs, Indels, etc.). Thereafter, the calculation of variant information content will help evaluate the potential importance of each variant, thus providing more accurate data support for subsequent analysis.

[0036] In this embodiment, the main tasks of step S2 include: performing quality control on the original sequencing data, aligning it with the reference genome, generating variant detection data, and screening out meaningful variants through information content calculation methods.

[0037] First, quality control is performed on the original sequencing data. In some embodiments, quality control mainly includes removing low-quality reads and removing sequencing errors. Specifically, for each sequencing fragment, the system will evaluate its reliability by calculating its quality score (Phred score), using the formula: ; where is the quality score; is the probability of sequencing error for this base. Generally, the quality score (i.e., the error rate is less than ) The data is regarded as high-quality data. In this way, low-quality sequencing data will be excluded to ensure the accuracy of subsequent analysis.

[0038] After removing the low-quality data, the next step is to align the quality-controlled data with the reference genome. Usually, common alignment tools (such as BWA, Bowtie, GATK, etc.) are used for genome alignment. The purpose of alignment is to identify potential gene variations, such as single nucleotide variations (SNPs), insertion / deletion variations (Indels), etc., by aligning the original sequencing data with the reference genome.

[0039] The alignment process is generally carried out through the following formula: ; where is the score representing the alignment of the entire sequencing fragment with the reference genome. This score can be used to evaluate the quality of the alignment. The higher the score, the better the matching degree between the sequencing data and the reference genome; is the length of the sequencing fragment, that is, the number of base pairs to be aligned; is the base of the th sequencing fragment; represents all the bases of this sequencing fragment; is the base at the th position in the reference genome.

[0040] After the alignment is completed, the system will output variant detection data. The variant detection data includes information on all the variant sites identified in the alignment, such as the variant type (SNP, Indel, etc.), the variant position, the variant frequency, etc. These variant sites will serve as the basic data for subsequent analysis.

[0041] Once the variant detection data is obtained, the next step is to calculate the information content of the variants. The calculation of the information content is based on the information entropy theory and usually uses the following formula: ; where represents the information entropy, also known as the information content, which measures the uncertainty of the random variable or the complexity of the information; represents the number of events or variant results, usually referring to the number of all possible variant types; represents the th possible result of the random variable (in genomic data, it may be a specific variant type or a variant occurring at a specific position); represents the probability of the th event (or variant type) occurring; Represents an event Probability of occurrence The binary logarithm of. The greater the information content of a variation, the stronger the unpredictability of the variation, which is usually considered a more significant variation. By calculating the entropy values of all variations, the system can assign an information content score to each variation.

[0042] As an option, during the quality control phase, more meticulous deduplication processing can be performed on the sequencing data to remove duplicate sequencing fragments to further improve the data quality. In addition, the coverage of the data can be increased by increasing the sequencing depth of the samples to ensure that each variant site can be accurately detected.

[0043] In the selection of the alignment algorithm, different alignment tools can be used according to specific requirements. For example, when dealing with low-complexity genomes, relatively simple alignment tools may be adopted to reduce the calculation time; while when dealing with high-complexity genomes, more accurate alignment algorithms may be selected to ensure the accuracy of the alignment results.

[0044] Through the quality control and alignment processing of the original sequencing data, the system can filter out low-quality data to ensure the accuracy of subsequent analysis. At the same time, the generation of variant detection data and the calculation of information content provide an effective screening mechanism, making the results of variant analysis more reliable. The variant information obtained in this step will provide a key basis for subsequent pathogenicity analysis, further improving the detection efficiency and accuracy of gene variants.

[0045] For step S3, suspicious variant sites are automatically screened out, and the information content of the variants is calculated and analyzed based on the entropy values of these variant sites. By screening for key variants based on a preset entropy threshold, the system can more accurately identify potential pathogenic variants. This step is a key link in the entire gene detection process and directly affects subsequent pathogenicity analysis and clinical decision support.

[0046] Generally, the screening process of variant data filters out low-information-content or irrelevant variant sites through preset thresholds. Through the calculation of information entropy, the system can evaluate the "uncertainty" of each variant site, that is, its potential clinical importance. In some embodiments, the system uses the entropy value as a way to measure the importance of variants, and the setting of the threshold can be adjusted according to specific application scenarios to achieve higher accuracy and reliability.

[0047] In this embodiment, the main objective of step S3 is to screen out key variants with clinical significance, which may be closely related to the occurrence of genetic diseases and therefore need further analysis and confirmation.

[0048] First, the system will screen based on the mutation detection data generated in the previous steps. During this process, the system will automatically detect mutation sites with high information content. Specifically, the mutation information content is evaluated by calculating the information entropy of each mutation site. The formula for information entropy is: ; where represents the information entropy, also known as the information content, which measures the uncertainty of the random variable or the complexity of the information; represents the number of events or mutation outcomes, usually referring to the number of all possible mutation types; represents the random variable 's th possible outcome (in genomic data, it could be a specific mutation type or a mutation occurring at a specific position); represents the th event (or mutation type) occurrence probability; represents the event 's occurrence probability 's binary logarithm.

[0049] Specifically, the information entropy calculation will be carried out according to the following steps: Determine the mutation type: For each mutation site, the system first identifies its mutation type. For example, whether it is a single nucleotide polymorphism (SNP), insertion - deletion (Indel), or other types of gene mutations.

[0050] Calculate the occurrence probability of the mutation: Calculate the occurrence probability of each mutation type based on the mutation frequency in the sample. This step is analyzed through statistical data to obtain the probability value of each mutation type .

[0051] Calculate the information entropy: Apply the information entropy formula to calculate the entropy value of each mutation. The larger the information entropy value, the stronger the unpredictability of the mutation site, and it may have higher biological significance.

[0052] After calculating the entropy value of each mutation, the system will screen out key mutations according to a preset entropy threshold. The setting of the threshold value is optimized according to clinical needs, the type of gene mutation, and the research needs of genetic diseases. For example, an entropy threshold can be set, and mutations with entropy values higher than this threshold are regarded as key mutations. These key mutations require further analysis and confirmation and may be pathogenic mutations.

[0053] ; where is a preset entropy threshold, and only the mutations that meet the conditions will be screened out.

[0054] As an option, the system can adaptively adjust the entropy threshold based on the clinical background of the mutations. For example, in certain specific genetic diseases, certain types of mutations may have higher pathogenicity. At this time, the entropy threshold can be adjusted according to the historical data of this type to ensure that the screened mutations are more in line with clinical needs.

[0055] In addition, the granularity of information entropy calculation can also be adjusted. In some embodiments, the way of entropy value calculation can be adjusted according to the region of the mutation (such as different regions of the genome). For example, mutations in certain functional regions may require a higher entropy threshold to avoid screening out irrelevant mutations.

[0056] Through this step, the system can automatically screen out mutations with higher information content according to the preset entropy threshold. These mutations usually have higher clinical significance and may be closely related to genetic diseases. The calculation of information entropy provides a quantitative assessment for each mutation site, reducing the interference of human factors and making the mutation screening process more objective and accurate.

[0057] Through this mechanism of automatic screening and entropy value analysis, the system can greatly improve the efficiency of mutation screening, ensure that the subsequent pathogenicity analysis can focus on the most likely relevant mutations, and thus provide strong support for clinical decision-making.

[0058] Verification of information entropy screening performance: Experimental design: Samples: 100 clinical samples (50 cases of cystic fibrosis, 50 cases of DMD), 100 healthy controls; Methods: Traditional manual screening, 3 people manually analyzed according to the ACMG guidelines; the method of the present invention, set the entropy threshold to 0.8, and automatically screen key mutations.

[0059] The test results are shown in Table 1: Table 1: Index Traditional method Method of the present invention Single-sample processing time 3.5 hours 0.4 hours Pathogenic variant recall rate 84% 96% False positive rate of healthy samples 12% 4% Through the dynamic threshold screening of information entropy of the present invention, the efficiency is increased by 8.7 times, the recall rate is increased by 12%, and the false screening rate is reduced by 67%.

[0060] For step S4, the system performs Bayesian inference analysis on the screened key mutations, calculates their pathogenicity probabilities, and determines their pathogenicity according to the ACMG mutation classification criteria.

[0061] In general, the Bayesian inference analysis method combines known clinical information and gene variant data, and calculates the posterior probability of the pathogenicity of each variant based on prior probabilities. This process can improve the accuracy of pathogenicity judgment by considering various types of evidence (such as functional experimental data, family history, literature reports, etc.). Then, in combination with the ACMG variant classification criteria, each variant is classified to obtain its pathogenicity determination result.

[0062] In this embodiment, the goal of step S4 is to calculate the pathogenicity probability of the variant based on Bayesian inference and determine the pathogenicity classification of the variant according to the ACMG standard. Through this process, the system can provide a clear pathogenicity conclusion for each variant, providing a scientific basis for clinical decision-making.

[0063] In this embodiment, first, Bayesian inference analysis is performed on the selected key variants. The core idea of Bayesian inference is to calculate the posterior probability of the variant based on known prior information and new observed data. Specifically, the Bayesian theorem formula is as follows: ; Where, represents the posterior probability that the variant belongs to pathogenic under the condition of the given variant ; represents the probability of observing the variant under the pathogenicity hypothesis, usually provided by experimental data or literature reports; is the prior probability of pathogenicity, that is, the probability that the variant itself belongs to a pathogenic variant; is the marginal probability of the variant, that is, the overall probability of observing this variant.

[0064] Through Bayesian inference, multiple pieces of evidence (such as family history, phenotypic manifestations, functional experiments, etc.) can be comprehensively considered to calculate the posterior probability that the variant is a pathogenic variant. The higher the posterior probability, the greater the likelihood that the variant is considered a pathogenic variant.

[0065] Next, the system classifies the variant into several types in the ACMG standard according to the calculated pathogenicity probability. The ACMG variant classification criteria classify variants into the following categories: Pathogenic (P): According to sufficient evidence, this variant is very likely to cause disease; Likely Pathogenic (LP): The variant has some evidence of pathogenicity, but the evidence is insufficient to determine it as pathogenic; Variant of Uncertain Significance (VUS): The pathogenicity of the variant cannot be determined and requires more evidence to support; Benign (B): The variant is unlikely to cause disease and is usually a common variant; Likely Benign (LB): There is insufficient evidence of the pathogenicity of the variant, and it is more likely to be benign.

[0066] The ACMG standard classifies variants by combining different types of evidence, such as functional experiments, family studies, literature reports, etc. The basis for specific classification usually includes: PVS1: Loss-of-function (LOF) variants are known pathogenic mechanisms; PS1: Having the same amino acid change as a known pathogenic variant; PS2: A de novo variant verified by both parents; PS3: In vitro and in vivo experiments clearly show that the variant impairs gene function; PS4: The frequency of the variant in the diseased population is significantly higher than that in the control population; PM1: The mutation is located in a known hot spot or functional region; PM2: The frequency of the mutation in the normal population is extremely low; PM3: In recessive genetic diseases, a pathogenic or suspected pathogenic variant is detected at the trans position of this variant; PM4: Protein length changes caused by in-frame insertions / deletions or loss of stop codons in non-repetitive regions; PM5: Having the same amino acid change position as a known pathogenic variant, but different variants; PM6: A de novo variant not verified by both parents; PP1: Evidence of co-segregation in the family; The mutation co-segregates with the disease in the family; PP2: A missense variant of a gene is the cause of a certain disease, and the proportion of benign variants in this gene is small. A new missense variant found in such a gene; PP3: Multiple statistical methods predict that the variant will have a harmful effect on the gene or gene product, including conservation prediction, evolutionary prediction, splicing site impact, etc.; PP4: The phenotype or family history of the variant carrier highly conforms to a certain monogenic genetic disease.

[0067] BA1: When the frequency in the population is greater than 5%, it is considered a benign variant.

[0068] BS1: Allele frequency is greater than the disease incidence; BS2: For diseases with early complete penetrance, the variant is found in healthy adults; BS3: Variants confirmed to have no effect on protein function and splicing in in vitro and in vivo experiments; BP1: A missense variant found in a gene known to cause a disease; BP3: Deletions / insertions within the repeat region of unknown function without causing changes in the gene coding frame; BP4: Multiple statistical methods predict that the variant will have no effect on the gene or gene product, including conservation prediction, evolutionary prediction, splicing site effect, etc. BP7: Synonymous variant and predicted not to affect splicing; These types of evidence and classification criteria will be used as a basis for decision making in this example to help determine the pathogenicity of each variant.

[0069] As an option, more clinical and experimental data can be introduced into the Bayesian inference analysis to further improve the accuracy of the pathogenicity probability. For example, by combining the patient's clinical phenotype characteristics and genotype data, adding more representative prior information, or combining more family history data.

[0070] In addition, in the application of ACMG standards, the system can adjust weights or thresholds according to specific pathological types (such as single-gene genetic diseases and multi-gene genetic diseases) to more accurately adapt to the classification requirements of specific diseases.

[0071] By combining Bayesian inference with ACMG standards, the system can integrate various evidences and provide pathogenicity judgments of variants. Bayesian inference analysis not only utilizes experimental data, clinical data, family history and other information, but can also dynamically update pathogenicity probabilities based on new evidence, thereby improving the accuracy of pathogenicity judgments. Combined with the ACMG variant classification standards, the pathogenicity of variants is subdivided into multiple levels, ensuring the scientificity and standardization of pathogenicity classification.

[0072] This step provides a more credible basis for the clinical diagnosis of gene mutations, helping doctors make more accurate diagnostic decisions and thus providing effective personalized treatment plans for patients with genetic diseases.

[0073] Step S5 is to evaluate these variants based on the patient's clinical phenotype characteristics and screen out variants that are highly correlated with the patient's disease manifestations.

[0074] In general, the patient's phenotypic information includes symptoms, medical history, family history, etc., which are an indispensable part of the diagnosis process. In this embodiment, by combining phenotypic information with variant data, the system can calculate the match between the variant and the patient's phenotype, further improving the accuracy of variant screening. Only when the variant is highly matched with the patient's phenotypic characteristics will it be regarded as a potential pathogenic variant and enter the subsequent verification and diagnosis stage.

[0075] In this embodiment, the goal of step S5 is to calculate the matching degree between the patient's clinical phenotype data and the selected variant information, so as to screen out the variants that conform to the patient's phenotypic characteristics, and these variants are more likely to be the pathogenic factors of the disease.

[0076] First, the system needs to obtain the patient's clinical phenotype information. This information can be obtained through various channels such as the patient's medical records, family history, and clinical examination data. Phenotype information usually includes but is not limited to symptoms, signs, the age of onset of the disease, whether there are patients with similar diseases in the family, etc.

[0077] Then, the system evaluates the relevance of each variant by calculating the matching degree between the variant and the patient's phenotype. When calculating the matching degree, the system uses the following matching function: ; where, represents the matching degree score between the variant and the phenotype information; represents the number of variant information to be matched; is the weight of the th variant feature matching the phenotype feature; is the variant feature matching the phenotype feature , and it is usually evaluated according to the degree of association between the variant and a specific symptom or disease; this function can be evaluated using boolean matching, distance metrics, or more complex machine learning models.

[0078] For each variant, the matching function evaluates whether the variant is related to the patient's specific phenotype (such as a specific symptom or family history). Specifically, if a variant is functionally considered to cause the appearance of certain symptoms, then its matching degree will be higher.

[0079] The weight reflects the strength of the association between the phenotype feature and the variant. Usually, symptoms or pathological states with higher clinical relevance will be given higher weights. For example, if a variant is associated with severe genetic symptoms, then the weight corresponding to this variant is larger.

[0080] After calculating the matching degree of each variant, the system will screen out the variants that conform to the patient's phenotypic characteristics according to a set matching degree threshold . Only when the matching degree between the variant and the patient's phenotype is greater than or equal to the threshold, the variant will be considered a potential pathogenic variant highly related to the disease and will then enter the subsequent analysis stage.

[0081] ; As an option, the matching degree function It can be adaptively optimized according to clinical needs. In some cases, the variation may only be related to some of the patient's symptoms, rather than all phenotypic characteristics. At this time, the matching function can be weighted according to the importance of different symptoms. For example, a higher weight may be given to the matching degree between the pathogenic variation and the lethal symptoms.

[0082] In another possible implementation, the matching of phenotypic information can be analyzed more complexly using deep learning or machine learning models. For example, a model can be trained to automatically identify which phenotypic characteristics have a strong association with specific variations, and such a model can optimize the matching degree calculation through a large amount of case data.

[0083] By calculating the matching degree between the key variations and the patient's phenotypic information, the system can effectively screen out the variations that match the patient's clinical characteristics. This process can significantly improve the relevance of variation screening and avoid misidentifying irrelevant variations as pathogenic variations. The system screens out the variations most likely related to the patient's disease according to the matching degree threshold, ensuring that the subsequent pathogenicity analysis is more targeted and accurate.

[0084] This screening method combined with clinical phenotypic information not only improves the accuracy of variation screening, but also enhances the clinical applicability of the system, enabling more personalized diagnosis and treatment recommendations based on the characteristics of individual patients.

[0085] HJB Resource Optimization Performance Verification: Experimental Design: Hardware: 16-core CPU / 64GB memory server; Load Scenarios: Low load (10 samples), high load (100 samples), extreme load (500 samples).

[0086] The test results are shown in Table 2: Table 2:

[0087] The HJB resource dynamic allocation strategy increases the high-load processing speed by 86% and still operates stably under extreme load.

[0088] For step S6, in the previous step S5, the system screens out the variations that highly match the patient's phenotype according to the patient's clinical phenotypic information. After the above screening, the system has obtained a set of key variations that may be related to the patient's disease. These variation information will be submitted to the doctor for final review in step S6, and the final screening results will be generated to provide a basis for clinical decision-making.

[0089] In general, the doctor's review work is a further confirmation based on the results of automated screening. Although the system has screened and analyzed mutations using a series of advanced algorithms, due to the complexity of the clinical environment and individual differences, the doctor's review step is particularly crucial. Doctors can judge the final screening results based on various information such as the pathogenicity classification of mutations, relevant evidence, and the specific condition of the patient. The finally generated screening results will be used for further clinical diagnosis, genetic counseling, or the formulation of treatment plans.

[0090] In this embodiment, the goal of step S6 is to submit the mutation information screened by the system to the doctor for review and generate the final screening results based on the doctor's feedback. This process needs to ensure the accuracy of the mutation information, clinical relevance, and comprehensive consideration of the individual needs of the patient.

[0091] In this embodiment, the process of doctor review mainly includes the following links: Data display and review: The system displays the screened mutation information to the doctor. The displayed information includes the detailed information of each mutation, pathogenicity classification (such as pathogenic, likely pathogenic, unclassified, etc.), relevant evidence (such as functional experimental data, family history, etc.), and the matching degree between the mutation and the patient's phenotype. Doctors can view the various analysis results of the mutation through the interface, including but not limited to: The type of mutation (SNP, Indel, etc.); The locus information of the mutation; The pathogenicity probability and ACMG classification; The matching degree score between the mutation and the patient's phenotype information.

[0092] In some embodiments, doctors can also use the tools provided by the system to annotate, edit, or comment on the mutation information to ensure the accuracy and integrity of the data.

[0093] Review process: Based on the data displayed by the system and combined with the patient's specific medical history, family history, and clinical manifestations, the doctor finally confirms the pathogenicity of the mutation. Specifically, the doctor may: Evaluate the clinical evidence of the mutation and confirm whether it conforms to clinical experience; Judge the pathogenicity of the mutation according to the ACMG classification standard; Consider the patient's specific phenotype information and judge whether the mutation conforms to the patient's clinical characteristics.

[0094] In some cases, doctors may require further experimental verification or additional evidence to confirm the pathogenicity of the mutation.

[0095] Feedback and Result Generation: The doctor's final review result of the variants will be fed back to the system, and a final screening report will be generated. This report will combine the pathogenicity classification of the variants, relevant evidence, and clinical relevance to support subsequent clinical decisions. The final screening results will be output in a standardized format for easy further analysis and reference by doctors. The report may include the following: The pathogenicity classification of the variants (such as pathogenic, likely pathogenic, benign, etc.); The evidence and support for the pathogenicity determination; The correlation and match degree between the variant and the patient's phenotype; Variants or samples that require further examination.

[0096] As an option, during the doctor's review, more detailed personalized evaluations can be performed on the screened variant information. For example, the doctor can adjust the pathogenicity of the variant based on information such as the patient's age, gender, and family genetic history. In addition, in some embodiments, the system can also provide personalized screening preference settings for doctors, allowing doctors to adjust the match degree threshold or weight to better meet the diagnostic needs of specific patients.

[0097] In another possible implementation, the doctor review module can integrate clinical databases or literature resources to automatically provide doctors with the latest variant-related literature and research results. This can help doctors make more accurate judgments based on more comprehensive information.

[0098] Through the doctor review in step S6, the system can effectively combine the automated screening results with the doctor's professional judgment to ensure that the final screening results meet the patient's clinical needs. The doctor review can not only improve the accuracy of the screening results but also provide personalized decision support based on the patient's individual circumstances and clinical characteristics.

[0099] The finally generated screening report provides standardized, detailed, and accurate variant information, providing strong support for the diagnosis of genetic diseases, the formulation of treatment plans, and genetic counseling. This process, through the combination of automation and artificial intelligence with the doctor's professional judgment, greatly improves the efficiency and accuracy of genetic testing.

[0100] Full-process Clinical Verification: Experimental Design: Samples: 500 clinically suspected genetic disease patients (covering 50 diseases).

[0101] Gold Standard: Sanger sequencing + clinical diagnosis results.

[0102] Process: Fully automated analysis (steps S1 - S6), doctor reviews the system results.

[0103] The test results are shown in Table 3: Table 3: Index Result Diagnostic accuracy rate 98%(490 / 500) Average report generation time 5 hours (traditional process: 24 hours) Doctor review and modification rate 5% (only adjust VUS annotations) The efficiency of the whole process is increased by 4.8 times, and the diagnostic accuracy rate is consistent with the gold standard.

[0104] For step S7, in the previous step S6, the key variants screened out are reviewed by doctors and the final screening results are generated. These results summarize the pathogenicity assessment of the variants, the matching degree with the patient phenotype, and the relevant clinical evidence. Based on these final screening results, the goal of step S7 is to generate a standardized genetic test report. This report will serve as an important communication tool between doctors and patients, providing a decision-making basis for further genetic disease diagnosis, treatment plan selection, and genetic counseling.

[0105] Generally, the generation of a genetic test report needs to follow certain format requirements to ensure the integrity, standardization, and clinical operability of the report content. The report not only needs to display the detailed information of the variants but also needs to provide personalized genetic analysis and suggestions according to the specific situation of the patient. Therefore, in this embodiment, the system can greatly improve the efficiency of report generation by automatically generating a standardized report while ensuring the accuracy and consistency of the report content.

[0106] In this embodiment, the key task of step S7 is to generate a genetic test report that meets the standardized requirements based on the final screening results. This report will list in detail each screened variant and provide information such as the pathogenicity classification of each variant, the relevant evidence, and the degree of association with the patient phenotype.

[0107] In this embodiment, the process of generating a standardized genetic test report includes the following steps: Report template design: The report is first formatted through a preset template to ensure that it meets the report format requirements of the medical industry. The report template will include basic patient information, test purpose, method description, variant test results, etc. The design of the report must consider the reading needs of clinicians and comply with relevant laws and medical norms.

[0108] Data integration and generation: According to the variant information screened out in the previous step S6, the system automatically fills this data into the report template. Specifically, the report includes the following parts: Basic patient information: Such as name, age, gender, family history, etc.

[0109] Test purpose and method: Briefly describe the purpose of the genetic test and the test method used (such as next-generation sequencing).

[0110] Variant information: List all the screened variant information, including the type, location, pathogenicity classification, matching degree with the patient phenotype, relevant evidence, etc.

[0111] Pathogenicity Classification: Based on the ACMG standard, each variant is classified. The pathogenicity of each variant will be finally classified according to the results of Bayesian inference and doctor review (such as pathogenic, likely pathogenic, benign, etc.).

[0112] Clinical Recommendations and Genetic Counseling: Based on the screening results, the system will give some clinical recommendations. For each pathogenic variant, the report will provide possible clinical impacts to help doctors make decisions. For patients with genetic diseases, the report may also include suggestions for genetic counseling, such as genetic risk assessment, screening of family members, etc.

[0113] Report Formatting and Output: After the report content is completed, the system will automatically format and output it. The format of the report follows a standardized medical report template so that doctors can directly use it in clinical practice. In some embodiments, the system also supports outputting the report in multiple formats, such as PDF or print version, to facilitate communication and archiving between doctors and patients.

[0114] Automated Proofreading and Verification: To ensure the accuracy of the report content, the system will perform automated proofreading on the data in the report. For example, the system will check whether the variant information is complete, whether it is consistent with the screening results, and whether the pathogenicity classification meets the legal standards, etc. This process helps to avoid errors caused by human negligence.

[0115] As an option, the gene test report can also provide visual charts and images to help doctors more intuitively understand the variant results. For example, charts can be used to show the pathogenicity classification of each variant and its corresponding clinical impacts, or heat maps can be used to display the matching degree between variants and patient phenotypes. These visualization tools help doctors quickly identify key variants and improve the efficiency of clinical decision-making.

[0116] In another possible implementation, the report generation process can combine the patient's historical medical records and other test data to automatically generate personalized treatment recommendations. For example, the system may speculate on possible treatment plans based on factors such as the patient's age, gender, family medical history, etc., and give relevant suggestions in the report.

[0117] By automatically generating a standardized gene test report in step S7, the system can significantly improve the output efficiency of gene test results and ensure the accuracy, integrity, and standardization of the report content. This not only reduces the workload of manual operations but also ensures that clinicians can quickly and accurately obtain the required information, thereby making scientific and reasonable diagnostic decisions.

[0118] The final generated report will provide detailed mutation information, pathogenicity assessment, and the degree of match with the patient's phenotype, providing personalized genetic counseling and treatment recommendations for the patient, helping doctors make more accurate clinical judgments, and promoting the development of personalized medicine.

[0119] Please refer to Figure 2 , the present invention also provides a data analysis system for genetic disease gene detection, including: A sample collection module for collecting patient samples and performing high-throughput sequencing to obtain raw sequencing data; The main function of the sample collection module is to collect biological samples from patients and perform high-throughput sequencing to obtain raw genomic data. According to different detection requirements, the collected samples can include blood, saliva, skin tissue, etc. The collected samples are subjected to high-throughput sequencing technology to read the genomic sequences and generate raw sequencing results containing large-scale gene data. These raw data provide basic information for subsequent mutation detection, genotype analysis, and disease diagnosis.

[0120] A data processing module for performing quality control and alignment processing on the raw sequencing data, generating mutation detection data, and calculating the mutation information content based on the mutation detection data; The data processing module is responsible for performing quality control and alignment processing on the raw sequencing data. The goal of quality control is to clean the data and remove low-quality or incorrect reads generated during the sequencing process. The alignment process aligns the raw data obtained by high-throughput sequencing with the reference genome to identify gene mutations (such as single nucleotide variations and insertions / deletions). This module not only generates mutation detection data but also evaluates the characteristics of each mutation based on these data, providing detailed mutation information for subsequent analysis.

[0121] A mutation screening module for automatically screening suspicious mutation sites, calculating and analyzing the mutation information content based on information entropy, and screening key mutations according to a preset entropy threshold; Based on the mutation detection data generated by the data processing module, the mutation screening module screens out potential meaningful mutations through an automated algorithm. This module uses techniques such as information entropy analysis to evaluate the unpredictability and potential importance of each mutation, automatically screening out key mutations that may be related to the patient's disease. Through a preset threshold, this module ensures that the screened mutations have high clinical attention value and can be further verified and analyzed.

[0122] A Bayesian inference analysis module for performing Bayesian inference analysis on the screened key mutations, calculating their pathogenicity probabilities, and making pathogenicity judgments based on the ACMG mutation classification criteria; The Bayesian inference analysis module conducts detailed statistical analysis on the selected key variants and calculates their pathogenic probabilities. By integrating the patient's clinical background, family history, and genetic variant databases, this module can deduce whether each variant is a pathogenic variant and classify it into categories such as "pathogenic", "likely pathogenic", or "benign". The module uses the Bayesian inference method to provide strong support for clinical decision-making, especially when judging the probability of variant pathogenicity, taking into account prior information and new evidence.

[0123] The phenotype matching module is used to calculate the matching degree between the key variants and the patient's phenotype information by integrating the patient's phenotype information, and screen out the variants that conform to the patient's phenotype characteristics; The phenotype matching module further evaluates the correlation between the variants and the patient's disease by matching the patient's phenotype information (such as symptoms, clinical signs, family history, etc.) with the selected variants. Based on the phenotype characteristics, this module calculates the matching degree of each variant with the patient and screens out the key variants that conform to the patient's clinical manifestations. For example, certain variants may have a strong association with specific genetic disease symptoms or disease types, and the system will give priority to screening these variants to provide more personalized diagnostic information for doctors.

[0124] The doctor review module is used to submit the selected variants to the doctor for review and generate the final screening results; The doctor review module submits the results after phenotype matching and variant screening to the doctor for final review. During the review process, the doctor will further confirm the pathogenicity of the variants based on the patient's specific medical history, clinical manifestations, and family history, and conduct clinical verification. The doctor can adjust the automatic classification results of the system according to their professional knowledge and combined with known genetic disease data to ensure that the selected variants conform to the actual disease manifestations of the patient. Finally, the screening results confirmed by the doctor will generate a report for clinical decision-making.

[0125] The report generation module is used to generate a standardized genetic testing report based on the final screening results; The report generation module automatically generates a standardized genetic testing report based on the screening results reviewed by the doctor. The report details the type, pathogenic classification, correlation with the patient's phenotype, and evidence supporting the pathogenicity of each selected variant. The report is formatted according to medical norms to ensure clear and accurate information and provide practical genetic information for clinicians. In some embodiments, the report may include personalized clinical suggestions, such as genetic risk assessment, family member screening, etc., to help doctors formulate appropriate treatment plans for patients.

[0126] Please refer to Figure 3, the present invention also provides a computer device, including: a processor 11 and a memory 12, where the memory 12 stores a computer program executable by the processor, and when the computer program is executed by the processor, the above method is executed.

[0127] The present invention also provides a storage medium. A computer program is stored on the storage medium 13, and when the computer program is run by the processor 11, the above method is executed.

[0128] Among them, the storage medium 13 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (abbreviated as SRAM), electrically erasable programmable read-only memory (abbreviated as EEPROM), erasable programmable read-only memory (abbreviated as EPROM), programmable read-only memory (abbreviated as PROM), read-only memory (abbreviated as ROM), magnetic memory, flash memory, magnetic disk or optical disc.

[0129] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A data analysis method for genetic disease gene detection, characterized in that: The following steps are involved: Collect patient samples and perform high-throughput sequencing to obtain raw sequencing data; Performing quality control and comparison processing on the raw sequencing data to generate variation detection data, and calculating variation information based on the variation detection data; Automatically screen suspicious mutation sites and calculate and analyze the amount of mutation information based on information entropy, and screen key mutations according to the preset entropy threshold; Bayesian inference analysis was performed on the selected key variants to calculate their pathogenicity probability, and pathogenicity was determined based on the ACMG variant classification criteria; Combine the patient's phenotypic information to calculate the matching degree between the key variants and the patient's phenotypic information, and screen out the variants that match the patient's phenotypic characteristics; Submit the screened variants to doctors for review and generate the final screening results; Generate a standardized genetic testing report based on the final screening results.

2. The data analysis method for genetic disease gene detection according to claim 1, characterized in that: The step of obtaining raw sequencing data comprises: Extract genomic DNA from patient samples; Sequence genomic DNA using a high-throughput sequencing platform to obtain raw sequencing files; Perform preliminary quality control on the original sequencing files, remove low-quality data and generate valid data files; Valid data files are converted into standardized formats and saved in a data storage system for subsequent analysis.

3. The data analysis method for genetic disease gene detection according to claim 1, characterized in that: The step of calculating the amount of variation information based on the variation detection data comprises: Annotate genomic regions for variant detection data and identify functional variant sites; Calculate the variation frequency, base change type and position in the reference genome for each variation site; Based on the mutation frequency and base substitution type functional impact score, the information entropy value of each mutation site is calculated; The information content of the variant site was evaluated based on the calculated information entropy value.

4. The data analysis method for genetic disease gene detection according to claim 1, characterized in that: The step of calculating and analyzing the amount of variation information based on information entropy includes: The information entropy value of each gene variant is calculated, which is used to measure the uncertainty of the variant in multiple pathogenicity classifications; According to the preset information entropy threshold, only variants with information entropy values ​​higher than the threshold are retained, and low-information variants are eliminated.

5. The data analysis method for genetic disease gene detection according to claim 1, characterized in that: The step of performing Bayesian inference analysis on the selected key variants includes: Each key variant was assigned a prior probability based on a database of known pathogenic variants; Calculate the likelihood of the variant under different evidence conditions, where the different evidence at least includes clinical manifestations, functional impact, and family genetic information; The posterior probability is calculated by combining the prior probability with the likelihood of the evidence through Bayes’ theorem; Based on the calculated posterior probability, the pathogenicity category of the variant is determined and classified.

6. The data analysis method for genetic disease gene detection according to claim 1, characterized in that: The step of determining pathogenicity based on the ACMG variation classification standard includes: According to the ACMG variant classification standard, each variant is assigned a relevant evidence type, and the relevant evidence type at least includes genetic evidence, functional evidence, and family information; The pathogenicity evidence score of the variant was calculated based on the weight of each evidence type; Based on the scores of each evidence item, the pathogenicity of the variant was determined according to the ACMG criteria and divided into pathogenic, possibly pathogenic, clinically unknown, possibly benign, and benign categories; Based on the judgment results, the variant is classified and its pathogenicity conclusion is reported.

7. The data analysis method for genetic disease gene detection according to claim 1, characterized in that: In the process of screening out the variants that meet the phenotypic characteristics of the patient, computing resource allocation optimization is performed based on optimal control theory, and the computing resource allocation optimization includes: Maximize the information increment of the selected variant set under the condition of limited computing resources; The optimal control strategy is solved by the Hamilton-Jacobi-Bellman equation to minimize the computational cost and maximize the value of the screened variation information.

8. The data analysis method for genetic disease gene detection according to claim 1, characterized in that: The steps of submitting the selected variants to doctors for review are as follows: Generate an electronic report of the screened variant information; Submit the electronic report to the doctor review system, and the doctor will view and analyze the variation information through the review interface; Doctors confirm or adjust the classification of variant information based on clinical experience and patient specific conditions; The doctor's review results are fed back to the system and form the final variant classification results.

9. A data analysis system for genetic disease gene detection, applied to a data analysis method for genetic disease gene detection as claimed in any one of claims 1 to 8, characterized in that: include: A sample collection module, used to collect patient samples and perform high-throughput sequencing to obtain raw sequencing data; A data processing module, used to perform quality control and comparison processing on the raw sequencing data, generate variation detection data, and calculate variation information based on the variation detection data; The mutation screening module is used to automatically screen suspicious mutation sites and calculate and analyze the amount of mutation information based on information entropy, and screen key mutations according to the preset entropy value threshold; The Bayesian inference analysis module is used to perform Bayesian inference analysis on the selected key variants, calculate their pathogenicity probability, and make pathogenicity determinations based on the ACMG variant classification standards; The phenotype matching module is used to calculate the matching degree between key variants and patient phenotype information in combination with the patient phenotype information, and screen out variants that meet the patient's phenotypic characteristics; The doctor review module is used to submit the screened variants to doctors for review and generate the final screening results; The report generation module is used to generate a standardized genetic testing report based on the final screening results.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, a data analysis method for genetic disease gene detection as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Method, device and terminal for detecting genome variations

    CN109074429A

  • Data analysis method for genetic disease gene testing, system thereof and storage medium

    CN109686439A

  • Analysis detection system for screening single gene hereditary disease pathogenic gene based on patient clinical symptom data and whole exome sequencing data

    CN110021364A

  • Microsatellite unstable site screening and analysis model construction method and device

    CN110797078A

  • Evidence optimization method and device based on Bayesian model, computer equipment and storage medium

    CN119719955A

Cited By

  • Identification method and device for variation-enriched region, and computer equipment

    CN120470419A

  • Method, device and computer equipment for identifying variant-enriched regions

    CN120470419B

  • Monogene disease genetic variation intelligent interpretation method, equipment and medium

    CN120998299A