Gene variation filtering and sequencing method, device and equipment

Through multi-dimensional comprehensive conditional filtering and sorting methods, the problem of inaccurate genetic variation screening in the existing technology is solved, efficient and accurate screening and sorting of variant data is achieved, and the accuracy of genomic research and clinical diagnosis is improved.

CN120496630APending Publication Date: 2025-08-15SHENZHEN LIUFENG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510150476.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing gene variant screening methods fail to effectively integrate ACMG rating, population frequency, gene Panel association and patient phenotype information, resulting in inaccurate screening results, increasing the burden and cost of manual interpretation, and affecting the accuracy of diagnosis.

Method used

Multi-dimensional comprehensive conditional filtering methods are adopted, including bioinformatics analysis quality control, population frequency filtering, phenotypic filtering, gene panel filtering, ACMG evidence rating filtering and literature search filtering. Combined with auxiliary interpretation tools, interpretation databases and phenotypic sorting, multi-dimensional priority setting and sorting of mutated data is achieved.

Benefits of technology

It improves the accuracy and efficiency of gene variant screening, reduces the burden of manual interpretation, and enhances the pertinence of variant screening and the accuracy of clinical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496630A_ABST
    Figure CN120496630A_ABST
Patent Text Reader

Abstract

The invention provides a gene variation filtering and sorting method, device and equipment. The gene variation filtering and sorting method comprises the following steps: S1, acquiring original variation data; s2, filtering by using a multi-dimensional comprehensive condition to screen the original variation data, and outputting the filtered variation data; s3, performing priority setting and sorting on the filtered variation data in the step S2 on the basis of a sorting condition of the variation data to obtain to-be-analyzed target variation data; according to the method, through the steps of biological information analysis quality control filtering, crowd frequency and phenotype filtering, gene Panel, ACMG evidence strip filtering and sorting and the like, gene variation is systematically screened and subjected to priority sorting, so that the accuracy and efficiency of variation screening are improved, the precision and comprehensiveness of gene variation screening are remarkably improved, and the method is suitable for popularization and application. And a more reliable basis is provided for genomics research and clinical diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of bioinformatics analysis technology, and more specifically, to a method, apparatus, and device for filtering and sorting gene variations. Background Art

[0002] In genomic research and clinical diagnosis, filtering and sorting gene variants are crucial steps. However, existing gene variant screening and filtering methods have the following technical problems:

[0003] Insufficient ACMG rating screening: Existing methods fail to further screen and sort variants after preliminary filtering based on ACMG evidence ratings, resulting in an excessive number of variants being transferred to manual interpretation, increasing the burden and cost of manual interpretation; Inaccurate population frequency filtering: Existing methods fail to dynamically adjust based on the frequency of variants in specific populations (such as race, region or disease groups), resulting in high-frequency variants being incorrectly retained or low-frequency variants being over-filtered, reducing the accuracy of screening results; Lack of gene panel association analysis: Existing methods fail to fully consider the association between specific gene panels and target diseases, resulting in insufficient correlation between the screened variants and clinical phenotypes, affecting the accuracy of diagnosis; Insufficient utilization of patient phenotypic information: Existing methods fail to effectively integrate patient phenotypic information (such as clinical symptoms, family history, etc.) to prioritize variants, resulting in a low match between screening results and the patient's actual condition, affecting the efficiency of clinical decision-making.

[0004] In addition, recent gene variation screening methods still suffer from the problem of lack of multidimensional filtering mechanisms. That is, existing methods usually rely on single-dimensional filtering, such as quality control or mutation detection of specific genes, and fail to comprehensively consider multidimensional information such as population frequency, gene panel, patient phenotype and ACMG rating, resulting in a lack of targeted screening conditions and inaccurate screening results.

[0005] Therefore, there is an urgent need for a multidimensional gene variation screening method that can integrate ACMG rating, population frequency, gene panel association and patient phenotypic information to solve the problems of low screening efficiency and insufficient accuracy in existing technologies.

[0006] To address the above issues, existing technologies for filtering and sorting gene variants still need to be improved. Summary of the Invention

[0007] The present application provides a method, device and equipment for filtering and sorting gene variations, which solves the technical problems of insufficient ACMG rating screening, low screening efficiency and insufficient accuracy in the prior art.

[0008] The technical solution of the present invention is achieved as follows:

[0009] In one aspect, the present invention provides a method for filtering and sorting gene variants, comprising the following steps:

[0010] S1, obtain original mutation data;

[0011] S2, filtering using multi-dimensional comprehensive conditions to screen the original variant data and output the filtered variant data;

[0012] S3, based on the sorting conditions of the variant data, priority setting and sorting are performed on the variant data filtered in step S2 to obtain target variant data to be analyzed.

[0013] On the basis of this technical solution, further, the methods for obtaining original variation data include sequencing data, user uploaded data, database imported data and manual input data.

[0014] On the basis of this technical solution, further, the multidimensional comprehensive condition filtering includes at least one of bioinformatics analysis quality control filtering, population frequency filtering, phenotype filtering, gene panel filtering, interpretation database filtering, ACMG evidence rating filtering and literature retrieval filtering.

[0015] On the basis of this technical solution, further, the sorting conditions of the variant data include at least one of auxiliary interpretation tool sorting, ACMG evidence rating sorting, interpretation database sorting, literature retrieval sorting and phenotype sorting.

[0016] On the basis of this technical solution, further, the phenotypic filtering includes the following steps:

[0017] Collect phenotypic data, describe and store phenotypes in a standardized manner;

[0018] Construct a phenotypic database and mark the source of phenotypic data;

[0019] Automated software was used for phenotypic matching;

[0020] Set filtering conditions to filter out unmatched variant data.

[0021] On the basis of this technical solution, the document retrieval and sorting further includes the following steps:

[0022] Select a database and collect relevant literature containing gene variation information;

[0023] Set screening criteria, screen and analyze the retrieved literature, and extract key information;

[0024] Perform comprehensive quality scoring on the retrieved literature, evaluate the relevance between the literature and the variant site based on the variant information, and perform relevance scoring;

[0025] Prioritize the literature based on the overall quality score and relevance score;

[0026] Evidence is output based on the key information, and the priority setting and sorting of the variant data are performed again based on the evidence.

[0027] On the basis of this technical solution, further, the priority setting and sorting include priority sorting and evidence weight sorting.

[0028] On the basis of this technical solution, further, the weight of evidence ranking is based on the highest level of evidence for variants of the same classification, and is ranked in sequence according to the level of the evidence items.

[0029] In another aspect, the present invention further provides a device for filtering and sorting gene variations, comprising:

[0030] Acquisition module, used to obtain original mutation data;

[0031] The filtering module is used to filter the original variant data in the acquisition module using multi-dimensional comprehensive conditions and output the filtered variant data;

[0032] The sorting module prioritizes the variant data filtered by the filtering module based on the sorting conditions of the variant data to obtain the target variant data to be analyzed.

[0033] In a third aspect, the present invention also provides an electronic device comprising at least one processor and a memory connected to the at least one processor, wherein the memory can execute instructions by the processor, and the instructions are executed by at least one processor so that the at least one processor can perform the method for genetic variation filtering and sorting as described in any one of the first aspects.

[0034] The method, device, and apparatus provided in this application for filtering and sorting gene variants have the following advantages over existing technologies:

[0035] This paper provides a more complex and flexible multi-dimensional, comprehensive conditional filtering method. It not only includes traditional quality control filtering and annotation, but also incorporates frequency filtering, phenotype filtering, gene panel filtering, and interpretation database filtering. These filtering conditions can be used individually or in combination to meet different research needs, thereby improving the accuracy and efficiency of variant screening.

[0036] Multi-dimensional prioritization: This study introduces multiple prioritization strategies, such as those based on auxiliary interpretation tools, interpretation databases, ACMG evidence ratings, and phenotypes. By combining data from different dimensions (e.g., clinical evidence, population frequency, functional predictions, etc.), a comprehensive assessment is performed on the filtered variant data, and each variant is assigned a priority accordingly. This ranking method can effectively identify variants that are more likely to be clinically significant, helping to accelerate the clinical decision-making process.

[0037] Automated Phenotype Matching and Sorting: This invention emphasizes the importance of automated software and standardized phenotypic vocabulary in phenotypic filtering and sorting. This enables patients or users to quickly and accurately match their phenotypic information to relevant genetic diseases, thereby enabling more personalized medical advice and services.

[0038] Integration of variant data from multiple sources: This system supports acquiring raw variant data from multiple sources, including but not limited to sequencing data, user-uploaded data, database import data, and manual data entry. This feature allows researchers and clinicians to select the most appropriate data source based on their specific circumstances, enhancing the system's flexibility and practicality. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0040] Figure 1 A flow chart of the method for filtering and sorting gene variations according to Example 1 of the present invention is shown;

[0041] Figure 2 A schematic diagram of a method for filtering and sorting 10 samples processed by WES in Example 2 of the present invention is shown;

[0042] Figure 3 A schematic diagram of a device for filtering and sorting gene variations according to an embodiment of the present invention is shown;

[0043] Figure 4 A schematic structural diagram of an electronic device in an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0044] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0045] Example 1

[0046] A method for filtering and ranking genetic variants, such as Figure 1 As shown, the following steps are included:

[0047] S1, obtain original mutation data;

[0048] Specifically, the raw variant data refers to data that needs to be processed by the variant screening, filtering, and sorting method. The sources and formats of the raw variant data are diverse, including but not limited to the following forms:

[0049] S1.1, Sequencing data: Raw sequencing data (such as FASTQ files) generated by high-throughput sequencing technologies, such as whole-genome sequencing, whole-exome sequencing, targeted sequencing, etc., and variant data (such as VCF files) generated after preliminary alignment and variant detection.

[0050] S1.2, User-uploaded data: The variation data files uploaded by users through the system interface support multiple formats (such as VCF, FASTQ, CSV, etc.) and can be automatically parsed and standardized.

[0051] S1.3, Database import data: Variant data imported from public or private databases (such as gnomAD, ClinVar, local databases), supporting batch import and real-time update.

[0052] S1.4, manually input data: variation data (such as specific gene loci, variation types, etc.) manually input by users through a graphical interface or command line tool.

[0053] S2, filtering using multi-dimensional comprehensive conditions to screen the original variant data and output the filtered variant data;

[0054] Specifically, the multidimensional comprehensive condition filtering includes at least one of bioinformatics analysis quality control filtering, population frequency filtering, phenotype filtering, gene panel filtering, interpretation database filtering, ACMG evidence rating filtering and literature retrieval filtering.

[0055] Multi-dimensional, comprehensive filtering encompasses multiple dimensions, including bioinformatics analysis quality control, population frequency, phenotype, gene panels, and interpretation databases. This multi-dimensional combination is similar to a multi-media filter, allowing for flexible adjustment of the types and combinations of filtering conditions to achieve optimal filtering results. For example, in some cases, users may only need to perform quality control filtering on variant data, while in other cases, comprehensive filtering may require a combination of multiple conditions, including population frequency and phenotype.

[0056] Furthermore, the various filtering conditions work together to gradually screen out valuable variant data. For example, quality control filtering is used to remove low-quality data, followed by population frequency filtering to remove common variants, and finally, gene panels and interpretation databases are used to identify variants associated with specific diseases.

[0057] Specifically, users can select one or more filtering conditions based on their needs and flexibly adjust the strategy to suit variant data of different sizes and types, whether it is whole genome data or specific gene panel sequencing data. It also supports user-defined rules to adjust filtering conditions according to preferences and goals to meet different application scenarios and research purposes, thus improving the flexibility of the filtering method and making it more personalized. Through multi-dimensional comprehensive condition filtering, more targeted variants and annotation results can be quickly screened out, thereby reducing the difficulty of interpretation and shortening delivery time. This filtering strategy can be widely used in genetic disease diagnosis, tumor gene testing, drug development and other fields, providing an important basis for the diagnosis, treatment and prevention of diseases.

[0058] More specifically, bioinformatics analysis quality control filtering refers to quality control filtering based on the base sequence data obtained by sequencing and annotation of variants to ensure the reliability and accuracy of the data. It specifically includes the following steps:

[0059] Raw data quality control and filtering: Quality control filtering is performed on FASTQ files to remove reads of low quality, containing adapter sequences, or with a high proportion of uncertain bases.

[0060] Sequence alignment and deduplication: Qualified reads after data filtering are mapped to the corresponding locations on the reference genome using alignment software to ensure accurate mapping to the target region. Identical duplicate sequences in the sequencing data are further counted and removed to prevent them from interfering with subsequent analysis.

[0061] Sequencing data quality control: Quality control filtering is performed on the data after sequence alignment and deduplication. Indicators include but are not limited to coverage of the target region, average depth, proportion of repeated sequences, proportion of reads aligned to the target region, sequencing depth of each base in the target region, etc.

[0062] Variant detection: Through sequence alignment, data different from the reference genome sequence is obtained, and the variants are classified into different variant types with the help of detection software. For example, GATK and SMAtools can distinguish variants into SNVs and INDELs.

[0063] Variant annotation: The process of associating variant name, type, and pathogenicity reference information based on genomic coordinate ranges, using open source software such as ANNOVAR or self-developed processes, covering basic variant information, related disease information, population frequency, disease database inclusion, and functional prediction results.

[0064] More specifically, population frequency filtering refers to screening variant sites based on base sequence data that has been filtered through bioinformatics analysis quality control to improve the specificity of the data and narrow the scope of interpretation. It specifically includes the following steps:

[0065] Select a population frequency database: First, select a frequency database suitable for the target population, such as ESP6500 (National Heart, Lung, and Blood Institute Exome Sequencing Project), 1000G (1000 Genomes Database), gnomAD (Global Population Variation Database), ExAC (Exome Aggregation Consortium Database), dbSNP (Single Nucleotide Polymorphism Data) or other databases targeting specific regions or populations; Second, calibrate the database version: ensure that the database version used is the latest and compatible with the genomic data being analyzed.

[0066] Set frequency thresholds: Set population frequency thresholds based on research or clinical needs, and combine them with allele frequency (AF) thresholds to distinguish common variants from rare variants to meet the analysis needs of different diseases.

[0067] Frequency screening: By setting a frequency threshold, common variants are excluded and rare variants are retained to screen out variants with higher disease relevance for subsequent analysis.

[0068] Output annotations: Output a list of variants and their related information (including gene, variant location (CPRA, rs, nucleotide position, amino acid position, exon or intron, etc.), allele frequency, number of homozygotes and frequency evidence, etc.), and mark high-priority variants to assist in subsequent ACMG evidence rating and clinical interpretation.

[0069] More specifically, phenotypic filtering involves combining databases, automated software, and clinical collaboration mechanisms, and using standardized terms to optimize data storage and utilization, enabling efficient comparison and analysis of patient or user-entered phenotypes with disease phenotypes caused by variant genes. Specifically, it involves the following steps:

[0070] Collect phenotypic data, describe and store the phenotype in a standardized manner; that is, collect phenotypic data input by patients or users, describe and store the phenotype using standardized phenotypic terms (such as HPO and CHPO) to ensure data consistency and comparability.

[0071] Build a phenotypic database and mark the source of phenotypic data; establish a gene-disease-phenotype database, mark the data source and version, and ensure the traceability and maintenance of the data.

[0072] Apply automated software for phenotypic matching, such as Phenolyzer and Exomiser, to improve efficiency, and sort mutations based on relevance to assist interpretation.

[0073] Set filtering conditions to filter out unmatched variant data; that is, set filtering conditions to compare and analyze the phenotypic data input by patients or users through the phenotypic database and with the help of automated software to filter out unmatched gene site data.

[0074] More specifically, gene panel filtering aims to screen variant sites using a specific gene set to improve the targeting and efficiency of analysis. It includes the following steps:

[0075] Identify gene panels: Gene panels can include sets of genes known to be associated with specific diseases, such as tumor-related gene panels (e.g., TMB panels) or gene sets for specific genetic diseases. Gene panel information may include, but is not limited to, genes, transcripts, variant information (e.g., chromosome location, nucleotide changes, amino acid changes, exons, introns, etc.), diseases (phenotypes), and inheritance patterns.

[0076] Calibrate panel versions: Ensure that the gene panel used is compatible with the genomic data being analyzed to avoid analysis errors caused by version differences (including gene, disease, and other information).

[0077] Variant site screening: Set filtering conditions to match the detected variant sites with the selected gene panel.

[0078] In a preferred embodiment, if the thalassemia gene panel is selected as the filtering condition, the variant data matching the gene panel is retained.

[0079] More specifically, interpretation database filtering refers to screening and filtering variant information of specific pathogenic types in the interpretation database to improve the accuracy and research efficiency of variant interpretation and optimize clinical diagnosis and scientific research. It specifically includes the following steps:

[0080] Identify the interpretation database: Select an interpretation database suitable for the target disease and research needs, such as ClinVar (Clinical Variation Database), HGMD (Human Gene Variation Database), ClinGen (Clinical Genomic Resource), or a locally built interpretation database. The database records information that may include, but is not limited to, genes, transcripts, variants (such as chromosome location, nucleotides, amino acids, exons, introns), diseases (phenotypes), inheritance patterns, ACMG evidence, pathogenicity classification, and references.

[0081] Calibrate the interpretation database version: Ensure that the interpretation database version used is compatible with the genomic data being analyzed to avoid analysis errors caused by version differences (including information on genes, transcripts, variants, inheritance patterns, diseases, etc.).

[0082] Variant site screening: Set filtering conditions according to the research purpose, and match the detected variant sites with the selected interpretation database for filtering.

[0083] In a preferred embodiment, if the purpose is to identify the cause of a disease, the database can be interpreted and filtered to filter out variants marked as benign, suspected benign, or benign / suspected benign in the database, while retaining variants that are pathogenic, suspected pathogenic, pathogenic / suspected pathogenic, of unknown clinical significance, or not recorded in the database.

[0084] In a preferred embodiment, if it is necessary to update the interpretation results of variants of a specific pathogenic classification in the interpretation library according to existing rules, it is possible to choose to output only the results that match the variant data included in the database, for example, only retaining pathogenic, suspected pathogenic, or pathogenic / suspected pathogenic variant data.

[0085] S3, based on the sorting conditions of the variant data, priority setting and sorting are performed on the variant data filtered in step S2 to obtain target variant data to be analyzed.

[0086] Specifically, the variant data ranking method prioritizes filtered variant data based on multi-dimensional, comprehensive criteria, including but not limited to ranking by auxiliary interpretation tools, ACMG evidence ratings, interpretation databases, and phenotypes. Through multi-dimensional, comprehensive evaluation, this setting aims to provide more instructive ranking results for clinical interpretation, thereby improving interpretation efficiency and ensuring priority analysis of key variants.

[0087] It should be noted that in specific implementations, users can select one or more conditions for sorting, such as selecting only "Assisted Interpretation Tool" for sorting, or they can select multiple sorting methods for scoring, or perform comprehensive scoring before sorting after completing all enumerated variant sorting methods. The specific implementation method or combination is uncertain.

[0088] More specifically, the assisted interpretation tool sorting aims to use existing interpretation tools to interpret and sort variant sites to improve the efficiency and accuracy of variant analysis. It specifically includes the following steps:

[0089] Identify interpretation tools: Select appropriate interpretation software (either publicly available or independently developed based on the ACMG framework, such as InterVar and Vasome). Ensure the software is up to date and fully functional. Prepare the variant data input format so that the software can accurately read and analyze it.

[0090] Software Interpretation and Evidence Generation: The selected software interprets the variant data and generates corresponding evidence and classification results. The software automatically identifies and labels the pathogenicity evidence type of the variant based on built-in algorithms and databases.

[0091] Evidence integration and classification: Integrate the evidence generated by the software with evidence from other sources (e.g., population frequency, functional data, etc.). According to the ACMG guidelines, the variant data are classified as "pathogenic", "probably pathogenic", "clinically of unknown significance", "probably benign", and "benign".

[0092] Priority setting and sorting: Priority setting and sorting include priority sorting and weight of evidence sorting.

[0093] Priority sorting: Sort by "pathogenic" > "likely pathogenic" > "clinically of unknown significance" > "likely benign" > "benign". Within the same category, sort by the strength of evidence interpreted by the software, giving priority to variants with a high level of evidence.

[0094] Weight of evidence ranking: For variants with the same pathogenicity classification, the highest level of evidence generated by the software is prioritized. For example, a variant with the highest level of evidence labeled PVS1 is ranked before a variant with the highest level of evidence labeled PM1. For variants with the same pathogenicity classification and the same highest level of evidence, the software-generated evidence number is used for ranking. For example, a variant with the highest level of evidence labeled PS1 is ranked before a variant with the highest level of evidence labeled PS2. If the evidence numbers are the same, the next highest level of evidence is used for ranking, and so on.

[0095] In preferred embodiments, variant data is interpreted using software such as InterVar or Vasome to generate an evidence rating for each variant. If some variants meet the "probable pathogenic" rating, while others are rated "clinically unsignificant" or "probably benign," these variants are assigned top priority. By prioritizing variants with a high likelihood of pathogenicity, this provides a key focus for subsequent analysis and clinical application.

[0096] In a preferred embodiment, variants that fail to meet the criteria for "pathogenic" or "suspected pathogenic" after software interpretation are eliminated or assigned a lower priority. This optimizes the screening strategy, reduces interference from invalid or low-value variants, and improves analysis efficiency and accuracy.

[0097] More specifically, interpretation database ranking aims to prioritize variant data by integrating variant interpretation information from multiple interpretation databases (such as ClinVar, HGMD, Clingen, and locally built interpretation libraries) to improve the accuracy and efficiency of variant interpretation. The specific implementation is as follows:

[0098] Determine the interpretation database: Follow the same steps as in the interpretation database filtering.

[0099] Calibrate the interpretation database version: Keep consistent with the steps in interpretation database filtering.

[0100] Prioritization and sorting:

[0101] Priority sorting: Sort by "pathogenic" > "likely pathogenic" > "clinically of unknown significance" > "likely benign" > "benign". Within the same category, variants are ranked based on their likelihood of being considered "pathogenic", with variants with a high level of evidence being prioritized.

[0102] Weight of evidence ranking: For variants in the same pathogenicity category, the highest level of evidence for the variant is prioritized, for example, variants with PVS1 evidence are ranked before variants with PM1 evidence. For variants in the same pathogenicity category with the same highest level of evidence, the highest level of evidence is ranked according to the evidence number, for example, variants with PS1 evidence are ranked before variants with PS2 evidence. If the highest level of evidence number is the same, the next highest level of evidence is ranked, and so on.

[0103] More specifically, the ACMG evidence ranking system aims to prioritize variant data by using (semi-)automated evidence and pathogenicity classification output by annotations such as population frequency, software prediction, and functional data before manually interpreting variant sites based on ACMG guidelines. The system includes the following steps:

[0104] Data collection and preparation: Collect variant data, including population frequency, software predictions, functional data, etc. Ensure the completeness and accuracy of the data for subsequent evidence annotation and classification.

[0105] Evidence annotation and classification: Use automated tools to annotate variant data to generate (semi-)automated evidence, and identify and label key evidence types according to ACMG guidelines, such as potent mutation evidence PVS1, functional evidence PM1, frequency evidence PM2, and computer prediction evidence PP3.

[0106] Evidence rating: Variant data were categorized as "pathogenic," "probably pathogenic," "clinically of unknown significance," "probably benign," or "benign" according to the pathogenicity classification criteria specified in the ACMG guidelines.

[0107] Prioritization and sorting:

[0108] Pathogenicity classification ranking: Variants in the study cohort were prioritized according to their pathogenicity classification in the order of "pathogenic" > "likely pathogenic" > "clinically of unknown significance" > "likely benign" > "benign".

[0109] Weight of evidence ranking: For a group of variants assigned to the same pathogenicity classification, the supporting evidence for pathogenicity is prioritized based on the highest level of evidence, while the reverse is true for benignity. For example, for variants with the same "pathogenic" classification, a variant with the highest level of evidence labeled PVS1 will be ranked before a variant with the highest level of evidence labeled PM1, and so on. For variants with the same pathogenicity classification and the same highest level of evidence, the ranking can continue based on the sequence number of the evidence, such as a variant with the highest level of evidence labeled PS1 will be ranked before a variant with the highest level of evidence labeled PS2. If the sequence number of the highest level of evidence is the same, the sequence number of the evidence with the next highest level of evidence will be used, and so on.

[0110] In a preferred embodiment, for pathogenicity classification and sorting, the variant dataset includes multiple variant sites (e.g., variant 1, variant 2, ..., variant 10). Evidence is (semi-)automatically output through population frequency analysis, software prediction, and functional data annotation, and the pathogenicity classification of the variants ("pathogenic," "suspected pathogenic," "clinically of unknown significance," "suspected benign," and "benign") is then achieved according to the ACMG guidelines.

[0111] In this variant dataset, the evidence rating for variants 1 and 2 meets the "pathogenic" standard, the evidence rating for variants 3 and 4 meets the "probably pathogenic" standard, the evidence rating for variants 5 and 6 meets the "clinically unknown" standard, the evidence rating for variants 7 and 8 meets the "probably benign" standard, and the evidence rating for variants 9 and 10 meets the "benign" standard. Based on the pathogenicity classification of each variant in the variant dataset, variants 1 and 2 can be assigned the first priority, variants 3 and 4 the second priority, variants 4 and 5 the third priority, variants 6 and 7 the fourth priority, variants 7 and 8 the fourth priority, and variants 9 and 10 the fifth priority. This prioritizes variant sites with a higher likelihood of pathogenicity, providing a focus for subsequent analysis and clinical application.

[0112] In a preferred embodiment, if (semi-)automatic output of evidence is achieved through population frequency analysis, software prediction, and functional data annotation, the pathogenicity classification of the variant ("pathogenic", "suspected pathogenic", "clinically of unknown significance", "suspected benign", and "benign") is then achieved according to the ACMG guidelines. If the "suspected pathogenic" category contains multiple variants (such as variant 1, variant 2, ..., variant 5), the highest level of evidence for variant 1 is PVS1, the highest level of evidence for variant 2 is PS1, the highest level of evidence for variant 3 is PM1, the highest level of evidence for variant 4 is PM2, and the highest level of evidence for variant 5 is PM4, then the output priority based on the highest level of evidence is that variant 1 takes precedence over other variants, and variant 2 takes precedence over other variants except variant 1. Among them, the highest level of evidence for variants 3, 4, and 5 is all at the PM (medium strength) level, but according to the evidence sequence number PM1>PM2>PM4, the priority of sorting according to the principle of evidence sequence under the same evidence level is variant 3>variation 4>variation 5. In summary, the ranking priority of the above mutations is mutation 1 > mutation 2 > mutation 3 > mutation 4 > mutation 5.

[0113] According to the above embodiment, the results of pathogenicity classification and weight of evidence ranking are shown in Table 1:

[0114] Table 1 ACMG evidence rating ranking

[0115]

[0116] More specifically, literature search and ranking aims to prioritize variants based on the academic literature and associated evidence associated with the variant site using automated tools without manual intervention, based on (semi-)automated evidence output from annotations such as population frequency, software predictions, and functional data. This method integrates data from multiple literature database platforms (such as PubMed, Google Scholar, LitVar, PubTator, Mastermind, etc.) to achieve efficient collection and evaluation of variant-related literature. Specifically, it includes the following steps:

[0117] Select a database and collect relevant literature containing gene variant information. When conducting a literature search, first determine the literature database platform to be used for the search. You can choose from a variety of platforms, including but not limited to PubMed, Google Scholar, LitVar, PubTator, Mastermind, etc. Next, based on the variant site information (such as gene name, variant type, functional impact, etc.), develop a precise set of search keywords to ensure the relevance and accuracy of the search results. Finally, by configuring automated search scripts or utilizing the platform API interface, regularly execute search tasks and download the latest literature data, thereby achieving efficient, accurate, and continuous literature search and data acquisition.

[0118] Screening criteria were set, and the retrieved literature was screened and analyzed to extract key information. During the literature screening phase, text mining techniques were used to conduct a preliminary screening of the retrieved literature based on pre-defined criteria (including publication year, journal impact, number of citations, and study design type). The content analysis phase then proceeded to conduct an in-depth analysis of the screened literature to extract key information, such as the impact mechanism of the variant, experimental verification methods, and clinical significance. Based on this key information, the literature was classified as highly relevant, moderately relevant, or lowly relevant.

[0119] The retrieved articles were evaluated for overall quality, and the relevance between the article and the variant site was assessed using variant information to create a relevance score. During the literature evaluation process, a comprehensive quality score was first calculated for each article based on multiple indicators (including journal impact factor, author team professional background, study sample size and diversity, etc.). Furthermore, the specific details of the variant site were considered to assess the direct relevance between the article and the variant, and a corresponding relevance score was assigned.

[0120] Literature is prioritized based on their combined quality and relevance scores. During the sorting process, literature is first sorted from high to low based on quality score. If quality scores are tied, they are further sorted based on relevance score. Literature with the same quality and relevance scores are sorted from newest to oldest based on publication date to highlight the latest research findings. Furthermore, articles with groundbreaking discoveries or widespread citations are given special tags to facilitate quick identification of important information.

[0121] After retrieving the highest-priority literature, the variant data is prioritized and sorted again. Based on the quality and number of literature output by automated variant retrieval, combined with (semi-)automated evidence output by annotations such as population frequency, software prediction, and functional data, if the classification criteria for suspected pathogenicity cannot be met under the default condition of the highest level of literature evidence output by literature retrieval, the variant can be eliminated or assigned a lower priority through optimized screening strategies.

[0122] Specifically, in a preferred embodiment, if the variant dataset contains multiple variant sites (such as variant 1, variant 2, ..., variant 10), first, the variant dataset is annotated based on population data (such as gnomAD, ESP6500, 1000G, ExAC, etc.), and evidence items such as BA1, BS1, BS2 and PM2 are output; secondly, computational prediction tools are used, such as the potent mutation prediction tool AutoPVS1, etc.; functional prediction tools PhyloP Vertebrates, PhyloP, Placental Mammals scores, GERP, SIFT, Condel, CADD, Mutation Taster, Polyphen2HVAR, dbsc SNV, etc., annotate the variant dataset and output evidence such as PVS1, PP3, BP4, and BP7. Next, use public or locally built data resources (such as Clinvar, OMIM, HGMD, etc.) to annotate the variant dataset and output evidence such as PS1, PM5, PM1, PM2, PM4, PP2, PP5, BP1, BP3, and BP6. Finally, use a literature retrieval platform to automatically obtain all relevant literature recording the variant, and automatically pre-output literature evidence based on the number of literature and literature content. If the literature content contains the term "confirmed to be ade novo variant" or a similar meaning, then literature evidence PS2 is output. Other literature evidence can also be assigned corresponding evidence items using the key words in the literature record. The results of the above embodiment are shown in Table 2.

[0123] Table 2 Literature search ranking results

[0124]

[0125] In addition, after completing the literature search and sorting process, users can further filter and sort the variant dataset based on the final results and the research objectives. For example, when constructing an interpretation database, variant sites classified as "pathogenic" or "suspected pathogenic" can be prioritized for screening and recording to improve the efficiency and specificity of database construction.

[0126] If the user's goal is to determine the cause of a disease, and if the test report only records variants classified as "pathogenic," "suspected pathogenic," and "clinically unresolved" related to the patient's phenotype, then filtering can be performed for variants classified as "benign," "suspected benign," and "clinically unresolved" that are not related to the disease. At the same time, manual interpretation can be performed based on variant priority and combined with literature priority. This method can quickly locate target variants, while lowering the threshold for literature reading and shortening interpretation time.

[0127] In a preferred embodiment, in the literature retrieval and sorting method, evidence is output based on the key information, and the variant data is prioritized and sorted again in combination with the evidence, that is, screening and sorting are performed based on pathogenicity classification and evidence type.

[0128] More specifically, phenotypic sorting aims to prioritize variant data by quantitatively evaluating the degree of match between the patient's clinical phenotype or user-input phenotypic data and the disease phenotype associated with the variant gene. The sorting criteria include but are not limited to the accuracy of phenotypic matching, the strength of evidence for the association between phenotype and gene, and the specificity of the phenotype in the disease. By combining standardized phenotypic vocabulary (such as HPO, CHPO) and automated phenotypic matching tools (such as Phenolyzer, Exomiser, etc.), the present invention can provide more targeted sorting results for clinical interpretation, thereby improving the efficiency of identifying pathogenic variants and optimizing the clinical diagnostic process.

[0129] Phenotypic data collection and standardization: consistent with the relevant steps in phenotypic filtering.

[0130] Phenotype database construction and maintenance: consistent with the relevant steps in phenotypic filtering.

[0131] Automated software application: consistent with the relevant steps in phenotypic filtering.

[0132] Phenotypic sorting: Set sorting conditions, use the phenotypic database and automated software to compare and analyze the phenotypic data entered by patients or users to score, and prioritize according to the score.

[0133] Example 2

[0134] This example includes 10 samples processed by WES (whole exome sequencing), and the variant data detected by the samples are filtered and sorted by the following steps: Figure 2 As shown:

[0135] Quality control filtering: Quality control software was used to perform quality analysis on the WES (whole exome sequencing) variant data of 10 samples. Variant data with a sequencing coverage depth of <20 (the proportion of the detection interval covered by the read sentence 20 times) were filtered out, and the following steps were continued for the remaining variant data;

[0136] Frequency filtering: For the variant data detected in the samples obtained after quality control filtering, continue to filter the variant data with a population frequency ≥ 0.05 and not recorded as "Pathogenic", "Pathogenic / Likely Pathogenic", or "Likely Pathogenic" in the Clinvar database through the population frequency database and the public database Clinvar. Continue to perform the following steps on the remaining variant data;

[0137] Phenotype filtering and sorting: Detect variant data from frequency-filtered samples, then continue to analyze the phenotypic data entered by patients or users through the phenotypic database and automated software Exomiser, filter variant data with unmatched phenotypes (e.g., patients with autism phenotypes), and continue to perform the following steps on the remaining variant data.

[0138] Gene panel filtering: Detect variant data for samples obtained after phenotypic filtering. If the detection target is a specific gene panel, filter based on whether the variant falls within the specific gene panel. Filter out variant data that is not within the specific gene panel, and continue to perform the following steps on the remaining variant data.

[0139] Interpretation database filtering and sorting: The variant data detected in the samples obtained after gene panel filtering are filtered with the help of the interpretation database. The variant data recorded as benign or suspected benign in the interpretation database are filtered out, and the following steps are continued for the remaining variant data.

[0140] Literature search, filtering, and ranking: For variant data detected in samples filtered through the interpretation database, automated tools are used to obtain literature data and generate ACMG evidence and pathogenicity classifications without manual intervention, building on (semi-)automated evidence output through annotations such as population frequency, software predictions, and functional data. Variants classified as "benign" or "suspected benign" are filtered out. For the remaining variant data, score weights are assigned based on multiple dimensions, including phenotype, interpretation database, and literature search, for comprehensive ranking. The target variant data for analysis and their priority are ultimately output.

[0141] Table 3 Comparison results between Example 3 and existing screening methods

[0142]

[0143] As can be seen from Table 3, the method of this embodiment significantly improves the analysis efficiency and effectively reduces the consumption of computing resources through the optimized filtering strategy. Specifically, this method can efficiently screen out benign and high-frequency variants, while significantly reducing the probability of false-negative pathogenic variant sites. After the optimized filtering process, the number of manual readings of literature for variant interpretation is greatly reduced, which not only saves time and energy, but also improves the accuracy of the overall processing. In addition, through the scoring and sorting function of the system, a clear priority is provided for the manual interpretation of variants, making the interpretation work more targeted and focused. This efficient filtering and sorting mechanism makes the overall processing process more accurate and efficient, and provides solid and reliable data support for subsequent clinical diagnosis.

[0144] In a preferred embodiment, a device for filtering and sorting gene variations is also provided, such as Figure 3 Shown, including:

[0145] Acquisition module, used to obtain original mutation data;

[0146] The filtering module is used to filter the original variant data in the acquisition module using multi-dimensional comprehensive conditions and output the filtered variant data;

[0147] The sorting module prioritizes the variant data filtered by the filtering module based on the sorting conditions of the variant data to obtain the target variant data to be analyzed.

[0148] In a preferred embodiment, Figure 4 As shown, an electronic device is also provided, comprising at least one processor and a memory connected to the at least one processor, wherein the memory stores instructions executable by the processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method for filtering and sorting genetic variations as described in any one of the first aspects.

[0149] Specifically, it may also include an input device and an output device. The input device is used to receive digital and character information input by the user, as well as user settings and function control signals related to the gene variation filtering and sorting device. The device supports multiple input methods, including but not limited to keyboard input, touch screen input, mouse operation or receiving remote commands through a network interface to enhance the flexibility and convenience of user interaction; the output device is used to display the results of sequence variation filtering and sorting, including but not limited to a display screen, printer or other display device. The device supports output in multiple data formats, such as text reports, graphical interface displays or transmitting the results to other devices through a network interface to meet the needs of different users. In addition, the output device can also provide visualization tools to help users more intuitively understand the screening and sorting results of variation data.

[0150] The present invention provides a method, device, and apparatus for filtering and ranking genetic variants. By combining multiple lines of evidence, including population frequency, computer predictions, and functional data, this method achieves (semi-)automated ACMG grading and classification. If, even with the highest level of literature evidence, a variant cannot be classified as "probably pathogenic," these variants can be filtered or prioritized. This information filtering and screening method, based on maximizing hypotheses based on available conditions, provides a new strategy for applications such as constructing variant interpretation databases and has broad application value.

[0151] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for filtering and sorting gene variations, characterized in that: The following steps are involved: S1, obtain original mutation data; S2, filtering using multi-dimensional comprehensive conditions to screen the original variant data and output the filtered variant data; S3, based on the sorting conditions of the variant data, priority setting and sorting are performed on the variant data filtered in step S2 to obtain target variant data to be analyzed.

2. The method for filtering and sorting gene variations according to claim 1, wherein: The methods for obtaining raw variation data include sequencing data, user uploaded data, database imported data and manual input data.

3. The method for filtering and sorting gene variations according to claim 2, characterized in that: The multidimensional comprehensive condition filtering includes at least one of bioinformatics analysis quality control filtering, population frequency filtering, phenotype filtering, gene panel filtering, interpretation database filtering, ACMG evidence rating filtering and literature search filtering.

4. The method for filtering and sorting gene variations according to claim 1, wherein: The sorting conditions of the variant data include at least one of auxiliary interpretation tool sorting, ACMG evidence rating sorting, interpretation database sorting, literature retrieval sorting and phenotype sorting.

5. The method for filtering and sorting gene variations according to claim 3, wherein: The phenotypic filtering comprises the following steps: Collect phenotypic data, describe and store phenotypes in a standardized manner; Construct a phenotypic database and mark the source of phenotypic data; Automated software was used for phenotypic matching; Set filtering conditions to filter out unmatched variant data.

6. The method for filtering and sorting gene variations according to claim 4, wherein: The document retrieval and sorting process comprises the following steps: Select a database and collect relevant literature containing gene variation information; Set screening criteria, screen and analyze the retrieved literature, and extract key information; Perform comprehensive quality scoring on the retrieved literature, evaluate the relevance between the literature and the variant site based on the variant information, and perform relevance scoring; Prioritize the literature based on the overall quality score and relevance score; Evidence is output based on the key information, and the priority setting and sorting of the variant data are performed again based on the evidence.

7. The method for filtering and sorting gene variations according to claim 4, wherein: The priority setting and sorting include priority sorting and evidence weight sorting.

8. The method for filtering and sorting gene variations according to claim 7, wherein: For variants of the same classification, the weight of evidence ranking is based on the highest level of evidence for the variant, and is ranked in order based on the level of the evidence item.

9. A device for filtering and sorting gene variations, characterized in that: include: Acquisition module, used to obtain original mutation data; The filtering module is used to filter the original variant data in the acquisition module using multi-dimensional comprehensive conditions and output the filtered variant data; The sorting module prioritizes the variant data filtered by the filtering module based on the sorting conditions of the variant data to obtain the target variant data to be analyzed.

10. An electronic device, characterized in that: The invention comprises at least one processor and a memory connected to the at least one processor, wherein the memory can be used to execute instructions by the processor, and the instructions are executed by at least one of the processors to enable the processor to perform the method for filtering and sorting gene variations as described in any one of claims 1 to 6.