A next-generation sequencing data-driven intelligent interpretation system and method for phenotypenegative genetic variations
The intelligent interpretation system for phenotypic genetic variations driven by second-generation sequencing data utilizes homologous sequence loss rate, family structure specificity, and sequence context complexity to construct a discriminant model, solving the accuracy problem of genetic variation interpretation under phenotypic conditions and achieving full-process automation and efficient variation identification.
Patent Information
- Application Number
- CN202511096709.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing genetic variation interpretation systems struggle to accurately distinguish between pathological and benign variations in the absence of clear phenotypic information, especially in asymptomatic carriers, early-risk individuals, or patients with atypical phenotypes, where the accuracy of variation identification declines.
A second-generation sequencing data-driven intelligent interpretation system for phenotypic genetic variations includes a sequencing data import terminal, a phenotypic variation annotation terminal, a rare variation structure modeling terminal, an intelligent inference and typing terminal, and a result visualization terminal. By statistically analyzing homologous sequence loss rate, family structure specificity, and sequence context complexity, a discriminant model is constructed to score pathogenicity potential and classify the discriminant results.
It achieves fully automated analysis without clinical phenotype input, improves the ability to interpret variants in unknown or subclinical individuals, enhances the sensitivity and specificity of variant identification, expands the system's applicability in latent variant analysis and preventive screening, and enhances the interpretability of variant scores and the efficiency of clinical pre-screening.
Smart Images

Figure CN120977381B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of genetic variation interpretation systems, and specifically to a second-generation sequencing data-driven intelligent interpretation system and method for phenotypic genetic variations. Background Technology
[0002] With the development of next-generation sequencing (NGS) technology, the detection of genetic variations at the genome level has become increasingly accurate and high-throughput, and is widely used in disease screening, pharmacogenomics, and personalized medicine. However, the interpretation of genetic variations in individuals without clear phenotypic indications (such as healthy carriers or asymptomatic variant carriers) still presents significant challenges. Most existing variant annotation and interpretation systems rely on existing phenotypic databases or manually generated annotation rule bases, making it difficult to accurately distinguish between pathological and benign variations when there is no prior phenotypic information.
[0003] Many genetic variation interpretation systems have been developed. Through extensive searching and reference, we found existing systems such as those disclosed in publications CN109243530A, CN104812947A, CN117373696B, and US20170286594A1. These systems generally include: a data access terminal, a variation annotation and classification terminal, and a result output terminal. The data access terminal receives sequencing result files or standardized variation data, completing data parsing and standardization. The variation annotation and classification terminal is used for variation annotation, pathogenicity classification, and visualization integration based on existing databases, phenotypic association rules, or expert systems. The result output terminal generates interpretation reports for user query or clinical review.
[0004] Because the aforementioned genetic variation interpretation systems generally rely on linear scoring models constructed from clinical phenotype information, existing database rules, or expert knowledge bases, they lack the ability to independently assess non-label information such as structural perturbations and sequence complexity in individuals without a phenotype. Therefore, they suffer from a decrease in the accuracy of variation judgment when screening for variations in asymptomatic carriers, early-risk individuals, or patients with atypical phenotypes. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of the aforementioned genetic variation interpretation systems by proposing a second-generation sequencing data-driven intelligent interpretation system and method for phenotypic genetic variations.
[0006] The present invention adopts the following technical solution:
[0007] A next-generation sequencing data-driven intelligent interpretation system for phenotypic genetic variants includes a sequencing data import terminal, a phenotypic variant annotation terminal, a rare variant structure modeling terminal, an intelligent inference and genotyping terminal, and a result visualization terminal. The sequencing data import terminal imports raw-format sequencing files and executes a preprocessing workflow to generate variant file information. The preprocessing workflow includes adapter removal, low-quality sequence filtering, alignment to a reference genome, and variant retrieval. The phenotypic variant annotation terminal performs basic annotation and non-phenotypic functional inference on each variant in the variant file information without clinical phenotype input. The rare variant structure modeling terminal is used to statistically analyze homologous sequence loss rate, family structure specificity, and sequence context complexity based on the results of basic annotation and non-phenotypic functional inference, extract non-phenotypic rare features, and construct a discriminant model; the intelligent inference and genotyping terminal is used to score the pathogenic potential of each variant in a non-phenotypic state using the discriminant model, and classify the discrimination results according to the scoring results; the result visualization terminal is used to display the original format sequencing file, the variant file information, the discriminant model, the pathogenic potential score, and the discrimination classification to support physician review and report export.
[0008] Optionally, the sequencing data import terminal includes a data reading module, a preprocessing module, and a variant retrieval module; the data reading module is used to read second-generation sequencing data files in their original format; the original format includes formats such as FASTQ, BAM, or CRAM; the preprocessing module is used to perform adapter sequence removal, low-quality read filtering, and sequence normalization operations; the variant retrieval module is used to align the preprocessed data to a reference genome and extract single nucleotide variants (SNVs) and small insertion / deletion variants (InDels) to generate variant file information; the variant file information includes VCF format variant files.
[0009] Optionally, the phenotypic variant annotation terminal includes a basic annotation module and a functional inference module; the basic annotation module is used to extract the gene location information, coding influence type and site conservation characteristics of each variant in the variant file information; the functional inference module is used to infer the functional consequences of the variant based on existing database cross-reference and heuristic rule sets without clinical phenotypic input.
[0010] Optionally, the rare variant structure modeling terminal includes a homology loss rate calculation module, a family structure specificity extraction module, a context complexity analysis module, and a discriminant model construction module. The homology loss rate calculation module is used to statistically analyze the missing information of variant sites in conserved regions of different species. The family structure specificity extraction module is used to evaluate the specific distribution of variants within families based on known family population genetic patterns. The context complexity analysis module is used to calculate the sequence complexity score, database hit density, and repetitive sequence index of the base region surrounding the variant, and construct a non-phenotypic feature vector adapted to the discriminant model. The discriminant model construction module is used to construct a discriminant model based on the non-phenotypic feature vector, the missing information, and the specific distribution.
[0011] Optionally, the intelligent inference and typing terminal includes a pathogenicity potential calculation module, a latent functional interference assessment module, and a classification judgment module; the pathogenicity potential calculation module is used to calculate a pathogenicity potential index based on the discriminant model; the latent functional interference assessment module is used to calculate the interference degree score of the variant on the latent functional area based on the sequence context complexity and functional database cross-analysis; the classification judgment module is used to classify the discrimination results according to the pathogenicity potential index and the interference degree score.
[0012] Optionally, the results visualization terminal includes a raw data display module, a variant path tracking module, and a report export module; the raw data display module is used to display the raw format sequencing file and variant file information for manual verification; the variant path tracking module is used to display the discriminant model, the pathogenicity potential score, the interference score, and the discriminant classification to support physician review; the report export module is used to integrate the discriminant results and risk scores into a standard clinical interpretation report format for access by the medical system and manual export.
[0013] A second-generation sequencing data-driven intelligent interpretation method for phenotypic genetic variations is applied to the second-generation sequencing data-driven intelligent interpretation system for phenotypic genetic variations as described in claim 6, wherein the intelligent interpretation method for phenotypic genetic variations includes:
[0014] S1: Import the raw format sequencing file and perform a preprocessing procedure to generate variant file information;
[0015] S2, without clinical phenotype input, perform basic annotation and non-phenotype functional inference on each variant in the variant file information;
[0016] S3, by statistically analyzing homologous sequence loss rate, family structure specificity and sequence context complexity, extracts non-phenotypic rare features and constructs a discriminative model;
[0017] S4. Using the discriminant model, the pathogenicity potential of each variant in the phenotypic state is scored, and the discriminant results are classified according to the scoring results.
[0018] S5 displays the original format sequencing file, the variant file information, the discrimination model, the pathogenicity potential score, and the discrimination classification to support physician review and report export.
[0019] The beneficial effects achieved by this invention are:
[0020] By setting up a sequencing data import terminal, a phenotypic variant annotation terminal, a rare variant structure modeling terminal, an intelligent inference and typing terminal, and a result visualization terminal, the system can complete a closed-loop analysis process from data import to pathogenicity determination without the input of clinical phenotypes. This facilitates fully automated processing and improves the ability to interpret variants in unknown or subclinical individuals. Consequently, it enables rapid and interpretable variant function prediction in scenarios such as genetic screening and disease risk assessment.
[0021] By dividing the sequencing data import process into a terminal into a data reading module, a preprocessing module, and a variant recall module, the system can progressively complete the decoding, quality filtering, and standardization of the raw data format. This helps ensure the accuracy and comparability of the input data, thereby improving the sensitivity and specificity of subsequent variant identification and ensuring the basic reliability of genetic variant analysis.
[0022] By setting the terminal for annotating phenotypic variations to include a basic annotation module and a functional inference module, the system can still reasonably annotate variations using gene loci, variation types, and database rules even in the absence of symptoms and phenotypic data. This helps to overcome the limitations of traditional systems that rely on phenotypes, thereby expanding the system's applicability in recessive variation analysis and preventive screening, and ultimately improving the system's generalization ability and interpretation coverage.
[0023] By setting up modules for calculating homology loss rate, extracting family structure specificity, analyzing contextual complexity, and constructing discriminant models, the rare variant structure modeling terminal can integrate non-phenotypic features from multiple dimensions to construct a multi-factor discriminant model. This is beneficial for capturing evolutionary conservation anomalies, population structure deviations, and sequence environment specificity of variants, thereby constructing a scoring mechanism with structural explanations of pathogenic potential. This, in turn, helps improve the accuracy and interpretability of label-free variant functional prediction.
[0024] By setting up a pathogenic potential calculation module, a latent functional interference assessment module, and a classification judgment module, the system can simultaneously assess the risk value of variants in both the structural damage dimension and the regulatory interference dimension. This facilitates comprehensive, nonlinear, and multidimensional variant risk classification based on multi-pathway information fusion, thereby supporting automated risk labeling and prioritization. This, in turn, promotes efficient variant interpretation and accurate support for intervention decisions.
[0025] By setting up modules for displaying raw data, tracking variant paths, and exporting reports, the results visualization terminal can comprehensively present the source of variants, processing procedures, and scoring criteria. This helps enhance physicians' transparent understanding and interpretation of the scoring process, thereby reducing the risk of clinical misinterpretation, increasing system trust, and ultimately facilitating the construction of an automated reporting output mechanism that meets clinical standards, thus improving the standardization and practicality of results delivery.
[0026] To further understand the features and technical content of the present invention, please refer to the following detailed description and drawings of the present invention. However, the drawings provided are for reference and illustration only and are not intended to limit the present invention. Attached Figure Description
[0027] Figure 1 This is a schematic diagram of the overall structure of the present invention;
[0028] Figure 2 This is a statistical distribution diagram of the interference score HFD in this invention;
[0029] Figure 3 This is a schematic diagram of the method flow for a second-generation sequencing data-driven intelligent interpretation method for phenotypic genetic variations in this invention.
[0030] Figure 4 This is a statistical distribution diagram of the interference score HFD in another embodiment of the present invention. Detailed Implementation
[0031] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can understand the advantages and effects of the present invention from the content disclosed in this specification. The present invention can be implemented or applied through other different specific embodiments, and various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the spirit of the present invention. Furthermore, the accompanying drawings of the present invention are for simple illustrative purposes only and are not depictions of actual dimensions; this is stated in advance. The following embodiments will further describe the relevant technical content of the present invention in detail, but the disclosed content is not intended to limit the scope of protection of the present invention.
[0032] Example 1: This example provides a second-generation sequencing data-driven intelligent interpretation system for phenotypic genetic variations. Combined with... Figure 1As shown, a second-generation sequencing data-driven intelligent interpretation system for phenotypic genetic variants includes a sequencing data import terminal, a phenotypic variant annotation terminal, a rare variant structure modeling terminal, an intelligent inference and genotyping terminal, and a result visualization terminal. The sequencing data import terminal imports raw-format sequencing files and performs a preprocessing workflow to generate variant file information. The preprocessing workflow includes adapter removal, low-quality sequence filtering, alignment to a reference genome, and variant retrieval. The phenotypic variant annotation terminal performs basic annotation and non-phenotypic functions on each variant in the variant file information without clinical phenotype input. The system can infer the following: the rare variant structure modeling terminal is used to statistically analyze homologous sequence loss rate, family structure specificity, and sequence context complexity based on the results of basic annotation and non-phenotypic functional inference, extract non-phenotypic rare features, and construct a discriminant model; the intelligent inference and genotyping terminal is used to score the pathogenic potential of each variant in a non-phenotypic state using the discriminant model, and classify the discrimination results according to the scoring results; the result visualization terminal is used to display the original format sequencing file, the variant file information, the discriminant model, the pathogenic potential score, and the discrimination classification to support physician review and report export.
[0033] Optionally, the sequencing data import terminal includes a data reading module, a preprocessing module, and a variant retrieval module. The data reading module is used to read second-generation sequencing data files in their original format. The original format includes formats such as FASTQ, BAM, or CRAM. FASTQ is the most original data format output by a second-generation sequencer. Each record contains a single base sequence and its corresponding sequencing quality score. The preprocessing module is used to perform adapter sequence removal, low-quality read filtering, and sequence standardization. BAM is a binary alignment result file generated after aligning the sequences in FASTQ to a reference genome. It is a binary compressed version of SAM (Sequence Alignment / Map) and is widely used to store aligned read information. CRAM is an advanced compressed version of the BAM format. By referencing the reference genome, it stores only information that differs from the reference, significantly compressing space. The variant retrieval module is used to align the preprocessed data to the reference genome and extract single nucleotide variants (SNVs) and small insertion / deletion variants (InDels) to generate variant file information. The variant file information includes VCF format variant files.
[0034] Optionally, the phenotypic variant annotation terminal includes a basic annotation module and a functional inference module; the basic annotation module is used to extract the gene location information, coding influence type and site conservation characteristics of each variant in the variant file information; the functional inference module is used to infer the functional consequences of the variant based on existing database cross-reference and heuristic rule sets without clinical phenotypic input.
[0035] Optionally, the rare variant structure modeling terminal includes a homology loss rate calculation module, a family structure specificity extraction module, a context complexity analysis module, and a discriminant model construction module. The homology loss rate calculation module is used to statistically analyze the missing information of variant sites in conserved regions of different species. The family structure specificity extraction module is used to evaluate the specific distribution of variants within families based on known family population genetic patterns. The context complexity analysis module is used to calculate the sequence complexity score, database hit density, and repetitive sequence index of the base region surrounding the variant, and construct a non-phenotypic feature vector adapted to the discriminant model. The discriminant model construction module is used to construct a discriminant model based on the non-phenotypic feature vector, the missing information, and the specific distribution.
[0036] Optionally, the intelligent inference and typing terminal includes a pathogenicity potential calculation module, a latent functional interference assessment module, and a classification judgment module. The pathogenicity potential calculation module calculates a pathogenicity potential index based on the discriminant model. The latent functional interference assessment module calculates the interference score of variations on latent functional regions based on sequence context complexity and functional database cross-analysis. A latent functional region refers to a region in the genome that does not directly participate in protein coding or cause domain destruction, but has potential regulatory functions in transcription, splicing, and expression regulation. This region typically includes promoters, enhancers, splice donor / receptor neighborhoods, untranslated regions (UTRs), non-coding RNA transcription regions, highly conserved sequence regions, and functionally enriched but low-repetition sequence context regions. The classification judgment module classifies the discrimination results based on the pathogenicity potential index and the interference score.
[0037] When the pathogenicity potential calculation module calculates the pathogenicity potential index, the following formula is satisfied:
[0038] ;
[0039] ;
[0040] Among them, DPI i S represents the pathogenicity potential index of the i-th variant; ij This indicates the degree of local alignment sequence perturbation caused by the i-th mutation located in the j-th functional domain. The larger the value, the greater the structural displacement or alignment loss in that region after mutation. It is obtained by comparing the sequence perturbation distance between the mutated sequence and the reference sequence after domain alignment. m represents the total number of functional domains affected by the mutation, which is obtained from a reference protein database. This represents the score for the difference in chemical properties corresponding to the i-th mutation located in the j-th functional domain. The difference in chemical properties can be, for example, a change from polar to nonpolar or a change in hydrophobicity. The score is estimated by looking up a table based on the comparison of the chemical properties of the amino acids before and after the mutation (e.g., hydrophobicity, charge changes). Specifically, the mutation site is first analyzed using the Ensembl VEP tool to determine the amino acid type before and after the mutation. Then, based on the amino acid physicochemical property dictionary, the corresponding hydrophilicity value, charge state, and relative volume value are extracted, and the difference in physicochemical properties is calculated using the Euclidean distance formula. The data sources for the property dictionary include the Kyte-Doolittle water solubility scale, the standard charge classification table, and standardized volume data from the Uniprot database. i R represents the maximum observed frequency of the i-th variant in a population database (such as gnomAD); max β represents the maximum background frequency limit set by the system (e.g., 1%), used for unified normalization; β represents the rarity modulator, the rarer the variant, the higher its pathogenic potential index growth rate.
[0041] L j This represents the length of the j-th functional domain, which is the total number of amino acids contained within the functional domain. The conservation score represents the p-th site corresponding to the j-th functional domain of the i-th variation; the p-th site represents the position of the p-th amino acid. This represents the conservation score of the p-th site in the j-th functional domain in the unmutated state; the conservation score is obtained through the following steps: preparing the input sequence; multiple sequence alignment; submitting to the ConSurf server to calculate the conservation score; and waiting for the server to return the score result.
[0042] When the pathogenicity potential index (DPI) i When the value is greater than the set threshold T1, the system will identify the variant as a high-risk variant that is "significantly structurally disturbed and rare in the population" and automatically trigger the following processing mechanisms: including generating a highlighted risk annotation label, suggesting to enter the functional verification process, prompting the possibility of family verification, and including it in the suspected pathogenic variant candidate database. At the same time, it will provide scoring details and clinical assistance prompts through the visual terminal to assist physicians in making accurate interpretations and intervention decisions.
[0043] Formula calculation example:
[0044] The following is a calculation example of the pathogenicity potential index mentioned above. The units of the known conditions in the calculation example correspond to the values that should be entered for each parameter when using the formula. Entering the parameter values in the corresponding units will improve the accuracy of the formula.
[0045] Given conditions: m=3, S 11 =0.9, S 12=0.3、S 13 =0.1、δ 11 =2.0、δ 12 =0.5、δ 13 =1.2, R1=0.0005, R max =0.01, β=1.5, T1=0.25. Substituting these values, we get:
[0046] ;
[0047] Since DPI1≈0.287>T1, the system classifies this variant as a high-risk variant characterized by "significant structural perturbation and rarity in the population," automatically triggering the aforementioned processing mechanism. Combined with... Figure 2 As shown, Figure 2 The statistical distribution chart of the pathogenicity potential index shows the distribution of DPI values calculated for 50 genetic variations, with the DPI values mainly concentrated in the range of 0.2 to 0.6.
[0048] Optionally, the results visualization terminal includes a raw data display module, a variant path tracking module, and a report export module; the raw data display module is used to display the raw format sequencing file and variant file information for manual verification; the variant path tracking module is used to display the discriminant model, the pathogenicity potential score, the interference score, and the discriminant classification to support physician review; the report export module is used to integrate the discriminant results and risk scores into a standard clinical interpretation report format for access by the medical system and manual export.
[0049] A second-generation sequencing data-driven intelligent interpretation method for phenotypic genetic variations is applied to the second-generation sequencing data-driven intelligent interpretation system for phenotypic genetic variations as described in claim 6, combined with... Figure 3 As shown, the intelligent interpretation method for phenotypic genetic variation includes:
[0050] S1: Import the raw format sequencing file and perform a preprocessing procedure to generate variant file information;
[0051] S2, without clinical phenotype input, perform basic annotation and non-phenotype functional inference on each variant in the variant file information;
[0052] S3, by statistically analyzing homologous sequence loss rate, family structure specificity and sequence context complexity, extracts non-phenotypic rare features and constructs a discriminative model;
[0053] S4. Using the discriminant model, the pathogenicity potential of each variant in the phenotypic state is scored, and the discriminant results are classified according to the scoring results.
[0054] S5 displays the original format sequencing file, the variant file information, the discrimination model, the pathogenicity potential score, and the discrimination classification to support physician review and report export.
[0055] In summary, by configuring a sequencing data import terminal, a phenotypic variant annotation terminal, a rare variant structure modeling terminal, an intelligent inference and genotyping terminal, and a result visualization terminal, the system can complete a closed-loop analysis process from data import to pathogenicity determination without clinical phenotypic input. This facilitates fully automated processing and improves the ability to interpret variants in unknown or subclinical individuals. Furthermore, by dividing the sequencing data import terminal into a data reading module, a preprocessing module, and a variant retrieval module, the system can progressively decode, filter, and standardize raw data, ensuring the accuracy and comparability of the input data. This improves the sensitivity and specificity of subsequent variant identification. By configuring the non-phenotypic variant annotation terminal to include a basic annotation module and a functional inference module, the system can still reasonably annotate variants using gene loci, variant types, and database rules even in the absence of symptoms and phenotypic data. This helps to overcome the limitations of traditional systems that rely on phenotypes, thereby expanding the system's applicability in recessive variant analysis and preventive screening. By setting up modules for homology loss rate calculation, family structure specificity extraction, context complexity analysis, and discriminant model construction, the rare variant structure modeling terminal can integrate non-phenotypic features from multiple dimensions. Constructing a multi-factor discriminant model is beneficial for capturing evolutionary conservation anomalies, population structure deviations, and sequence-environment specificities of variants, thereby enabling the development of a scoring mechanism with a structural explanation of pathogenic potential. By setting up modules for pathogenic potential calculation, latent functional interference assessment, and classification, the system can simultaneously evaluate the risk values of variants in both structural disruption and regulatory interference dimensions. This facilitates comprehensive, non-linear, and multi-dimensional variant risk classification based on multi-pathway information fusion, supporting automated risk labeling and prioritization. Finally, by including modules for raw data display, variant path tracking, and report export, the results are visualized. The platform can comprehensively present the source of variants, processing procedures, and scoring criteria, which helps enhance physicians' transparent understanding and interpretation of the scoring process, thereby reducing the risk of clinical misinterpretation and increasing system trust. By constructing a nonlinear pathogenic potential index calculation model that combines the intensity of structural perturbation with the frequency of variants in the population, it is possible to quantitatively assess the potential destructiveness of each variant in the functional structural domain in the absence of phenotypic information. This allows for the identification of variants that cause serious perturbation to conserved structural domains and are extremely rare in the population, thereby improving the ability to identify potential pathogenic variants in phenotyped individuals and significantly enhancing the interpretive transparency of variant scores and the efficiency of clinical pre-screening.
[0056] Example 2: This example includes all the content of Example 1 and provides a second-generation sequencing data-driven intelligent interpretation system for phenotypic genetic variations. The recessive functional interference assessment module is used to calculate the interference score of the variation on the recessive functional region based on the cross-analysis of sequence context complexity and functional database.
[0057] When the latent function interference assessment module calculates the interference score, the following formula is satisfied:
[0058] ;;
[0059] Among them, HFD i Z represents the interference score of the i-th variant on the recessive functional region; the higher the sequence complexity, the higher the database overlap, and the lower the frequency of repeated sequences, the greater the potential functional interference risk of the variant; i C represents the normalization factor, specifically the maximum value in the sequence complexity score; i+l represents the sequence complexity score of site i+l within the nucleotide sequence of the i-th variant within the window range, measuring whether the region is a low-complexity repetitive region; higher complexity indicates greater functional potential; l represents the context offset distance relative to the variant site of the i-th variant (i.e., the relative position index from the center site); w represents the window range; θ represents the neighborhood decay parameter, based on empirical presets, which users or administrators can manually adjust based on protein domain characteristics or research needs; P i+l The score representing the frequency of repeat sequences at site i+l; M i+l The function database hit count for site i+l indicates how many function database annotations mark the base at that site as a "functional region"; the sequence complexity score C i+l Base regions are extracted using a sliding window and scored using Shannon entropy or Dustmasker; the neighborhood decay parameter θ is set by the administrator based on experience and is used to construct spatial decay weights; the repeat sequence frequency score P... i+l Based on the results of reviewing RepeatMasker annotations, RepeatMasker is a widely used bioinformatics tool for identifying, annotating, and masking repetitive elements in genomic sequences; the functional database hit count M i+l The number of functional evidence hits for each locus is calculated based on functional genomic databases such as ENCODE and RegulomeDB; all parameters can be automatically extracted by scripts and used as inputs into the interference scoring formula.
[0060] When the interference score HFD exceeds the set threshold T2, the system determines the variant as a candidate site that "may be located in a key region of functional regulation and has a strong risk of non-phenotypic interference". It automatically triggers mechanisms such as visual warnings, functional annotation tracing, experimental suggestions and manual interpretation prompts to help genetic experts identify recessive variants that appear benign but may actually cause disease, and improve the breadth and accuracy of interpreting variants in non-coding regions.
[0061] Formula calculation example:
[0062] The following is a calculation example of the interference score mentioned above. The units of the known conditions in the calculation example correspond to the input values of each parameter when using the formula. Inputting the parameter values in the corresponding units will improve the accuracy of the formula.
[0063] Given conditions: i=1, w=1, C i-1 =1.2、C i =1.8, C i+1 =1.5, M i-1 =3、M i =7、M i+1 =1、P i-1 =0.3, P i =0.1, P i+1 =0.8, θ=0.5, T2=0.9; substituting these values into the calculation yields:
[0064] ;
[0065] Since HFD1≈0.965 > the threshold of 0.9, the system classifies this variant as a candidate site that "may be located in a key region of functional regulation and has a strong risk of non-phenotypic interference." Combined with DPI1≈0.287 > T1, it is classified as a two-dimensional high-risk variant. Therefore, it is recommended to classify this variant as a key suspected pathogenic candidate variant and suggest proceeding to subsequent clinical family co-segregation analysis or functional experimental validation pathways. Figure 4 As shown, Figure 4 The HFD (High Disturbance Factor) statistical distribution chart shows the distribution of HFD scores for 50 variant samples. Most variants have HFD values concentrated in the range of 0.4 to 0.8, exhibiting a left-leaning unimodal distribution.
[0066] In summary, by constructing a nonlinear combination model based on sequence complexity, database functional annotation hit count, and repetition region frequency to calculate the latent functional interference degree of each variant, it is beneficial to identify latent variants with pathogenic risk due to disruption of regulatory mechanisms in regions without significant protein structural changes, such as non-coding regions and near splice sites. This supplements the shortcomings of traditional structural models under phenotypic conditions, thereby comprehensively improving the system's ability to identify and predict non-dominant functional site variants.
[0067] The content disclosed above is only a preferred and feasible embodiment of the present invention, and is not intended to limit the scope of protection of the present invention. Therefore, all equivalent technical changes made based on the content of the present invention specification and drawings are included within the scope of protection of the present invention. Furthermore, the elements therein can be updated as technology develops.
Claims
1. A second-generation sequencing data-driven intelligent interpretation system for phenotypic genetic variations, characterized in that, It includes a sequencing data import terminal, a non-phenotypic variant annotation terminal, a rare variant structure modeling terminal, an intelligent inference and genotyping terminal, and a result visualization terminal; the sequencing data import terminal is used to import raw format sequencing files and perform preprocessing to generate variant file information; The preprocessing workflow includes adapter removal, low-quality sequence filtering, alignment to a reference genome, and variant recall. The non-phenotypic variant annotation terminal is used to perform basic annotation and non-phenotypic functional inference on each variant in the variant file information without clinical phenotypic input. The rare variant structure modeling terminal is used to statistically analyze homology sequence loss rate, family structure specificity, and sequence context complexity based on the results of basic annotation and non-phenotypic functional inference, extract non-phenotypic rare features, and construct a discriminant model. The intelligent inference and typing terminal is used to score the pathogenic potential of each variant in a non-phenotypic state using the discriminant model, and classify the discrimination results based on the scoring results. The results visualization terminal is used to display the original format sequencing file, the variant file information, the discrimination model, the pathogenicity potential score, and the discrimination classification, so as to support physicians in reviewing and exporting reports.
2. The next-generation sequencing data-driven intelligent interpretation system for phenotypic genetic variations as described in claim 1, characterized in that, The sequencing data import terminal includes a data reading module, a preprocessing module, and a variant recall module; the data reading module is used to read second-generation sequencing data files in raw format; the preprocessing module is used to perform adapter sequence removal, low-quality read filtering, and sequence normalization operations; The variant invocation module is used to align the preprocessed data to the reference genome, extract single nucleotide variants and small fragment insertion / deletion variants, and generate variant file information.
3. The next-generation sequencing data-driven intelligent interpretation system for phenotypic genetic variations as described in claim 2, characterized in that, The phenotypic variant annotation terminal includes a basic annotation module and a functional inference module. The basic annotation module is used to extract the gene location information, coding influence type, and site conservation characteristics of each variant in the variant file information. The functional inference module is used to infer the functional consequences of the variant based on existing database cross-references and heuristic rule sets without clinical phenotypic input.
4. The next-generation sequencing data-driven intelligent interpretation system for phenotypic genetic variations as described in claim 3, characterized in that, The rare variant structure modeling terminal includes a homology loss rate calculation module, a family structure specificity extraction module, a context complexity analysis module, and a discriminant model construction module. The homology loss rate calculation module is used to statistically analyze the missing information of variant sites in conserved regions of different species. The family structure specificity extraction module is used to evaluate the specific distribution of variants within families based on known family population genetic patterns. The context complexity analysis module is used to calculate the sequence complexity score, database hit density, and repetitive sequence index of the base region surrounding the variant, and construct a non-phenotypic feature vector adapted to the discriminant model. The discriminant model construction module is used to construct a discriminant model based on the non-phenotypic feature vector, the missing information, and the specific distribution.
5. The next-generation sequencing data-driven intelligent interpretation system for phenotypic genetic variations as described in claim 4, characterized in that, The intelligent inference and typing terminal includes a pathogenicity potential calculation module, a latent functional interference assessment module, and a classification judgment module. The pathogenicity potential calculation module is used to calculate a pathogenicity potential index based on the discrimination model. The latent functional interference assessment module is used to calculate the interference score of the variant on the latent functional area based on the sequence context complexity and the functional database cross-analysis. The classification judgment module is used to classify the discrimination results according to the pathogenicity potential index and the interference score.
6. The next-generation sequencing data-driven intelligent interpretation system for phenotypic genetic variations as described in claim 5, characterized in that, The results visualization terminal includes a raw data display module, a variant path tracking module, and a report export module. The raw data display module is used to display the raw format sequencing file and variant file information for manual verification. The variant path tracking module is used to display the discriminant model, the pathogenicity potential score, the interference score, and the discriminant classification to support physician review. The report export module is used to integrate the judgment results and risk scores into a standard clinical interpretation report format, which can be accessed by the medical system and exported manually.
7. A second-generation sequencing data-driven intelligent interpretation method for phenotypic genetic variations, applied to the second-generation sequencing data-driven intelligent interpretation system for phenotypic genetic variations as described in claim 6, characterized in that, The intelligent interpretation method for non-phenotypic genetic variations includes: S1: Import the raw format sequencing file and perform a preprocessing procedure to generate variant file information; S2, without clinical phenotype input, perform basic annotation and non-phenotype functional inference on each variant in the variant file information; S3, by statistically analyzing homologous sequence loss rate, family structure specificity and sequence context complexity, extracts non-phenotypic rare features and constructs a discriminative model; S4. Using the discriminant model, score the pathogenicity potential of each variant in the phenotypic state, and classify the discriminant results based on the scoring results; S5 displays the original format sequencing file, the variant file information, the discrimination model, the pathogenicity potential score, and the discrimination classification to support physician review and report export.
Citation Information
Patent Citations
System and methods for detecting genetic variation
CN104812947A
Genetic variation determination method and system and storage medium
CN109243530A
A genetic disease automatic interpretation system and method based on literature evidence library
CN117373696B
Genetic Variant-Phenotype Analysis System And Methods Of Use
US20170286594A1
Phenotype inference based on incomplete genetic data
EP3859739A1