A method, device and medium for intelligent interpretation of genetic variations in single-gene diseases

By generating an enhanced similarity matrix using disease-specific phenotypic templates and a cosine similarity algorithm, and combining it with a multi-level evidence fusion algorithm inspired by quantum field theory, the problem of weakened phenotypic-genotype association and insufficient depth of multi-level biological evidence fusion in the diagnosis of monogenic diseases is solved, achieving efficient pathogenicity scoring and personalized discrimination.

CN120998299BActive Publication Date: 2026-05-26MINNAN NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MINNAN NORMAL UNIV
Filing Date
2025-10-21
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies are insufficient in phenotype-genotype association strength in the diagnosis of monogenic diseases, making it difficult to effectively handle redundancy, conflict, and hierarchical dependence among multi-level biological evidence, resulting in insufficient biological interpretability and clinical credibility of the interpretation results.

Method used

An enhanced similarity matrix is ​​generated using disease-specific phenotypic templates and a cosine similarity algorithm. A dynamic weight vector is obtained through a dynamic adjustment function. A multi-level evidence fusion algorithm inspired by quantum field theory is used to calculate the pathogenicity score and generate a comprehensive report, which is then integrated into the doctor's workstation to update the case database.

Benefits of technology

It enables deep semantic analysis and quantitative modeling of individualized phenotypic information of patients, enhances the personalized discrimination ability of variant interpretation, and improves the interpretability and clinical credibility of pathogenicity judgment in complex genetic backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998299B_ABST
    Figure CN120998299B_ABST
Patent Text Reader

Abstract

This invention discloses a method, device, and medium for intelligent interpretation of genetic variations in single-gene diseases, relating to the field of bioinformatics. The method includes: generating an enhanced similarity matrix based on standardized phenotypic feature vectors using disease-specific phenotypic templates and a cosine similarity algorithm; obtaining a dynamic weight vector through a dynamic adjustment function; integrating the dynamic weight vector into a normalized multi-omics data matrix; calculating a preliminary perturbation score using a path integral formal algorithm; obtaining a pathway perturbation score; fusing the pathway perturbation score and the dynamic weight vector; calculating a pathogenicity score using a quantum field theory-inspired multi-level evidence fusion algorithm; and generating a comprehensive report of variant pathogenicity grading and clinical recommendations based on a thermodynamic partition function model. This invention achieves nonlinear, high-dimensional collaborative modeling and thermodynamic stability optimization of heterogeneous biological evidence, and also improves the interpretability of pathogenicity judgment in complex genetic contexts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics technology, and in particular to a method, device and medium for intelligent interpretation of genetic variations in single-gene diseases. Background Technology

[0002] With the rapid development of precision medicine and genomics, the interpretation of genetic variations in monogenic diseases has become a core component of clinical molecular diagnostics. Monogenic diseases are caused by mutations in a single gene, exhibiting high genetic heterogeneity and phenotypic complexity; their diagnosis relies on accurate assessment of the functional impact of gene variations. In recent years, the widespread application of high-throughput sequencing technologies (such as whole-exome sequencing and whole-genome sequencing) has enhanced variation detection capabilities and promoted the integrated application of multi-omics data (including genomics, transcriptomics, and epigenomics) in clinical practice. Simultaneously, the accumulation of unstructured clinical text in electronic medical records (EMRs) provides a data foundation for the systematic extraction of phenotypic information.

[0003] Current technical solutions reveal two key shortcomings when facing the complex biological background of monogenic diseases: First, traditional phenotypic feature extraction relies heavily on keyword matching or shallow semantic models, making it difficult to accurately capture disease-specific phenotypic combinations from unstructured electronic medical records, resulting in insufficient phenotypic-genotype association strength. Second, mainstream evidence fusion frameworks are mostly based on Bayesian rules or weighted scoring models, lacking deep modeling mechanisms for the nonlinear interactions between multi-level biological evidence (such as pathway perturbations, functional annotations, population frequencies, etc.), making it difficult to effectively handle redundancy, conflict, and hierarchical dependencies among evidence, thus affecting the biological interpretability and clinical credibility of the final interpretation results. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides an intelligent interpretation method for genetic variations in single-gene diseases to solve the problems of weakened phenotype-genotype association and insufficient depth of multi-level biological evidence fusion in existing technologies.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] In a first aspect, this invention provides an intelligent interpretation method for genetic variations in single-gene diseases, comprising: collecting patient electronic medical record data and multi-omics data; extracting standardized phenotypic feature vectors using natural language processing tools and performing quality control correction on the multi-omics data to generate a normalized multi-omics data matrix; generating an enhanced similarity matrix based on the standardized phenotypic feature vectors using disease-specific phenotypic templates and a cosine similarity algorithm, and obtaining a dynamic weight vector through a dynamic adjustment function; integrating the dynamic weight vectors into the normalized multi-omics data matrix, calculating a preliminary perturbation score through a path integral formal algorithm, and obtaining a pathway perturbation score; fusing the pathway perturbation score and the dynamic weight vectors, calculating a pathogenicity score through a multi-level evidence fusion algorithm inspired by quantum field theory, and generating a comprehensive report of variant pathogenicity grading and clinical recommendations based on a thermodynamic partition function model; converting the comprehensive report into a standardized clinical document, pushing it to the doctor's workstation through a medical data interface, and updating the case database.

[0008] As a preferred embodiment of the intelligent interpretation method for genetic variations in single-gene diseases described in this invention, the method involves: collecting patient electronic medical record data and multi-omics data; extracting standardized phenotypic feature vectors using natural language processing tools; performing quality control correction on the multi-omics data; and generating a normalized multi-omics data matrix. The specific steps are as follows:

[0009] Collect patients' electronic medical record data and multi-omics data to generate a raw dataset;

[0010] Natural language processing tools are used to extract standardized phenotypic feature vectors from the original dataset, and quality control correction is performed on the multi-omics data to generate quality control corrected multi-omics data.

[0011] The federated learning node initialization process performs node allocation and data formatting on the quality-controlled and corrected multi-omics data to generate distributed node local data.

[0012] Based on local data from distributed nodes, a federated aggregated temporary matrix is ​​generated through an aggregation protocol of a federated learning architecture.

[0013] The variance adjustment and noise filtering of the federated aggregate temporary matrix are performed through data calibration optimization to generate a federated aggregate multi-omics data matrix.

[0014] The federated multi-omics data is dynamically normalized using a dynamic normalization algorithm to generate a normalized multi-omics data matrix.

[0015] As a preferred embodiment of the intelligent interpretation method for genetic variations in single-gene diseases described in this invention, the step of generating an enhanced similarity matrix based on standardized phenotypic feature vectors using disease-specific phenotypic templates and a cosine similarity algorithm comprises the following specific steps:

[0016] Calculate the matching similarity score between the standardized phenotypic feature vector and the disease-specific phenotypic template to generate an initial similarity matrix;

[0017] An enhanced similarity matrix is ​​generated by weighting and integrating multi-dimensional features of the initial similarity matrix using the cosine similarity algorithm.

[0018] As a preferred embodiment of the intelligent interpretation method for genetic variations in single-gene diseases described in this invention, the specific steps for obtaining the dynamic weight vector through a dynamic adjustment function are as follows:

[0019] The enhanced similarity matrix is ​​optimized by a dynamic adjustment function to generate a dynamic weight matrix;

[0020] Principal component analysis is used to extract the main weight features from the dynamic weight matrix to obtain the dynamic weight vector.

[0021] As a preferred embodiment of the intelligent interpretation method for genetic variations in single-gene diseases described in this invention, the steps of integrating the dynamic weight vector into the normalized multi-omics data matrix and calculating the preliminary perturbation score using a path integral formal algorithm are as follows:

[0022] The product of the dynamic weight vector and the normalized multi-omics data matrix is ​​used as the enhanced data tensor;

[0023] Based on the enhanced data tensor, a preliminary perturbation score is calculated using a path integral formal algorithm.

[0024] As a preferred embodiment of the intelligent interpretation method for genetic variations in single-gene diseases described in this invention, wherein: the acquisition of pathway perturbation scores is expressed as:

[0025] Multi-scale wavelet transform analysis is performed on the preliminary perturbation score to generate a multi-scale perturbation index;

[0026] Biological validation and scoring optimization of multi-scale perturbation indices were performed using cross-domain knowledge transfer algorithms to obtain pathway perturbation scores.

[0027] As a preferred embodiment of the intelligent interpretation method for genetic variations in single-gene diseases described in this invention, the fusion pathway perturbation score and dynamic weight vector are used to calculate a pathogenicity score through a multi-level evidence fusion algorithm inspired by quantum field theory. Based on a thermodynamic partition function model, a comprehensive report of variant pathogenicity grading and clinical recommendations is generated. The specific steps are as follows:

[0028] By integrating the pathway perturbation score and dynamic weight vector through data fusion processing, fused evidence data is generated.

[0029] A first-level evidence fusion algorithm inspired by quantum field theory is used to perform preliminary fusion processing on the fused evidence data to generate a preliminary fusion score.

[0030] The primary fusion score is optimized and calculated using a second-level evidence fusion algorithm inspired by quantum field theory to generate a pathogenicity score;

[0031] Based on the pathogenicity score, evidence is optimized and integrated using a thermodynamic partition function model to generate an optimized pathogenicity score.

[0032] The optimized pathogenicity score is mapped to a grade using a clinical decision rule engine to generate a clinical grading conclusion.

[0033] A comprehensive report is generated by standardizing and logically verifying clinical grading conclusions through a multi-evidence fusion framework.

[0034] The comprehensive report includes a classification of the pathogenicity of the variant and clinical recommendations.

[0035] As a preferred embodiment of the intelligent interpretation method for genetic variations in single-gene diseases described in this invention, the specific steps for converting the comprehensive report into standardized clinical documents, pushing them to the doctor's workstation via a medical data interface, and updating the case database are as follows.

[0036] The comprehensive report is converted into a standardized clinical document and encrypted before being pushed to the doctor's workstation, and a document transmission status confirmation signal is obtained;

[0037] Based on the document transmission status confirmation signal, the case database is updated through the database transaction management mechanism.

[0038] In a second aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, wherein when the computer program is executed by the processor, it implements any step of the intelligent interpretation method for genetic variations of single-gene diseases as described in the first aspect of the present invention.

[0039] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the intelligent interpretation method for genetic variations in single-gene diseases as described in the first aspect of the present invention.

[0040] The beneficial effects of this invention are as follows: By utilizing disease-specific phenotypic templates and cosine similarity algorithms to generate an enhanced similarity matrix, deep semantic analysis and quantitative modeling of individualized phenotypic information of patients are achieved. This not only overcomes the shortcomings of static weighting methods in ignoring individual differences, but also enhances the personalized discrimination ability of variant interpretation. Furthermore, by employing a multi-level evidence fusion algorithm inspired by quantum field theory to calculate pathogenicity scores, nonlinear, high-dimensional collaborative modeling and thermodynamic stability optimization of heterogeneous biological evidence are achieved, which also improves the interpretability of pathogenicity judgment in complex genetic backgrounds. Attached Figure Description

[0041] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a flowchart of a method for intelligent interpretation of genetic variations in single-gene diseases.

[0043] Figure 2 This is a flowchart of the multi-omics data preprocessing and normalization process.

[0044] Figure 3 The flowchart for phenotypic feature matching and dynamic weight generation.

[0045] Figure 4 This is a flowchart for evidence fusion and generating a comprehensive report. Detailed Implementation

[0046] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0047] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0048] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0049] Reference Figures 1-4This is one embodiment of the present invention, which provides a method for intelligent interpretation of genetic variations in single-gene diseases, including the following steps:

[0050] S1. Collect patient electronic medical record data and multi-omics data, extract standardized phenotypic feature vectors using natural language processing tools, perform quality control correction on the multi-omics data, and generate a normalized multi-omics data matrix.

[0051] Collect patients' electronic medical record data and multi-omics data to generate a raw dataset.

[0052] The specific process involves collecting patients' electronic medical record data and multi-omics data through a medical data interface. This data is then integrated to generate a raw dataset containing textual phenotypic information and molecular characteristic information. This raw dataset serves as the foundational input for subsequent analysis processes, providing a structured data source for natural language processing tools to extract standardized phenotypic feature vectors and for multi-omics data quality control and correction. The construction of this raw dataset ensures unified storage and management of medical information from different sources and supports the integrity and traceability of downstream genetic variation interpretation tasks.

[0053] Patient electronic medical record data includes basic demographic information, clinical diagnostic records, medical history and surgical records, laboratory test results, medical imaging reports, medication records and prescription information, physiological monitoring data, clinical assessments, and disease progress records.

[0054] Multi-omics data includes genomics data, transcriptomics data, proteomics data, metabolomics data, and microbiome data.

[0055] Natural language processing tools are used to extract standardized phenotypic feature vectors from the original dataset, and quality control correction is performed on the multi-omics data to generate quality-controlled and corrected multi-omics data.

[0056] The specific process includes: based on the original dataset, natural language processing tools perform entity recognition and semantic parsing on the unstructured text in the electronic medical record data to extract standardized phenotypic feature vectors. Simultaneously, the ComBat algorithm is used to perform quality assessment and batch effect correction on the multi-omics data, removing low-quality measurements and outliers. The standardized phenotypic feature vectors are then integrated with the quality-controlled multi-omics data through feature alignment and numerical standardization to finally generate quality-controlled and corrected multi-omics data. This process ensures consistency in feature dimensions and numerical scales across different data types, providing high-quality input for subsequent analysis.

[0057] The federated learning node initialization process performs node allocation and data formatting on the quality-controlled and corrected multi-omics data to generate distributed node local data.

[0058] The specific process includes: First, the federated learning node initialization process performs node allocation on the quality-controlled and corrected multi-omics data. This allocation splits the multi-omics data into multiple subsets based on sample features, with each subset assigned to an independent federated learning node. Second, data formatting transforms each node subset into a standard matrix structure, where rows correspond to sample columns and columns correspond to feature dimensions. This formatting process ensures consistency in the dimensions of the multi-omics data and numerical standardization. Finally, distributed node-local data is generated, providing the input basis for subsequent federated aggregation. The entire process maintains the consistency of the multi-omics data distribution and avoids centralized transmission of raw data, complying with the privacy protection principles of federated learning.

[0059] Based on local data from distributed nodes, a federated aggregate temporary matrix is ​​generated through the aggregation protocol of the federated learning architecture.

[0060] The specific process includes taking the local data of distributed nodes as input, processing the integration of multi-node data through the aggregation protocol of the federated learning architecture. The aggregation protocol of the federated learning architecture uses secure multi-party exchange of encrypted parameters instead of the original data. The encrypted parameters contain the statistical characteristics and gradient information of the local data of each node. The aggregation protocol fuses the multi-node parameters through a weighted average algorithm. The weighted average algorithm assigns weight ratios according to the quality of node data and sample size. The fused encrypted parameters are reconstructed into a temporary federated aggregation matrix. The temporary federated aggregation matrix retains the characteristics of multi-node data while ensuring that the original data does not leave the local node.

[0061] The variance of the federated aggregate temporary matrix is ​​adjusted and noise is filtered by data calibration optimization to generate a federated aggregate multi-omics data matrix.

[0062] The specific process includes taking the federated aggregated temporary matrix as input, optimizing it through data calibration, using a variance adjustment method to balance the distribution differences of data from different nodes, analyzing the coefficient of variation of each feature dimension and implementing scale normalization, and simultaneously performing noise filtering to identify and remove outlier data points. The noise filtering operation uses a statistical outlier detection method to remove data that exceeds the dynamic discrimination criteria. The processed federated aggregated temporary matrix retains valid biological signals and eliminates technical biases, ultimately generating a federated aggregated multi-omics data matrix. The federated aggregated multi-omics data matrix has higher data quality and biological consistency, providing reliable input for subsequent analysis.

[0063] It should be noted that the dynamic discrimination criteria are set through an automated analysis process based on the data distribution characteristics (including statistics such as mean, standard deviation, and interquartile range) of each feature dimension in the federated aggregate temporary matrix. The setting process uses a dynamic identification method to generate discrimination boundaries in real time. First, the statistical indicators of each feature dimension are identified, and then data-driven discrimination rules are established based on the statistical indicators (such as the 3σ principle or Tukey's fence method). Finally, the boundary values ​​are adjusted according to the data distribution pattern.

[0064] The federated multi-omics data is dynamically normalized using a dynamic normalization algorithm to generate a normalized multi-omics data matrix.

[0065] The specific process includes taking federated multi-omics data as input, using a dynamic normalization algorithm to standardize the scale of the federated multi-omics data, analyzing the distribution characteristics of the federated multi-omics data and deriving dynamic adjustment parameters, determining the scaling ratio based on data volatility based on the dynamic adjustment parameters, applying the scaling ratio to each dimension of the federated multi-omics data to achieve personalized normalization, eliminating dimensional differences between data from different sources, and finally generating a normalized multi-omics data matrix. The normalized multi-omics data matrix has a unified numerical scale and comparability characteristics, providing standardized input for subsequent analysis.

[0066] S2. Based on standardized phenotypic feature vectors, an enhanced similarity matrix is ​​generated using disease-specific phenotypic templates and a cosine similarity algorithm, and a dynamic weight vector is obtained through a dynamic adjustment function.

[0067] Calculate the matching similarity score between the standardized phenotypic feature vector and the disease-specific phenotypic template to generate an initial similarity matrix, expressed as:

[0068] ;

[0069] in, Indicates the first Similarity score between a disease-specific phenotypic template and a patient's phenotypic features. An index representing a disease-specific phenotypic template. This represents the natural exponential function. This indicates the scaling adjustment parameter. This represents the total dimension of the phenotypic feature vector. Index variables representing phenotypic characteristics, The first phenotype feature vector representing the standardized phenotypic feature vector of a patient 1 eigenvalue, Indicates the first The first disease-specific phenotypic template 1 eigenvalue, Indicates the first The variance of a phenotypic characteristic in a population.

[0070] The specific process includes: based on standardized phenotypic feature vectors, the disease-specific phenotypic template matching process calculates similarity scores using mathematical expressions to generate an initial similarity matrix. The difference between each feature value in the patient's standardized phenotypic feature vector and the corresponding feature value in the disease-specific phenotypic template is calculated, and this difference is squared. These squared differences are then normalized using the variance of each feature. The normalized results of all features are summed. Finally, a scaling parameter is used to control the exponential decay rate of the overall result, thus converting the summed value into a similarity score ranging from 0 to 1. The similarity score reflects the degree of matching between the patient's phenotype and a specific disease template, with a value range between zero and one; a higher value indicates a higher degree of matching. The entire calculation process ensures that the similarity assessment considers both the absolute value of feature differences and the natural variability of features in the population, providing a quantitative basis for subsequent analysis.

[0071] It should be noted that the disease-specific phenotype template is a standardized vector representation constructed based on clinically validated disease phenotype features from authoritative medical knowledge bases (such as OMIM, HPO, Orphanet, etc.). The acquisition process is as follows: extract all relevant phenotype terms for a specific disease from the authoritative medical knowledge base, standardize them through medical ontology, and derive phenotype weights based on large-scale population data to finally form the disease-specific phenotype template vector.

[0072] An enhanced similarity matrix is ​​generated by weighting and integrating multi-dimensional features of the initial similarity matrix using the cosine similarity algorithm.

[0073] The specific process includes: using the initial similarity matrix as input data, processing the angle relationship between multidimensional feature vectors through the cosine similarity algorithm, obtaining the directional consistency in the feature space and generating a directional consistency metric, and combining the directional consistency metric with the feature weight coefficients in the multidimensional feature vector weighting integration process. The feature weight coefficients are allocated according to the importance of the features. The weighting integration process enhances the contribution of the feature space and weakens the influence of noise, finally outputting an enhanced similarity matrix. The enhanced similarity matrix provides a more accurate representation of feature relationships and provides an optimized similarity metric basis for subsequent analysis steps.

[0074] The enhanced similarity matrix is ​​optimized by using a dynamic adjustment function to generate a dynamic weight matrix.

[0075] The specific process includes taking the enhanced similarity matrix as input data, processing the similarity values ​​in the enhanced similarity matrix through a dynamic adjustment function, adjusting the parameter configuration of the dynamic adjustment function according to the real-time data characteristics, performing adaptive weight optimization on each element in the enhanced similarity matrix, taking into account the importance of features and the characteristics of data distribution, recalibrating the similarity relationship after optimization, and finally generating a dynamic weight matrix. The dynamic weight matrix reflects the updated feature association strength and provides a precise weighted similarity basis for subsequent analysis steps.

[0076] It should be noted that the dynamic adjustment function is the core algorithm component for processing the enhanced similarity matrix. Its main function is to weight and optimize the similarity values. By considering feature importance and clinical relevance from multiple dimensions, the dynamic adjustment function achieves intelligent adjustment of the similarity matrix.

[0077] Principal component analysis is used to extract the main weight features from the dynamic weight matrix to obtain the dynamic weight vector.

[0078] The specific process includes: based on the dynamic weight matrix, processing the weight distribution pattern in the matrix through principal component analysis, obtaining eigenvectors and eigenvalues ​​through principal component analysis, identifying the main component directions with the largest variance contribution in the dynamic weight matrix, extracting and retaining the main weight features and performing dimensionality reduction processing, the main weight features constitute a new feature space, and finally obtaining the dynamic weight vector, which represents the compressed core weight information and provides a dimensionality-reduced feature representation for subsequent analysis.

[0079] It should be noted that the main weight features refer to the eigenvectors extracted from the dynamic weight matrix through principal component analysis that have the largest variance contribution rate. These eigenvectors define the main direction of data variation and constitute the core feature subspace after dimensionality reduction.

[0080] S3. Integrate the dynamic weight vector into the normalized multi-omics data matrix, calculate the preliminary perturbation score through the path integral formal algorithm, and obtain the pathway perturbation score.

[0081] The product of the dynamic weight vector and the normalized multi-omics data matrix is ​​used as the enhanced data tensor.

[0082] The specific process includes using a dynamic weight vector and a normalized multi-omics data matrix as input data. The integration process of the dynamic weight vector and the normalized multi-omics data matrix is ​​processed through tensor product operation. The tensor product operation expands each element of the dynamic weight vector by performing an outer product with the corresponding dimension of the normalized multi-omics data matrix, generating a high-dimensional tensor structure. The integration process retains the feature relationships of the original data and combines them with the weight dimension to finally generate an enhanced data tensor. The enhanced data tensor contains weighted multi-omics feature information, providing a high-dimensional data representation basis for subsequent path integral analysis.

[0083] It should be noted that the weight dimension refers to the additional dimension added to the generated enhanced data tensor after expanding the elements of the dynamic weight vector with each dimension of the normalized multi-omics data matrix through tensor product operation. This dimension is used to characterize the weight of the importance of different features. The value of the weight dimension comes directly from the assignment of each weight component in the dynamic weight vector, reflecting the relative importance differences of different features in the multi-omics data.

[0084] Based on the enhanced data tensor, a preliminary perturbation score is calculated using a path integral formal algorithm, expressed as:

[0085] ;

[0086] in, This indicates the initial disturbance score. Represents the hyperbolic tangent function. Indicates the sensitivity parameter. Represents the weight parameter matrix of the first element. line, number Column, No. Layer element values, In the enhanced data tensor, the first... line, number Column, No. Layer element values, Represents the normalization parameter. Indicates the baseline value. Indicates the threshold offset. Index indicating the row direction. Index indicating the column direction, Indicates the index in the depth direction.

[0087] The specific process involves a path integral formal algorithm performing weighted aggregation on the enhanced data tensor, followed by numerical mapping via a nonlinear transformation function, ultimately generating a preliminary perturbation score characterizing the degree of data perturbation. The path integral formal algorithm employs a composite function structure combining an exponential function and a hyperbolic tangent function, where the exponential function handles data scaling and the hyperbolic tangent function implements saturated nonlinear mapping. The output of the path integral formal algorithm is limited to a numerical range of zero to one, where zero represents an unperturbed state, one represents a fully perturbed state, and intermediate values ​​represent different levels of perturbation. This preliminary perturbation score provides a standardized evaluation basis for subsequent analysis steps.

[0088] Multi-scale wavelet transform analysis is performed on the preliminary disturbance score to generate a multi-scale disturbance index.

[0089] The specific process includes: using the initial perturbation score as input data, analyzing and processing the time series features through multi-scale wavelet transform; decomposing the initial perturbation score using basis functions at different scales, extracting the energy distribution of each frequency component, controlling the scaling range of the basis functions using scale parameters, adjusting the time position of the basis functions using translation parameters, generating a series of detail coefficients and approximation coefficients during the decomposition process; capturing high-frequency local features using detail coefficients, and preserving low-frequency trend features using approximation coefficients; the energy values ​​of each scale coefficient constitute an energy spectrum; the energy spectrum is normalized to form a multi-scale perturbation index; and the multi-scale perturbation index characterizes the intensity distribution of the perturbation at different time scales, providing a multi-resolution feature representation for subsequent analysis.

[0090] It should be noted that in wavelet transform for signal processing, the scaling parameter controls the scaling range of the basis function. Specifically, it refers to adjusting the width of the wavelet basis function through mathematical transformation, thereby enabling the extraction and analysis of different frequency components of the signal: when the scaling parameter is large, the basis function is broadened to capture the low-frequency general features of the signal; when the scaling parameter is small, the basis function is contracted to capture the high-frequency details of the signal.

[0091] Biological validation and scoring optimization of multi-scale perturbation indices were performed using cross-domain knowledge transfer algorithms to obtain pathway perturbation scores.

[0092] The specific process includes: based on a multi-scale perturbation index, a cross-domain knowledge transfer algorithm is used to process biological validation and score optimization. The cross-domain knowledge transfer algorithm aligns the multi-scale perturbation index with a database of known biological pathways. The feature alignment process matches perturbation patterns with known biological effects. The biological validation process confirms the pathological relevance of the perturbation patterns. The score optimization process adjusts the weight allocation of the multi-scale perturbation index. The optimized score reflects a more accurate biological impact. Finally, a pathway perturbation score is obtained. The pathway perturbation score quantifies the network-level impact of variations on biological pathways, providing biologically validated assessment results for subsequent analysis.

[0093] It should be noted that the biological pathway database is a structured knowledge base built by integrating biomedical knowledge from multiple sources. It includes pathway information from authoritative databases such as KEGG, Reactome, and BioCyc. It can be obtained through methods such as downloading from public databases, API calls, and literature mining. It stores the interaction relationships of genes, proteins, and metabolites in biological pathways in a standardized format.

[0094] Furthermore, cross-domain knowledge transfer algorithms are intelligent computing methods based on deep neural networks. Their core function is to map the knowledge structure of a source domain (such as a known biological pathway database) to a target domain (multi-scale perturbation index data), achieving cross-domain knowledge transfer through feature space alignment and distribution matching. Cross-domain knowledge transfer algorithms employ adversarial training and maximum mean difference (MMD) optimization to ensure consistent distributions in the latent feature spaces of the two domains, thereby achieving accurate identification and functional annotation of biological perturbation patterns.

[0095] S4 integrates pathway perturbation scores and dynamic weight vectors, calculates pathogenicity scores using a multi-level evidence fusion algorithm inspired by quantum field theory, and generates a comprehensive report of variant pathogenicity grading and clinical recommendations based on a thermodynamic partition function model.

[0096] By integrating the pathway perturbation score and dynamic weight vector through data fusion processing, fused evidence data is generated.

[0097] The specific process includes integrating information through data fusion processing based on the perturbation score and dynamic weight vector. The data fusion processing adopts a weighted combination method to multiply and sum the pathway perturbation score and dynamic weight vector. The weighted combination method assigns fusion weights according to the importance of features. The fusion process retains key biological signals (such as abnormal gene expression signals, protein interaction perturbation signals, and metabolic pathway activity change signals) and eliminates redundant information (such as technical noise, batch effects, and low-quality measurement data). Finally, fused evidence data is generated, which contains optimized feature representations and provides unified evidence input for subsequent analysis.

[0098] A first-level evidence fusion algorithm inspired by quantum field theory is used to perform preliminary fusion processing on the fused evidence data to generate a preliminary fusion score.

[0099] The specific process includes inputting the fused evidence data into a quantum field theory-inspired first-level evidence fusion algorithm for initial fusion. The quantum field theory-inspired first-level evidence fusion algorithm uses path integrals to perform field theory transformations on the fused evidence data. The path integrals identify the weighted average of all possible evidence paths. The weighted averaging process uses an exponential weighting function to model the probability distribution. The result of the probability distribution modeling captures the nonlinear interactions between evidence, and finally generates a primary fusion score. The primary fusion score characterizes the pathogenicity tendency intensity after the initial integration of evidence, providing a basic score for subsequent optimization.

[0100] The primary fusion score is optimized using a second-level evidence fusion algorithm inspired by quantum field theory to generate a pathogenicity score, expressed as:

[0101] ;

[0102] in, Indicates pathogenicity score, Represents the field function of biological pathways Perform functional integration on all possible configurations. Represents the biological pathway field function. Represents a four-dimensional spacetime volume element. Represents the biological pathway field function at four-dimensional spacetime coordinates. The specific field strength value at that location, Represents four-dimensional spacetime coordinates.

[0103] in, Represents the field function of biological pathways The free field action functional is expressed as:

[0104] ;

[0105] It should be noted that, This represents the four-dimensional partial derivative operator. Indicates field mass parameters, This represents the self-coupling constant.

[0106] in, The external coupling term is represented by the following expression:

[0107] ;

[0108] It should be noted that, This represents the coupling coefficient for the path perturbation score. This indicates the pathway disturbance score. Represents the coupling coefficient of the weight vector. This represents the total dimension of the dynamic weight vector. Position index of dynamic weight vector, Represents the first in the dynamic weight vector The specific values ​​of each component.

[0109] in, Represents the field function of biological pathways The observable functional is expressed as:

[0110] ;

[0111] The specific process includes using pathway perturbation scores and dynamic weight vectors as input data, calculating pathogenicity scores through a multi-level evidence fusion algorithm inspired by quantum field theory. This algorithm employs path integrals to handle all possible configurations of the biological pathway field function. The path integrals are used to calculate the expected value of observables through exponential function weighted averaging. Observable functionals characterize the global activity of the biological pathway field, while action functionals describe the dynamic behavior of the field. Exogenous field terms are used to fuse the coupling contribution of pathway perturbation scores and dynamic weight vectors. A four-dimensional spatiotemporal volume element provides the integral measure. The field strength of the biological pathway field function at spatiotemporal coordinates reflects the local activity level, ultimately generating a pathogenicity score. This pathogenicity score quantifies the overall impact of variations on the biological network, providing a final assessment result for clinical decision-making.

[0112] It should be noted that the biological pathway field function is a mathematical framework based on quantum field theory. It abstracts the molecular interactions in biological pathways into a continuous spatiotemporal field. By solving the field equations at specific spatiotemporal coordinates, it quantifies the local biological activity level and finally generates a pathogenicity score based on the field intensity distribution pattern.

[0113] Based on the pathogenicity score, evidence is optimized and integrated using a thermodynamic partition function model to generate an optimized pathogenicity score.

[0114] The specific process includes: based on the pathogenicity score, evidence optimization and integration are processed through a thermodynamic partition function model. The thermodynamic partition function model uses the Boltzmann distribution principle to reweight the pathogenicity score. The Boltzmann distribution principle realizes the probability weight allocation in the form of an exponential function. The evidence optimization process adjusts the contribution ratio of different evidence sources. The evidence integration process integrates multi-source information and eliminates contradictions. The optimized score reflects a more reliable pathogenicity assessment result. Finally, an optimized pathogenicity score is generated. The optimized pathogenicity score provides a calibrated measure of the impact of variation, providing the final judgment basis for clinical decision-making.

[0115] Furthermore, the construction of the thermodynamic partition function model begins with the mathematical definition of state space. First, based on the topological characteristics and dynamic parameters of the molecular interaction network, the phase space of the biological pathway field configuration is defined, establishing the mapping relationship between field variables and energy states. Then, based on the variational principle and Noether's theorem in quantum field theory, an energy functional expression containing kinetic energy, potential energy, and external field coupling terms is constructed. Finally, the extremum characteristics of the energy functional are determined through the variational principle to ensure that the equilibrium conditions of physical constraints are met. The energy functional provides the core mathematical description for subsequent partition function calculations.

[0116] Furthermore, the pre-training process of the thermodynamic partition function model is as follows: First, training samples are extracted from a known pathogenic variant database, containing clinically validated pathogenic and benign variant data. Then, the parameters of the thermodynamic partition function model are initialized, including mass parameters, coupling constants, and temperature coefficients. The parameter configuration is optimized using maximum likelihood estimation to maximize the statistical consistency between the pathogenicity score output by the thermodynamic partition function model and the clinically labeled results. The optimization process uses a gradient descent algorithm to iteratively adjust the parameter values, with each iteration yielding the cross-entropy loss between the predicted value of the thermodynamic partition function model and the true label. The loss function updates the parameters using a backpropagation algorithm until the thermodynamic partition function model converges to a stable state. Finally, the pre-trained thermodynamic partition function model parameter set is obtained, which serves as the basis for subsequent pathogenicity score calculations. This pre-training process ensures that the model possesses clinically usable discriminative capabilities.

[0117] It should be noted that the pathogenic variant database is a standardized knowledge base constructed by integrating experimentally validated and clinically confirmed genetic variant data from authoritative global biomedical databases (such as ClinVar, OMIM, and HGMD) and combining this with literature mining. The construction process of the pathogenic variant database includes multi-source data collection, standardized annotation of variant sites, clinical evidence level classification, population frequency integration, and regular updates and maintenance. Ultimately, it forms a structured database containing classification information such as pathogenicity, benignity, and ambiguous significance, providing a standard reference for the diagnosis and interpretation of genetic diseases and variants.

[0118] The optimized pathogenicity score is then mapped to a clinical grading system using a clinical decision rule engine to generate a clinical grading conclusion.

[0119] The specific process includes taking the optimized pathogenicity score as input data, processing the grading mapping through a clinical decision rule engine, dividing the pathogenicity level according to the intelligent optimization threshold, and converting the continuous score into discrete classification results during the grading mapping process. The classification results include five levels: benign, possibly benign, ambiguous, possibly pathogenic, and pathogenic. Finally, a clinical grading conclusion is generated, which provides a standardized clinical interpretation and provides a classification basis for report generation.

[0120] The Clinical Decision Rule Engine is an automated classification framework built on international clinical guidelines and machine learning algorithms. It provides interpretable decision support for interpreting genetic variations by mapping continuous pathogenicity scores to standardized clinical classification levels.

[0121] The intelligent optimization threshold is set as a classification standard dynamically generated through large-scale clinical validation data and machine learning algorithms. Its core value range is based on the distribution characteristics of pathogenicity scores among known variants. Density clustering and statistical optimization methods are used to determine the optimal classification boundary. The specific numerical range is divided into five continuous probability intervals according to the clinical interpretation requirements of the ACMG / AMP guidelines: benign variant discrimination interval (0.0-0.2) for identifying clearly benign variants; possibly benign interval (0.2-0.4); clinically insignificant interval (0.4-0.6); possibly pathogenic interval (0.6-0.8); and pathogenicity determination interval (0.8-1.0).

[0122] A comprehensive report is generated by standardizing and logically verifying clinical grading conclusions through a multi-evidence fusion framework.

[0123] The specific process includes inputting the clinical grading conclusions into a multi-evidence fusion framework for information integration. The multi-evidence fusion framework uses a structured template to align the clinical grading conclusions with multi-source evidence. The structured template follows the HL7 clinical document architecture standard to ensure standardized formatting. The alignment process verifies the logical consistency between the clinical grading conclusions and the chain of evidence. Logical consistency checks are performed through a rule engine to detect conflicts and resolve contradictions. The verified content is organized according to standard chapters, and finally a comprehensive report is generated. The comprehensive report contains complete clinical interpretation content and a standardized format, providing an authoritative document for medical decision-making.

[0124] S5. Transform the comprehensive report into a standardized clinical document, push it to the doctor's workstation through the medical data interface, and update the case database.

[0125] The comprehensive report is converted into a standardized clinical document and encrypted before being pushed to the physician's workstation, and a document transmission status confirmation signal is obtained.

[0126] The specific process includes taking the comprehensive report as input and converting it into a standardized clinical document through a structured conversion engine. The structured conversion engine reorganizes the report content according to the HL7 clinical document architecture standard. The encrypted push process uses a transport layer security protocol to encapsulate the standardized clinical document. After encapsulation, it is sent to the doctor's workstation through a medical data interface. After receiving the data, the doctor's workstation verifies its integrity and parses the content. When the verification is successful, it returns a document transmission status confirmation signal, which indicates that the data has been safely delivered and can be used clinically.

[0127] Based on the document transmission status confirmation signal, the case database is updated through the database transaction management mechanism.

[0128] The specific process includes: the database transaction management mechanism verifies the integrity and validity of the document transmission status confirmation signal; after the confirmation signal is valid, atomic transaction operation is initiated; the atomic transaction operation writes the variant pathogenicity classification and clinical recommendation information in the standardized clinical document into the corresponding fields of the case database; the writing process maintains data consistency and transaction isolation; after the transaction is committed, a database update completion flag is generated; the database update completion flag confirms that the case database has successfully synchronized the latest clinical decision information, providing data integrity assurance for medical record management.

[0129] It should be noted that the database transaction management mechanism is the core technical framework in database management components that ensures the atomicity, consistency, isolation, and durability (ACID properties) of data operations. The core objective of the database transaction management mechanism is to ensure that the database remains in a logically correct state under concurrent access and failure scenarios.

[0130] This embodiment also provides a computer device applicable to the intelligent interpretation method of genetic variations in single-gene diseases, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the intelligent interpretation method of genetic variations in single-gene diseases as proposed in the above embodiment.

[0131] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0132] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the intelligent interpretation method for genetic variations in single-gene diseases as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0133] In summary, this invention achieves deep semantic analysis and quantitative modeling of individualized phenotypic information of patients by using disease-specific phenotypic templates and cosine similarity algorithms to generate enhanced similarity matrices. This not only overcomes the shortcomings of static weighting methods in ignoring individual differences but also enhances the personalized discrimination ability of variant interpretation. Furthermore, by employing a multi-level evidence fusion algorithm inspired by quantum field theory to calculate pathogenicity scores, this invention achieves nonlinear, high-dimensional collaborative modeling and thermodynamic stability optimization of heterogeneous biological evidence and improves the interpretability of pathogenicity judgment in complex genetic contexts.

[0134] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for intelligent interpretation of genetic variations in single-gene diseases, characterized in that: include, Collect patients' electronic medical record data and multi-omics data, extract standardized phenotypic feature vectors using natural language processing tools, perform quality control correction on the multi-omics data, and generate a normalized multi-omics data matrix; Based on standardized phenotypic feature vectors, an enhanced similarity matrix is ​​generated using disease-specific phenotypic templates and a cosine similarity algorithm, and a dynamic weight vector is obtained through a dynamic adjustment function. The dynamic weight vector is integrated into the normalized multi-omics data matrix, and the preliminary perturbation score is calculated using a path integral formal algorithm to obtain the pathway perturbation score. The specific steps are as follows: The product of the dynamic weight vector and the normalized multi-omics data matrix is ​​used as the enhanced data tensor; Based on the enhanced data tensor, a preliminary perturbation score is calculated using a path integral formal algorithm; Multi-scale wavelet transform analysis is performed on the preliminary perturbation score to generate a multi-scale perturbation index; Biological validation and scoring optimization of multi-scale perturbation indices were performed using cross-domain knowledge transfer algorithms to obtain pathway perturbation scores. The pathogenicity score is calculated using a multi-level evidence fusion algorithm inspired by quantum field theory, which integrates pathway perturbation scores and dynamic weight vectors. Based on a thermodynamic partition function model, a comprehensive report including variant pathogenicity grading and clinical recommendations is generated. The specific steps are as follows: By integrating the pathway perturbation score and dynamic weight vector through data fusion processing, fused evidence data is generated. A first-level evidence fusion algorithm inspired by quantum field theory is used to perform preliminary fusion processing on the fused evidence data to generate a preliminary fusion score. The primary fusion score is optimized and calculated using a second-level evidence fusion algorithm inspired by quantum field theory to generate a pathogenicity score; Based on the pathogenicity score, evidence is optimized and integrated using a thermodynamic partition function model to generate an optimized pathogenicity score. The optimized pathogenicity score is mapped to a grade using a clinical decision rule engine to generate a clinical grading conclusion. A comprehensive report is generated by standardizing and logically verifying clinical grading conclusions through a multi-evidence fusion framework. The comprehensive report includes a classification of variant pathogenicity and clinical recommendations; The comprehensive report is transformed into a standardized clinical document, which is then pushed to the physician's workstation via a medical data interface and the case database is updated.

2. The intelligent interpretation method for genetic variations in single-gene diseases as described in claim 1, characterized in that: The process involves collecting patient electronic medical record data and multi-omics data, extracting standardized phenotypic feature vectors using natural language processing tools, performing quality control correction on the multi-omics data, and generating a normalized multi-omics data matrix. The specific steps are as follows: Collect patients' electronic medical record data and multi-omics data to generate a raw dataset; Natural language processing tools are used to extract standardized phenotypic feature vectors from the original dataset, and quality control correction is performed on the multi-omics data to generate quality control corrected multi-omics data. The federated learning node initialization process performs node allocation and data formatting on the quality-controlled and corrected multi-omics data to generate distributed node local data. Based on local data from distributed nodes, a federated aggregated temporary matrix is ​​generated through an aggregation protocol of a federated learning architecture. The variance adjustment and noise filtering of the federated aggregate temporary matrix are performed through data calibration optimization to generate a federated aggregate multi-omics data matrix. The federated multi-omics data is dynamically normalized using a dynamic normalization algorithm to generate a normalized multi-omics data matrix.

3. The intelligent interpretation method for genetic variations in single-gene diseases as described in claim 2, characterized in that: The process of generating an enhanced similarity matrix based on standardized phenotypic feature vectors, using disease-specific phenotypic templates and a cosine similarity algorithm, involves the following steps: Calculate the matching similarity score between the standardized phenotypic feature vector and the disease-specific phenotypic template to generate an initial similarity matrix; An enhanced similarity matrix is ​​generated by weighting and integrating multi-dimensional features of the initial similarity matrix using the cosine similarity algorithm.

4. The intelligent interpretation method for genetic variations in single-gene diseases as described in claim 3, characterized in that: The specific steps for obtaining the dynamic weight vector through a dynamic adjustment function are as follows. The enhanced similarity matrix is ​​optimized by a dynamic adjustment function to generate a dynamic weight matrix; Principal component analysis is used to extract the main weight features from the dynamic weight matrix to obtain the dynamic weight vector.

5. The intelligent interpretation method for genetic variations in single-gene diseases as described in claim 4, characterized in that: The process of converting the comprehensive report into standardized clinical documents, pushing them to the doctor's workstation via a medical data interface, and updating the case database involves the following steps. The comprehensive report is converted into a standardized clinical document and encrypted before being pushed to the doctor's workstation, and a document transmission status confirmation signal is obtained; Based on the document transmission status confirmation signal, the case database is updated through the database transaction management mechanism.

6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the intelligent interpretation method for genetic variations of single-gene diseases as described in any one of claims 1 to 5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the intelligent interpretation method for genetic variations of single-gene diseases as described in any one of claims 1 to 5.