Genetic disease gene detection data analysis method and system based on big data

By using technical means such as multi-dimensional quality assessment, dual-threshold dynamic quality control and multi-level variation detection in genetic disease gene detection data analysis, the problems of inefficiency of existing methods and difficulty in dealing with complex data relationships are solved, and efficient and accurate genetic detection data analysis and individual risk assessment are achieved.

CN120164522AInactive Publication Date: 2025-06-17CHANGZHOU CHILDRENS HOSPITAL (CHANGZHOU SIXTH PEOPLES HOSPITAL)
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202411799116.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-06-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing genetic disease gene detection data analysis methods are inefficient, susceptible to subjective factors, and are difficult to effectively deal with sequencing errors, noise data, and capture complex nonlinear relationships and multi-scale information in gene data.

Method used

A high-quality standard variation detection report is generated through technical means such as multi-dimensional quality assessment, dual-threshold dynamic quality control, locally sensitive hash algorithm deduplication, distributed computing framework, multi-level variation detection strategy and Bayesian statistical model, and risk assessment is carried out through multi-scale feature sampling, God-frequent differential equation network and integrated predictor.

Benefits of technology

The quality and processing speed of gene sequencing data are improved, the accuracy and reliability of variant detection are enhanced, and the generated reports provide a high-quality data basis for subsequent analysis, achieving accurate assessment of individual genetic disease risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164522A_ABST
    Figure CN120164522A_ABST
Patent Text Reader

Abstract

The invention provides a genetic disease gene detection data analysis method and system based on big data, and relates to the technical field of gene detection data analys.The genetic disease gene detection data analysis method comprises the steps that quality control and duplicate removal are conducted on gene sequencing original data, sequence comparison and variation detection are conducted, and a standard variation detection report is generated; performing multi-scale feature sampling and optimization by using an improved random field diffusion model, inputting high-dimensional feature distribution into a dual self-activation iterative network for feature extraction and integration to obtain a feature mapping matrix and a feature evolution trajectory, and inputting the feature mapping matrix and the feature evolution trajectory into an integrated predictor to obtain an integrated predictor; and in combination with Gaussian mixture process regression analysis and an improved Bayesian reasoning network, establishing a risk association network and training a deep hierarchical decision tree, and outputting a multi-dimensional risk assessment report.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gene detection data analysis, and particularly to a method and system for analyzing genetic disease gene detection data based on big data. Background Art

[0002] Genetic disease gene detection is an important part of precision medicine, aiming to evaluate the risk of an individual having a genetic disease by detecting variant sites in the individual's genome. With the rapid development of high-throughput sequencing technology, the acquisition cost of gene data has been greatly reduced, making it possible to conduct gene detection on a large scale. How to efficiently and accurately extract effective information from massive gene data and perform disease risk prediction has become a hot and difficult issue in current research.

[0003] Traditional methods for analyzing genetic disease gene detection data mainly rely on expert experience and manual interpretation, which are inefficient and easily affected by subjective factors. Existing analysis methods based on machine learning are not fine enough in filtering sequencing errors and noise data, and it is difficult to effectively capture complex non-linear relationships and multi-scale information in gene data;

[0004] Therefore, there is an urgent need for a method to solve the problems existing in the prior art. Summary of the Invention

[0005] An embodiment of the present invention provides a method and system for analyzing genetic disease gene detection data based on big data, which can at least solve some problems existing in the prior art.

[0006] In the first aspect of the embodiment of the present invention, a method for analyzing genetic disease gene detection data based on big data is provided, including:

[0007] Obtaining the original gene sequencing data of the individual to be detected, performing multi-dimensional quality assessment on the original gene sequencing data to generate a quality assessment index set, filtering the quality assessment index set through a dual-threshold dynamic quality control algorithm to obtain high-quality sequence data, and performing deduplication operation on the high-quality sequence data through an improved locality-sensitive hashing algorithm to obtain non-duplicate sequence data. Building a sequence comparison system based on a distributed computing framework, generating a reference genome, and comparing the non-duplicate sequence data with the reference genome by combining the Burrows-Wheeler transform algorithm and a fast comparison algorithm based on a suffix array to generate a sequence position mapping file, and performing variant detection through a multi-level variant detection strategy to generate an initial variant detection result. Scoring the credibility of the initial variant detection result through a Bayesian statistical model to generate a standard variant detection report including credibility scores;

[0008] Based on the standard variant detection report, multi-scale feature sampling is performed according to the improved random field diffusion model in the pre-set self-activation depth analysis system, combining Langevin dynamics and Hamiltonian Monte Carlo methods to generate an initial feature distribution. The initial feature distribution is optimized based on an adaptive adjustment mechanism and Wasserstein distance to generate a high-dimensional feature distribution, which is then input into the pre-set dual self-activation iterative network. Spatial features are extracted through the dynamic routing algorithm and attention guidance mechanism in the excitation path based on the capsule network. The spatial features are added to the inhibition path for contrast learning and competitive collaboration to obtain balanced features. The optimal features are selected according to the Boltzmann machine and annealing algorithm in the metastable state, and the optimal features are added to the dynamic system based on Lie differential equations. Combining neural ordinary differential equation networks to model the temporal evolution and perform feature integration to obtain a feature mapping matrix and a feature evolution trajectory;

[0009] The feature mapping matrix and the feature evolution trajectory are added to the pre-set ensemble predictor, which includes a deep belief network and a hypergraph neural network. The parameters of the ensemble predictor are optimized according to the alternating training strategy and the initial risk score is calculated. The feature evolution trajectory is analyzed by a mixture Gaussian process regression to predict the risk change trend. The initial risk score and the risk change trend are added to the improved Bayesian inference network. The Markov blanket model is executed to update the risk probability, and the probability estimation bias is corrected by the particle filter algorithm to establish a risk association network between mutation sites. Based on the risk association network, a deep hierarchical decision tree is trained, and node division is performed in combination with the adaptive splitting threshold algorithm to obtain the risk level and output the multi-dimensional risk assessment report corresponding to the current individual.

[0010] In an alternative embodiment,

[0011] Obtain the original gene sequencing data of the individual to be detected, perform multi-dimensional quality assessment on the original gene sequencing data to generate a quality assessment index set, filter the data of the quality assessment index set through a double-threshold dynamic quality control algorithm to obtain high-quality sequence data, and perform deduplication operations through an improved locality-sensitive hashing algorithm to obtain non-repetitive sequence data. Based on the distributed computing framework, a sequence comparison system is constructed to generate a reference genome, and the non-repetitive sequence data and the reference genome are compared by combining the Burrows-Wheeler transform algorithm and the fast comparison algorithm based on the suffix array to generate a sequence position mapping file and perform variant detection through a multi-level variant detection strategy to generate an initial variant detection result. The initial variant detection result is scored for credibility through a Bayesian statistical model to generate a standard variant detection report including credibility scores:

[0012] Obtain the original gene sequencing data of the individual to be detected measured by a high-throughput sequencing instrument, perform multi-dimensional quality assessment on the original gene sequencing data through the FastQC tool, screen the original gene sequencing data according to the pre-set assessment indicators, and obtain the quality assessment indicator set;

[0013] According to the quality assessment indicator set, set the first quality threshold and the second quality threshold according to the double-threshold dynamic quality control algorithm, where the first quality threshold is less than the second quality threshold. For each gene sequence in the quality assessment indicator set, calculate the average quality value and compare it with the first quality threshold and the second quality threshold. If the average quality value is less than the first quality threshold, filter the current gene sequence. If the average quality value is greater than the second quality threshold, retain the current gene sequence. If the average quality value is greater than the first quality threshold and less than the second quality threshold, excise the bases at both ends of the current gene sequence with quality values lower than the first quality threshold and retain the middle fragment. Output the retained middle fragment and gene sequence as high-quality sequence data;

[0014] For the high-quality sequence data, extract the base segments in each gene sequence as key values through the locality-sensitive hashing algorithm, use the gene sequence as the corresponding value to construct a hash table, and judge the sequence repeatability by looking up the hash table. Remove the repeated sequences in the high-quality sequences to obtain non-repeated sequence data;

[0015] For the non-repeated sequence data, perform sharding according to the pre-set sharding rules, generate multiple data shards and send them to a pre-constructed sequence comparison system based on a distributed computing framework. On the computing nodes in the sequence comparison system, generate a reference genome and perform a rough comparison based on the Burrows-Wheeler transform algorithm. Determine candidate regions on the reference genome according to the rough comparison results, and perform an accurate comparison in combination with the fast comparison algorithm based on the suffix array. Output the initial comparison result, perform quality assessment on the initial comparison result, generate a read quality score, and judge whether to save the current initial comparison result according to the pre-set read quality threshold. Merge the retained initial comparison results to generate a sequence position mapping file;

[0016] Based on the sequence position mapping file, map the gene sequences in the non-repeated sequence data to the reference genome, perform local reassembly on the sequences mapped to the reference genome, detect single nucleotide polymorphism sites and insertion-deletion sites in combination with the global debiasing strategy, evaluate the confidence of each site, and filter out false positive sites to obtain the initial variant detection result;

[0017] Based on the initial mutation detection results, evaluate the sequence depth, quality value, and mutation frequency of each mutation site in the gene sequence, calculate the probability score corresponding to each mutation site, sort the mutation sites according to the probability score, and select the mutation site with the highest probability score to generate a standard mutation detection report.

[0018] In an alternative embodiment,

[0019] Based on the sequence position mapping file, map the gene sequences in the non-redundant sequence data to the reference genome, perform local reassembly on the sequences mapped to the reference genome, combine the global debiasing strategy to detect single nucleotide polymorphism sites and insertion / deletion sites, evaluate the confidence of each site, and filter false positive sites, including:

[0020] Read the non-redundant sequence data, reference genome sequence, and sequence position mapping file. The sequence position mapping file contains the mapping position information and matching degree information of the reads on the reference genome. Extract the read sequences according to the sequence position mapping file, align the read sequences to the corresponding positions on the reference genome, and parse the alignment string to determine the base matching relationship between the read sequences and the reference genome;

[0021] Perform local reassembly on the read sequences aligned to the same region with a ten-base window, use the minimum edit distance algorithm to find the optimal alignment method, correct the alignment offset of the read sequences, obtain the reassembled read sequences, calculate the likelihood probability of the site bases using the Bayesian statistical model according to the base types and quality values supported by the reassembled read sequences, and combine the reference genome sequence to judge the mutation sites;

[0022] Establish a correction model, which includes sequencing quality parameters, base content parameters, and read depth parameters. Correct the likelihood probability of the mutation sites according to the correction model to obtain the corrected likelihood probability. Screen the mutation sites according to the corrected likelihood probability. When the mutation site is a heterozygous mutation, retain the mutation sites with the mutation base frequency between 20% and 80%. When the mutation site is a homozygous mutation, retain the mutation sites with the mutation base frequency greater than 80%;

[0023] Calculate the quality value, sequencing depth, and number of reads supporting the mutation of the mutation sites. Use the single nucleotide polymorphism sites with the quality value greater than 30 and the sequencing depth greater than ten times, and the insertion / deletion sites with the quality value greater than 50 and the sequencing depth greater than twenty times as high-confidence mutation sites, annotate the high-confidence mutation sites, and generate a mutation annotation file including genomic position information, mutation type information, allele frequency information, and functional impact information.

[0024] In an alternative embodiment,

[0025] Based on the standard mutation detection report, according to the improved random field diffusion model in the pre-set self-activation depth analysis system, combined with Langevin dynamics and Hamiltonian Monte Carlo methods, multi-scale feature sampling is performed to generate an initial feature distribution. Based on the adaptive adjustment mechanism and Wasserstein distance, the initial feature distribution is optimized to generate a high-dimensional feature distribution and input it into the pre-set dual self-activation iterative network. Spatial features are extracted through the dynamic routing algorithm and attention guidance mechanism in the excitation path based on the capsule network. The spatial features are added to the inhibition path and contrast learning and competitive collaboration are performed to obtain balanced features. The optimal features are selected under metastable states according to the Boltzmann machine and annealing algorithm. The optimal features are added to the dynamic system based on Lie differential equations, and the time series evolution is modeled and feature integration is performed in combination with the neural ordinary differential equation network to obtain a feature mapping matrix and a feature evolution trajectory, including:

[0026] Based on the standard mutation detection report, multi-scale feature sampling is performed according to the improved random field diffusion model in the pre-set self-activation depth analysis system. Among them, the improved random field diffusion model uses the Langevin dynamics method to describe the movement trajectory of particles on the potential energy surface, combined with the Hamiltonian Monte Carlo method to accept sub-optimal sampling points with a preset probability, and samples mutation features at the chromosome level, gene level, and exon level to generate an initial feature distribution;

[0027] Based on the adaptive adjustment mechanism and Wasserstein distance, the initial feature distribution is optimized. The adaptive adjustment mechanism dynamically adjusts the temperature parameter and step size parameter according to the sampling acceptance rate. The Wasserstein distance is used to measure the difference between the current distribution and the ideal distribution, and the distribution is optimized by minimizing the Wasserstein distance to generate a high-dimensional feature distribution;

[0028] The high-dimensional feature distribution is input into the pre-set dual self-activation iterative network. The dual self-activation iterative network includes an activation path and an inhibition path. In the activation path, the dynamic routing algorithm based on the capsule network is used to adaptively adjust the connection strength between capsules, and the attention mechanism is introduced to dynamically adjust the weights according to feature correlation to extract spatial features. The spatial features are input into the inhibition path for perturbation, and the original features and perturbed features are discriminated through contrast learning. The balanced features are obtained through the competition and collaboration of the dual paths;

[0029] A Boltzmann machine is constructed, the weight matrix is defined based on the feature correlation, and the energy function is defined based on the disease risk. The annealing algorithm is used to iteratively sample the balanced features under metastable states to obtain the optimal features with the lowest energy;

[0030] Add the optimal features to the dynamic system based on Lie differential equations, and combine the neural ordinary differential equation network to perform temporal modeling on the dynamic system. Consider each network layer as a time step for solving ordinary differential equations, sample at different time steps to obtain the feature mapping matrix, and generate the feature evolution trajectory.

[0031] In an alternative implementation,

[0032] Adding the optimal features to the dynamic system based on Lie differential equations, and combining the neural ordinary differential equation network to perform temporal modeling on the dynamic system. Considering each network layer as a time step for solving ordinary differential equations, sampling at different time steps to obtain the feature mapping matrix, and generating the feature evolution trajectory includes:

[0033] Obtain the values of the optimal features at different time steps and construct the system state vector. Based on the system state vector, establish a Lie differential equation dynamic system, calculate the interaction strength between state variables to generate the state transition matrix, and calculate the influence degree of external input on state variables to generate the control matrix;

[0034] Fuse the set corresponding to the optimal features with the system state vector, calculate the interaction strength between the fused features and the original features, and update the state transition matrix and the control matrix according to the interaction strength;

[0035] Build a neural ordinary differential equation network including an input layer, multiple hidden layers, and an output layer. In each hidden layer, multiply the previous hidden state by the weight matrix and add the bias vector, input the calculation result into the activation function to obtain the time change rate of the current hidden state, update the current hidden state based on the time change rate through numerical integration methods, and map the state transition matrix and the control matrix to the network parameter space to complete network initialization;

[0036] Use the backpropagation algorithm to calculate the parameter gradients and iteratively update the network parameters. Execute the forward propagation process at a preset number of time steps, obtain and save the hidden state corresponding to each time step, extract the hidden state at a preset sampling interval in the time series, splice the extracted hidden states in chronological order to generate the feature mapping matrix, and smooth the feature mapping matrix to obtain the feature evolution trajectory.

[0037] In an alternative implementation,

[0038] Add the feature mapping matrix and the feature evolution trajectory to a pre-set integrated predictor, which includes a deep belief network and a hypergraph neural network. Optimize the parameters of the integrated predictor according to an alternating training strategy and calculate an initial risk score. Analyze the feature evolution trajectory based on a mixture Gaussian process regression to predict the risk change trend. Add the initial risk score and the risk change trend to an improved Bayesian inference network, execute a Markov blanket model to update the risk probability, correct the probability estimation bias through a particle filtering algorithm, and establish a risk association network between mutation sites. Train a deep hierarchical decision tree based on the risk association network, perform node division in combination with an adaptive splitting threshold algorithm, obtain the risk level, and output a multi-dimensional risk assessment report corresponding to the current individual, including:

[0039] Add the feature mapping matrix and the feature evolution trajectory as input samples to a pre-set integrated predictor, which includes a deep belief network and a hypergraph neural network. The deep belief network is composed of multiple stacked restricted Boltzmann machines, and the hypergraph neural network maps features to a high-order hypergraph space and models the interaction between features through hyperedge connections;

[0040] Optimize the parameters of the integrated predictor using an alternating training strategy. Fix the parameters of the deep belief network and update the parameters of the hypergraph neural network, then fix the parameters of the hypergraph neural network and update the parameters of the deep belief network. Iteratively train until convergence and then predict the input samples to obtain an initial risk score;

[0041] Construct a mixture Gaussian process regression model composed of multiple sub-Gaussian processes. Use the expectation maximization algorithm to iteratively optimize the parameters of the mixture Gaussian process regression model, calculate the posterior probability that the input samples belong to different sub-Gaussian processes, and update the hyperparameters of each sub-Gaussian process. Predict the feature evolution trajectory based on the trained mixture Gaussian process regression model to obtain the risk change trend;

[0042] Construct an improved Bayesian inference network, which includes multiple time slices. Each time slice corresponds to an independent Bayesian network, and adjacent time slices are connected through conditional probability distributions. Set the prior probability of the risk nodes in the initial time slice according to the initial risk score. Calculate the second posterior probability corresponding to each node using the Markov blanket model. Correct the second posterior probability according to the particle filtering algorithm, approximate the posterior probability distribution through weighted sampling particles, update the particles according to the state transition model, and adjust the particle weights according to the observation model. Calculate the risk correlation between mutation sites based on the corrected second posterior probability and establish a risk association network;

[0043] Construct a deep hierarchical decision tree, optimize node splitting using an adaptive splitting threshold algorithm, dynamically adjust the splitting threshold according to node purity and sample distribution, train the deep hierarchical decision tree based on the risk association network, predict the risk level of the input sample through the deep hierarchical decision tree and analyze key mutation combinations, obtain the risk level corresponding to the input sample, and combine the risk level of the input sample and the key mutation combinations to obtain a multi-dimensional risk assessment report.

[0044] In an alternative embodiment,

[0045] Update the hyperparameters of each sub-Gaussian process as shown in the following formula

[0046]

[0047] where, θ k represents the hyperparameter set of the k-th sub-Gaussian process, represents the updated hyperparameter set of the k-th sub-Gaussian process, γ(z nk ) represents the posterior probability that the n-th input sample belongs to the k-th sub-Gaussian process, represents the likelihood function of the sample y n under the k-th sub-Gaussian process, y n represents the observed value of the n-th input sample, δ represents the mean of the sub-Gaussian process distribution, with a value of 0, represents the covariance matrix of the k-th sub-Gaussian process at sample n, σ 2 represents the noise variance, I represents the identity matrix, v represents the degree-of-freedom parameter of the probability density function, represents the probability density function, and N represents the total number of input samples.

[0048] In the second aspect of the embodiments of the present invention, a genetic disease gene detection data analysis system based on big data is provided, including:

[0049] A first unit for obtaining the original gene sequencing data of an individual to be detected, generating a quality assessment index set for multi-dimensional quality assessment of the original gene sequencing data, filtering the quality assessment index set through a dual-threshold dynamic quality control algorithm to obtain high-quality sequence data and performing a deduplication operation through an improved locality-sensitive hashing algorithm to obtain non-repetitive sequence data, constructing a sequence comparison system based on a distributed computing framework, generating a reference genome, and comparing the non-repetitive sequence data and the reference genome by combining the Burrows-Wheeler transform algorithm and a fast comparison algorithm based on a suffix array to generate a sequence position mapping file and performing mutation detection through a multi-level mutation detection strategy to generate an initial mutation detection result, and performing a credibility score on the initial mutation detection result through a Bayesian statistical model to generate a standard mutation detection report containing credibility scores;

[0050] The second unit is configured to generate an initial feature distribution based on the standard variant detection report, perform multi-scale feature sampling according to the improved random field diffusion model in the preset self-activation depth analysis system, combine Langevin dynamics and Hamiltonian Monte Carlo method, optimize the initial feature distribution based on the adaptive adjustment mechanism and Wasserstein distance to generate a high-dimensional feature distribution and input it into the preset dual self-activation iterative network, extract spatial features through the dynamic routing algorithm and attention guidance mechanism in the excitation path based on the capsule network, add the spatial features to the suppression path and perform contrast learning and competitive collaboration to obtain balanced features, select the optimal features according to the Boltzmann machine and annealing algorithm in the metastable state, add the optimal features to the dynamic system based on Lie differential equations, combine neural ordinary differential equation networks to model the temporal evolution and perform feature integration to obtain a feature mapping matrix and a feature evolution trajectory;

[0051] The third unit is configured to add the feature mapping matrix and the feature evolution trajectory to the preset ensemble predictor, which includes a deep belief network and a hypergraph neural network, optimize the parameters of the ensemble predictor according to the alternating training strategy and calculate the initial risk score, analyze the feature evolution trajectory according to the mixture Gaussian process regression and predict the risk change trend, add the initial risk score and the risk change trend to the improved Bayesian inference network, execute the Markov blanket model to update the risk probability and correct the probability estimation deviation through the particle filter algorithm and establish a risk association network between mutation sites, train a deep hierarchical decision tree based on the risk association network, and perform node division in combination with the adaptive splitting threshold algorithm to obtain the risk level and output the multi-dimensional risk assessment report corresponding to the current individual.

[0052] In the third aspect of the embodiments of the present invention,

[0053] A kind of electronic device is provided, including:

[0054] A processor;

[0055] A memory for storing instructions executable by the processor;

[0056] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0057] In the fourth aspect of the embodiments of the present invention,

[0058] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0059] In the present invention, through operations such as multi-dimensional quality assessment, dual-threshold dynamic quality control, and duplicate removal using an improved locality-sensitive hashing algorithm, the quality of gene sequencing data is effectively improved. By using a distributed computing framework and an efficient sequence alignment algorithm, the data processing speed is accelerated. The application of a multi-level mutation detection strategy and a Bayesian statistical model improves the accuracy and reliability of mutation detection. Finally, a standard mutation detection report containing credibility scores is generated, providing a high-quality data basis for subsequent analysis. Extract multi-scale features from the standard mutation detection report, and combine with neural ordinary differential equation network to model the temporal evolution, obtaining a feature mapping matrix and a feature evolution trajectory, which can more comprehensively reflect the genetic information and risk status of an individual. Integrate the feature information and the risk change trend, and correct the probability estimation deviation through a particle filter algorithm. Finally, establish a risk association network between mutation sites, which can more accurately evaluate the genetic disease risk of an individual and output a multi-dimensional risk assessment report including risk levels, providing personalized health guidance for the individual. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 is a schematic flowchart of the data analysis method for genetic disease gene detection based on big data according to an embodiment of the present invention;

[0061] Figure 2 is a schematic structural diagram of the data analysis system for genetic disease gene detection based on big data according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0063] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0064] Figure 1 is a schematic flowchart of the data analysis method for genetic disease gene detection based on big data according to an embodiment of the present invention, as Figure 1 shown, the method includes:

[0065] S1. Obtain the original gene sequencing data of the individual to be detected, perform multi-dimensional quality assessment on the original gene sequencing data to generate a quality assessment index set, filter the quality assessment index set through a double-threshold dynamic quality control algorithm to obtain high-quality sequence data, and perform deduplication operations through an improved locality-sensitive hashing algorithm to obtain non-repetitive sequence data. Construct a sequence comparison system based on a distributed computing framework, generate a reference genome, and combine the Burrows-Wheeler transform algorithm and a fast comparison algorithm based on a suffix array to compare the non-repetitive sequence data with the reference genome, generate a sequence position mapping file, and perform mutation detection through a multi-level mutation detection strategy to generate an initial mutation detection result. Perform credibility scoring on the initial mutation detection result through a Bayesian statistical model to generate a standard mutation detection report containing credibility scores;

[0066] The original gene sequencing data refers to the original sequence data obtained through gene sequencing technology without any processing, usually containing short sequence fragments, which are used for further genomic analysis and research after subsequent processing. The double-threshold dynamic quality control algorithm is a data quality control method that dynamically adjusts the range of data quality control in gene data processing by setting two thresholds to help identify and remove low-quality sequence data. The locality-sensitive hashing algorithm is an algorithm for approximate similarity search that can quickly find similar gene sequences in a large-scale dataset, thus accelerating sequence alignment and reducing computational complexity. The Burrows-Wheeler transform algorithm is an efficient string compression algorithm commonly used for the storage and compression of gene sequencing data, which improves the data compression rate and retrieval efficiency by transforming the arrangement of gene sequence data. The fast comparison algorithm based on a suffix array is a data structure based on a suffix array and a suffix tree that can efficiently perform matching and comparison in gene sequences and is usually used for sequence alignment and genetic mutation detection. The sequence position mapping file is a file that records the correspondence between each position in the gene sequence and the reference genome, usually used for aligning and annotating mutation positions to help with accurate sequence analysis. The Bayesian statistical model is a statistical model that infers unknown parameters by combining prior knowledge with observed data through Bayes' theorem and is commonly used in genomics for mutation prediction and uncertainty processing in sequence analysis.

[0067] In an alternative embodiment,

[0068] Obtain the original gene sequencing data of the individual to be detected, perform multi-dimensional quality assessment on the original gene sequencing data to generate a quality assessment index set, filter the quality assessment index set through a double-threshold dynamic quality control algorithm to obtain high-quality sequence data, and perform deduplication operations through an improved locality-sensitive hashing algorithm to obtain non-redundant sequence data. Based on a distributed computing framework, construct a sequence comparison system, generate a reference genome, and combine the Burrows-Wheeler transform algorithm and a fast comparison algorithm based on suffix arrays to compare the non-redundant sequence data with the reference genome, generate a sequence position mapping file, and perform variant detection through a multi-level variant detection strategy to generate an initial variant detection result. Perform credibility scoring on the initial variant detection result through a Bayesian statistical model to generate a standard variant detection report including credibility scores, including:

[0069] Obtain the original gene sequencing data of the individual to be detected measured by a high-throughput sequencing instrument, perform multi-dimensional quality assessment on the original gene sequencing data through the FastQC tool, and screen the original gene sequencing data according to pre-set assessment indicators to obtain the quality assessment index set;

[0070] According to the quality assessment index set, set a first quality threshold and a second quality threshold according to the double-threshold dynamic quality control algorithm, where the first quality threshold is less than the second quality threshold. For each gene sequence in the quality assessment index set, calculate the average quality value and compare it with the first quality threshold and the second quality threshold. If the average quality value is less than the first quality threshold, filter the current gene sequence. If the average quality value is greater than the second quality threshold, retain the current gene sequence. If the average quality value is greater than the first quality threshold and less than the second quality threshold, excise the bases at both ends of the current gene sequence whose quality values are lower than the first quality threshold and retain the middle segment. Output the retained middle segment and gene sequence as high-quality sequence data;

[0071] For the high-quality sequence data, extract the base segments in each gene sequence as key values through the locality-sensitive hashing algorithm, use the gene sequence as the corresponding value to construct a hash table, judge the sequence repeatability by searching the hash table, and remove the repeated sequences in the high-quality sequences to obtain non-redundant sequence data;

[0072] For the non-repetitive sequence data, it is fragmented according to a pre-set fragmentation rule, generating multiple data fragments and sending them to a pre-constructed sequence comparison system based on a distributed computing framework. On the computing nodes in the sequence comparison system, a reference genome is generated and a rough comparison is performed based on the Burrows-Wheeler transform algorithm. Candidate regions are determined on the reference genome according to the rough comparison results, and an accurate comparison is performed in combination with a fast comparison algorithm based on a suffix array, outputting an initial comparison result. The initial comparison result is quality-evaluated, generating a read quality score and determining whether to save the current initial comparison result according to a pre-set read quality threshold. The retained initial comparison results are merged to generate a sequence position mapping file;

[0073] Based on the sequence position mapping file, the gene sequences in the non-repetitive sequence data are mapped to the reference genome, local reassembly is performed on the sequences mapped to the reference genome, and single nucleotide polymorphism sites and insertion / deletion sites are detected in combination with a global debiasing strategy. The confidence level of each site is evaluated and false positive sites are filtered to obtain an initial variant detection result;

[0074] Based on the initial variant detection result, the sequence depth, quality value, and mutation frequency of each variant site in the gene sequence are evaluated, a probability score corresponding to each variant site is calculated, and the variant sites are sorted based on the probability score. The variant sites with the highest probability scores are selected to generate a standard variant detection report.

[0075] The high-throughput sequencing instrument is an instrument capable of simultaneously sequencing a large number of DNA or RNA samples, usually having a high sequencing efficiency and being used to quickly obtain genomic or transcriptomic data. The FastQC tool is a software tool for quality control of high-throughput sequencing data, capable of quickly evaluating the quality of raw sequencing data and generating a report to help analyze data quality problems. The base slice refers to a short sequence fragment obtained in gene sequencing, consisting of a series of bases, usually used for assembling and aligning genomes. The hash table is a common data structure for storing key-value pairs, capable of efficiently performing data lookup and indexing, and is widely used in the storage and retrieval of gene sequence data. Fragmentation is the process of splitting a long gene sequence into smaller fragments for sequencing in gene sequencing, usually helping to improve the efficiency and accuracy of sequencing. The global debiasing strategy is a method for removing data bias, commonly used in genomic data processing, and systematically adjusts or corrects biases through algorithms to ensure data accuracy. False positive sites refer to sites that are erroneously identified as variants in gene variant detection, and such errors are usually reduced through quality control and algorithm optimization. The sequence depth refers to the number of times a certain position is sequenced in gene sequencing, and usually a higher sequence depth can provide more accurate variant detection results.

[0076] Obtain the original gene sequencing data of the individual to be detected. The sequencing data can be sourced from various high-throughput sequencing platforms, such as Illumina, PacBio, or Nanopore, etc. The data format can be FASTQ, BAM, or CRAM, etc. For example, obtain a FASTQ file named sample.fastq from the Illumina platform, which contains the original gene sequencing data of the individual to be detected.

[0077] Conduct multi-dimensional quality assessment on the original gene sequencing data to generate a quality assessment index set. Use quality assessment tools such as FastQC to analyze the original data. The assessment indexes include the distribution of sequencing quality values, GC content distribution, adapter sequence content, repetitive sequence content, etc. For example, use FastQC to analyze the sample.fastq file and generate a quality report in HTML format, which contains the statistical results of various quality indexes, such as the average quality value of each sequencing cycle, the quality value distribution of each base position, etc. These indexes constitute the quality assessment index set.

[0078] Filter the data of the quality assessment index set through a double-threshold dynamic quality control algorithm to obtain high-quality sequence data. Set the first quality threshold, such as 15, and the second quality threshold, such as 30. Traverse each sequencing sequence and calculate its average quality value. If the average quality value is less than the first quality threshold, discard the sequence; if the average quality value is greater than the second quality threshold, retain the sequence; if the average quality value is between the first quality threshold and the second quality threshold, start checking each base from both ends of the sequence, excise the bases with quality values lower than the first quality threshold, and retain the middle fragments with qualified quality. For example, the average quality value of a sequence is 25, which is between 15 and 30. Some bases at both ends of the sequence have quality values lower than 15, so these low-quality bases are excised, and the middle fragments with qualified quality are retained. These fragments and the sequences that pass completely are output to the clean.fastq file as high-quality sequence data.

[0079] Duplication removal operation is performed through an improved locality-sensitive hashing algorithm to obtain duplicate-free sequence data. The high-quality sequence data is processed by locality-sensitive hashing (LSH). Each sequence is split into several short fragments (k-mers) of length k, for example, k = 10. A part of the k-mers is selected as the hash key value, and the sequence itself is used as the value to construct a hash table. If two sequences have the same hash key value, they are considered potentially duplicate, and further sequence alignment is performed for confirmation. If the sequences are exactly the same, one of the duplicate sequences is removed. The deduplicated sequence data is output to the dedup.fastq file. For example, if two sequences both contain the same k-mer "ATCGGCTAGC", they are marked as potentially duplicate sequences and further aligned. If the sequences are exactly the same, one of them is removed.

[0080] A sequence comparison system is constructed based on a distributed computing framework. A reference genome is generated and combined with the Burrows-Wheeler transform algorithm and a fast comparison algorithm based on the suffix array to compare the duplicate-free sequence data with the reference genome, generating a sequence position mapping file. The sequence data in the dedup.fastq file is split into multiple data shards and distributed to each node of the distributed computing cluster for processing. On each node, the reference genome data is loaded, such as the human reference genome hg38. The reference genome is indexed using the Burrows-Wheeler transform algorithm, and then the sequencing sequences are roughly aligned with the reference genome to determine the candidate alignment regions. In the candidate regions, a fast comparison algorithm based on the suffix array is used for fine alignment to obtain accurate alignment position information. The alignment results on each node are merged to generate a sequence position mapping file align.bam, which records information such as the alignment position and alignment quality of each sequence on the reference genome.

[0081] Variant detection is performed through a multi-level variant detection strategy to generate an initial variant detection result. According to the sequence alignment information in the align.bam file, the sequences mapped to the reference genome are locally reassembled to improve the accuracy of variant detection. A multi-level variant detection strategy is adopted, including single nucleotide polymorphism (SNP) detection and insertion-deletion (Indel) detection. The confidence of each potential variant site is evaluated, and the false positive sites with low confidence are filtered out. The detected variant information is output to the variants.vcf file.

[0082] The initial variant detection results are scored for credibility using a Bayesian statistical model to generate a standard variant detection report containing credibility scores. For each variant site in the variants.vcf file, its credibility score is calculated using a Bayesian statistical model based on information such as sequencing depth, quality value, and mutation frequency. The credibility scores are added to the variants.vcf file to generate the final standard variant detection report report.vcf, which contains detailed information for each variant site, such as location, type, credibility score, etc.

[0083] In this embodiment, through technical means such as multi-dimensional quality control, locality-sensitive hashing deduplication, distributed computing, multi-level variant detection strategies, and Bayesian statistical models, the false positive and false negative rates can be effectively reduced, and the accuracy of variant detection can be improved. The sequence alignment system based on the distributed computing framework can process a large amount of sequencing data in parallel, significantly improving the efficiency of variant detection. The generated standard variant detection report contains detailed information and credibility scores for each variant site, providing more comprehensive variant information for subsequent genomic analysis and clinical applications.

[0084] In an alternative embodiment,

[0085] Based on the sequence position mapping file, the gene sequences in the non-redundant sequence data are mapped to the reference genome, and local reassembly is performed on the sequences mapped to the reference genome. Combining the global debiasing strategy, single nucleotide polymorphism sites and insertion / deletion sites are detected, and the confidence of each site is evaluated and false positive sites are filtered, including:

[0086] Read the non-redundant sequence data, the reference genome sequence, and the sequence position mapping file. The sequence position mapping file contains the mapping position information and matching degree information of the reads on the reference genome. Extract the read sequences according to the sequence position mapping file, align the read sequences to the corresponding positions on the reference genome, and parse the alignment string to determine the base matching relationship between the read sequences and the reference genome;

[0087] Perform local reassembly on the read sequences aligned to the same region with a ten-base window, use the minimum edit distance algorithm to find the optimal alignment method, correct the alignment offset of the read sequences, obtain the reassembled read sequences, and calculate the likelihood probability of the site bases using a Bayesian statistical model based on the base types and quality values supported by the reassembled read sequences. Combine the reference genome sequence to judge the variant sites;

[0088] A calibration model is established. The calibration model includes sequencing quality parameters, base content parameters, and read depth parameters. The likelihood probability of the variant site is calibrated according to the calibration model to obtain a calibrated likelihood probability. The variant sites are screened according to the calibrated likelihood probability. When the variant site is a heterozygous variant, the variant sites with variant base frequencies between 20% and 80% are retained. When the variant site is a homozygous variant, the variant sites with variant base frequencies greater than 80% are retained;

[0089] Calculate the quality value, sequencing depth, and number of reads supporting the variant of the variant site. The single nucleotide polymorphism sites with a quality value greater than 30 and a sequencing depth greater than ten times, and the insertion and deletion sites with a quality value greater than 50 and a sequencing depth greater than twenty times are used as high-confidence variant sites. The high-confidence variant sites are annotated to generate a variant annotation file including genomic position information, variant type information, allele frequency information, and functional impact information.

[0090] The minimum edit distance algorithm is an algorithm used to calculate the minimum number of edit operations (such as insertion, deletion, or substitution) required between two sequences to transform one sequence into another sequence. It is widely used in gene sequence alignment. The reassembled read sequence refers to the process of reassembling multiple short gene sequence fragments into a complete gene sequence through the reassembly method in genomics, which is usually used to solve the fragmentation problem in high-throughput sequencing.

[0091] Read the non-redundant sequence data, reference genome sequence, and sequence position mapping file. The sequence position mapping file contains mapping position information, matching degree information, and other information that may be helpful for variant detection of the reads on the reference genome, such as read length, direction, etc. Taking the variant detection of the human genome as an example, the reference genome can use the human genome reference sequence hg38, and the non-redundant sequence data refers to the sequencing data after removing PCR duplicate reads. Suppose the information of a read in the sequence position mapping file is: chr1, 1000, +, 95%, which means that this read is aligned to the chr1 chromosome of the hg38 reference genome, the starting position is 1000, the direction is the positive strand, and the matching degree is 95%. According to this information, the corresponding read sequence is extracted from the sequencing data file.

[0092] Align the read sequences to the corresponding positions on the reference genome. Use alignment tools such as BWA and Bowtie2 to align the extracted read sequences to the specified positions on the reference genome. For example, align the above reads to the vicinity of the starting position 1000 on chromosome 1 of the hg38 reference genome. After alignment, parse the alignment string, such as "2M1I3M1D2M", where M represents a match, I represents an insertion, and D represents a deletion. Determine the base matching relationship between the read sequence and the reference genome according to the alignment result. For example, assume that the sequence of the corresponding region of the reference genome is "AGCTAGCT" and the read sequence is "AGTCAGCT", then the alignment result may be "2M1I3M".

[0093] Perform local reassembly with a ten-base window. Reassemble the read sequences aligned to the same region of the reference genome locally. For example, assume that three reads are aligned to the same region of the reference genome. With a ten-base sliding window, take out ten bases and their corresponding read sequences each time, and use the minimum edit distance algorithm to find the optimal alignment method. If there is an insertion or deletion at a certain position in one of the reads, the alignment offset of this read can be corrected through alignment with other reads to obtain the reassembled read sequence. Assume that the reference sequence within the window is "AGCTAGCTAG", and the three read sequences are "AGCTAGCTAG", "AGCTCAGCTAG", and "AGCTAGCTAG" respectively. Then the "C" at position 2 in the second read sequence will be recognized as an insertion.

[0094] Calculate the likelihood probability of the site base using a statistical model based on the base types and quality values supported by the reassembled read sequences, and combine with the reference genome sequence to judge the variant sites. Assume that at a certain site, the base in the reference genome is A, and among the reassembled read sequences, 8 reads support base A and 2 reads support base T, and the quality values of the reads supporting base T are relatively high. Then it can be preliminarily judged that there may be an A>T variant at this site.

[0095] Build a correction model. The correction model includes sequencing quality parameters, base content parameters, and read depth parameters. The sequencing quality parameter can be represented by the average sequencing quality value, the base content parameter can be represented by the GC content, and the read depth parameter can be represented by the coverage of this site. Correct the likelihood probability of the variant site according to the correction model. For example, if the sequencing depth of a certain site is very low, the probability of a variant occurring at this site will decrease. Assume that the uncorrected likelihood probability of a site is 0.9, and after correction, the likelihood probability drops to 0.7. Screen the variant sites according to the corrected likelihood probability. When the variant site is a heterozygous variant, retain the variant sites with variant base frequencies between 20% and 80%; when the variant site is a homozygous variant, retain the variant sites with variant base frequencies greater than 80%.

[0096] Calculate the quality value, sequencing depth, and the number of reads supporting the variant at the variant site. Single nucleotide polymorphism (SNP) sites with a quality value greater than 30 and a sequencing depth greater than 10-fold, as well as insertion-deletion (Indel) sites with a quality value greater than 50 and a sequencing depth greater than 20-fold, are considered high-confidence variant sites. For example, if a single nucleotide polymorphism site has a quality value of 40, a sequencing depth of 20, and 18 reads supporting the variant, then this site is considered a high-confidence single nucleotide polymorphism site.

[0097] In this embodiment, through local reassembly and global debiasing strategies, false positive variants caused by sequencing errors and alignment errors can be effectively reduced, and the accuracy of variant detection can be improved. By combining multiple parameters to establish a correction model, low-frequency variants and complex variants can be effectively identified, and the sensitivity of variant detection can be improved. By annotating high-confidence variant sites, rich variant information can be provided, facilitating subsequent analysis and research by users.

[0098] S2. Based on the standard variant detection report, according to the improved random field diffusion model in the pre-set self-activation depth analysis system, combined with Langevin dynamics and Hamiltonian Monte Carlo methods for multi-scale feature sampling, generate an initial feature distribution. Optimize the initial feature distribution based on the adaptive adjustment mechanism and Wasserstein distance to generate a high-dimensional feature distribution and input it into the pre-set dual self-activation iterative network. Extract spatial features through the dynamic routing algorithm and attention guidance mechanism in the excitation path based on the capsule network, add the spatial features to the suppression path and perform contrast learning and competitive collaboration to obtain balanced features. Select the optimal features according to the Boltzmann machine and annealing algorithm in the metastable state, add the optimal features to the dynamic system based on Lie differential equations, combine neural ordinary differential equation networks to model the temporal evolution and perform feature integration to obtain a feature mapping matrix and a feature evolution trajectory;

[0099] The self-activating depth analysis system is a deep learning system that continuously adjusts and optimizes model parameters through a self-activating mechanism, aiming to enhance the model's analysis ability through a dynamic feedback mechanism, especially showing higher adaptability when dealing with complex data analysis tasks. The improved random field diffusion model is an algorithm that improves the traditional random field model to more precisely simulate the data diffusion process, commonly used in image processing, data transmission, and the modeling of propagation processes, and can effectively capture the spatial or temporal dependencies of data. The Langevin dynamics is a kinetic model that describes the motion of particles under the action of random forces, commonly used in statistical physics and computational biology to simulate the evolution process of a system in a random environment. The Hamiltonian Monte Carlo method is a method that combines Hamiltonian dynamics with Monte Carlo sampling, mainly used in high-dimensional integration problems, and through simulating the evolution of particles in Hamiltonian dynamics, it conducts effective sampling and inference, widely applied in physics, statistics, and Bayesian inference. The Wasserstein distance is a distance metric method used to measure the difference between two probability distributions, commonly used in generative models and data matching problems, especially widely applied in image processing and generative adversarial networks. The dual self-activating iterative network is an iterative optimization structure that combines a self-activating mechanism with a multi-level network structure, and through dual feedback and adaptive adjustment, enhances the robustness and processing ability of the model to input data. The excitation pathway based on the capsule network is a network structure based on the capsule network, aiming to extract feature information from data through the dynamic routing mechanism of the capsule network and conduct feature transfer through the excitation pathway, usually used to handle complex structure and spatial transformation problems, such as image recognition and classification. The Boltzmann machine is a randomly generated neural network model, usually used for unsupervised learning, and can learn the latent representation of data by simulating the state transition between neurons, widely applied in pattern recognition and image processing. The Lie differential equation is a class of differential equations used to describe the change of a system over time, especially used in physics and engineering to model the behavior of dynamic systems, widely applied in control theory and signal processing. The neural ordinary differential equation network is a network structure that combines a neural network with an ordinary differential equation, aiming to simulate a dynamic system by solving the ordinary differential equation, widely applied in fields such as time series prediction, physical system modeling, and neural control.

[0100] In an alternative embodiment,

[0101] Based on the standard variant detection report, multi-scale feature sampling is performed according to the improved random field diffusion model in the pre-set self-activation depth analysis system, combining Langevin dynamics and Hamiltonian Monte Carlo methods to generate an initial feature distribution. Based on the adaptive adjustment mechanism and Wasserstein distance, the initial feature distribution is optimized to generate a high-dimensional feature distribution and input it into the pre-set dual self-activation iterative network. Spatial features are extracted through the dynamic routing algorithm and attention guidance mechanism in the excitation path based on the capsule network. The spatial features are added to the inhibition path and contrast learning and competitive collaboration are performed to obtain balanced features. The optimal features are selected according to the Boltzmann machine and annealing algorithm in the metastable state. The optimal features are added to the dynamic system based on Lie differential equations, and the time series evolution is modeled and feature integration is performed through the neural ordinary differential equation network to obtain a feature mapping matrix and a feature evolution trajectory, including:

[0102] Based on the standard variant detection report, multi-scale feature sampling is performed according to the improved random field diffusion model in the pre-set self-activation depth analysis system. Among them, the improved random field diffusion model uses the Langevin dynamics method to describe the movement trajectory of particles on the potential energy surface, combines the Hamiltonian Monte Carlo method to accept sub-optimal sampling points with a preset probability, and samples variant features at the chromosome level, gene level, and exon level to generate an initial feature distribution;

[0103] Based on the adaptive adjustment mechanism and Wasserstein distance, the initial feature distribution is optimized. Among them, the adaptive adjustment mechanism dynamically adjusts the temperature parameter and step size parameter according to the sampling acceptance rate. The Wasserstein distance is used to measure the difference between the current distribution and the ideal distribution, and the distribution is optimized by minimizing the Wasserstein distance to generate a high-dimensional feature distribution;

[0104] The high-dimensional feature distribution is input into the pre-set dual self-activation iterative network. The dual self-activation iterative network includes an activation path and an inhibition path. In the activation path, the dynamic routing algorithm based on the capsule network is used to adaptively adjust the connection strength between capsules, and the attention mechanism is introduced to dynamically adjust the weight according to the feature correlation to extract spatial features. The spatial features are input into the inhibition path for perturbation, and the original features and perturbed features are discriminated through contrast learning. The balanced features are obtained through the competition and collaboration of the dual paths;

[0105] A Boltzmann machine is constructed, the weight matrix is defined according to the feature correlation, and the energy function is defined according to the disease risk. The annealing algorithm is used to iteratively sample the balanced features in the metastable state to obtain the optimal features with the lowest energy;

[0106] Add the optimal features to the dynamic system based on Lie differential equations, and combine the neural ordinary differential equation network to perform temporal modeling on the dynamic system. Consider each network layer as a time step for solving ordinary differential equations, sample at different time steps to obtain the feature mapping matrix, and generate the feature evolution trajectory.

[0107] The sub-optimal sampling points refer to the points sampled during the optimization process that are close to the optimal solution but not necessarily completely optimal. They are often used in solving large-scale optimization problems to accelerate convergence and reduce computational complexity.

[0108] Preset a self-activation depth analysis system, and build an improved random field diffusion model on this basis. The model simulates the movement trajectory of particles on the potential energy surface, uses the Langevin dynamics method to describe the movement process of particles, and combines the Hamiltonian Monte Carlo method to accept non-optimal sampling points with a certain probability. Taking a variant detection report containing 100 samples as an example, 50, 30, and 20 variant features are collected at the chromosome level, gene level, and exon level respectively to form the initial feature distribution. For example, at the chromosome level, chromosome copy number variant features can be collected; at the gene level, gene mutation frequency features can be collected; at the exon level, exon expression level features can be collected.

[0109] Optimize the obtained initial feature distribution. Adopt an adaptive adjustment mechanism to dynamically adjust the temperature parameter and step size parameter in Langevin dynamics according to the sampling acceptance rate. For example, if the sampling acceptance rate is too low, reduce the temperature parameter and decrease the step size parameter. At the same time, use the Wasserstein distance to measure the difference between the current distribution and the ideal distribution (e.g., the standard normal distribution), and optimize the feature distribution by minimizing the Wasserstein distance, finally generating a high-dimensional feature distribution. Suppose the Wasserstein distance between the initial distribution and the ideal distribution is 0.8, and after optimization, the distance is reduced to 0.2.

[0110] Input the optimized high-dimensional feature distribution into a pre-set dual self-activation iterative network. This network includes an excitation pathway and an inhibition pathway. In the excitation pathway, adopt a dynamic routing algorithm based on the capsule network to adaptively adjust the connection strength between capsules according to the similarity between features, and introduce an attention mechanism to dynamically adjust the weights according to feature correlation to extract spatial features. For example, if two features are highly correlated, both the connection strength and the weights between them will increase. Input the extracted spatial features into the inhibition pathway for perturbation, for example, adding random noise. Through contrastive learning, compare the original features and the perturbed features, and through the competition and cooperation of the two pathways, obtain balanced features.

[0111] Construct a Boltzmann machine, define the weight matrix using feature correlation, and define the energy function using disease risk. Use the annealing algorithm to iteratively sample the equilibrium features in the metastable state to obtain the optimal features with the lowest energy. For example, set the initial temperature to 100, reduce the temperature in each iteration until the temperature reaches 1. At each temperature, perform multiple samplings and select the feature with the lowest energy.

[0112] Input the optimal features into a dynamic system based on the Lie differential equation. Combine the neural ordinary differential equation network to perform temporal modeling on this dynamic system. Consider each network layer as a time step for solving the ordinary differential equation, sample at different time steps to obtain the feature mapping matrix and generate the feature evolution trajectory. For example, 10 time steps can be set, and features are collected at each time step to form a feature mapping matrix with 10 rows and n columns, where n is the dimension of the optimal features.

[0113] In this embodiment, through multi-scale feature sampling and the adaptive optimization mechanism, it is possible to capture variant features more comprehensively and accurately, avoid information loss, and improve the efficiency and accuracy of feature extraction. The competitive cooperation mechanism of the dual self-activation iterative network and the metastable state selection strategy of the Boltzmann machine can effectively suppress the influence of noise and outliers, enhance the robustness and generalization ability of the model. By using the neural ordinary differential equation network to model the feature evolution trajectory, it is possible to more accurately depict the dynamic process of disease development and achieve accurate prediction of disease risk.

[0114] In an alternative embodiment,

[0115] Adding the optimal features to the dynamic system based on the Lie differential equation, and combining the neural ordinary differential equation network to perform temporal modeling on the dynamic system. Considering each network layer as a time step for solving the ordinary differential equation, sampling at different time steps to obtain the feature mapping matrix and generate the feature evolution trajectory includes:

[0116] Obtain the values of the optimal features at different time steps and construct a system state vector. Based on the system state vector, establish a Lie differential equation dynamic system, calculate the interaction strength between state variables to generate a state transition matrix, and calculate the influence degree of external input on state variables to generate a control matrix;

[0117] Fuse the set corresponding to the optimal features with the system state vector, calculate the interaction strength between the fused features and the original features, and update the state transition matrix and the control matrix according to the interaction strength;

[0118] Build a neural ordinary differential equation network that includes an input layer, multiple hidden layers, and an output layer. In each hidden layer, multiply the previous hidden state by a weight matrix and add a bias vector, then input the calculation result into an activation function to obtain the time derivative of the current hidden state. Update the current hidden state based on the time derivative through a numerical integration method. Map the state transition matrix and the control matrix to the network parameter space to complete network initialization;

[0119] Use the backpropagation algorithm to calculate the parameter gradients and iteratively update the network parameters. Perform the forward propagation process at a preset number of time steps, obtain and save the hidden states corresponding to each time step, extract the hidden states at a preset sampling interval in the time series, concatenate the extracted hidden states in chronological order to generate a feature mapping matrix, and smooth the feature mapping matrix to obtain a feature evolution trajectory.

[0120] The numerical integration method refers to the process of approximately solving definite integrals through numerical algorithms, which is often used for complex integral problems that cannot be analytically solved, such as numerical methods like Monte Carlo integration, Simpson's method, and trapezoidal method, and is widely used in scientific computing and engineering analysis.

[0121] Select the features that contribute the most to the target task from the dataset, and these features are called optimal features. For example, in an image classification task, the optimal features can be the edges, textures, colors, etc. of the image. Suppose we select three optimal features: edge intensity, texture complexity, and color saturation. Record the values of these features at different time steps. For example, in video analysis, we can record the values of these three features every second. Combine these feature values into a vector, called the system state vector. For example, at the t-th second, the system state vector can be represented as [edge intensity(t), texture complexity(t), color saturation(t)]. Suppose at the three time steps of 0 second, 1 second, and 2 seconds, the values of these three features are [1, 0.5, 0.8], [1.2, 0.6, 0.9], and [1.1, 0.7, 0.7] respectively.

[0122] Construct a Lie differential equation dynamic system based on the system state vector. By analyzing the changes in the system state vector at different time steps, the interaction strength between state variables can be calculated. For example, if the edge intensity increases while the texture complexity also increases, it indicates a positive correlation between these two features. Represent these interaction strengths as a matrix, called the state transition matrix. At the same time, the influence of external inputs on state variables also needs to be considered. For example, in video analysis, the external input can be the movement of the camera or the change in lighting. By analyzing the relationship between the external input and the state variables, the influence degree of the external input on the state variables can be calculated. Represent these influence degrees as a matrix, called the control matrix.

[0123] Fuse the set corresponding to the optimal features with the system state vector. For example, the optimal feature set can be represented as a vector, and then this vector is concatenated with the system state vector. Calculate the interaction strength between the fused features and the original features. For example, methods such as correlation analysis or mutual information calculation can be used to calculate the interaction strength between features. Update the state transition matrix and the control matrix according to the calculated interaction strength. For example, if there is a strong interaction between the fused features and the original features, the elements in the state transition matrix and the control matrix need to be adjusted accordingly.

[0124] Build a neural ordinary differential equation network including an input layer, multiple hidden layers, and an output layer. In each hidden layer, multiply the previous hidden state by the weight matrix and add the bias vector, and input the calculation result into the activation function to obtain the time derivative of the current hidden state. Based on the time derivative, update the current hidden state through a numerical integration method (such as Euler's method or Runge-Kutta method). Map the state transition matrix and the control matrix to the network parameter space to complete network initialization. For example, the elements of the state transition matrix can be used as the initial values of some weight parameters in the network.

[0125] Use the backpropagation algorithm to calculate the parameter gradients and iteratively update the network parameters. Perform the forward propagation process at multiple preset time steps, and obtain and save the hidden states corresponding to each time step. For example, assume the preset time steps are 0 seconds, 1 second, and 2 seconds, then the hidden states corresponding to these three time steps need to be saved. Extract the hidden states at the preset sampling interval in the time series. For example, if the sampling interval is 1 second, then the hidden states corresponding to 0 seconds, 1 second, and 2 seconds need to be extracted. Concatenate the extracted hidden states in chronological order to generate a feature mapping matrix. Smooth the feature mapping matrix, for example, using methods such as moving average or Gaussian filtering, to obtain the feature evolution trajectory.

[0126] In this embodiment, by introducing the Lie differential equation dynamic system, the time-dependent relationship between features can be better captured, thereby improving the model's ability to model time series data, and further enhancing the model's prediction accuracy. The Lie differential equation dynamic system can provide a clear description of the feature evolution process, making the model's decision-making process more transparent, thereby enhancing the interpretability of the model. By treating each network layer as a time step for solving ordinary differential equations, the problem of needing to calculate step by step at each time step in traditional recurrent neural networks can be avoided, thereby reducing the computational complexity of the model and improving the training efficiency of the model.

[0127] S3. Add the feature mapping matrix and the feature evolution trajectory to a pre-set integrated predictor, where the integrated predictor includes a deep belief network and a hypergraph neural network. Optimize the parameters of the integrated predictor according to an alternating training strategy and calculate an initial risk score. Analyze the feature evolution trajectory based on a mixture Gaussian process regression and predict the risk change trend. Add the initial risk score and the risk change trend to an improved Bayesian inference network, execute a Markov blanket model to update the risk probability, correct the probability estimation bias through a particle filter algorithm, and establish a risk association network among mutation sites. Train a deep hierarchical decision tree based on the risk association network, perform node division in combination with an adaptive splitting threshold algorithm, obtain the risk level, and output a multi-dimensional risk assessment report corresponding to the current individual.

[0128] The deep belief network is a generative neural network composed of multiple layers, usually used for unsupervised learning. By training layer by layer and learning the probability distribution of input data, it is widely applied in fields such as image recognition, speech recognition, and data dimensionality reduction. The hypergraph neural network is a structure that extends the traditional graph neural network. By introducing hyperedges (i.e., edges connecting multiple nodes) in the graph structure, it can effectively capture more complex relationships and interactions, and is particularly suitable for processing data with multiple relationships and complex networks. The mixture Gaussian process regression is a regression model that combines Gaussian processes. By mixing multiple Gaussian process models, it enhances the fitting ability for complex data, and is particularly suitable for function modeling with high uncertainty, such as time series prediction and spatial data analysis. The Bayesian inference network is a probability inference network based on Bayes' theorem. By inferring the probability distribution of unobserved variables through the conditional probability relationships between nodes, it is widely applied in medical diagnosis, decision support systems, and inference problems in machine learning. The Markov blanket model is a graph model used to describe the conditional independence structure of a variable. The Markov blanket refers to the independence relationship between a node and the remaining nodes under certain conditions, and is used for probabilistic graph models and causal inference analysis. The particle filter algorithm is a recursive algorithm for state estimation of nonlinear and non-Gaussian dynamic systems. By representing the system state as a set of particles and updating the particle weights according to the observed data, it is widely applied in fields such as target tracking, navigation, and signal processing. The risk association network is a network model used to analyze and model the relationships between different risk factors. By identifying the mutual influences and propagation paths of risk factors, it helps decision-makers evaluate and manage potential risks in complex systems, and is often used in the fields of finance, engineering, and security.

[0129] In an optional implementation manner,

[0130] Add the feature mapping matrix and the feature evolution trajectory to a pre-set integrated predictor, which includes a deep belief network and a hypergraph neural network. Optimize the parameters of the integrated predictor according to an alternating training strategy and calculate an initial risk score. Analyze the feature evolution trajectory based on a mixture Gaussian process regression to predict the risk change trend. Add the initial risk score and the risk change trend to an improved Bayesian inference network, execute a Markov blanket model to update the risk probability, correct the probability estimation bias through a particle filter algorithm, and establish a risk association network among mutation sites. Train a deep hierarchical decision tree based on the risk association network, perform node division in combination with an adaptive splitting threshold algorithm, obtain the risk level, and output a multi-dimensional risk assessment report corresponding to the current individual, including:

[0131] Add the feature mapping matrix and the feature evolution trajectory as input samples to a pre-set integrated predictor, which includes a deep belief network and a hypergraph neural network. The deep belief network is composed of multiple stacked restricted Boltzmann machines, and the hypergraph neural network maps features to a high-order hypergraph space and models the interaction between features through hyperedge connections;

[0132] Optimize the parameters of the integrated predictor using an alternating training strategy. Fix the parameters of the deep belief network and update the parameters of the hypergraph neural network, then fix the parameters of the hypergraph neural network and update the parameters of the deep belief network. Iteratively train until convergence and then predict the input samples to obtain an initial risk score;

[0133] Construct a mixture Gaussian process regression model composed of multiple sub-Gaussian processes. Use the expectation-maximization algorithm to iteratively optimize the parameters of the mixture Gaussian process regression model. Calculate the posterior probability that the input samples belong to different sub-Gaussian processes and update the hyperparameters of each sub-Gaussian process. Predict the feature evolution trajectory based on the trained mixture Gaussian process regression model to obtain the risk change trend;

[0134] Construct an improved Bayesian inference network, which includes multiple time slices. Each time slice corresponds to an independent Bayesian network, and adjacent time slices are connected through conditional probability distributions. Set the prior probability of risk nodes in the initial time slice according to the initial risk score. Calculate the second posterior probability corresponding to each node using the Markov blanket model. Correct the second posterior probability according to the particle filter algorithm. Approximate the posterior probability distribution through weighted sampling particles. Update the particles according to the state transition model and adjust the particle weights according to the observation model. Calculate the risk correlation among mutation sites based on the corrected second posterior probability and establish a risk association network;

[0135] Construct a deep hierarchical decision tree, optimize node splitting using an adaptive splitting threshold algorithm, dynamically adjust the splitting threshold according to node purity and sample distribution, train the deep hierarchical decision tree based on the risk association network, predict the risk level of the input sample through the deep hierarchical decision tree and analyze key variant combinations, obtain the risk level corresponding to the input sample, and synthesize the risk level of the input sample and the key variant combinations to obtain a multi-dimensional risk assessment report.

[0136] The restricted Boltzmann machine is an unsupervised learning model and a simplified form of the Boltzmann machine. It consists of a visible layer and a hidden layer, where the layers are connected by symmetric weights. The high-order hypergraph space is an extension in hypergraph theory, which not only considers the direct relationships between nodes but also the multi-tuple relationships of high-order associations. It is suitable for representing more complex multi-dimensional and multi-variate relationships and is commonly used in fields such as social networks, knowledge graphs, and bioinformatics. The hyperedge is a concept in hypergraphs, representing an edge that connects multiple nodes, different from the ordinary edge in a traditional graph (usually connecting two nodes). A hyperedge can contain any number of nodes and is used to express more complex many-to-many relationships and is widely applied in high-dimensional data analysis, relationship modeling, and graph mining. The time slice is a commonly used technique in data analysis, used to divide continuous time series data into discrete time periods for time series analysis or modeling. The adaptive splitting threshold algorithm is an algorithm for dynamically adjusting thresholds and is usually used in data segmentation, classification, and clustering tasks.

[0137] Construct an ensemble predictor composed of a deep belief network and a hypergraph neural network. The deep belief network is stacked by multiple restricted Boltzmann machines and is used to learn the deep representation of features. The hypergraph neural network maps features to the high-order hypergraph space and models the complex interactions between features through hyperedge connections. For example, take gene loci as nodes and the set of gene loci with interactions as hyperedges to construct a gene interaction hypergraph. Use individual genotype data and physiological index data as inputs, and alternately train the deep belief network and the hypergraph neural network. Fix the parameters of the deep belief network and update the parameters of the hypergraph neural network using the gradient descent method; then fix the parameters of the hypergraph neural network and update the parameters of the deep belief network using the contrastive divergence algorithm. Iteratively train until the model converges to obtain an initial risk score. For example, the initial risk score of individual A is 0.6, indicating that the initial risk of this individual suffering from a certain disease is 60%.

[0138] Utilize the mixture Gaussian process regression to analyze the feature evolution trajectory and predict the risk change trend. The mixture Gaussian process regression model consists of multiple sub-Gaussian processes, and each sub-Gaussian process represents a risk change pattern. The expectation-maximization algorithm is used to iteratively optimize the model parameters, calculate the posterior probability that the input sample belongs to different sub-Gaussian processes, and update the hyperparameters of each sub-Gaussian process (such as the mean function, covariance function, etc.). For example, input the blood pressure feature evolution trajectory into the trained mixture Gaussian process regression model to predict the blood pressure change trend in the next year, and further predict the change trend of an individual's risk of cardiovascular disease.

[0139] Construct an improved Bayesian inference network, which contains multiple time slices. Each time slice corresponds to an independent Bayesian network, and adjacent time slices are connected by conditional probability distributions. Set the prior probability of the risk node in the initial time slice according to the initial risk score. For example, set the initial risk score 0.6 of individual A as the prior probability of the "diseased" node in the initial time slice. Use the Markov blanket model to calculate the second posterior probability corresponding to each node, and use the particle filter algorithm to correct the second posterior probability. Approximate the posterior probability distribution by weighted sampling particles, update the particles according to the state transition model, and adjust the particle weights according to the observation model. Calculate the risk correlation between mutation sites according to the corrected second posterior probability, and establish a risk association network. For example, if it is found that the mutations of gene locus A and gene locus B exist simultaneously, which significantly increases the risk of suffering from a certain disease, then a strong connection between A and B is established in the risk association network.

[0140] Construct a deep hierarchical decision tree. Use the adaptive splitting threshold algorithm to optimize node splitting, and dynamically adjust the splitting threshold according to node purity and sample distribution. Train the deep hierarchical decision tree based on the risk association network. For example, use the gene locus combination with a higher connection strength in the risk association network as the basis for splitting the root node of the decision tree. Predict the risk level of the input sample through the deep hierarchical decision tree and analyze the key mutation combinations. For example, according to the genotype and physiological index data of individual A, predict that the risk level of suffering from a certain disease is high risk, and identify that the mutation combination of gene locus A and gene locus B is the key factor leading to high risk. Generate a multi-dimensional risk assessment report by integrating the risk level and key mutation combinations. For example, the report contains information such as the disease risk level of individual A, key mutation combinations, and personalized health management suggestions.

[0141] In this embodiment, by combining a deep belief network, a hypergraph neural network, a mixture of Gaussian processes regression, and an improved Bayesian inference network, it is possible to more comprehensively capture feature information and the law of feature evolution, improve the accuracy of risk prediction. By analyzing the feature evolution trajectory and constructing an improved Bayesian inference network, the risk change trend can be predicted, realizing the dynamic assessment and early warning of risks. By analyzing key mutation combinations through a deep hierarchical decision tree, a personalized risk assessment report can be provided, offering more accurate health management suggestions for individuals.

[0142] In an alternative embodiment,

[0143] The hyperparameters of each sub-Gaussian process are updated as shown in the following formula

[0144]

[0145] where θ k represents the set of hyperparameters of the k-th sub-Gaussian process, represents the updated set of hyperparameters of the k-th sub-Gaussian process, γ(z nk ) represents the posterior probability that the n-th input sample belongs to the k-th sub-Gaussian process, represents the likelihood function of the sample y n under the k-th sub-Gaussian process, y n represents the observed value of the n-th input sample, δ represents the mean of the sub-Gaussian process distribution, with a value of 0, represents the covariance matrix of the k-th sub-Gaussian process at sample n, σ 2 represents the noise variance, I represents the identity matrix, v represents the degree-of-freedom parameter of the probability density function, represents the probability density function, and N represents the total number of input samples.

[0146] In this embodiment, by dividing the dataset into multiple subsets and using sub-Gaussian processes for modeling respectively, it is possible to better capture the local features in the data, thereby improving the fitting accuracy of the model. Since each sub-Gaussian process is only responsible for fitting a part of the data, overfitting can be effectively avoided, thus enhancing the generalization ability of the model. This means that the model can still maintain good prediction performance when facing new and unseen data. Compared with using a single complex Gaussian process model, using multiple simple sub-Gaussian processes can reduce the computational complexity. Especially in the case of a large dataset scale, this method can significantly reduce the time for model training and prediction, thereby improving efficiency.

[0147] Figure 2 is a schematic structural diagram of the genetic disease gene detection data analysis system based on big data according to the embodiment of the present invention. As Figure 2 shown, the system includes:

[0148] The first unit is used to obtain the original gene sequencing data of an individual to be detected, perform multi-dimensional quality assessment on the original gene sequencing data to generate a quality assessment index set, filter the data of the quality assessment index set through a double-threshold dynamic quality control algorithm to obtain high-quality sequence data, and perform deduplication operations through an improved locality-sensitive hashing algorithm to obtain non-repetitive sequence data. A sequence comparison system is constructed based on a distributed computing framework to generate a reference genome, and the non-repetitive sequence data and the reference genome are compared by combining the Burrows-Wheeler transform algorithm and a fast comparison algorithm based on a suffix array to generate a sequence position mapping file, and mutation detection is performed through a multi-level mutation detection strategy to generate an initial mutation detection result. The initial mutation detection result is scored for credibility through a Bayesian statistical model to generate a standard mutation detection report containing credibility scores;

[0149] The second unit is used to, based on the standard mutation detection report, perform multi-scale feature sampling according to an improved random field diffusion model in a pre-set self-activation depth analysis system, in combination with Langevin dynamics and Hamiltonian Monte Carlo methods, to generate an initial feature distribution. The initial feature distribution is optimized based on an adaptive adjustment mechanism and the Wasserstein distance to generate a high-dimensional feature distribution and input it into a pre-set dual self-activation iterative network. Spatial features are extracted through a dynamic routing algorithm and an attention guidance mechanism in an excitation path based on a capsule network, and the spatial features are added to an inhibition path and subjected to contrast learning and competitive collaboration to obtain balanced features. The optimal features are selected according to a Boltzmann machine and an annealing algorithm in a metastable state, and the optimal features are added to a dynamic system based on Lie differential equations, and the time series evolution is modeled and feature integration is performed in combination with a neural ordinary differential equation network to obtain a feature mapping matrix and a feature evolution trajectory;

[0150] The third unit is used to add the feature mapping matrix and the feature evolution trajectory to a pre-set integrated predictor, which includes a deep belief network and a hypergraph neural network. The parameters of the integrated predictor are optimized according to an alternating training strategy and an initial risk score is calculated. The feature evolution trajectory is analyzed by a mixture Gaussian process regression to predict the risk change trend. The initial risk score and the risk change trend are added to an improved Bayesian inference network, the Markov blanket model is executed to update the risk probability, and the probability estimation deviation is corrected through a particle filter algorithm and a risk association network between mutation sites is established. Based on the risk association network, a deep hierarchical decision tree is trained, and node division is performed in combination with an adaptive splitting threshold algorithm to obtain a risk level and output a multi-dimensional risk assessment report corresponding to the current individual.

[0151] In the third aspect of the embodiments of the present invention,

[0152] There is provided an electronic device, including:

[0153] A processor;

[0154] A memory for storing processor-executable instructions;

[0155] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0156] In a fourth aspect of the embodiments of the present invention,

[0157] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0158] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium on which computer-readable program instructions for performing various aspects of the present invention are loaded.

[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A genetic disease gene detection data analysis method based on big data, characterized in that: include: Obtaining the original gene sequencing data of the individual to be tested, performing a multi-dimensional quality assessment on the original gene sequencing data to generate a quality assessment indicator set, filtering the quality assessment indicator set through a dual-threshold dynamic quality control algorithm to obtain high-quality sequence data, and performing a deduplication operation through an improved local sensitive hashing algorithm to obtain non-repetitive sequence data, building a sequence comparison system based on a distributed computing framework, generating a reference genome, and combining the Burrows-Wheeler transform algorithm and a fast comparison algorithm based on a suffix array to compare the non-repetitive sequence data with the reference genome, generating a sequence position mapping file, and performing variation detection through a multi-level variation detection strategy to generate an initial variation detection result, performing a credibility score on the initial variation detection result through a Bayesian statistical model, and generating a standard variation detection report containing a credibility score; Based on the standard variation detection report, according to the improved random field diffusion model in the pre-set self-activation deep analysis system, multi-scale feature sampling is performed in combination with Langevin dynamics and Hamiltonian Monte Carlo method to generate an initial feature distribution, the initial feature distribution is optimized based on the adaptive adjustment mechanism and Wasserstein distance, a high-dimensional feature distribution is generated and input into a pre-set dual self-activation iterative network, spatial features are extracted through a dynamic routing algorithm and an attention guidance mechanism in an excitation pathway based on a capsule network, the spatial features are added to the inhibition pathway and contrast learning and competitive collaboration are performed to obtain balanced features, the optimal features are selected in a metastable state according to a Boltzmann machine and an annealing algorithm, the optimal features are added to a dynamic system based on a Lie differential equation, the time series evolution is modeled in combination with a neural ordinary differential equation network and features are integrated to obtain a feature mapping matrix and a feature evolution trajectory; The feature mapping matrix and the feature evolution trajectory are added to a pre-set integrated predictor, which includes a deep belief network and a hypergraph neural network. The parameters of the integrated predictor are optimized according to an alternating training strategy and the initial risk score is calculated. The feature evolution trajectory is analyzed according to a mixed Gaussian process regression and the risk change trend is predicted. The initial risk score and the risk change trend are added to an improved Bayesian inference network. The Markov blanket model is executed to update the risk probability, and the probability estimation bias is corrected through a particle filtering algorithm to establish a risk association network between mutation points. A deep hierarchical decision tree is trained based on the risk association network, and node division is performed in combination with an adaptive split threshold algorithm to obtain a risk level and output a multi-dimensional risk assessment report corresponding to the current individual.

2. The method according to claim 1, characterized in that The original gene sequencing data of the individual to be tested is obtained, and a multi-dimensional quality assessment is performed on the original gene sequencing data to generate a quality assessment indicator set. The quality assessment indicator set is filtered through a dual-threshold dynamic quality control algorithm to obtain high-quality sequence data and perform deduplication operation through an improved local sensitive hashing algorithm to obtain non-repetitive sequence data. A sequence comparison system is constructed based on a distributed computing framework, a reference genome is generated, and the non-repetitive sequence data is compared with the reference genome in combination with a Burrows-Wheeler transformation algorithm and a fast comparison algorithm based on a suffix array, a sequence position mapping file is generated, and variation detection is performed through a multi-level variation detection strategy to generate an initial variation detection result, and a credibility score is performed on the initial variation detection result through a Bayesian statistical model, and a standard variation detection report containing a credibility score is generated, including: Obtaining raw gene sequencing data of the individual to be tested measured by a high-throughput sequencing instrument, performing multi-dimensional quality assessment on the raw gene sequencing data by using a FastQC tool, and screening the raw gene sequencing data according to pre-set assessment indicators to obtain the quality assessment indicator set; According to the quality assessment indicator set, a first quality threshold and a second quality threshold are set according to a dual-threshold dynamic quality control algorithm, wherein the first quality threshold is less than the second quality threshold, for each gene sequence in the quality assessment indicator set, an average quality value is calculated and compared with the first quality threshold and the second quality threshold, if the average quality value is less than the first quality threshold, the current gene sequence is filtered, if the average quality value is greater than the second quality threshold, the current gene sequence is retained, if the average quality value is greater than the first quality threshold and less than the second quality threshold, the bases at both ends of the current gene sequence whose quality values ​​are lower than the first quality threshold are cut off and the middle fragment is retained, and the retained middle fragment and gene sequence are output as high-quality sequence data; For the high-quality sequence data, a base fragment in each gene sequence is extracted as a key value by a local sensitive hash algorithm, a hash table is constructed with the gene sequence as a corresponding value, sequence repeatability is determined by searching the hash table, and repeated sequences in the high-quality sequence are removed to obtain non-repeating sequence data; For the non-repetitive sequence data, fragmentation is performed according to a preset fragmentation rule to generate multiple data fragments and send them to a pre-built sequence comparison system based on a distributed computing framework, a reference genome is generated on a computing node in the sequence comparison system, and a rough comparison is performed based on a Burrows-Wheeler transform algorithm, candidate regions are determined on the reference genome according to the rough comparison result, and a precise comparison is performed in combination with a fast comparison algorithm based on a suffix array, an initial comparison result is output, a quality assessment is performed on the initial comparison result, a read quality score is generated, and it is determined whether to save the current initial comparison result according to a preset read quality threshold, the retained initial comparison results are merged, and a sequence position mapping file is generated; Based on the sequence position mapping file, the gene sequence in the non-repetitive sequence data is mapped to the reference genome, the sequence mapped to the reference genome is locally reassembled, single nucleotide polymorphism sites and insertion and deletion sites are detected in combination with a global debiasing strategy, the confidence of each site is evaluated and false positive sites are filtered to obtain an initial variation detection result; Based on the initial variation detection results, the sequence depth, quality value and mutation frequency of each variation site in the gene sequence are evaluated, the probability score corresponding to each variation site is calculated, the variation sites are sorted according to the probability scores, and the variation sites with the highest probability scores are selected to generate a standard variation detection report.

3. The method according to claim 2, characterized in that Based on the sequence position mapping file, the gene sequence in the non-repetitive sequence data is mapped to the reference genome, the sequence mapped to the reference genome is locally reassembled, and the single nucleotide polymorphism sites and insertion and deletion sites are detected in combination with the global debiasing strategy, and the confidence of each site is evaluated and false positive sites are filtered, including: Reading non-repetitive sequence data, reference genome sequence and sequence position mapping file, wherein the sequence position mapping file contains mapping position information and matching degree information of the read segment on the reference genome, extracting the read segment sequence according to the sequence position mapping file, aligning the read segment sequence to the corresponding position of the reference genome, parsing the alignment string to determine the base matching relationship between the read segment sequence and the reference genome; The read sequences aligned to the same region are locally reassembled with a window of ten bases, and the optimal alignment method is found by using the minimum edit distance algorithm, and the alignment offset of the read sequences is corrected to obtain the reassembled read sequences. According to the base types and quality values ​​supported by the reassembled read sequences, the likelihood probability of the bases at the sites is calculated by using the Bayesian statistical model, and the variant sites are determined in combination with the reference genome sequence; Establishing a correction model, wherein the correction model includes a sequencing quality parameter, a base content parameter, and a read depth parameter, correcting the likelihood probability of the variant site according to the correction model to obtain a corrected likelihood probability, screening the variant site according to the corrected likelihood probability, retaining a variant site with a variant base frequency between 20% and 80% when the variant site is a heterozygous variant, and retaining a variant site with a variant base frequency greater than 80% when the variant site is a homozygous variant; The quality value, sequencing depth and number of supporting variant reads of the variant site are calculated, and the single nucleotide polymorphism sites with a quality value greater than thirty and a sequencing depth greater than ten times and the insertion and deletion sites with a quality value greater than fifty and a sequencing depth greater than twenty times are taken as high-confidence variant sites. The high-confidence variant sites are annotated to generate a variant annotation file including genomic location information, variant type information, allele frequency information and functional impact information.

4. The method according to claim 1, characterized in that Based on the standard variation detection report, according to the improved random field diffusion model in the pre-set self-activation deep analysis system, multi-scale feature sampling is performed in combination with Langevin dynamics and Hamiltonian Monte Carlo method to generate an initial feature distribution, and the initial feature distribution is optimized based on the adaptive adjustment mechanism and Wasserstein distance to generate a high-dimensional feature distribution and input it into the pre-set dual self-activation iterative network. The spatial features are extracted through the dynamic routing algorithm and attention guidance mechanism in the excitation pathway based on the capsule network, and the spatial features are added to the inhibition pathway and contrast learning and competitive collaboration are performed to obtain balanced features. The optimal features are selected in the metastable state according to the Boltzmann machine and annealing algorithm, and the optimal features are added to the dynamic system based on the Lie differential equation. The neural ordinary differential equation network is combined to model the time series evolution and perform feature integration to obtain the feature mapping matrix and feature evolution trajectory including: Based on the standard mutation detection report, multi-scale feature sampling is performed according to an improved random field diffusion model in a pre-set self-activated deep analysis system, wherein the improved random field diffusion model adopts the Langevin dynamics method to describe the motion trajectory of particles on the potential energy surface, and combines the Hamiltonian Monte Carlo method to accept sub-optimal sampling points with a preset probability, and samples the mutation features at the chromosome level, gene level and exon level to generate an initial feature distribution; The initial feature distribution is optimized based on an adaptive adjustment mechanism and a Wasserstein distance, wherein the adaptive adjustment mechanism dynamically adjusts a temperature parameter and a step size parameter according to a sampling acceptance rate, and the Wasserstein distance is used to measure the difference between a current distribution and an ideal distribution, and distribution optimization is performed by minimizing the Wasserstein distance to generate a high-dimensional feature distribution; The high-dimensional feature distribution is input into a pre-set dual self-activation iterative network, which includes an activation path and an inhibition path. In the activation path, a dynamic routing algorithm based on a capsule network is used to adaptively adjust the connection strength between capsules, and an attention mechanism is introduced to dynamically adjust the weight according to feature correlation to extract spatial features. The spatial features are input into the inhibition path for perturbation, and the original features and the perturbation features are distinguished by comparative learning, and balanced features are obtained through competition and cooperation of the dual paths; Constructing a Boltzmann machine, defining a weight matrix based on the feature correlation, defining an energy function based on the pathogenic risk, and using an annealing algorithm to iteratively sample the equilibrium feature in a metastable state to obtain an optimal feature with the lowest energy; The optimal features are added to a dynamic system based on the Lie differential equation, and the dynamic system is modeled in time series in combination with a neural ordinary differential equation network. Each network layer is regarded as a time step for solving the ordinary differential equation, and feature mapping matrices are sampled at different time steps to generate feature evolution trajectories.

5. The method according to claim 4, characterized in that The optimal features are added to the dynamic system based on Lie differential equations, and the dynamic system is modeled in time series by combining the neural ordinary differential equation network. Each network layer is regarded as a time step for solving the ordinary differential equation. The feature mapping matrix is ​​sampled at different time steps and the feature evolution trajectory is generated, including: Obtaining the values ​​of the optimal feature at different time steps and constructing a system state vector, establishing a Lie differential equation dynamic system based on the system state vector, calculating the interaction strength between state variables to generate a state transfer matrix, and calculating the influence of external inputs on state variables to generate a control matrix; Fusing the set corresponding to the optimal feature with the system state vector, calculating the interaction strength between the fused feature and the original feature, and updating the state transfer matrix and the control matrix according to the interaction strength; Construct a neural ordinary differential equation network including an input layer, multiple hidden layers and an output layer, multiply the hidden state of the previous layer by a weight matrix in each hidden layer and superimpose a bias vector, input the calculation result into an activation function to obtain the time change rate of the hidden state of the current layer, update the hidden state of the current layer based on the time change rate by a numerical integration method, map the state transfer matrix and the control matrix to the network parameter space to complete network initialization; A back-propagation algorithm is used to calculate parameter gradients and iteratively update the network parameters. A forward propagation process is performed at a preset number of time steps to obtain and save the hidden state corresponding to each time step. The hidden state is extracted at a preset sampling interval in a time series. The extracted hidden states are concatenated in chronological order to generate a feature mapping matrix. The feature mapping matrix is ​​smoothed to obtain a feature evolution trajectory.

6. The method according to claim 1, characterized in that The feature mapping matrix and the feature evolution trajectory are added to a pre-set integrated predictor, which includes a deep belief network and a hypergraph neural network. The parameters of the integrated predictor are optimized according to an alternating training strategy and an initial risk score is calculated. The feature evolution trajectory is analyzed according to a mixed Gaussian process regression and the risk change trend is predicted. The initial risk score and the risk change trend are added to an improved Bayesian inference network. The Markov blanket model is executed to update the risk probability and the probability estimation deviation is corrected through a particle filter algorithm to establish a risk association network between mutation points. A deep hierarchical decision tree is trained based on the risk association network, and node division is performed in combination with an adaptive split threshold algorithm to obtain a risk level and output a multi-dimensional risk assessment report corresponding to the current individual, including: Adding the feature mapping matrix and the feature evolution trajectory as input samples to a pre-set integrated predictor, wherein the integrated predictor includes a deep belief network and a hypergraph neural network, wherein the deep belief network is composed of a plurality of stacked restricted Boltzmann machines, and the hypergraph neural network maps features to a high-order hypergraph space and models the interaction between features through hyperedge connections; Adopting an alternating training strategy to optimize the parameters of the integrated predictor, fixing the parameters of the deep belief network and updating the parameters of the hypergraph neural network, then fixing the parameters of the hypergraph neural network and updating the parameters of the deep belief network, iterating the training until convergence, and predicting the input sample to obtain an initial risk score; Construct a mixed Gaussian process regression model composed of multiple sub-Gaussian processes, use the expectation-maximization algorithm to iteratively optimize the parameters of the mixed Gaussian process regression model, calculate the posterior probability that the input sample belongs to different sub-Gaussian processes and update the hyperparameters of each sub-Gaussian process, and predict the feature evolution trajectory based on the trained mixed Gaussian process regression model to obtain the risk change trend; Constructing an improved Bayesian inference network, wherein the improved Bayesian inference network includes multiple time slices, each time slice corresponds to an independent Bayesian network, adjacent time slices are connected by conditional probability distribution, the prior probability of the risk node in the initial time slice is set according to the initial risk score, the second posterior probability corresponding to each node is calculated using the Markov blanket model, the second posterior probability is corrected according to the particle filter algorithm, the posterior probability distribution is approximated by weighted sampling particles, the particles are updated according to the state transition model and the particle weights are adjusted according to the observation model, the risk correlation between the mutation points is calculated based on the corrected second posterior probability and a risk association network is established; A deep hierarchical decision tree is constructed, and an adaptive splitting threshold algorithm is used to optimize node splitting. The splitting threshold is dynamically adjusted according to node purity and sample distribution. The deep hierarchical decision tree is trained based on the risk association network. The risk level of the input sample is predicted and the key variation combination is analyzed through the deep hierarchical decision tree to obtain the risk level corresponding to the input sample. The risk level of the input sample and the key variation combination are combined to obtain a multi-dimensional risk assessment report.

7. The method according to claim 6, characterized in that The hyperparameters of each sub-Gaussian process are updated as shown in the following formula Among them, θ k represents the hyperparameter set of the kth sub-Gaussian process, represents the updated hyperparameter set of the k-th sub-Gaussian process, γ(z nk ) represents the posterior probability that the nth input sample belongs to the kth sub-Gaussian process, Represents sample y n The likelihood function under the kth sub-Gaussian process, y n represents the observed value of the nth input sample, δ represents the mean of the sub-Gaussian process distribution, which is 0. represents the covariance matrix of the k-th sub-Gaussian process at sample n, σ 2 represents the noise variance, I represents the identity matrix, v represents the degree of freedom parameter of the probability density function, represents the probability density function, and N represents the total number of input samples.

8. A genetic disease gene detection data analysis system based on big data, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to obtain the original gene sequencing data of the individual to be detected, perform multi-dimensional quality assessment on the original gene sequencing data to generate a quality assessment indicator set, perform data filtering on the quality assessment indicator set through a dual-threshold dynamic quality control algorithm to obtain high-quality sequence data and perform deduplication operation through an improved local sensitive hashing algorithm to obtain non-repetitive sequence data, build a sequence comparison system based on a distributed computing framework, generate a reference genome and compare the non-repetitive sequence data with the reference genome in combination with a Burrows-Wheeler transform algorithm and a fast comparison algorithm based on a suffix array, generate a sequence position mapping file and perform variation detection through a multi-level variation detection strategy to generate an initial variation detection result, perform a credibility score on the initial variation detection result through a Bayesian statistical model, and generate a standard variation detection report containing a credibility score; The second unit is used to perform multi-scale feature sampling based on the standard variation detection report and the improved random field diffusion model in the pre-set self-activation deep analysis system, combined with Langevin dynamics and Hamiltonian Monte Carlo method, to generate an initial feature distribution, optimize the initial feature distribution based on an adaptive adjustment mechanism and Wasserstein distance, generate a high-dimensional feature distribution and input it into a pre-set dual self-activation iterative network, extract spatial features through a dynamic routing algorithm and an attention guidance mechanism in an excitation pathway based on a capsule network, add the spatial features to an inhibitory pathway and perform contrastive learning and competitive collaboration to obtain a balanced feature, select the optimal feature in a metastable state according to a Boltzmann machine and an annealing algorithm, add the optimal feature to a dynamic system based on a Lie differential equation, model the time series evolution in combination with a neural ordinary differential equation network and perform feature integration, and obtain a feature mapping matrix and a feature evolution trajectory; The third unit is used to add the feature mapping matrix and the feature evolution trajectory to a pre-set integrated predictor, wherein the integrated predictor includes a deep belief network and a hypergraph neural network, optimize the parameters of the integrated predictor according to an alternating training strategy and calculate the initial risk score, analyze the feature evolution trajectory according to a mixed Gaussian process regression and predict the risk change trend, add the initial risk score and the risk change trend to an improved Bayesian inference network, execute the Markov blanket model to update the risk probability, correct the probability estimation bias through a particle filtering algorithm, and establish a risk association network between mutation points, train a deep hierarchical decision tree based on the risk association network, perform node division in combination with an adaptive split threshold algorithm, obtain the risk level, and output a multi-dimensional risk assessment report corresponding to the current individual.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • HIV molecular network construction method based on gene distance analysis and analysis system

    CN120690294A

  • Whole genome genetic diversity cloud analysis method based on workflow automation

    CN120849416A

  • A cloud-based approach to genome-wide genetic diversity analysis based on workflow automation

    CN120849416B

  • Abnormity identification traceability method based on coronavirus high-throughput detection

    CN121148469A

  • Network link reliability monitoring method and device, medium and program product

    CN121173711A