Big model technology-based biological information analysis system
Through a biological information analysis system based on large model technology, genome comparison and clinical phenotype data are dynamically calibrated, combined with protein network abnormal screening, the problem of data association failure in the existing technology is solved, and the unified benchmark and three-dimensional spatial distribution of multimodal data are realized, which improves the accuracy and efficiency of bioinformatic analysis.
Patent Information
- Application Number
- CN202510317994.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-18
AI Technical Summary
When processing biological information data, it is difficult for the existing technology to match the dynamic evolution characteristics of clinical phenotypes, cross-period data associations are invalid, cross-sample global association rules cannot be effectively mined, and heterogeneous characteristics of variant distribution are ignored, resulting in delays in identification of key biomarkers and waste of resources. The two-dimensional model cannot analyze protein cyberspace aggregation characteristics, resulting in bias in functional annotation and experimental verification.
The biological information analysis system based on large model technology is adopted, and the time stamp of genomic sequence alignment information and clinical phenotype data is obtained through the data analysis calibration module, the complementary pairing mode of cross-sample bases is extracted, and the discrete distribution of genotype-phenotype association intensity is calculated, and the standard value of multimodal data is generated. The problem disassembly module calls the semantic weight of the pathogenic gene annotation, adjusts the execution order of the analysis task, generates a three-dimensional spatial distribution, and the feedback iteration module adjusts the judgment threshold, and eliminates abnormal samples.
Dynamic calibration of genome comparison eliminates timing dislocation bias, enhances low-abundance synergistic capture, solves the loss of standardized data of multi-source heterogeneity, balances statistical thresholds and biological functions, integrates gene expression clustering characteristics in three-dimensional distribution, and reduces the false positive rate in population genetic heterogeneity scenarios.
Smart Images

Figure CN120260674A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics, and particularly to a bioinformatics analysis system based on large model technology. Background Art
[0002] The technical field of bioinformatics involves cross - applications of multiple disciplines such as computer science, statistics, and biology, and is mainly used for analyzing and interpreting biological data. The core content of this technical field includes genomics, proteomics, metabolomics, and biological system modeling. Biological data is obtained through high - throughput sequencing technology, and biological information is extracted from massive data by combining machine learning, mathematical modeling, and statistical analysis methods. The overall technical system of bioinformatics includes links such as data acquisition, data storage, data analysis, and data visualization, and is widely applied in fields such as disease research, drug development, personalized medicine, and agricultural breeding.
[0003] Among them, a bioinformatics analysis system based on large model technology refers to constructing a biological data analysis model using deep learning and natural language processing technologies to improve data parsing ability and automation level. This patent theme covers technical matters such as biological sequence alignment, gene function prediction, protein structure analysis, and biological network inference. By constructing a biological data parsing framework based on neural networks, optimizing data mining ability through large - scale parameter training methods, adopting self - supervised learning strategies to reduce the need for data annotation, and combining attention mechanisms to enhance feature extraction ability.
[0004] The prior art relies on static time windows to process time - series data, making it difficult to match the dynamic evolution characteristics of clinical phenotypes and prone to cross - period data association failure. Traditional variant detection adopts local feature extraction strategies, lacking cross - sample global association rule mining, resulting in the omission of synergistic mechanisms. The linear task execution mode ignores the heterogeneity characteristics of variant distribution, causing delays in identifying key biomarkers and wasting resources. Isolated parameter analysis cannot reflect the spatial coupling effect of gene expression and protein interaction in three - dimensional topological relationships, restricting the interpretation of molecular pathway mechanisms. Fixed thresholds and artificial rules are difficult to adapt to changes in sample scale and type, and error filtering or noise residue is likely to occur in population genetic diversity scenarios. Two - dimensional models cannot analyze the spatial aggregation characteristics of protein networks, leading to deviations in function annotation and experimental verification. Summary of the Invention
[0005] The purpose of the present invention is to solve the deficiencies existing in the prior art, and a bioinformatics analysis system based on large model technology is proposed.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A bioinformatics analysis system based on large model technology includes:
[0007] The data parsing and calibration module obtains the genomic sequence alignment information and the time stamps of clinical phenotype data, detects the overlapping range between the coordinates of gene mutation sites and the sequence coding regions, extracts the cross-sample base complementary pairing patterns, performs abnormal identification and elimination operations on the nodes of the protein interaction network, calculates the discrete distribution quantity of the genotype-phenotype association strength, and generates the multi-modal data standard values;
[0008] Based on the multi-modal data standard values, the problem decomposition module calls the semantic weights of pathogenic gene annotations and the logical dependency relationships of mutation function predictions, divides the numerical differences of the structural variation classification boundaries, and generates the analysis task priority sequence;
[0009] The analysis task scheduling module calls the variation classification boundaries in the analysis task priority sequence, screens the distribution ranges of promoter region mutations and exon splicing site mutations, adjusts the analysis execution order, and generates the task execution optimization coefficient;
[0010] Based on the task execution optimization coefficient, the result mapping module extracts the clustering center coordinates of the gene expression profile, maps the mutation site enrichment region and the topological parameters of the protein interaction network into a three-dimensional space distribution, and generates the association resolution set;
[0011] The feedback iteration module detects the matching degree between the coverage rate of the structural variation hotspots and the functional annotation association matrix in the association interpretation set, adjusts the determination threshold of the sequence coding region, updates the abnormal sample elimination rules for cross-modal data, and generates the process optimization parameter set.
[0012] As a further solution of the present invention, the multi-modal data standard values include the genomic sequence alignment information standard value, the clinical phenotype data standard value, the gene mutation site coordinate standard value, the sequence coding region overlapping range standard value, the cross-sample base complementary pairing pattern standard value, the protein interaction network abnormal identification standard value, and the genotype-phenotype association strength discrete distribution quantity standard value. The analysis task priority sequence includes the semantic weights of pathogenic gene annotations, the logical dependency relationships of mutation function predictions, and the numerical differences of the structural variation classification boundaries. The task execution optimization coefficient includes the variation classification boundary, the promoter region mutation distribution range, the exon splicing site mutation distribution range, and the analysis execution order adjustment parameter. The association resolution set includes the clustering center coordinates of the gene expression profile, the mutation site enrichment region, and the three-dimensional space distribution of the topological parameters of the protein interaction network.
[0013] As a further solution of the present invention, the data parsing and calibration module includes:
[0014] The genomic sequence alignment sub-module obtains genomic sequence alignment information and clinical phenotype data timestamps, analyzes the genomic sequence alignment data, extracts the alignment results between samples, calculates the alignment score, obtains the mutation sites in the genomic alignment sequence, screens the regions with high alignment coverage, and associates the genomic variation information of each sample according to the phenotype data timestamp to generate high-coverage gene sequence alignment data;
[0015] The gene variant functional region detection sub-module, based on the high-coverage gene sequence alignment data, detects the coordinates of gene variant sites, analyzes the overlap between the variant sites and the coding regions of the genomic sequence, determines whether the variant sites fall into functional gene regions, calculates the overlap ratio of the variant sites, screens the mutation sites with high overlap ratios, and generates a set of variant sites in functional regions;
[0016] The genotype-phenotype association calculation sub-module calls the set of variant sites in functional regions, extracts the cross-sample base complementary pairing patterns, analyzes the distribution of different genotype combinations in the phenotype data, calculates the discrete distribution quantity of the association strength between genotypes and phenotypes, and uses the formula:
[0017]
[0018] Performs operations to obtain the genotype-phenotype association strength, screens the association patterns that meet the distribution characteristics, standardizes the multi-modal data, and obtains the standard values of the multi-modal data;
[0019] where S represents the distribution value of the genotype-phenotype association strength, represents the value of the i g th genotype variable, represents the phenotypic variable value corresponding to V, represents the value of the j g th irrelevant variant variable, represents the phenotypic variable value corresponding to V jg n g represents the number of genotype variables, and m g represents the number of irrelevant variant variables.
[0020] As a further solution of the present invention, the problem decomposition module includes:
[0021] The standard value calculation sub-module, based on the standard values of the multi-modal data, analyzes the standard ranges of the differential gene feature parameters, calls the gene expression levels, mutation frequencies, and epigenetic regulators of the target individual, calculates the deviation of each feature parameter from the standard value, and analyzes the overall variation situation through cumulative deviation, using the formula:
[0022]
[0023] Obtain the standard value through calculation, calculate the deviation amount, and establish the standard value deviation result;
[0024] Among them, D represents the standard value calculation deviation amount, represents the value of the i s th gene feature parameter, represents the corresponding multi-modal data standard value, n s represents the number of gene feature parameters, represents the individual environmental factor value, represents the environmental factor reference value, m s represents the number of environmental factors;
[0025] The pathogenic gene annotation sub-module, based on the standard value deviation result, analyzes the functional characteristics of pathogenic genes, calls the known gene database to extract gene association characteristics, analyzes the mutation sites and influence ranges of gene sequences, calculates the semantic weight value of the mutation function prediction logic, obtains the gene semantic weight calculation value, and establishes the pathogenic gene function annotation result;
[0026] The gene structure variation classification sub-module calls the pathogenic gene function annotation result, screens the structural differences corresponding to gene variations, calculates the classification boundary value of multi-structural variations, and compares the standard value deviation result to generate the numerical difference of the structural variation classification boundary, and establishes the analysis task priority sequence.
[0027] As a further solution of the present invention, the analysis task scheduling module includes:
[0028] The variation classification boundary recognition sub-module, based on the analysis task priority sequence, extracts the variation category information corresponding to multiple analysis tasks, calculates the distribution interval of mutation types, and compares the reference value of the mutation influence degree to identify the classification boundary of mutations, calculates the deviation amount of multiple category boundaries, and generates the variation classification boundary value;
[0029] The mutation distribution screening sub-module calls the variation classification boundary value, screens the distribution ranges of promoter region mutations and exon splicing site mutations, calculates the distribution ratio of multiple mutation categories in the differential region, and classifies and sorts the distribution according to the mutation function influence coefficient to generate the mutation distribution classification interval;
[0030] The task execution order adjustment sub-module calls the mutation distribution classification interval, adjusts the analysis task execution order according to the task priority, and uses the formula:
[0031]
[0032] Calculate to obtain the task execution optimization coefficient;
[0033] Among them, T opt represents the task execution optimization coefficient, represents the adjusted task distribution ratio, represents the classification boundary value before adjustment, represents the priority weight of the task, m x represents the number of analysis tasks.
[0034] As a further solution of the present invention, the result mapping module includes:
[0035] The task optimization coefficient calculation sub-module extracts the mutation site distribution parameter, the protein interaction network connectivity parameter, and the gene expression level parameter based on the task execution optimization coefficient, and uses the formula:
[0036]
[0037] performs operations to obtain the optimization adjustment coefficient;
[0038] Among them, C opt represents the optimization adjustment coefficient, V p represents the distribution density weighting coefficient of the mutation site, represents the distribution density of the mutation site in the protein interaction network, α represents the adjustment coefficient, L d represents the deviation degree of the gene expression level, V n represents the network connectivity, S c represents the gene expression mean deviation adjustment coefficient;
[0039] The gene expression profile clustering sub-module extracts the gene expression profile data based on the optimization adjustment coefficient, calculates the gene expression mean, and identifies the deviation of the mutant gene in the expression profile, screens the genes whose expression deviation value exceeds the set threshold, calculates the expression profile clustering center coordinates of the mutant gene, and obtains the gene expression clustering center coordinates;
[0040] The gene network three-dimensional topology mapping sub-module calculates the topological parameters of the mutant gene enrichment region in the protein interaction network based on the gene expression clustering center coordinates, maps its three-dimensional spatial distribution, and generates an association resolution set according to the connection situation of multiple genes in the topological network.
[0041] As a further solution of the present invention, the system further includes:
[0042] The feedback iteration module detects the matching degree between the structure variation hot spot coverage rate and the functional annotation correlation matrix in the association interpretation degree set, adjusts the sequence coding region determination threshold, updates the cross-modal data abnormal sample elimination rule, and generates a process optimization parameter set;
[0043] The process optimization parameter set includes the structure variation hot spot coverage rate, the functional annotation correlation matrix matching degree, the sequence coding region determination threshold, and the cross-modal data abnormal sample elimination rule.
[0044] As a further solution of the present invention, the feedback iteration module includes:
[0045] Based on the coverage rate of structural variation hotspots in the correlation interpretation degree set, the structural variation hotspot coverage sub-module calculates the coverage rate of multi-hotspots in the differential sequence coding region, calls the determination threshold to screen the hotspot regions that meet the requirements, and uses the formula:
[0046]
[0047] Perform operations to obtain the hotspot coverage metric value;
[0048] where C hot represents the hotspot coverage metric value, represents the coverage probability of the i t th structural variation hotspot, represents the correlation interpretation degree score of the i t th structural variation hotspot, δ s represents the adjustment factor, V max represents the maximum correlation interpretation degree score, N s represents the number of hotspots;
[0049] The functional annotation association matrix matching sub-module calls the hotspot coverage metric value, calculates the matching degree of multi-coding regions based on the functional annotation association matrix, extracts the matching scores of multi-regions, screens the regions with matching scores higher than the set threshold, eliminates the coding regions with low matching degrees, and obtains the functional matching coefficient;
[0050] The cross-modal data anomaly elimination sub-module calls the functional matching coefficient, adjusts the anomaly elimination rules of cross-modal data, calculates the deviation coefficient of each anomaly sample, screens the anomaly samples according to the deviation coefficient threshold, eliminates the data with deviations exceeding the standard range, and obtains the process optimization parameter set.
[0051] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0052] In the present invention, by dynamically calibrating genomic sequence alignment information and clinical phenotype timestamp data, the association bias caused by time series misalignment is eliminated. The extraction of cross-sample base complementary pairing patterns combined with the screening of abnormal nodes in the protein network enhances the capture accuracy of low-abundance co-variations. The quantification of the discrete distribution of genotype-phenotype associations establishes a unified benchmark for multi-modal data, solving the problem of the lack of standardization of traditional multi-source heterogeneous data. The numerical differences in the classification boundaries of structural variations are used to integrate the semantic weights of pathogenic genes, balancing the statistical significance threshold and the biological function rationality. The analysis execution order is dynamically adjusted to optimize resource allocation, synchronously covering key mutation regions of promoters and exons, and improving the functional annotation efficiency of non-coding regions. The three-dimensional spatial distribution mapping integrates gene expression clustering features and protein network topology parameters, breaking through the spatial association expression limitations of two-dimensional models. The closed-loop feedback mechanism real-time corrects the decision threshold and iteratively eliminates rules, reducing the false positive rate in the scenario of population genetic heterogeneity. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is the system flow chart of the present invention;
[0054] Figure 2 is the flow chart of the data analysis and calibration module of the present invention;
[0055] Figure 3 is the flow chart of the problem decomposition module of the present invention;
[0056] Figure 4 is the flow chart of the analysis task scheduling module of the present invention;
[0057] Figure 5 is the flow chart of the result mapping module of the present invention;
[0058] Figure 6 is the flow chart of the feedback iteration module of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0059] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0060] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "a plurality of" is two or more, unless otherwise specifically defined.
[0061] Embodiment 1
[0062] Please refer to Figure 1 , the bioinformatics analysis system based on large model technology includes:
[0063] The data parsing and calibration module obtains the genomic sequence alignment information and the clinical phenotype data timestamp, detects the overlapping range between the gene mutation site coordinates and the sequence coding region, extracts the cross-sample base complementary pairing pattern, performs abnormal identification and elimination operations on the protein interaction network nodes, calculates the discrete distribution quantity of the genotype-phenotype association strength, and generates the multi-modal data standard value;
[0064] The problem decomposition module, based on the multi-modal data standard value, calls the semantic weights of pathogenic gene annotations and the logical dependency relationship of mutation function prediction, divides the numerical differences of the structural variation classification boundaries, and generates the analysis task priority sequence;
[0065] The analysis task scheduling module calls the mutation classification boundaries in the analysis task priority sequence, screens the distribution ranges of promoter region mutations and exon splicing site mutations, adjusts the analysis execution order, and generates the task execution optimization coefficient;
[0066] The result mapping module, according to the task execution optimization coefficient, extracts the gene expression profile clustering center coordinates, maps the mutation site enrichment region and the protein interaction network topological parameters into a three-dimensional space distribution, and generates the association resolution set;
[0067] The feedback iteration module detects the matching degree between the structural variation hot spot coverage rate and the functional annotation association matrix in the association interpretation set, adjusts the determination threshold of the sequence coding region, updates the cross-modal data abnormal sample elimination rule, and generates the process optimization parameter set.
[0068] The standard values of multi-modal data include the standard values of genomic sequence alignment information, clinical phenotype data, gene variant site coordinate, sequence coding region overlap range, cross-sample base complementary pairing pattern, protein interaction network anomaly identification, genotype-phenotype association strength discrete distribution quantity. The analysis task priority sequence includes the semantic weight of pathogenic gene annotation, the logical dependence relationship of mutation function prediction, and the numerical difference of structural variant classification boundary. The task execution optimization coefficient includes the variant classification boundary, the mutation distribution range in the promoter region, the mutation distribution range in the exon splicing site, and the analysis execution order adjustment parameter. The association resolution set includes the gene expression profile clustering center coordinate, the mutation site enrichment region, and the three-dimensional spatial distribution of protein interaction network topology parameters. The process optimization parameter set includes the structural variant hot spot coverage rate, the functional annotation association matrix matching degree, the sequence coding region determination threshold, and the cross-modal data abnormal sample elimination rule.
[0069] Please refer to Figure 2 , and the data parsing and calibration module includes:
[0070] The genomic sequence alignment sub-module obtains the time stamps of genomic sequence alignment information and clinical phenotype data, parses the genomic sequence alignment data, extracts the alignment results between samples, calculates the alignment score, obtains the mutation sites in the genomic alignment sequence, screens the regions with high alignment coverage, and associates the genomic variation information of each sample according to the phenotype data time stamp to generate high-coverage gene sequence alignment data;
[0071] First, obtain the genomic sequencing data of multiple samples and perform standardized preprocessing on it. The sequence data of each sample undergoes quality filtering to remove low-quality base regions to ensure data reliability. During the sample alignment process, the input genomic sequence is segmented using the sliding window method, with each window size set to 100bp. Each segment is aligned with the reference genome in turn, and the alignment score of each segment is recorded. The calculation of the alignment score is based on the weighted processing of the number of matching bases and the mismatch situation. The matching base score is set to +2, the mismatch score is set to -3, and the insertion / deletion (indel) penalty is set to -5. The calculation example is as follows: If a 100bp segment contains 80 matching bases, 15 mismatch bases, and 5 deletions, the alignment score is calculated as follows:
[0072] S = (80×2)+(15×(-3))+(5×(-5)) = 160 - 45 - 25 = 90;
[0073] After summarizing all alignment results, the alignment consistency score is calculated based on the difference in alignment scores between different samples. The overall alignment level between samples is obtained by calculating the mean value. For the genomic alignment sequences of each sample, variant sites are extracted, including single nucleotide variants (SNPs) and insertions and deletions (Indels). For each variant site, its base substitution information (e.g., A→T) is recorded, and the variant frequency of this site in all samples is statistically analyzed. The screening of variant sites is based on the calculation of coverage, where coverage is defined as:
[0074]
[0075] When the coverage is higher than 0.8, this site is regarded as a high-confidence site. For example, if the sequencing depth of a site is 200X and 160 sequences contain this mutation, the coverage is calculated as follows:
[0076]
[0077] This result indicates that this mutant site has a high alignment coverage in the sequencing samples, so it has a high credibility and can be used for subsequent detection of gene variant functional regions and genotype-phenotype association analysis.
[0078] The gene variant functional region detection sub-module is based on the high-coverage gene sequence alignment data to detect the coordinates of gene variant sites, analyze the overlap between the variant sites and the coding regions of the genomic sequences, determine whether the variant sites fall into the functional gene regions, calculate the overlap ratio of the variant sites, screen the mutant sites with high overlap ratios, and generate a set of functional region variant sites;
[0079] Query one by one whether this coordinate is located in the known gene functional region. For the coordinate position of each mutant site, it is compared with the genomic annotation database to determine whether this variant falls into the coding region, promoter region or regulatory region. For the mutations falling into the coding region, further calculate the amino acid substitution caused by the variant. For example, if the codon GAG corresponding to the variant site chr1:12345 (G→A) encodes glutamic acid (E) and becomes AAG encoding lysine (K) after the mutation, then this mutation causes an amino acid substitution E→K. Record this change and statistically analyze the proportion of variant sites affecting protein function. The functional impact calculation of variant sites is based on the overlap ratio:
[0080]
[0081] If among 100 variant sites, 80 are located in the functional region, the calculation is as follows:
[0082]
[0083] The results indicate that in this gene alignment data, the vast majority of variant sites have significant overlap with known gene functional regions, suggesting that these variants may directly affect gene function and can thus be further used for genotype-phenotype association calculations to determine their correlation with clinical phenotypes.
[0084] The genotype-phenotype association calculation sub-module calls the set of variant sites in the functional region, extracts the cross-sample base complementary pairing patterns, analyzes the distribution of differential genotype combinations in the phenotype data, and calculates the discrete distribution quantity of the genotype-phenotype association strength using the formula:
[0085]
[0086] Perform operations to obtain the genotype-phenotype association strength, screen for association patterns that meet the distribution characteristics, standardize the multi-modal data, and obtain the standard values of the multi-modal data;
[0087] Among them, S represents the distribution value of the genotype-phenotype association strength, represents the value of the i g th genotype variable, represents the corresponding phenotype variable value, represents the value of the j g th irrelevant variant variable, represents the corresponding phenotype variable value, n g represents the number of genotype variables, m g represents the number of irrelevant variant variables.
[0088] First, for all samples, extract the allelic genotypes of each locus and form a genotype matrix. Each genotype variable value is calculated from the allelic frequencies. For example, if in 100 samples, the allelic frequency of a certain mutation locus is 0.4, then The phenotype variable value is obtained after normalizing the clinical phenotype data, using min-max normalization:
[0089]
[0090] Assume that the minimum value of the phenotype variable P in the dataset is 50 and the maximum value is 200. If the phenotype value of a certain sample is 100, then the normalization calculation is as follows:
[0091]
[0092] Subsequently, calculate the genotype-phenotype association strength S. For all genotype variables and phenotype variables calculate according to the following formula:
[0093]
[0094] Let n g = 3 genotype variables, m g = 2 unrelated variant variables, and the data are as follows:
[0095] G1 = 0.4, P1 = 0.33, G2 = 0.5, P2 = 0.50, G3 = 0.3, P3 = 0.20, V1 = 0.2, Q1 = 0.15, V2 = 0.1, Q2 = 0.10;
[0096] Calculate the numerator part:
[0097] (0.4 × 0.33) + (0.5 × 0.50) + (0.3 × 0.20) - [(0.2 × 0.15) + (0.1 × 0.10)] = 0.132 + 0.25 + 0.06 - (0.03 + 0.01) = 0.402;
[0098] Calculate the denominator part:
[0099]
[0100] Final calculation:
[0101]
[0102] The result shows that the genotype-phenotype association strength distribution value is 0.54. This value indicates that there is a certain association between genotype variables and phenotype variables, but the association strength is at a medium level. In subsequent analyses, genotype classification can be further optimized based on the distribution of different phenotype data to improve the association between genotype and phenotype, thereby enhancing the accuracy of disease prediction or functional research.
[0103] Please refer to Figure 3 , the problem decomposition module includes:
[0104] The standard value calculation sub-module, based on the multi-modal data standard value, analyzes the standard range of differential gene feature parameters, invokes the gene expression level, mutation frequency, and epigenetic regulatory factors of the target individual, calculates the deviation of each feature parameter from the standard value, and analyzes the overall mutation situation through cumulative deviation, using the formula:
[0105]
[0106] Operate to obtain the standard value calculation deviation amount and establish the standard value deviation result;
[0107] where D represents the standard value calculation deviation amount, represents the value of the i s th gene feature parameter, Represents the corresponding multi-modal data standard value, n s Represents the number of gene characteristic parameters Represents the individual environmental factor value Represents the environmental factor reference value, m s Represents the number of environmental factors;
[0108] The main calculation steps are as follows:
[0109] 1. Obtain the standard value and the target individual data
[0110] (1) Extract the standard value from the database This standard value usually comes from gene expression datasets of healthy populations, such as TCGA, GTEx, or 1,000Genomes.
[0111] (2) Obtain the gene characteristic parameter values and environmental factor values These data can be obtained by high-throughput sequencing (RNA-seq), whole-genome sequencing (WGS), or epigenetic detection (such as ATAC-seq).
[0112] For example, set the partial gene data of an individual as follows:
[0113] Table 1 Target individual gene expression values and standard values
[0114] Gene Target individual expression level (TPM) Standard value (TPM) A 30 50 B 80 90 C 25 20
[0115] Table 2 Target individual environmental factor deviation:
[0116] Environmental factor Target individual value Standard value Air pollution index (μg / m3) 80 50
[0117] 2. Calculate the gene expression deviation
[0118] The expression level deviation of each gene is calculated as follows:
[0119]
[0120] That is:
[0121] Gene A:
[0122] Gene B:
[0123] Gene C:
[0124] This calculation indicates that the expression level of Gene A is 40% lower than the standard value, the expression level of Gene B is 11.1% lower, and the expression level of Gene C is 25% higher.
[0125] 3. Calculate the environmental factor deviation
[0126] The environmental factor is calculated as follows:
[0127]
[0128] That is:
[0129] Air pollution index: |80 - 50| = 30
[0130] 4. Cumulative deviation calculation
[0131] Sum the squares of all gene deviations and add the deviation of the environmental factor:
[0132]
[0133] Substitute the data:
[0134] D = (-0.4) 2 + (-0.111) 2 +(0.25) 2 + 30;
[0135] = 0.16 + 0.0123 + 0.0625 + 30;
[0136] = 30.2348;
[0137] 5. Result analysis
[0138] Finally, the calculated D = 30.2348, which represents the overall variation of the target individual. To determine whether this value is abnormal, it can be compared with the data of healthy individuals. Assume the mean D of healthy individuals healthy is 3.2 and the standard deviation σ = 1.5, then:
[0139] If D > D healthy + 2σ = 3.2 + 2(1.5) = 6.2, it is considered that the individual has significant genomic abnormalities.
[0140] The calculated value D = 30.2348 far exceeds 6.2, indicating that the gene expression and environmental factors of the target individual deviate significantly from the normal range.
[0141] Finally, both the gene expression level and environmental factors of this individual show significant deviations. Subsequently, it is necessary to further analyze whether there are specific pathogenic gene mutations or epigenetic regulation abnormalities and enter the pathogenic gene annotation sub-module.
[0142] Based on the results of the standard value deviation, the pathogenic gene annotation sub-module analyzes the functional characteristics of pathogenic genes, calls the known gene database to extract gene-associated features, analyzes the variant sites and influence ranges of gene sequences, calculates the semantic weight value of the mutation function prediction logic, obtains the gene semantic weight calculation value, and establishes the pathogenic gene function annotation results.
[0143] For example, if the mRNA expression level of a certain gene B is abnormal and its variant site rsID234567 is annotated as "Pathogenic" in the ClinVar database, it is necessary to further analyze the influence range of the gene variant. When analyzing the gene sequence, first locate the exon or intron region where the gene variant is located. For example, if a certain mutation occurs in exon 3 of gene B and this mutation results in a non-synonymous substitution in the protein amino acid sequence (such as alanine A being replaced by valine V), it is necessary to calculate its potential functional impact. When calculating the mutation function prediction, multiple bioinformatics scoring metrics are used, such as PolyPhen-2 and SIFT scores. These scores predict the impact of the mutation on protein structure and function through machine learning models. For example, if the PolyPhen-2 score is 0.95 (the value range is 0 - 1, and the closer the value is to 1, the more harmful the mutation), then this mutation may have a greater functional impact. In addition, it is also necessary to calculate the semantic weight value of the gene. The semantic weight value can be obtained by the database statistics of the functional correlation of genes. For example, in the GeneOntology (GO) database, if a certain gene is involved in multiple key biological processes, its semantic weight can be obtained by calculating its frequency in different GO classifications. For example, the semantic weight value of gene B is calculated to be 2.3. Finally, integrate the calculated semantic weight value and the pathogenic mutation information to establish the pathogenic gene function annotation results.
[0144] The gene structure variant classification sub-module calls the pathogenic gene function annotation results, screens the structural differences corresponding to gene variants, calculates the classification boundary values of multiple structural variants, and compares the results of the standard value deviation to generate the numerical differences of the structural variant classification boundaries, and establishes the analysis task priority sequence.
[0145] First, extract the variant data of the target individual, such as copy number variation (CNV), insertion-deletion variation (Indel), chromosomal translocation, etc., and call the reference genome for alignment. For example, if there is a deletion in the chromosomal position chr12:345678 - 345890 of gene C in an individual, after comparison with the genomic reference sequence, it is confirmed that it belongs to the CNV type variation. Then, calculate the impact of this variation on gene function and classify it using the variant frequency data.
[0146] For example, in the 1000 Genomes database, the population frequency of this CNV is 0.001, indicating that it belongs to rare variants. Then, the classification boundary values of multi-structural variants are calculated. The calculation of classification boundary values depends on the statistical distribution in genomic data. For example, in the TCGA database, if the quantile boundaries of the structural variants of a certain gene affecting function are 0.05, 0.5, and 0.95, then the variant samples can be divided into low impact (<0.05), medium impact (0.05 - 0.95), and high impact (≥0.95). For example, the impact score calculation of the variant of gene C in the target individual is 0.98, exceeding the 95% quantile, and it is determined as a high-impact variant. Finally, by comparing the deviation results with the standard values, the numerical differences in the structural variant classification boundaries are generated, and an analysis task priority sequence is established. For example, if the target individual carries 3 high-impact variants, 2 medium-impact variants, and 5 low-impact variants, then the functional impacts of the high-impact variants are analyzed first, and the mutation sites with the highest priority are sorted out.
[0147] Table 1: Example Data
[0148] Gene Target individual expression level (TPM) Standard value (TPM) Deviation calculation A 30 50 -0.4 B 80 90 -0.111 C 25 20 0.25
[0149] As shown in Table 1, the expression level of gene A is on the low side, gene B is close to the standard value, and the expression level of gene C is slightly higher. Combining with the deviation calculation values, the deviation degree of the expression patterns of these genes in the target individual can be further analyzed.
[0150] Suppose the deviations of genes A, B, and C in the target individual are -0.4, -0.111, and 0.25 respectively, the air pollution exposure index of the environmental factor is 80, and the standard value is 50. Then:
[0151] D = (-0.4) 2 + (-0.111) 2 + (0.25) 2 + |80 - 50|;
[0152] = 0.16 + 0.0123 + 0.0625 + 30;
[0153] = 30.2348;
[0154] This result indicates that the comprehensive variation degree of this individual is relatively large, and there may be significant genomic abnormalities, and the impacts of its specific pathogenic genes need to be further analyzed.
[0155] Please refer to Figure 4 , the analysis task scheduling module includes:
[0156] The mutation classification boundary recognition sub-module extracts the mutation category information corresponding to multiple analysis tasks based on the analysis task priority sequence, calculates the distribution interval of the mutation types, compares with the benchmark value of the mutation impact degree, identifies the classification boundary of the mutation, calculates the deviation amount of multiple category boundaries, and generates the mutation classification boundary value;
[0157] First, based on the analysis task priority sequence, extract the mutation category information involved in the current task respectively. This information includes the specific type of mutation (such as point mutation, insertion / deletion mutation, copy number variation, etc.), as well as the corresponding chromosome region number and the gene sequence segment affected. And perform data cleaning according to the known mutation distribution characteristics recorded in the database, only retaining the mutation categories that meet the scope of the current analysis task, and excluding the mutation categories that are not within the task scope. Secondly, calculate the distribution interval of the mutation types, that is, count the occurrence frequency of the same mutation category in the sample and calculate its relative distribution ratio in the whole genome range. For example, if the number of occurrences of a certain mutation category detected in a certain gene region is 1200 times, and the total number of mutations of this category in the whole genome is 18000 times, then the relative distribution ratio in this region can be calculated as:
[0158]
[0159] For each mutation type, calculate its impact degree, using the impact grading benchmark value in the reference database or experimental data, such as the mutation impact grading criteria provided by scoring methods such as SIFT and PolyPhen. For example, an SIFT value lower than 0.05 is considered to have a greater impact, and a PolyPhen score greater than 0.85 indicates a greater mutation impact. By comparing with the benchmark value of the mutation impact degree, judge whether the current mutation belongs to the high-impact category, and set the classification boundary accordingly. For example, set the mutation category with an SIFT value in [0, 0.05] as high impact, and the category with a PolyPhen score in [0.85, 1.0] as high-risk mutation. In addition, to quantify the deviation amount of multiple category boundaries, it is necessary to calculate the deviation degree between the mutation distribution of the current task and the benchmark value. The calculation formula is as follows:
[0160]
[0161] Among them, P baseline is the average genome-wide distribution value of the corresponding mutation category in the reference dataset. For example, if the distribution ratio of this category in the benchmark database is 0.08, then:
[0162]
[0163] The result shows that the mutation category distribution in the current sample has a 16.6% deviation compared to the baseline data, indicating that the mutation characteristics of the current sample have certain individual differences. If this deviation value exceeds the set threshold (such as 0.15), it is necessary to adjust the mutation classification boundary to ensure that the classification criteria adopted in subsequent analysis tasks can accurately reflect the mutation characteristics of the current sample, thereby reducing classification errors caused by baseline deviations. Finally, the calculated classification boundary value is stored in the task scheduling module for subsequent task calls.
[0164] The mutation distribution screening sub-module calls the mutation classification boundary value to screen the distribution ranges of promoter region mutations and exon splicing site mutations, calculates the distribution ratio of multiple mutation categories in the differential region, and classifies and sorts the distribution according to the mutation function influence coefficient to generate a mutation distribution classification interval.
[0165] First, obtain the classified mutation categories from the mutation classification boundary value generated by the previous module, and screen out the mutations located in the promoter region and exon splicing sites. The screening criteria are based on the genomic coordinate database. For example, the promoter region is usually defined as the range 2000bp upstream of the transcription start site, and exon splicing site mutations mainly focus on splicing donor and acceptor mutations within the range of ±2bp. Therefore, extract the mutations that meet this range and count their distribution ratios in different mutation categories. For example, in the promoter region of a certain gene, mutation category A appears 300 times, while mutation category B appears 500 times. At the exon splicing site, mutation category A appears 100 times and category B appears 400 times. Then calculate their distribution ratios in the differential region as follows:
[0166]
[0167] The result shows that the mutation ratio of category B is higher in both the promoter region and the splicing site, and it may have a more significant regulatory effect. Therefore, in subsequent task scheduling, the mutation of category B should be given priority for detailed functional analysis.
[0168] The task execution order adjustment sub-module calls the mutation distribution classification interval and adjusts the execution order of analysis tasks according to the task priority, using the formula:
[0169]
[0170] Calculate the task execution optimization coefficient;
[0171] where, T opt represents the task execution optimization coefficient, represents the adjusted task distribution ratio, represents the classification boundary value before adjustment, represents the priority weight of the task, mx Represents the number of analysis tasks.
[0172] First, call the mutation distribution classification interval to obtain the distribution ranking of different mutation categories in the target area, and adjust the tasks in combination with the analysis task priority. The priority weight is assigned according to the importance of the mutation category. For example, for high-impact mutation categories, a higher priority is set, such as U = 5, while for medium-impact categories, U = 3 is set. Calculate the optimization coefficient T opt Use the following formula:
[0173]
[0174] Assume the classification boundary value before adjustment The adjusted task distribution ratio Then calculate the deviation:
[0175]
[0176] Assume the current number of analysis tasks m x = 3, and the sum of its priority weights is:
[0177]
[0178] Substitute into the calculation:
[0179] T opt = 0.166 × log(12) = 0.166 × 1.079 = 0.179;
[0180] This result indicates that the adjustment degree of the task execution order calculated based on the optimization coefficient is small, indicating that the current task priority ranking is relatively reasonable. If the optimization coefficient exceeds a certain set threshold (such as 0.3), it indicates that the task execution order needs to be further adjusted to improve the overall execution efficiency of the mutation screening and analysis tasks.
[0181] Please refer to Figure 5 , and the result mapping module includes:
[0182] The task optimization coefficient calculation sub-module extracts the mutation site distribution parameter, protein interaction network connectivity parameter, and gene expression level parameter based on the task execution optimization coefficient, and uses the formula:
[0183]
[0184] Perform operations to obtain the optimization adjustment coefficient;
[0185] Among them, C opt represents the optimization adjustment coefficient, V p represents the distribution density weighting coefficient of the mutation site, represents the distribution density of mutation sites in the protein - interaction network, α represents the adjustment coefficient, L d represents the deviation degree of gene expression level, V n represents the network connectivity, S c represents the adjustment coefficient for the deviation of the mean gene expression;
[0186] First, it is necessary to statistically analyze the distribution of mutation sites. By parsing the gene - sequencing results of sample data, calculate the distribution frequency of each mutation site within the entire genome range, and assign the distribution - density weight coefficient V p , for example, in 1000 gene samples, if a certain gene mutation occurs 150 times, then its mutation frequency can be expressed as 150 / 1000 = 0.15. Then, set the distribution - density weight V p = 0.15×10 (set 10 as the normalization factor), that is, V p = 1.5. Subsequently, calculate the distribution density of mutation sites in the protein - interaction network , first construct the protein - interaction network, count the total number N of protein nodes connected to this gene p , and calculate its mutation frequency in this network. Suppose there are 50 protein nodes involved in this gene and the mutation occurs 10 times, then Then calculate the deviation degree L d of gene expression level. Based on the gene - expression - profile data, count the difference between the expression level of the target gene and the average expression level. Suppose the normal - expression value mean of a certain gene is 8.0, and the expression value of the mutant sample is 12.0, then L d = |12.0 - 8.0| = 4.0. After determining the adjustment coefficient α (for example, set α = 0.5), calculate α·L d = 0.5×4.0 = 2.0. Next, calculate the network connectivity V n , this value can be obtained by counting the number of adjacent nodes of the target gene in the protein - interaction network. For example, if the target gene has 20 connection nodes in the network, then V n = 20. Finally, calculate the adjustment coefficient S c for the deviation of the mean gene expression, and its calculation method is the mean of the gene - expression deviation values. For example, in 5 different samples, the gene - expression deviation values of the target gene are 3.0, 4.0, 2.5, 3.5, 4.2 respectively, then S c = (3.0 + 4.0 + 2.5 + 3.5 + 4.2) / 5 = 3.44. Substitute the values calculated above into the optimized - adjustment - coefficient calculation formula:
[0187]
[0188] This result shows that the optimized adjustment coefficient C optA negative value indicates a lower distribution density of the gene mutation, a higher connectivity in the protein interaction network, and a larger deviation in gene expression level. The calculation result of this optimization adjustment coefficient will be used to further determine whether the gene belongs to the group of genes with greater impact mutations, and whether higher weights or screening processes need to be assigned in subsequent calculations. This value can be used to determine gene screening criteria or construct a risk assessment model.
[0189] The gene expression profile clustering sub-module extracts gene expression profile data based on the optimization adjustment coefficient, calculates the gene expression mean, and identifies the deviation of mutant genes in the expression profile, screens genes with expression deviation values exceeding the set threshold, calculates the expression profile clustering center coordinates of mutant genes, and obtains the gene expression clustering center coordinates.
[0190] First, obtain the gene expression profile data matrix, record the expression values of different genes under different experimental conditions. For example, extract the expression values of the target gene in 100 samples to form a 100×1 expression matrix, and calculate the gene expression mean. Suppose the expression mean of a certain gene in 10 normal samples is 8.5, while in 10 mutant samples it is 12.1. Then the degree of deviation in expression level is calculated as follows:
[0191] ΔE = 12.1 - 8.5 = 3.6;
[0192] Then, perform gene screening. Set the expression deviation threshold T to screen genes with expression deviation values exceeding the threshold. For example, set T = 2.0. Then the expression deviation of this gene, 3.6 > 2.0, meets the screening conditions. Subsequently, calculate the expression profile clustering center coordinates of the mutant gene. By the K-means clustering method, divide the expression data into multiple clustering centers. Suppose the center coordinates of the clustering of the expression data of this gene are (8.4, 12.2). Then the expression profile clustering center coordinates of this gene are (8.4, 12.2). This result indicates that the expression data of this gene shows an obvious separation trend in the mutant and non-mutant states, indicating that it may play a key role in the regulatory network and may be used as a candidate biomarker for further research or disease prediction. This value can be used for the topological mapping analysis of the subsequent gene network.
[0193] The three-dimensional topological mapping sub-module of the gene network calculates the topological parameters of the mutant gene enrichment region in the protein interaction network based on the gene expression clustering center coordinates, maps its three-dimensional spatial distribution, and generates an association resolution set according to the connection situation of multiple genes in the topological network.
[0194] First, obtain the protein interaction network data, count the number of connections between the target gene and other genes. Suppose the number of connections of the target gene is 15. Its local clustering coefficient is calculated as follows:
[0195]
[0196] Assume that there are 10 adjacent genes of the target gene, and the number of edges formed by their connection is 25, then
[0197]
[0198] Then calculate the betweenness centrality B of the gene in the network by counting the frequency of it as an intermediate node in the shortest path. Assume that the betweenness centrality of the target gene in the network shortest path is 0.35, then its three-dimensional topological parameters can be expressed as (15, 0.56, 0.35). Subsequently, analyze the connection situation of multiple genes to generate an association resolution set, such as counting the connection characteristics of multiple mutant genes to form a data set
[0199] {(15, 0.56, 0.35), (20, 0.63, 0.42), (18, 0.59, 0.39)}. Finally, obtain the gene association analysis data of the topological network. The result shows that the topological structure of the target gene in the protein interaction network is relatively tight, with a low betweenness centrality but a high local clustering coefficient, indicating that the role of this gene in the network is more likely to be related to local regulation rather than global information transmission. Further analysis can combine the three-dimensional topological parameters of other genes to jointly construct a gene regulatory network to explore the role it plays in specific biological processes. This data can be used for subsequent bioinformatics analysis to optimize the screening strategy of mutant genes.
[0200] Please refer to Figure 6 , the feedback iteration module includes:
[0201] The structural variation hot spot coverage sub-module calculates the coverage rate of multiple hot spots in the differential sequence coding region based on the structural variation hot spot coverage rate in the association interpretation degree set, calls the determination threshold to screen the hot spot regions that meet the requirements, and uses the formula:
[0202]
[0203] Perform operations to obtain the hot spot coverage metric value;
[0204] Among them, C hot represents the hot spot coverage metric value, represents the coverage probability of the i t th structural variation hot spot, represents the association interpretation degree score of the i t th structural variation hot spot, δ s represents the adjustment factor, V max represents the maximum association interpretation degree score, N s represents the number of hot spots;
[0205] First, extract its corresponding coverage probability value This value represents the distribution probability of the hotspot in the current sequence coding region, and at the same time extracts the associated interpretability score. This score is given based on historical data or experimental results and characterizes the significance of the hotspot in terms of gene mutation impact or functional annotation. To ensure that the calculation of coverage can effectively measure the importance of the hotspot region, it is necessary to call the coverage probabilities and interpretability scores of all N s hotspots and perform the calculation. To obtain the cumulative weighted coverage value of the hotspots. After completing the calculation of the cumulative value, divide it by the total number of hotspots N s , to obtain the standardized hotspot coverage benchmark value. This benchmark value is used to adjust the coverage measurement range in subsequent calculations. Further call the adjustment factor δ s for deviation correction. This adjustment factor can be determined from historical data. For example, by analyzing the hotspot distribution of multiple groups of sequences, taking the standard deviation of their coverage rates as the value of δ s . Multiply the corrected result by the maximum interpretability score V max , which is determined by the maximum value among all hotspots , and finally divide by the highest coverage probability value among all hotspots This value is used to normalize the coverage measurement value C hot , to ensure that the measurement results are applicable to hotspot regions of different scales. Finally, obtain the coverage measurement value of the structural variant hotspots, and based on this, screen the hotspot regions with a coverage rate higher than the determination threshold.
[0206] Formula parameter assignment and calculation:
[0207] Suppose there are currently 4 structural variant hotspots, and their coverage probabilities and interpretability scores are as follows:
[0208]
[0209] Calculate
[0210] Let δ s = 10 (average adjustment factor calculated based on multiple datasets), then:
[0211]
[0212] Let V max = 75, Final calculation:
[0213]
[0214] The result shows that the overall metric value covered by the current structural variation hotspots is 210.94, which can be used to screen hotspots with higher coverage rates.
[0215] The function annotation correlation matrix matching sub-module calls the hotspot coverage metric value, calculates the matching degree of multiple coding regions based on the function annotation correlation matrix, extracts the matching scores of multiple regions, screens the regions with matching scores higher than the set threshold, eliminates the coding regions with low matching degrees, and obtains the function matching coefficient;
[0216] First, obtain the corresponding function annotation score from the function annotation matrix. The function annotation score is derived from database annotation information or experimental measurement data. The function score of each coding region represents its importance in structural variation impact or gene expression. Subsequently, calculate the matching score between this coding region and all structural variation hotspots. The calculation method of the matching score is to traverse all hotspot regions, extract their relevant scores in the function annotation matrix, and calculate the cosine similarity or correlation score between the coding region and the hotspot region. For example, assume that a coding region has 3 hotspot associations, and its matching score is calculated as follows:
[0217] Assume that the function scores of the coding region are 60, 75, 80, and the corresponding hotspot coverage scores are 130, 150, 180. The matching score calculation:
[0218]
[0219] Calculate the numerator:
[0220] (60 × 130) + (75 × 150) + (80 × 180) = 7800 + 11250 + 14400 = 33450;
[0221] Calculate the denominator:
[0222]
[0223] The final matching score:
[0224]
[0225] This result indicates a relatively high matching degree between the coding region and the hotspot region. If the matching score exceeds the preset threshold (such as 0.95), then retain this coding region; otherwise, eliminate it.
[0226] The cross-modal data anomaly elimination sub-module calls the function matching coefficient, adjusts the anomaly elimination rules for cross-modal data, calculates the deviation coefficient of each abnormal sample, screens abnormal samples according to the deviation coefficient threshold, eliminates the data with deviations beyond the standard range, and obtains the set of process optimization parameters.
[0227] First, calculate the deviation coefficient for each abnormal sample. The deviation coefficient is calculated using the method of normalizing the deviation of the sample from the matching mean. For example, for a sample with matching scores of 0.85, 0.88, and 0.92, calculate the matching mean:
[0228]
[0229] Calculate the deviation:
[0230]
[0231] Assume that the set deviation threshold is 0.05. Since 0.0287 is less than the threshold, this sample is retained; otherwise, it is excluded.
[0232] Finally, exclude the data with deviations outside the standard range to obtain the set of process optimization parameters to optimize the final result.
[0233] The above are only the preferred embodiments of the present invention and do not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A bioinformatics analysis system based on large model technology, characterized in that, The system includes: The data parsing and calibration module obtains genomic sequence alignment information and clinical phenotype data timestamps, detects the overlapping range between gene variant site coordinates and sequence coding regions, extracts cross-sample base complementary pairing patterns, performs anomaly identification and elimination operations on protein interaction network nodes, calculates the discrete distribution quantity of genotype-phenotype association strength, and generates multi-modal data standard values; The problem decomposition module, based on the multi-modal data standard values, calls the semantic weights of pathogenic gene annotations and the logical dependency relationships of mutation function predictions, divides the numerical differences in the classification boundaries of structural variations, and generates an analysis task priority sequence; The analysis task scheduling module calls the variant classification boundaries in the analysis task priority sequence, screens the distribution ranges of promoter region mutations and exon splicing site mutations, adjusts the analysis execution order, and generates a task execution optimization coefficient; The result mapping module, according to the task execution optimization coefficient, extracts the gene expression profile clustering center coordinates, maps the mutation site enrichment region and the topological parameters of the protein interaction network into a three-dimensional spatial distribution, and generates an association resolution set.
2. The bioinformatics analysis system based on large model technology according to claim 1, wherein The multi-modal data standard values include genomic sequence alignment information standard values, clinical phenotype data standard values, gene variant site coordinate standard values, sequence coding region overlapping range standard values, cross-sample base complementary pairing pattern standard values, protein interaction network anomaly identification standard values, genotype-phenotype association strength discrete distribution quantity standard values. The analysis task priority sequence includes pathogenic gene annotation semantic weights, mutation function prediction logical dependency relationships, and structural variation classification boundary numerical differences. The task execution optimization coefficient includes variant classification boundaries, promoter region mutation distribution ranges, exon splicing site mutation distribution ranges, and analysis execution order adjustment parameters. The association resolution set includes gene expression profile clustering center coordinates, mutation site enrichment regions, and three-dimensional spatial distributions of protein interaction network topological parameters.
3. The bioinformatics analysis system based on large model technology according to claim 2, wherein The data parsing and calibration module includes: The genomic sequence alignment sub-module obtains genomic sequence alignment information and clinical phenotype data timestamps, parses the genomic sequence alignment data, extracts the alignment results between samples, calculates the alignment scores, obtains the mutation sites in the genomic alignment sequences, screens the regions with high alignment coverage, and associates the genomic variant information of each sample according to the phenotype data timestamp to generate high-coverage gene sequence alignment data; The gene variant functional region detection sub-module, based on the high-coverage gene sequence alignment data, detects the gene variant site coordinates, analyzes the overlapping situation between the variant sites and the genomic sequence coding regions, determines whether the variant sites fall into functional gene regions, calculates the overlapping ratio of the variant sites, screens the mutation sites with high overlapping ratios, and generates a set of variant sites in functional regions; The genotype-phenotype association calculation sub-module calls the set of variant sites in functional regions, extracts the cross-sample base complementary pairing patterns, analyzes the distribution of different genotype combinations in the phenotype data, calculates the discrete distribution quantity of genotype-phenotype association strength, and uses the formula: Calculate the genotype-phenotype association strength through operations, screen the association patterns that meet the distribution characteristics, standardize the multi-modal data, and obtain the standard values of the multi-modal data; Among them, S represents the genotype-phenotype association strength distribution value, represents the value of the i g th genotype variable, represents the phenotype variable value corresponding to , represents the value of the j g th uncorrelated variant variable, represents the phenotype variable value corresponding to , n g represents the number of genotype variables, m g represents the number of uncorrelated variant variables.
4. The bioinformatics analysis system based on large model technology according to claim 3, wherein The problem decomposition module includes: Based on the standard values of the multi-modal data, the standard value calculation sub-module analyzes the standard range of the differential gene feature parameters, invokes the gene expression level, mutation frequency, and epigenetic regulatory factors of the target individual, calculates the deviation of each feature parameter from the standard value, and analyzes the overall variation through cumulative deviation. The formula is as follows: Calculate the deviation amount of the standard value calculation through operations and establish the standard value deviation result; Among them, D represents the deviation of the standard value calculation, represents the value of the i s -th gene feature parameter, represents the corresponding standard value of the multimodal data, n s represents the number of gene feature parameters, represents the value of the individual environmental factor, represents the reference value of the environmental factor, m s represents the number of environmental factors; Based on the standard value deviation result, the pathogenic gene annotation sub-module analyzes the functional characteristics of the pathogenic genes, invokes the known gene database to extract gene association characteristics, analyzes the mutation sites and influence ranges of the gene sequences, calculates the semantic weight value of the mutation function prediction logic, obtains the gene semantic weight calculation value, and establishes the pathogenic gene function annotation result; The gene structure variation classification sub-module invokes the pathogenic gene function annotation result, screens the structural differences corresponding to the gene variations, calculates the classification boundary values of the multi-structural variations, compares with the standard value deviation result, generates the numerical difference of the structural variation classification boundary, and establishes the analysis task priority sequence.
5. The bioinformatics analysis system based on large model technology according to claim 4, characterized in that, The analysis task scheduling module includes: Based on the analysis task priority sequence, the variation classification boundary identification sub-module extracts the variation category information corresponding to the multi-analysis tasks, calculates the distribution interval of the mutation types, compares with the benchmark value of the mutation influence degree, identifies the classification boundary of the mutation, calculates the deviation amount of the multi-category boundary, and generates the variation classification boundary value; The mutation distribution screening sub-module invokes the variation classification boundary value, screens the distribution ranges of the promoter region mutations and exon splicing site mutations, calculates the distribution ratios of the multi-mutation categories in the differential regions, and classifies and sorts the distributions according to the mutation function influence coefficient, generating the mutation distribution classification interval; The task execution order adjustment sub-module invokes the mutation distribution classification interval, adjusts the analysis task execution order according to the task priority, using the formula: Calculate the task execution optimization coefficient; Among them, T opt represents the task execution optimization coefficient, represents the adjusted task distribution ratio, represents the classification boundary value before adjustment, represents the priority weight of the task, m x represents the number of analysis tasks.
6. The bioinformatics analysis system based on large model technology according to claim 5, wherein The result mapping module includes: Based on the task execution optimization coefficient, the task optimization coefficient calculation sub-module extracts the mutation site distribution parameters, protein interaction network connectivity parameters, and gene expression level parameters, using the formula: Calculate the optimization adjustment coefficient through operations; Among them, C opt represents the optimization adjustment coefficient, V p represents the distribution density weighted coefficient of mutation sites, represents the distribution density of mutation sites in the protein interaction network, α represents the adjustment coefficient, L d represents the deviation degree of gene expression level, V n represents the network connectivity, S c represents the gene expression mean shift adjustment coefficient; Based on the optimization adjustment coefficient, the gene expression profile clustering sub-module extracts the gene expression profile data, calculates the gene expression mean value, identifies the deviation of the mutant genes in the expression profile, screens the genes with the expression deviation value exceeding the set threshold, and calculates the expression profile clustering center coordinates of the mutant genes to obtain the gene expression clustering center coordinates; Based on the gene expression clustering center coordinates, the gene network three-dimensional topology mapping sub-module calculates the topological parameters of the mutant gene enrichment region in the protein interaction network, maps its three-dimensional spatial distribution, and generates an association resolution set according to the connection conditions of the multi-genes in the topological network.
7. The bioinformatics analysis system based on large model technology according to claim 6, wherein The system also includes: The feedback iteration module detects the matching degree between the coverage rate of structural variation hotspots and the functional annotation correlation matrix in the associated interpretation degree set, adjusts the determination threshold of the sequence coding region, updates the abnormal sample elimination rule for cross-modal data, and generates a set of process optimization parameters; The set of process optimization parameters includes the coverage rate of structural variation hotspots, the matching degree of the functional annotation correlation matrix, the determination threshold of the sequence coding region, and the abnormal sample elimination rule for cross-modal data.
8. The bioinformatics analysis system based on large model technology according to claim 7, wherein The feedback iteration module includes: The structural variation hotspot coverage sub-module calculates the coverage rate of multiple hotspots in the differential sequence coding region based on the coverage rate of structural variation hotspots in the associated interpretation degree set, calls the determination threshold to screen the hotspot regions that meet the requirements, and uses the formula: Calculate the hotspot coverage metric value through operation; Among them, C hot represents the hotspot coverage metric value, represents the coverage probability of the i t -th structural variation hotspot, represents the associated interpretability score of the i t -th structural variation hotspot, δ s represents the adjustment factor, V max represents the maximum associated interpretability score, N s represents the number of hotspots; The functional annotation correlation matrix matching sub-module calls the hotspot coverage metric value, calculates the matching degree of multiple coding regions based on the functional annotation correlation matrix, extracts the matching scores of multiple regions, screens the regions with matching scores higher than the set threshold, eliminates the coding regions with low matching degrees, and obtains the functional matching coefficient; The cross-modal data anomaly elimination sub-module calls the functional matching coefficient, adjusts the anomaly elimination rule for cross-modal data, calculates the deviation coefficient of each abnormal sample, screens the abnormal samples according to the deviation coefficient threshold, eliminates the data with deviations exceeding the standard range, and obtains a set of process optimization parameters.
Citation Information
Patent Citations
Convolutional neural network-based variation clinical effect analysis and identification method and equipment
CN115482879A
Biomedicine data statistical system and analysis method thereof
CN119091108A
Methods and systems for genome analysis
US20160092631A1
Inter-model prediction score recalibration
US20230223100A1
Cited By
Multi-omics data integration plant gene function inference system and method based on large language model
CN120853669A
Bioinformation comparative analysis method for gastric cancer pathogenic gene sequence
CN121811972A
Method for bioinformatics alignment analysis of gastric cancer pathogenic gene sequence
CN121811972B