The invention relates to the technical field of
bioinformatics, in particular to a biological
information analysis system based on a
large model technology, which comprises a
data analysis calibration module, a problem disassembly module, an analysis task arrangement module, a result mapping module and a feedback iteration module. According to the method,
genome comparison and
clinical phenotype timestamps are dynamically calibrated,
time sequence dislocation deviation is eliminated, base complementary
pairing is combined with
protein network anomaly screening, low-abundance collaborative variation capture is enhanced,
genotype-
phenotype discrete distribution quantifies and unifies multi-
modal data benchmark, and the problem of multi-source heterogeneous
standardization deficiency is solved; the method comprises the following steps: classifying and integrating pathogenic
gene semantic weights by
structural variation, balancing a statistical threshold and a biological function, dynamically optimizing an analysis sequence, synchronously covering a key
mutation region, improving function
annotation of a non-
coding region, integrating
gene expression clustering and
protein network topology in a three-dimensional distribution manner, breaking through two-dimensional space limitation, performing closed-loop feedback to correct a threshold iteration
elimination rule, and finally obtaining a high-quality
gene expression cluster. And the genetic heterogeneity
false positive rate is reduced.