A model for identifying allele imbalance markers driving tumor evolution based on a hierarchical bayesian framework and a construction method thereof

By constructing a model using a hierarchical Bayesian framework and utilizing whole-genome and multi-omics sequencing data, the challenges of handling genomic bias and clonal heterogeneity in existing technologies have been addressed. This has enabled accurate identification and robust comparison of tumor allele imbalance events, improving the accuracy and reliability of identification.

CN122435987APending Publication Date: 2026-07-21ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2026-04-27
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively correct for genomic and human biases, handle clonal and regional heterogeneity, and lack robust inter-sample statistical comparison frameworks, resulting in insufficient accuracy and reliability in identifying tumor allele imbalances.

Method used

A hierarchical Bayesian framework was used to construct a model, and tumor genome evolution map and allele expression profile were constructed using whole genome sequencing and multi-omics sequencing data. The hierarchical Bayesian model was then used to perform posterior distribution inference, correct genomic bias, and identify allele imbalance events.

Benefits of technology

It enables accurate identification and robust comparison of allele imbalance events during tumor evolution, improving the accuracy and reliability of identification. It can systematically correct multi-level biases and adapt to regional data sparsity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122435987A_ABST
    Figure CN122435987A_ABST
Patent Text Reader

Abstract

The application provides a model for identifying allele imbalance markers driving tumor evolution based on a hierarchical Bayesian framework and a construction method thereof, and belongs to the field of information technology.The application provides a construction method of a model for identifying allele imbalance markers driving tumor evolution based on a hierarchical Bayesian framework, a three-level Bayesian hierarchical model is constructed, global noise, subclone specificity, regional discreteness and allele single nucleotide mutation (SNV) site observation are jointly modeled, and a reparameterization correction allele copy number deviation is introduced.On this basis, a Markov Monte Carlo (MCMC) sampling is used to obtain a posterior distribution of parameters, and a sample comparison method of a posterior distribution probability of each parameter including a true allele imbalance coefficient and a Kullback-Leibler divergence is provided, which can be used for identifying allele imbalance and analyzing markers driving evolution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information technology, specifically relating to a model for identifying allele imbalance markers driving tumor evolution based on a hierarchical Bayesian framework and its construction method. Background Technology

[0002] Tumor adaptive evolution provides the cellular and molecular basis for tumor spread, metastasis, drug resistance, and immune escape. Tumor evolution is spatiotemporally specific and subject to various systemic and biological noises. Therefore, studying key drivers of tumor evolution is of significant clinical importance but also presents considerable challenges. In tumor genomics research, accurately identifying clonal allele imbalance (AI) events is crucial for understanding tumor heterogeneity, clonal evolution, and drug resistance mechanisms. Humans are diploid organisms, meaning that each normal cell contains two copies of the same gene. During tumorigenesis and development, short-sequence mutations often occur in the intracellular genome, accompanied by large-scale structural and copy number variations. These variations provide the molecular driving force for tumor evolution and lead to a selection shift in gene expression or epigenetic modifications controlling gene expression between two different alleles. Allele imbalance events have been reported in various cancers, and these events play important roles in tumorigenesis, metastasis, and evolution. Currently, algorithms combining tumor genomic evolution with the identification of allele imbalance events are very limited and face the following technical challenges: (1) Genomic and human bias interference: Traditional allele ratio analysis is susceptible to two types of bias. One is the bias caused by the inherent allele-specific copy number variation in the genome; the other is the human bias introduced by the experimental and sequencing processes, such as insufficient sequencing depth and reference genome alignment bias. These biases are coupled with each other and are transmitted from the whole genome level to local regions, resulting in false positive or false negative results.

[0003] (2) Difficulty in handling clonal and regional heterogeneity: Tumor samples contain complex clonal hierarchies (such as subclonal copy number variations), normal cell contamination, and sparse heterozygous SNV density in different genomic regions. Existing methods are difficult to handle both the overall shift at the subclonal level and the discreteness at the regional level at the same time, often leading to overfitting to low-information regions or excessive sensitivity to high-noise regions.

[0004] (3) Lack of statistical support for inter-sample comparisons: For the dynamic changes of allele imbalance before and after treatment (such as primary lesions and drug-resistant lesions), existing methods mostly only compare point estimates (such as the mean), ignoring the full picture of the posterior distribution, making it difficult to accurately quantify changes in distribution morphology (such as shifts and changes in dispersion), and lacking robust statistical inference methods.

[0005] Therefore, there is an urgent need for a method for identifying allele imbalances that can systematically correct multi-level biases, adapt to regional data sparsity, and provide a comprehensive statistical comparison framework. Summary of the Invention

[0006] This invention provides a model and its construction method for identifying allelic imbalance markers driving tumor evolution based on a hierarchical Bayesian framework. It can solve the problems in the prior art that it is difficult to effectively correct genomic and technological biases, cannot reasonably handle clonal and regional heterogeneity, and lacks a robust statistical comparison framework between samples. It provides a statistical modeling method for identifying allelic imbalance events that drive drug resistance in tumor clonal evolution analysis and performing inter-sample comparisons.

[0007] This invention provides a method for constructing a model based on a hierarchical Bayesian framework to identify allele imbalance markers that drive tumor evolution, comprising the following steps: (1) obtaining tumor tissues from different regions and / or treatment nodes in humans, and performing whole-genome sequencing or whole-exome sequencing on spatiotemporally specific tumor tissues, and constructing a tumor genome evolution map based on the sequencing results; (2) Using the spatiotemporally specific tumor tissues in step (1) to perform several omics sequencing, and construct allele expression profiles based on the obtained omics sequencing results; (3) Using the tumor genome evolution map described in step (1) and the allele expression profile described in step (2) as input data, construct a hierarchical Bayesian model; (4) Solve the hierarchical Bayesian model described in step (3) and infer the posterior distribution to obtain the actual allele imbalance ratio that occurs during tumor evolution and construct an allele imbalance event map. (5) Based on the allele imbalance ratio obtained in step (4), a model for identifying key driver biomarkers in tumor evolution is constructed by comparing spatiotemporally specific samples, and key driver biomarkers (genes or regulatory elements) in tumor evolution are obtained by comparing the differences between spatiotemporally specific samples.

[0008] In one specific embodiment of the present invention, in step (1), based on the whole genome sequencing results, heterozygous site mutations and genome copy number variations in the tumor genome are calculated, thereby constructing a tumor genome evolution map.

[0009] In one specific embodiment of the present invention, the omics sequencing in step (2) includes at least one of the following: transcriptome sequencing, open chromatin sequencing, three-dimensional genome sequencing and chromatin immunoprecipitation sequencing.

[0010] In one specific embodiment of the present invention, in step (2), the allele sequence depth of the corresponding heterozygous mutation site is calculated based on the omics sequencing results, thereby constructing an allele expression profile.

[0011] In one specific embodiment of the present invention, the hierarchical Bayesian model in step (3) includes three levels: upper, middle and lower, with information flow transmitted from top to bottom; The upper layer is responsible for handling alignment bias and allele ratio bias of the whole genome, providing a prior distribution for the mean at the subcloning level; The prior distribution of the middle and upper layers is used as input, and the variance of the regional and global allele ratios is measured by combining the heterozygous mutation density of the segment. The output is the shape parameter of the gamma function, which is used as the prior distribution of the lower layer. The lower layer uses the shape parameters of the middle layer and the allele ratio bias of the upper layer, and combines them with the obtained allele expression profile to obtain the prior distribution of single heterozygous site mutations and the corresponding θ distribution; θ is the observed proportion of variant alleles.

[0012] In one specific embodiment of the present invention, the obtained θ is converted into the actual allele imbalance ratio a: ; Where b represents the proportion of variant alleles with mutations at the same heterozygous site in the genome.

[0013] In one specific embodiment of the present invention, the model solving and posterior distribution inference in step (4) include implementing the model in parallel using the probabilistic programming language Stan, sampling the model parameters using Hamiltonian Monte Carlo and its No-U-Turn sampler algorithm, and then diagnosing convergence by using the Gelman-Rubin statistic and effective sample size.

[0014] In one specific embodiment of the present invention, after the identification model is constructed in step (5), the method further includes single-sample allele imbalance identification and multi-sample comparison.

[0015] The present invention also provides a recognition model constructed using the above-described construction method.

[0016] The present invention also provides a computer-readable storage medium on which a program is stored, and when the program is executed by a processor, the above-mentioned recognition model is implemented.

[0017] Beneficial Effects: This invention provides a method for constructing a model based on a hierarchical Bayesian framework to identify allelic imbalance markers driving tumor evolution. First, tumor tissues from different regions or treatment nodes in humans are obtained. Whole-genome sequencing is performed on spatiotemporally specific tumor tissues, along with matching sequencing of several other omics sequences corresponding to the same spatiotemporal location. Somatic mutation analysis is conducted on spatiotemporally specific tumors using whole-genome sequencing data to construct a tumor clonal evolutionary genomic region map. Expression profiles of spatiotemporally specific alleles or heterozygous mutation sites of epigenetic regulatory elements are constructed using multi-omics data. Combining spatiotemporally specific multi-omics data, global systematic errors, such as reference sequence alignment errors and allelic ratio bias, are used as hyperparameters at the highest level of the Bayesian hierarchical model. After establishing the global prior, the allelic frequency of heterozygous mutation sites in each subclonal region follows a certain order. The algorithm employs a normal distribution as a baseline hyperparameter. Simultaneously, it utilizes an adaptive weighted strategy with a logistic scaling sigmoid function to model the transitional discreteness of genomic fragments, incorporating the density of regional heterozygous mutations and calculating the shape parameter of the gamma function. Finally, it models single heterozygous mutation sites using a Beta-binomial distribution, correcting for genomic allele copy number shifts. Markov Monte Carlo sampling is used to sample model parameters, and a posterior distribution of the true allele imbalance coefficient is constructed. Combining population genetics haplotype phase analysis results, different alleles of heterozygous mutations are weighted and integrated based on sequencing depth to obtain allele imbalance coefficients at the gene or epigenetic regulatory element level. The algorithm also provides a Kullback-Leibler method-based analysis of the posterior distribution differences of allele imbalance coefficients among tumors in different time and space.

[0018] This invention constructs a three-level Bayesian hierarchical model to jointly model global noise, subclone specificity, regional discreteness, and single nucleotide variant (SNV) site observations, and introduces reparameterization to correct for allele copy number bias. Based on this, Markov Monte Carlo (MCMC) sampling is used to obtain the posterior distribution of parameters, and a method for inter-sample comparison based on the posterior distribution probability and Kullback-Leibler divergence is provided. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the hierarchical Bayesian model framework of the present invention (showing the information flow, distribution assumptions, and parameter relationships of the global layer, segment layer, and single SNV layer); Figure 2 The figure shows the results of the benchmark test of the simulated data. In the figure, HBOmics is the name of the model of this invention, and BaalChIP is one of the few existing algorithms that considers the genomic allele copy number shift but cannot integrate the tumor subclonal structure. Figure 3 The graph shows the results of benchmark tests using real data. Detailed Implementation

[0020] This invention provides a method for constructing a model for identifying allele imbalance markers in tumor evolution based on a hierarchical Bayesian framework, comprising the following steps: (1) obtaining tumor tissues from different regions and / or treatment nodes in humans, and performing whole-genome sequencing or whole-exome sequencing on spatiotemporally specific tumor tissues, and constructing a tumor genome evolution map based on the sequencing results; (2) Using the spatiotemporally specific tumor tissues in step (1) to perform several omics sequencing, and construct allele expression profiles based on the obtained omics sequencing results; (3) Using the tumor genome evolution map described in step (1) and the allele expression profile described in step (2) as input data, construct a hierarchical Bayesian model; (4) Solve the hierarchical Bayesian model described in step (3) and infer the posterior distribution to obtain the actual allele imbalance ratio that occurs during tumor evolution and construct an allele imbalance event map. (5) Based on the allele imbalance ratio obtained in step (4), a model for identifying key driver biomarkers in tumor evolution is constructed by comparing spatiotemporally specific samples, and key driver biomarkers in tumor evolution are obtained by comparing the differences between spatiotemporally specific samples.

[0021] This invention collects tumor tissues from different regions, stages, before and after treatment, and before and after metastasis from the same patient, and performs whole-genome sequencing or whole-exome sequencing on the collected spatiotemporally specific tumor tissues to obtain multiple sets of spatiotemporally specific tumor whole-genome data. Based on the tumor whole-genome data, a tumor genome evolution map is constructed.

[0022] This invention does not specifically limit the construction process of the tumor genome evolution map, but requires that the tumor genome evolution map contain at least the cancer cell fraction of each heterozygous SNV, and that the whole genome be divided into subclonal segments, i.e., different genomic regions correspond to different cancer cell fractions, thereby obtaining a spatiotemporally specific clonal evolution structure of the tumor. Specifically, this is manifested in segmenting the tumor genome and assigning different cancer cell fractions to these segments, with different cancer cell fractions corresponding to their origin from different tumor subclonal characteristics. In one embodiment of this invention, based on whole tumor genome data, heterozygous site mutations (het.SNVs) and genomic copy number variations in the tumor genome are calculated using the GATK tool; the TitanCNA algorithm is used to construct the tumor genome evolution map by utilizing genomic copy number variations and heterozygous SNVs, and the allele ratios can be obtained from genome sequencing.

[0023] This invention performs spatiotemporally specific tumor tissue sequencing on collected tumor tissues, corresponding to several omics sequences in the same spatiotemporal context as whole-genome sequencing or whole-exome sequencing. The omics sequencing can be at least one of the following: transcriptome sequencing (RNA-seq), chromatin immunoprecipitation sequencing (ChIP-seq), three-dimensional genome sequencing (ChIA-PET / HiChIP), and chromatin open sequencing (ATAC-seq) data, etc. This allows for comparison of tumor tissues from different regions, at different stages, before and after treatment, and before and after metastasis from the same patient.

[0024] This invention utilizes multi-omics data from spatiotemporally specific tumor tissues and employs the Samtools tool to calculate the allele sequence depth at corresponding heterozygous mutation sites, thereby constructing allele expression profiles of heterozygous SNV transcriptomes or epigenomes. These allele expression profiles are then used to calculate the allele ratios of gene expression or epigenetic modification signals.

[0025] This invention utilizes the two sets of data obtained above as data inputs to a hierarchical Bayesian model framework to obtain the true allele imbalance ratio occurring during tumor evolution. Specifically, it combines the allele ratio and subclonal origin of genomic heterozygous SNVs with allele expression profile data obtained from multi-omics data as input. For example, for a heterozygous mutation site, the corresponding genomic allele ratio is 1:1 (two identical allele copies), and the corresponding subclonal tumor cell ratio is 1 (original clonal population). However, in the transcriptome or epigenome (i.e., during gene expression or histone modification), this allele ratio shifts to 10:1, thus obtaining the true allele imbalance ratio occurring during tumor evolution.

[0026] The hierarchical Bayesian model described in this invention comprises three levels, with information flowing from top to bottom: Upper layer (global level): To address systematic errors affecting the entire sequencing library, a global alignment bias hyperparameter μ is set. g and global allele ratio variance hyperparameter c g This shrinkage estimation strategy prevents overfitting to subclones with large deviations in allele proportions. The average allele proportion for each subclone is calculated. m t Modeling is based on sampling from a global distribution: , Formula I; Mid-level (genomic segment level): Modeling the excessive dispersion of genomic segments (such as chromosomal segments). An adaptive weighting strategy is introduced to dynamically adjust the weights of variance sources based on the density of heterozygous SNVs within a segment: Formula II; Formula III; In Equations II and III, α j For fragments j The shape parameters of the gamma distribution, v j For local segment variance, F t For subclonal variance, w Based on the number of SNVs n Density-dependent weights. In SNV-rich regions ( w Approaching 1) Prioritize the use of local data; in sparse SNV regions ( w If the variance is close to 0, it depends on the subclonal level variance; Lower layer (single SNV level): The observed allele counts are modeled using a reparameterized Beta-binomial distribution, and the fragment discreteness estimated in the middle layer is passed to a single locus. Formula IV; , formula V; , VI; Equation VII; In equations IV to VII, A To count the variant alleles, N The total depth of this allele, i The observed proportion of variant alleles, m and c Then it is i The mean and accuracy.

[0027] In this invention, global and local noise from the genome is passed down from top to bottom to the likelihood of the allele ratio at the final single SNV locus. The upper-level model handles alignment bias and allele ratio bias across the entire genome, providing a prior distribution for the mean at the subclonal level. The middle layer handles local noise from different genomic segments obtained during the construction of the tumor genome evolution map. The middle layer takes the prior distribution from the upper layer as input, combines the heterozygous SNV density of the segment to measure the variance of the regional and global allele ratios, and outputs the shape parameters of a gamma function, which are then used as the basis for the lower layer. c The prior distribution of the upper layer is used; the lower layer uses the shape parameters of the middle layer, and the mean and variance of the allele proportions of the upper layer are used. Combined with the allele expression profiles of the heterozygous SNV transcriptome or epigenome, the single SNV is obtained. c Prior distribution and corresponding i distributed.

[0028] This invention is based on the parameters in the hierarchical Bayesian framework obtained above. c , m The method involves correcting for genomic allele copy number bias through parameter transposition to obtain the true allele imbalance ratio 'a'. This invention introduces a reparameterization formula in this step to correct for biases caused by genomic allele-specific copy number variations, thus accurately determining the observed allele ratio. i Converted to true allele imbalance ratio a: Formula VIII; In formula VIII, b This indicates the proportion of alleles for the same SNV variant in the genome. This step removes false-positive imbalance signals caused purely by copy number variations.

[0029] Then, by substituting the parameter transpose formula into equations IV to VII, the complete parameter relationships and distributions of the hierarchical Bayesian framework can be obtained. After the hierarchical Bayesian model is constructed, the present invention performs model solving and posterior distribution inference. The above model is implemented in parallel using the probabilistic programming language Stan, and the Hamiltonian Monte Carlo (HMC) algorithm and its No-U-Turn Sampler (NUTS) algorithm are used to sample the model parameters (including the joint posterior distribution of the true allele imbalance ratio α). The Gelman-Rubin statistic is then used. The effective sample size (ESS) is used to diagnose convergence.

[0030] After constructing the model, this invention further processes it by integrating the obtained SNV unit point results into gene-level and multi-gene-level applications, such as single-sample allele imbalance identification. Specifically, based on the posterior distribution of α for each SNV locus obtained in the above steps, the biallelic hypothesis (α=0.5) and the imbalance hypothesis (α close to 0 or 1) are tested. The nominal... p Value: Calculate the proportion of α greater than or less than 0.5 from the posterior distribution chain, and take the smaller value multiplied by 2 (or one-sided test) as the significance criterion.

[0031] Formula IX; In equation IX, s Let be the posterior sample size, and 1(·) be the indicator function.

[0032] This invention determines the value of 'a' with the highest posterior probability as the true allele imbalance ratio at that locus. Thus, the method of this invention obtains the posterior distribution of the allele imbalance ratio 'a' occurring in gene expression or epigenetic modification at each heterozygous SNV locus, corresponding to... p Values, and the true allele imbalance ratio.

[0033] This invention also requires multi-sample comparisons, that is, using the posterior distribution of the allele imbalance ratio α generated in the above steps, to compare the α posterior distribution of the same SNV locus between two samples (e.g., primary sample C and relapsed / drug-resistant sample T): (1) Direct comparison method: Compare the true allele imbalance ratio (the a value corresponding to the highest posterior probability) and its confidence interval (CIs) of the two samples, and calculate the ratio difference.

[0034] (2) KL divergence comparison method: Calculate the Kullback-Leibler divergence between the posterior distributions of two samples to quantify the difference in distribution shape: Formula X; For feature levels (such as gene, chromatin accessibility peaks, or histone modification peaks), a weighted integration is performed combining haplotype phase information and allele depth: , Formula XI; in depth Ti SNV site i Total allele depth. A higher KL divergence value indicates a more significant difference in the distribution pattern of allele imbalance among samples.

[0035] The present invention also provides a recognition model constructed using the above-described construction method.

[0036] The hierarchical Bayesian model described in this invention has Figure 1 The schematic diagram of the framework is shown and named HBomics.

[0037] The present invention also provides a computer-readable storage medium on which a program is stored, and when the program is executed by a processor, the above-mentioned recognition model is implemented.

[0038] To further illustrate the present invention, the following detailed description, in conjunction with embodiments, of a model and its construction method for identifying allele imbalance markers driving tumor evolution based on a hierarchical Bayesian framework, but these descriptions should not be construed as limiting the scope of protection of the present invention.

[0039] Example 1: Simulated Data Benchmark Test This invention primarily evaluates the performance of the constructed model HBOmics on simulated data and compares it with BaalChIP. BaalChIP is one of the few algorithms specifically designed for tumor genomes and performs allele copy number correction, but it cannot systematically consider global and local errors and incorporate tumor evolution structure. The simulated allele ratios are derived from a Beta-binomial distribution model, and the simulated true allele imbalance ratio is normalized to a randomized genomic allele copy number ratio. The simulated true allele imbalance ratio is set to a range of 0.05 to 0.95. Sequencing depth and allele ratio bias are introduced as the main confounding factors in the simulation. The simulated sequencing depths are set to 10×, 20×, and 30×, and the allele ratio variance is set to 0.01, 0.05, 0.10, and 0.15. BaalChIP is run using default parameters, with the biallelic ratio range set to [0.4, 0.6]. By reparameterizing the observed enriched allele ratios, this invention fixes the genomic allele ratios and incorporates the simulated systematic allele ratio bias into a Beta-binomial distribution model to calculate the observed enriched allele ratios.

[0040] The results are as follows Figure 2 As shown, the HBOmics described in this invention, compared to BaalChIP, can stably and accurately identify the true allele imbalance ratio under different allele frequency errors and sequencing depths. Specifically, under different conditions, the F1 score of this invention is above 0.9, which is a significant improvement compared to BaalChIP (around 0.7). Furthermore, it shows significant improvements over BaalChIP in both prediction precision and recall, two evaluation criteria with different focuses.

[0041] Example 2: Real Data Benchmarking To comprehensively evaluate the performance of HBomics in identifying allele imbalance events in tumor clonal evolution, this invention designed a pseudoclonal sample by mixing drug-resistant cell lines with their parental cell lines in a known ratio. Since the true allele imbalance ratio of tumors is almost impossible to know, a clonal structure with a known cell clone ratio was constructed. This invention screened SNVs from RNA-seq data using the following criteria: At the genomic level, the coverage depth of SNVs in both samples must exceed 30 sequencing sequences.

[0042] The variant alleles of SNV should be expressed in drug-resistant cells, but completely silenced in parental cells.

[0043] In this way, the RNA-seq allele ratio of drug-resistant cells can be regarded as the true allele imbalance ratio of the pseudoclonal tumor, while the reference allele of parental cells serves as interference from normal cells on the true allele frequency, affecting the final observed allele ratio. Therefore, the core task of the benchmark test is to examine whether the HPomics of this invention can reliably calculate the true subclonal allele imbalance ratio under the interference of confounding allele bias from biological sample sources. In each iteration, the pseudoclonal mixing ratio is randomly sampled from 0.7 to 0.9 (i.e., the parental cell ratio is 0.1 to 0.3), and five random ratios are used for testing in each iteration. BaalChIP is run with default parameters, and the biallelic ratio range is set to [0.4, 0.6].

[0044] The results are as follows Figure 3 As shown, the true allele imbalance ratio predicted by HBOmics in this invention has a higher good of fit than the true allele imbalance ratio in the actual data, especially when the true allele imbalance ratio is large. The allele imbalance ratio predicted by this invention is closer to the true value than BaalChIP. The true allele imbalance ratio is divided into biased, reference allelic, and alternative allelic events, and a confusion matrix is ​​constructed. The results show that, without causing misclassification of biased alleles, this invention can accurately identify more allele imbalance events, and the accuracy is not biased in the direction of allele shift.

[0045] Although the above embodiments have provided a detailed description of the present invention, they are only some embodiments of the present invention, and not all embodiments. People can obtain other embodiments based on these embodiments without creative effort, and these embodiments all fall within the protection scope of the present invention.

Claims

1. A method for constructing a model based on a hierarchical Bayesian framework for identifying allelic imbalance markers driving tumor evolution, characterized in that, Includes the following steps: (1) Obtain tumor tissues from different regions and / or treatment nodes in humans, and perform whole-genome sequencing or whole-exome sequencing on spatiotemporally specific tumor tissues, and construct a tumor genome evolution map based on the sequencing results; (2) Using the spatiotemporally specific tumor tissues in step (1) to perform several omics sequencing, and construct allele expression profiles based on the obtained omics sequencing results; (3) Using the tumor genome evolution map described in step (1) and the allele expression profile described in step (2) as input data, construct a hierarchical Bayesian model; (4) Solve the hierarchical Bayesian model described in step (3) and infer the posterior distribution to obtain the actual allele imbalance ratio that occurs during tumor evolution, and construct an allele imbalance event map; (5) Based on the allele imbalance ratio obtained in step (4), an identification model of the allele imbalance event map in the tumor evolution process is constructed, and key driving markers in tumor evolution are obtained by comparing the differences between spatiotemporally specific samples.

2. The construction method according to claim 1, characterized in that, In step (1), based on the whole genome sequencing results, heterozygous site mutations and genome copy number variations in the tumor genome are calculated to construct a tumor genome evolution map.

3. The construction method according to claim 1, characterized in that, The omics sequencing in step (2) includes at least one of the following: transcriptome sequencing, open chromatin sequencing, three-dimensional genome sequencing and chromatin immunoprecipitation sequencing.

4. The construction method according to claim 1 or 3, characterized in that, In step (2), the allele sequence depth of the corresponding heterozygous mutation site is calculated based on the omics sequencing results, thereby constructing the allele expression profile.

5. The construction method according to claim 1, characterized in that, The hierarchical Bayesian model described in step (3) consists of three levels: upper, middle, and lower, with information flowing from top to bottom. The upper layer is responsible for handling alignment bias and allele ratio bias of the whole genome, providing a prior distribution for the mean at the subcloning level; The prior distribution of the middle and upper layers is used as input, and the variance of the regional and global allele ratios is measured by combining the heterozygous mutation density of the segment. The output is the shape parameter of the gamma function, which is used as the prior distribution of the lower layer. The lower layer uses the shape parameters of the middle layer and the allele ratio bias of the upper layer, and combines them with the obtained allele expression profile to obtain the prior distribution of single heterozygous site mutations and the corresponding θ distribution; θ is the observed proportion of variant alleles.

6. The construction method according to claim 5, characterized in that, Convert the obtained θ into the true allele imbalance ratio a: ; Where b represents the proportion of variant alleles with mutations at the same heterozygous site in the genome.

7. The construction method according to claim 1, characterized in that, Step (4) involves solving the model and inferring the posterior distribution, including implementing the model in parallel using the probabilistic programming language Stan, sampling the model parameters using Hamiltonian Monte Carlo and its No-U-Turn sampler algorithm, and then diagnosing convergence using the Gelman-Rubin statistic and effective sample size.

8. The construction method according to claim 1, characterized in that, After the identification model is constructed in step (5), the process also includes single-sample allele imbalance identification and multi-sample comparison.

9. The recognition model constructed using the construction method described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the operation of the recognition model of claim 9.