A method for verifying function of BRCA1 / 2 gene unknown significance variation
By designing single-point editable gRNAs and HDR templates, and combining them with haploid host cells for gene editing and dynamic abundance detection, the problem of low throughput and high cost in functional verification of BRCA1/2 gene variants with unclear significance has been solved. This enables flexible and accurate functional verification, making it suitable for clinical laboratories.
Patent Information
- Application Number
- CN202511709833.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-03-17
- Estimated Expiration
- 2045-11-20
AI Technical Summary
Existing methods for functional verification of BRCA1/2 gene variants with unclear significance suffer from low throughput, high cost, and incomplete coverage, making it difficult to meet high-throughput requirements. Furthermore, these methods are complex to operate and result in significant resource waste.
We designed gRNAs and HDR templates for single-point editing, combined with haploid host cells for gene editing, and precisely verified the function of unsigned variants in the BRCA1/2 genes through dynamic abundance detection and functional scoring.
It enables flexible and precise functional validation of BRCA1/2 gene variants of unknown significance, reduces operational complexity and cost, and improves the accuracy and throughput of results, making it suitable for clinical laboratories.
Smart Images

Figure CN121171336B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology, specifically relating to a... BRCA1 / 2 Methods for verifying the function of variants with unclear gene significance. Background Technology
[0002] BRCA1 / 2 Genes BRCA1 Gene or BRCA2 Gene, BRCA1 The gene is breast cancer gene 1. BRCA2 The gene is breast cancer gene number 2. BRCA1 / 2 These are two genes that suppress the development of malignant tumors and play important roles in regulating cell replication, DNA damage repair, and normal cell growth. Families with this gene mutation tend to have a higher incidence of breast cancer, usually occurring at a younger age, and also have a higher risk of contralateral breast and ovarian cancer.
[0003] BRCA1 / 2 Genetic testing has been widely applied in clinical practice. However, variants of uncertain significance (VUS) pose significant challenges to patient consultation, clinical decision-making, and the allocation of medical resources due to the uncertainty of their clinical meaning. Asian populations BRCA1 / 2 The detection rate of variants of unknown significance was as high as 47.8% (Kurian A W, Ward KC, Abrahamse P, et al. Time Trends in Receipt of Germline Genetic Testing and Results for Women Diagnosed With Breast Cancer or Ovarian Cancer, 2012-2019[J]. J Clin Oncol , 2021,39(15):1631-1640.), the existing functional data is far from meeting clinical needs.
[0004] Traditional methods for validating the function of variants of unknown significance are often only applicable to the assessment of specific functional domains or types of variants, and are difficult to meet high-throughput requirements. CRISPR-Cas9-based gene editing technology can introduce testable variants into disease genes while preserving the natural genomic background, thereby improving the accuracy of functional studies by maintaining endogenous cellular transcription and splicing mechanisms, and greatly simplifying and accelerating the site-directed mutagenesis process. However, generating stable cell lines carrying specific target variants using CRISPR-based editing technology remains a time-consuming and labor-intensive process, typically requiring over a month for single-clone screening. Although saturation genome editing (SGE) can achieve parallel analysis of multiple sites and high-throughput evaluation of variants, it has three shortcomings: 1. It requires the synthesis of complex saturation mutation libraries and the enrichment of cells through techniques such as fluorescence-activated cell sorting, which is complicated and costly; 2. The frequency of each variant in the cell pool is low, which is greatly affected by the background noise of second-generation sequencing, affecting the accuracy of the results; 3. The realization of high-throughput evaluation depends on the targeting of continuous genomic regions (about 100-150bp), while in reality, variants of unknown significance are often scattered across genes. Therefore, if SGE is used for functional verification after obtaining variants of unknown significance through sequencing, it cannot give full play to its high-throughput advantage and may instead lead to a waste of resources ([1] Sahu S, Galloux M, Southon E, et al. Saturationgenome editing-based clinical classification of BRCA2 variants[J]. Nature ,2025,638(8050):538-545. [2]Huang H, Hu C, Na J, et al. Functional evaluation and clinical classification of BRCA2 variants[J]. Nature , 2025,638(8050):528-537. [3]Findlay GM, Daza RM, Martin B, et al. Accurate classification of BRCA1 variants with saturation genome editing[J]. Nature , 2018, 562(7726):217-222.). Summary of the Invention
[0005] This invention addresses the aforementioned shortcomings of the prior art by providing a solution. BRCA1 / 2 A method for functional validation of variants of unknown significance in genes. This method overcomes the problems of low throughput, high cost, and incomplete coverage in existing techniques for functional validation of scattered variants of unknown significance in the BRCA1 / 2 genes. It is flexible, precise, and suitable for clinical laboratories.
[0006] A sort of BRCA1 / 2 A method for functional validation of variants of unknown significance includes the following steps: S1, for each variant of unknown significance to be tested, design one gRNA and two HDR templates for single-point editing, one of the two HDR templates containing the variant of unknown significance sequence and the other containing the control variant sequence; S2, for each variant of unknown significance to be tested or the control variant, co-transfer the gRNA, HDR template and Cas9 enzyme into haploid host cells for gene editing; S3, dynamic abundance detection, after culturing the gene-edited cells, the abundance of cells carrying the variant of unknown significance to be tested is detected; S4, the functional score of the variant of unknown significance is calculated based on the results of the dynamic abundance detection.
[0007] Preferably, the haploid host cell is a haploid HAP1 cell. More preferably, the haploid HAP1 cell contains... LIG4 Gene knockout.
[0008] Preferably, the control variant is a synonymous mutation that is as close as possible to the variant of unknown significance to be tested. More preferably, the control variant is a known benign or possibly benign variant within ±3 bp of the variant of unknown significance to be tested.
[0009] Preferably, gRNAs with cleavage sites within 10 nt of the variant of unknown significance, high targeting efficiency scores, and low off-target risk are selected.
[0010] Preferably, 0-3 blocking mutations located near the PAM site or adjacent to the 3' end of the gRNA are introduced into the HDR template, and all blocking mutations are confirmed to be benign or potentially benign variants; the non-complementary strand to the gRNA is selected as the HDR template strand; for missense variants, splice site variants, and small fragment insertion or deletion variants <6bp, the HDR template length is designed to be 80-100nt; for variants of unknown significance, the two HDR templates used are of the same length and contain the same blocking mutations.
[0011] Preferably, before functional validation, multiple computer prediction tools are used to filter out variants with low pathogenicity and unclear significance, thus narrowing down the validation targets. Among them, only missense variants, splice site variants, and small fragment insertion or deletion variants with a length ≤6bp that are predicted to be harmful, pathogenic, or affect function by at least one computer prediction tool are included in further functional validation trials.
[0012] Preferably, in step S3, after the gene-edited cells are cultured for 3 to 5 days, a portion of the cells are harvested for gene sequencing, and the remaining cells are cultured for at least 2 passage cycles before being harvested for gene sequencing. The sequence data are statistically analyzed, and the sequences carrying the target variant and the blocking mutation, or only carrying the target variant, are counted to determine the abundance of cells carrying variants of unknown significance.
[0013] The purpose of the first sample collection is to preliminarily assess the proportion of edited cells in the gene-edited cell population. According to the gene-editing reagent supplier's instructions, testing can usually be carried out 48-72 hours after transfection, at which point the editing effect has tended to stabilize. However, in practice, if samples are collected on the 3rd day after transfection, some samples may not have enough cells to support subsequent experiments. Therefore, considering both the reagent instructions and practical feasibility, we determined the first sample collection time to be 3-5 days after transfection, preferably the 4th day as a representative time point, which can ensure sufficient cell quantity while fully demonstrating the editing effect.
[0014] The second sampling aims to analyze the change in the proportion of edited cells over time, thereby assessing the functional differences between different genotypes under competitive growth conditions. Theoretically, extending the culture time helps to more accurately reflect the compositional changes of the cell population under steady-state conditions, but it also increases the experimental cycle and cost. If the interval between the two samplings is too short, it may be difficult to effectively distinguish the functional effects of different variants due to insignificant changes in cell proportions. Based on practical experience, we have found that an interval of 2-3 passage cycles (approximately 4-6 days) between the first and second samplings is sufficient to achieve good differentiation. Therefore, we determined the second sampling time to be 4-6 days after the first sampling; in this application, day 10 is preferably selected as an example.
[0015] Next-generation sequencing is preferred for gene sequencing.
[0016] More preferably, in step S4, the formula for calculating the functional score S of the unsigned variant function score is: S = (A vus, Dy ×A Control, Dx ) / (A vus, Dx ×A Control, Dy ), where A vus, Dx For the first gene sequencing to detect the abundance of cells carrying variants of unknown significance, A vus, Dy To detect the abundance of cells carrying variants of unknown significance in a second gene sequencing, A Control, Dx For the first gene sequencing to detect the abundance of cells carrying control variants, A Control, Dy For the second gene sequencing test to detect the abundance of cells carrying control variants, a low functional score (S) reflects the undetermined significance of the variant being tested. BRCA1 / 2 Loss of gene function.
[0017] In the case where the first sampling is on day 4 and the second sampling is on day 10, the formula for calculating the functional score S of the unsigned variant in step S4 is: S = (A vus, D10 ×A Control, D4 ) / (A vus, D4 ×A Control, D10 ), where A vus, D4 The abundance of cells carrying variants of unknown significance on day 4, A vus, D10 The abundance of cells carrying variants of unknown significance on day 10, A Control, D4 A represents the abundance of cells carrying the control variant on day 4. Control, D10 The abundance of cells carrying the control variant on day 10; a low functional score (S) reflects the undetermined significance of the variant. BRCA1 / 2 Loss of gene function. A function score S < 0.5 is considered a loss-of-function variant.
[0018] When we validated this method based on known clinical classifications of variants, we found that it had good discriminative power, with benign / probably benign functional scores close to or greater than 1, while pathogenic / probably pathogenic variants were much less than 1. Therefore, we defined S≥0.5 as a functional variant and S<0.5 as a variant that leads to loss of function.
[0019] Compared with the prior art, the beneficial effects of the present invention are:
[0020] This method allows for the precise introduction of single target variants, including single nucleotide variants and small insertion / deletion variants, at any selected genomic location, enabling customized functional validation of clinically unsigned variants.
[0021] By introducing only one variant in each experiment and combining optimized gRNA and template design, our method ensures high initial abundance of the target variant in the post-edited cell population, thereby significantly reducing background noise in next-generation sequencing and phenotypic analysis and improving the sensitivity and reliability of functional signal detection.
[0022] Because each experiment targets a single variant, this method simplifies the experimental process, eliminating the complex steps of constructing saturated mutant libraries and enriching cells. Compared to the SGE method, it significantly reduces the demand for synthesis, sequencing, and computational resources, and has lower requirements for experimental equipment, data processing capabilities, and personnel expertise, making it easier to implement and promote in routine clinical research laboratories or institutions with basic molecular biology capabilities.
[0023] This method provides a more flexible, accurate, and easy-to-operate functional verification scheme for interpreting BRCA1 / 2 variants of unknown significance, and is expected to solve the bottleneck of difficulty in acquiring and interpreting functional data of variants of unknown significance. Attached Figure Description
[0024] Figure 1 The diagram shown is a flowchart of the verification method for this function.
[0025] Figure 2 The results of ploidy analysis in Example 1 are shown below. Figure 2 In this context, A represents a haploid cell; Figure 2 In this context, B represents a mixture of haploid and diploid cells; Figure 2 C in the text represents a diploid cell.
[0026] Figure 3 The image shows HAP1 cells in Example 1. LIG4 Gene knockout results.
[0027] Figure 4 The figure shows the cell abundance carrying known taxonomic variants on days 4 and 10 in Example 2. Each mutation was detected on day 4 (D4) and day 10 (D10). Variants are displayed according to their known pathogenicity (benign / possibly benign, pathogenic / possibly pathogenic variants). The fill pattern and fill color of each bar represent time and taxonomy, respectively. In some cases, the cell abundance is very low, and the bar height is not obvious in the figure.
[0028] Figure 5 The figure shows the functional score differences of known categorical variations in Example 2.
[0029] Figure 6 The diagram shows the identification process for undetermined variants in Example 3.
[0030] Figure 7 The figure shows the functional scores of the variants of unknown significance in Example 3.
[0031] Figure 8 The figure shows a comparison between the functional verification results of the variant with unclear significance in Example 3 and the results of existing large-scale functional verification experiments. Detailed Implementation
[0032] Example 1: HAP1- LIG4 Construction and validation of knockout (KO) cell lines
[0033] This functional validation method relies on the haploid HAP1 cell line (human chronic myeloid leukemia cells). Haploid cells contain only one set of chromosomes, avoiding the masking of mutant functional effects by heterozygous states, making them an ideal model for functional studies. Previous studies have confirmed that BRCA1 / 2 genes are crucial for the survival of haploid cells, and loss of function of either gene can lead to cell death (Blomen VA, Majek P, Jae LT, et al. Gene essentiality and syntheticlethality in haploid human cells[J]. Science , 2015, 350(6264):1092-1096.). Furthermore, inhibition of the non-homologous end joining (NHEJ) pathway's key molecule, DNA ligase IV (by... LIG4 (Gene encoding) can significantly improve the efficiency of homology-directed repair (HDR) mediated precise editing in mammalian cells (Chu VT, Weber T, Wefers B, et al. Increasing the efficiency of homology-directed repair for CRISPR-Cas9-induced precise gene editing in mammalian cells[J]. Nat Biotechnol , 2015,33(5):543-548.). Therefore, to improve the efficiency of mutation introduction, we first used CRISPR technology to construct LIG4 Gene knockout monoclonal haploid HAP1 cell line (HAP1- LIG4 KO).
[0034] To enrich haploid HAP1 cells, cells from T75 culture flasks with a confluence of 60%-70% were digested. The digested cell suspension was resuspended in complete medium containing 7 µg / mL Hoechst 34580 (Selleck) and incubated at 37°C in the dark for 45 minutes. After incubation, cells were collected by centrifugation and resuspended in phosphate-buffered saline containing 1% fetal bovine serum, Hoechst 34580, and propidium iodide (MedChemExpress). Haploid cells were sorted using a MoFlo Astrios EQ (Beckman) flow cytometer with a sample pressure of 1-4 (speed 1000-1500 cells / second) and excitation sources of 405 nm and 561 nm lasers, respectively. The sorting strategy was as follows: excluding propidium iodide-positive (dead cell) signals, and selecting cell populations with relatively low Hoechst 34580 fluorescence intensity (haploid peak) from the two main cell populations under 405 nm excitation light. The sorted haploid cells were collected in complete culture medium and seeded into culture dishes of appropriate size according to the cell count.
[0035] A specific gRNA sequence (5'-GCATAATGTCACTACAGATC-3') was designed, a single-stranded DNA oligonucleotide was synthesized, annealed to form a double-stranded fragment, and then ligated into a linearized CRISPR / Cas9 vector (Jiman Gene) containing ZsGreen fluorescent protein and a puromycin resistance gene after enzyme digestion. The ligation product was transformed into E. coli Stbl3 competent cells, and positive clones were screened by colony PCR. The sequence correctness was verified by bidirectional sequencing to construct a targeted... LIG4 CRISPR / Cas9 plasmid of the gene. Endotoxin-free plasmid (<0.005 EU / μg) was extracted for subsequent experiments.
[0036] Healthy HAP1 (Horizon) cells were seeded in T25 culture flasks and transfected when confluence reached 50%-60% (approximately 16-20 hours). Lipofectamine 3000 (Thermo Fisher Scientific) transfection reagent was used: 5 μL of the reagent was diluted in 250 μL of Opti-MEM; 5 μg of plasmid DNA was diluted in 250 μL of Opti-MEM, and 10 μL of P3000 enhancer was added and mixed well. Equal volumes of the two solutions were mixed, and after standing at room temperature for 10-15 minutes, 500 μL of the complex was added to the cell culture system. Forty-eight hours after transfection, the culture medium was replaced with fresh medium, and ZsGreen expression was observed under 488 nm blue light excitation using an inverted fluorescence microscope. Approximately 24 hours after transfection, single ZsGreen-positive cells were sorted into 96-well plates containing complete culture medium using flow cytometry. After culturing monoclonal cells for 1-2 weeks, green fluorescent positive monoclonal cells were selected and cultured for another 1-2 weeks to identify ploidy and knockout efficiency.
[0037] Cell ploidy was analyzed using an Attune NxT (Thermo Fisher) flow cytometer. Cells were collected, stained with propidium iodide, and then analyzed. The maximum number of cells collected was set to 20,000, and the flow rate was 12.5 cells / second. Since haploid HAP1 cells can spontaneously transform from haploid to diploid, wild-type HAP1 cells passaged for more than 30 generations were used as a control. Figure 2 (C) By analyzing the scatter plot of cell number versus propidium iodide fluorescence intensity, the haploid cell peaks (G0 / G1 phase and G2 / M phase) were determined by comparing them with control cells. Haploid cell lines should have only two main peaks (corresponding to the G0 / G1 and G2 / M phases of the cell cycle, respectively), and the fluorescence intensity of the two main peaks should be half that of the two main peaks in the corresponding high-generation cell lines. Figure 2 A in the text. Cells that are both haploid and diploid show three cellular peaks (A). Figure 2 (B) The FlowJo software was used to analyze and visualize the ploidy analysis experimental data.
[0038] Genomic DNA was extracted from haploid cell clones using the EZNA Tissue DNA Extraction Kit (Omega Bio-tek). Specific PCR primers (approximately ±600 bp flanking the editing site) were designed (forward primer: 5'-TTTTGCAGATTTGTGTTCAACTTTAGAACG-3'; reverse primer: 5'-TTGTTTTACCATAATTCCCTCTTCTCTTTT-3'). The PCR products were amplified and Sanger sequenced to validate the results in HAP1 cells. LIG4The effect of gene knockout. The results showed that... LIG4 A 26bp frameshift deletion mutation was generated in the gene, suggesting... LIG4 Gene knockout successful ( Figure 3 ).
[0039] Example 2: Validation of Known Clinical Classification Variations
[0040] To verify the effectiveness of this method, we selected 11 known taxonomic variants (data from the ClinVar database, accessed on August 7, 2023) reviewed by an expert panel accredited by the Clinical Genome Resource Center (ClinGen) for testing: 5 BRCA1 Variants (2 benign / probably benign variants: c.3418A>G, c.1233T>G; 3 pathogenic / probably pathogenic variants: c.390C>A, c.5089T>C, c.5513T>A) and 6 BRCA2 Variants (3 benign / probably benign variants: c.5552T>G, c.8182G>A, c.2803G>A; 3 pathogenic / probably pathogenic variants: c.8486A>G, c.5966C>G, c.5946del). Testing methods are as follows: Figure 1 As shown.
[0041] Based on the aforementioned design principles, we used the Alt-R™ HDR design tool (Integrated DNA Technologies) to assist in the design of gRNA and HDR template oligonucleotides, summarized in Table 1. The crRNA was synthesized by Integrated DNA Technologies, and the HDR template oligonucleotides were synthesized by Sangon Biotech (Shanghai) Co., Ltd.
[0042] Table 1. crRNAs and HDR template oligonucleotides used for validation of known clinical classification variants (VUS template and control template in the table).
[0043]
[0044] 200 μM crRNA was mixed with an equal volume of 200 μM tracrRNA (Integrated DNA Technologies), resulting in a final concentration of 100 μM for each. The mixture was heated at 95 °C for 5 minutes and then cooled to room temperature to form gRNA. Subsequently, 62 μM Cas9 enzyme (Integrated DNA Technologies) was added, and the mixture was incubated at room temperature for 20 minutes to form a ribonucleoprotein (RNP) complex.
[0045] HAP1- LIG4 KO cells were seeded at a high density 16-20 hours prior to nuclear transfection. Cells were digested, collected, counted, and adjusted to a cell number of 2 × 10⁶. 5 Cells were washed once with 1× phosphate-buffered saline and centrifuged to remove the supernatant. Cells were resuspended in 20 μL of SE nuclear transfection reagent (Lonza), and pre-assembled RNP complex, 100 pmol HDR template oligonucleotide, and 100 pmol electroporation enhancer (Integrated DNA Technologies) were added. Nuclear transfection was performed using the 4D-Nucleofector™ system (Lonza) with program EO-100 selected. Immediately after transfection, cells were transferred to 96-well plates containing 150 μL of pre-warmed medium and incubated.
[0046] On day 4 of cell culture, approximately 3 / 4 of the cells from each sample were harvested; the remaining 1 / 4 of cells were cultured until day 10, at which point all cells were harvested. Genomic DNA was extracted using the EZNA® Tissue DNA Kit (Omega Bio-tek). For each target variant, PCR primers were designed upstream and downstream, amplifying fragments of 200-280 bp in length, ensuring the target variant was not located within the first or last 20 bp region of the amplified product. PCR amplification was performed using a high-fidelity KAPA HiFi HotStart ReadyMix (Roche). The quality of the PCR products was assessed by agarose gel electrophoresis. DNA concentration was precisely quantified using Qubit; samples with a concentration ≥0.5 μg were used for library construction. DNA samples underwent end repair, A-tailing, and sequencing adapter ligation at both ends of the fragments, followed by purification (without intermediate PCR amplification) to construct PCR-free libraries. After quality control, the libraries were sequenced at 150 bp (PE150) ends on an Illumina platform according to their effective concentration. Samples with a Q30 score ≥80% were selected for subsequent analysis. The R1 and R2 sequencing reads were assembled using FLASH software. The assembled sequences were then aligned to a reference sequence using BLAST software, followed by multiple sequence alignment using MAFFT software. All sequence data were analyzed, and sequences carrying the target variant, blocking mutations, or only the target variant were counted.
[0047] The results showed that cell abundance on day 4 was not directly correlated with whether the mutation was pathogenic / potentially pathogenic. Figure 4 This is related to the differences in CRISPR editing efficiency at different gene locations. Therefore, the abundance at a single time point cannot directly determine the functional impact of the variant. Compared to day 4, the relative abundance of cells carrying pathogenic / potentially pathogenic variants decreased significantly on day 10; while the abundance of cells carrying benign / potentially benign variants changed very little or increased slightly. Figure 4 This phenomenon stems from the fact that, under population selection pressure, cells carrying pathogenic / potentially pathogenic variants (loss of function) are at a disadvantage in competitive growth, while benign / potentially benign variants (normal function) do not affect cell fitness. By comparing the abundance changes of the tested variant cells and control variant cells, the functional score of each variant is calculated according to the following formula: S=(A vus, D10 ×A Control, D4 ) / (A vus, D4 ×A Control, D10 ).
[0048] The results showed that the functional scores of benign / potentially benign variants were approximately equal to or greater than 1, while the functional scores of pathogenic / potentially pathogenic variants were much less than 1. Figure 5 Wilcoxon rank-sum test P=0.004329, effect size δ=-1), indicating that the method can effectively distinguish between pathogenic / potentially pathogenic and benign / potentially benign variants.
[0049] Example 3: Application of computer-based pre-screening to breast cancer patients BRCA1 / 2 Interpretation and Verification of Variations with Unclear Significance
[0050] In Example 3, we applied this method to 63 breast cancer patients identified in a real-world population. BRCA1 / 2 Interpretation of variants of unknown significance. We first used ten different computer prediction tools to exclude some variants of unknown significance with low pathogenicity. Seven tools (AlphaMissense, REVEL, EVE, gMVP, MetaRNN, DeepSAV, and MutScore) were used for missense variant prediction, and three tools (dbscSNV, MaxEntScan, and SpliceAI) were used for splice site variant prediction. Most of the selected tools are based on machine learning algorithms, while the rest are ensemble methods or classic methods widely used in clinical practice. The prediction results are summarized in Table 2. In short, 28 (44.4%) variants were identified. BRCA1 / 2 Variations of unknown significance were considered to have an extremely low probability of being pathogenic or possibly pathogenic and were excluded from functional validation.
[0051] Table 2. Computer prediction results for variants of unknown significance.
[0052]
[0053] * P: pathogenic, VUS: variant of uncertain significance, B: benign; ** D: deleterious, T: tolerable; *** A: affected, U: unaffected; ****Missense / splicing site variants that meet the following criteria are defined as pred_P: predicted as "deleterious (D)", "pathogenic (P)" or "affected (A)" by at least one algorithmic tool; missense / splicing site variants that are not predicted as "deleterious (D)", "pathogenic (P)" or "affected (A)" by any tool are defined as pred_B.
[0054] We screened 33 patients from a high-hereditary-risk breast cancer cohort. BRCA1 / 2 Further functional validation was performed on variants of unknown significance (see identification process). Figure 6 The remaining specific technical details are the same as in Example 2. The results showed that 33 variants of unknown significance were significantly distinguished (…). Figure 7 Of these, 5 variants (2 of them) BRCA1 : c.4211T>G, c.100C>T; 3 BRCA2 The variants c.475+3A>G, c.476-3C>A, and c.7871A>C exhibited low functional scores (<0.5), and the abundance of cells carrying these variants in the cell pool decreased significantly over time. P <0.01), indicating that it leads to a significant loss of gene function. The above-mentioned loss-of-function variants account for all BRCA1 9.1% (2 / 22) of the variation of unknown significance and all BRCA2 Of the 7.3% (3 / 41) of variants of unknown significance, 3 were missense variants and 2 were splice site variants. No small insertion / deletion variants were identified as loss-of-function variants. Of all the missense variants identified as loss-of-function, at least 5 / 7 of the computer prediction tools predicted them as "pathogenic" or "harmful"; one loss-of-function splice site variant was predicted as "pathogenic" or "VUS" by 2 / 3 of the tools, and another was predicted as "pathogenic" or "affects function" by 3 / 3 of the tools (Table 2). BRCA1 c.100C>T is located in the RING structure domain. BRCA2 c.7871A>C is located in the helical domain, while the other three loss-of-function variants are not located in specific functional domains. Comparison with multiple large, independent functional studies revealed a high degree of consistency in the functional classification of all previously reported unsigned variants (100% consistency, 22 / 22). Figure 8 The results indicate that this method has good reproducibility, supporting its potential as a tool for assessing the function of variants of unknown significance.
Claims
1. A kind BRCA1 / 2 A method for functional verification of variants of unknown gene significance, characterized in that, The method comprises the following steps: S1, for each variant of unknown significance, design one gRNA and two HDR templates for single-point editing, one of the two HDR templates contains the variant of unknown significance sequence, and the other contains the control variant sequence; the control variant is selected as a synonymous mutation as close as possible to the variant of unknown significance; S2, for each to-be-tested unknown meaning variation or control variation, gRNA, HDR template and Cas9 enzyme are respectively co-transferred into haploid host cells for gene editing; the haploid host cells are haploid HAP1 cells, and the haploid HAP1 cells are modified to have a gene knockout at the target site of the to-be-tested unknown meaning variation or control variation LIG4 gene knockout; S3, dynamic abundance detection, after culturing the cells after gene editing, the abundance of cells carrying the variant of unknown significance is detected; the cells after gene editing are cultured for 3-5 days, part of the cells are collected for the first gene sequencing, the remaining cells are continuously cultured for at least 2 passages, and then the cells are collected for the second gene sequencing, the sequence data is counted, and the sequences carrying the target variation and the blocking mutation or only carrying the target variation are counted, i.e. the abundance of cells carrying the variant of unknown significance; S4. Based on the dynamic abundance detection results, the functional score of variants of unknown significance is calculated. The formula for calculating the functional score S of variants of unknown significance is: S = (A vus, Dy ×A Control, Dx ) / (A vus, Dx ×A Control, Dy ), wherein A vus, Dx is the first gene sequencing to detect the abundance of cells carrying the variant of unknown significance to be tested, A vus, Dy To detect the abundance of cells carrying the variant of interest for the second genetic sequencing, A Control, Dx To detect the abundance of cells carrying control variants for the first time, A Control, Dy To detect the abundance of cells carrying control variants for the second genetic sequencing, Low fraction S reflects that the variant under test is not causally associated with the trait of interest BRCA1 / 2 loss of gene function.
2. The method of claim 1 BRCA1 / 2 Method for functional verification of genetic variants of unknown significance, characterized in that, The control variant is selected as a known benign or possibly benign variant within ±3bp of the variant of unknown significance.
3. The method of claim 1 BRCA1 / 2 Method for functional verification of genetic variants of unknown significance, characterized in that, The gRNA is selected to be within 10nt of the variant of unknown significance, with high target efficiency score and low off-target risk.
4. The method of claim 1 BRCA1 / 2 Method for functional verification of genetic variants of unknown significance, characterized in that, 0-3 blocking mutations are introduced in the HDR template, which are located near the PAM site or the adjacent region of the 3' end of the gRNA, and all the blocking mutations are confirmed as benign or possibly benign variants; The non-complementary strand of the gRNA is selected as the HDR template strand; For missense variants, splice site variants, small fragment insertion or deletion variants with a length of ≤6bp, the length of the HDR template is designed to be 80-100nt. The two HDR templates used for the same variant of unknown significance have the same length and contain the same blocking mutations.
5. The method of claim 1 BRCA1 / 2 Method for functional verification of genetic variants of unknown significance, characterized in that, Before functional verification, low pathogenicity variants of unknown significance are filtered by using various computer prediction tools to reduce the verification target, wherein only missense variants, splice site variants, small fragment insertion or deletion variants with a length of ≤6bp predicted as harmful, pathogenic or affecting function by at least one computer prediction tool are included in the further functional verification test.
6. The method of claim 1 BRCA1 / 2 Method for functional verification of genetic variants of unknown significance, characterized in that, In the case that the first sampling is at day 4 and the second sampling is at day 10, the formula for calculating the functional score S of the significance-unknown variant function in step S4 is: S = (A vus, D10 × A Control, D4 ) / (A vus, D4 × A Control, D10 ), wherein A vus, D4 is the cell abundance carrying the to-be-tested significance-unknown variant at day 4, A vus, D10 is the cell abundance carrying the to-be-tested significance-unknown variant at day 10, A Control, D4 is the cell abundance carrying the control variant at day 4, and A Control, D10 is the cell abundance carrying the control variant at day 10. A low functional score S indicates that the to-be-tested significance-unknown variant leads to loss of function of the gene. BRCA1 / 2 7. The method of claim 1 BRCA1 / 2 Method for functional verification of genetic variants of unknown significance, characterized in that, In step S3, the gene sequencing uses second-generation sequencing.
Citation Information
Patent Citations
Products and methods for annotating gene function using locally haploid, human non-cancer cells
WO2023034704A1