Method, apparatus and program product for screening of potential drug target regions

CN122676902APending Publication Date: 2026-09-01NORTHEAST NORMAL UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610880257.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0004]然而,面对大规模测序发现的海量编码变异,仅依赖实验手段进行功能表征面临着成本高昂、周期漫长的技术瓶颈,难以实现系统化筛选

Benefits of technology

1.针对现有变异注释工具(如AlphaMissense, V2P)仅能提供通用致病性评分、缺乏特定疾病背景解释的痛点,本发明提出将疾病相关编码变异映射至蛋白质结构锚定的热点区域的计算框架COMPASS,该计算框架整合疾病相关编码变异与实验测定及人工智能预测的蛋白质结构,评估变异对蛋白结构与功能的影响,从而促进遗传驱动的靶点筛选与药物设计。这种从“通用评分”到“疾病特异性机制”的跨越,填补了遗传学发现与精准临床应用之间的鸿沟,为理解特定疾病的发病机理提供结构证据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122676902A_ABST
    Figure CN122676902A_ABST
Patent Text Reader

Abstract

This application relates to the field of intelligent healthcare, specifically to a method, device, and program product for screening potential drug target regions. It includes: acquiring coding variant data of the target gene; for cases where multiple nucleotide mutations exist at a single amino acid position, screening candidate variants based on association statistics to determine the dominant amino acid sequence and perform functional annotation of the amino acid sequence; structural hotspot localization: constructing wild-type and mutant protein structures of the target gene based on experimental assays or artificial intelligence prediction tools. By defining local spatial neighborhoods centered on amino acid residues, traversing the protein structure, and screening based on aggregated association signals, structural hotspot regions significantly associated with diseases are obtained; functional hotspot annotation: integrating multidimensional biological database resources and using large language models for automated literature mining, performing functional annotation and druggability assessment of the structural hotspot regions, thereby obtaining potential therapeutic targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent healthcare, specifically to a method, device, program product, and computer-readable storage medium for screening potential drug target regions. Background Technology

[0002] Despite significant progress in drug development in recent years, the overall success rate of clinical trials remains low, at only 9.6%, highlighting the urgent need to improve the accuracy of target identification and validation efficiency in the early stages of development. Studies have shown that drug targets supported by human genetic evidence have a 2.6 times higher success rate in clinical trials compared to targets lacking such evidence. For example, the JAK2 p.V617F mutation has been identified as a driver of myeloproliferative neoplasms, suggesting that abnormal activation of the JAK-STAT pathway is a core pathogenesis, directly driving the development and approval of inhibitors such as Ruxolitinib. This confirms the crucial role of genetic insights in target discovery.

[0003] The development of high-throughput sequencing technology and statistical genetics has made it possible to systematically study disease-related variants, providing a foundation for target selection and drug discovery. Large-scale population cohort studies, exemplified by the UK Biobank and All of Us, support genetic research on complex diseases by combining high-quality whole-genome sequencing data with in-depth phenotypic analysis. In genetic association analysis, although analyses of common variants are widely used, these signals typically have weak effects and are mostly located in non-coding regions, making it difficult to elucidate their underlying molecular mechanisms. In contrast, rare variant signals usually have larger effect sizes and are more directly linked to gene function and biological pathways, thus facilitating mechanism interpretation and functional validation.

[0004] However, faced with the massive amounts of coding variants discovered through large-scale sequencing, relying solely on experimental methods for functional characterization faces significant technical bottlenecks, including high costs and lengthy processing times, hindering systematic screening. While computational methods such as AlphaMissense, which combine predicted protein structure to assess mutation impact, have emerged, these annotation tools primarily provide general pathogenicity scores, lacking structural interpretations specific to disease contexts and failing to offer insights into disease biological mechanisms. Currently, a systematic computational framework is lacking to assess the impact of disease-related coding variants on protein structure and function to guide target screening and drug design. Summary of the Invention

[0005] To address the aforementioned problems, this invention proposes COMPASS, a computational framework that maps coding variations to hotspot regions anchored in protein structures. This framework combines disease-related coding variations with experimentally determined or AI-predicted protein structures to assess the impact of mutations on structure and function, thereby facilitating gene-driven target screening and drug design. Specifically, it includes: Obtain gene variant data that are statistically associated with the target disease; The dominant amino acid sequence was obtained by screening the gene variation data; Obtain the protein structure predicted from the dominant amino acid sequence or obtain the wild-type protein structure of the target gene; Traverse each amino acid residue in the predicted protein structure or wild-type protein structure, construct a local spatial neighborhood (patch) centered on the residue, identify variants falling into each local spatial neighborhood, and calculate the aggregated disease association statistic of the variant set within the local spatial neighborhood; The aggregated disease association statistics are sorted, and the top K local spatial neighborhoods are selected as potential drug target regions, where K is a natural number greater than or equal to 1.

[0006] Optionally, the process of obtaining the target gene is as follows: extract all coding variation data of the associated gene under a specific disease background based on whole genome sequencing data, map the coding variation to the amino acid site corresponding to the reference transcript, all amino acid sites containing the variation constitute a set to be screened, and screen the set to be screened to obtain the dominant amino acid sequence. Optionally, the screening includes identification and classification, classifying the variant data of the target gene into single variants and multiple nucleotide variants, screening the multiple nucleotide variants to obtain a first variant, combining the first variant with the single variant to obtain a dominant variant, and constructing a dominant amino acid sequence through the dominant variant. Optionally, the screening process for each of the multiple nucleotide variants is as follows: S1. Keeping the variations at other positions in the gene unchanged, remove a single candidate variation from the candidate variation set at the current position, and calculate the overall disease association statistic of the remaining variation set after removing the variation. Repeat this step until all candidate variations at the current position are traversed to obtain a set of statistics after removal. S2. Compare and sort the set of statistical data after removal, select the variant with the largest disease association statistic in the set of remaining variants after removal, and obtain the dominant variant at the current position. S3. Repeat steps S1-S2 to obtain a variant or set of variants with multiple nucleotide multiple variation positions. Optionally, the positions of multiple mutations in all multiple nucleotide variants are screened to obtain a first variant or a set of first variants. The first variant or set of first variants is combined with a single variant to obtain a dominant variant. A dominant amino acid sequence is constructed through the dominant variant. Optionally, the step of selecting the variant with the largest disease association statistic in the remaining variant set after removal is achieved by sequentially removing individual candidate variants at the current position, calculating the disease association statistic of the remaining variant set for each variant, and directly comparing the calculated remaining set statistics to select the variant corresponding to the largest statistic.

[0007] Optionally, instead of screening multiple nucleotide variants to obtain a first variant, a second screening of multiple nucleotide variants is performed to obtain a second variant or a set of second variants. The second variant is combined with a single variant to obtain a dominant variant, and a dominant amino acid sequence is constructed through the dominant variant. The second screening process is as follows: the set of multiple variant sites of multiple nucleotides is defined as an initial current set; in the initial current set, a removal operation is performed on each candidate variant of the screening site, and the overall disease association statistic of the remaining variant set after removing the variant is calculated; the variant with the largest overall disease association statistic of the remaining variant set after removal is selected, and the variant is marked as the dominant variant at its amino acid position. Remove other candidate variants from the initial current set that have been labeled with the dominant variant amino acid position, update the current set to obtain an updated set, and repeat the above steps based on the updated set until the amino acid position of each multiple nucleotide multiple variant site is labeled with a dominant variant, to obtain a second variant or a second variant set.

[0008] Optionally, instead of screening multiple nucleotide variants to obtain a first variant, a third screening of multiple nucleotide variants is performed to obtain a third variant or a set of third variants. The third variant is combined with a single variant to obtain a dominant variant, and a dominant amino acid sequence is constructed through the dominant variant. The third screening process is as follows: for each multiple nucleotide variant site, one variant is selected from the candidate variant set, and together with the variant at the single variant site, a candidate variant combination is generated through full permutation and combination. For each generated variant combination, the overall association statistic with the target disease is calculated, the overall association statistic is compared and ranked, and the variant combination with the smallest overall association statistic is selected to obtain the dominant amino acid sequence.

[0009] Optionally, the local spatial neighborhood is constructed by forming a spherical region with a predetermined length radius centered on the Cα atom of an amino acid residue.

[0010] Optionally, the preset length radius is set according to the length of the target protein sequence.

[0011] Optionally, the method further includes functional annotation. Potential drug target regions are extracted from biological databases to extract relevant biological pathways, drug targets, and functional annotation information to construct basic structured data and provide detailed biological background annotations. Based on this, the potential target regions, structured data, and biological background annotations are input into a large language model to obtain a biological functional summary covering structural hotspots, an interpretation of pathogenic molecular mechanisms in specific disease contexts, and a druggability assessment. Automated literature mining technology using the large language model is used to parse unstructured textual evidence from massive biomedical literature, thereby outputting a biological functional summary covering structural hotspots, an interpretation of pathogenic molecular mechanisms in specific disease contexts, and a druggability assessment.

[0012] Optionally, the reasoning process of the large language model is as follows: basic structured data is obtained by extracting biological pathways, drug targets and functional annotation information based on biological databases, and unstructured text data is extracted based on biological literature; the structured data and unstructured text data are input into the large language model for deep fusion and logical reasoning, and output biological functional summaries of structural hotspots, interpretation of pathogenic molecular mechanisms in specific disease backgrounds, and drugability assessment.

[0013] The purpose of this invention is to provide a computer program product comprising a computer program or instructions, wherein the computer program or instructions are executed by a processor to implement the above-described method for screening potential drug target regions.

[0014] The purpose of this invention is to provide a computer device comprising a memory, a processor, and a computer program or instructions stored in the memory, wherein the computer program or instructions are executed by the processor to implement the above-described method for screening potential drug target regions.

[0015] The purpose of this invention is to provide a computer-readable storage medium having a computer program or instructions stored thereon, which are executed by a processor to implement the above-described method for screening potential drug target regions.

[0016] Advantages of this invention: 1. Addressing the shortcomings of existing variant annotation tools (such as AlphaMissense, V2P), which only provide general pathogenicity scores and lack specific disease-specific contextual explanations, this invention proposes COMPASS, a computational framework that maps disease-related coding variants to hotspot regions anchored in protein structures. This framework integrates disease-related coding variants with experimentally determined and AI-predicted protein structures to assess the impact of variants on protein structure and function, thereby facilitating genetically driven target screening and drug design. This leap from "general scoring" to "disease-specific mechanisms" bridges the gap between genetic discovery and precision clinical applications, providing structural evidence for understanding the pathogenesis of specific diseases.

[0017] 2. To address the genetic heterogeneity issue of multiple nucleotide variations potentially existing on the same amino acid residue, this invention proposes an innovative dominant amino acid sequence screening algorithm (including leave-one-out method and optimal subset method). This method can accurately identify the dominant amino acid variation that contributes the most to disease-associated signals. This process effectively improves the statistical detection power of rare variations by integrating scattered nucleotide variation signals into a unified amino acid-level signal, and eliminates the granularity difference between genetic data (nucleotide dimension) and structural analysis (amino acid dimension), providing the most pathogenic and representative input data for subsequent structural analysis.

[0018] 3. To address the challenge of precise localization in drug target discovery based on genetic evidence, this invention proposes a fragment scanning technique based on the three-dimensional structure of proteins. By constructing local spatial neighborhoods, this invention can aggregate association signals of discrete variations in three-dimensional space. Compared to traditional one-dimensional sequence analysis, this method can more sensitively identify three-dimensional structural hotspots, which often correspond to catalytic centers, allosteric regulatory sites, or drug-binding pockets in proteins. By ranking and prioritizing the association statistics of hotspot regions, this invention can directly identify potential target regions with high drugability based on human genetic evidence, thereby significantly improving the success rate of clinical drug development.

[0019] 4. Addressing the challenge of elucidating the mechanisms of potential drug targets, this invention integrates a comprehensive annotation module and creatively introduces a generative large language model. This module not only integrates multidimensional biological databases (such as KEGG pathways and DrugBank drug targets), but also processes heterogeneous data and automatically mines literature, generating coherent functional summaries, pathogenic mechanism interpretations, and therapeutic hypotheses within specific disease contexts. This ability to transform static annotations into actionable biomedical insights significantly deepens researchers' understanding of the mechanisms underlying sequence variations, protein structures, and potential target regions, providing intelligent decision support for precision medicine and drug design. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the screening method for potential drug target regions provided in an embodiment of the present invention; Figure 2 The COMPASS process provided for embodiments of the present invention.

[0022] Figure 3 This is a schematic diagram of the dominant amino acid sequence screening strategy provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the one-step leave-one selection method provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the stepwise leave-one-out selection method provided in an embodiment of the present invention; Figure 6 The correlation between JAK2 structural hotspots and treatment in polycythemia vera provided by embodiments of the present invention; Figure 7 The rare and disruptive missense variants provided in this embodiment of the invention are JAK2 structural hotspots in leukemia, non-Hodgkin's lymphoma, and other lymphocytic and histiocytic carcinomas. Detailed Implementation

[0023] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0024] Figure 1 A method for screening potential drug target regions provided in this embodiment of the invention includes: S101. Obtain gene variation data that are statistically associated with the target disease; In one embodiment, the process of acquiring the variant data is as follows: based on whole-genome sequencing data, acquire all coding variant data of genes statistically associated with the target disease, and map the coding variants to the amino acid positions corresponding to the reference transcripts; identify all amino acid sites carrying nucleotide variants to construct a set to be screened, and screen the set to be screened to obtain the dominant amino acid sequence.

[0025] In one specific embodiment, the preprocessing step for mapping disease-related coding variations to amino acid sequences specifically includes: S1: Based on whole genome sequencing (WGS) data, obtain all coding variations of the target gene and map the variations to the corresponding amino acid positions of the reference transcript; S2: Traverse all mapped amino acid positions, identify amino acid positions with multiple nucleotide variations, and mark these positions as the set to be screened; for amino acid positions containing only a single variation, directly retain that variation.

[0026] S102. The dominant amino acid sequence is obtained by screening the gene variation data; In one embodiment, the screening includes identification and classification, identifying and classifying variant data of the target gene into single variants and multiple nucleotide variants, screening multiple nucleotide variants to obtain a first variant, combining the first variant with a single variant to obtain a dominant variant, and constructing a dominant amino acid sequence through the dominant variant. Optionally, the screening process for each of the multiple nucleotide variants is as follows: S1. Keeping the variations at other positions in the gene unchanged, remove a single candidate variation from the candidate variation set at the current position, and calculate the overall disease association statistic of the remaining variation set after removing the variation. Repeat this step until all candidate variations at the current position are traversed to obtain a set of statistics after removal. S2. Compare and sort the set of statistical data after removal, select the variant with the largest disease association statistic in the set of remaining variants after removal, and obtain the dominant variant at the current position. S3. Repeat steps S1-S2 to obtain a variant or set of variants with multiple nucleotide multiple variation positions. Optionally, the multiple variant positions of all multiple nucleotide variants are screened to obtain a first variant or a set of first variants. The first variant or set of first variants is combined with a single variant to obtain a dominant variant. A dominant amino acid sequence is constructed through the dominant variant.

[0027] In one embodiment, selecting the variant with the largest disease association statistic in the remaining variant set after removal is achieved by sequentially removing individual candidate variants at the current position, calculating the disease association statistic of the remaining variant set for each variant, and directly comparing the calculated remaining set statistics to select the variant corresponding to the largest statistic.

[0028] In another embodiment, instead of screening multiple nucleotide variants to obtain a first variant, a second screening of multiple nucleotide variants is performed to obtain a second variant or a set of second variants. The second variant is combined with a single variant to obtain a dominant variant, and a dominant amino acid sequence is constructed through the dominant variant. The second screening process is as follows: the set of multiple variant sites of multiple nucleotides is defined as an initial current set; in the initial current set, a removal operation is performed on each candidate variant of the site to be screened, and the overall disease association statistic of the remaining variant set after removing the variant is calculated; the variant with the largest overall disease association statistic of the remaining variant set after removal is selected, and the variant is marked as the dominant variant at its amino acid position. Remove other candidate variants from the initial current set that have been labeled with the dominant variant amino acid position, update the current set to obtain an updated set, and repeat the above steps based on the updated set until the amino acid position of each multiple nucleotide multiple variant site is labeled with a dominant variant, to obtain a second variant or a second variant set.

[0029] In another embodiment, instead of screening multiple nucleotide variants to obtain a first variant, a third screening is performed on multiple nucleotide variants to obtain a third variant or a set of third variants. The third variant is combined with a single variant to obtain a dominant variant, and a dominant amino acid sequence is constructed through the dominant variant. The third screening process is as follows: for each multiple nucleotide variant site, one variant is selected from the candidate variant set, and together with the variant at the single variant site, a candidate variant combination is generated through full permutation and combination. For each generated variant combination, the overall association statistic with the target disease is calculated, the overall association statistic is compared and ranked, and the variant combination with the smallest overall association statistic is selected to obtain the dominant amino acid sequence.

[0030] In another embodiment, variant data of the target gene are identified and classified into single variants and multiple nucleotide variants. Based on the number of multiple nucleotide variant positions and / or evaluation requirements, when the number of multiple nucleotide variant positions is small, multiple nucleotide variant screening is performed through a third screening; when the number of multiple nucleotide variants is large and / or the evaluation requirement is rapid evaluation, multiple nucleotide variant screening is performed through a first screening; when the number of multiple nucleotide variants is large and / or the evaluation requirement is high-precision evaluation, multiple nucleotide variant screening is performed through a second screening.

[0031] Among them, the limited number of positions of multiple nucleotide variations means a smaller search space, with only a limited number of possible permutations and combinations.

[0032] In one specific embodiment, the filtering method includes: Implementation Method 1: Best Subset Selection.

[0033] This embodiment uses enumeration and combination, which is suitable for situations where there are few multiple mutation sites (small search space). S4: For each multiple mutation site, select one mutation from its candidate mutation set, and together with the mutation of the single mutation site, generate all possible candidate amino acid sequence combinations through full permutation and combination; S5: For each generated candidate amino acid sequence combination, calculate its association statistic with the disease; S6: Compare the correlation statistics of all combinations, select the combination with the smallest correlation statistics (i.e. the strongest correlation), and determine the corresponding sequence as the final dominant amino acid sequence.

[0034] Implementation Method 2: Leave-One-Out One-Step Selection.

[0035] This strategy is suitable for scenarios where computing resources are limited or where rapid evaluation is required.

[0036] S4: For each tagged multiple variant site, perform the following operations: Keep the variants at other locations in the gene unchanged, remove individual candidate variants one by one from the candidate variant set at that location using the leave-one-out method, and calculate the overall disease association statistic of the remaining variant set after removing the variant; repeat this step until all candidate variants at that location have been traversed to obtain a set of statistics after removal. S5: Compare the post-removal statistics obtained in step S4. If a variant contributes the most to the disease association signal, removing that variant will lead to the most significant decrease in the association significance of the remaining set. Therefore, select the variant that results in the largest association statistic of the remaining set after removal and identify it as the dominant variant at that position; S6: Combine the dominant variants selected from all multiple variant sites with the variants from single variant sites to construct the dominant amino acid sequence.

[0037] Implementation Method 3: Leave-One-Out Stepwise Selection This strategy, through iterative optimization, is suitable for processing complex genetic signals and achieves higher accuracy.

[0038] S4: The set of multiple variant sites to be marked is defined as the initial current set; S5: Enter the iteration loop, locking in an optimal mutation in each round: In the initial current set, perform a "removal" operation for each candidate variant of the site to be screened, and calculate the association statistics of the remaining set after removing the variant; The variant that causes the greatest loss of associated signal after removal is selected and identified as the dominant variant at its amino acid position. Remove all other candidate variants (i.e. unselected variants) at that amino acid position from the current set, and update the current set for the next iteration; Step S6: Repeat the above iterative steps until a dominant mutation is locked for each amino acid position with multiple mutations; finally, combine all the mutations in the set with the single mutation site to construct the dominant amino acid sequence.

[0039] S103. Obtain the protein structure predicted by the dominant amino acid sequence or obtain the wild-type protein structure of the target gene. In one embodiment, the predicted protein structure or wild-type protein structure is obtained through computer-aided prediction, including but not limited to any one or more of the following: AlphaFold 2, AlphaFold 3, DI-TASSER model, and homology modeling model. In another embodiment, the wild-type protein structure is obtained through any one or more of the following techniques: X-ray crystallography, nuclear magnetic resonance spectroscopy, and cryo-electron microscopy.

[0040] S104. Traverse each amino acid residue in the predicted protein structure or wild-type protein structure, construct a local spatial neighborhood centered on the residue, identify variants falling into each local spatial neighborhood, and calculate the aggregated disease association statistic of the variant set within the local spatial neighborhood. In one embodiment, the local spatial neighborhood is constructed by forming a spherical region with a predetermined length radius centered on the Cα atom of an amino acid residue; Optionally, the preset length radius is set according to the length of the target protein sequence.

[0041] In one embodiment, each local spatial region also includes a set of output residues.

[0042] S105. Sort the aggregated disease association statistics and select the top K local spatial neighborhoods as potential drug target regions, where K is a natural number greater than or equal to 1.

[0043] In one embodiment, the aggregated disease association statistics are sorted, and the first local spatial neighborhood is selected as a potential drug target region.

[0044] In one embodiment, the method further includes functional annotation, in which the potential drug target region is input into a large language model for analysis to obtain functional annotation of the potential drug target. The functional annotation includes a biological functional summary of structural hotspots, interpretation of pathogenic molecular mechanisms in a specific disease context, and druggability assessment. Optionally, the large language model training and inference involves: extracting basic structured data from biological databases by extracting biological pathways, drug targets, and functional annotation information; and extracting unstructured text data from biological literature. The structured and unstructured text data are then input into the large language model for deep fusion and logical inference, outputting a biological functional summary of structural hotspots, an interpretation of pathogenic molecular mechanisms in specific disease contexts, and a druggability assessment. The large language model can be any one or more of Deepseek, GPT, QWen3, and Gemini.

[0045] In summary, in one embodiment, the leading amino acid sequence screening involves: screening common and rare variants associated with the target disease based on whole-genome sequencing data to obtain coding variant data of the target gene; for cases where multiple nucleotide mutations exist at a single amino acid position, candidate variants are screened based on association statistics to determine the leading amino acid sequence, and the amino acid sequence is functionally annotated. Structural hotspot localization: Based on experimental determination or AI tools such as AlphaFold3, wild-type and mutant protein structures of target genes are constructed. By defining local spatial neighborhoods centered on amino acid residues, the protein structure is traversed, and structural hotspot regions significantly associated with the disease are screened based on aggregation correlation signals. Functional hotspot annotation: By integrating multidimensional biological database resources and using large-scale language models for automated literature mining, druggability assessment and functional annotation are performed on the structural hotspot regions to obtain potential therapeutic targets.

[0046] Based on large-scale validation using whole-genome sequencing data from approximately 500,000 participants in the UK Biobank, this invention not only successfully reproduced established drug targets in an analysis covering 83 cancer phenotypes, but also discovered novel structural hotspots with potential druggability, demonstrating the effectiveness and versatility of this method in identifying potential therapeutic targets.

[0047] In one specific embodiment, the technical solution of the present invention includes three modules. First, the COMPASS-Seq module maps disease-related coding variations to the amino acid sequence of a reference transcript and uses statistical strategies to screen for the most pathogenic leading amino acid sequence. Then, the COMPASS-Struct module performs structure-based patch scanning, identifying disease-significant local substructures (i.e., structural hotspots) in the protein by defining local spatial neighborhoods (patches). Furthermore, the present invention introduces the COMPASS-Anno annotation module, which integrates multidimensional biological databases and utilizes a generative large-scale language model to perform functional annotation and druggability assessment on the structural hotspots, generating disease-specific mechanistic interpretations and therapeutic hypotheses.

[0048] In one specific embodiment, for functional hotspots obtained from experimentally determined or AI-predicted protein structures, this invention proposes a potential target region screening method COMPASS-Struct based on structural fragment scanning, which specifically includes: S1: Obtain the wild-type protein structure of the target gene or the protein structure predicted based on the dominant amino acid sequence; S2: Traverse every amino acid residue in the target protein structure and construct a local spatial neighborhood centered on that residue; identify variants falling into each patch and calculate the aggregated disease association statistic of the variant set within that patch; S3: Sort the aggregated disease association statistics and select the first patch with the most significant association statistics as the potential target area (structural hotspot).

[0049] In one specific embodiment, the COMPASS-Anno function is commented as follows: To address the mechanism interpretation and drug development value assessment of potential target regions, this invention proposes COMPASS-Anno, a target functional annotation method based on large language model enhancement, which specifically includes: S1: Combine multidimensional biological databases (such as UniProt, DrugBank, KEGG) to extract biological pathways, drug targets and functional annotation information; S2: Introduce an embedded large-scale language model and use a context-aware generation framework to deeply integrate and logically reason with the basic structured data extracted in step S1 and the unstructured text data obtained through automated document mining. S3: The model outputs a biological functional summary of the structural hotspots, an interpretation of the pathogenic molecular mechanisms in specific disease contexts, and a druggability assessment.

[0050] In one specific embodiment, COMPASS is a framework designed to integrate biobank-scale whole-genome sequencing data with protein structure computation, such as... Figure 2As shown, its core design principle lies in mapping disease-related coding variations to protein structural hotspots and functionally operable sites. COMPASS's main goal is to bridge the gap between statistical genetic associations and mechanistic structural insights, thereby elucidating complex disease mechanisms and prioritizing clinically promising drug targets. To achieve this goal, COMPASS systematically integrates three core analysis modules. First, the sequence analysis module (COMPASS-Seq) is responsible for mapping variation association signals to amino acid changes. When genetic heterogeneity arises due to multiple nucleotide variations at the same amino acid position, this module employs an optimal subset strategy and a leave-one-out strategy to screen for the dominant amino acid sequence with the most significant association signal for subsequent structural analysis. Second, the structure analysis module (COMPASS-Struct) performs fragment scanning based on experimentally determined and AI-predicted protein structures. By constructing local spatial neighborhoods centered on the Cα atom of the residue, it locates disease-related regions. This module calculates the aggregate association statistics for each fragment and marks the statistically most significant fragments as disease-related structural hotspots. These hotspots typically correspond to key functional sites and can guide drug development. Finally, the annotation module (COMPASS-Anno) not only integrates static annotations from databases, such as biological pathways, functional domains, tissue-specific expression, and drug target status, but also incorporates AI-assisted automated literature mining technology. It uses large language models to summarize specific disease backgrounds, generate mechanism hypotheses, and assess treatment relevance, thereby providing researchers with actionable biomedical insights.

[0051] In one specific embodiment, the sequence analysis module (COMPASS-Seq) is responsible for selecting the dominant amino acid sequence. Addressing the genetic heterogeneity issue of multiple nucleotide variations at the same amino acid position, the COMPASS-Seq module aims to screen for the most representative dominant amino acid sequence for subsequent structural analysis. This module first maps disease-related variations to amino acid positions in the reference transcript, and then employs two complementary statistical optimization strategies (e.g., ...). Figure 3 As shown in the diagram, in the optimal subset selection, the module systematically evaluates all possible candidate sequence combinations, identifies and selects the combination with the smallest association statistic (i.e., the strongest association signal) as the dominant amino acid sequence. In the leave-one-out selection, for each amino acid position with multiple nucleotide variations, the module performs iterative evaluation, removing individual variations one by one and recalculating the association statistic of the remaining sequence set. Based on the principle of maximum signal loss, it selects those variations that result in the largest association statistic of the remaining set after removal (i.e., the most significant decrease in significance), and identifies them as the dominant variation at that position. This process effectively solves the complex heterogeneity problem caused by different amino acid substitutions resulting from different nucleotide mutations of the same codon, ensuring that the sequences input for structural analysis have the highest representativeness of pathogenic association.

[0052] Optimal subset selection: Let This indicates the total number of amino acid sites in the reference transcript. (and This indicates the number of sites carrying amino acid alterations caused by disease-related coding variations. For the... Variable residue sites ( ),set up The set of variants at a given locus is defined as the number of disease-related coding variants mapped to that locus. . Candidate variant combinations Defined as the set generated by selecting one mutation from each variable residue site, i.e. , and . This sequence selection method generates a set of candidate variants by applying all substitutions at each variable site. For each candidate combination... COMPASS-Seq calculates the corresponding disease association statistics. And select the combination that minimizes the correlation statistic as the optimal solution: . While maintaining the positions of all other non-variant residues consistent with the reference transcript, the optimal mutation set is used. The induced amino acid sequence is defined as the dominant amino acid sequence used for structural analysis. It should be noted that optimal subset selection is only applicable when the search space is small, because the total number of possible combinations is limited. It will grow exponentially as the search space expands.

[0053] Leave-one-out-of-one selection method: Employs the notation of optimal subset selection, let... amino acid position The set of mutations at a given location is defined. This is the union of all disease-related coding variations. For location... Choose the method of leaving one step behind:

[0054] This involves selecting the variant that, after removal, produces the largest association statistic in the remaining variant set, thus resulting in the most significant loss of association signal. Applying this strategy independently to each variable residue position yields the dominant variant set.

[0055] Combined with all other positions in the preserved reference transcript, by The induced amino acid sequence is the dominant amino acid sequence used for structural analysis. This single-round selection needs to be performed. This evaluation, while being computationally efficient, retains the most critical substitutions for the overall signal at each residue position.

[0056] Example (e.g.) Figure 4 As shown): In the leave-one-out selection method, variants are first grouped according to their mapped amino acid positions. For sites containing only a single variant, the variant is retained by default without further selection. Specifically, Variant1 at position 1, Variant4 at position 3, and Variant5 at position 4 are all retained. For sites with multiple candidate variants, COMPASS-Seq performs a leave-one-out evaluation to determine the dominant variant. Two variants (Variant2 and Variant3) are present at site 2. After removing one variant each time (leaving the others unchanged), the gene association statistic is recalculated. Removing Variant3 resulted in a greater increase in the association statistic than removing Variant2, indicating a more significant loss of association signal. Therefore, Variant3 was selected as the dominant variant at site 2. Three variants (Variant6, Variant7, and Variant8) were observed at site 5. Using the same leave-one-out evaluation, removing Variant6 resulted in the largest increase in the association statistic among the three candidate variants. Therefore, Variant6 was selected as the dominant variant at this site. After resolving all positions with multiple nucleotide variations, the final dominant variant set includes Variant1, Variant3, Variant4, Variant5, and Variant6. These selected variants, together with the reference amino acid sequences for all other positions, define the dominant amino acid sequences used for subsequent structural analyses.

[0057] The leave-one-out selection method is used iteratively, unlike the one-step method, to determine the dominant mutation through optimization. The same notation is used to initialize the current mutation set. and the selected dominant variant set In the first The next iteration ( When COMPASS-Seq calculates... Evaluate each candidate variant in the current set And select the variant that maximizes the association statistic:

[0058] set up Selected variants The location of the residue. Update the set state:

[0059]

[0060] Thus Set as The final selection of the position is made, and all other variants corresponding to that residue position are removed from the set to be screened. This process is repeated until every variable residue position is assigned a variant, resulting in the following set:

[0061] This is the final set of dominant variants selected. Combined with all other positions in the retained reference transcript, the resulting amino acid sequence is the dominant amino acid sequence used for structural analysis.

[0062] Example (e.g.) Figure 5 As shown): In the stepwise leave-one-out selection method, variants are first grouped according to their mapped amino acid positions. Positions containing only a single variant (such as Variant1 at position 1, Variant4 at position 3, and Variant5 at position 4) are retained and fixed throughout the process. For positions containing multiple candidate variants, the stepwise leave-one-out selection method is applied.

[0063] Iteration 1: Starting from the complete variant set, remove a single candidate variant each time (keeping the rest), and recalculate the gene association statistics. Removing Variant2 resulted in the largest increase in the association statistics, so it was selected as the core variant at the second locus, while the candidate variant Variant3 at this locus was removed.

[0064] Iteration 2: Repeat the above process on the reduced variant set. At this point, there are still three candidate variants (Variant6, Variant7, and Variant8) at the 5th locus. Removing Variant6 results in the largest increase in the association P-value among the remaining candidate variants, so it is selected as the dominant variant at the 5th locus, while Variant7 and Variant8 are removed.

[0065] After processing all positions with multiple variants, the final master variant set includes Variant1, Variant2, Variant4, Variant5, and Variant6. These, along with the retained variants and reference amino acid sequences for all other positions, define the dominant amino acid sequence used in subsequent structural analyses.

[0066] In one specific embodiment, the structure analysis module (COMPASS-Struct) focuses on neighborhood scanning and structural hotspot identification. The COMPASS-Struct module aims to accurately identify disease-related hotspots using fragment scanning technology based on three-dimensional structures. This module supports a dual-channel structural input mode: it includes both wild-type protein structures and mutant structures generated from dominant amino acid sequences screened by the COMPASS-Seq module. These structural models can originate from experimental methods (such as cryo-electron microscopy or X-ray crystallography) or artificial intelligence prediction tools; for AI prediction models, the structural model with the highest predicted template modeling (pTM score) score is preferentially selected for analysis. This dual-input mode allows fragment scanning to not only capture hotspots reflecting the intrinsic function of proteins but also identify structural hotspots directly affected by specific amino acid alterations.

[0067] For each residue in the protein structure COMPASS-Struct defines a patch as... A spherical region centered on the Cα atom, encompassing all spatially adjacent residues within a user-specified radius. The set of adjacent residues for this patch is defined as:

[0068] in Represents residues and The Euclidean distance between them. Represents the user-defined radius (in Å).

[0069] In practical applications, to ensure statistically comparable local coverage of proteins of different sizes, COMPASS-Struct employs a method based on the target protein sequence length. Recommended radius based on experience:

[0070] For each patch's defined set of neighboring residues, COMPASS-Struct maps and extracts the corresponding set of neighboring variants, directly guided by the mapping relationships in the dominant amino acid sequences determined by COMPASS-Seq. The module evaluates each patch by calculating disease association statistics aggregated based on this set of neighboring variants and ranks them according to the association statistics. The patch with the most significant association statistics (Top 1) is reported as a disease-related hotspot, while the top ten patches serve as supplementary structural background to support robustness assessment and spatial consistency verification of the hotspot signal.

[0071] In one specific embodiment, COMPASS-Anno: a web implementation with an interactive interface.

[0072] The COMPASS web server is built as an integrated platform that achieves comprehensive annotation of protein functions by coordinating high-throughput biological data mining and artificial intelligence technologies. The system first integrates multidimensional biological attributes from authoritative databases: specifically, it maps signaling pathways and related disease phenotypes from KEGG; obtains tissue-specific expression profiles and subcellular localization data from UniProt; and incorporates the domain architecture of Pfam and known drug-target interaction information from DrugBank. These standardized data layers are not only presented intuitively through interactive visualization dashboards, but also provide structured "fact benchmark" contexts for generative large-scale language models (LLMs, including GPT-5-nano and DeepSeek-V3). Based on a context-aware generative framework, LLMs perform deep processing on the above heterogeneous inputs to generate coherent functional summaries, interpret the pathogenic relevance of functional hotspots in specific disease contexts, and propose evidence-based mechanistic hypotheses and treatment strategies, thereby transforming static database entries into practically guiding (actionable) biomedical insights.

[0073] In this embodiment, the present invention applied COMPASS to whole-genome sequencing data from approximately 500,000 participants in the UK Biobank, covering 83 cancer phenotypes, to assess variant-structure mapping in heterogeneous cancer types. We used STAAR-Burden P values ​​to quantify disease associations of rare variants (minimum allele frequency <1%) and AlphaFold3 to predict protein structures. COMPASS reproduced known structural hotspots and revealed shared hotspots across diseases, suggesting the potential for drug retargeting. A total of 252 gene-trait pairs with genome-wide significance were identified across the 83 cancer phenotypes (STAAR-Burden P < 2.5 × 10⁻⁶). -6= 0.05 / 20,000), including 9 genes already identified as drug targets. Since protein truncation and loss-of-function variants typically result in overall loss of function rather than causing interpretable local structural perturbations, this invention applies a structural analysis module to missense variants, including those annotated as destructive. Of these 9 drug target genes, 2 were driven solely by protein truncation or loss-of-function signals and were therefore excluded from structure-based analysis, leaving 7 genes for subsequent analysis. Of these 7 drug target genes, COMPASS detected structural hotspots overlapping with the binding pockets of approved drugs in 4 of them. These included the JAK2 hotspot, which spans the ATP-binding pocket targeted by the approved inhibitor; and an XPO1 hotspot, which overlaps with the nuclear output signaling groove targeted by a selective inhibitor. The sharing of multiple hotspots across different phenotypes indicates a convergent mechanism across diseases and provides a structural basis for therapeutic retargeting.

[0074] In one embodiment, whole-genome sequencing analysis of polycythemia vera (PV) revealed a rare disruptive missense variant in the JAK2 gene that was significantly associated with the disease across the entire genome. ; Figure 6 As shown in a). According to the annotation module COMPASS-Anno, JAK2 belongs to the Janus kinase (JAK) family and transduces cytokine receptor signals through phosphorylation of cytokine receptors and STAT transcription factors. JAK2 contains a catalytically active JH1 tyrosine kinase domain and a regulatory JH2 pseudokinase domain, the latter limiting JH1 activity through autoinhibition. Figure 6 As shown in d). Applying COMPASS-Seq to integrate multiple nucleotide changes at the same residue into a single dominant amino acid substitution further strengthens the association ( ; Figure 6 (As shown in b). The resulting amino acid sequence includes the V617F mutation, a key driver of PV pathogenesis, detected in over 95% of PV patients. The V617F mutation introduces a larger hydrophobic aromatic side chain, leading to steric hindrance and enabling it to undergo π-π stacking interactions with aromatic residues near the αC helix (as shown in b). Figure 6 (as shown in c). This structural alteration disrupts the stability of the JH2-JH1 self-inhibition interface, leading to persistent activation of the JH1 kinase domain, abnormal JAK2-EpoR signaling, and excessive erythrocyte proliferation.

[0075] COMPASS-Struct aggregated variant-phenotypic signals in the three-dimensional local space of proteins and located a more significant PV-related hotspot. This hotspot is centered on the Met600 residue and has a radius of [missing information]. and with clinical and investigational drugs ( Figure 6 The binding site (shown in image e) highly overlaps with the ATP-binding pocket within the JH1 domain. Specifically, this hotspot contains the ATP-binding pocket and highly overlaps with the binding sites of the clinical drugs ruxolitinib and fedratinib. Both drugs occupy this pocket and act as competitive inhibitors, thereby inhibiting JAK2 kinase activity. The pyrrolopyrimidine core of ruxolitinib forms stable hydrogen bonds with Glu930 and Leu932, and van der Waals forces and hydrophobic interactions with the P-loop (Leu855, Gly856) and DFG motif (Asp994) firmly anchor the drug within the pocket, preventing ATP binding and inhibiting JAK2 kinase activity. Figure 6 (As shown in e). Furthermore, this hotspot extends to a subregion of JH2, consistent with the target site of the investigational drug flonoltinib maleate in a Phase II clinical trial, which binds to the JH2 non-catalytic site of the V617F mutant JAK2 (as shown in e). Figure 6 As shown in e), it stabilizes the pseudokinase domain and restores the autoinhibitory regulation of the JH1 kinase domain.

[0076] In addition to PV, whole-genome sequencing analysis also detected rare disruptive missense variants in JAK2 that were significantly associated with leukemia (including myeloid and acute myeloid leukemia), non-Hodgkin's lymphoma, and other lymphocytic and histiocytic carcinomas across the entire genome. Figure 7 (As shown in a). In these phenotypes, COMPASS-Seq selected the dominant amino acid sequence, including V617F. V617F mutations are frequently detected in secondary and new-onset leukemias, consistent with the shared pathogenic mechanisms among these diseases (see figure a). Figure 7 As shown in b). COMPASS-Struct further localized hotspot regions of these diseases, finding that they highly overlapped with PV hotspots, including the V617F neighborhood and ATP-binding pockets, supporting shared structural and functional drivers ( Figure 7 (As shown in c).

[0077] COMPASS-Anno summarized pathway evidence consistent with these findings. In JAK2-driven leukemia, persistent activation of STAT5 constitutes a core oncogenic signaling axis, driving leukemia cell proliferation and survival. Simultaneously, aberrant JAK2 signaling activates the PI3K-AKT-mTOR pathway, leading to upregulation of RNA-binding La protein. Increased La protein expression elevates MDM2 levels, resulting in weakened p53 tumor-suppressive function, thereby promoting leukemia progression. In certain lymphomas, JAK2 amplification or aberrant activation leads to persistent phosphorylation and nuclear translocation of downstream STAT molecules (including STAT3 and STAT5), driving transcriptional activation of pro-proliferative and anti-apoptotic genes, enhancing lymphoma cell proliferation and survival. In summary, these results provide a mechanistic basis for repositioning JAK2 inhibitors from PV to leukemia and lymphoma, consistent with clinical benefits observed in certain patient subgroups.

[0078] In one embodiment, whole-genome sequencing analysis of chronic lymphocytic leukemia (CLL) revealed a rare disruptive missense variant in the XPO1 gene that was significantly associated with the disease across the entire genome. COMPASS-Anno studies show that XPO1 is the primary nuclear export receptor, responsible for maintaining nucleoplasmic transport and signal homeostasis in eukaryotic cells. In the nucleus, XPO1 binds to RAN-GTP and recognizes the nuclear export signal (NES) of cargo proteins, forming a stable ternary complex. This complex is then transported to the cytoplasm via nuclear pore complexes. Hydrolysis of RAN-GTP in the cytoplasm to RAN-GDP triggers the disintegration of the complex and release of cargo proteins. Overexpression of XPO1 leads to aberrant nuclear export of key tumor suppressor factors (such as p53, p21, and FOXO), reducing their intranuclear activity and thus promoting the development and malignant progression of leukemia.

[0079] COMPASS-Seq integrates multiple nucleotide changes at the same residue into a single dominant amino acid substitution. The resulting sequence contains the pathogenic E571K amino acid substitution. E571K is one of the most frequent hotspot mutations in CLL (approximately 3%), altering substrate selectivity through a charge-complementary mechanism. The substitution of a negatively charged glutamate with a positively charged lysine reshapes the local electrostatic potential at the edge of the NES binding groove, enhancing the affinity of XPO1 for cargo proteins containing negatively charged residues, such as the NFκB inhibitor IκBα. This charge-dependent export-specific shift interferes with nucleocytoplasmic transport of key regulatory proteins and promotes the progression of CLL.

[0080] COMPASS-Struct aggregates variant-phenotypic signals within the local spatial neighborhood of a protein, locating a more significantly associated CLL-related hotspot. This hotspot, centered on the Lys534 residue and with a radius of 25 Å, covers the NES-binding groove, a hydrophobic cleft where XPO1 performs its core functions and is crucial for NES recognition, nuclear export, and inhibitor interactions. The hotspot includes the E571K neighborhood and the key residue Cys528, the covalent binding site for selective nuclear export inhibitors. SINE-type drugs, such as selinexor, occupy the NES-binding groove by covalently binding to Cys528, blocking the binding of cargo proteins carrying nuclear export signals. This covalent inhibition prevents XPO1-dependent nuclear export, causing tumor suppressor proteins to remain in the nucleus, thereby promoting transcriptional and apoptotic activity and ultimately inhibiting tumor growth. CLL cells carrying the E571K amino acid substitution exhibit higher sensitivity to XPO1 inhibitors such as selinexor, which may be related to increased drug binding affinity and accelerated degradation of the mutant protein, thus amplifying the drug's effect.

[0081] In addition to validating known targets, COMPASS identified 22 structural hotspots in genes not targeted by drugs. These hotspots clustered within conserved motifs, catalytic sites, and other functionally critical domains. These regions may represent potential therapeutic sites in genes for which currently lack approved therapies. For example, the OB-fold hotspot in POT1 is involved in telomere protection and leukemia susceptibility; and the catalytic hotspot in SAMHD1 is associated with impaired dNTP hydrolysis and potential genomic instability and altered chemotherapy response. Overall, these findings extend the framework from validation to discovery and provide a structural map to guide experimental investigations and genetically information-based target prioritization.

[0082] In one embodiment, whole-genome sequencing analysis of leukemia revealed a significant association between rare disruptive missense variants in the POT1 gene and the disease. COMPASS-Anno shows that POT1 is the only sheelterin subunit in the telomere protection protein complex that utilizes its tandem OB-fold domain to form a DNA-binding domain. This domain directly binds to the 3'-terminal guanine-rich single-stranded telomere overhang, helping to protect chromosome ends from being mistakenly identified as DNA damage. Through the binding of the C-terminus of XPO1 to TPP1, the POT1-TPP1 complex regulates telomerase recruitment and activity to maintain telomere length homeostasis. Mutations in POT1 or other sheelterin complex components (such as ACD / TPP1 and TERF2IP / RAP1) disrupt telomere protection, leading to chromosome end exposure, genomic instability, and an increased risk of malignant transformation in chronic lymphocytic leukemia. COMPASS-Seq integrates multiple nucleotide changes at the same residue into a single dominant amino acid substitution (…). ).

[0083] COMPASS-Struct identified hotspots associated with leukemia. The hotspot is located within the N-terminal OB1 and OB2 domains of POT1, centered on the Phe32 residue with a radius of 20 Å. These two consecutive OB-fold domains constitute the DNA-binding domain of POT1, specifically binding to the 3' single-stranded overhang containing the TTAGGG repeat sequence in telomeric DNA. At the molecular level, the OB1 domain establishes hydrogen bonds and base stacking interactions with the first 6 nucleotides of the telomeric repeat sequence, while the OB2 domain specifically locks onto the guanine at the 3' end, achieving high affinity and sequence-specific binding. Furthermore, the DNA backbone at the junction of the OB1 and OB2 binding grooves exhibits significant bending, promoting bridging between the two OB fold binding sites. COMPASS-Struct further precisely located key aromatic residues within these OB fold domains, including Phe31, Tyr161, Tyr223, and Tyr271, which are involved in base stacking interactions. This conformation effectively protects the ssDNA ends, preventing them from being misrecognized by DNA damage mechanisms and thus preventing the activation of abnormal damage responses. In summary, these results link genetic associations to a substructure crucial for telomere protection and are consistent with increased leukemia susceptibility when POT1 binding is impaired. This hotspot provides a concrete entry point for experimental investigations into how POT1 variants disrupt telomere DNA binding and for assessing the druggable potential of therapeutic modulation of nearby pockets or interfaces.

[0084] In one embodiment, whole-genome sequencing analysis of chronic lymphocytic leukemia revealed a significant association between a rare disruptive missense variant in the SAMHD1 gene and the disease. COMPASS-Anno shows that SAMHD1, as a deoxynucleoside triphosphate (dNTP) triphosphate hydrolase, is responsible for regulating intracellular dNTP levels and maintaining genomic stability. Loss-of-function mutations in SAMHD1 impair tetramerization and disrupt the catalytic pocket, leading to loss of dNTPase activity. This results in dNTP accumulation, causing DNA damage and genomic instability, and is associated with aberrant innate immune activation and tumorigenesis. COMPASS-Seq integrates multiple nucleotide changes at the same residue into a single dominant amino acid substitution, further enhancing the associated signaling. ).

[0085] COMPASS-Struct identified a hotspot region located within the catalytic domain. The SAMHD1 enzyme activity is centered at residue 304 with a radius of 20 Å. The enzyme activity depends on the conformational remodeling of its catalytic pocket during tetramerization. In the inactive dimer state, the substrate-binding pocket is loose and cannot effectively bind dNTP substrates. When dGTP binds to the allosteric site, SAMHD1 tetramerizes, promoting catalytic activation and dNTP hydrolysis. This hotspot region overlaps with the catalytic center and contains residues essential for substrate recognition and catalytic function, including Arg366, His370, Tyr374, Lys312, and Tyr315. During tetramerization, the α-helix shifts approximately 3 Å towards the catalytic pocket, enabling Arg366, His370, and Tyr374 to form hydrogen bonds and π-π stacking interactions with the phosphate groups, deoxyribose sugars, and bases of the substrate. Simultaneously, under the stabilization of the ring region (residues 303-308), Lys312 transforms from a disordered conformation to an ordered conformation, and synergistically with Tyr315 and Arg366 anchors γ-phosphoric acid in dGTP, while coordinating Mg²⁺ + The ions further stabilized the geometry of the active site. This conformational remodeling and residue repositioning stabilized the substrate within the catalytic pocket, forming an important structural basis for the efficient hydrolysis of dNTPs by SAMHD1.

[0086] In summary, these results pinpoint the genetic association to a specific catalytic substructure associated with CLL, providing a mechanistic framework for explaining how SAMHD1 mutations disrupt dNTP homeostasis and genome stability. This localization hotspot provides specific targets for experimental validation and may inspire novel strategies for regulating catalytic function or allosteric activation in therapeutically significant contexts.

[0087] COMPASS, the invention, is a computational framework designed to map disease-related coding variations to amino acid sequences and three-dimensional protein structures, ultimately linking them to disease mechanisms and therapeutic targets. The framework comprises three core modules: a sequence module, performing sequence-based amino acid selection; a structure module, performing structure-based fragment scanning to locate disease-related structural hotspots; and an annotation module, integrating a large language model for protein functional annotation and literature mining to simplify the variation interpretation process.

[0088] Applying this invention to 500,000 whole-genome sequencing data from the UK Biobank, covering 83 cancer phenotypes, COMPASS not only successfully reproduced established drug-related hotspots but also revealed cross-disease shared hotspots suggesting drug repositioning potential and identified potential drug targets. The cases of JAK2 and XPO1 demonstrate successful reproduction of known drug targets: the hotspot in JAK2 covers the V617F region and ATP-binding pocket, while the hotspot in XPO1 overlaps with the nuclear export signal binding groove. The sites within these two hotspots are spatially highly consistent with the binding sites of approved or investigational inhibitors. Furthermore, POT1 and SAMHD1 provide examples of novel target discovery; COMPASS successfully linked rare variant signals to structural mechanisms involved in telomere protection and deoxynucleoside triphosphate metabolism, which may provide direction for identifying potential druggable targets and guiding therapeutic regulation.

[0089] In summary, COMPASS provides a scalable and reproducible pathway from sequence variation to structural mechanisms. By integrating disease-related coding variations from sequencing data with experimentally determined and AI-predicted protein structures, this invention generates hotspot-level hypotheses that can be prioritized for orthogonal validation and therapeutic exploration. With its wider application in more traits and populations, COMPASS holds promise for facilitating the translation of human genetics findings into mechanism-driven target discovery.

[0090] The present invention also discloses a computer program product or system, including a computer program that, when executed by a processor, implements the above-described screening method steps for potential drug target regions.

[0091] This invention provides a screening system for potential drug target regions, comprising: Acquisition module: Acquires gene variant data that are statistically associated with the target disease; Sequence module: Filters the gene variation data to obtain the dominant amino acid sequence; Structural module: Obtain the protein structure predicted by the dominant amino acid sequence or obtain the wild-type protein structure of the target gene; traverse each amino acid residue in the predicted protein structure or wild-type protein structure, construct a local spatial neighborhood centered on the residue, identify variants falling into each local spatial neighborhood, and calculate the aggregated disease association statistic of the variant set within the local spatial neighborhood; Target module: Sort the aggregated disease association statistics and select the top K local spatial neighborhoods as potential drug target regions, where K is a natural number greater than or equal to 1; In one embodiment, the system further includes an annotation module for providing sequence-level functional annotations and / or druggability assessments of structural hotspots.

[0092] An embodiment of the present invention provides a computer device, specifically comprising: The system includes a memory and a processor; the memory is used to store program instructions; the processor is used to invoke the program instructions when the program instructions are executed to perform the above-described variant screening method or to execute the above-described screening method for potential drug target regions.

[0093] The present invention also discloses a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, is used to perform the above-described method for screening potential drug target regions.

[0094] The verification results of this verification embodiment show that assigning inherent weights to indications can improve the performance of this method compared to the default settings. Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, indirect coupling or communication connection of devices or units, and may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated; the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of this embodiment. Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. This program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0095] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0096] The computer device provided by the present invention has been described in detail above. For those skilled in the art, there will be changes in the specific implementation and application scope based on the ideas of the embodiments of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for screening potential drug target regions, characterized in that, include: Obtain gene variant data that are statistically associated with the target disease; The dominant amino acid sequence was obtained by screening the gene variation data; Obtain the protein structure predicted from the dominant amino acid sequence or obtain the wild-type protein structure of the target gene; Traverse each amino acid residue in the predicted protein structure or wild-type protein structure, construct a local spatial neighborhood centered on the residue, identify variants falling into each local spatial neighborhood, and calculate the aggregated disease association statistic of the variant set within the local spatial neighborhood. The aggregated disease association statistics are sorted, and the top K local spatial neighborhoods are selected as potential drug target regions, where K is a natural number greater than or equal to 1.

2. The method for screening potential drug target regions according to claim 1, characterized in that, The screening includes identification and classification, classifying the variant data of the target gene into single variants and multiple nucleotide variants, screening the multiple nucleotide variants to obtain the first variant, combining the first variant with the single variant to obtain the dominant variant, and constructing the dominant amino acid sequence through the dominant variant. Optionally, the screening process for each of the multiple nucleotide variants is as follows: S1. Keeping the variations at other positions in the gene unchanged, remove a single candidate variation from the candidate variation set at the current position, and calculate the overall disease association statistic of the remaining variation set after removing the variation. Repeat this step until all candidate variations at the current position are traversed to obtain a set of statistics after removal. S2. Compare and sort the set of statistical data after removal, select the variant with the largest disease association statistic in the set of remaining variants after removal, and obtain the dominant variant at the current position. S3. Repeat steps S1-S2 to obtain a variant or set of variants with multiple nucleotide multiple variation positions. Optionally, the positions of multiple mutations in all multiple nucleotide variants are screened to obtain a first variant or a set of first variants. The first variant or set of first variants is combined with a single variant to obtain a dominant variant. A dominant amino acid sequence is constructed through the dominant variant. Optionally, the step of selecting the variant with the largest disease association statistic in the remaining variant set after removal is achieved by sequentially removing individual candidate variants at the current position, calculating the disease association statistic of the remaining variant set for each variant, and directly comparing the calculated remaining set statistics to select the variant corresponding to the largest statistic.

3. The method for screening potential drug target regions according to claim 2, characterized in that, The process of screening multiple nucleotide variants to obtain a first variant is replaced by: performing a second screening on multiple nucleotide variants to obtain a second variant or a set of second variants; combining the second variant with a single variant to obtain a dominant variant; and constructing a dominant amino acid sequence through the dominant variant; the second screening process is as follows: defining the set of multiple variant sites of multiple nucleotides as the initial current set; In the initial current set, a removal operation is performed on each candidate variant of the screening site, and the overall disease association statistic of the remaining variant set after removing the variant is calculated; the variant with the largest overall disease association statistic of the remaining variant set after removal is selected, and the variant is marked as the dominant variant at its amino acid position; Remove other candidate variants from the initial current set that have been labeled with the dominant variant amino acid position, update the current set to obtain an updated set, and repeat the above steps based on the updated set until the amino acid position of each multiple nucleotide multiple variant site is labeled with a dominant variant, to obtain a second variant or a second variant set.

4. The method for screening potential drug target regions according to claim 2, characterized in that, The process of screening multiple nucleotide variants to obtain the first variant is replaced by: a third screening of multiple nucleotide variants to obtain a third variant or a set of third variants. The third variant is combined with a single variant to obtain a dominant variant, and a dominant amino acid sequence is constructed through the dominant variant. The third screening process is as follows: for each multiple nucleotide variant site, one variant is selected from the candidate variant set, and together with the variant at the single variant site, a candidate variant combination is generated through full permutation and combination. For each generated variant combination, the overall association statistic with the target disease is calculated, the overall association statistic is compared and ranked, and the variant combination with the smallest overall association statistic is selected to obtain the dominant amino acid sequence.

5. The method for screening potential drug target regions according to claim 1, characterized in that, The process of obtaining the target gene is as follows: based on whole-genome sequencing data, all coding variation data of the associated gene in a specific disease background are extracted, the coding variation is mapped to the amino acid site corresponding to the reference transcript, all amino acid sites containing the variation constitute the screening set, and the screening set is screened to obtain the dominant amino acid sequence.

6. The method for screening potential drug target regions according to claim 1, characterized in that, The construction of the local spatial neighborhood is a spherical region formed with a predetermined length and radius, centered on the Cα atom of the amino acid residue. Optionally, the preset length radius is set according to the length of the target protein sequence.

7. The method for screening potential drug target regions according to claim 1, characterized in that, The method also includes functional annotation. Potential drug target regions are extracted from biological databases to extract relevant biological pathways, drug targets, and functional annotation information to construct basic structured data and provide detailed biological background annotations. On this basis, potential target regions, structured data, and biological background annotations are input into a large language model to obtain biological functional summaries covering structural hotspots, interpretations of pathogenic molecular mechanisms in specific disease contexts, and druggability assessments.

8. A computer screening system having computer programs or instructions, characterized in that, include: Acquisition module: Acquires variant data of the target gene; Sequence screening module: Screens variant data of the target gene to obtain the dominant amino acid sequence; Protein structure module: Obtain the protein structure predicted from the dominant amino acid sequence or obtain the wild-type protein structure of the target gene; Variant structure module: Traverse each amino acid residue in the predicted protein structure or wild-type protein structure, construct a local spatial neighborhood centered on the residue, identify variants falling into each local spatial neighborhood, and calculate the aggregated disease association statistic of the variant set within the local spatial neighborhood; Target module: Sort the aggregated disease association statistics and select the top K local spatial neighborhoods as potential drug target regions, where K is a natural number greater than or equal to 1; Optionally, the system further includes an annotation module for providing sequence-level functional annotations and / or druggability assessments of structural hotspots.

9. A computer device comprising a memory, a processor, and a computer program or instructions stored in the memory, characterized in that, The computer program or instructions are executed by a processor to implement the screening method for potential drug target regions as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, The computer program or instructions are executed by a processor to implement the screening method for potential drug target regions as described in any one of claims 1-7.