Method for identifying hot-spot anchor sites of thermosensitive udg based on dual-pathway evolutionary analysis

CN122337334BActive Publication Date: 2026-09-29PULUOMAIGE BIOLOGICAL PRODS SHANGHAI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610428855.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-02
Publication Date
2026-09-29
Estimated Expiration
2046-04-02

AI Technical Summary

Technical Problem

[0004]然而,对于热敏UDG这一兼具保守催化功能基序(motif)与显著谱系分化特征的酶系统,上述单一路径方法均存在固有的技术局限:

Benefits of technology

1、序列空间边界约束与谱系平衡设计:通过家族闸口机制将分析目标精准限定于UDG家族1型催化核心域一致的序列空间,同时以跨原核与真核的多源热敏种子作为分析锚点,从方法源头确保结果的功能相关性并有效规避单一谱系偏倚;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122337334B_ABST
    Figure CN122337334B_ABST
Patent Text Reader

Abstract

The application discloses a heat-sensitive UDG anchor site identification method based on double-path evolution analysis. The application adopts two paths for site identification, wherein the first path constructs a double-layer multiple sequence alignment of a full family layer and a cold adaptation subset layer in a sequence space defined by a UDG family 1 type family gate, and obtains a functional motif anchor site set by combining column-level conservation, branch difference degree and a conserved functional motif neighborhood window constraint; the second path constructs a homologous multiple sequence alignment with a seed sequence as a center, and obtains a seed homologous core anchor site set by fusing alignment statistics and a sequence spectrum hidden Markov model for joint scoring. Finally, the two path results are fused in a seed sequence residue coordinate system, and site index and corresponding amino acid types are output, which can be directly used as a fixed site constraint input for heat-sensitive UDG protein engineering modification or generative protein design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of bioinformatics, evolutionary biology and protein engineering, and specifically relates to a method for identifying key anchor points of thermosensitive uracil-DNA glycosylase (UDG) based on a dual evolutionary pathway. Background Technology

[0002] Uracil-DNA glycosylase (UDG) is one of the core enzymes in the Base Excision Repair (BER) pathway. Its catalytic function is to specifically recognize and hydrolyze the N-glycosidic bond between uracil residues and deoxyribose in the DNA strand, releasing free uracil and generating a base-free site (AP site), thereby initiating downstream repair cascade reactions. Due to the common problem of carry-over contamination in polymerase chain reaction (PCR) amplification and nucleic acid detection processes, the "dUTP / UDG system" is widely used in industry and research for prevention. The principle of this system is as follows: in the amplification reaction, deoxyuridine triphosphate (dUTP) replaces deoxythymidine triphosphate (dTTP), so that the thymine position in the historical amplification product is replaced by uracil; before the start of a new round of amplification, UDG is added to specifically degrade the uracil-containing contaminating template, and then UDG is inactivated by temperature treatment to ensure that the normal amplification of the target template is not interfered with. The aforementioned application scenarios place specific requirements on the thermostability of UDGs, necessitating a type of "thermally sensitive" or "thermally unstable" UDG. This type of enzyme needs to maintain sufficient catalytic activity at relatively low reaction temperatures (typically room temperature to 37°C) while being rapidly and irreversibly inactivated when the temperature rises to the PCR denaturation step temperature (typically above 50°C), thus avoiding interference with subsequent target amplification. Currently available commercially available thermosensitive UDG products are mainly derived from psychrophilic bacteria or other low-temperature adapted microorganisms. UDGs from different sources exhibit significant differences in primary sequence, three-dimensional structure, and thermostability parameters. For example, UDG derived from the psychrophilic bacterium BMTU3346 has been reported in the literature to possess the characteristic of complete inactivation at relatively low temperatures.

[0003] When conducting rational protein engineering for thermosensitive undefined genes (UDGs) or designing de novo proteins based on generative deep learning models, it is usually necessary to predetermine a set of evolutionarily constrained anchor residues as "rigid constraints" in the backbone or sequence generation process to prevent the design process from disrupting catalytic active sites, substrate recognition frameworks, or functionally related flexible loop structures. Traditional methods for determining constraint sites often rely on a single path: the first type of method constructs multiple sequence alignment (MSA) based on the entire family sequence set and selects conserved sites based on the statistical characteristics of each column of the alignment matrix. Common evaluation metrics include column occupancy, Shannon entropy, and their derived residue conservation scores. The theoretical basis of these metrics originates from Shannon information theory and has mature applications in the field of bioinformatics. The method of combining MSA with information entropy has been widely used for the evolutionary conservation analysis of protein families. The second type of method starts with a specific target sequence (i.e., the "seed sequence") and uses iterative homology search combined with probabilistic modeling of Hidden Markov Models (HMMs) to identify conserved sites within a context of homologous sequences that are evolutionarily closer to the seed sequence. For example, the Jackhmmer tool in the HMMER software suite is a publicly available implementation of iterative homology search. It gradually expands the range of homologous sequences through multiple rounds of search and builds and optimizes a profile HMM during the iteration process to obtain a more statistically robust assessment of site conservation.

[0004] However, for the thermosensitive UDG enzyme system, which possesses both conserved catalytic motifs and significant lineage differentiation characteristics, the aforementioned single-pathway methods all have inherent technical limitations: (1) The whole family conservation analysis has the problem of “signal averaging”: Although some sites do not show high conservation at the whole family level, they play a key role in the regulation of adaptive functions in specific psychrophilic adaptation lineages or specific taxa (such as UDG of fish origin); if site screening is based solely on the global conservation threshold, such sites with branch-specific functional significance are easily missed.

[0005] (2) Single seed homology analysis has the problem of 'pedigree bias': the homology sequence set obtained based on a single seed sequence search may have insufficient pedigree coverage or be biased towards a specific evolutionary branch, resulting in the failure to fully identify and emphasize core skeletal sites that are generally conserved and functionally essential across the entire family.

[0006] (3) Thermosensitive UDGs (especially UDG family type 1) have multiple highly conserved functional motifs across species in their primary structure, including sequence segments related to catalytic activity, substrate recognition specificity and structural flexibility; however, existing methods only divide sites based on a single statistical conservation or difference threshold, which makes it difficult to achieve the hierarchical constraint strategy of "rigid constraint of core catalytic region" and "moderate tunability of functional motif neighborhood". Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a dual-pathway evolutionary constraint anchor point identification method for thermosensitive uracil-DNA glycosylase (UDG family type 1). This method uses prokaryotic and eukaryotic thermosensitive UDG seed sequences obtained from publicly available literature and databases as the starting point for analysis. It synergistically utilizes two complementary evolutionary information pathways: the "family gate double-layer multiple sequence alignment pathway" and the "single seed homology background analysis pathway," to obtain a set of evolutionary constraint anchor points that can be directly mapped to the residue coordinates of the seed sequence. The output includes site indexes and corresponding amino acid types, thus providing a biologically interpretable and methodologically reproducible evolutionary constraint basis for protein engineering and generative protein design of thermosensitive UDGs.

[0008] Specifically, this method can synergistically utilize two complementary evolutionary information pathways at the family and seed levels: First, using the family domain model of UDG family type 1 as the sequence admission criterion, three types of cold adaptation lineage subsets are independently constructed and merged to form a whole family layer panel. On this basis, a two-layer multi-sequence alignment is constructed with the subset layer and the whole family layer sharing a unified alignment coordinate system, thereby simultaneously capturing common conserved signals across lineages and differential signals of cold adaptation-specific branches; Second, with a specific seed sequence as the center, multi-sequence alignment is constructed in the set of orthologous sequences that are closer to it in evolution, and site-level joint probability scoring is performed in combination with the sequence spectrum hidden Markov model, thereby obtaining a set of seed homology core anchor points that are statistically robust, biologically interpretable, and can be directly mapped to the seed sequence residue numbering system. This invention, by integrating the outputs of the two analytical paths mentioned above, yields a set of evolutionary constraint anchor points that can provide systematic evolutionary evidence for the subsequent protein engineering of thermosensitive UDGs. This not only locks in the stability information of the catalytic framework and key functional motif neighborhoods at the site level, effectively reducing the risk of accidental damage to the core functional structure during engineering, but also identifies branching differential sites related to cold adaptation. This provides actionable site-level guidance for subsequent rational optimization around goals such as thermosensitive inactivation characteristics, activity maintenance, and stability balance, thereby systematically improving the reliability and engineering reproducibility of thermosensitive UDGs as controllable contamination removal tool enzymes in molecular biology reaction systems.

[0009] To achieve the above-mentioned objectives, the technical solution of the present invention is as follows: I. A Thermosensitive UDG Constraint Anchor Point Identification Method Based on Dual-Path Evolutionary Analysis (1) Obtain thermosensitive UDG family type 1 sequences from both eukaryotic and prokaryotic sources from publicly available literature and / or publicly available protein sequence databases, and designate them as seed sequences; that is, the seed sequences should contain at least two sequences, namely one thermosensitive UDG family type 1 sequence from eukaryotic sources and one thermosensitive UDG family type 1 sequence from prokaryotic sources. Prioritize screening UDG family type 1 sequences with clear evidence of thermosensitive characteristics or cold adaptation phenotypes from publicly available literature and publicly available sequence databases such as UniProtKB and NCBI as the starting point for analysis. In this way, the selected seed sequences cover at least two lineages from both prokaryotic and eukaryotic sources, so as to reduce the bias of a single lineage.

[0010] (2) Construct a subset of cold adaptation sequences based on the obtained seed sequences, and assemble each subset of cold adaptation sequences into a target family sequence library; then generate a set of first path function motif constraint anchor points for the seed sequences based on the target family sequence library and the subset of cold adaptation sequences. (3) Construct a seed homology sub-library based on the obtained seed sequence, and then generate a set of second path seed homology core anchor points based on the seed homology sub-library; The sequence library construction strategies of the first and second paths differ significantly, as follows: The first path focuses on characterizing the "commonalities of family skeletons and differences in cold adaptation branches." Based on ecological adaptation types or temperature adaptation information, it independently constructs cold adaptation-related branch subsets with preset lineage boundaries, and merges these subsets into a complete family sequence set after screening at the family gate. The second path focuses on characterizing the "direct homology background of single seeds." Instead of using cold adaptation characteristics as a prerequisite for sequence screening, it selects a classification homology sequence library that matches the lineage of each seed sequence and has a broader coverage as the analysis background. For example, for seeds from fish, a homology sequence library covering the vertebrate range can be selected, that is, the scope is broadened to the higher taxonomic level to which the seed belongs, to ensure the stability and lineage representativeness of single-seed homology statistical analysis.

[0011] (4) An evolutionary constraint anchor point set is generated based on the first-path functional motif constraint anchor point set and the second-path seed homologous core anchor point set of the seed sequence. The evolutionary constraint anchor point set includes the site residue number and the corresponding amino acid type. Specifically, the first-path functional motif constraint anchor point set and the second-path seed homologous core anchor point set of the seed sequence are fused in the seed sequence coordinate system. The former is used to characterize the conservation of the family backbone and the differences in cold adaptation branches, while the latter is used to characterize the stable core sites in the homologous background of the individual. The evolutionary constraint anchor point set can be directly used as the fixed site constraint input in the subsequent protein engineering targeted modification or generative protein design process.

[0012] In step (2), the construction of a target family sequence library and a cold-adapted sequence subset based on the obtained seed sequences includes: Based on the lineage origin of multiple seed sequences, a preset lineage boundary is determined. The UDG family type 1 family domain model is used as the sequence admission criterion. Multiple cold-adapted sequence subsets are independently constructed according to the preset lineage boundary. After family gate screening, redundancy removal and balanced sampling are performed on each subset, the subsets are merged and assembled into a target family sequence library.

[0013] The family gate screening is implemented using the PF03167 family domain model in the Pfam database. Specifically, it includes: using the PF03167 family domain model to perform a hidden Markov model scan on candidate sequences to obtain hit intervals; screening based on coverage indicators, retaining candidate sequences with sequence coverage ≥ 0.50 and model coverage ≥ 0.70 to enter the target family sequence library, that is, only retaining sequences with the same structure as the catalytic core domain of the UDG family type 1, thereby eliminating the interference of non-target family sequences or sequences containing non-homologous structural domains on the quality of multiple sequence alignment and downstream statistical analysis results; wherein, sequence coverage is defined as the proportion of the hit interval to the total length of the sequence, and model coverage is defined as the coverage ratio of the hit interval to the PF03167 model.

[0014] The cold-adapted sequence subset includes at least one cold-adapted eukaryotic subset and one cold-adapted prokaryotic subset. Specifically, a preset lineage boundary is determined based on the lineage origin of the seed sequence, and candidate sequences are screened and merged according to the preset lineage boundary, such as a subset of marine psychrophilic γ-proteobacteria, a subset of cold-adapted or low-temperature active Gram-positive bacteria, and a subset of cold-adapted fish. Each subset can be further subdivided.

[0015] Furthermore, the cold-adapted sequence subsets were clustered and redundancy removed according to the sequence consistency threshold; then, balanced sampling was performed on each of the redundancy-removed cold-adapted sequence subsets to control the scale of multiple sequence alignments and improve the representativeness of the lineage, thus obtaining the final cold-adapted sequence subsets.

[0016] In step (2), based on the target family sequence library and the cold-adapted sequence subset, a set of first path functional motif constraint anchor points for the seed sequence is generated, including: Construct a full family layer multi-sequence alignment on the target family sequence library, and extract the sequence rows corresponding to each cold-adapted sequence subset from the full family layer multi-sequence alignment results to form a subset layer multi-sequence alignment corresponding to that cold-adapted sequence subset, so that the full family layer and the subset layer share the same alignment column coordinate system; Based on whole-family layer multiple sequence alignment, the loci are classified according to column-level statistical features to obtain family-level locus classification results including backbone conserved loci, branch-differentiated loci, and background stable loci; based on subset layer multiple sequence alignment corresponding to each cold-adapted sequence subset, the loci of each cold-adapted sequence subset are classified according to column-level statistical features to obtain subset layer classification results for each cold-adapted sequence subset. By combining the functional motif neighborhood window set W, the functional motif constraints of the family-level site classification results and the subset-level classification results are further classified to obtain the family-level three-level anchor point set and the subset-level three-level anchor point set corresponding to each cold-adapted sequence subset. Then, the family-level three-level anchor point set and the subset-level three-level anchor point set corresponding to each cold-adapted sequence subset are merged according to hierarchical categories to generate the bilayer joint anchor point set corresponding to each cold-adapted sequence subset. The bilayer joint anchor point set is used to supplement differential sites. According to the phylogenetic affiliation of each seed sequence, the bilayer joint anchor point set of its corresponding cold-adapted sequence subset is mapped to the residue coordinates of the seed sequence to obtain the first path functional motif constraint anchor point set corresponding to each seed sequence.

[0017] Furthermore, in the full family layer multi-sequence alignment results, the column occupancy of the alignment columns is not less than a preset threshold, in order to perform quality control on the alignment columns.

[0018] The classification methods for skeletal conserved sites, branched differential sites, and background stable sites are as follows: Category A loci are conserved loci with a complete family skeleton. The selection criteria are Occ ≥ 0.85 and column conservation ranking in the top 10% of all aligned columns.

[0019] Category C loci are branch-differential loci, selected based on Occ ≥ 0.60 and Div located in the top 10% of all alignment columns. If the number of Category C loci obtained according to this rule is insufficient, they are supplemented according to Div from high to low to a minimum of 30 columns or a minimum of 18% of the total number of alignment columns, and the larger of the two values ​​is taken.

[0020] Category B loci are background stable loci. After excluding categories A and C, the screening criteria are Occ ≥ 0.75 and column conservatism not lower than the median of all aligned columns.

[0021] The functional motif neighborhood window set W is obtained by locating conserved functional motifs shared across lineages in the UDG family type 1 primary sequence. These conserved functional motifs include sequence segments associated with catalytic active centers, structurally flexible ring regions, and substrate recognition interfaces. The functional motif neighborhood window is formed by extending a predetermined number of residues to both sides of the functional motif segment. By introducing this prior window, a verifiable and reproducible correspondence is established between the anchoring point screening results and the aforementioned key functional regions. This preserves interpretable differential signals related to cold adaptation characteristics within the functional motif neighborhood while protecting the overall framework stability.

[0022] The construction rules for the three-level anchor point set are as follows: The set of anchor points for Class A is the set of conservative sites in the skeleton, A; the set of anchor points for Class A+ is the union of the set of anchor points for Class A and the stable background sites located within W; the set of anchor points for Class A++ is the union of the set of anchor points for Class A+ and the branching sites located within W.

[0023] The set of dual-layer joint anchor points is an A++ level set of dual-layer joint anchor points, A++_union, as shown in the following formula: A++_union=A++_fam ∪ A++_sub ∪ A+_union; A+_union=A+_fam ∪ A+_sub ∪ A_union; A_union = A_fam ∪ A_sub; Among them, A+_union is the set of A+ level two-layer joint anchor points, A_union is the set of A level two-layer joint anchor points, A_fam is the set of family level A-level anchor points, A+_fam is the set of family level A+ level anchor points, A++_fam is the set of family level A++ level anchor points, A_sub is the set of subset level A-level anchor points, A+_sub is the set of subset level A+ level anchor points, and A++_sub is the set of subset level A++ level anchor points.

[0024] In step (3), constructing a seed homology sub-library based on the obtained seed sequence includes: Using each seed sequence as the query center, iterative homology retrieval is performed in the background of homologous sequences that are compatible with its classification, thereby constructing a seed homology sub-library. The classification level of the homologous sequence background is higher than the genus-level classification to which the seed sequence belongs. In other words, cold adaptation characteristics are not used as a prerequisite for sequence screening, so as to ensure the phylogenetic representativeness of single seed homology statistical analysis.

[0025] In (3), the set of second-path seed homology core anchor points for generating seed sequences based on the seed homology sub-library includes: Based on the seed homology sub-library, seed homology multiple sequence alignment is generated. Then, based on the seed homology multiple sequence alignment, the alignment observation statistics and the column-level information of the seed-specific sequence spectrum statistical model are generated. Next, the loci are scored and sorted according to the alignment observation statistics and the column-level information of the seed-specific sequence spectrum statistical model to obtain the loci sorting results. Based on the loci sorting results, the set of second-path seed homology core anchor points corresponding to each seed sequence is generated.

[0026] The iterative homology search includes: An iterative homology search is performed in a sequence database compatible with the seed sequence classification, using the seed sequence as the query. High-confidence sequences are extracted from the hit results to form an initial homology set, and representative sampling is performed to reduce sampling bias and control sequence redundancy to construct a seed homology sub-library. Subsequently, a secondary refinement search and multi-sequence alignment are performed on the seed homology sub-library to obtain seed homology multi-sequence alignments for subsequent screening.

[0027] Representative sampling employs a genus-level redundancy removal strategy, retaining a maximum of a preset number of sequences for each genus; the upper limit of the number of sequences in the initial homologous set is set to a preset value to obtain a sufficient sortable candidate space while maintaining a controllable scale.

[0028] The seed-specific sequence spectrum statistical model is a sequence spectrum hidden Markov model; the matching state column of the model is used to form a unified core column coordinate system and to stably map column-level features back to the seed sequence residue coordinates.

[0029] The column-level information of the seed-specific sequence profile statistical model includes the model dominance bias (HMM_pmax) and the information content index (KL). HMM_pmax is the maximum emission probability among the 20 amino acids in the model column, characterizing the strength of the dominant residue preference. KL is the Kullback-Leibler divergence of the emission probability distribution of the model column relative to a uniform background distribution, characterizing the strength of the identifiable pattern. Each site's score is a weighted combination of the sequence profile statistical model information content and the alignment observation statistics, as shown in the following formula: Score=1.0×KL+0.5×HMM_pmax+0.4×(1-H_norm)+0.2×pmax+0.1×Occ.

[0030] Optionally, a two-level seed homologous core anchor point set is constructed, comprising a high-confidence core anchor point set and an extended core anchor point set. The high-confidence core anchor point set consists of sites ranking in the top first preset proportion in the site ranking results, and the extended core anchor point set consists of sites ranking in the top second preset proportion in the site ranking results, with the second preset proportion being greater than the first preset proportion. Either the high-confidence core anchor point set or the extended core anchor point set can be selected as the seed homologous core anchor point set for the second path according to actual needs.

[0031] II. A Thermosensitive UDG Constraint Anchor Point Identification Device Based on Dual-Path Evolutionary Analysis The seed sequence acquisition unit is used to acquire and record the type 1 thermosensitive UDG family sequence containing eukaryotic and prokaryotic sources as the seed sequence. The sequence library and subset construction unit is used to determine the preset lineage boundary based on the lineage source of the seed sequence, independently construct multiple cold adaptation sequence subsets according to the preset lineage boundary, and combine the cold adaptation sequence subsets after family gate screening to obtain the target family sequence library. The first anchor point set construction unit is used to generate the first path function motif constraint anchor point set of the seed sequence based on the target family sequence library and the cold adaptation sequence subset. The second anchor point set construction unit is used to construct a seed homology sub-library based on the obtained seed sequence, and then generate the second path seed homology core anchor point set of the seed sequence based on the seed homology sub-library. The third anchor point set construction unit is used to generate an evolutionary constraint anchor point set based on the first path functional motif constraint anchor point set and the second path seed homologous core anchor point set of the seed sequence.

[0032] III. A computer device The device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the thermal UDG constraint anchor point identification method based on dual-path evolution analysis.

[0033] IV. A computer-readable storage medium The medium stores a computer program, which, when executed by a processor, implements the steps of the thermal UDG constraint anchor point identification method based on dual-path evolutionary analysis.

[0034] V. A computer program product The product includes a computer program / instruction that, when executed by a processor, implements the steps of the thermal UDG constraint anchor point identification method based on dual-path evolutionary analysis.

[0035] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Sequence space boundary constraint and spectral balance design: The analysis target is precisely limited to the sequence space consistent with the catalytic core domain of the UDG family type 1 by the family gate mechanism. At the same time, multi-source thermosensitive seeds across prokaryotes and eukaryotes are used as analytical anchors to ensure the functional relevance of the results from the source of the method and effectively avoid single spectral bias. 2. A complementary fusion strategy of multi-level evolutionary information: The first path simultaneously captures cross-lineage skeleton conservation and cold adaptation branch difference signals at the family level, while the second path locks individualized core conserved features at the seed homology level. The fusion of the results of the two paths effectively overcomes the inherent information bias and statistical bias of the traditional single-path method. 3. Synergistic integration mechanism of prior functional knowledge and data-driven statistics: On the basis of pure data-driven statistical hierarchies, conservative functional motif neighborhoods are explicitly introduced as prior constraint windows, so that the final anchor point system can establish a traceable and verifiable correspondence with key functional regions such as catalysis, identification, and flexible regulation. 4. Standardized output interface for downstream engineering applications: The final output is directly mapped to the seed sequence residue coordinate system and presented in a structured format of site index and amino acid type, which can be used as a plug-and-play constraint input for generative protein design or directed evolutionary modification processes. 5. Openness and cross-sample transferability of the method: All core computational steps are implemented based on publicly available algorithms and general bioinformatics tools, ensuring the independent reproducibility of the method and supporting consistent transfer applications between different thermally sensitive UDG seed sequences. Attached Figure Description

[0036] To make the technical solution and beneficial effects of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings. The following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope of protection of the present invention.

[0037] Figure 1 This is a schematic diagram of the overall process of the thermal UDG dual-path evolution constraint anchor point identification method of the present invention; Figure 2 This is a schematic diagram showing the multi-source thermosensitive UDG seed sequence and temperature adaptation characteristics used in the embodiment. Figure 3 This is a schematic diagram illustrating the number of sequences retrieved from candidate sequences in Path 1 and grouped by cold-adapted relevant subsets after deduplication. Figure 4 This is a schematic diagram comparing the number of sequences in each group before and after the gate screening of the PF03167 family in Path 1 (based on HMM scanning); Figure 5This is a diagram showing the comparison of the number of sequences in each group before and after redundancy removal (CD-HIT clustering) based on the 90% sequence consistency threshold in Path 1; Figure 6 This is a schematic diagram of the preliminary classification results of the three levels (A / B / C) of the full family layer multiple sequence alignment in Path 1; where (A) is the number statistics of the three types of sites, and (B) is a two-dimensional distribution of residue conservation and branch difference and a schematic diagram of the classification threshold boundary. Figure 7 A schematic diagram showing the functional annotations and reference sources of the four types of conserved functional motifs of type 1 in the UDG family in Path 1; Figure 8 This is a schematic diagram showing the distribution of functional sequence constraint anchor points (Level A / A+ / A++) and functional sequence neighborhood windows in the alignment column coordinates for the entire family layer in Path 1. Figure 9 A schematic diagram showing the union results (A_union, A+_union, A++_union) of the anchor points of the whole family layer and the supplementary anchor points of the subset layer in Path 1, and their distribution on the column coordinates of the multi-sequence alignment in each cold adaptation branch subset layer; Figure 10 A schematic diagram showing the mapping of the final functional motif constraint anchor point (A++_union) of path one to the coordinates of residues in each thermal UDG seed sequence; where (A) is the distribution of anchor points on the four seed sequences, and (B) is an example of site labeling; Figure 11 This is a statistical comparison diagram of the size of the seed homology sub-database before and after the single-seed iterative homology retrieval refinement in Path 2. Figure 12 This is a schematic diagram comparing the size of the two-level seed homology core anchor point sets (CORE_HARD level / CORE_PLUS level) obtained by joint scoring and screening based on multiple sequence alignment and sequence spectrum hidden Markov model in Path 2. Figure 13 This is a schematic diagram showing the mapping of the extended core anchor points (CORE_PLUS level) in Path 2 to the coordinates of the residues in each thermal UDG seed sequence; where (A) is a comparison of the distribution of anchor points on the four seed sequences, and (B) is an example of site labeling. Figure 14 The diagram shows the dual-path fusion result and the output of the final evolutionary constraint anchor point; where (A) shows the distribution and overlap of A++_union and CORE_PLUS on the seed sequence residue coordinates, and (B) shows an example of the sequence labeling of the final evolutionary constraint anchor point (A++_union∪CORE_PLUS). Detailed Implementation

[0038] To make the technical solution and effects of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described below are only used to explain the present invention and do not constitute a limitation on the scope of protection of the present invention; other implementation methods obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the scope of protection of the present invention.

[0039] This embodiment uses thermosensitive UDGs as the research object, selecting four type 1 proteins of the thermosensitive UDG family with clear origins as seed sequences to conduct dual-pathway evolutionary constraint anchor site identification. The amino acid sequence information of the seed sequences can be directly obtained from publicly available literature and databases (UniProtKB, NCBI, etc.), and simultaneously covers prokaryotic sources (marine psychrophilic bacterium BMTU3346, Bacillus subtilis). Bacillus sp. HJ171) and eukaryotic sources (Atlantic cod) Gadus morhua、 rainbow trout Oncorhynchus mykiss This reduces single-lineage bias. The four seed sequences mentioned above play a dual role in this embodiment: serving as coordinate mapping references for the first path (family gate double-layer multi-sequence alignment) and as query centers for the second path (single-seed homology background analysis). The final anchor point sets obtained from both paths are output in the format of "seed sequence residue number (starting from 1) + corresponding amino acid single-letter code," for example, "22V" indicates that the 22nd residue of the seed sequence is valine (V).

[0040] See Figure 1 This embodiment's dual-path evolutionary constraint anchor point identification method constructs two complementary paths: "family-level two-layer multiple sequence alignment analysis" and "single-seed homology background analysis." The first path constructs full-family-level multiple sequence alignments within the sequence space defined by the UDG family type 1 gate, and further constructs subset-level multiple sequence alignments for cold adaptation-related lineages. This simultaneously captures two types of signals: "cross-lineage backbone conservation" and "cold adaptation branch differences," and combines this with a conserved functional motif neighborhood window to form a set of functional motif constraint anchor points that can be mapped to seed residue coordinates. The second path uses each seed sequence as the query center, constructing seed homology multiple sequence alignments within a broader range of classification homology sequences that match its lineage affiliation. It then integrates alignment observation statistics with column-level information from the sequence spectrum statistical model for joint weighted scoring, obtaining a set of core seed homology anchor points. Finally, the anchor point sets from the two paths are fused to output a final evolutionary constraint anchor point set that can be directly used for subsequent protein engineering or generative protein design.

[0041] 1. Example: Source and description of thermosensitive UDG seeds The specific sources of the seed sequences used in this embodiment are as follows: CodUDG (Atlantic cod UDG) corresponds to UniProtKB entry Q9I982; OmUDG (rainbow trout UDG) corresponds to NCBI protein entry XP_021477375.1; Bacillus_HJ171 UDG corresponds to NCBI protein entry 167077379; the sequence information of BMTU3346 UDG comes from publicly published research literature on thermosensitive UDGs of marine psychrophilic bacteria. See also... Figure 2 The seed sequences mentioned above all have clear biological origins, complete sequence information, and verifiable temperature adaptation characteristics, which is conducive to transforming the engineering goal of "thermal sensitivity / cold adaptation" into a traceable basis for evolutionary screening.

[0042] 2. First Path: Family Gate Double-Layer Multi-Sequence Alignment and Functional Sequence Constraint Anchor Point Identification See Figure 1 The core objective of the first approach is to construct a two-layer multiple sequence alignment framework. The cold adaptation branch subset layer aims to construct a representative subset of sequences related to low-temperature adaptation, used to capture differential signals between the cold adaptation branch and the entire family backbone. The entire family layer, composed of the merged subsets of each cold adaptation branch, aims to cover the sequence space of UDG family type 1 as much as possible, used to identify highly conserved catalytic core residues shared across species. To ensure the stability of subsequent multiple sequence alignment and statistical analysis, this embodiment performs the following preprocessing steps sequentially before constructing the multiple sequence alignment: constrained construction of the cold adaptation branch subset, family gate screening, remedial processing of anomalous subsets, sequence redundancy removal and balanced sampling, and merging and assembling of the entire family layer panels.

[0043] 2.1 Restricted Construction of Cold-Adapted Branch Subsets and Assembly of the Entire Family of Layered Panels (1) Restricted construction of cold-adapted branch subsets.

[0044] This embodiment constructs three independent subsets of cold adaptation-related sequences from publicly available protein databases by limiting specific taxonomic groups, ensuring that the boundaries of each subset are clear and reproducible. Each subset uses the family domain model corresponding to the Pfam database entry PF03167 as the entry criterion for the family gate, i.e., searching for candidate sequences containing PF03167 family domain annotations or matching the PF03167 family domain model. Simultaneously, to reduce interference from fragmented sequences or anomalous fusion proteins in subsequent alignments, a selection range is set for sequence length; this embodiment uses a length range of 150–400 amino acid residues, which can be adjusted appropriately for different subsets based on actual conditions. After deduplication of duplicate records, the initial candidate sets for each subset are obtained.

[0045] The constraints for the three types of cold adaptation-related subsets are as follows: a) Subset of marine psychrophilic γ-proteobacteria: preferably limited to γ-proteobacterial lineages and orders / families / genus that survive in low-temperature marine environments (e.g., Vibrionales , Alteromonadales , Oceanospirillales as well as Vibrio , Shewanella , Colwellia , Moritella , Pseudoalteromonas wait); b) A subset of cold-adapted or low-temperature-active Gram-positive bacteria: preferably representative genera limited to the Firmicutes or Actinobacteria phyla (e.g. Bacillus , Paenibacillus , Planococcus , Exiguobacterium , Carnobacterium , Listeria , Arthrobacter wait); c) A subset of cold-adapted fish: preferably limited to representative genera of cold-water or polar fish (e.g. Gadus , Oncorhynchus , Salmo , Notothenia wait).

[0046] like Figure 3 As shown, the initial candidate sequence counts for the three subsets are as follows: 2190 for the marine psychrophilic γ-proteobacteria subset, 1870 for the cold-adapted or low-temperature active Gram-positive bacteria subset, and 54 for the cold-adapted fish subset.

[0047] (2) Secondary confirmation and coverage screening of family gates.

[0048] Since relying solely on database retrieval criteria may still result in sequences with only local similarity or inconsistent domain boundaries, this embodiment performs secondary family domain verification on each subset of candidate sequences. This secondary verification employs a family domain scanning method based on a Hidden Markov Model (HMM), specifically implemented using the HMMER software suite. The PF03167 family domain model is used to scan each candidate sequence, obtaining a structured scan result that includes the hit significance index (E-value) and the start and end coordinates of the hit interval.

[0049] To ensure that sequences entering subsequent multiple sequence alignments have sufficiently complete core domains and reliable alignment, this embodiment introduces two coverage quality control metrics: sequence coverage, which measures the proportion of the family domain hit interval to the full length of the sequence; and model coverage, which measures the degree to which the sequence covers the PF03167 family domain model. The coverage threshold is set according to conventional empirical principles of family analysis: aiming to eliminate sequences that only hit short segments or only cover local regions of the model, thereby reducing the interference of weak hits on the stability of multiple sequence alignments. The screening criteria used in this embodiment are: sequence coverage ≥ 0.50 and model coverage ≥ 0.70.

[0050] like Figure 4 As shown, after the above-mentioned family gate screening, the number of sequences retained in this embodiment is as follows: 2190 sequences from the marine psychrophilic γ-proteobacteria subset, 1851 sequences from the cold-adapted or low-temperature active Gram-positive bacteria subset, and 32 sequences from the cold-adapted fish subset. It can be seen that a relatively abnormal "large proportion of deletions" occurred in the cold-adapted fish subset.

[0051] (3) Rescue the subset of cold-adapted fish that has been abnormally removed in large numbers.

[0052] Analysis revealed that fish UDG homologous sequences typically contain long non-core appendages or low-complexity regions, resulting in a systematic suppression of the proportion of core family domain hit intervals within the total length. Consequently, these sequences are mistakenly deleted during quality screening, where coverage is the primary metric. To avoid this systematic deletion problem, this embodiment employs a two-level remedial strategy for fish subsets to ensure that fish candidate sequences are not biased against due to differences in length and structure.

[0053] a) First Remedial Strategy – Functional Motif Anchor Plugging and Reversal: For fish candidate sequences that failed the gate screening, a conserved short peptide functional motif of UDG family type 1 was located in the sequence as a pruning anchor. In this embodiment, a conserved short peptide with “GQDPYH” as its core was used, allowing for conserved amino acid substitutions at individual sites (such as Y being replaced by F). A fragment covering the core domain was truncated around this anchor. In this embodiment, a truncation range of 35 residues upstream and 210 residues downstream of the anchor was retained. Subsequently, the truncated fragment was re-scanned for the PF03167 family domain for confirmation. After processing with this strategy, all 22 fish candidate sequences that originally failed the gate screening were recovered, restoring the fish subset from 32 to 54 sequences.

[0054] b) Second Remedial Strategy – Supplementing with a Fish-Specific Family Domain Model: After the above-mentioned retrieval, this embodiment further constructs a fish-specific family domain model using the confirmed fish core domain fragments to more stably determine the core domain boundaries of fish sequences. The specific operations are as follows: The family domain hit interval is extracted from the confirmed fish sequences, and 5 residues are extended at each end of the interval as safety boundaries; multiple sequence alignment is performed on the above fragments, and a fish-specific HMM model is trained based on the HMMER suite; simultaneously, the four thermosensitive seed sequences of this embodiment are added to the training set as anchoring references to ensure the consistency between the model and the seed coordinate system. After processing with this strategy, another 10 fish candidate sequences that previously failed the gate screening were retrieved. Combining the two remedial strategies, the cold-adapted fish subset is finally restored to a complete set of 54 sequences.

[0055] (4) Redundancy removal and balanced sampling.

[0056] To reduce bias in statistical analysis caused by closely related repeating sequences, this embodiment performs sequence deredundancy processing on each subset using the CD-HIT sequence clustering tool. CD-HIT's function is to cluster highly similar sequences into a single cluster while retaining representative sequences, given a sequence consistency threshold. This embodiment sets the sequence consistency threshold to 90% and adds alignment coverage constraints to avoid mis-clustering of short segments. Figure 5 As shown, after 90% redundancy removal, the number of sequences in each subset changed as follows: the marine psychrophilic γ-proteobacteria subset decreased from 2190 to 1190; the cold-adapted or low-temperature active Gram-positive bacteria subset decreased from 1847 to 789; and the cold-adapted fish subset (referred to as cold-water fish) decreased from 54 to 27.

[0057] Based on this, to form a controllable-scale sequence panel for multi-sequence alignment at the whole family level, this embodiment performs balanced sampling on each subset. Balanced sampling follows these principles: without sacrificing core domain quality, it reasonably allocates the proportion of each subset and covers as many different genus-level sources as possible to improve phylogenetic representativeness. The planned sampling quantity in this embodiment is: 100 sequences from the marine psychrophilic γ-proteobacteria subset, 40 sequences from the cold-adapted or low-temperature active Gram-positive bacteria subset, and 60 sequences from the cold-adapted fish subset; if the number of usable sequences in a subset after redundancy removal is less than the planned number, all sequences are included. The final actual number of sequences obtained is: 100 sequences from the marine psychrophilic γ-proteobacteria subset, 40 sequences from the cold-adapted or low-temperature active Gram-positive bacteria subset, and 27 sequences from the cold-adapted fish subset.

[0058] (5) Merging and assembling the core domain panels of the entire family layer.

[0059] This embodiment merges the three types of preprocessed cold-adapted branch subsets to form the whole family layer core domain panel. During merging, cross-subset deduplication is performed to ensure that only one sequence repeating in different subsets is retained. Furthermore, to ensure that subsequent anchor points can be mapped to the coordinates of the thermally sensitive seed sequences, this embodiment includes four thermally sensitive seed sequences (CodUDG, OmUDG, Bacillus_HJ171, BMTU3346) in the whole family layer panel, making them reference sequences to participate in multiple sequence alignment along with other family sequences. After merging, the whole family layer panel contains a total of 167 sequences (including the four seed sequences).

[0060] 2.2 Construction of whole-family layer multiple sequence alignment and classification of whole-family loci After completing the construction of the whole family core domain panel and the cold adaptation-related branch subset panel, this embodiment first constructs multiple sequence alignment for the whole family core domain panel, and performs column-level quality control and site classification on the alignment results to form a three-level site set of the whole family A / B / C, which lays the foundation for subsequent anchor point integration by combining conserved functional motif neighborhoods.

[0061] (1) Construction of full family layer multiple sequence alignment This embodiment uses the obtained whole-family core domain panel as input and performs multiple sequence alignment using the protein multiple sequence alignment program MAFFT. The function of MAFFT is to establish a column-by-column correspondence between multiple homologous amino acid sequences, ensuring that homologous positions fall on the same alignment column as much as possible, thus supporting subsequent column-by-column statistical analysis. To improve the alignment stability of cross-family sequences, this embodiment adopts a high-precision alignment strategy of "local pairing alignment + iterative refinement," with specific parameters set as follows: enable local pairing strategy (--localpair), set the number of iterations to 1000 (--maxiterate1000), and rearrange the output according to the input order (--reorder). The whole-family core domain panel already contains four thermosensitive seed sequences (CodUDG, OmUDG, Bacillus_HJ171, BMTU3346), which can be directly used as alignment input.

[0062] (2) Alignment trimming and quality control Because whole-family alignments include distantly related sequences, some alignment columns may exhibit excessively high gap ratios or unstable correspondences. This embodiment employs the trimAl alignment trimming program to perform column-level quality control on the whole-family alignment results. The trimAl program automatically removes alignment columns that are missing in most sequences and have low reliability based on the proportion of gap characters in each alignment column, thereby improving the reproducibility and interpretability of subsequent statistical analysis.

[0063] In this embodiment, column occupancy is calculated for each alignment column, defined as the proportion of non-gap characters in that column out of all sequences, and only alignment columns with a column occupancy ≥ 0.55 are retained. This threshold is achieved through the gap threshold parameter (-gt0.55) of trimAl. The rationale for setting this threshold is that when more than half of the sequences in a column are gaps, the residue correspondence in that column is easily dominated by the insertion or deletion events of a few sequences and becomes unstable. Therefore, retaining a column occupancy ≥ 0.55 can achieve a balance between removing obviously unreliable columns and retaining family common information.

[0064] (3) Calculation of column-level statistical characteristics Based on the full-family alignment results after column trimming, this embodiment calculates column-level statistical features for each alignment column to simultaneously characterize cross-family skeleton commonalities and branch differences. The following statistical indicators are all conventional column-level statistical methods in the field of sequence alignment analysis; this embodiment applies them to site classification and anchor point selection.

[0065] a) Column occupancy (Occ): Column occupancy can be calculated as "number of non-gap characters in the column / total number of sequences". The higher the column occupancy, the more comparable amino acid residues the column has in more sequences.

[0066] b) Column Conservation: In this embodiment, Shannon entropy is used to characterize the dispersion of the amino acid distribution in the comparison column. The specific calculation method is as follows: After removing missing characters, the frequency distribution of the 20 standard amino acids in the column is statistically analyzed, and the entropy H is calculated as H = -∑(p i ·log p i ), where p i Let H be the frequency of the i-th amino acid in this column. To ensure comparability between different alignment columns, this embodiment uses the theoretical maximum entropy log(20) (when the probability of occurrence of the 20 standard amino acids in a certain alignment column is completely equal, the dispersion of the column reaches its maximum) to normalize the information entropy H, resulting in the normalized entropy H_norm = H / log(20), with a value range of [0,1]. The column conservatism is defined as Cons = 1 - H_norm, and the larger the Cons value, the more conservative the residue composition of the column.

[0067] c) Branch divergence (Div): To identify the location where a cold-adaptation-related subset exhibits a significantly different amino acid preference relative to the rest of the sequence in a specific alignment column, this embodiment divides the whole family sequence into three established source subsets and compares the amino acid frequency distribution differences between a subset and the merged subsets in each alignment column. This embodiment uses Jensen-Shannon divergence (JSD) as a measure of distribution difference. In the same alignment column, the JSD values ​​of the three subsets relative to the rest of the sequence are calculated, and the maximum value is taken as the difference-pointing subset for that column, characterizing which type of cold-adaptation-related branch the difference signal in that column mainly originates from.

[0068] d) Dominant amino acid frequency pmax: defined as the frequency of the amino acid with the highest frequency in the alignment column, used to characterize the concentration of dominant residues; (4) Family A / B / C loci grading rules Based on three statistical measures—column occupancy, column conservatism, and branch difference—this embodiment divides the whole-family alignment columns into three levels of loci: A, B, and C. The specific classification rules are as follows: Category A loci (conserved loci of the whole family skeleton): The screening criteria are column occupancy ≥ 0.85 and column conservation ranking in the top 10 of all aligned columns.

[0069] Category C loci (branching difference loci): The selection criteria are a column occupancy ≥ 0.60 and a branching difference ranking in the top 10% of all aligned columns. If the number of Category C loci obtained according to this rule is insufficient, they are supplemented according to branching difference from high to low to a minimum of 30 columns or at least 18% of the total number of aligned columns (whichever is greater) to ensure sufficient coverage of cold adaptation branching difference signals. To avoid the same aligned column being classified into both Category A and Category C, this embodiment prioritizes classifying columns with higher differences into Category C, and then selects Category A from the remaining columns.

[0070] Category B loci (stable background loci): After excluding categories A and C, the selection criteria are column occupancy ≥ 0.75 and column conservatism not lower than the median of all aligned columns. These are used to provide a background reference set for statistical analysis.

[0071] Based on the above hierarchical rules, this embodiment obtains a complete family of A / B / C level three loci, and the quantity statistics are as follows: Figure 6 As shown in (A): there are 12 Class A loci, 26 Class B loci, and 134 Class C loci. Figure 6As shown in (B), all candidate sites are plotted as a two-dimensional distribution chart of conservation and branch dissimilarity (JSD), with the horizontal axis representing column conservation and the vertical axis representing JSD. The dashed line indicates the boundary of the hierarchical threshold used in this embodiment. Figure 6 As shown in (B), type A loci are mainly clustered in highly conserved, low-difference regions, reflecting stable skeletal characteristics shared across species; type B loci are mostly located in moderately to highly conserved, low-difference regions, serving as a background reference for subsequent statistical analysis; type C loci are mainly distributed in significantly different regions, indicating that these loci are more likely to be related to adaptive changes in specific cold adaptation branches. This embodiment also records the column number, column occupancy, column conservatism, branch dissimilarity, and dissimilarity-directing subset information for each alignment column in the A / B / C loci sets, laying the foundation for subsequent anchor point integration using conserved functional motif neighborhood windows.

[0072] 2.3 Cold Adaptation Branch Subset Layer Multi-Sequence Alignment and Functional Motif Constraint Anchor Point Output After completing the full family-level multiple sequence alignment and A / B / C site classification, this embodiment further verifies and supplements the cold adaptation-related branch subset layer, and integrates and classifies the site set by combining the conserved functional motif neighborhood window, thereby obtaining the first path functional motif constraint anchor point set that can be directly mapped to the thermally sensitive seed sequence residue coordinates. In this embodiment, the subset layer analysis and the full family layer share the same alignment column coordinate system, enabling the full family layer classification results, subset layer verification results, and subsequent seed mapping results to correspond to each other within the same coordinate framework, facilitating review and verification.

[0073] (1) Construction and column-level verification of multiple sequence alignments in cold-adapted branch subset layers Subset-level multiple sequence alignments are constructed as follows: In the full-family-level multiple sequence alignment results, corresponding subset sequence rows are extracted according to subset membership, and their alignment column numbers in the full-family-level are retained to obtain the subset-level multiple sequence alignments. This method ensures that the subset layer and the full-family-level naturally share the same alignment column-residue position coordinate system, avoiding difficulties in backtracking due to inconsistent column numbers caused by re-alignment. For subset-level multiple sequence alignments, this embodiment verifies the alignment columns according to column-level quality control rules, removing columns with excessively high gap ratios, only a few sequences participating in the alignment, or unstable alignments. The purpose of subset-level verification is to check the alignment stability of the core domain within a more closely related and homogeneous sequence set, reducing the interference of alignment drift caused by excessively large full-family spans on the interpretation of differential sites, and laying the foundation for subsequent functional motif window localization, anchor point integration, and seed mapping.

[0074] (2) Determination of the neighborhood window of the conservative functional motif To establish a verifiable correspondence between anchor points and key functional regions of the UDG family type 1, this embodiment locates conserved functional motifs shared across lineages on the whole-family consensus sequence or thermally sensitive seed sequence, and constructs a functional motif neighborhood window set W centered on the functional motif. This embodiment identifies the following four types of functional motif windows, the structure and function of which are based on [reference needed]. Figure 7 : QDPYH functional motif window: This is the region containing the "QDPY" fragment and where "H" appears within 0-4 residues thereafter; GV(L / M)LLN functional motif window: Positions a conservative segment that begins with "GV", is followed by two "L" or "M", and ends with "LN". HPSPL[AS] Functional sequence window: Positions a conservative segment that begins with "HPSP", is followed by L or I, and ends with A or S. Pro-rich functional motif window: Defined as the interval between the QDPYH functional motif and the HPSPL[AS] functional motif, prioritizing the identification of short fragments containing continuous proline features (such as PPPPS, PPPSL, KPPSL, PPAS*, PPHSL, etc.); if the priority mode is not hit, candidate fragments with a length of 6, a proline number ≥2 and an ending position of S / A / H are selected within the interval, and the final fragment is determined by combining the degree of proline enrichment and the proximity to the target distance.

[0075] The window is formed by extending a fixed number of residues to both sides of the functional motif fragment, covering the functional motif itself and its adjacent key structural regions. The extension length remains consistent in both the whole family layer and the subset layer to ensure the comparability of window positions between different seeds. In this embodiment, after locating four types of functional motifs on the whole family consensus sequence, the window extends three residues to both sides and is mapped to the whole family layer alignment column coordinate system. The specific range of the window set W in the alignment columns is: columns 17-27 of the QDPYH window, columns 37-48 of the proline-rich window, columns 71-82 of the GV(L / M)LLN window, and columns 140-151 of the HPSPL[AS] window.

[0076] (3) Functional sequence constraint re-grading of the whole family layer A / B / C sites to A / A+ / A++ anchor system This embodiment uses the obtained family-wide A / B / C locus classification results as a basis, and combines them with the aforementioned functional motif neighborhood window set W to further classify the locus set using functional motif constraints, forming a three-level anchor point system. For ease of description, the A / B / C locus sets are denoted as A, B, and C, respectively. The construction rules for the three-level anchor points are as follows: Grade A (Conservative anchor points of the whole family skeleton): Directly take the set of Class A sites to constrain the catalytic core and structural skeleton shared across the spectrum; A+ level (functional motif neighborhood stable anchor point): Based on the A level, B type sites located within the functional motif window W are incorporated, that is, A+ = A ∪ (B ∩ W); A++ level (functional motif neighborhood adaptation anchor point): Based on the A+ level, the C-class site located within the functional motif window W is incorporated, that is, A++ = A+ ∪ (C ∩ W).

[0077] In this embodiment, the number of anchor points for the entire family layer obtained after the above re-grading are as follows: 12 columns for Grade A, 16 columns for Grade A+, and 33 columns for Grade A++. Their distribution on the comparison column coordinates is as follows: Figure 8 As shown.

[0078] (4) Subset layer supplementation and determination of dual-layer joint anchor points and seed mapping output After obtaining the three-level anchor point system of the entire family layer (A / A+ / A++), this embodiment further supplements the anchor point set with the subset layer verification results and forms a joint set, which is then mapped to the coordinate output of each thermal seed sequence. For ease of description, the following notation is used: the three-level anchor point set of the entire family layer is denoted as A_fam, A+_fam, and A++_fam, respectively; the three-level anchor point set obtained by verification and supplementation of a certain cold-adapted subset layer is denoted as A_sub, A+_sub, and A++_sub, respectively; and the joint set after merging the two sets is denoted as A_union, A+_union, and A++_union, respectively.

[0079] a) Subset layer differential site supplementation and final joint anchor point set determination In this embodiment, for subset-level alignment, the column occupancy and conservatism of each aligned column are calculated, and combined with the branch difference of that column to the target subset, a set of subset-level difference supplementary sites is obtained. To reduce statistical fluctuations caused by gaps and small samples, a basic quality threshold is set during screening: column occupancy ≥ 0.60 and column conservatism ≥ 0.20. Among the columns that pass the quality threshold, columns with difference intensity in the top 15% are selected, and the difference is required to be more biased towards this subset relative to the other sequences (i.e., positive difference). Columns already covered by A+_fam are removed to reduce duplication. The above thresholds are conventional empirical settings in column-level statistical analysis of multiple sequence alignments, used to reduce the risk of misjudgment caused by gaps and random fluctuations while retaining interpretable difference signals.

[0080] For each cold-adapted subset, this embodiment performs a union of the anchor points of the entire family layer and the supplementary anchor points of the subset layer, while maintaining the hierarchical nesting relationship between the sets. The construction rules are as follows: A_union = A_fam ∪ A_sub A+_union=A+_fam ∪ A+_sub ∪ A_union A++_union=A++_fam ∪ A++_sub ∪ A+_union The distribution of the joint results on the alignment column coordinates is as follows Figure 9 As shown.

[0081] In this embodiment, the number of joint anchor points corresponding to the three types of cold adaptation subsets are as follows: Subsets of marine psychrophilic γ-proteobacteria: A_union has 18 columns, A+_union has 26 columns, and A++_union has 52 columns; Subsets of cold-adapted or low-temperature active Gram-positive bacteria: A_union has 13 columns, A+_union has 21 columns, and A++_union has 45 columns; Subsets of cold-adapted fish: A_union has 19 columns, A+_union has 25 columns, and A++_union has 54 columns.

[0082] b) Mapping and outputting the joint anchor point to thermal seed coordinates After obtaining the joint anchor point set for each subset, this embodiment maps the anchor points from the alignment column coordinates to the residue number coordinates of each thermosensitive seed sequence, recording the residue numbers and amino acid types belonging to the joint anchor points in each subsequence. This embodiment uses the A++_union hierarchy, which balances stable mapping of key functional regions with preservation of cold adaptation-related differential signals, as the first path output set. The specific results after mapping to each thermosensitive seed are as follows: BMTU3346 has 52; Bacillus_HJ171 has 45; CodUDG and OmUDG both have 54. Figure 10 As shown in (A), the distribution of A++_union sites of the four thermosensitive seeds on their respective sequence coordinates; as Figure 10 As shown in (B), this is an example of the annotation of A++_union sites in each thermosensitive seed sequence. The above mapping output constitutes the functional motif constraint anchor point results for the first path, which serves as the input for subsequent fusion with the homologous core anchor points of the second path seed.

[0083] 3. Second Path: Single Seed-Driven Evolutionary Constraint Anchor Point Identification 3.1 Construction of Single-Seed Homologous Sublibraries The first approach focuses on characterizing the commonalities of the entire family and the differences in cold-adaptation branches. However, some sites that are more constraining for a particular thermosensitive UDG seed are easily averaged out and not highlighted in the whole-family statistics. To further characterize the conservative patterns of a single seed in its direct homology background, this embodiment designs a second approach: using each thermosensitive UDG seed sequence as the query center, homologous sequences are obtained within its broad and class-compatible homology background, and a seed homology sub-library is constructed. This provides a more stable and verifiable set of comparison inputs for subsequent seed-level core anchor point identification.

[0084] (1) Selection of classification database To ensure that the homology background is both broad enough and matches the seed classification, this embodiment selects different taxonomic databases for different seed sequences: Bacillus_HJ171 uses the Swiss-Prot high-quality reference library; BMTU3346 uses the Actinobacteria related library to cover its bacterial lineage homology diversity; and CodUDG and OmUDG use the Vertebrate related library to cover a wider range of fish and closely related vertebrate homology sets. The above selection logic is not limited to a specific genus-level classification, but deliberately maintains a broad phylum range to avoid insufficient homology coverage due to an overly narrow database.

[0085] (2) Iterative homology retrieval and representative sampling After the database is determined, this embodiment uses the Jackhmmer tool from the HMMER suite for iterative homology retrieval, with 5 iteration rounds. The top 5000 high-confidence sequences are extracted from the hit results as the initial homology set, followed by genus-level deduplication (keeping a maximum of 5 sequences per genus) to reduce sampling bias and control redundancy, ultimately forming a seed homology sub-database. For example... Figure 11 As shown, the homologous sub-library construction results for various sub-sub ...

[0086] (3) Secondary Jackhmmer refining To obtain high-quality alignment results that more closely resemble the local structure of the seeds, this embodiment further performs Jackhmmer refinement search and multiple sequence alignment construction on the seed homologous sub-library to obtain seed homologous multiple sequence alignments for subsequent column-level statistics and model analysis.

[0087] 3.2 Seed homologous multiple sequence alignment In the alignment of the obtained seed homologous multiple sequences, this embodiment calculates the column-level features of the alignment columns, including: column occupancy (Occ), dominant amino acid frequency (pmax), and normalized entropy (H_norm). The calculation method of the above statistics is consistent with the column-level statistical method, so as to compare the measurement scope with the results of the first path.

[0088] 3.3 Seed-specific statistical model and determination of core anchor points (1) Construction of seed-specific statistical model and unified column coordinates To improve the robustness of candidate selection and obtain a unified, mappable column coordinate system, this embodiment constructs a seed-specific sequence profile Hidden Markov Model (HMM) based on seed homology multiple sequence alignment. The model's matching state columns (i.e., the core column set used to express homology correspondences) can form a unified core column coordinate system, which is used to stably map column-level features back to the seed sequence. Model quality testing showed that the column mapping coverage of various sub-models reached a high level (above 97%), indicating complete model coverage and reliable mapping.

[0089] (2) Representation of information content at the column level of the model This embodiment extracts the emission probability distribution of each core column for 20 amino acids from the statistical model and constructs the following information content and preference characterization indicators: a) Model Dominance Preference (HMM_pmax): The maximum emission probability, used to characterize the strength of the dominant residue preference. The larger the HMM_pmax, the more dominant residues exist in that column. b) Information content index (KL): Kullback-Leibler divergence relative to a relatively uniform background distribution, used to characterize the intensity of the recognizable pattern in the column relative to the background. The larger the KL, the richer the information content and the more recognizable the pattern in the column.

[0090] The above indicators are commonly used metrics in the fields of information theory and sequence model analysis, and can form complementary evidence with observational statistics of multiple sequence alignments.

[0091] (3) Joint scoring and two-level core set Within the candidate site set, this embodiment uses a joint weighted scoring and ranking of multiple sequence alignment observation statistics and the column-level information content of the sequence spectrum Hidden Markov Model to form a unified, screenable benchmark. The joint score can be weighted by an information content term, a model preference term, and a multiple sequence alignment conservatism / coverage term to simultaneously suppress low-coverage columns and preferentially retain columns with high information content and strong consistency. The scoring formula used in this embodiment is: Score=1.0×KL+0.5×HMM_pmax+0.4×(1-H_norm)+0.2×pmax+0.1×Occ The weighting is based on the following: KL directly measures information content and serves as the sovereign weight; HMM_pmax reflects the model's dominant residue preference and serves as the secondary sovereign weight; the conservation of multiple sequence alignment sides and pmax provide supplementary observational support; and Occ is used to suppress low-coverage columns. This scoring method unifies evidence from different sources onto the same scale, facilitating ranking and screening within the candidate set.

[0092] Regarding the sorting results, this embodiment defines a two-level seed homologous core anchor point set: a) High-confidence core anchors (CORE_HARD level): The top 35% of sites in the candidate set, used to provide core constraints with high consistency; b) Expanded core anchor points (CORE_PLUS level): The top 55% of sites in the candidate set, which improve coverage while maintaining reliability and reflecting functional diversity.

[0093] like Figure 12 As shown, the size of the two-level set in the embodiment is: a) CORE_HARD level: Bacillus_HJ171=15, BMTU3346=31, CodUDG=32, OmUDG=30; b) CORE_PLUS level: Bacillus_HJ171=24, BMTU3346=48, CodUDG=49, OmUDG=47.

[0094] The final output scheme for the second path: In this embodiment, the CORE_PLUS level is used as the set of seed homologous core anchor points for the second path. The selection criteria are: the CORE_PLUS level, while including high-confidence cores of the CORE_HARD level, can provide more sufficient site coverage while maintaining a controllable size, making it more suitable as a reliable core constraint on the seed side when fusing with anchor points of the first path. The CORE_PLUS level set is mapped to each thermosensitive seed according to the alignment column-seed residue number mapping rule. The specific results are as follows: BMTU3346 has 48, Bacillus_HJ171 has 24, CodUDG has 49, and OmUDG has 47. Figure 13 As shown in (A), the distribution of CORE_PLUS level sites of the four heat-sensitive seeds on their respective sequence coordinates; as Figure 13 As shown in (B), this is an example of the annotation of CORE_PLUS level sites in each thermosensitive seed sequence. The above mapping output constitutes the seed homology core anchor location results for the second path, which serves as the input for subsequent fusion with the functional motif constraint anchor locations of the first path.

[0095] 4. Dual-path fusion and final evolution constraint anchor point output This embodiment combines the set of functional motif constraint anchor points (A++_union) output from the first path with the set of homologous core anchor points (CORE_PLUS level) output from the second path in the seed sequence coordinate system to obtain the final set of evolutionary constraint anchor points. The A++_union set provides anchor points related to family-level backbone constraints and functional motif neighborhood adaptation information; the CORE_PLUS level set provides anchor points related to stable core constraints in a single-seed homology context. Through this union fusion, the final set of anchor points can simultaneously cover key constraint sites shared across lineages and stable core sites at the seed level, making it suitable as constraint input for subsequent protein design or engineering. The technical solution of this invention focuses on the identification and standardized output of evolutionary constraint anchor points, without limiting the specific implementation of subsequent design algorithms or sequence generation methods.

[0096] like Figure 14 As shown in (A), the distribution relationship of A++_union sites, CORE_PLUS level sites, and overlapping sites of the four thermosensitive seeds on their respective sequence coordinates is used to demonstrate the fusion logic of dual-path complementarity and overlap-enhanced confidence; as Figure 14 As shown in (B), examples of the final anchor points in each thermosensitive seed sequence are used to visually demonstrate the distribution of the final sites on the seed sequence.

[0097] For ease of use, this embodiment outputs the site index and its corresponding amino acid type in the format of "residue number + amino acid single-letter code", and also provides the number of sites and their proportion of the total sequence length. The final evolutionary constraint anchor point set for each type of protein in this embodiment is as follows: (1) The seed sequence of Bacillus_HJ171 is 245 aa in length, with 57 final anchor sites, accounting for 23.27% of the total sequence length. The sites are: 22V, 60P, 79V, 82L, 83G, 84Q, 85D, 86P, 87Y, 88H, 93A, 94M, 95G, 97S, 99S, 104I, 105P, 106K, 107P, 108P, 109S, 110L, 112N, 118A, 128H, 130D, 131L, 138G, 139V, 140L, 141L, 142L, 143N, 148V, 150E, 153P, 156H, 175Q,177E, 184W, 185G, 186S, 188A, 191K, 202I, 206V, 207H, 208P, 209S, 210P, 211L,212A, 215R, 217G, 219F, 222K, 228N.

[0098] (2) The BMTU3346 seed sequence is 227 aa long, with 80 final anchor sites, accounting for 35.24% of the total sequence length. The sites are: 32R, 34V, 36G, 37E, 41P, 43P, 48R, 49S, 53P, 62L, 63G, 64Q, 65D, 66P, 67Y, 68P, 69T, 70P, 71G, 74I, 75G, 77S, 84V, 86P, 87L, 88P, 89R, 90S, 93N, 97E, 98L, 99S, 101D, 109H, 110G, 111D, 113T, 119G, 120V, 121L. 122M, 123L,124N, 125R, 128T, 129V, 130R, 131A, 133A, 137H, 141G, 143E, 146T, 149A, 152A,155A, 156R, 159P, 165W, 166G, 167N, 169A, 187S, 188P, 189H, 190P, 191S, 192P, 193L, 194S, 195A, 197R, 198G, 200F, 201E, 202S, 204P, 206S, 209N, 216G.

[0099] (3) The CodUDG seed sequence is 268 aa long, with 82 final anchor sites, accounting for 30.60% of the total sequence length. The sites are: 58A, 59A, 62E, 69L, 70M, 76E, 80H, 81T, 85P, 93T, 97D, 106L, 107G, 108Q, 109D, 110P, 111Y, 112H, 113G, 114P, 115N, 118H, 119G, 121C, 125Q, 126K, 128V, 129P, 130P, 131P, 132P, 133S, 136N, 140E, 144D, 149K, 153H, 154G, 155D, 159W,163G, 164V, 165L, 166L, 167L, 168N, 169A, 172T, 173V, 174R, 175A, 182K, 185G,187E, 190T, 193V, 200N, 203G, 209W, 210G, 213A, 228L, 229Q, 230A, 231V, 232H,233P, 234S, 235P, 236L, 237S, 238A, 240R, 241G, 243L, 244G, 245C, 247H, 249S, 252N, 259G, 262P.

[0100] (4) The OmUDG seed sequence is 267 aa long, with 81 final anchor sites, accounting for 30.34% of the total sequence length. The sites are: 43N, 47A, 50E, 56L, 62K, 63P, 75E, 79N, 84P, 96D, 105L, 106G, 107Q, 108D, 109P, 110Y, 111H, 112G, 113P, 114N, 117H, 118G, 120C, 124Q, 125R, 127I, 128P, 129P, 130P, 131P, 132S, 135N, 139E, 143D, 148Q, 152H. 153G, 154D, 158W, 162G,163V, 164L, 165L, 166L, 167N, 168A, 171T, 172V, 173R, 174A, 181K, 184G, 186E,189T, 192V, 199N, 202G, 208W, 209G, 212A, 227L, 228P, 229T, 230V, 231H, 232P, 233S, 234P, 235L, 236S, 237A, 239R, 240G, 242F, 243G, 244C, 246H, 248S, 251N, 258G, 261P.

[0101] The computational process in this embodiment can be implemented by a computer program. Specifically, a script can be written in Python in a Linux server environment, using pandas for tabular data processing, numpy for numerical calculations, matplotlib for visualization, the re and collections modules for functional sequence matching and frequency statistics, and pathlib for path management; multiple sequence alignment is completed by calling the MAFFT program or its equivalent; alignment column pruning is completed by implementing column filtering rules equivalent to trimAl; sequence redundancy removal is performed by calling CD-HIT; iterative homology retrieval and sequence spectrum statistical model construction are performed by calling the HMMER suite.

[0102] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for identifying thermally sensitive UDG constraint anchor points based on dual-path evolutionary analysis, characterized in that, Includes the following steps: (1) Obtain the type 1 sequence of the thermosensitive UDG family containing eukaryotic and prokaryotic sources and record it as the seed sequence; (2) Construct a subset of cold adaptation sequences based on the obtained seed sequences, and assemble each subset of cold adaptation sequences into a target family sequence library; then generate a set of first path function motif constraint anchor points for the seed sequences based on the target family sequence library and the subset of cold adaptation sequences. In step (2), based on the target family sequence library and the cold-adapted sequence subset, a set of first path functional motif constraint anchor points for the seed sequence is generated, including: Construct a full family layer multi-sequence alignment on the target family sequence library, and extract the sequence rows corresponding to each cold-adapted sequence subset from the full family layer multi-sequence alignment results to form a subset layer multi-sequence alignment corresponding to that cold-adapted sequence subset, so that the full family layer and the subset layer share the same alignment column coordinate system; Based on whole-family layer multiple sequence alignment, the loci are classified according to column-level statistical features to obtain family-level locus classification results including backbone conserved loci, branch-differentiated loci, and background stable loci; based on subset layer multiple sequence alignment corresponding to each cold-adapted sequence subset, the loci of each cold-adapted sequence subset are classified according to column-level statistical features to obtain subset layer classification results for each cold-adapted sequence subset. By combining the functional motif neighborhood window set W, the functional motif constraints of the family-level site classification results and the subset-level classification results are further classified to obtain the family-level three-level anchor point set and the subset-level three-level anchor point set corresponding to each cold-adapted sequence subset. Then, the family-level three-level anchor point set and the subset-level three-level anchor point set corresponding to each cold-adapted sequence subset are merged according to hierarchical categories to generate the two-level joint anchor point set corresponding to each cold-adapted sequence subset. According to the phylogenetic affiliation of each seed sequence, the two-level joint anchor point set of its corresponding cold-adapted sequence subset is mapped to the residue coordinates of the seed sequence to obtain the first-path functional motif constraint anchor point set corresponding to each seed sequence. (3) Construct a seed homology sub-library based on the obtained seed sequence, and then generate a set of second path seed homology core anchor points based on the seed homology sub-library; In (3), the set of second-path seed homology core anchor points for generating seed sequences based on the seed homology sub-library includes: Based on the seed homology sub-library, seed homology multiple sequence alignment is generated. Then, based on the seed homology multiple sequence alignment, the alignment observation statistics and the column-level information of the seed-specific sequence spectrum statistical model are generated. Next, the loci are scored and sorted according to the alignment observation statistics and the column-level information of the seed-specific sequence spectrum statistical model to obtain the loci sorting results. Based on the loci sorting results, a set of second-path seed homology core anchor points corresponding to each seed sequence is generated. (4) Generate an evolutionary constraint anchor point set based on the first path functional motif constraint anchor point set and the second path seed homologous core anchor point set.

2. The thermal UDG constraint anchor point identification method based on dual-path evolutionary analysis according to claim 1, characterized in that, In step (2), the construction of a target family sequence library and a cold-adapted sequence subset based on the obtained seed sequences includes: Based on the pedigree origin of the seed sequence, the preset pedigree boundary is determined. The UDG family type 1 family domain model is used as the sequence admission criterion. Multiple cold-adapted sequence subsets are independently constructed according to the preset pedigree boundary. After family gate screening of each cold-adapted sequence subset, the selected subsets are merged and assembled into the target family sequence library.

3. The thermal UDG constraint anchor point identification method based on dual-path evolutionary analysis according to claim 1, characterized in that, The set of functional motif neighborhood windows W is obtained by locating conserved functional motifs shared across lineages on the type 1 primary sequence of the UDG family. The conserved functional motifs include sequence segments related to catalytic active centers, structurally flexible ring regions, and substrate recognition interfaces. The functional motif neighborhood window is formed by extending a predetermined number of residues to both sides of the functional motif segment.

4. The thermal UDG constraint anchor point identification method based on dual-path evolutionary analysis according to claim 1, characterized in that, In step (3), constructing a seed homology sub-library based on the obtained seed sequence includes: Using each seed sequence as the query center, iterative homology retrieval is performed in the background of homologous sequences that are compatible with its classification, thereby constructing a seed homology sub-library. The classification level of the homologous sequence background is higher than the genus-level classification to which the seed sequence belongs.

5. A thermally sensitive UDG constraint anchor point identification device based on dual-path evolutionary analysis, characterized in that, include: The seed sequence acquisition unit is used to acquire and record the type 1 thermosensitive UDG family sequence containing eukaryotic and prokaryotic sources as the seed sequence. The sequence library and subset construction unit is used to determine the preset lineage boundary based on the lineage source of the seed sequence, independently construct multiple cold adaptation sequence subsets according to the preset lineage boundary, and combine the cold adaptation sequence subsets after family gate screening to obtain the target family sequence library. The first anchor point set construction unit is used to generate the first path function motif constraint anchor point set of the seed sequence based on the target family sequence library and the cold adaptation sequence subset. Specifically, it includes: Construct a full family layer multi-sequence alignment on the target family sequence library, and extract the sequence rows corresponding to each cold-adapted sequence subset from the full family layer multi-sequence alignment results to form a subset layer multi-sequence alignment corresponding to that cold-adapted sequence subset, so that the full family layer and the subset layer share the same alignment column coordinate system; Based on whole-family layer multiple sequence alignment, the loci are classified according to column-level statistical features to obtain family-level locus classification results including backbone conserved loci, branch-differentiated loci, and background stable loci; based on subset layer multiple sequence alignment corresponding to each cold-adapted sequence subset, the loci of each cold-adapted sequence subset are classified according to column-level statistical features to obtain subset layer classification results for each cold-adapted sequence subset. By combining the functional motif neighborhood window set W, the functional motif constraints of the family-level site classification results and the subset-level classification results are further classified to obtain the family-level three-level anchor point set and the subset-level three-level anchor point set corresponding to each cold-adapted sequence subset. Then, the family-level three-level anchor point set and the subset-level three-level anchor point set corresponding to each cold-adapted sequence subset are merged according to hierarchical categories to generate the two-level joint anchor point set corresponding to each cold-adapted sequence subset. According to the phylogenetic affiliation of each seed sequence, the two-level joint anchor point set of its corresponding cold-adapted sequence subset is mapped to the residue coordinates of the seed sequence to obtain the first-path functional motif constraint anchor point set corresponding to each seed sequence. The second anchor point set construction unit is used to construct a seed homology sub-library based on the obtained seed sequence, and then generate the second path seed homology core anchor point set of the seed sequence based on the seed homology sub-library. Specifically, it includes: Based on the seed homology sub-library, seed homology multiple sequence alignment is generated. Then, based on the seed homology multiple sequence alignment, the alignment observation statistics and the column-level information of the seed-specific sequence spectrum statistical model are generated. Next, the loci are scored and sorted according to the alignment observation statistics and the column-level information of the seed-specific sequence spectrum statistical model to obtain the loci sorting results. Based on the loci sorting results, a set of second-path seed homology core anchor points corresponding to each seed sequence is generated. The third anchor point set construction unit is used to generate an evolutionary constraint anchor point set based on the first path functional motif constraint anchor point set and the second path seed homologous core anchor point set of the seed sequence.

6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the thermal UDG constraint anchor point identification method based on dual-path evolutionary analysis as described in any one of claims 1 to 4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the thermal UDG constraint anchor point identification method based on dual-path evolutionary analysis as described in any one of claims 1 to 4.

8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the thermal UDG constraint anchor point identification method based on dual-path evolutionary analysis as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Chrysocystin species genome Homebox gene family identification method and device, electronic equipment and storage medium

    CN118918963A