Screening method for rational mutation sites of heat-sensitive udg based on strain culture temperature

By employing a thermosensitive UDG rational mutation site screening method based on strain culture temperature, combined with multi-source evidence from sequence and structure, the problems of site selection uncertainty and redundancy in existing methods are solved, achieving efficient and reproducible UDG modification site screening, which is applicable to the engineering optimization of different UDG backbones.

CN122337320APending Publication Date: 2026-07-03PULUOMAIGE BIOLOGICAL PRODS SHANGHAI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PULUOMAIGE BIOLOGICAL PRODS SHANGHAI
Filing Date
2026-04-02
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing methods for screening rational mutation sites in thermosensitive uracil-DNA glycosylase (UDG) rely excessively on empirical judgment or large-scale random mutation screening. They lack a systematic control data framework driven by temperature adaptability, making it difficult to simultaneously consider the intensity of temperature response effects and the safety of catalytic function. Candidate sites tend to exhibit redundant aggregation in three-dimensional space, resulting in insufficient reproducibility of screening results and inadequate ability to migrate across the UDG backbone.

Method used

The method for screening rational mutation sites of thermosensitive UDG based on strain culture temperature integrates temperature difference characteristics at the sequence level and multi-source evidence at the structural level to output a set of preferred mutation sites for the engineering modification of thermosensitive UDG. This includes constructing a control dataset of heat-resistant group and cold-adapted group, combining temperature difference characteristics at the sequence level and temperature difference characteristics at the structural level, and performing multi-dimensional evaluation and cluster analysis to output a small number of preferred mutation sites.

Benefits of technology

It significantly narrows the experimental exploration space, improves the reproducibility and interpretability of site screening results, enhances the mobility across the UDG backbone, provides a systematic site optimization method, and reduces the cost and time investment of random mutation screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122337320A_ABST
    Figure CN122337320A_ABST
Patent Text Reader

Abstract

This invention discloses a method for screening thermosensitive UDG rational mutation sites based on strain culture temperature. The method first divides strains into heat-tolerant and cold-adapted groups according to temperature thresholds based on the culture temperature of the strain preservation database, constructing corresponding UDG sequence sets and three-dimensional structure sets. Then, a first candidate mutation site set is generated through multiple sequence alignment and amino acid distribution difference analysis. A second candidate mutation site set is generated through rigid body alignment and functional protection rules under a unified coordinate system. Finally, multiple-source indicators such as structural dispersion, contact and salt bridge networks, and kinetic coupling are calculated for the second candidate mutation site set. Combined with effect-safety dual-dimensional scoring, spatial clustering, and cross-validation of the first candidate mutation site set, a preferred mutation site set is output. This invention achieves a systematic transformation from temperature ecotags to engineered sites, improving the reproducibility and interpretability of site screening without requiring a large-scale mutation library.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of bioinformatics, protein engineering and computational structural biology, and specifically relates to a method for screening rational mutation sites of thermosensitive uracil-DNA glycosylase (UDG) based on ecological data of strain culture temperature. Background Technology

[0002] Uracil-DNA glycosylase (UDG) is a key initiating enzyme in the Base Excision Repair (BER) pathway. It specifically recognizes and removes mismatched uracil bases in the DNA strand, forming apurinic / apyrimidinic sites (AP sites), thereby initiating downstream repair mechanisms. In nucleic acid amplification and molecular detection systems, UDG is widely used as a contamination control tool: by introducing uracil-containing nucleotides to replace thymine in the reaction system (e.g., dUTP instead of dTTP), and adding UDG before a new round of amplification, previously amplified products containing uracil can be selectively degraded, effectively preventing false positive results caused by cross-contamination. Such applications place clear engineering requirements on the temperature behavior of UDG: it must maintain sufficient activity during pretreatment or low-temperature reaction stages to complete contamination removal, while rapidly inactivating or significantly reducing its activity before initiating subsequent high-temperature amplification steps to avoid continuous degradation of newly synthesized products, which could lead to decreased amplification efficiency or abnormal detection results. Therefore, the "thermal UDG" with predictable thermal deactivation characteristics has become a key functional element in such systems.

[0003] Currently available thermosensitive UDGs mainly originate from natural variants of low-temperature adapted microorganisms or some eukaryotic UDGs and their recombinant expression products. However, due to the limited availability of natural sources, the cost of expression and purification, and batch stability, relying solely on natural resources is insufficient to meet the diverse needs of different application scenarios regarding heat inactivation temperature windows, residual activity levels, and activity-stability balance. Therefore, it is necessary to utilize protein engineering techniques to rationally modify and functionally optimize UDGs that already possess mature expression and purification systems and application validation foundations. In protein engineering practice, the selection of mutation sites is a core factor determining the efficiency and success rate of modification: overly conservative site selection makes it difficult to effectively regulate thermosensitivity; overly aggressive site selection can easily disrupt the catalytic active site, substrate recognition interface, or overall folding stability, leading to significant loss of enzyme activity or even expression failure.

[0004] Existing mutation site selection strategies mainly fall into two categories: one is based on a small number of known samples, using empirical rules, local structural analysis, or single-sequence conservation statistics to determine candidate sites; the other is to construct large-scale mutation libraries through random mutation or directed evolution for high-throughput screening. However, random mutation and directed evolution paths typically require significant experimental investment and screening costs, and the resulting mutation combinations are highly dependent on the starting template and screening conditions, making them difficult to transfer and apply across different UDG backbones. Furthermore, rational design methods based on a single information source (e.g., relying solely on sequence conservation analysis or static structural features) often struggle to simultaneously consider both the "intensity of temperature response effect" and "catalytic functional safety," easily leading to problems such as redundant spatial distribution of candidate sites, insufficient reproducibility of results, and unstable cross-UDG backbone migration effects. Especially for thermosensitivity, a complex trait involving overall conformational flexibility, local interaction networks, and long-range dynamic coupling, it is difficult to reliably determine the priority of site modification based on single-dimensional evidence alone. Summary of the Invention

[0005] The purpose of this invention is to address the following technical problems in the existing screening process for rational mutation sites in thermosensitive uracil-DNA glycosylase (UDG): site selection relies excessively on empirical judgment or large-scale random mutation screening, lacking a systematic control data framework driven by temperature adaptability; methods based on a single information source (e.g., sequence conservation analysis or static structural features) struggle to simultaneously consider the intensity of temperature response effects and the safety of catalytic function; candidate sites tend to exhibit redundant aggregation in three-dimensional space, resulting in insufficient reproducibility of screening results and inadequate cross-UDG backbone migration ability. Therefore, this invention proposes a method for screening rational mutation sites in thermosensitive UDG based on strain culture temperature. This method, without requiring the construction of a large-scale mutation library, integrates sequence-level temperature difference characteristics and multi-source evidence at the structural level to output a set of preferred mutation sites for thermosensitive UDG engineering, which is highly operable, interpretable, and reproducible across systems. It is particularly suitable for targeted modification of UDG with existing mature expression and purification systems or application validation foundations as engineering seeds.

[0006] Specifically, this invention uses strain culture temperature ecological data as a grouping driving signal. A standardized temperature labeling system is constructed by systematically analyzing temperature records in the strain preservation database, establishing a control dataset for the heat-resistant group and the cold-adapted group. Simultaneously, the corresponding UDG family type 1 homologous sequences and three-dimensional structural data for each group are obtained. Based on this, the cold-temperature differences in amino acid composition sequences between the two groups are mined at the sequence level, and a first set of candidate mutation sites is constructed based on these sequence-level differences. Simultaneously, core folding structure alignment and residue numbering are unified under a unified reference UDG structural coordinate system. By setting avoidance rules for high-risk regions such as catalytic or functional pocket neighborhoods and structurally buried core regions, combined with surface or loop region preference criteria, a second set of candidate mutation sites is initially screened. Furthermore, the candidate sites in the second set of candidate mutation sites are quantified using multi-source structural indicators, including structural set dispersion, residue contact networks, salt bridge networks, and dynamic coupling based on elastic network models and molecular dynamics simulations, forming a multi-dimensional evaluation index system. Finally, through multi-evidence fusion scoring, global ranking, and spatial proximity-based clustering analysis, a small number of preferred mutation sites covering different structural regions are output, which can be directly used for subsequent rational mutation design and experimental verification. This invention can transform temperature ecological information systems into site-level actionable evidence, significantly reducing the experimental exploration space without the need to construct large-scale mutation libraries, improving the reproducibility, interpretability, and transferability across the UDG backbone of site screening results, and providing a systematic site selection method for the rational design and engineering optimization of thermosensitive UDGs.

[0007] To achieve the above objectives, the technical solution of the present invention is as follows: I. A method for screening thermosensitive UDG rational mutation sites based on strain culture temperature (1) Based on the strain preservation database, heat-resistant strains and cold-adapted strains were constructed according to the strain culture temperature. Then, a set of heat UDG sequences and corresponding set of heat three-dimensional structures were constructed based on the heat-resistant strains, and a set of cold UDG sequences and corresponding set of cold three-dimensional structures were constructed based on the cold-adapted strains. (2) Combining the hot and cold differences in sequence, a first candidate mutation site set is generated based on the hot UDG sequence set and the cold UDG sequence set; (3) Align the three-dimensional structures in the hot three-dimensional structure set and the cold three-dimensional structure set to the reference structure coordinate system, and perform preliminary screening of the residue positions in the aligned three-dimensional structures on the reference structure coordinate system to obtain the second candidate mutation site set; (4) Combining the characteristics of thermal differences at the structural level, a set of mutation sites for rational mutation design of thermosensitive UDG is selected from the first set of candidate mutation sites and the second set of candidate mutation sites.

[0008] In (1), based on the strain preservation database, heat-resistant strains and cold-adapted strains are constructed according to the strain culture temperature, including: Strains in the strain preservation database whose maximum growth temperature is not lower than the first preset temperature are designated as heat-resistant strains, and strains in the strain preservation database whose maximum growth temperature does not exceed the second preset temperature are designated as cold-adapted strains.

[0009] A set of thermal UDG sequences and corresponding thermal three-dimensional structures were constructed based on thermostable strains, including: Based on the thermotolerant strains, the corresponding UDG family type 1 homologous sequences were retrieved and screened from the protein sequence database to obtain the thermal UDG sequence set; the three-dimensional structure corresponding to each UDG family type 1 homologous sequence in the thermal UDG sequence set was obtained to obtain the thermal three-dimensional structure set.

[0010] Based on cold-adapted strains, a set of cold UDG sequences and corresponding sets of cold three-dimensional structures were constructed, including: Based on the cold-adapted strains, the corresponding UDG family type 1 homologous sequences were retrieved and screened from the protein sequence database to obtain the cold UDG sequence set; the three-dimensional structure corresponding to each UDG family type 1 homologous sequence in the cold UDG sequence set was obtained to obtain the cold three-dimensional structure set.

[0011] The (2) includes: Based on the hot and cold UDG sequence sets, multiple sequence alignment was performed on all UDG sequences to obtain the alignment results. Combining the amino acid information corresponding to the hot and cold UDG sequence sets, cold-hot difference characteristics were generated for each alignment column in the sequence alignment results. Sites were selected based on the cold-hot difference characteristics of the alignment columns to obtain the first set of candidate mutation sites. The information for each alignment column (i.e., the evidence set) includes the alignment column number, the proportion of non-deletion sequences in the alignment column, the number of valid sequences in the column for the heat-resistant group and the cold-adapted group, and the cold-hot difference characteristics.

[0012] The thermal difference feature is obtained by calculating the difference in amino acid probability distribution between the heat-resistant group and the cold-adapted group in each alignment column. The thermal difference feature of each alignment column includes information such as the temperature difference score (JSD), the dominant amino acid of the heat-resistant group and its frequency.

[0013] The temperature difference fraction (JSD) is calculated as follows: The number of 20 standard amino acids in each alignment column of the heat-resistant group and the cold-adapted group were counted separately to obtain the amino acid frequency distribution of the heat-resistant group and the amino acid frequency distribution of the cold-adapted group in the current alignment column. Then, the smoothed amino acid frequency distribution of the two groups was quantified by using Jensen-Shannon divergence (JSD) to obtain the temperature difference score (JSD) of the current alignment column.

[0014] The site selection was based on an intensity threshold for the difference measure and the difference in the dominant amino acid type between the two groups.

[0015] Optionally, sites are selected based on the cold-heat difference characteristics of the aligned columns, including: The alignment column of the selected site must simultaneously meet the following conditions: a) Difference Intensity Threshold: Sites in the core region column with a temperature difference score (JSD) not lower than a preset threshold are selected. In this embodiment, the threshold is set to 0.21.

[0016] b) Different dominant amino acids between groups: The dominant amino acid types in the current alignment are different between the heat-resistant group and the cold-adapted group.

[0017] c) Intragroup consistency constraint: The dominant amino acid frequency of both the heat-resistant group and the cold-adapted group is not less than 0.50, in order to ensure that the differences come from a relatively consistent distribution preference within the group rather than random fluctuations of individual samples.

[0018] In (3), the reference structural coordinate system is the structural coordinate system corresponding to one of the three-dimensional structures in the thermal three-dimensional structure set and the cold three-dimensional structure set. Generally, the three-dimensional structure with a higher structural confidence score (such as pTM) is selected as the reference structure by default.

[0019] The alignment of the three-dimensional structures is a rigid body alignment. Rigid body alignment uses the conservative core folding region of the UDG as the alignment reference, unifying the three-dimensional structures of each UDG to the coordinate system of the reference UDG structure, and using the residue numbering of the reference UDG structure as the unified residue numbering system. The conservative core folding region is determined by structural occupancy rate screening. Structural occupancy rate is defined as the percentage of samples with valid Cα coordinates at that location within the group's structural set; residue locations in both the heat-resistant and cold-adapted groups with structural occupancy rates reaching a preset threshold are retained, where the preset threshold is an occupancy rate ≥ 80%.

[0020] In step (3), the residue positions in the aligned three-dimensional structure are initially screened on the reference structural coordinate system to obtain a second set of candidate mutation sites, including: Residue sites with structural alignment occupancy below a threshold in the aligned 3D structure are removed; residue sites in the neighborhood of the catalytic functional pocket in the aligned 3D structure are removed; residue sites in the structurally buried core region in the aligned 3D structure are removed; and surface or loop region residue sites in the aligned 3D structure are retained, thereby obtaining a set of second candidate mutation sites.

[0021] In (4), the thermal difference characteristics at the structural level include structural set dispersion difference index, residue contact network difference index, salt bridge network difference index and dynamic coupling difference index.

[0022] The (4) includes: The process involves generating thermal difference features at the structural level for each site in the second candidate mutation site set, calculating a score for each site based on these features, and ranking the sites in the second candidate mutation site set accordingly. Then, based on the spatial distance between sites (e.g., Cα-Cα distance), the sites in the second candidate mutation site set are clustered to obtain site clustering results. Combining the site ranking results, site clustering results, and the first candidate mutation site set, a set of mutation sites for rational mutation design of thermally sensitive UDGs is selected from the second candidate mutation site set. Specifically, the set of mutation sites for rational mutation design of thermally sensitive UDGs includes sites that overlap with the first and second candidate mutation site sets, and these sites are located in different site clusters with higher scores.

[0023] II. A thermosensitive UDG rational mutation site screening device based on strain culture temperature Grouping units are used to construct heat-resistant strains and cold-adapted strains based on strain preservation databases and according to strain culture temperature. The dataset construction unit is used to construct a set of thermal UDG sequences and corresponding thermal three-dimensional structures based on heat-resistant strains, and to construct a set of cold UDG sequences and corresponding cold three-dimensional structures based on cold-adapted strains. The first site screening unit is used to combine the hot and cold difference features at the sequence level to generate a first candidate mutation site set based on the hot UDG sequence set and the cold UDG sequence set. The second site screening unit is used to align the three-dimensional structures in the hot three-dimensional structure set and the cold three-dimensional structure set to the reference structural coordinate system, and to perform preliminary screening of the residue positions in the aligned three-dimensional structures on the reference structural coordinate system to obtain the second candidate mutation site set. The third site screening unit, combining the thermal difference characteristics at the structural level, selects a set of mutation sites for rational mutation design of thermosensitive UDG from the first and second candidate mutation site sets.

[0024] III. A computer device The computer device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the thermosensitive UDG rational mutation site screening method based on strain culture temperature.

[0025] IV. A computer-readable storage medium The medium stores a computer program, which, when executed by a processor, implements the steps of the thermosensitive UDG rational mutation site screening method based on strain culture temperature.

[0026] V. A computer program product The product includes a computer program / instructions that, when executed by a processor, implement the steps of the thermosensitive UDG rational mutation site screening method based on strain culture temperature.

[0027] Compared with the prior art, the present invention has the following beneficial effects: 1. A systematic control framework driven by temperature ecological data was established: Based on the culture temperature ecological data in the publicly available strain preservation database, a control system of heat-resistant group and cold-adapted group was constructed through standardized analysis and grouping. This enabled the screening of mutation sites to have a clear temperature-adaptation driven direction and reproducible data sources, overcoming the limitations of existing methods that rely too much on empirical judgment or small sample inference.

[0028] 2. Achieved joint evaluation of multi-level evidence from sequence and structure: Simultaneously introducing both sequence-level and structural-level thermal difference features. The structural-level thermal difference features include multi-dimensional evidence such as structural set dispersion, residue contact networks, salt bridge networks, and kinetic coupling, enabling comprehensive evaluation and global ranking of candidate sites. This significantly improves the stability, reproducibility, and interpretability of the screening results, avoiding the difficulty in balancing the intensity of temperature response effects and the safety of catalytic function in single-source information methods.

[0029] 3. Achieving structured organization and redundancy control of sites through spatial clustering: Candidate sites are clustered according to three-dimensional spatial proximity and divided into multiple spatial clusters, so that the final output mutation sites are reasonably distributed in different structural regions, avoiding excessive concentration of sites in the same local area; at the same time, the quantitative index system based on multi-source evidence can flexibly adjust the weights according to different engineering focuses, such as a differentiated list of sites that emphasizes temperature response effects or structural safety, thereby improving the operability and adaptability of rational design of thermal UDG without expanding the mutation space.

[0030] 4. Provides a complete transformation path from temperature ecological data to engineering sites: This invention establishes a systematic method and process from strain temperature ecological tags, construction of UDG homology sequence and structure datasets, mining of sequence and structural difference features, candidate site avoidance and preference rule screening, multi-source evidence quantification and fusion evaluation, to final site output and seed UDG mapping. It provides a reproducible and transferable site optimization technology solution for the rational modification of thermosensitive UDGs, effectively reducing the experimental cost and time investment of random mutation screening. Attached Figure Description

[0031] Figure 1 This is a schematic diagram of the overall process of the method proposed in this invention.

[0032] Figure 2 This is a schematic diagram of the structured results of a strain-temperature independent database constructed based on the original temperature records from BacDive.

[0033] Figure 3 This is a schematic diagram of the structured results of the cold adaptation group strain dataset.

[0034] Figure 4 This is a schematic diagram of the structured results of the heat-resistant strain dataset.

[0035] Figure 5 This is a schematic diagram showing the distribution of temperature difference fractions (Jensen-Shannon divergence, JSD) along the reference UDG residue number between the heat-resistant group and the cold-adapted group.

[0036] Figure 6 This is a schematic diagram showing the statistical and preliminary analysis results of the dominant amino acid preferences among different sites at the sequence level.

[0037] Figure 7 This is a schematic diagram showing the distribution of UDG structure occupancy (Occ) along the reference UDG residue number in the heat-resistant and cold-adapted groups.

[0038] Figure 8 This is a schematic diagram of the spatial distribution of 38 candidate sites in the UDG three-dimensional structure for reference.

[0039] Figure 9 The figure shows the statistical results of the ensemble RMSF (eRMSF) of the UDG structures in the heat-resistant group and the cold-adapted group; where (A) is a schematic diagram of the distribution of the two eRMSFs along the reference UDG residue number, and (B) is a schematic diagram of the distribution of the difference between the two eRMSFs along the reference UDG residue number.

[0040] Figure 10 This is a schematic diagram of the mapping of the contact network of heat-resistant and cold-adapted residues on the reference UDG three-dimensional structure.

[0041] Figure 11This is a schematic diagram of the mapping of the salt bridge networks of the heat-resistant group and the cold-adapted group onto the reference UDG three-dimensional structure.

[0042] Figure 12 This is a schematic diagram showing the mapping of the final rational mutation site onto the three-dimensional structure of the engineered seed UDG.

[0043] Figure 13 The figures show the statistical results of the kinetic indices of the heat-resistant group and the cold-adapted group based on the Elastic Network Model (ENM); where (A) is a schematic diagram of the distribution of the mean ENM-RMSF of the two groups along the reference UDG residue number, and (B) is a schematic diagram of the distribution of the mean ENM-coupling of the two groups along the reference UDG residue number.

[0044] Figure 14 The figures show the statistical results of the kinetic indices based on explicit solvent molecular dynamics (MD) simulations for the heat-resistant group and the cold-adapted group; where (A) is a schematic diagram of the distribution of the mean MD-RMSF values ​​of the two groups along the reference UDG residue number, and (B) is a schematic diagram of the distribution of the mean MD-coupling values ​​of the two groups along the reference UDG residue number.

[0045] Figure 15 This is a schematic diagram of the multi-evidence fusion scoring and global ranking results for 38 candidate sites.

[0046] Figure 16 This is a schematic diagram of the spatial clustering results for 38 candidate sites. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this invention and are not intended to limit the scope of protection of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.

[0048] This embodiment focuses on thermosensitive uracil-DNA glycosylase (UDG; preferably UDG family type 1) as the research object. A prokaryotic UDG protein with a clear source was selected as the reference sequence and structural benchmark to systematically screen for rational mutation sites. The screening method is driven by the ecological data of strain culture temperature. A cold and hot control dataset was constructed using culture temperature-related records in the BacDive database, and the corresponding UDG homologous sequences and three-dimensional structural data were obtained by linking the UniProt database. Based on this, a set of candidate sites was formed by mining temperature difference features at the sequence level and quantifying multi-source evidence at the structural level, and the preferred sites for subsequent rational mutation design were output. The sites are represented in the form of "residue number (starting from 1) + amino acid letter code", for example, "40L" indicates that the 40th residue of the UDG sequence is leucine (L). When applied to a specific seed UDG, the reference residue number can be further mapped to the residue numbering system of the corresponding seed UDG before output.

[0049] See Figure 1 , Figure 1 This is a schematic diagram of the overall process of the method corresponding to this embodiment. The process includes the following main steps: First, based on the BacDive database, strain culture temperature ecological data is obtained and a strain-temperature independent database is constructed. Based on the preset temperature threshold, strain screening and cold / hot grouping are completed, and the UDG family type 1 sequence and structure datasets corresponding to the cold and hot groups are constructed in conjunction with the protein database. Then, multiple sequence alignment of the two groups of UDG sequences is constructed and the temperature difference feature evidence at the sequence level is calculated. At the same time, the reference residue mapping number of the temperature difference feature site is determined. Further, the core folding structure alignment and residue numbering are completed under the unified reference UDG structural coordinate system. Based on avoidance and preference rules, the sites are initially screened and the three-dimensional spatial distribution and reference mapping number of the candidate sites are determined. Then, the candidate sites are subjected to multi-source structural evidence quantification, including structural set dispersion difference evidence, residue contact and salt bridge network difference evidence, and dynamic coupling difference evidence, forming a multi-dimensional quantitative index matrix of candidate sites. Finally, the candidate sites are subjected to multi-evidence fusion evaluation and comprehensive ranking, and spatial clusters are obtained by performing spatial proximity clustering analysis. Based on the optimal screening within the cluster, the final mutation sites and their combination design suggestions for rational design of thermosensitive UDG are output.

[0050] 1. Construction of a control dataset driven by temperature ecological data In this embodiment, a consistent dataset of "temperature ecological data - strain grouping - UDG family type 1 sequence and three-dimensional structure" is first constructed as the input basis for subsequent extraction of temperature difference features at the sequence level, initial screening for structural avoidance, and quantification and fusion evaluation of multi-source structural evidence. The dataset construction in this embodiment includes: acquisition of strain culture temperature ecological data, parsing and cleaning of temperature data to form a strain-temperature independent database, and the linkage construction of cold and hot grouping based on temperature thresholds with the UDG family type 1 sequence and structure dataset.

[0051] (1) Acquisition of ecological data on strain culture temperature Public resources such as microbial strain preservation databases have accumulated a wealth of microbial culture temperature ecology data. This data systematically reflects the temperature adaptation characteristics of microorganisms at the population level, providing a reliable basis for constructing heat-resistant and cold-adapted control populations. By integrating temperature ecology tags with UDG homologous sequences and three-dimensional structural information, multi-level differential characteristics related to heat sensitivity can be systematically mined from a "temperature ecology data-driven" perspective, providing a new data foundation and methodological support for rational mutation sites.

[0052] In a preferred embodiment, the strain culture temperature ecological data is obtained from the publicly available strain preservation database BacDive (Bacterial Diversity Metadatabase). BacDive uses "strain entries" as the organizational unit, with each entry corresponding to a unique BacDive-ID identifier. The BacDive-ID of the target strain entry can be obtained through the BacDive official website or an equivalent search interface; furthermore, the BacDive-ID can be precisely retrieved through the application programming interface (API) provided by BacDive to obtain the complete structured information object of the strain entry.

[0053] It should be noted that the strain entry information returned by the API is usually a hierarchical structured data object (e.g., dictionary or JSON format). Fields cover multiple modules such as "General Information," "Taxonomic Information," "Cultivation and Growth Conditions," "Isolation Information," "Sequence Information," and "Literature Information." Multiple records may exist within the same module (e.g., multiple temperature records for the same strain under different culture conditions or from different literature sources). This invention does not require listing all fields of this structured object, but only extracts the key fields necessary for temperature range reconstruction and subsequent sequence-structure linkage, including but not limited to: BacDive-ID, NCBI taxonomy ID (NCBI taxid), scientific name, and temperature-related records in the "Culture and Growth Conditions" module (e.g., culture temp, growth temperature, optimal temperature, maximum temperature, minimum temperature, or equivalent fields).

[0054] For example, when retrieving the strain entry with BacDive-ID 175115, the API returns the temperature-related records in its "Culture and growth conditions" module, showing a culture temp of 50 ℃, and also returns the strain's NCBI taxid and taxonomic information. Through this mechanism, it is possible to batch retrieve strain entries from the entire database using BacDive-ID as the entry point, and provide raw data input for subsequent temperature field parsing and standardization.

[0055] (2) Temperature field parsing and cleaning, strain-temperature independent database construction and cold and hot grouping Since the temperature information in the original BacDive entries may appear in different field names and different numerical formats (such as single-value temperature, interval temperature, multi-point temperature measurement values, descriptions of gronomic and non-gronomic conditions, etc.), this embodiment performs systematic parsing, cleaning and standardization of the original temperature records, reconstructs the scattered records into calculable "strain growth temperature interval" information, and forms a strain-temperature independent database accordingly.

[0056] In a preferred embodiment, temperature field parsing and cleaning includes at least the following processing steps: a) Temperature value standardization: All temperature expressions are uniformly converted into Celsius temperature values; for temperature ranges in the form of "25-37 ℃", they are resolved to a lower limit of 25 ℃ and an upper limit of 37 ℃; for single-value temperatures in the form of "50 ℃", they are resolved to a single-point temperature value of 50 ℃. b) Merging of positive and negative growth information: When the original record contains descriptions such as "growth positive" and "growth negative", it is merged into "growth temperature range or set" and "growth temperature point, range or set" respectively. c) Merging multiple records: When there are multiple temperature-related records for the same strain, it is preferable to merge the records with "growing" as the main record to obtain the lower and upper limits of the growing temperature of the strain; if necessary, retain the non-growing information as a supplementary constraint or a basis for consistency checks. d) Derivation index calculation: The maximum growth temperature of the strain is derived from the "growable temperature range" (defined as the upper limit of the growable range or the maximum value of the growable temperature point), and is used for subsequent threshold-based cold and hot grouping.

[0057] See Figure 2 , Figure 2 This illustration schematically demonstrates the structured results of a "strain-temperature independent database" constructed based on the original temperature records of BacDive entries. In this embodiment, the strain-temperature independent database includes the following fields: BacDive-ID, NCBI taxid, species name, standardized growth temperature information (including the range of achievable and non-achievable temperature records), and the derived maximum growth temperature, etc. As an example, some records in the independent database are represented as follows: for strain BacDive-ID 135894, after parsing, "achievable temperature: 25.0-37.0 ℃; non-achievable temperature: 10.0 ℃ and 45.0 ℃"; for strain BacDive-ID 2728, after parsing, "achievable temperature: 60.0 ℃; non-achievable temperature: none".

[0058] After completing the construction of the independent strain-temperature database, this embodiment further divides the strains into a heat-resistant group and a cold-adapted group based on preset temperature thresholds. Preferably, the maximum growth temperature is used as the grouping criterion: strains with a maximum growth temperature not exceeding 25 ℃ are classified into the cold-adapted group, and strains with a maximum growth temperature not lower than 60 ℃ are classified into the heat-resistant group; strains with a maximum growth temperature between 25 ℃ and 60 ℃ may be excluded from the above two groups or separately labeled as an intermediate group to avoid interference from boundary samples in subsequent difference statistics.

[0059] See Figure 3 and Figure 4 , Figure 3The structured results of the cold-adapted strain dataset are illustrated schematically. Figure 4 The diagram illustrates the structured results of the heat-resistant strain dataset. Thus, this embodiment yields two sets of control strains with clearly defined temperature labels, providing an input basis for subsequently linking temperature-related ecological data to UDG sequence and structural data.

[0060] (3) Constructing a dataset of UDG family type 1 sequences and three-dimensional structures based on hot and cold grouping linkage After obtaining the cold-acclimatized and heat-resistant strain sets, this embodiment further uses the NCBI taxid of each strain as a bridging identifier to retrieve and obtain its corresponding UDG family type 1 homologous sequence in a protein sequence database (preferably the UniProt database), thereby establishing a correspondence between "temperature grouping - UDG sequence". In addition, to supplement the representativeness and coverage of the thermosensitive UDG samples on the cold-acclimatized side, this embodiment can also supplement a small number of thermosensitive UDG entries from psychrophilic bacteria reported in the literature. Their amino acid sequences and three-dimensional structural information are organized according to a unified standard and incorporated into the cold-acclimatized dataset, with the temperature labels of the entries based on publicly reported data.

[0061] In a preferred embodiment, the retrieval and screening of UDG family type 1 homologous sequences includes at least the following processing steps: a) Limiting the species scope based on NCBI taxid: Using the NCBI taxid of each strain as the search criteria, obtain protein entries within that taxonomic scope; b) UDG family type 1 limited by annotation and domain: Use protein name or functional annotation (e.g. uracil-DNA glycosylase) and domain annotation (e.g., corresponding UDG family type 1 entries in databases such as InterPro and Pfam) for joint screening to avoid misinclude non-family type 1 or functionally similar but different family glycosylation enzymes. c) Sequence quality control: Remove sequences with obvious truncated sequences, abnormal length sequences, or sequences containing more than a preset threshold of unknown amino acids; when there are multiple highly similar UDG entries under the same NCBI taxid, representative sequences can be retained for subsequent statistical analysis to reduce statistical bias caused by close repetitions.

[0062] After obtaining the cold and hot UDG sequences, this embodiment further acquires the corresponding UDG three-dimensional structure for subsequent site avoidance and multi-source structural evidence quantification under a unified structural coordinate system. Preferably, the acquisition of the three-dimensional structure follows the following strategy: a) When a publicly available experimental parsing structure exists for a UDG entry, the corresponding structure file should be downloaded and used directly first. b) When publicly available experimental structures are unavailable, protein structure prediction models (such as AlphaFold2) can be used to predict the monomeric structure of the UDG sequence, thereby obtaining a three-dimensional structure model of UDG and a structure confidence evaluation index (such as predicted TM-score, pTM). The three-dimensional structure model that passes the structure confidence evaluation (such as pTM not lower than 0.8) can be used as input for subsequent analysis.

[0063] Therefore, this embodiment ultimately constructs two sets of UDG family type 1 sequence and three-dimensional structure datasets, one for heat tolerance and the other for cold adaptation. The heat-tolerant UDG family type 1 dataset contains 35 samples, and the cold-adapted UDG family type 1 dataset contains 35 samples. Furthermore, two publicly reported thermosensitive UDG samples from psychrophilic bacteria are added, bringing the total number of cold-adapted samples to 37. These datasets provide a unified and reproducible data input for subsequent steps in this invention, such as "sequence-level temperature difference feature extraction" and "structural alignment and preliminary screening of candidate sites driven by avoidance and preference rules."

[0064] 2. Extraction of temperature difference features at the sequence level After constructing the UDG family type 1 sequence and 3D structure datasets for the heat-resistant and cold-adapted groups, this embodiment further conducts a comparative analysis of the two groups of UDG sequences at the sequence level to uncover differential features related to temperature adaptation and form sequence-level temperature difference evidence that can be directly used for subsequent site evaluation and fusion. In this embodiment, the heat-resistant UDG family type 1 sequence consists of 35 sequences, and the cold-adapted UDG family type 1 sequence consists of 37 sequences. The cold-adapted group includes 35 sequences obtained by linking temperature group datasets and 2 sequences from psychrophilic bacteria reported in the literature (BMTU3346 and...). Bacillus sp. HJ171) thermosensitive UDG sequence. Furthermore, to facilitate mapping sequence difference signals to a unified residue numbering system and integrating them with structural evidence, this embodiment selects from the heat-resistant group and the cold-adapted group. Aeribacillus pallidus The UDG amino acid sequence (UniProt accession number A0A161WTR3) was used as a reference sequence for subsequent fusion evaluation with structural evidence and for plotting differential distribution curves.

[0065] (1) Construction of multiple sequence alignment of two sets of UDG sequences In a preferred embodiment, the heat-resistant and cold-adapted UDG family type 1 sequences are merged into a unified input set, and a Multiple Sequence Alignment (MSA) is constructed. The MSA tool can be MAFFT, Clustal Omega, MUSCLE, or other equivalent MSA programs; in this embodiment, MAFFT is used to obtain the alignment results of the two sets of sequences. After the MSA is completed, the consistency of the column lengths of the alignment matrix is ​​checked to ensure that subsequent column-level statistics can be executed stably; when an abnormal alignment length is found in individual sequences, their sequence quality is reviewed, and obviously truncated or abnormal entries are removed before re-alignment.

[0066] (2) Alignment quality control and core area retention Multiple sequence alignment results often exhibit alignment columns with a high proportion of missing (gap) segments at the N- or C-terminus. Directly including these in differential statistics can easily lead to differential scores being misled by missing information and deviating from the true amino acid differences. Therefore, this embodiment performs column-level quality control on alignment columns, retaining core regions with high alignment occupancy for differential feature calculation. In a preferred embodiment, a "non-missing percentage" is calculated for each alignment column, defined as the ratio of the number of non-missing characters in that column to the total number of sequences. When the non-missing percentage of an alignment column falls below a preset threshold, it is removed from subsequent differential feature calculations. This embodiment sets the threshold at 70%, meaning that sequence-level temperature difference evidence is calculated only in regions where the non-missing percentage of alignment columns is not less than 70%, thus reducing the impact of high-missing-terminal regions on the statistical results.

[0067] Furthermore, to avoid instability in the differential measurement due to insufficient effective sample count within a group, this embodiment also counts the number of non-missing effective sequences in each column for both the heat-resistant group and the cold-adapted group, and sets a threshold for effective samples within the group; when the number of effective sequences in a column for either the heat-resistant group or the cold-adapted group is lower than the threshold, that column is not included in the differential feature calculation. In this embodiment, the threshold for the number of effective sequences in both the heat-resistant group and the cold-adapted group is set to be no less than 5.

[0068] (3) Calculation of evidence for temperature difference characteristics After completing the alignment quality control, this embodiment calculates the amino acid distribution differences between the heat-resistant group and the cold-adapted group for the retained core region alignment, and forms cold and heat difference characteristics at the sequence level.

[0069] In a preferred embodiment, the counts of 20 standard amino acids in each retained aligned column are statistically analyzed for both the heat-resistant and cold-adapted groups, yielding the amino acid frequency distributions for both groups. To avoid zero probability, which could render the difference measurement incalculable or unstable, this embodiment introduces a pseudo-count smoothing process. Specifically, a pseudo-count is added to the counts of the 20 amino acids in each column for each group before normalization to a probability distribution. The pseudo-count is set to 0.5. Based on the smoothed amino acid distributions, this embodiment uses Jensen-Shannon divergence (JSD) to quantify the difference between the two distributions, obtaining the temperature difference score for that column. The closer the temperature difference score is to 0, the closer the amino acid preferences of the two groups are; the larger the temperature difference score, the more significant the difference in amino acid composition between the two groups. See also... Figure 5 , Figure 5 The diagram illustrates the distribution curve of temperature difference fractions along the sequence position based on the reference UDG residue number, which is used to visually represent the aggregation location and intensity of the sequence-level temperature difference signal on the reference UDG.

[0070] In addition to the temperature difference score, this embodiment also outputs the dominant amino acid type and frequency of each alignment in both groups, to help explain the source of the difference and provide more readable feature information for subsequent comprehensive evaluation. Thus, this embodiment forms a sequence-level temperature difference feature evidence set, which includes at least: alignment column number, the proportion of non-deleted alignment columns, the number of valid sequences in that column in the heat-resistant and cold-adapted groups, the temperature difference score (JSD), and the dominant amino acids and their frequencies in both groups. Through the above method, this embodiment obtains sortable and reproducible sequence-level temperature difference feature evidence, providing input for subsequent fusion evaluation with structural-level evidence.

[0071] (4) Preliminary screening of differentially expressed sites at the sequence level After obtaining the temperature difference score, this embodiment further screens a group of preliminary differential sites with "strong differential signals and different dominant amino acids between groups" at the sequence level, which are used as important sequence evidence input for subsequent multi-evidence fusion evaluation.

[0072] In one feasible implementation, selecting sites based on the temperature difference characteristics of the aligned columns includes: The alignment column of the selected site must simultaneously meet the following conditions: a) Difference Intensity Threshold: Sites in the core region column with a temperature difference score (JSD) not lower than a preset threshold are selected. In this embodiment, the threshold is set to 0.21.

[0073] b) Different dominant amino acids between groups: The dominant amino acid types in the current alignment are different between the heat-resistant group and the cold-adapted group.

[0074] c) Intragroup consistency constraint: The dominant amino acid frequency of both the heat-resistant group and the cold-adapted group is not less than 0.50, in order to ensure that the differences come from a relatively consistent distribution preference within the group rather than random fluctuations of individual samples.

[0075] Following the above criteria, this embodiment screened out four sequence-level differential sites, and compared them with the reference UDG ( Aeribacillus pallidus -A0A161WTR3) corresponds to 41D, 176N, 185S and 223S respectively; the four sites are representative sites with strong temperature difference signals at the sequence level. The above four sequence level difference sites constitute the first candidate mutation site set described in this invention, which is used for subsequent fusion evaluation with structural level evidence.

[0076] Furthermore, to support the determination of replacement direction in subsequent rational mutation design, this embodiment statistically analyzes the amino acid composition of the four sites in the heat-resistant and cold-adapted groups, providing at least the dominant amino acid types and their proportions in both groups, and also providing information on the proportions of minor amino acids, to characterize the preference differences and concentration of the site in the two groups. See [link to documentation]. Figure 6 , Figure 6 The diagram illustrates the statistical results of the amino acid composition of the four sites in the heat-resistant and cold-adapted groups, along with corresponding preliminary analysis. Therefore, when a subsequent mutation scheme includes these sites, the amino acid preferences shown in the cold-adapted group can be referenced to preferentially allocate the replacement amino acid type for that site.

[0077] 3. A preliminary screening method for candidate sites based on a unified structural coordinate system After constructing the type 1 sequence and three-dimensional structure datasets of the heat-resistant and cold-adapted UDG families, and extracting temperature difference features at the sequence level, this embodiment establishes a unified reference UDG structural coordinate system and performs structural alignment and residue numbering standardization to achieve quantitative comparability analysis of UDG structures from different sources at the residue level. Based on this, avoidance and preference rules that comprehensively consider catalytic function protection, structural stability, and engineering operability are constructed to initially screen mutable sites, obtaining a set of candidate mutation sites for subsequent multi-source structural evidence quantification and comprehensive evaluation. The technical features of this method are: by rigidly aligning the conserved regions of the core folds, heterogeneous UDG structures are unified to the same coordinate reference, enabling quantitative comparison of cross-species structural differences under a unified reference system; through spatial geometric distance constraints and neighborhood contact density analysis, functional protection regions and preferred modification regions are systematically defined, enriching surface ring sites with temperature regulation potential while protecting catalytic activity.

[0078] (1) Establishment of a unified reference UDG structural coordinate system and determination of residue numbering benchmark In a preferred embodiment, consistent with the sequence-level temperature difference feature extraction step, the reference UDG is still selected. Aeribacillus pallidus Source: UDG (UniProt accession number A0A161WTR3). All subsequent structural samples were aligned, numbered uniformly, and had structural features such as residue distances, contact relationships, and salt bridge relationships calculated in this reference UDG structural coordinate system to ensure that the structural evidence between the cold and hot groups is comparable and corresponds one-to-one.

[0079] (2) Rigid structure alignment based on the core folding conservative region Type 1 UDG family exhibits a conservative α / β domain topology at the core fold level, while the N-terminus, C-terminus, and connecting loop regions show length variations and conformational diversity. To avoid perturbation of alignment by highly variable regions, this embodiment uses the conservative core fold region as the rigid body alignment object and performs core fold alignment on each structural sample.

[0080] In a preferred embodiment, a mapping relationship between reference residue numbers and sample residues is first established based on the results of multiple sequence alignment: taking the non-deleted position of the reference UDG sequence in the multiple sequence alignment as a benchmark, a mapping from the alignment position to the reference residue number is established, and the corresponding position of the sample sequence in the alignment is further mapped to the residue index in the sample structure, thereby obtaining a set of coordinate points in each sample structure that can be mapped to the reference residue numbering system.

[0081] Furthermore, to improve alignment stability, this embodiment calculates the structural coverage of each reference residue number position in the two sets of structures (hot and cold), and defines the core site set for alignment accordingly: when the structural coverage of a reference site in both the heat-resistant group and the cold-adapted group is not lower than a preset threshold, the site is determined to belong to the core fold comparable region and included in the core site set; the coverage threshold is preferably 0.8, which ensures that the alignment benchmark comes from the structural regions that are common and comparable to the vast majority of samples in the two groups.

[0082] In a preferred embodiment, rigid structure alignment is performed on each sample UDG structure to unify the three-dimensional structures of UDGs from different sources into the same reference UDG structure coordinate system. Rigid alignment refers to repositioning the sample structure using only a global rotation matrix and translation vector, without altering the protein's internal bond lengths, bond angles, and dihedral angles, to achieve optimal spatial overlap between its core folded region and the reference structure. Specific process: a) Corresponding point selection: Extract the Cα atom coordinates of the corresponding residues of the core site set from the reference UDG structure and each sample structure to form a one-to-one matching point pair; b) Alignment Transformation Solution: Since the origin and overall orientation of the 3D coordinates of UDG structures from different sources are arbitrary when saved or predicted, if this placement difference caused by the coordinate system is not eliminated first, the coordinates of the same reference residue in different structures cannot be directly compared, and subsequent structural screening methods will be amplified by errors. Therefore, this step uses the corresponding Cα atom coordinate pairs of the core site set as a benchmark to find a set of overall rotation and translation parameters for the sample structure, so that the core corresponding points in the sample structure are spatially closest to the corresponding points in the reference structure, minimizing the overall deviation between all corresponding points. This "closest" principle is a classic concept in the comparison of different protein structures. Its significance lies in achieving optimal overlap between the two structures in the core folding region without changing the internal geometry, thereby placing the sample structure uniformly in the reference UDG structure coordinate system and ensuring the comparability and consistency of subsequent residue-level spatial feature calculations.

[0083] c) Measurement and Application of Overlap Degree: The average deviation of the matched point pairs after alignment is used as the evaluation index of overlap degree. In this embodiment, the root mean square deviation (RMSD) is used for measurement. RMSD can be understood as the average level of the distance (in Å) between corresponding Cα atoms after alignment. The smaller the value, the higher the overlap degree of the two structures in the core region. After obtaining the overall rotation and overall translation parameters, they are applied to the coordinates of all mapped residues in the sample structure (optionally including the coordinates of all atoms in the structure), so that the entire sample structure falls into the reference UDG structural coordinate system.

[0084] In this way, the UDG structures of the heat-resistant group and the cold-adapted group are unified under the same reference coordinate system, thereby ensuring that the subsequent calculations of structural features such as inter-residue distance, contact relationship, salt bridge relationship and site spatial proximity are comparable, and that statistical and fusion analysis can be performed under a unified reference residue numbering system.

[0085] (3) Screening of comparable residue positions based on structural occupancy After unifying the structures of the hot and cold UDGs to a reference coordinate system, due to potential length differences or incomplete modeling at the ends or local loop regions of UDGs from different sources, not all reference residue positions have valid structural coordinates in all samples. To ensure the statistical robustness and inter-group comparability of subsequent quantitative structural indicators, this embodiment introduces the occupancy (Occ) index to screen comparable residue positions with sufficient structural data support in both the hot and cold groups.

[0086] In a preferred embodiment, using the reference UDG sequence length as a benchmark (in this embodiment, the non-deleted length of the reference UDG is 227 bits), for each reference residue number position i, the number of structural samples at that position whose Cα coordinates can be successfully mapped and which are not missing in both the heat-resistant group and the cold-adapted group are counted, and the structural occupancy rate at that position in the two groups is calculated accordingly: a) Occ-hot(i) of heat-resistant group = number of effective structures in heat-resistant group divided by total number of structures in heat-resistant group; b) Occ-cold(i) of cold adaptation group = number of effective structures in cold adaptation group divided by total number of structures in cold adaptation group.

[0087] Furthermore, to ensure comparability between the two groups, the smaller of the two occupancy rates is used as the minimum occupancy rate for that position: min-Occ(i) = min{Occ-hot(i), Occ-cold(i)} min-Occ(i) was used as the basic comparability threshold for subsequent initial screening.

[0088] In a preferred embodiment, the minimum occupancy rate threshold is set to 0.8, meaning that only reference residue positions with a min-Occ(i) of not less than 0.8 are retained for subsequent structural evidence quantification and initial screening of candidate sites. This threshold ensures that at least 80% of the samples at this position provide valid structural data in both groups, resulting in a sufficient statistical sample size. See also Figure 7 , Figure 7 The structural occupancy distribution of the two groups is shown. Furthermore, the first 5 residues at the N-terminus and the last 5 residues at the C-terminus of the reference UDG are avoided even if they meet the occupancy threshold, in order to reduce the interference of high terminal flexibility and incomplete modeling on the statistical analysis of structural differences.

[0089] Through the above screening, this embodiment ultimately identified approximately 200 residue positions that met the comparability requirements, providing a high-quality range of residue positions for subsequent application of avoidance rules and initial screening of candidate sites.

[0090] (4) Catalytic functional pocket neighborhood avoidance based on DNA substrate complex structure The catalytic activity of UDG depends on the precise recognition of DNA substrates and uracil bases at its active site. Mutations in this active site and its adjacent regions can directly affect substrate binding, catalytic efficiency, or enzyme activity, and are considered high-risk regions that must be strictly protected in thermosensitive mutation design. To systematically define this functional protection region, this embodiment develops a spatial geometry avoidance method based on the UDG-DNA complex structure. In a preferred embodiment, the establishment of the catalytic functional pocket neighborhood avoidance domain includes the following steps: a) Geometric marker template selection and coordinate system alignment: A UDG-DNA complex structure containing DNA ligands (e.g., PDB ID: 1EMH) was selected from the PDB database as a geometric marker template. This complex structure contains a co-crystal structure of UDG and a uracil-containing double-stranded DNA substrate, fully demonstrating the binding conformation of the substrate at the active site. The UDG protein portion of the complex structure was used as the alignment object, and the core fold rigid body alignment method consistent with step (2) was used to align and unify the UDG protein of the complex to the reference UDG structure coordinate system; the same rigid body transformation was synchronously applied to the DNA coordinates in the complex, thereby enabling the DNA to achieve stable positioning in the reference UDG structure coordinate system.

[0091] b) DNA Marker Set Determination: After coordinate system alignment, all heavy atoms (carbon, nitrogen, oxygen, and phosphorus atoms) of the DNA substrate in the complex are defined as the DNA marker set to describe the spatial extent of the functional pocket neighborhood. Further, when the complex contains identifiable uracil sites, all heavy atoms of the uracil bases are defined separately as the uracil marker set to more precisely define the core neighborhood of the substrate recognition pocket.

[0092] c) Calculation of the minimum distance from residues to DNA neighborhoods: For each residue position i in the reference UDG residue numbering system, calculate the Euclidean distance between its Cα atomic coordinates and the atomic coordinates of all atoms in the DNA marker set, and take the minimum value to obtain the distance index dist. DNA (i) When the uracil label set is enabled, the minimum distance between the Cα atom of the residue and the uracil label set is further calculated to obtain dist. U (i).

[0093] d) Avoidance threshold setting and avoidance domain generation: When dist DNA (i) When the value is less than a preset threshold (preferably 8 Å), the location is marked as a DNA neighborhood avoidance site; when the uracil marker set is enabled and dist U (i) When the distance is less than a preset threshold (preferably 10 Å), the location is marked as a uracil site neighborhood avoidance site. The distance threshold is set based on the typical amino acid residue interaction distance range; 8.0 Å can cover the first and second layer residues that are in direct contact with or adjacent to the DNA substrate, and a 10.0 Å threshold ensures the complete protection range of the substrate recognition pocket. The resulting avoidance site merge set (dist pocket (i) =dist DNA (i) ∪ dist U (i) constitutes the catalytic function pocket neighborhood avoidance domain in the reference UDG coordinate system.

[0094] This method transforms biochemical functional constraints into computable spatial geometric constraints, enabling quantitative and reproducible protection of the catalytic core region.

[0095] (5) Structural core area avoidance and surface ring area preference rules based on neighborhood contact density To avoid uncontrollable disturbances to the folded buried core and to improve the operability of site engineering, this embodiment constructs buried core region avoidance and surface and ring region preference rules, which are used to further screen regions more suitable for rational mutation exploration after functional region avoidance.

[0096] In a preferred embodiment, the number of neighboring contacts is used as a structural proxy for the degree of burial: taking the Cα atom at position i of each reference residue in the reference UDG structure as the center, the number of neighboring Cα residues within a preset radius (preferably set to 10 Å) is counted and denoted as contact10(i). Generally, a larger contact10(i) indicates more neighbors around the position and a tighter burial; a smaller contact10(i) indicates a more surface or loop region. This 10 Å radius is based on the typical mid-range interaction distance in protein structures and can effectively distinguish between buried residues and surface residues.

[0097] Furthermore, this embodiment adaptively determines the surface threshold using a quantile approach to adapt to the overall contact distribution of different reference UDG structures: the 25th percentile of contact10(i) is used as the surface threshold. thr Only contact10(i) below surface is retained. thr The location of contact10(i) is used as the preferred site for the surface loop region, and locations where contact10(i) is significantly higher than the threshold are excluded from the initial screening as burial core region avoidance sites. This adaptive threshold strategy enables the avoidance rule to automatically adapt to the overall contact density distribution of the reference UDG, avoiding the applicability problem of a fixed threshold across different protein backbones.

[0098] In the above manner, an executable structural preference rule for surface ring regions and buried core avoidance can be formed under the reference UDG residue numbering system.

[0099] (6) Joint preliminary screening and determination of candidate site set based on multi-layered avoidance and preference rules In a preferred embodiment, the initial screening of candidate sites is performed according to the following procedure: a) Determination of the set of comparable positions: The minimum occupancy rate min-Occ(i) obtained in step (3) is not less than 0.8 as the comparability threshold. Positions that meet this condition are first retained for initial screening. Several residues at the end of the reference UDG (e.g., 5 residues each at the N-terminus and C-terminus) are avoided to reduce the impact of high plasticity and deletion fluctuations at the end on the screening stability.

[0100] b) Functional region avoidance: Within the set of comparable locations mentioned above, the catalytic functional pocket neighborhood avoidance domain from step (4) is applied to exclude dist. pocket (i) At the corresponding position, the initial screening set far from the functional area is obtained. Where dist pocket (i) by dist DNA (i) and dist U (i) Identify and exclude residues when the minimum distance between the residue and the DNA substrate or uracil recognition site is less than a preset threshold.

[0101] c) Buried core avoidance and surface ring region preference: Within the initial screening set far from functional areas, the contact number preference rule from step (5) is applied to retain positions that satisfy the surface ring region preference and avoid buried core region positions, resulting in a candidate set more suitable for engineering exploration. Specifically, contacts10(i) below surface are retained. thr Surface ring region residues, excluding buried core residues where contact10(i) is significantly higher than the threshold.

[0102] d) Candidate Site Set Determination and Indicator Set Registration: After completing steps a to c, the reference residue numbers that simultaneously meet the comparability threshold, functional region avoidance, buried core region avoidance, and surface ring region preference are summarized to form a candidate site set. In this embodiment, 38 candidate sites are obtained after initial screening according to the above avoidance and preference rules. Further, the structural characteristic indicators (including but not limited to contact10(i), dist) of the candidate site set are registered site by site. DNA (i), dist U (i), dist pocket (i) etc.) and form a set of candidate site indicators, which serve as the data basis for subsequent multi-source structural evidence quantification, fusion evaluation and comprehensive ranking.

[0103] Furthermore, the set of 38 candidate sites is preferably output using the reference UDG residue numbering system and represented in the form of residue number plus amino acid letter. In this embodiment, the positions of the 38 candidate sites on the reference UDG (A0A161WTR3) and their corresponding amino acids are: 11G, 12L, 23K, 24E, 27E, 28F, 30K, 31Q, 33Y, 34A, 35H, 36H, 41D, 42M, 43Y, 55E, 83G, 101L, 102G, 138G, 146D, 150S, 151C, 154E, 155R, 173E, 174M, 176N, 177M, 178N, 199G, 208Q, 211K, 212S, 213I, 214G, 216E, 219D.

[0104] See Figure 8 , Figure 8 The distribution of 38 candidate sites in the three-dimensional structure of the reference UDG is shown. These 38 candidate sites form the second candidate mutation site set, which will serve as the candidate input set for subsequent multi-source structural evidence quantification, comprehensive evaluation, spatial clustering, and final output.

[0105] 4. Quantification of multi-source structural evidence for candidate sites After obtaining 38 candidate sites under the reference UDG residue numbering system in step 3, this embodiment further performs quantitative calculations of multi-source structural evidence for the candidate sites in a unified coordinate system to obtain a set of structural indicators that can be used for subsequent comprehensive evaluation and ranking. Multi-source structural evidence includes structural set dispersion correlation indicators, residue contact network difference indicators, salt bridge network difference indicators, and dynamic correlation indicators based on the Elastic Network Model (ENM) and Molecular Dynamics (MD). All of these indicators are output under the reference UDG residue numbering system and correspond one-to-one with the candidate sites, forming a multi-indicator quantitative dataset for the candidate sites. The technical feature of this multi-source evidence quantification method is that it systematically characterizes the temperature-related features of candidate sites from two dimensions—static structural differences and dynamic behavioral differences—by combining structural set statistics with dynamic simulation, providing multi-level, cross-verifiable quantitative evidence for subsequent comprehensive evaluation.

[0106] (1) Quantification of structural set discreteness related indicators In a preferred embodiment, ensemble RMSF (eRMSF) is used to characterize the spatial dispersion of the same reference residue number position within a set of structures. This ensemble RMSF is not a temporal fluctuation in the sense of molecular dynamics trajectories, but rather a coordinate dispersion across the structure set, used to establish a comparable difference signal between the heat-resistant group and the cold-adapted group. Its calculation follows the statistical idea of ​​root mean square fluctuation (RMSF) in molecular dynamics analysis: after uniform alignment to the reference UDG coordinate system, the Cα coordinates of the same reference residue number position in all valid structural samples within the set are considered as a set of observations. The fluctuation amplitude relative to the average position is statistically analyzed according to the RMSF calculation method, thus obtaining the dispersion (in Å) of that position. A larger dispersion indicates that the position is more dispersed or variable within the set of structures, reflecting that the position is more likely to be located in a more plastic ring region or flexible region, or has more significant conformational differences.

[0107] In a preferred embodiment, the heat resistance group dispersion eRMSF is calculated for each candidate site i. hot (i) Dispersion of cold adaptation group eRMSF cold(i). Further, calculate the difference in dispersion between the two sets ΔeRMSF(i) = eRMSF cold (i) -eRMSF hot (i) is used to characterize the direction and intensity of the difference in plasticity at this site under cold and hot conditions. To facilitate subsequent integration with other structural evidence, the eRMSF... hot (i), eRMSF cold (i) and ΔeRMSF(i) are included in the candidate site index set as candidate site structural fluctuation indexes, and their distribution curves along the reference residue number positions are plotted to visually show the aggregation location and intensity of structural plasticity difference signals at candidate sites.

[0108] See Figure 9 , Figure 9 (A) shows the structural dispersion distribution of the two groups. Figure 9 (B) shows the distribution of the structural dispersion difference between the two groups.

[0109] (2) Quantification of evidence of differences in residue contact networks Whether a site is on the surface or in a loop region is often insufficient to reflect its potential impact on overall stability and structural network. To characterize the differences in structural network of candidate sites under cold and hot controls from the perspective of neighborhood connectivity, this embodiment constructs two sets of residue contact probability matrices and extracts site-level contact reconnection, network coreness, and bridging properties based on these matrices.

[0110] In a preferred embodiment, the establishment and quantification of residue contact relationships preferably includes the following steps: a) Single-structure contact determination: In the aligned structure, using the residue position under the reference residue numbering system as the node, for each protein monomer structure, the Cα-Cα distance is preferably used as the geometric criterion. Specifically, when the Cα distance between two residues does not exceed 8.0 Å, they are considered to be in contact. To avoid miscounting necessarily nearest neighbors of adjacent residues in the main chain as structural network contacts, this embodiment preferably avoids residue pairs with excessively close sequence distances, only counting residue pairs with a residue number difference greater than 3.

[0111] b) Construction of the intragroup contact probability matrix: For both the heat-resistant and cold-adapted groups, count the number of times each residue pair (i, j) is identified as having contact in the structures of different proteins within the group, and divide by the number of protein structures in that group to obtain the contact probability P corresponding to the heat-resistant and cold-adapted groups. hot (i, j) and P cold (i, j). This probability can be understood as what proportion of the structural samples in the heat-resistant or cold-adapted group show that (i, j) is a pair of contact residues.

[0112] c) Contact difference matrix and node metric: In obtaining P hot (i, j) and P cold After (i, j), from the perspective of candidate site i as a node, the probability information of residue pairs is summarized into site-level indicators for subsequent comprehensive evaluation: (i) Contact degree: For a site i in both groups, calculate the sum of its contact probabilities with all other residues to obtain the contact degree deg of the heat-resistant group. hot (i) = ∑ j P hot (i, j) and the contact degree between the cold adaptation group (deg) cold (i)= ∑ j P cold (i, j). This index can be understood as: how much contact connection or how stable the contact is on average at site i in this set of structures. A larger value usually means that the site is in a tighter structural enclosure environment, that is, it is more inclined to be buried deeper or to bear denser neighborhood constraints.

[0113] (ii) Contact rewiring amplitude: To characterize how much the neighborhood contact pattern of site i changes between the hot and cold groups, the contact probability difference between site i and each residue j is compared one by one, and these differences are summed into a total, resulting in the contact rewiring amplitude rewire(i) = ∑ j | P hot (i, j) - P cold (i, j) |. Its intuitive meaning is: compare the proportion of contact between site i and all potential neighbors j in the heat-resistant group and the cold-adapted group item by item; the greater the change, the greater the contribution. Finally, summing these proportions gives the overall change magnitude of site i. Therefore, even if deg hot (i) with deg cold (i) If the total contact strength appears to change little, and if site i mainly contacts A, B, and C in the heat-resistant group and mainly contacts D, E, and F in the cold-adapted group (i.e., the contact objects are systematically replaced), or if the proportion of the same batch of contact edges in the two groups changes significantly, then rewire(i) will still be large, thus indicating that the site may be in a temperature-related structural rearrangement position, which can be used as a supplementary correction parameter for the contact strength index.

[0114] (iii) The core and bridging nature of contact networks: To transform the probability matrix into more easily interpretable structural network topology features, the contact probabilities are thresholded for mapping: when P hot When (i, j) is not lower than the contact edge threshold (e.g., 0.30), residues i and j are connected in the heat-resistant contact network to form an edge (i, j); when Pcold When (i, j) is not lower than the contact edge threshold (e.g., 0.30), residues i and j are connected to form edges (i, j) in the cold-adapted group contact network. The resulting network can be understood as retaining only the more common or stable contact connections within this set of structures, used to describe the network skeleton and the roles of key nodes. See also Figure 10 , Figure 10 The diagram illustrates an example of the mapping of the core residue contact networks of the hot and cold groups onto the reference UDG structure, which can intuitively locate the spatial clusters of temperature-related contact reconnections.

[0115] Through the above steps, this embodiment can obtain a set of contact relationship difference indicators for each candidate site, including deg hot (i), deg cold (i) and rewire(i), etc.

[0116] (3) Quantification of Salt Bridge Network Difference Indicators In protein structures, salt bridges can form between oppositely charged side chains. These interactions are directional and have strict distance requirements. When multiple salt bridges are spatially interconnected, they often form an electrical constraint network on local structures or across secondary structure elements, thus affecting structural stability, local flexibility, and temperature-adaptive conformational distribution. To further capture the rearrangement differences of the electrical constraint network of candidate sites under cold and hot controls, this embodiment performs structural ensemble statistics on salt bridge relationships and forms a salt bridge network difference index.

[0117] In a preferred embodiment, the quantification of salt bridge network differences includes the following steps: a) Single-structure salt bridge determination: In each aligned structure sample, the set of charged residues is screened. For example, aspartic acid (Asp, D) and glutamic acid (Glu, E) are considered as sources of negatively charged residues, and lysine (Lys, K) and arginine (Arg, R) are considered as sources of positively charged residues. Optionally, histidine (His, H) is included in the statistics as a weakly charged residue. For any negatively charged residue and any positively charged residue, if there is an atomic pair of no more than 4.0 Å between the charged heavy atoms of their side chains, then the residue pair is determined to form a salt bridge in the structure sample.

[0118] b) Construction of the intra-group salt bridge probability matrix: Similar to contact relationship statistics, in both the hot and cold groups, if a certain residue pair (i,j) and its distance meet the salt bridge determination rules, then a salt bridge edge (i,j) for that residue pair can be determined to exist. Further, the frequency of each salt bridge edge (i,j) appearing in the structures of different proteins within each group is counted, and this frequency is normalized by dividing by the number of proteins in that group to obtain the salt bridge probability matrix P-salt. hot (i, j) and P-saltcold (i, j). This probability can be understood as what proportion of structural samples in the heat-resistant or cold-adapted group show that (i, j) forms a salt bridge.

[0119] c) Salt Bridge Occupancy and Node Metrics: After obtaining the salt bridge probability matrix, the following node-level salt bridge metrics are extracted for candidate sites: (i) Salt occupancy (salt-occ): Within the same temperature group, for all available structural samples at coordinate i of site i, it is determined whether the site forms a salt bridge with at least one residue of opposite charge. The proportion of samples forming a salt bridge to the total number of samples is used as the occupancy, and the salt occupancy of the heat-resistant group is obtained. hot (i) Salt bridge occupancy (salt-occupancy) compared to the cold-adapted group cold (i). This index reflects the frequency and stability of the site in the form of salt bridges in this temperature group (the higher the occupancy, the more structural samples the site participates in salt bridge formation).

[0120] (ii) Salt degree: In both the cold and hot groups, the salt degree (salt-deg) of the heat-resistant group is obtained by summing the salt bridge probabilities of site i with all other residues. hot (i) = ∑ j P-salt hot (i, j) and the salt-deg salt bridge degree of the cold adaptation group cold (i) = ∑ j P-salt cold (i, j) can be understood as: how many salt bridge connections can site i maintain on average in this set of structures, or how stable the salt bridge connections are. The larger the value, the more active the site is in the salt bridge network, the stronger the electrical constraints it carries, and the more significant its contribution to local or overall stability.

[0121] (iii) Salt rewiring amplitude: To characterize how much the salt bridge neighborhood of site i changes between the hot and cold groups, the differences in the salt bridge formation probability between site i and each residue j are compared between the hot and cold groups, and the absolute values ​​of each difference are summed to obtain the salt bridge rewiring amplitude: rewire-salt(i) = ∑ j | P-salt hot (i, j) – P-salt cold(i, j) |. Its intuitive meaning is: compare the proportion of salt bridge occurrences at site i with all potential salt bridge partners j in both groups; the greater the change, the greater the contribution. Finally, summing these proportions yields the overall change in the salt bridge neighborhood of site i. Therefore, even if the salt bridge reconnection amplitude in the heat-resistant group is salt-deg hot (i) Salt bridge reconnection amplitude with the cold adaptation group (salt-deg) cold (i) If the total salt bridge intensity appears to change little, and the main salt bridge object at site i is replaced between the two groups, or the proportion of the same batch of salt bridge edges shifts significantly between the two groups, then rewire-salt(i) will still be large, thus indicating that the site may be in a temperature-related electrical network rearrangement position, which can be used as a supplementary correction parameter for the salt bridge intensity index.

[0122] (iv) Coreness and Bridging of Salt Bridge Networks: To transform the salt bridge probability matrix into more easily interpretable electrical network topology features, the salt bridge formation probability is also thresholded. Specifically, in the cold and hot groups, when the salt bridge formation probability P-salt of the heat-resistant group... hot (i, j) or the probability of salt bridge formation in the cold adaptation group, P-salt cold When (i, j) is not lower than the salt bridge edge threshold (e.g., 0.10), residues i and j are connected in both sets of salt bridge networks, and a salt bridge edge (i, j) is introduced. The resulting network can be understood as retaining only the more common or stable salt bridge connections within that set of structures, thus characterizing the skeletal structure of the salt bridge network and the core and bridging role of key nodes in electrical constraints. Since salt bridges are sparser and more temperature-sensitive than general contacts, using a lower edge threshold (e.g., 0.10) helps improve the stability and comparability of network topology metrics while ensuring interpretability.

[0123] See Figure 11 , Figure 11 The diagram illustrates an example of the mapping of the core salt bridge networks of the cold and hot groups onto the reference UDG structure, which can intuitively locate the spatial clusters of temperature-related salt bridge reconnections.

[0124] Through the above steps, this embodiment can obtain a set of salt bridge network differential indicators for each candidate site, including salt-occ hot (i) salt-occ cold (i) salt-deg hot (i) salt-deg cold (i) and rewire-salt(i), etc.

[0125] (4) Quantification of dynamic coupling difference index based on elastic network model In a preferred embodiment, this example uses ProDy (ENM analysis suite) to construct a Cα-based Elastic Network Model (ENM) for protein structure samples aligned to the reference UDG coordinate system in both the cold and hot groups. Based on this, the predicted flexibility of candidate sites and their coupling strength to the functional core are quantified. It should be noted that ENM reflects the relative flexibility trend and cooperative motion relationship of the structure near its equilibrium conformation, and does not rely on long-time trajectory integration. It is a conventional method for consistent coarse-grained comparative statistics of a large number of structural samples in molecular dynamics analysis.

[0126] In a preferred embodiment, the quantification of the ENM kinetic index includes the following steps: a) ENM network construction: Using the Cα atoms of each residue as network nodes, when the Cα-Cα distance between any two nodes does not exceed 10.0 Å, a spring connection is introduced between them to form an elastic network.

[0127] b) Low-frequency collective motion and correlation estimation: In this embodiment, 325 K (approximately 50 °C) was used as the calibration temperature. Rigid body modes (zero-frequency modes) corresponding to overall translation and rotation were first eliminated on the aforementioned elastic network. Then, the remaining modes were sorted by frequency from low to high, and the top 20 lowest-frequency non-zero modes were selected as approximate representations of the protein's main collective motion. It should be noted that a mode refers to a characteristic overall motion pattern that a protein may exhibit under ENM prediction; it can be understood as a motion pattern of a group of residues coordinating displacements in a fixed proportion. Different modes correspond to different overall deformation patterns. Low-frequency modes typically represent collective motions with larger spans and more structural segments involved, and are better able to reflect long-range coupling and overall linkage. Rigid body modes refer to the motion patterns in which the protein as a whole undergoes translation or rotation. Such motions do not change the relative configuration between residues or reflect internal deformation, and are therefore eliminated in the analysis.

[0128] c) ENM Predictive Flexibility (ENM-RMSF): For both the cold and hot groups, obtain the ENM-RMSF based on the ENM output for each group. hot (i) with ENM-RMSF cold(i) Index. ENM-RMSF(i) represents the theoretical fluctuation amplitude of a site i under the low-frequency collective motion mode predicted by ENM. This index is calculated based on the superposition effect of the top 20 lowest-frequency non-zero modes, reflecting the intrinsic flexibility of the site near the equilibrium conformation dominated by collective motion. Unlike eRMSF, which is based on structural set statistics, ENM-RMSF is derived from the mechanical model prediction of a single structure. It can capture the intrinsic dynamic characteristics determined by protein topological connectivity and spatial constraints, and can complement the aforementioned cross-structure set dispersion eRMSF: eRMSF reflects the degree of static coordinate dispersion among different structural samples, while ENM-RMSF reflects the dynamic fluctuation trend determined by the topological network within a single structure.

[0129] d) Functional Core Coupling (ENM-coupling): To quantitatively assess the kinetic association between candidate sites and the catalytic functional core region, this embodiment preferentially defines the functional core set, Core, in the reference UDG coordinate system. Specifically, the UDG-DNA complex crystal structure (e.g., PDB:1EMH) is aligned to the reference structural coordinate system, and the reference residue position with a minimum spatial distance of no more than 6.0 Å from any atom of DNA or uracil bases is defined as Core. This definition ensures that the Core set encompasses the active site residues directly involved in substrate recognition and catalytic reactions, as well as their immediately adjacent first-layer functionally related residues.

[0130] Furthermore, for both the cold and hot groups, the ENM-coupling values ​​based on ENM calculations were obtained for each group. hot (i) and ENM-coupling cold (i) Index. ENM-coupling(i) represents the kinetic coupling strength between a site i and the functional core during low-frequency collective motion. The physical meaning of this index is: when site i is displaced during collective motion, the direction and amplitude of its motion are correlated with the motion of residues in the Core region. Sites with high coupling strength are more likely to have conformational perturbations that affect the conformational state of the functional core region through the mechanotransmission pathway within the protein, thus possessing a stronger potential regulatory ability on catalytic activity or substrate binding; conversely, sites with low coupling strength are relatively decoupled from the functional core kinetically, and their mutations pose a lower risk of long-range perturbation to the functional core.

[0131] e) Summary and Quantification of Differences in Hot and Cold Comparisons: To ensure comparability of the two groups of ENM indices under the reference residue numbering system, this embodiment performs intra-group statistical summaries for the heat-resistant group and the cold-adapted group respectively. Specifically, for any reference site i, the ENM-RMSF value and ENM-coupling value calculated for each structural sample in the heat-resistant or cold-adapted group are collected. After removing structural samples that are missing or unusable at that site, the arithmetic mean of the remaining valid samples is taken to obtain the group mean ENM-RMSF-mean for that site. hot (i) and ENM-RMSF-mean cold (i), and ENM-coupling-mean hot (i) and ENM-coupling-mean cold (i). The in-group averaging strategy can eliminate the random fluctuations of individual structural samples and obtain a more robust estimate of the dynamic characteristics at the level of the structure set in this temperature group.

[0132] See Figure 13 , Figure 13 (A) shows the distribution of the mean curves of ENM-RMSF along the reference residue number in the heat-resistant group and the cold-adapted group. Figure 13 (B) shows the distribution of the mean curves of ENM-coupling in the two groups. By comparing the two sets of curves, the trends of flexibility differences and core linkage differences between the cold and hot groups at different sites can be intuitively identified.

[0133] After obtaining the group mean, in order to further characterize the direction and intensity of the differences in predicting flexibility and core coupling among candidate sites under cold and hot control, this embodiment calculates the difference: (i) ΔENM-RMSF(i) is taken as ENM-RMSF-mean cold (i) - ENM-RMSF-mean hot (i) This differential value is used to characterize the magnitude and direction of changes in the predicted flexibility of site i in the context of collective motion. A positive ΔENM-RMSF(i) indicates that the cold-adapted group exhibits higher intrinsic flexibility at this site, meaning that this site is more prone to fluctuations in collective motion patterns among cold-adapted proteins; a negative value indicates that the heat-resistant group exhibits higher intrinsic flexibility at this site or that the cold-adapted group is more rigid. This index can be used to identify sites where flexibility undergoes systematic changes during temperature adaptation.

[0134] (ii) ΔENM-coupling(i) is taken as ENM-coupling-mean cold (i) - ENM-coupling-mean hot (i) This differential value is used to characterize the temperature-related changes in the strength of the kinetic coupling between site i and the functional core. A positive ΔENM-coupling(i) indicates a stronger kinetic coupling between the cold-adapted group and the functional core, meaning that perturbations at this site are more easily transmitted to the functional core in cold-adapted proteins; a negative value indicates a stronger core coupling at this site in the heat-resistant group or greater decoupling in the cold-adapted group. This indicator is of significant reference value for assessing the safety of mutations: sites with larger absolute values ​​of ΔENM-coupling(i) show significant changes in the functional transduction pathway between the cold and heat groups, suggesting that this site may be a key node in the kinetic network remodeling during temperature adaptation.

[0135] Through the above steps, this embodiment can obtain a set of dynamic coupling difference indices based on ENM for each candidate site, mainly including ΔENM-RMSF(i) and ΔENM-coupling(i).

[0136] (5) Quantification of dynamic coupling difference index based on explicit solvent molecular dynamics simulation To supplement the limitations of the ENM linear approximation model and obtain dynamic evidence of candidate sites on a time sampling scale, this embodiment uses the OpenMM molecular simulation suite to perform explicit solvent molecular dynamics (MD) simulations on protein structure samples aligned to the reference UDG coordinate system in both the hot and cold groups. This allows for the quantification of the dynamic flexibility of candidate sites and their time-dependent coupling strength to the functional core. Unlike the harmonic approximation of ENM based on equilibrium conformation, MD simulations, through explicit calculations of interatomic interactions and numerical integration of Newton's equations of motion, can capture the true thermal trajectories of proteins at finite temperatures, including anharmonic dynamic features that ENM cannot describe, such as local conformational transitions, side-chain rearrangements, and solvation effects. To ensure the comparability of results between the hot and cold groups, the heat-resistant group and the cold-adapted group preferably use the same force field, solvent model, simulation temperature, and sampling procedure for MD sampling, and differences are amplified under uniform thermal stress conditions.

[0137] In a preferred embodiment, the MD dynamic index quantification includes the following steps: a) System Construction and Simulation Setup: Using the UDG monomer conformation of each structural sample as the initial structure, protein force fields (e.g., AMBER ff19SB) were used for parameterization, and an explicit water model (e.g., OPC water model) was used to construct the solvent environment. The protein was placed in a periodic boundary condition water box, with the minimum buffer distance from the water box boundary to any atom of the protein set to 10 Å, and ions were added as needed to achieve electroneutrality. In the non-bonded interaction treatment, the Ewald method (Particle Mesh Ewald, PME) was used for long-range electrostatics, and the short-range cutoff distance was set to 1.0 nm; constraints were applied to the hydrogen bond length to support a larger integration step size, which was set to 2.0 fs. To provide a control under uniform thermal stress conditions, the production simulation temperature was set to 325 K (approximately 50 °C) in this embodiment. This temperature is between the high-temperature stress region of cold-adapted proteins and the suitable temperature region of heat-resistant proteins, which can simultaneously apply moderate thermal perturbation to both groups of proteins, thereby amplifying their structural dynamic differences. The simulation was performed under NPT conditions for equilibration and sampling, with a target pressure set at 1.0 bar. The simulation process for each structural sample included energy minimization, short-term equilibration (approximately 100 ps), and production sampling (approximately 300 ps); the trajectory saving interval was preferably 1 ps to obtain sufficient temporal resolution for subsequent statistical analysis.

[0138] b) Repeated Simulation and Trajectory Alignment: To reduce the impact of initial velocity randomness on statistical results, each structural sample undergoes multiple independent repeated simulations (e.g., 5 repetitions), with different random seeds used to generate initial velocities in each repetition. The mean or median of each repetition's results is taken as the representative MD index value for that structural sample. To ensure the comparability of trajectories across samples and repetitions, this embodiment performs rigid body alignment of the production stage trajectory frame by frame with the core folding region to the reference UDG coordinate system to eliminate the influence of overall translation and rotation; and a preheating burn-in time (e.g., 50 ps) is set before index statistics are performed to exclude unsteady fluctuations in the initial equilibrium stage.

[0139] c) MD Dynamic Flexibility (MD-RMSF): For both the cold and hot groups, the MD-RMSF based on the MD trajectory output is obtained for each group. hot (i) with MD-RMSF cold(i) Index. MD-RMSF(i) represents the fluctuation amplitude of a site i relative to its time-averaged position in the MD simulation time series. This index is calculated based on the time variance of the Cα atom coordinates of that site in the trajectory, directly reflecting the true thermal motion intensity of that site at finite temperatures. Compared with ENM-RMSF (Collective Mode Predicted Flexibility), MD-RMSF includes motion contributions at all time scales, including not only slow collective motion but also rapid local fluctuations and side-chain conformational transitions, thus providing a more comprehensive characterization of the site's dynamic flexibility. The complementarity between the two lies in the fact that ENM-RMSF reflects the inherent flexibility trend determined by the structural topology, while MD-RMSF reflects the dynamic behavior actually observed under explicit solvent and finite temperature conditions.

[0140] d) Functional Core Coupling Strength (MD-coupling): To focus on functionally relevant regions of the UDG, this embodiment adopts the definition of the functional core set (Core) in the ENM analysis described above (reference residue positions whose minimum distance from any atom of the DNA or uracil site does not exceed 6.0 Å). Further, the temporal correlation strength between candidate site i and the functional core set (Core) is calculated based on the MD trajectory to obtain the MD-coupling strength. hot (i) with MD-coupling cold (i). MD-coupling(i) represents the motion correlation between a site i and the functional core in the MD simulation time series. This index is obtained by calculating the cross-correlation coefficient or covariance between the Cα atomic shift vector of site i and the Cα atomic shift vector of residues in the Core region, reflecting the degree of synchronous motion between site i and the functional core in the time dimension. A high MD-coupling value indicates that the thermal motion of the site is highly coordinated with the functional core, and its perturbation is more likely to affect the conformational state of the functional core through kinetic transmission; a low MD-coupling value indicates that the site and the functional core are relatively independent in kinetics, and its mutation has a lower risk of long-range perturbation of the functional core.

[0141] e) Cold / Hot Comparison Summary and Difference Quantification: To ensure comparability of the two groups of MD indices under the reference residue numbering system, this embodiment performs intra-group statistical summaries for the heat-resistant group and the cold-adapted group respectively. Specifically, for any reference site i, the MD-RMSF value and MD-coupling value of each structural sample (and its repeated summaries) within the heat-resistant or cold-adapted group are collected. After removing structural samples that are missing or unusable at that site, the arithmetic mean of the remaining valid samples is taken to obtain the group mean MD-RMSF-mean for that site. hot (i) and MD-RMSF-mean cold (i), and MD-coupling-mean hot(i) and MD-coupling-mean cold (i).

[0142] See Figure 14 , Figure 14 (A) shows the distribution of the mean curves of MD-RMSF along the reference residue number in the heat-resistant and cold-adapted groups. Figure 14 (B) illustrates the distribution of the mean curves for MD-coupling in the two groups. By comparing the two sets of curves, the site distribution patterns of dynamic flexibility differences and functional core linkage differences between the cold and hot groups under MD simulation conditions can be identified.

[0143] After obtaining the group mean, in order to further characterize the direction and intensity of differences in dynamic flexibility and core coupling of candidate sites under cold and hot control, this embodiment calculates the amount of difference: (i) ΔMD-RMSF(i) is taken as MD-RMSF-mean cold (i) - MD-RMSF-mean hot (i) This difference measure is used to characterize the amplitude and direction of temporal fluctuations at this site. A positive ΔMD-RMSF(i) indicates that the cold-adapted group exhibits stronger thermal motion fluctuations at this site, meaning that the site is more flexible at finite temperatures among cold-adapted proteins. A negative value indicates that the heat-resistant group exhibits stronger fluctuations at this site, or that the cold-adapted group is more rigid. Combining the flexibility indices at both the ENM and MD levels allows for a more comprehensive assessment of the consistency of site flexibility changes: if ΔENM-RMSF(i) and ΔMD-RMSF(i) have the same sign and comparable amplitude, it indicates that the flexibility differences at this site are mutually corroborated by static topological predictions and dynamic simulation observations, thus possessing higher reliability.

[0144] (ii) ΔMD-coupling(i) is taken as MD-coupling-mean cold (i) - MD-coupling-mean hot (i) This differential measure is used to characterize the temperature-dependent changes in the temporal linkage strength between site i and the functional core. A positive ΔMD-coupling(i) indicates that the cold-adapted group has stronger temporal linkage between the site and the functional core, meaning that perturbations at this site are more easily transmitted to the functional core via kinetic pathways in cold-adapted proteins; a negative value indicates that the thermotolerant group has stronger core linkage at this site or that the cold-adapted group is more decoupled. Similar to ΔENM-coupling(i), ΔMD-coupling(i) is of significant reference value for assessing the functional impact risk of mutations, and calculations based on MD trajectories can capture anharmonic coupling and solvent-mediated long-range interactions that ENM cannot describe.

[0145] Through the above steps, this embodiment can obtain a set of dynamic coupling difference indicators based on MD for each candidate site, mainly including ΔMD-RMSF(i) and ΔMD-coupling(i), which together with the ENM index in step (5) serve as dynamic evidence to support the subsequent selection and ranking of key sites.

[0146] (6) Construction of the multi-source structural index matrix of candidate sites Through steps (1) to (6), this embodiment generates structural evidence quantification results for each candidate site, which are then summarized into a multi-source structural index matrix for candidate sites. This matrix uses candidate sites as row indexes and various structural indicators as column indexes, systematically organizing the multi-dimensional quantitative difference characteristics of 38 candidate sites between the hot and cold groups.

[0147] In a preferred embodiment, the multi-source structural index matrix includes at least the following index categories: a) Site identification: reference residue number, amino acid letter; b) Structural set dispersion class: heat-resistant group structural dispersion eRMSF hot (i) Cold adaptation group structural dispersion eRMSF cold (i) The structural dispersion difference ΔeRMSF(i) between the hot and cold groups; c) Contact Relationships and Networks: Contact Degree of Heat-Resistant Residues (deg) hot (i) Cold adaptation group residue contact degree deg cold (i) Reconnection amplitude of cold and hot residues; d) Salt bridge network type: Salt-occupancy of heat-resistant group salt bridges hot (i) Salt bridge occupancy in the cold adaptation group (salt-occupancy) cold (i) Salt bridge degree of heat-resistant group (salt-deg) hot (i) Salt bridge degree (salt-deg) of the cold adaptation group cold (i) Rewire-salt(i) amplitude of salt bridge reconnection in cold and hot groups; e) ENM dynamics class: Cold and hot elastic network model predicts flexibility difference ΔENM-RMSF(i), cold and hot elastic network model functional core coupling strength difference ΔENM-coupling(i); f) MD dynamics class: dynamic flexibility difference ΔMD-RMSF(i) in cold and hot molecular dynamics simulation, functional core coupling strength difference ΔMD-coupling(i) in cold and hot molecular dynamics simulation.

[0148] The structural evidence quantification matrix will be used in subsequent steps for fusion evaluation and ranking. Multi-source indicators provide a basis for generating lists of preferred sites based on different engineering focuses, such as emphasizing structural safety, temperature-related plasticity differences, or dynamic coupling with functional cores. This improves the operability and adaptability of rational design for thermosensitive UDGs without expanding the mutation space. The technical value of this matrix lies in transforming a temperature-ecologically driven set of hot and cold control structures into site-level, rankable, interpretable, and cross-validable structural difference indicators through multi-level quantitative calculations under a unified coordinate system, providing a systematic data foundation for subsequent comprehensive evaluation.

[0149] 5. Multi-evidence fusion evaluation, spatial clustering, and rational mutation site output After obtaining the multi-source structural index matrix of 38 candidate sites in step (4), this step further normalizes, hierarchically weights and merges and sorts the structural evidence from different sources, and clusters the candidate sites in the three-dimensional structural space to form a priority list and site combination output that can be directly used for mutation design.

[0150] (1) Unification and normalization of multi-source indicators Because different indicators have different dimensions and value ranges, this embodiment uses rank-normalization to map each indicator to a comparable expected score space. Specifically, for the values ​​of 38 candidate sites on a certain indicator, they are first sorted by numerical value, and then the ranking position is converted into a normalized score between 0 and 1. The highest-ranked site receives a score of 1, the lowest-ranked site receives a score of 0, and the scores of intermediate sites are linearly interpolated according to their ranking. The advantages of this method are: it can automatically eliminate the differences in dimensions and numerical scales between different indicators, making all indicators comparable in subsequent weighted fusion; at the same time, it is robust to outliers, avoiding excessive influence of individual extreme values ​​on the overall ranking.

[0151] For indicators where "the larger the value, the better" (such as temperature difference magnitude indicators like rewire(i) and ΔeRMSF(i), the closer the normalized score is to 1, the better. For indicators where "the smaller the value, the better" (such as coupling / linkage indicators with the functional core, degree of burial, or number of contacts contact10(i), the normalization direction is reversed (i.e., 1 minus the normalization result), and they are uniformly converted into a score form where "the larger the value, the better" for subsequent weighted summation.

[0152] (2) Hierarchical weighted fusion scoring system based on effect-safety dual dimensions To simultaneously consider both "temperature-driven controllability (effect)" and "mutation safety (safety)," this embodiment constructs two sub-scores and fuses them at the top level to form a comprehensive evaluation score. This two-dimensional scoring system reflects the core principles of rational design for thermosensitive UDGs: on the one hand, candidate sites should exhibit significant and consistent structural or kinetic differences under cold and hot controls, indicating that the site has a strong potential regulatory capacity for temperature response performance (effect dimension); on the other hand, candidate sites should be far from the catalytic core and relatively decoupled kinetically from the functional region to ensure that mutations do not cause unacceptable damage to enzyme activity and substrate recognition (safety dimension).

[0153] a) Construction of Effect-score The effector score is used to characterize whether candidate sites exhibit consistent and significant structural or kinetic differences under cold and hot controls. The more significant the difference, the greater the potential regulatory leverage for engineering modification. The effect score consists of three parts: network effect, kinetic effect, and structural dispersion supplement, each of which is obtained by weighting the scores after normalization of the aforementioned indicators.

[0154] This embodiment adopts the following top-level fusion method: Effect-score(i) = 0.55×Effect-net(i) + 0.40×Effect-dyn(i) + 0.05×Effect-design(i) in: (i) Network effect: reflects temperature-related rearrangement at the network level, specifically: Effect-net(i) = 0.55×rewire(i) + 0.2×rewire-salt(i) + 0.25×[salt-occ hot (i) - salt-occ cold (i)] In this section, rewire(i) and rewire-salt(i) reflect the reconnection magnitude of the residue contact network and the salt bridge network, respectively. Their higher weights reflect the core role of network rearrangement in temperature adaptation; the difference in salt bridge occupancy [salt-occ]... hot (i) - salt-occ cold [(i)] reflects the change in the frequency of salt bridge formation at this site between the two groups. A positive value indicates that the heat-resistant group is more likely to form salt bridges (which is consistent with the mechanism by which heat-resistant proteins improve stability by enhancing surface salt bridges), while a negative value indicates that the cold-adapted group is more likely to form salt bridges.

[0155] (ii) Kinetic effect - dyn(i): reflects temperature-related differences at the kinetic level, specifically: Effect-dyn(i) = 0.45 × ΔMD-RMSF + 0.3 × ΔMD-coupling + 0.25 × ΔENM-RMSF (structural mobility) In this category, ΔMD-RMSF(i) has the highest weight, reflecting the advantage of MD simulation in capturing real thermal motion; ΔMD-coupling(i) and ΔENM-RMSF(i) provide supplementary evidence from the perspectives of time-dependent coupling and collective modal flexibility, respectively. The combined use of these three indicators can comprehensively evaluate the temperature response characteristics of a site from both static topology and dynamic behavior perspectives.

[0156] (iii) Structural Discreteness Supplement Effect-Design(i): This is used to provide supplementary signals for "temperature-dependent plasticity differences" in addition to the network and dynamics, specifically: Effect-design(i) = | ΔeRMSF | The absolute value of this item is taken because regardless of whether the cold-adapted group or the heat-resistant group is more flexible, as long as there is a significant difference, it indicates that the plasticity of this site has undergone a systematic change during the temperature adaptation process, and has the potential for regulation.

[0157] b) Construction of the Safety-score Safety sub-scores are used to characterize whether candidate sites pose a low potential perturbation risk to the functional core or substrate binding as mutation sites. This embodiment comprehensively evaluates safety from both kinetic and geometric perspectives: kinetic safety emphasizes that the weaker the coupling with the functional core, the safer it is; geometric safety emphasizes that the farther away from the DNA substrate and the more surface-oriented it is, the safer it is.

[0158] This embodiment adopts the following top-level fusion method: Safety-score(i) = 0.60×Safety-dyn(i) + 0.40×Safety-geom(i) in: (i) Dynamic safety-dyn(i): reflects the risk of this site acting as a mutation point to the dynamic perturbation of the functional core (the lower the coupling, the safer), specifically: Safety-dyn(i) = 0.65×ΔMD-coupling(i) + 0.35×ΔENM-coupling(i) (ii) Geometric safety (i): Reflects the safety of the site being farther away from the DNA substrate and more surface-level, specifically: Safety-geom(i) = 0.6×dist DNA (i) + 0.4×contact10(i) c) Final score and global ranking The comprehensive score is used to form the final priority list. In this embodiment, the top-level fusion is based primarily on effects and secondarily on safety to obtain the final ranking criteria. Final-score(i) = 0.70×Effect-score(i) + 0.30×Safety-score(i) This weighting reflects the "temperature response effect priority, safety assurance" strategy in thermosensitive UDG design: 70% of the weight is allocated to the effector score to ensure that the selected sites have strong temperature regulation potential; 30% of the weight is allocated to the safety score to ensure that the selected sites will not cause serious damage to enzyme function. This weighting ratio can be adjusted according to specific engineering goals: if more emphasis is placed on mutation safety, the weight of the Safety-score can be increased; if more emphasis is placed on temperature response effect, the weight of the Effect-score can be increased.

[0159] The 38 loci were ranked in descending order based on their combined scores obtained from the Final-score(i) to form a comprehensive priority list of candidate loci. Specifically, the comprehensive scores of the 38 loci, from highest to lowest, are as follows: 43Y, 11G, 55E, 174M, 155R, 154E, 173E, 178N, 42M, 12L, 176N, 177M, 34A, 151C, 150S, 102G, 214G, 146D, 83G, 216E, 27E, 30K, 35H, 36H, 23K, 219D, 24E, 213I, 212S, 31Q, 33Y, 211K, 41D, 138G, 199G, 208Q, 101L, 28F.

[0160] See Figure 15 , Figure 15 The diagram illustrates the overall score ranking of the 38 sites.

[0161] (3) Cross-validation integration of sequence-level temperature difference evidence and structural-level candidate sites In the process of extracting temperature difference features at the sequence level, this embodiment initially screened out four candidate sites with strong difference signals, namely 41D, 176N, 185S, and 223S. Furthermore, after initial screening of candidate sites driven by structure avoidance and preference rules, and quantification and comprehensive scoring of multi-source structural evidence, this embodiment obtained a comprehensive ranking result of 38 candidate sites.

[0162] Cross-integration of the two sets of results reveals that two of the four sequence-level loci overlap with the set of 38 loci: 41D and 176N. The remaining two loci (185S and 223S) failed to enter the set of 38 loci because they did not simultaneously meet the structural designability constraints (e.g., low occupancy or location in the terminal avoidance region). Therefore, 41D and 176N can be considered cross-candidate loci supported by both the "sequence difference signal" and "structural designability or quantifiable evidence" chains of evidence.

[0163] The technical value of this cross-validation lies in the fact that the amino acid composition differences at the sequence level reflect the systematic preference differentiation between the hot and cold groups of proteins at this site on an evolutionary scale, while the multi-source evidence at the structural level reflects the temperature-related differences in the three-dimensional space and dynamic behavior of this site. The cross-support of the two chains of evidence significantly improves the credibility of this site as a temperature-regulating target and reduces the risk of false positives. Therefore, in the subsequent final site output, 41D and 176N should be given priority and included in the preferred mutation site set.

[0164] (4) Cluster analysis of candidate sites based on spatial proximity Because high-scoring sites may be highly clustered in three-dimensional space (multiple residues within the same local structural patch simultaneously score high), selecting only the top N sites based on the overall score can easily lead to spatial redundancy. This results in the final output mutation sites being overly concentrated in a certain structural region, while other potential temperature-regulating regions remain uncovered. To address this issue, this embodiment performs spatial clustering on 38 sites to characterize their clustering relationships in three-dimensional structure, and takes into account the coverage of different spatial clusters when outputting subsequent sites.

[0165] In a preferred embodiment, spatial clustering is performed according to the following process: a) Spatial distance matrix calculation: Extract the Cα atom coordinates of each candidate site from the reference UDG structure, and calculate the Cα-Cα Euclidean distance d(i, j) between any two candidate sites (i, j) to form a 38×38 spatial distance matrix.

[0166] b) Proximity determination: A spatial distance threshold of 8.0 Å is set. When d(i, j) does not exceed 8.0 Å, site i and site j are determined to be a spatially adjacent pair, and a connection edge is established between them. This 8.0 Å threshold is based on the typical interaction distance range in protein structures and can effectively identify residues located in the same local structural patch.

[0167] c) Connectivity Clustering: An undirected graph is constructed based on the above proximity relationships, where candidate sites are nodes and the connections between neighboring pairs are edges. Clustering is performed using connectivity rules: if site i can be reached via one or more nearest neighbors to site j (i.e., i and j are connected in the graph), then i and j belong to the same spatial cluster; if there is no reachable path between two groups of sites (i.e., they are not connected in the graph), then they are divided into different spatial clusters.

[0168] After clustering all 38 sites according to the above rules, 9 spatial clusters were finally obtained. See also Figure 16 , Figure 16 The distribution of the 38 candidate sites in each spatial cluster is illustrated. Furthermore, the site with the highest overall score within each spatial cluster can be defined as the representative site of that cluster, and the composition of the cluster members (listed in descending order of overall ranking) can be summarized for subsequent optimal selection within the cluster.

[0169] (5) Determination of rational mutation sites and mapping output to the engineering seed UDG After completing the comprehensive sorting of 38 sites and the partitioning of 9 spatial clusters, this embodiment further determined the set of sites for the rational design of thermosensitive UDG mutations. The determination of this set of sites comprehensively considered the following factors: a) Prioritize comprehensive score: Prioritize candidate sites with the highest Final-score(i) in the comprehensive ranking to ensure that the selected sites have a strong temperature response effect and high mutation safety; b) Spatial cluster coverage: Taking into account the coverage of different spatial clusters, avoiding the excessive aggregation of selected sites in a certain local area in three-dimensional space, and ensuring that the output sites are distributed in multiple functional regions of the UDG structure; c) Prioritize cross-validation sites: Prioritize the inclusion of cross-validation sites (i.e., 41D and 176N) that simultaneously receive dual support from both "sequence-level temperature difference evidence" and "structural-level multi-source evidence" to improve the credibility and theoretical support of site selection.

[0170] Based on the above principles, this embodiment finally identified five rational mutation sites, which are 41D, 43Y, 55E, 102G and 176N in the reference UDG (A0A161WTR3) residue numbering system.

[0171] The combined output of the five sites reflects a comprehensive optimization strategy of "efficacy priority, safety assurance, cross-validation, and spatial coverage," which can provide multi-target and multi-region mutation schemes for the rational design of thermosensitive UDGs.

[0172] (6) Site mapping to the engineering seed UDG To utilize the rational mutation sites obtained under the aforementioned "reference UDG residue numbering system" for practical engineering modification, this embodiment further selects a mesophilic UDG as the engineering seed protein and completes site mapping output. In this embodiment, the engineering seed UDG is *Escherichia coli* (E. coli). Escherichia coli The UDG source (UniProtKB accession number: P12295) was chosen as the starting template for low-temperature modification. This choice was based on the following considerations: First, E. coli UDG has a mature recombinant expression and purification system, which is easy to verify experimentally; second, the optimal catalytic temperature of this UDG is approximately 37 °C, classifying it as a mesophilic enzyme. By introducing cold-adaptation-related structural or kinetic features, it is expected to achieve enhanced low-temperature activity and thermal inactivation characteristics while maintaining basic catalytic activity; third, this UDG has a broad validation base in applications such as nucleic acid amplification contamination control, and has good engineering prospects.

[0173] In a preferred embodiment, site mapping is performed as follows: the amino acid sequence of the engineered seed UDG (P12295) is paired with the sequence of the reference UDG (A0A161WTR3) or the results of multiple sequence alignment constructed in step 2 are used to establish a correspondence between "reference UDG residue number and engineered seed UDG residue number" during the alignment; then the five rational mutation sites that are finally determined are projected from the reference numbering system to the numbering system of the engineered seed UDG and output in the form of residue numbers to facilitate subsequent mutation design and experimental construction.

[0174] In this embodiment, after mapping the rational mutation sites under the reference UDG numbering system to the engineering seed UDG (P12295), the corresponding output site numbers are 40, 42, 54, 101, and 175.

[0175] The rational mutation sites of the engineering seed UDG are: position 40 (corresponding to 41D of reference UDG), position 42 (corresponding to 43Y of reference UDG), position 54 (corresponding to 55E of reference UDG), position 101 (corresponding to 102G of reference UDG), and position 175 (corresponding to 176N of reference UDG).

[0176] See Figure 12 , Figure 12 The mapping of the final rational mutation sites on the three-dimensional structure of the engineered seed UDG (P12295) is shown, and the spatial distribution of the five sites in the UDG structure and their relative position to the catalytic pocket can be observed intuitively.

[0177] Through the above site mapping, this invention realizes a complete transformation path from temperature-ecological data-driven cold-heat comparison analysis, multi-source evidence quantification under a unified structural coordinate system, comprehensive evaluation ranking and spatial clustering, to the final engineering seed UDG mutation sites that can be directly used for experimental verification. It provides a systematic, reproducible and operable site selection method for the rational design and engineering optimization of thermosensitive UDGs.

[0178] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for screening thermosensitive UDG rational mutation sites based on strain culture temperature, characterized in that, Includes the following steps: (1) Based on the strain preservation database, heat-resistant strains and cold-adapted strains were constructed according to the strain culture temperature. Then, a set of heat UDG sequences and corresponding set of heat three-dimensional structures were constructed based on the heat-resistant strains, and a set of cold UDG sequences and corresponding set of cold three-dimensional structures were constructed based on the cold-adapted strains. (2) Combining the hot and cold differences in sequence, a first candidate mutation site set is generated based on the hot UDG sequence set and the cold UDG sequence set; (3) Align the three-dimensional structures in the hot three-dimensional structure set and the cold three-dimensional structure set to the reference structure coordinate system, and perform preliminary screening of the residue positions in the aligned three-dimensional structures on the reference structure coordinate system to obtain the second candidate mutation site set; (4) Combining the characteristics of thermal differences at the structural level, a set of mutation sites for rational mutation design of thermosensitive UDG is selected from the first set of candidate mutation sites and the second set of candidate mutation sites.

2. The method for screening thermosensitive UDG rational mutation sites based on strain culture temperature according to claim 1, characterized in that, In (1), based on the strain preservation database, heat-resistant strains and cold-adapted strains are constructed according to the strain culture temperature, including: Strains in the strain preservation database whose maximum growth temperature is not lower than the first preset temperature are designated as heat-resistant strains, and strains in the strain preservation database whose maximum growth temperature does not exceed the second preset temperature are designated as cold-adapted strains.

3. The method for screening thermosensitive UDG rational mutation sites based on strain culture temperature according to claim 1, characterized in that, The (2) includes: Based on the hot UDG sequence set and the cold UDG sequence set, multiple sequence alignment is performed on all UDG sequences to obtain sequence alignment results. Combining the amino acid information corresponding to the hot UDG sequence set and the cold UDG sequence set, the hot and cold difference features corresponding to each alignment column in the sequence alignment results are generated. Sites are selected according to the hot and cold difference features of the alignment column to obtain the first candidate mutation site set.

4. The method for screening thermosensitive UDG rational mutation sites based on strain culture temperature according to claim 1, characterized in that, In step (3), the residue positions in the aligned three-dimensional structure are initially screened on the reference structural coordinate system to obtain a second set of candidate mutation sites, including: Residue sites with structural alignment occupancy below a threshold in the aligned 3D structure are removed; residue sites in the neighborhood of the catalytic functional pocket in the aligned 3D structure are removed; residue sites in the structurally buried core region in the aligned 3D structure are removed; and surface or loop region residue sites in the aligned 3D structure are retained, thereby obtaining a set of second candidate mutation sites.

5. The method for screening thermosensitive UDG rational mutation sites based on strain culture temperature according to claim 1, characterized in that, In (4), the thermal difference characteristics at the structural level include structural set dispersion difference index, residue contact network difference index, salt bridge network difference index and dynamic coupling difference index.

6. The method for screening thermosensitive UDG rational mutation sites based on strain culture temperature according to claim 1, characterized in that, The (4) includes: The system generates thermal difference features at the structural level for each site in the second candidate mutation site set. It then calculates a score for each site based on these features and sorts the sites in the second candidate mutation site set accordingly, obtaining the site sorting results. Based on the spatial distance between sites, it clusters the sites in the second candidate mutation site set, obtaining the site clustering results. Combining the site sorting results, site clustering results, and the first candidate mutation site set, it selects a set of mutation sites from the second candidate mutation site set for rational mutation design of thermally sensitive UDG.

7. A thermosensitive UDG rational mutation site screening device based on strain culture temperature, characterized in that, include: Grouping units are used to construct heat-resistant strains and cold-adapted strains based on strain preservation databases and according to strain culture temperature. The dataset construction unit is used to construct a set of thermal UDG sequences and corresponding thermal three-dimensional structures based on heat-resistant strains, and to construct a set of cold UDG sequences and corresponding cold three-dimensional structures based on cold-adapted strains. The first site screening unit is used to combine the hot and cold difference features at the sequence level to generate a first candidate mutation site set based on the hot UDG sequence set and the cold UDG sequence set. The second site screening unit is used to align the three-dimensional structures in the hot three-dimensional structure set and the cold three-dimensional structure set to the reference structural coordinate system, and to perform preliminary screening of the residue positions in the aligned three-dimensional structures on the reference structural coordinate system to obtain the second candidate mutation site set. The third site screening unit, combining the thermal difference characteristics at the structural level, selects a set of mutation sites for rational mutation design of thermosensitive UDG from the first and second candidate mutation site sets.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the thermosensitive UDG rational mutation site screening method based on strain culture temperature as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the thermosensitive UDG rational mutation site screening method based on strain culture temperature as described in any one of claims 1 to 6.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the thermosensitive UDG rational mutation site screening method based on strain culture temperature as described in any one of claims 1 to 6.