DNA G-quadruplex structure prediction method based on sequence information

Through decision tree classifier and energy minimization technology, based on the known G4 structure data of the PDB database, a complete modeling method from sequence to structure is constructed, solving the problem of DNA G-four-strand full atomic structure prediction, and achieving efficient G4 structure prediction and disease diagnosis and treatment applications.

CN120452545AActive Publication Date: 2025-08-08SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510567159.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-08
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the full-atomic structure model of DNA G-quadrilateral directly from sequence information, and the experimental acquisition process is complex and inefficient, limiting the large-scale systematic research of G4 structure and the application of disease diagnosis and treatment.

Method used

Through the decision tree classifier, a complete modeling method from sequence to structure is constructed based on the known G4 structure data of the PDB database, including predicting base position parameters in the G-stem region, constructing G-loop region structure, generating alternative models using fragment assembly method, and screening stable models through energy minimization.

Benefits of technology

It realizes a complete prediction of the entire atomic structure from sequence information to DNA G-four strands, provides a deeper understanding of biological functions and the basis for disease diagnosis and treatment, and improves the availability and research efficiency of G4 structure data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452545A_ABST
    Figure CN120452545A_ABST
Patent Text Reader

Abstract

The invention discloses a DNA G-quadruplex structure prediction method based on sequence information. The DNA G-quadruplex structure prediction method comprises the following steps: S1) predicting a structure framework type based on the sequence information of DNA G4; s2) determining a mapping relation of base position parameters in the G-stem region based on the structure framework type, predicting base coordinates in the G-stem region, and constructing a full-atom structure model of the G-stem region; s3) based on the sequence information, carrying out modeling on the G-loop structure connected with each G-fragment, and constructing a plurality of alternative G4 full-atom models with complete structures by using a fragment assembly method; and S4) minimizing the energy of the alternative G4 full-atom model by simulating a relaxation process, and obtaining the G4 full-atom model with a stable structure through energy fraction screening. According to the invention, the full-atomic structure of G4 is predicted according to sequence information for the first time, and an important basis is provided for better understanding of biological functions of G4 and research on drug target discovery, nucleic acid probe design and the like based on the G4 structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of biological information technology and relates to a DNA G-quadruplex structure prediction method based on sequence information. Background Art

[0002] The DNA G4 (DNA G-quadruplex, hereinafter referred to as G4) structure is a hot topic in biological research. Existing research progress indicates that G4 plays an important role in regulating gene expression, maintaining genome stability, and participating in disease pathogenesis. Although the importance of G4 structure is widely recognized, detailed understanding of its structure is relatively limited. Furthermore, G4 structure is polymorphic and can be generated by polyguanylate formation within or between strands.

[0003] G4 is a non-standard nucleic acid secondary structure formed by the folding of a DNA sequence rich in guanine bases (G), and is mostly located at the end of telomere DNA and the promoter region of oncogenes. Figure 1 As shown, G4 is composed of four consecutive guanine segments (G-tracts, G-segments), and the four guanine bases from different segments form a stable planar arrangement - a guanine tetrad (G-tetrad, G-tetrad) through Hoogsteen hydrogen bonds; two or more G-tetrads form the core of the G4 structure - the stem region (G-stem, G-stem) through π-π stacking; there is a negatively charged cavity inside the stacked G-tetrads, and monovalent cations such as Na+ or K+ pass through the center of the G4 structure and coordinate with the O6 atom of guanine, thereby stabilizing the G4 structure; G-segments are connected to each other by a sequence of variable length - the loop region (G-loop, G-loop region), which is very important for the folding and structural stability of G4. Generally, the fewer bases in the G-loop region, the more stable the G4 structure; the DNA sequence that forms the G4 structure contains four G-segments separated by three G-loop regions.

[0004] Currently, the main methods for detecting G4 include circular dichroism (CD), X-ray diffraction by crystals, nuclear magnetic resonance (NMR), fluorescence spectroscopy (FS), and fluorescence resonance energy transfer (FRET). Although there are many methods available for obtaining structural information about G4, G4 structural data is still very limited. Currently, the PDB database contains only 517 G4 structural data. The Protein Data Bank (PDB) is a global scientific research database that mainly collects three-dimensional structural information of proteins, nucleic acids, and their complexes. These data are mainly derived from experimental techniques such as X-ray crystallography, nuclear magnetic resonance spectroscopy, and cryo-electron microscopy.

[0005] Internationally, recent advances in G4 structure prediction have primarily focused on the development and application of deep learning algorithms. While algorithms like G4detector and G4mismatch focus on learning and predicting the potential for G4 formation from sequence information, the ASC-G4 algorithm extracts detailed structural features from known structures, providing computational tools for G4 structure prediction and analysis. However, challenges remain in the field of G4 structure prediction and computational modeling. For example, how can the full-atom model of a G4 be predicted directly from sequence information? How can the spatial relative positions of the bases within a G4 be accurately restored? If the G4 structures formed by specific sequences could be accurately predicted, specific small molecule drugs or nucleic acid probes could be designed to target these structures, potentially leading to breakthroughs in areas such as cancer treatment and genetic disease diagnosis.

[0006] However, the experimental analysis of G4 structure is not only complex and technically demanding, but also inefficient. Currently, there are only about 500 G4 structural information items available in the PDB database. The scarcity of such data severely limits large-scale systematic research on G4 structure. Existing algorithms mostly remain at the classification prediction stage of G4 formation potential or topological structure. Currently, there is no method that can directly predict the full atomic structure model of G4 based on sequence information, let alone restore the details of the spatial structure. Summary of the Invention

[0007] Purpose of the invention: The technical problem to be solved by the present invention is to conduct an in-depth analysis of the G4 structural characteristics based on the known G4 structural data in the PDB database, and to construct a complete modeling method from sequence to structure. This can not only provide a new perspective for understanding the function of the G4 structure in the organism, but also has important significance for the development of new disease diagnosis and treatment methods based on the G4 structure. In view of the current situation that the experimental acquisition of G4 structure is not only complicated and inefficient, but also has little G4 structural data, and in view of the application needs of obtaining G4 structural information in the study of major diseases, the present invention provides a G4 structure prediction algorithm based on DNA sequence information, which conducts an in-depth study of the G4 structural characteristics through the known G4 structures in the PDB database, constructs a complete modeling method from sequence to structure, and realizes a complete prediction from G4 sequence to structure.

[0008] Technical solution: In order to solve the above technical problems, the present invention provides a method for predicting DNA G-quadruplex structure based on sequence information, comprising the following steps:

[0009] S1) Predicting the structural framework type based on the sequence information of DNA G4;

[0010] S2) determining a mapping relationship of base position parameters within the G-stem region based on the structural framework type, predicting base coordinates within the G-stem region, and constructing a full-atom structural model of the G-stem region;

[0011] S3) Modeling the structure of the G-loop region connecting each G-segment based on sequence information, and constructing multiple structurally complete candidate G4 all-atom models using the fragment assembly method;

[0012] S4) By simulating the relaxation process, the energy of the alternative G4 all-atom model is minimized, and a structurally stable G4 all-atom model is obtained through energy score screening.

[0013] The predicted structural framework types in step S1) include parallel structure, antiparallel structure-positive, negative, and negative, antiparallel structure-positive, negative, and negative, or 3+1 mixed structure.

[0014] The steps for obtaining the predicted structure framework type in step S1) are as follows:

[0015] S1-1) Collect all known DNA G4 structural data from the PDB database, including the DNA sequence of G4 and the atomic coordinates in the structure;

[0016] S1-2) Eliminate missing or incomplete records, extract the sequence characteristics and structural framework types of G4, use the sequence characteristics as the classification attributes of the samples, and the structural framework types as the classification results to construct a training sample set;

[0017] S1-3) Establishing a decision tree classifier model: The classifier model uses the sequence characteristics of DNA G4 as classification attributes, and outputs the classification result as the type of structural framework.

[0018] Among them, the sequence characteristics described in S1-2) include one or more of the number of G-quadruplexes, the number of G-loop regions, the length of the G-loop region, the proportion of A in the G-loop region, the proportion of T in the G-loop region, the proportion of C in the G-loop region, and the proportion of G in the G-loop region.

[0019] The model in S1-3) performs a five-fold cross-validation on the training sample set obtained in step S1-2): the training sample set is randomly divided into five equal parts, four of which are used as the training set for each training session, and the remaining one is used as the test set, and validation is performed five times. The performance of the decision tree model is evaluated for each type of structural framework. Let TP, TN, FP, and FN be the number of true positive samples, true negative samples, false positive samples, and false negative samples, respectively. Four metrics are obtained: accuracy, precision, recall, and F1 score:

[0020]

[0021] Among them, one of the purposes in step S2) is to determine the base position parameters in the G-stem region based on the structural framework type, and the steps for obtaining the mapping relationship of the base position parameters in the G-stem region based on the structural framework type are as follows:

[0022] S2-1) Collecting the relative position parameters of bases in the G-stem region of G4 structures from the PDB database according to the structural framework type;

[0023] S2-2) calculating the empirical distribution function of each location parameter corresponding to each structural frame type;

[0024] S2-3) The mean, median, or sampling value that follows the empirical distribution of the parameters at each position corresponding to each structural framework type is taken as the common parameter of the base structure within the G-tetrad and between adjacent G-tetrads.

[0025] Among them, one of the purposes of step S2) is to determine the structural parameter geometric system of G4, and the relative position parameters of the bases in the G-stem region include the parameters of the relative positions of the bases G in the G-tetrad and the positional relationship between the G-tetrads. Preferably, they specifically include 6 parameters for describing the relative positions of the bases G in the G-tetrad using the local coordinate system of the base G and 6 parameters for describing the relative positions between the G-tetrads using the local coordinate system of the G-tetrad: the parameters of the relative positions of the bases G in the G-tetrad include: displacement δx along the x-axis, displacement δy along the y-axis, displacement δz along the z-axis, rotation α around the x-axis, rotation β around the y-axis, and rotation γ around the z-axis; the parameters of the relative positions between the G-tetrads include: displacement Δx along the x-axis, displacement Δy along the y-axis, displacement Δz along the z-axis, rotation А around the x-axis, rotation В around the y-axis, and rotation Г around the z-axis. The 12 parameters defined in the present invention are the common parameters of the base structure within a G-tetrad and between adjacent G-tetrads.

[0026] The function of the local coordinate system of base G is to determine the position of base G so as to describe the positional relationship between adjacent bases G in the G-tetrad. The establishment of the local coordinate system of base G includes the following steps: setting the geometric center of the base or any atomic position in the base as the origin; setting the line from the origin to any atomic position in the base that does not coincide with the origin as the x-axis; selecting an atomic position in the base that is not on the x-axis, and the point and the x-axis together determine the base plane; setting the straight line on the base plane that is perpendicular to the x-axis as the y-axis; setting a line that passes through the origin and is perpendicular to the x-axis and y-axis at the same time, and satisfies the right-hand rule with the x-axis and y-axis. The straight line is the z-axis; the function of the G-tetrad local coordinate system is to determine the position of the G-tetrad so as to describe the positional relationship between adjacent G-tetrads, and its establishment includes the following steps: setting the geometric center of the G-tetrad as the origin; setting the plane passing through the origin and fitting the geometric centers of the four bases in the G-tetrad by the least squares method as the tetrad plane; setting the normal of the tetrad plane passing through the origin as the z-axis; setting any straight line perpendicular to the z-axis in the tetrad as the x-axis; setting the straight line perpendicular to both the z-axis and the x-axis in the tetrad and satisfying the right-hand rule with the z-axis and the x-axis as the y-axis.

[0027] Among them, one of the purposes in step S2) is to predict the base coordinates in the G-stem region and construct an all-atom structure model of the G-stem region. The predicted base coordinates in the G-stem region and the construction of the all-atom structure model of the G-stem region include the following steps: obtaining multiple base G structures from the PDB database, aligning them and calculating the average coordinates of each atom, relaxing the averaged coordinates using molecular simulation software, and obtaining the optimized base G structure as a standard structure; constructing a rectangular coordinate system as the local coordinate system of the first base G inside the G-tetrad with an arbitrary point in space as the origin, and calculating the origin coordinates and basis vectors of the local coordinate system of the second base G inside the G-tetrad based on the first local coordinate system through the relative position relationship of the base G within the G-tetrad , and so on, the local coordinate systems of the third and fourth base G in the G-tetrad are obtained in turn, the standard base G structure is embedded in each local coordinate system, and by relaxation, the G-tetrad full-atom model containing four bases G is obtained; still with any point in space as the origin, a rectangular coordinate system is constructed as the local coordinate system of the first G-tetrad in the G-stem region structure, and through the relative position relationship between the G-tetrads, based on the first G-tetrad local coordinate system, the origin coordinates and basis vectors of the second G-tetrad local coordinate system are calculated, and so on, the local coordinate systems of all subsequent G-tetrads are obtained, the calculated G-tetrad full-atom model is embedded in each G-tetrad local coordinate system, and relaxed to obtain the G-stem region full-atom model.

[0028] The G-loop structure modeling in step S3) includes the following steps:

[0029] S3-1) Download all single-stranded DNA and DNA G4 structural models from the PDB database, extract all fragment structures of single-base, double-base, and triple-base lengths, and construct a fragment structure library;

[0030] S3-2) If the target G-loop region is longer than two, the sequence is randomly split into two-base and three-base fragments, with one base overlap between adjacent fragments; otherwise, the G-loop region is treated as a single fragment;

[0031] S3-3) For each fragment split in S3-2), sample the fragment structure library generated in S3-1) according to its sequence, and splice the sampled fragments according to the overlapping parts;

[0032] S3-4) extracting distance constraints for the start and end atoms of the target G-loop region based on the G-stem region structure, performing energy minimization through molecular dynamics simulation, and optimizing the target G-loop region all-atom structure model under the constraints;

[0033] S3-5) aligning the start and end atomic coordinates to achieve the splicing of the target G-loop region structure and the G-stem region structure to form a structurally complete candidate G4 all-atom model;

[0034] S3-6) Repeat (S3-2)-(S3-5) multiple times to form a set of candidate models.

[0035] Among them, the steps of step S4) are as follows: structural relaxation is performed based on a dynamic simulation tool to minimize the energy of each alternative G4 all-atom model generated in step S3), optimize the atomic structure of the model, identify the structural group with the best performance through cluster analysis, and select the structure with the lowest energy and the smallest RMSD as the final model to complete the fine construction of the G4 all-atom model.

[0036] The present invention provides a method for predicting G4 structure based on sequence information, which specifically comprises the following steps:

[0037] (1) First, a decision tree algorithm was used to input sequence features such as the number of G-tetrads, the number of G-loops, the length of the G-loops, and the proportion of bases A, T, C, and G in the G-loops to achieve the four structural types of G4: parallel structure, antiparallel structure-positive reverse reverse reverse, antiparallel structure-positive reverse reverse reverse, and 3+1 mixed structure (see Figure 3 ) for classification prediction and define it as the G4 structure framework;

[0038] (2) Secondly, based on the mapping relationship between the structural framework type and the base position parameters in the G-stem region, the coordinates of each base G in the G-stem region are predicted and the full-atom structure model of the G-stem region is constructed (see Figure 2 b and Figure 4 Twelve structural parameters were defined to describe the relative positions of G bases within a G-tetrad and between G-tetrads. These parameters include six "intra-G-tetrad structural parameters," which use the local coordinate system of the G bases to describe the relative positions of the four G bases within a G-tetrad, and six "inter-G-tetrad structural parameters," which use the local coordinate system of the G bases to describe the relative positions between G-tetrads. By calculating, statistically analyzing, and analyzing known G4 structures in the PDB database, a set of common structural parameters was designed for each type of G4 structural framework. These parameters can be used to calculate the coordinates of all other G bases in the G-stem region from a single standard base G, thereby generating a predicted structure for the G-stem region.

[0039] (3) Then, the G-loop structure connecting each G-segment is modeled based on the sequence information to generate all possible G-loop structures (see Figure 2c) Download all single-stranded DNA and DNA G4 structural models from the PDB database, extract all fragment structures of single-base, double-base, and triple-base lengths, and construct a fragment structure library; develop a G-loop region structure sampling algorithm. If the target G-loop region length is greater than two, the sequence is randomly split into double-base and triple-base fragments, with one base overlap between adjacent fragments; otherwise, the G-loop region is treated as a single fragment; for each split fragment, sample it in the fragment structure library according to its sequence, and splice the sampled fragments according to the overlapping parts; extract the distance constraints of the start and end atoms of the target G-loop region based on the G-stem region structure, and then extract the distance constraints of the start and end atoms of the target G-loop region based on the G-stem region structure. Energy minimization is performed through molecular dynamics simulation to optimize the target G-ring region all-atom structure model under constrained conditions; the starting and ending atomic coordinates are aligned to realize the splicing of the target G-ring region structure and the G-stem region structure, and the target G-ring region is rotated with the line connecting the starting and ending atoms as the axis so that the dihedral angle formed by the target G-ring region center of mass, the line connecting the starting and ending atoms, and the G-stem region center of mass is as large as possible to reduce the atomic conflict between the target G-ring region and the G-stem region in the model, and form a structurally complete alternative G4 all-atom model; the above steps are repeated multiple times to form a set of alternative models.

[0040] (4) Finally, a structurally stable G4 all-atom model is constructed. Based on energy minimization, each candidate G4 all-atom model is optimized and screened to obtain an optimized stable G4 all-atom model. Structural relaxation is performed based on dynamic simulation tools to minimize the energy of each candidate G4 all-atom model generated in (3) and optimize the atomic structure of the model. Through cluster analysis, the structural group with the best performance is identified, and the structure with the lowest energy and smallest RMSD is selected as the final model, completing the fine construction of the G4 all-atom model.

[0041] The concept of the present invention is as follows: for a G4 sequence of unknown structure, sequence features are first extracted, and the G4 structural framework is predicted using the extracted sequence features. On this basis, the structures of the G-stem region and the G-loop region are predicted. Finally, the predicted structures of the G-stem region and the G-loop region are integrated. After structural optimization, an all-atom prediction model of G4 can be generated.

[0042] Beneficial effects: Compared with the existing technology, the present invention has the following advantages: Based on the known DNA G4 structures in the PDB database and the structure-related data obtained through experiments, the present invention explores and summarizes the correspondence between sequence features, structural frameworks, and structural parameters, and constructs a complete modeling method from sequence to structure, which provides an important foundation for better understanding the biological functions of these G4s, as well as for drug target discovery, nucleic acid probe design and other research based on G4 structures. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 Schematic diagram of the G4 structure;

[0044] Figure 2 Flowchart of the steps for G4 structure prediction;

[0045] Figure 3 It is the G4 structural framework and its classification diagram;

[0046] Figure 4 It is the geometric system diagram of G4 structural parameters;

[0047] Figure 5 This is the full-atom structure model diagram of the G-stem region of 2LOD;

[0048] Figure 6 This is a comparison diagram of the 2LOD all-atom structure model;

[0049] Figure 7 This is the full-atom structure model of the G-stem region of 5NYS;

[0050] Figure 8 This is a comparison chart of the 5NYS full-atom structure model. DETAILED DESCRIPTION

[0051] The present invention is described in further detail below.

[0052] Example 1 Structure prediction method based on G4 sequence information

[0053] 1. A structural framework classification prediction method based on DNA G4 sequence information, specifically comprising the following steps:

[0054] (1) Data collection: All known DNA G4 data were collected from the PDB database, including sequence information and atomic coordinate structural information for each G4. The length of these G4 sequences ranged from 12 to 84 nt, some of which were redundant, but all information was retained because the same G4 sequence can form different structures and topologies.

[0055] (2) Data preprocessing: The raw data were cleaned and missing or incomplete records were removed. Each G4 file was parsed to extract sequence features and structural framework types. Sequence features included the number of G-tetrads (num_G4tetrads), the number of G-loops (num_loops), the length of the G-loops (G_loop_length), the proportion of A in the G-loops (loop.A), the proportion of T in the G-loops (loop.T), the proportion of C in the G-loops (loop.C), and the proportion of G in the G-loops (loop.G). The structural framework was determined by the orientation of the four G-segments in the G4 structure, which were parallel, antiparallel (UDDU), antiparallel (UDUD), and 3+1 hybrid. Sequence features were used as sample classification attributes, and structural framework types were used as classification results to construct a training sample set.

[0056] (3) Establish a decision tree classifier model: The decision tree model is constructed based on the DecisionTreeClassifier class in the Scikit-learn library. The classifier model uses the sequence characteristics of DNA G4 as the classification attribute and outputs the classification result as the type of structural framework. The model performs a five-fold cross-validation on the training sample set obtained above, that is, the training sample set is randomly divided into five equal parts, four of which are taken as training sets for each training, and the remaining one is taken as the test set, and five validations are performed. For each type of structural framework, the performance of the decision tree model is evaluated separately. Let TP, TN, FP, and FN be the number of true positive samples, true negative samples, false positive samples, and false negative samples, respectively. The calculation of the four evaluation indicators of accuracy, precision, recall, and F1 score are shown in the following formula:

[0057]

[0058] The decision tree classifier model achieved an average accuracy of 0.9643, precision of 0.9612, recall of 0.9639, and F1 score of 0.9619.

[0059] 2. Based on the mapping relationship between the structural framework type and the base position parameters in the G-stem region, the base coordinates in the G-stem region are predicted and the full-atom structure model of the G-stem region is constructed:

[0060] (1) Defining the structural parameters of the G-stem region: To explore the details of the G4 structure in the genome, this study used a multi-degree-of-freedom method to describe these structures (see Figure 4To address the unique characteristics of the G4 structure, six "G-tetrad intra-base G structural parameters" were selected to describe the relative positions of the G bases within a G-tetrad, and six "inter-G-tetrad structural parameters" were selected to describe the relative positions of adjacent G-tetrads. The six "G-tetrad intra-base G structural parameters" focus on describing the spatial relationship between the four G bases within a G-tetrad. These parameters include six parameters that describe the relative positions of adjacent G bases within a G-tetrad using the local coordinate system of the bases G: displacement δx along the x-axis, displacement δy along the y-axis, displacement δz along the z-axis, rotation α around the x-axis, rotation β around the y-axis, and rotation γ around the z-axis; and six parameters that describe the relative positions of adjacent G-tetrads using the local coordinate system of the G-tetrad: displacement Δx along the x-axis, displacement Δy along the y-axis, displacement Δz along the z-axis, rotation A around the x-axis, rotation B around the y-axis, and rotation Γ around the z-axis.

[0061] (2) Define the internal structural parameters of the G-tetrad

[0062] a. Define the local coordinate system for base G. Defining the base coordinate system involves the following steps: set the geometric center of the base as the origin; set the line from the origin to any atomic position within the base that does not coincide with the origin as the x-axis; select an atomic position within the base that is not on the x-axis; this point and the x-axis together define the base plane; set the line perpendicular to the x-axis on the base plane as the y-axis; set the line passing through the origin, perpendicular to both the x-axis and the y-axis, and satisfying the right-hand rule with the x-axis and y-axis as the z-axis.

[0063] b. Based on the local coordinate system of the bases within the G-tetrad defined above, calculate the six structural parameters between the four bases G within the G-tetrad. Assume that there are two bases G within the G-tetrad, and the local coordinate system of each base G is defined by three orthogonal basis vectors, whose local coordinate systems are (X1, Y1, Z1) and (X2, Y2, Z2), respectively. The position vectors of the origin of the base local coordinate system are O1 and O2, respectively. The linear displacement parameters between two adjacent bases G are calculated using the following formula:

[0064] δx=(O2-O1)×X1

[0065] δy=(O2-O1)×Y1

[0066] δz=(O2-O1)×Z1

[0067] Among them, δx is the displacement of the local coordinate system of two adjacent bases G along the x-axis, δy is the displacement of the local coordinate system of two adjacent bases G along the y-axis, and δz is the displacement of the local coordinate system of two adjacent bases G along the z-axis.

[0068] The rotation matrix R is calculated from the basis vectors of the local coordinate system of the two bases by the following formula:

[0069] R=B2B1 -1

[0070] Where B1 and B2 are the basis vectors of the local coordinate systems of the two bases.

[0071] Extracting the Euler angle from the rotation matrix R can get the rotation parameters. The calculation process is as follows:

[0072]

[0073] Among them, R ij Represents the element in the i-th row and j-th column of the rotation matrix R, α is the rotation of the local coordinate system of two adjacent bases G around the x-axis, β is the rotation of the local coordinate system of two adjacent bases G around the y-axis, and γ is the rotation of the local coordinate system of two adjacent bases G around the z-axis.

[0074] (3) Define the structural parameters between G-tetrads

[0075] a. Define the local coordinate system of the G-tetrad. The specific steps are as follows: Set the geometric center of the G-tetrad as the origin; set the plane passing through the origin and fitting the geometric centers of the four bases in the G-tetrad using the least squares method as the tetrad plane; set the normal of the tetrad plane passing through the origin as the z-axis; set any line perpendicular to the z-axis within the tetrad as the x-axis; set the line perpendicular to both the z-axis and the x-axis within the tetrad as the y-axis.

[0076] b. Based on the local coordinate system of the G-tetrad defined above, calculate the six structural parameters between the G-tetrads. Assume there are two G-tetrads, their local coordinate systems are (X1, Y1, Z1) and (X2, Y2, Z2), and the centers of the G-tetrads are C1 and C2. The linear displacement parameters between the G-tetrads are calculated using the following formula:

[0077] Δx=(C2-C1)×X1

[0078] Δy=(C2-C1)×Y1

[0079] Δz=(C2-C1)×Z1

[0080] Wherein, Δx is the displacement of the local coordinate system of two adjacent G-tetrads along the x-axis, Δy is the displacement of the local coordinate system of two adjacent G-tetrads along the y-axis, and Δz is the displacement of the local coordinate system of two adjacent G-tetrads along the z-axis.

[0081] The rotation matrix R can be constructed by the dot product and cross product of two sets of basis vectors, as follows:

[0082]

[0083] Extracting the Euler angle from the rotation matrix R can get the rotation parameters. The calculation process is as follows:

[0084] А=arctan2(R 32 ,R 33 )

[0085] В=arcsin(-R 31 )

[0086] Г=arctan2(R 21 ,R 11 )

[0087] Among them, R ij represents the element in the i-th row and j-th column of the rotation matrix R, where A is the rotation of the local coordinate system of two adjacent G-tetrads around the x-axis, B is the rotation of the local coordinate system of two adjacent G-tetrads around the y-axis, and Г is the rotation of the local coordinate system of two adjacent G-tetrads around the z-axis.

[0088] (4) Determine the commonality of G4 structural parameters: 517 G4 structural data extracted from the PDB database were classified according to parallel structure (parallel), anti-parallel structure-positive reverse reverse (anti-parallel UDDU), anti-parallel structure-positive reverse reverse reverse (anti-parallel UDUD), and 3+1 hybrid structure (3+1 hybrid). 12 structural parameters of each G4 structure in each category were calculated and analyzed. In order to reveal the distribution pattern of structural parameters of each type of G4, the empirical distribution function of each position parameter corresponding to each structural framework type was calculated. The mean, median, or sample value that obeyed the empirical distribution of each position parameter corresponding to each structural framework type was taken to determine the commonality of base structural parameters within G-tetrads (see Table 1) and the commonality of base structural parameters between adjacent G-tetrads (see Table 2).

[0089] Table 1 Common parameters of the relative positions of base G within the G-tetrad

[0090]

[0091] Table 2 Common parameters of relative positions between G-tetrads

[0092]

[0093] (5) Predicting the G-stem region structure: Using the common structural parameters summarized above, these structural frameworks are converted into specific structural parameters, and the G-stem region structure is generated accordingly.

[0094] a. First, obtain multiple base G structures from the PDB database, align them and calculate the average coordinates of each atom. Specifically, let the spatial coordinates of an atom in the i-th structure be (x i ,y i , z i ), a total of N structure samples are collected, then the average coordinate of the atom The calculation is as follows:

[0095]

[0096] Subsequently, the averaged coordinates are relaxed using the molecular simulation software OPENMM to obtain the optimized base G structure as the standard structure. The molecular simulation software GROMACS / AMBER can also be used for this purpose.

[0097] b. With any point in space as the origin, construct a rectangular coordinate system as the local coordinate system of the first base G within the G-tetrad. Based on the relative positional relationships of the bases within the G-tetrad described by the common structural parameters calculated in (4), the origin coordinates and basis vectors of the local coordinate system of the second base within the G-tetrad are calculated based on the first local coordinate system. Similarly, the local coordinate systems of the third and fourth bases within the G-tetrad are obtained in sequence. The standard base G structure is embedded in each local coordinate system according to the definition of 3(a), and through relaxation, a full-atom model of the G-tetrad containing four bases G is obtained.

[0098] c. Still using any point in space as the origin, construct a rectangular coordinate system as the local coordinate system of the first G-tetrad in the G-stem region structure. Based on the relative positional relationships between the G-tetrads described by the common structural parameters counted in (4), calculate the origin coordinates and basis vectors of the local coordinate system of the second G-tetrad based on the local coordinate system of the first G-tetrad, and so on, to obtain the local coordinate systems of all subsequent G-tetrads. The calculated G-tetrad full-atom model is embedded in each G-tetrad local coordinate system according to the definition of 3(b), and relaxed to obtain the G-stem region full-atom model.

[0099] 3. Model the G-loop region connecting each G-segment based on sequence information to generate an alternative G4 all-atom model

[0100] (1) Construction of a DNA fragment structure database: First, all required single-stranded DNA and DNAG4 structural models were downloaded from the PDB database. The downloaded files were preprocessed to remove water molecules and impurities, retaining only the DNA strand information. This step ensured the accuracy and relevance of subsequent analysis.

[0101] Next, the processed PDB file is parsed to extract detailed information for each base, including sequence information and the types and three-dimensional coordinates of all atoms that make up the base. Based on this information, single-base, double-base, triple-base and other fragments are generated programmatically. For single-base fragments, they are obtained directly from the extracted base data; for double-base and triple-base fragments, two or three consecutive bases are combined into fragments by traversing the base sequence. Finally, the information of each fragment (including sequence, atomic coordinates, and source PDB ID) is stored in the database. A DNA fragment structure database is constructed, using MySQL as the database management software.

[0102] The DNA Fragment Structure Database consists of two main data tables: the fragment sequence table (Fragments) and the atomic coordinate table (Atoms). The fragment sequence table is the core of the database, used to systematically store and manage DNA fragment information extracted from PDB files. The atomic coordinate table meticulously records detailed information about each atom in the DNA fragment to support accurate structural analysis. This table is designed to ensure data integrity and ease of management.

[0103] (2) G-loop structural sampling and modeling

[0104] a. First, the target G-loop sequence is split into possible single-base, double-base, and triple-base fragments. The splitting mechanism is that if the target G-loop length is greater than two, the sequence is randomly split into double-base and triple-base fragments, with one base overlap between adjacent fragments. Otherwise, the G-loop is treated as a single fragment.

[0105] b. For each split fragment, sample the fragment structure library generated in (1) according to its sequence, and splice the sampled fragments according to the overlapping parts; extract the distance constraints of the start and end atoms of the target G-loop region according to the G-stem region structure, and optimize the target G-loop region full-atom structure model under the constraint conditions through energy minimization and molecular dynamics simulation;

[0106] c. Align the coordinates of the start and end atoms to achieve the splicing of the target G-loop region structure and the G-stem region structure. Rotate the target G-loop region with the line connecting the start and end atoms as the axis so that the dihedral angle formed by the target G-loop region center of mass, the line connecting the start and end atoms, and the G-stem region center of mass is as large as possible to reduce the atomic conflict between the target G-loop region and the G-stem region in the model and form a structurally complete alternative G4 all-atom model;

[0107] d. Repeat ac multiple times to form a set of alternative models.

[0108] 4. Construct a structurally stable G4 all-atom model. Based on energy minimization, optimize and screen each candidate G4 all-atom model to obtain an optimized and stable G4 all-atom model.

[0109] (1) Screening for structurally stable G4 all-atom models: Structural relaxation was performed using dynamic simulation tools to minimize the energy of each candidate G4 all-atom model generated in step 3 and optimize the atomic structure of the model. Cluster analysis was used to identify the structural group with the best performance, and the structure with the lowest energy and smallest RMSD was selected as the final model, completing the detailed construction of the G4 all-atom model.

[0110] (2) Structural evaluation: RMSD was used to evaluate the structure of the predicted G4 structure model and the experimentally measured G4 native conformation.

[0111] Example 2

[0112] The structure of G4 with PDB ID 2LOD was modeled, and its sequence is: GGGATGGGACACAGGGGACGGG. The specific implementation steps are as follows:

[0113] The number of G-tetrads of G4 is 3, the number of G-loop regions is 3, the length of the G-loop region is 10, the proportion of A in the loop region is 0.6, the proportion of T in the loop region is 0.1, the proportion of C in the loop region is 0.3, and the proportion of G in the loop region is 0.1. These sequence features are input into the structural framework prediction model, and the structural framework is predicted to be a 3+1 mixed structure. The predicted structure of the G-stem region is calculated using the common structural parameters of G4 obtained in Example 1, and analyzed using Pymol software, as shown in FIG. Figure 5 As shown in the figure, the green portion represents the topmost G-tetrad; the blue portion represents the middle G-tetrad; and the yellow portion represents the bottom G-tetrad. The arrows indicate the direction of each G-segment chain. Three G-segment chains are aligned, while one chain is oriented in the opposite direction, indicating a 3+1 hybrid structure.

[0114] On this basis, according to the G-loop region structure splitting mechanism, its G-loop region is split into single-base, double-base and three-base fragments. For example, if the AT length is equal to 2, no splitting is required, ACACA is split into AC+CA+AC+CA or AC+CA+ACA or AC+CAC+CA or ACA+AC+CA or ACA+ACA, and GAC is split into GA+AC or GAC. Sampling is performed in the fragment structure database constructed in Example 1, and the sampled fragments are spliced according to the overlapping parts. The distance constraints of the start and end atoms of the target G-loop region are extracted according to the G-stem region structure. By energy minimization and molecular dynamics simulation, the target G-loop region full-atom structure model under the constraint conditions is optimized; the start and end atomic coordinates are aligned to realize the splicing of the G-loop region structure and the G-stem region structure to form multiple structurally complete alternative G4 full-atom models; structural relaxation is performed based on the dynamics simulation tool to achieve energy minimization of each alternative G4 full-atom model and optimize the atomic structure of the model. Through cluster analysis, the structural group with the best performance was identified, and the structure with the lowest energy and smallest RMSD was selected as the final model to complete the detailed construction of the G4 all-atom structure model.

[0115] The structural model was compared with the natural conformation obtained by experimental methods and analyzed using Pymol software. The structural comparison diagram is shown in the figure below. Figure 6 As shown, the yellow part is the 2LOD native conformation obtained by the experimental method, and the blue part is the G4 full-atom structure model predicted by this method. The results show that the results obtained by this method have high accuracy.

[0116] Example 3

[0117] The structure of G4 with PDB ID 5NYS was modeled, and its sequence structure is: TAGGGACGGGCGGGCAGGGT. The specific implementation steps are as follows:

[0118] The number of G-tetrads of G4 is 3, the number of G-loop regions is 5, the length of the G-loop region is 8, the proportion of A in the loop region is 0.375, the proportion of T in the loop region is 0.25, the proportion of C in the loop region is 0.375, and the proportion of G in the loop region is 0. These sequence features are input into the structure framework prediction model, and the structure framework is predicted to be a parallel structure. The predicted structure of the G-stem region is calculated using the common structural parameters of G4 obtained in Example 1, and analyzed using Pymol software, as shown in FIG. Figure 7As shown in the figure, the green portion represents the topmost G-tetrad; the blue portion represents the middle G-tetrad; and the yellow portion represents the bottom G-tetrad. The arrows indicate the direction of each G-segment chain. The four G-segment chains all have the same direction, indicating a parallel structure.

[0119] On this basis, according to the G-ring region structure splitting mechanism, its G-ring region is split into single-base, double-base and three-base fragments. The G-ring region sequence length in 5NYS is less than or equal to 2 and does not need to be split. All fragments of the G-ring region are TA, AC, C, CA, T. Sampling is carried out in the fragment structure database constructed in Example 1, and the sampled fragments are spliced according to the overlapping parts. The distance constraints of the start and end atoms of the target G-ring region are extracted according to the G-stem region structure. By energy minimization and molecular dynamics simulation, the target G-ring region full-atom structure model under the constraint conditions is optimized; the start and end atomic coordinates are aligned to realize the splicing of the G-ring region structure and the G-stem region structure to form multiple alternative G4 full-atom models with complete structures; structural relaxation is performed based on dynamic simulation tools to achieve energy minimization of each alternative G4 full-atom model and optimize the atomic structure of the model. Through cluster analysis, the structural group with the best performance is identified, and the structure with the lowest energy and the smallest RMSD is selected as the final model to complete the fine construction of the G4 full-atom structure model.

[0120] The structural model was compared with the natural conformation obtained by experimental methods and analyzed using Pymol software. The structural comparison diagram is shown in the figure below. Figure 8 As shown, the yellow part is the native conformation of 5NYS obtained by experimental methods, and the blue part is the G4 all-atom model predicted by this method. Subsequently, the RMSD calculated is It also has higher accuracy.

Claims

1. A method for predicting DNA G-quadruplex structure based on sequence information, characterized in that: The following steps are involved: S1) Predicting the structural framework type based on the sequence information of DNA G4; S2) determining a mapping relationship of base position parameters within the G-stem region based on the structural framework type, predicting base coordinates within the G-stem region, and constructing a full-atom structural model of the G-stem region; S3) Modeling the structure of the G-loop region connecting each G-segment based on sequence information, and constructing multiple structurally complete candidate G4 all-atom models using the fragment assembly method; S4) By simulating the relaxation process, the energy of the alternative G4 all-atom model is minimized, and a structurally stable G4 all-atom model is obtained through energy score screening.

2. The method for predicting DNA G-quadruplex structure based on sequence information according to claim 1, wherein The predicted structural framework types in step S1) include parallel structure, antiparallel structure-positive, negative, and negative, antiparallel structure-positive, negative, and negative, or 3+1 mixed structure.

3. The method for constructing a DNA G-quadruplex structure prediction model based on sequence information according to claim 2, wherein: The steps for obtaining the predicted structure framework type in step S1) are as follows: S1-1) Collect all known DNA G4 structural data from the PDB database, including the DNA sequence of G4 and the atomic coordinates in the structure; S1-2) Eliminate missing or incomplete records, extract the sequence characteristics and structural framework types of G4, use the sequence characteristics as the classification attributes of the samples, and the structural framework types as the classification results to construct a training sample set; S1-3) Establishing a decision tree classifier model: The classifier model uses the sequence characteristics of DNA G4 as classification attributes, and outputs the classification result as the type of structural framework.

4. The method for predicting DNA G-quadruplex structure based on sequence information according to claim 3, characterized in that: The sequence characteristics include one or more of the number of G-quadruplexes, the number of G-loop regions, the length of the G-loop region, the proportion of A in the G-loop region, the proportion of T in the G-loop region, the proportion of C in the G-loop region, and the proportion of G in the G-loop region.

5. The method for predicting DNA G-quadruplex structure based on sequence information according to claim 1, wherein In step S2), the steps for obtaining the mapping relationship of the base position parameters in the G-stem region based on the structural framework type are as follows: S2-1) Collecting the relative position parameters of bases in the G-stem region of G4 structures from the PDB database according to the structural framework type; S2-2) calculating the empirical distribution function of each location parameter corresponding to each structural frame type; S2-3) The mean, median, or sampling value that follows the empirical distribution of the parameters at each position corresponding to each structural framework type is taken as the common parameter of the base structure within the G-tetrad and between adjacent G-tetrads.

6. The method for predicting DNA G-quadruplex structure based on sequence information according to claim 5, characterized in that: The relative position parameters of the bases in the G-stem region include parameters of the relative positions of the bases G in the G-tetrad and the positional relationship between the G-tetrads. Preferably, they specifically include 6 parameters for describing the relative positions of the bases G in the G-tetrad using the local coordinate system of the base G and 6 parameters for describing the relative positions between the G-tetrads using the local coordinate system of the G-tetrad: the parameters of the relative positions of the bases G in the G-tetrad include: displacement δx along the x-axis, displacement δy along the y-axis, displacement δz along the z-axis, rotation α around the x-axis, rotation β around the y-axis, and rotation γ around the z-axis; the parameters of the relative positions between the G-tetrads include: displacement Δx along the x-axis, displacement Δy along the y-axis, displacement Δz along the z-axis, rotation А around the x-axis, rotation В around the y-axis, and rotation Г around the z-axis.

7. The method for predicting DNA G-quadruplex structure based on sequence information according to claim 6, characterized in that: The function of the local coordinate system of base G is to determine the position of base G so as to describe the positional relationship between adjacent bases G in the G-tetrad. The establishment of the local coordinate system of base G includes the following steps: setting the geometric center of the base or any atomic position in the base as the origin; setting the line from the origin to any atomic position in the base that does not coincide with the origin as the x-axis; selecting an atomic position in the base that is not on the x-axis, and the point and the x-axis together determine the base plane; setting the straight line on the base plane that is perpendicular to the x-axis as the y-axis; setting a line that passes through the origin and is perpendicular to both the x-axis and the y-axis, and satisfies the right-hand rule with the x-axis and the y-axis. The straight line is the z-axis; the function of the G-tetrad local coordinate system is to determine the position of the G-tetrad so as to describe the positional relationship between adjacent G-tetrads, and its establishment includes the following steps: setting the geometric center of the G-tetrad as the origin; setting the plane passing through the origin and fitting the geometric centers of the four bases in the G-tetrad by the least squares method as the tetrad plane; setting the normal of the tetrad plane passing through the origin as the z-axis; setting any straight line perpendicular to the z-axis in the tetrad as the x-axis; setting the straight line perpendicular to both the z-axis and the x-axis in the tetrad and satisfying the right-hand rule with the z-axis and the x-axis as the y-axis.

8. The method for predicting DNA G-quadruplex structure based on sequence information according to claim 1, wherein In step S2), the base coordinates in the G-stem region are predicted, and the all-atom structure model of the G-stem region is constructed, which includes the following steps: obtaining multiple base G structures from the PDB database, aligning them and calculating the average coordinates of each atom, relaxing the averaged coordinates using molecular simulation software, and obtaining the optimized base G structure as the standard structure; taking any point in space as the origin, constructing a rectangular coordinate system as the local coordinate system of the first base G inside the G-tetrad, and calculating the origin coordinates and basis vectors of the local coordinate system of the second base G inside the G-tetrad based on the first local coordinate system through the relative position relationship of the base G inside the G-tetrad, and so on, obtaining the third, The local coordinate system of the fourth base G, embeds the standard base G structure into each local coordinate system, and through relaxation, obtains the G-tetrad full-atom model containing four bases G; still with any point in space as the origin, constructs a rectangular coordinate system as the local coordinate system of the first G-tetrad in the G-stem region structure, and through the relative position relationship between the G-tetrads, based on the first G-tetrad local coordinate system, calculates the origin coordinates and basis vectors of the second G-tetrad local coordinate system, and so on, obtains the local coordinate systems of all subsequent G-tetrads, embeds the calculated G-tetrad full-atom model into each G-tetrad local coordinate system, and relaxes to obtain the G-stem region full-atom model.

9. The method for predicting DNA G-quadruplex structure based on sequence information according to claim 1, wherein: The G-loop structure modeling in step S3) includes the following steps: S3-1) Download all single-stranded DNA and DNA G4 structural models from the PDB database, extract all fragment structures of single-base, double-base, and triple-base lengths, and construct a fragment structure library; S3-2) If the target G-loop region is longer than two, the sequence is randomly split into two-base and three-base fragments, with one base overlap between adjacent fragments; otherwise, the G-loop region is treated as a single fragment; S3-3) For each fragment split in S3-2), sample the fragment structure library generated in S3-1) according to its sequence, and splice the sampled fragments according to the overlapping parts; S3-4) extracting distance constraints for the start and end atoms of the target G-loop region based on the G-stem region structure, performing energy minimization through molecular dynamics simulation, and optimizing the target G-loop region all-atom structure model under the constraints; S3-5) aligning the start and end atomic coordinates to achieve the splicing of the target G-loop region structure and the G-stem region structure to form a structurally complete candidate G4 all-atom model; S3-6) Repeat (S3-2)-(S3-5) multiple times to form a set of candidate models.

10. The method for predicting DNA G-quadruplex structure based on sequence information according to claim 1, wherein: The steps of step S4) are as follows: structural relaxation is performed based on a dynamic simulation tool to minimize the energy of each alternative G4 all-atom model generated in step S3), optimize the atomic structure of the model, identify the structural group with the best performance through cluster analysis, and select the structure with the lowest energy and the smallest RMSD as the final model to complete the fine construction of the G4 all-atom model.

Citation Information

Patent Citations

  • Cell specific genome G-quadruplex prediction method

    CN113160877A

  • Molecular interaction sites of coronavirus RNA and methods of modulating the same

    WO2004110386A2