Method for analyzing higher-order structure of nucleic acid

The integration of NMR, PCA, and CD/UV analyses allows for precise identification and analysis of nucleic acid structures, addressing the limitations of existing methods in distinguishing G-quadruplexes and i-motifs, enhancing our understanding of their structural dynamics.

JP2026010685APending Publication Date: 2026-01-22NAT UNIV CORP TOKYO UNIV OF AGRI & TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025116170
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-09
Filing Date
2025-07-09
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing methods struggle to accurately distinguish and analyze the three-dimensional structure of nucleic acids, particularly non-canonical structures like G-quadruplexes and i-motifs, which are crucial for understanding biological functions and drug development.

Method used

A method combining nuclear magnetic resonance (NMR) spectroscopy with principal component analysis (PCA) and clustering analysis of circular dichroism (CD) spectra, along with CD and UV temperature change measurements, to accurately determine the topology and structure of nucleic acids.

Benefits of technology

Enables precise discrimination of non-canonical nucleic acid structures such as G-quadruplexes and i-motifs, providing a more accurate understanding of their conformational changes in response to environmental conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026010685000019
    Figure 2026010685000019
  • Figure 2026010685000020
    Figure 2026010685000020
  • Figure 2026010685000021
    Figure 2026010685000021
Patent Text Reader

Abstract

To provide a method for analyzing the higher-order structure of a nucleic acid.SOLUTION: Detecting a peak of imino proton within a range of chemical shift values of 10.0 to 16. 0ppm in a proton nuclear magnetic resonance spectrum; Determining the types of base pairs, optionally performing principal component and clustering analyses of the circular dichroism spectra, and optionally performing circular dichroism temperature shift measurements at single wavelengths in the range of 220 to 320nm; A step of performing curve fitting using a thermodynamic equilibrium model and performing evaluation, a step of performing ultraviolet temperature change measurement at a single wave length within a range of 290 to 300nm in some cases, and a step of determining that two or more types of nucleic acid higher-order structures including a guanine quadruplex structure are formed in a case where fitting to a state model indicating equilibrium of two or more states is shown, A step of determining that a guanine quadruplex structure is not formed when no fit to any state model indicating an equilibrium of two or more states is shown.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for analyzing the higher-order structure of a nucleic acid. [Background technology]

[0002] The well-known standard conformation of DNA is the right-handed double helix formed by Watson-Crick base pairs. However, it is known that DNA and other nucleic acids can also adopt non-standard conformations, such as left-handed Z-shaped structures, parallel structures, i-motifs, triplex structures, and G-quadruplex structures. In recent years, the formation of these non-standard structures has been thought to play an important role in biological processes.

[0003] The G-quadruplex (G4) structure is one of the higher-order structures formed by nucleic acids (DNA and RNA) with guanine-rich sequences. The G4 structure is formed when four guanine molecules are bonded to two adjacent guanine molecules via Hoogsteen base pairs to form a guanine planar structure (called a G-quartet), and the G-quartets are further stacked by π-π bonds. Cations (especially K + and Na +The coordination of the nucleoside at the center of the G-quartet neutralizes the repulsion between the negative charges of the nucleic acid, stabilizing the G4 structure. Various types of G4 structures are known, including monomolecular structures formed intramolecularly, bimolecular structures formed between multiple molecules, tetrad structures, and higher-order structures (higher-order quartets) formed from multiple G4 structures. A typical G4 structure consists of four guanine-containing oligonucleotide strands that form stacked G-quartets and three loops (usually 1–7 nucleotides long). G4 structures are also known to have different topologies. Depending on the orientation of the four oligonucleotide strands (the 3' and 5' ends), G4 topologies include parallel (all four oligonucleotide strands are in the same orientation), hybrid (three oligonucleotide strands are in the same orientation, one in an opposite orientation), and antiparallel (two oligonucleotide strands are in the same orientation, two in an opposite orientation). These G4 topologies undergo dynamic structural changes depending on the surrounding environmental conditions. On the other hand, the complementary strand of guanine-rich nucleic acids is cytosine-rich. In such complementary strands, cytosines are protonated to form base pairs and intercalate to form a unique quadruplex structure, the i-motif. Like G4 structures, the formation of the i-motif is dynamically controlled by environmental conditions, and the cleavage of the nucleic acid duplex due to i-motif formation is known to affect the formation of G4 structures in the complementary strand. G4 structures are present in the promoter regions of genes encoding growth-related factors such as insulin, c-MYC, and VEGF, and have been reported to be involved in the suppression or promotion of gene expression by regulating the binding of transcription factors. Accurate techniques for distinguishing higher-order structures such as G4 structures and i-motifs in nucleic acids are needed to control biological functions in drug development and other areas.

[0004] Techniques are known for detecting G4 structures and i-motifs in nucleic acids by using nuclear magnetic resonance (NMR) to detect imino protons derived from Hoogsteen base pairs that constitute G quartets and cytosine-cytosine (C:C+) base pairs (e.g., Non-Patent Documents 1 and 2). However, NMR cannot distinguish the topology of G4 structures, and it is difficult to evaluate conformational changes in nucleic acids.

[0005] A technique (CD PCA / HCA assay) is known that combines circular dichroism (CD) spectroscopy with principal component analysis and hierarchical clustering (PCA / HCA) to detect the orientation of nucleic acid bases and thereby determine the topology of nucleic acids (e.g., Non-Patent Document 3). However, this method can sometimes misjudge nucleic acids that do not actually form a G4 structure as forming a G4 structure, or misjudge an i-motif as antiparallel.

[0006] There is a need for a technology that can more accurately determine and analyze the three-dimensional structure of nucleic acids and structural changes that occur in response to the surrounding environment. [Prior art documents] [Non-patent literature]

[0007] [Non-Patent Document 1] Adrian M., et al., (2012) Methods, 57, 11-24. [Non-patent document 2] Dai J., et al., (2010) PLoS ONE 5(7):e11647. [Non-patent document 3] Del Villar-Guerra R., et al., (2018) Angew. Chem. Int. Ed. Engl., 57, 7171-7175. Summary of the Invention [Problem to be solved by the invention]

[0008] An objective of the present invention is to provide a method for analyzing the higher-order structure of nucleic acids, particularly nucleic acids that are suspected of forming or already form non-canonical structures such as G-quadruplex structures. [Means for solving the problem]

[0009] As a result of extensive research to solve the above problems, the present inventors have discovered that by combining nuclear magnetic resonance spectrum analysis, principal component analysis and clustering analysis of circular dichroism spectra, and optionally circular dichroism (CD) temperature change measurement data analysis and ultraviolet (UV) temperature change measurement data analysis, it is possible to more accurately distinguish non-canonical nucleic acid structures such as G-quadruplex structures in addition to the canonical nucleic acid double-stranded structures, and have completed the present invention.

[0010] That is, the present invention includes the following.

[0011] [1] A method for analyzing the higher-order structure of a test nucleic acid in a sample solution, comprising: a) detecting imino proton peaks within a chemical shift value range of 10.0 to 16.0 ppm in the proton nuclear magnetic resonance spectrum of the test nucleic acid, and determining the type of base pair detected in the test nucleic acid based on the distribution of chemical shift values ​​of the detected peaks; b) when it is determined in step a) that a guanine-guanine Hoogsteen base pair has been detected, determining whether the topology of the G-quadruplex structure belongs to a parallel type, a hybrid type, or an antiparallel type by performing principal component analysis and clustering analysis on the circular dichroism spectrum of the test nucleic acid; c) if the topology of the G-quadruplex structure in step b) belongs to the parallel type, the hybrid type, or the antiparallel type, it is determined that the test nucleic acid forms a G-quadruplex structure having that topology; and if the topology of the G-quadruplex structure in step b) does not belong to any of the parallel type, the hybrid type, and the antiparallel type, it is determined that the test nucleic acid forms a G-quadruplex structure having that topology; and performing circular dichroism temperature shift measurement on the test nucleic acid at a single wavelength in the range of 220 to 320 nm; d) performing curve fitting using a thermodynamic equilibrium model on the circular dichroism temperature change measurement data obtained in step c) and evaluating the fit to a state model showing equilibrium of three or more states; e) if step d) indicates a fit to a state model showing equilibrium of three or more states, measuring the ultraviolet temperature change of the test nucleic acid at a single wavelength in the range of 290 to 300 nm; f) a step of performing curve fitting using a thermodynamic equilibrium model on the ultraviolet temperature change measurement data obtained in step e) and evaluating the fit to a state model showing equilibrium of two or more states; g) determining that the test nucleic acid forms two or more nucleic acid higher-order structures including a G-quadruplex structure if step f) indicates a match to a state model showing equilibrium in two or more states, and determining that the test nucleic acid does not form a G-quadruplex structure if step f) does not indicate a match to any state model showing equilibrium in two or more states; A method comprising:

[0012] [2] The method according to [1] above, wherein the type of base pair detected in step a) is a Hoogsteen base pair between guanine and guanine, a cytosine-cytosine base pair, or a Watson-Crick base pair.

[0013] [3] The method according to [1] or [2] above, wherein, when it is determined that a cytosine-cytosine base pair is detected in step a), the i-motif is identified by further performing principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid.

[0014] [4] The method according to any one of [1] to [3] above, further comprising measuring the proton nuclear magnetic resonance spectrum and / or the circular dichroism spectrum of the test nucleic acid.

[0015] [5] The method according to any one of [1] to [4] above, wherein the circular dichroism spectrum is measured at wavelengths in the range of 220 to 320 nm.

[0016] [6] The method according to any one of [1] to [5] above, wherein the principal component analysis and clustering analysis in step b) are performed using principal component models and clusters created from reference data of circular dichroism spectra of G-quadruplex-forming nucleic acids whose G-quadruplex topology is known.

[0017] [7] The method according to any one of the above [1] to [6], wherein the circular dichroism temperature change measurement is carried out at a single wavelength within the range of 260 to 265 nm.

[0018] [8] The method according to any one of the above [1] to [7], wherein the ultraviolet temperature change measurement is carried out at a single wavelength of 295 nm.

[0019] [9] The method according to any one of the above [1] to [8], wherein the sample solution contains a cation.

[0020]

[10] The method according to any one of [1] to [9] above, wherein the sample solution contains two or more types of the test nucleic acids.

[0021]

[11] The method according to

[10] above, comprising calculating an estimated value of the mixing ratio of two or more types of the test nucleic acids in the sample solution based on the results of principal component analysis and clustering analysis of the circular dichroism spectra.

[0022]

[12] The method according to any one of [1] to

[11] above, wherein the sample solution further contains a substance that interacts or has the potential to interact with the test nucleic acid.

[0023]

[13] A data receiving unit that receives proton nuclear magnetic resonance spectrum data, circular dichroism spectrum data, circular dichroism temperature change measurement data, and ultraviolet temperature change measurement data of the test nucleic acid in the sample solution, respectively; a nuclear magnetic resonance spectrum analysis processor that detects imino proton peaks within a chemical shift value range of 10.0 to 16.0 ppm in the proton nuclear magnetic resonance spectrum of the test nucleic acid, and determines the type of base pair detected in the test nucleic acid based on the distribution of chemical shift values ​​of the detected peaks; a G-quadruplex topology determination processor that, when it is determined that a Hoogsteen base pair between guanine and guanine has been detected in the nuclear magnetic resonance spectrum analysis processor, determines whether the topology of the G-quadruplex belongs to a parallel type, a hybrid type, or an antiparallel type by performing principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid; and a circular dichroism temperature change measurement data evaluation processing unit that performs curve fitting using a thermodynamic equilibrium model on circular dichroism temperature change measurement data of the test nucleic acid at a single wavelength in the range of 220 to 320 nm, and evaluates the conformity to a state model that shows equilibrium in three or more states; an ultraviolet temperature change measurement data evaluation processing unit that performs curve fitting using a thermodynamic equilibrium model on ultraviolet temperature change measurement data of the test nucleic acid at a single wavelength within a range of 290 to 300 nm, and evaluates the conformity to a state model that shows equilibrium in two or more states; a higher-order structure determination processor that determines the higher-order structure of the test nucleic acid based on result data generated by the nuclear magnetic resonance spectrum analysis processor, the G-quadruplex structure topology discrimination processor, the circular dichroism temperature change measurement data evaluation processor, and the ultraviolet temperature change measurement data evaluation processor; and A nucleic acid higher-order structure analysis device comprising:

[0024]

[14] The apparatus described in

[13] above, wherein the type of base pair detected by the nuclear magnetic resonance spectrum analysis processing unit is a Hoogsteen base pair between guanine and guanine, a cytosine-cytosine base pair, or a Watson-Crick base pair.

[0025]

[15] The apparatus described in

[13] or

[14] above, further comprising an i-motif discrimination processing unit that, when it is determined that a cytosine-cytosine base pair has been detected in the nuclear magnetic resonance spectrum analysis processing unit, discriminates the i-motif by further performing principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid.

[0026]

[16] The apparatus according to any one of

[13] to

[15] above, wherein the G-quadruplex topology discrimination processor performs principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid using principal component models and clusters created from reference data of circular dichroism spectra of G-quadruplex-forming nucleic acids whose G-quadruplex topology is known.

[0027]

[17] A circular dichroism temperature change measurement unit that performs circular dichroism temperature change measurement on the test nucleic acid at a single wavelength in the range of 220 to 320 nm; an ultraviolet temperature change measurement unit that performs ultraviolet temperature change measurement on the test nucleic acid at a single wavelength in the range of 290 to 300 nm; The device according to any one of

[13] to

[16] above, further comprising:

[0028]

[18] The apparatus according to any one of

[13] to

[17] above, wherein the G-quadruplex topology discrimination processor further calculates an estimated value of the mixing ratio of two or more types of the test nucleic acids in the sample solution based on the results of principal component analysis and clustering analysis of the circular dichroism spectra.

[0029]

[19] A program for analyzing the higher-order structure of a test nucleic acid in a sample solution, comprising: a) detecting imino proton peaks within a chemical shift value range of 10.0 to 16.0 ppm in the proton nuclear magnetic resonance spectrum of the test nucleic acid, and determining the type of base pair detected in the test nucleic acid based on the distribution of chemical shift values ​​of the detected peaks; b) when it is determined in step a) that a guanine-guanine Hoogsteen base pair has been detected, determining whether the topology of the G-quadruplex structure belongs to a parallel type, a hybrid type, or an antiparallel type by performing principal component analysis and clustering analysis on the circular dichroism spectrum of the test nucleic acid; c) determining that the test nucleic acid forms a G-quadruplex structure having a topology that is parallel, hybrid, or antiparallel in step b), and evaluating the conformity of the test nucleic acid to a state model showing equilibrium with three or more states by performing curve fitting using a thermodynamic equilibrium model on data obtained by measuring the temperature change of the test nucleic acid by circular dichroism at a single wavelength within a range of 220 to 320 nm in step b) if the topology of the G-quadruplex structure does not belong to any of parallel, hybrid, or antiparallel. d) if step c) indicates a fit to a state model showing equilibrium of three or more states, performing curve fitting using a thermodynamic equilibrium model on the ultraviolet temperature change measurement data of the test nucleic acid at a single wavelength in the range of 290 to 300 nm, and evaluating the fit to a state model showing equilibrium of two or more states; e) determining that the test nucleic acid forms two or more nucleic acid higher-order structures including a G-quadruplex structure if step d) indicates a match to a state model showing equilibrium in two or more states, and determining that the test nucleic acid does not form a G-quadruplex structure if step d) does not indicate a match to any state model showing equilibrium in two or more states; A program that causes a computer to execute a process including the steps of:

[0030]

[20] The program according to

[19] above, wherein the type of base pair detected in step a) is a Hoogsteen base pair between guanine and guanine, a cytosine-cytosine base pair, or a Watson-Crick base pair.

[0031]

[21] The program described in

[19] or

[20] above, wherein, when it is determined that a cytosine-cytosine base pair is detected in step a), the processing further comprises identifying the i-motif by performing principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid.

[0032]

[22] The program according to any one of

[19] to

[21] above, wherein the principal component analysis and clustering analysis in step b) are performed using principal component models and clusters created from reference data of circular dichroism spectra of G-quadruplex-forming nucleic acids whose G-quadruplex topology is known.

[0033]

[23] The program according to any one of

[19] to

[22] above, wherein the processing further comprises calculating an estimated value of the mixing ratio of two or more types of the test nucleic acids in the sample solution based on the results of principal component analysis and clustering analysis of the circular dichroism spectra. [Effects of the Invention]

[0034] According to the present invention, the higher-order structure of nucleic acids, in particular non-canonical structures such as G-quadruplex structures, can be determined more accurately. [Brief explanation of the drawings]

[0035] [Figure 1] FIG. 1 shows a typical example of a decision tree for determining the higher-order structure of a nucleic acid according to the method of the present invention. [Figure 2] Figure 2 shows a PCA score plot created based on PCA / HCA analysis of the reference data, showing antiparallel, parallel, and hybrid clusters (with 95% confidence intervals). [Figure 3] Figure 3 shows CD spectra of IGA3 and the c-MYC promoter sequence in water or various concentrations of PBS. A: IGA3, B: c-MYC promoter sequence (c-MYC), C: T-IGA3. [Figure 4] Figure 4 shows the PCA score plot based on the reference data, plotting the principal component scores of IGA3 and the c-MYC promoter sequence (c-MYC) in water or various concentrations of PBS. The PCA score plot based on the reference data shows antiparallel, parallel, and hybrid clusters (with 95% confidence intervals). Individual PCA scores from the reference data are omitted in Figure 4. [Figure 5] FIG. 5 shows the results of UV temperature shift measurements at 295 nm of IGA3 and the c-MYC promoter sequence (c-MYC). [Figure 6] Figure 6 shows CD spectra of IGA3 in PBS containing different concentrations of cations (Na+ and / or K+): A: 0.01x PBS + 0-150 mM NaCl, B: 0.01x PBS + 0-4 mM KCl, C: 0.01x PBS + 0-160 mM NaCl + 0-4.5 mM KCl, D: 1x PBS. [Figure 7] Figure 7 shows the H-NMR spectra of the c-MYC promoter sequence (c-MYC), IGA3, T-IGA3, and dGMP in water or 0.01x PBS. A: Solvent H2O, B: Solvent DO. [Figure 8] Figure 8 shows the CD temperature change measurements and fitting results for the c-MYC promoter sequence at 262 nm. A: 0.01xPBS (2-state model, R2 = 0.9977), B: 0.05xPBS (2-state model, R2 = 0.9988), C: 0.2xPBS (2-state model, R2 = 0.9988), D: 0.5xPBS (2-state model, R2 = 0.9987), E: 1xPBS (2-state model, R2 = 0.9987), F: Water (2-state model, R2 = 0.9799). Fit: Fitting results, Raw: CD temperature change measurements. [Figure 9] Figure 9 shows the CD temperature change measurements and fitting results for IGA3 at 262 nm. A: 0.01xPBS (3-state model, R2 = 0.9921), B: 0.05xPBS (3-state model, R2 = 0.9933), C: 0.2xPBS (3-state model, R2 = 0.9926), D: 0.5xPBS (3-state model, R2 = 0.9948), E: 1xPBS (4-state model, R2 = 0.9967), F: Water (3-state model, R2 = 0.9807). Fit: Fitting results, Raw: CD temperature change measurements. [Figure 10]FIG. 10 is a photograph showing electrophoretic images of IGA3, T-IGA3, and the c-MYC promoter sequence in PBS at various concentrations. [Figure 11] Figure 11 shows the UV temperature change measurements and fitting results for IGA3 (A) and c-MYC (B). No model fit for IGA3 (A), whereas a 2-state model fit for c-MYC (B). [Figure 12] Fig. 12 shows PCA score plots obtained by CD PCA / HCA analysis using samples containing oligonucleotides with parallel G4 structures and oligonucleotides with hybrid G4 structures mixed at different ratios. In Fig. 12, the individual PCA scores derived from the reference data are not shown. [Figure 13] FIG. 13 shows the 1H-NMR spectrum of the ParaG4 and HybrG4 oligonucleotide mixture. [Figure 14] Figure 14 shows the plot of principal component scores of the oligonucleotide mixture on a PCA score plot based on the reference data, where individual PCA scores from the reference data are omitted. [Figure 15] FIG. 15 shows the CD temperature change measurements and fitting results for an oligonucleotide mixture at 262 nm. [Figure 16] FIG. 16 shows UV temperature change measurements and fitting results for oligonucleotide mixtures. DETAILED DESCRIPTION OF THE INVENTION

[0036] The present invention will be described in detail below.

[0037] The present invention relates to a technique for analyzing the higher-order structure of nucleic acids, preferably nucleic acids that are predicted to form or already form non-canonical structures such as G-quadruplex structures.

[0038] In the present invention, the nucleic acid (test nucleic acid) whose higher-order structure is analyzed may be DNA or RNA, or a hybrid thereof, and may be a single-stranded polynucleotide or a double-stranded polynucleotide, but is not limited thereto. The test nucleic acid may have any base sequence. In one embodiment, the test nucleic acid may be a nucleic acid containing a guanine-rich sequence that is likely to form a G-quadruplex structure. A guanine-rich sequence is a base sequence that is rich in guanine, and in the present invention, may be, for example, a base sequence containing four or more consecutive guanine sequences. In one embodiment, the guanine-rich sequence is, for example, a 5'-G ≧2 N 1-7 G ≧2 N 1-7 G ≧2 N 1-7 G ≧2 -3' (wherein N: A, C, G, T, or U). Alternatively, in another embodiment, the test nucleic acid may be a nucleic acid containing a cytosine-rich sequence that is likely to form an i motif. A cytosine-rich sequence is a base sequence that is rich in cytosine, and in the present invention, it may be, for example, a base sequence that contains multiple (e.g., four or more) sequences of two or more consecutive cytosines. In one embodiment, the test nucleic acid is a nucleic acid that contains or consists of an aptamer or promoter sequence.

[0039] In the present invention, nuclear magnetic resonance (NMR) spectroscopy, typically proton NMR spectroscopy, is performed on a test nucleic acid in a sample solution. In the present invention, based on the results of the proton NMR spectroscopy, proton NMR spectroscopy is combined with principal component analysis and clustering analysis of circular dichroism spectra. In a more preferred embodiment of the present invention, proton NMR spectroscopy, principal component analysis and clustering analysis of circular dichroism spectra, and further analysis of the circular dichroism and ultraviolet color temperature shift measurement data are combined for the test nucleic acid in the sample solution. This combination of analyses allows for more accurate discrimination of non-standard nucleic acid structures, such as G-quadruplex structures and i-motifs, in addition to standard double-stranded structures.

[0040] In one embodiment, the method of the present invention is a method for analyzing the higher-order structure of a test nucleic acid (e.g., a nucleic acid containing a guanine-rich sequence and / or a cytosine-rich sequence) in a sample solution, comprising: a) detecting imino proton peaks, typically within a chemical shift range of 10.0 to 16.0 ppm, in the proton nuclear magnetic resonance spectrum of the test nucleic acid, and determining the type of base pair detected in the test nucleic acid based on the distribution of chemical shift values ​​of the detected peaks; b) when it is determined that a guanine-guanine Hoogsteen base pair or a cytosine-cytosine base pair (C:C+ base pair) is detected in step a), a step of performing principal component analysis and clustering analysis of the circular dichroism spectrum of the nucleic acid The method may be a method for analyzing the higher-order structure of a nucleic acid, comprising:

[0041] In one embodiment, the method of the present invention is a method for analyzing the higher-order structure of a test nucleic acid in a sample solution, comprising the steps of: a) detecting imino proton peaks within a chemical shift value range of 10.0 to 16.0 ppm in the proton nuclear magnetic resonance spectrum of the test nucleic acid, and determining the type of base pair detected in the test nucleic acid based on the distribution of chemical shift values ​​of the detected peaks; b) when it is determined in step a) that a guanine-guanine Hoogsteen base pair has been detected, determining whether the topology of the G-quadruplex structure belongs to a parallel type, a hybrid type, or an antiparallel type by performing principal component analysis and clustering analysis on the circular dichroism spectrum of the test nucleic acid; c) if the topology of the G-quadruplex structure in step b) belongs to the parallel type, the hybrid type, or the antiparallel type, it is determined that the test nucleic acid forms a G-quadruplex structure having that topology; and if the topology of the G-quadruplex structure in step b) does not belong to any of the parallel type, the hybrid type, and the antiparallel type, it is determined that the test nucleic acid forms a G-quadruplex structure having that topology; and performing circular dichroism temperature shift measurement on the test nucleic acid at a single wavelength in the range of 220 to 320 nm; d) performing curve fitting using a thermodynamic equilibrium model on the circular dichroism temperature change measurement data obtained in step c) and evaluating the fit to a state model showing equilibrium of three or more states; e) if step d) indicates a fit to a state model showing equilibrium of three or more states, measuring the ultraviolet temperature change of the test nucleic acid at a single wavelength in the range of 290 to 300 nm; f) a step of performing curve fitting using a thermodynamic equilibrium model on the ultraviolet temperature change measurement data obtained in step e) and evaluating the fit to a state model showing equilibrium of two or more states; g) determining that the test nucleic acid forms two or more nucleic acid higher-order structures including a G-quadruplex structure if step f) indicates a match to a state model showing equilibrium in two or more states, and determining that the test nucleic acid does not form a G-quadruplex structure if step f) does not indicate a match to any state model showing equilibrium in two or more states; The method may include:

[0042] In the method of the present invention, NMR spectrum data of the nucleic acid to be analyzed (test nucleic acid) may be obtained from a public or private database, and the data may be used for analysis including peak detection. Alternatively, the method of the present invention may include a step of measuring the spectrum of the test nucleic acid by proton NMR (nuclear magnetic resonance spectroscopy). The proton NMR spectrum may be measured by a conventional method. In a preferred embodiment, the proton NMR spectrum is 1 The proton NMR spectrum may be a H-NMR spectrum. Examples of the measurement conditions for proton NMR are as follows: Observation frequency: 500.1 MHz Observation kernel: 1 H nucleus Solvent: H2O : D2O = 95 : 5 Temperature: 25℃ Solvent elimination: ES (Excitation sculpting) method (4.70 ppm)

[0043] In proton NMR spectrum analysis of nucleic acids, imino proton peaks (also referred to as imino proton signals) within a chemical shift range of 10.0 to 16.0 ppm in the proton nuclear magnetic resonance spectrum are detected. Based on the distribution of chemical shift values ​​of the detected peaks, the type of base pair detected in the nucleic acid can be determined. Examples of types of base pairs detected in nucleic acids include Hoogsteen base pairs between guanine and guanine (GG), cytosine-cytosine base pairs, and Watson-Crick base pairs.

[0044] When peaks are detected only within a chemical shift range of 10.0 to 12.0 ppm in the NMR spectrum of a nucleic acid, it can be determined that a Hoogsteen base pair between guanine and guanine in the nucleic acid has been detected. In the present invention, the term "Hoogsteen base pair" refers to a base pair formed with the participation of atoms present on the Hoogsteen edges of bases constituting the nucleic acid. Peaks detected only within a chemical shift range of 10.0 to 12.0 ppm in the NMR spectrum of a nucleic acid indicate the presence of a Hoogsteen base pair that may form a G4 structure. Therefore, in the present invention, it can be determined based on the peak that a Hoogsteen base pair between guanine and guanine in the nucleic acid has been detected (or that it has been formed).

[0045] When peaks are detected only within the chemical shift value range of 12.0 to 14.0 ppm in the proton NMR spectrum of a nucleic acid, it can be determined that Watson-Crick base pairs of the nucleic acid (formation of the base pairs) have been detected.

[0046] When no peak is detected within the chemical shift range of 10.0 to 16.0 ppm in the proton NMR spectrum of a nucleic acid, it can be determined that there is no imino proton signal, and in this case, it is determined that the nucleic acid does not form a specific higher-order structure.

[0047] When a peak is detected within a chemical shift value range of 15.0 to 16.0 ppm in the proton NMR spectrum of a nucleic acid, it can be determined that a cytosine-cytosine base pair (C:C+ base pair) of the nucleic acid (formation of the cytosine-cytosine base pair) has been detected.

[0048] When peaks other than those mentioned above are detected within the chemical shift value range of 10.0 to 16.0 ppm in the proton NMR spectrum of a nucleic acid, it can be determined that multiple types of base pairs (formation of) are detected in the nucleic acid.

[0049] In the present invention, a peak may be determined to be "detected" when the signal-to-noise ratio (s / n ratio) of the imino proton peak (signal) in the proton NMR spectrum of a nucleic acid is 10 or more. However, the reference s / n ratio can be changed as appropriate depending on the NMR measurement conditions.

[0050] In the present invention, when it is determined that Watson-Crick base pairing (formation thereof) has been detected as described above based on the distribution of chemical shift values ​​of imino proton peaks in the nuclear magnetic resonance spectrum, the test nucleic acid is determined to be double-stranded or to have a hairpin structure.

[0051] In the present invention, when it is determined that multiple types of base pairs (formation of base pairs) are detected as described above based on the distribution of chemical shift values ​​of imino proton peaks in the nuclear magnetic resonance spectrum, the test nucleic acid can be determined to be a mixture containing two or more structures selected from the group consisting of a G4 structure, an i-motif, a double strand, and a hairpin structure.

[0052] In the present invention, when it is determined that a guanine-guanine Hoogsteen base pair (formation thereof) has been detected as described above based on the distribution of chemical shift values ​​of imino proton peaks in the nuclear magnetic resonance spectrum, principal component analysis (PCA) and clustering analysis are performed on the circular dichroism spectrum of the test nucleic acid.

[0053] In the present invention, when it is determined that a cytosine-cytosine base pair (formation thereof) has been detected as described above based on the distribution of chemical shift values ​​of the imino proton peak in the nuclear magnetic resonance spectrum, principal component analysis (PCA) and clustering analysis are performed on the circular dichroism spectrum of the test nucleic acid.

[0054] In the present invention, circular dichroism spectral data of the test nucleic acid may be obtained from a public or private database, etc., and used in principal component analysis (PCA) and clustering analysis. Alternatively, the present invention may include a step of measuring the circular dichroism spectrum of the test nucleic acid by circular dichroism (CD) spectroscopy. Circular dichroism spectrum measurement can be performed by a conventional method. In a preferred embodiment, the CD spectrum of the test nucleic acid can be measured at a wavelength range of 220 to 320 nm, including wavelengths in the range of 200 to 340 nm. Examples of conditions for measuring the CD spectrum are as follows: Observation range: 200 to 340 nm Cell path length: 1 mm Temperature: 25℃ Response: 4 seconds Data acquisition interval: 0.5 nm Scanning speed: 50 nm / min

[0055] In one embodiment, principal component analysis (PCA) and clustering analysis of the circular dichroism spectra of nucleic acids may be performed using principal component models and clusters created from reference data (e.g., reference library data) of circular dichroism spectra of G-quadruplex-forming nucleic acids with known G-quadruplex topology. The method of the present invention may comprise the steps of performing principal component analysis (PCA) and clustering analysis of reference data (e.g., reference library data) of circular dichroism spectra of G-quadruplex-forming nucleic acids with known G-quadruplex topology, and creating principal component models and clusters (clusters of G-quadruplex topology). Reference data (e.g., reference library data) of circular dichroism spectra of G-quadruplex-forming nucleic acids with known G-quadruplex topology is not particularly limited, and examples include at least a portion or all of the nucleic acid data shown in Table 1 below (e.g., a data set including at least Nos. 1 to 22 in Table 1). This data can be obtained from public or private databases, or newly measured circular dichroism spectra can be used. Alternatively, principal component analysis (PCA) and clustering analysis of nucleic acid circular dichroism spectra can be performed using principal component models and clusters created from reference data (e.g., reference library data) of circular dichroism spectra of i-motif-forming nucleic acids known to form an i-motif. Reference data (e.g., reference library data) of circular dichroism spectra of i-motif-forming nucleic acids known to form an i-motif is also not particularly limited, and can be obtained from public or private databases, or newly measured circular dichroism spectra can be used.

[0056] When a test nucleic acid is determined to have a guanine-guanine Hoogsteen base pair (formation thereof), the topology of the G-quadruplex structure of the test nucleic acid can be determined by principal component analysis and clustering analysis of the nucleic acid's circular dichroism spectrum using principal component models and clusters created from reference data (e.g., data from a reference library) of circular dichroism spectra of G-quadruplex-forming nucleic acids with known G-quadruplex structure topologies. In a preferred embodiment, the G-quadruplex structure (G4 structure) topology determined is parallel, hybrid, or antiparallel. In one embodiment, a nucleic acid determined to have a guanine-guanine Hoogsteen base pair is classified (determined) into a parallel, hybrid, or antiparallel structure, or other structure, by principal component analysis and clustering analysis of the circular dichroism spectrum. Specifically, for example, for a nucleic acid determined to have a guanine-guanine Hoogsteen base pair, if the principal component scores obtained by applying the principal component model of the reference data are clustered into a parallel, hybrid, or antiparallel cluster created from the reference data within a 95% confidence interval, the topology of the G-quadruplex structure of the nucleic acid is determined to be parallel, hybrid, or antiparallel (the nucleic acid has a parallel, hybrid, or antiparallel G4 structure), respectively. On the other hand, if the nucleic acid is not clustered into any of the parallel, hybrid, or antiparallel clusters created from the reference data, the topology of the G-quadruplex structure of the nucleic acid is determined to be neither parallel, hybrid, nor antiparallel.

[0057] Furthermore, when a cytosine-cytosine base pair (C:C+ base pair) is determined to be detected in a test nucleic acid, the i-motif of the test nucleic acid can be identified by principal component analysis and clustering analysis of the nucleic acid's circular dichroism spectrum using a principal component model and clusters created from reference data (e.g., reference library data) of the circular dichroism spectra of i-motif-forming nucleic acids known to form i-motifs. For example, for a nucleic acid determined to have a cytosine-cytosine base pair, if the principal component score obtained by applying the principal component model of the reference data is clustered into a cluster within the 95% confidence interval of the i-motif created from the reference data, the nucleic acid is determined to be an i-motif type (forming an i-motif). On the other hand, if the nucleic acid is not clustered into a cluster within the 95% confidence interval of the i-motif created from the reference data, the nucleic acid is determined to be a mixture of i-motif nucleic acids and nucleic acids forming other structures, such as a G4 structure, a duplex, and a hairpin structure (Figure 1). Alternatively, when a cytosine-cytosine base pair (C:C+ base pair) is detected in a test nucleic acid, the nucleic acid may be determined to be of the i-motif type (i-motif formation) based on the circular dichroism spectrum of the nucleic acid. In the circular dichroism spectrum, the i-motif of a nucleic acid typically exhibits a positive peak at 290-295 nm and a negative peak at 260-270 nm.

[0058] In the present invention, when the principal component analysis and clustering analysis of the circular dichroism spectrum of a nucleic acid in which (the formation of) a guanine-guanine Hoogsteen base pair is determined to be detected reveals that the topology of the G-quadruplex structure does not belong to any of the parallel, hybrid, and antiparallel types, it is preferable to perform circular dichroism temperature shift measurements (CD melting measurements) on the test nucleic acid. Circular dichroism temperature shift measurements can be performed by conventional methods. Circular dichroism temperature shift measurements may be performed at a single wavelength within the range of 220 to 320 nm. In circular dichroism temperature shift measurements, it is preferable to measure the circular dichroism spectrum of the nucleic acid at a detectable single wavelength, for example, a single wavelength corresponding to a peak characteristic of each topology of the G-quadruplex structure (parallel, hybrid, or antiparallel), and calculate the CD value. For example, in the CD spectrum, a parallel G-quadruplex structure exhibits a positive peak at 260 to 265 nm and a negative peak at 240 to 245 nm, an antiparallel G-quadruplex structure exhibits a positive peak at 290 to 295 nm and a negative peak at 260 to 270 nm, and a hybrid G-quadruplex structure exhibits positive peaks near 295 nm and 260 nm and a negative peak near 240 nm. In one embodiment, when investigating the temperature change of a parallel G-quadruplex structure, circular dichroism temperature change measurement may be performed at a single wavelength within the range of 260 to 265 nm, for example, a single wavelength of 262 nm. The circular dichroism temperature change measurement is performed while changing the measurement temperature over time. The measurement temperature is not particularly limited and may be, for example, from 4°C to 95°C.

[0059] Examples of measurement conditions for the circular dichroism temperature change measurement are as follows. Measurement wavelength: 262 nm Cell path length: 1 mm Evaluation temperature: 4℃~95℃ Response: 1 second (sec) Data capture interval: 0.1°C or 0.2°C Temperature gradient: 1°C / min or 2°C / min

[0060] Furthermore, it is preferable to perform curve fitting using a thermodynamic equilibrium model on the data obtained by measuring the temperature change of circular dichroism of the test nucleic acid.

[0061] In the present invention, it is also preferable to perform ultraviolet temperature change measurements (UV melting measurements) on test nucleic acids in which it has been determined that a guanine-guanine Hoogsteen base pair (formation) has been detected and the topology of the G-quadruplex structure has been determined not to belong to any of the parallel, hybrid, or antiparallel types. The ultraviolet temperature change measurements can be performed by a conventional method. In the ultraviolet temperature change measurements, absorbance at a single wavelength in the range of 290 to 300 nm (preferably 295 nm) can be measured. UV absorption at 295 nm is characteristic of the G-quadruplex structure. The ultraviolet temperature change measurements are performed while changing the measurement temperature over time. The measurement temperature is not particularly limited, but may be, for example, from 4°C to 95°C. As described below, ultraviolet temperature change measurements are preferably performed when curve fitting using a thermodynamic equilibrium model to the circular dichroism temperature change measurement data of the test nucleic acid indicates a fit to a state model showing equilibrium in three or more states. Although circular dichroism temperature shift measurements may detect nucleic acid structures other than G-quadruplex structures, for example, around 260 nm, combining them with ultraviolet temperature shift measurements allows for more accurate determination of the higher-order structure of nucleic acids.

[0062] An example of the measurement conditions for measuring the ultraviolet temperature change is as follows. Measurement wavelength: 295 nm Cell path length: 1 mm Evaluation temperature: 4℃~95℃ Response: 1 second (sec) Data capture interval: 0.1°C or 0.2°C Temperature gradient: 1°C / min or 2°C / min

[0063] Furthermore, it is preferable to perform curve fitting using a thermodynamic equilibrium model on the ultraviolet temperature change measurement data of the test nucleic acid.

[0064] Curve fitting using a thermodynamic equilibrium model for nucleic acid circular dichroism temperature change measurement data and ultraviolet temperature change measurement data can be performed by standard methods. Curve fitting using a thermodynamic equilibrium model can evaluate the fit to a state model that shows equilibrium between reactants and products, or between reactants, one or more intermediates, and products. For example, the number of states (number of components) of nucleic acid in a sample solution can be identified based on the curve fitting. Examples of nucleic acid states (components) in a sample solution include, but are not limited to, multiple types of nucleic acids that react with each other, intermediates and products formed by the reaction, and the same nucleic acid forming different higher-order structures (intermediates and products) depending on the temperature. In one embodiment, for example, the polymerization state of nucleic acid molecules can be evaluated / determined based on the curve fitting.

[0065] In order to curve fit the circular dichroism temperature change measurement data and the ultraviolet temperature change measurement data using a thermodynamic equilibrium model, the method of the present invention may set an equation for at least one equilibrium constant according to the thermodynamic equilibrium model.

[0066] Curve fitting using a thermodynamic equilibrium model for the circular dichroism temperature change measurement data and the ultraviolet temperature change measurement data can be performed by a conventional method, but for example, a 2-state model (2-state model; [1] below), a 3-state model (3-state model; [2] below), or a 4-state model (4-state model; [3] below) may also be used as the equilibrium model. Search parameters may be Tm and ΔH ([4] and [5] below).

[0067]

number

[0068]

number

[0069]

number

[0070]

number

[0071]

number

[0072] In the above formula, θ obs is the measured CD value (circular dichroism temperature change measurement data) or absorbance (ultraviolet temperature change measurement data).

[0073]

number

[0074]

number

[0075] R is the gas constant, exp is the exponential function, ΔH is the enthalpy, and ΔS is the entropy.

[0076] In the above formulas [4] and [5], T is the measurement temperature, and Tm,i is the denaturation midpoint of each phase transition.

[0077] In one embodiment, the Levenberg-Marquardt algorithm may be used for fitting.

[0078] Curve fitting can be performed on circular dichroism temperature change measurement data at a single wavelength at which circular dichroism spectra can be detected, preferably a single wavelength within the range of 220 to 320 nm (e.g., 260 to 265 nm), for example, 262 nm, to evaluate the fit to a state model showing equilibrium in three or more states. For example, if the results converge with a 3-state model or a 4-state model, the fit to a state model showing equilibrium in three or more states is indicated, indicating the presence of multiple G4 structures or multiple higher-order structures. Alternatively, if the curve fitting is performed on circular dichroism temperature change measurement data and the results of evaluating the fit to a state model showing equilibrium in three or more states converge with a 2-state model, or if the fit to any state model showing equilibrium in two or more states is not indicated, the structure is classified as "other." Note that if the fit to multiple state models is converged, the state model with the higher coefficient of determination is determined to be the one that fits.

[0079] In a preferred embodiment, if the curve fitting of the circular dichroism temperature change measurement data indicates a fit to a state model showing equilibrium in three or more states, it is preferable to further perform curve fitting of the ultraviolet temperature change measurement data of the test nucleic acid to evaluate the fit to a state model showing equilibrium in two or more states. The ultraviolet temperature change measurement data of the test nucleic acid may be acquired at a single wavelength within the range of 290 to 300 nm, for example, 295 nm. For example, if the curve fitting of the ultraviolet temperature change measurement data converges using either a 2-state model, a 3-state model, or a 4-state model, the test nucleic acid can be determined to correspond to a "nucleic acid mixture containing a G4 structure." A "nucleic acid mixture containing a G4 structure" refers to the presence of multiple G4 structures or a mixture of G4 structures and other higher-order structures in the test nucleic acid. Alternatively, if none of the 2-state model, 3-state model, or 4-state model converges, the test nucleic acid can be determined to have a "non-G4 structure," i.e., no G4 structure has been formed. Note that if convergence is achieved using multiple state models, the test nucleic acid is determined to conform to the state model with the higher coefficient of determination.

[0080] The coefficient of determination for curve fitting can be calculated using the following formula:

[0081]

number

[0082] An example of a decision tree for determining the higher-order structure of a nucleic acid for carrying out the above-described method of the present invention is shown in Figure 1. In one embodiment, the method of the present invention comprises determining the higher-order structure of a nucleic acid using a decision tree such as the decision tree shown in Figure 1.

[0083] In the present invention, it is preferable that test nucleic acids in a solution (sample solution) are the subject of measurement and data analysis. It is known that the higher-order structure of nucleic acids (e.g., G4 structure, i-motif, double strand, hairpin, etc.) changes depending on the ambient environmental conditions. Therefore, when analyzing the higher-order structure of nucleic acids, it is preferable that nucleic acids in a sample solution having the desired ambient environmental conditions are the subject of measurement and data analysis.

[0084] In one embodiment, the sample solution containing the test nucleic acid preferably contains cations, more preferably monovalent cations. The cations include sodium ions (Na + ), potassium ions (K + ), and / or lithium ion (Li + Examples of cations for stabilizing the G4 structure include, but are not limited to, sodium ions (Na + ), potassium ions (K + ) is particularly preferred.

[0085] In one embodiment, the sample solution containing nucleic acids may contain water or a buffer solution. Examples of buffer solutions include, but are not limited to, phosphate-buffered saline, potassium phosphate buffer, citrate buffer, Tris-HCl buffer, HEPES buffer, and borate buffer. In one embodiment, the sample solution containing nucleic acids may contain phosphate-buffered saline in a concentration range from 0.01x PBS to 1x PBS. The composition of 1x PBS (phosphate-buffered saline) is as follows: 10 mM NaHPO, 1.76 mM KHPO, 137 mM NaCl, 2.68 mM KCl, pH 7.4.

[0086] In one embodiment, the sample solution may contain two or more nucleic acids (test nucleic acids; for example, nucleic acids that are expected to form or form non-canonical structures such as G-quadruplex structures or i-motifs, or nucleic acids that are expected to form duplexes or hairpins or form non-canonical structures). In one embodiment, the sample solution may contain two or more nucleic acids that are expected to form or form G-quadruplex structures. For example, when a sample solution containing two or more nucleic acids is used for measurement, in principal component analysis (PCA) and clustering analysis of the nucleic acid circular dichroism spectra, the equilibrium state of the two or more nucleic acids may shift between clusters of topologies while maintaining the linearity of the mixture ratio (concentration ratio). In this case, the mixture ratio (concentration ratio) of the two or more nucleic acids can also be estimated by principal component regression using the nucleic acid circular dichroism spectral data. More specifically, for example, the circular dichroism spectrum of a nucleic acid-containing sample (nucleic acid sample) with known topology and mixture ratio (concentration ratio) is measured, and PC1 and PC2 scores, which are principal component scores for the sample, are calculated. Next, a coefficient vector (

[10] below) is obtained by multiple regression ([8] and [9] below) using the topology mixture ratio (concentration ratio) as the objective variable matrix ([6] below) and the principal component scores (PC1, PC2) as the explanatory variable matrix ([7] below). An estimate of the mixture ratio (concentration ratio) of the unknown sample (

[12] below) can be obtained by multiplying the vector (

[11] below) consisting of PC1 and PC2 obtained by performing principal component analysis on the CD spectrum of the unknown sample by the coefficient vector β. That is, the method of the present invention may also include calculating an estimate of the mixture ratio of two or more test nucleic acids in a sample solution, as described above, based on the results of principal component analysis and clustering analysis of the circular dichroism spectra obtained above for a sample solution containing two or more test nucleic acids.In one embodiment, the calculation of the estimated value of the mixing ratio of two or more test nucleic acids in the sample solution may be performed when it is determined in step g) above that the test nucleic acid forms two or more nucleic acid higher-order structures including a G-quadruplex structure, or it may be performed in other cases where the sample solution can be determined to be a mixture of two or more test nucleic acids, for example, when it is determined that a cytosine-cytosine base pair has been detected or multiple types of base pairs have been detected (for example, but not limited to, when the test nucleic acid is a mixture containing two or more structures selected from the group consisting of a G4 structure, an i-motif, a double strand, and a hairpin structure).

[0087]

number

[0088] In one embodiment, the sample solution may further contain a substance that interacts with or may interact with nucleic acids. For example, when analyzing the higher-order structure of a nucleic acid that is an aptamer, the sample solution may contain the aptamer and its binding partner molecule. For example, the sample solution may contain the aptamer IGA3 and its binding partner insulin.

[0089] The present invention also provides a program for causing a computer to execute the method of the present invention as described above. The program of the present invention may be a program for determining the higher-order structure of nucleic acid.

[0090] In one embodiment, the present invention provides a program for analyzing the higher-order structure of a test nucleic acid in a sample solution, comprising: a) detecting imino proton peaks within a chemical shift value range of 10.0 to 16.0 ppm in the proton nuclear magnetic resonance spectrum of the test nucleic acid, and determining the type of base pair detected in the test nucleic acid based on the distribution of chemical shift values ​​of the detected peaks; b) when it is determined in step a) that a guanine-guanine Hoogsteen base pair has been detected, determining whether the topology of the G-quadruplex structure belongs to a parallel type, a hybrid type, or an antiparallel type by performing principal component analysis and clustering analysis on the circular dichroism spectrum of the test nucleic acid; c) determining that the test nucleic acid forms a G-quadruplex structure having a topology that is parallel, hybrid, or antiparallel in step b), and evaluating the conformity of the test nucleic acid to a state model showing equilibrium with three or more states by performing curve fitting using a thermodynamic equilibrium model on data obtained by measuring the temperature change of the test nucleic acid by circular dichroism at a single wavelength within a range of 220 to 320 nm in step b) if the topology of the G-quadruplex structure does not belong to any of parallel, hybrid, or antiparallel. d) if step c) indicates a fit to a state model showing equilibrium of three or more states, performing curve fitting using a thermodynamic equilibrium model on the ultraviolet temperature change measurement data of the test nucleic acid at a single wavelength in the range of 290 to 300 nm, and evaluating the fit to a state model showing equilibrium of two or more states; e) determining that the test nucleic acid forms two or more nucleic acid higher-order structures including a G-quadruplex structure if step d) indicates a match to a state model showing equilibrium in two or more states, and determining that the test nucleic acid does not form a G-quadruplex structure if step d) does not indicate a match to any state model showing equilibrium in two or more states; The present invention provides a program that causes a computer to execute a process including the steps of:

[0091] In one embodiment, in the program of the present invention, the type of base pair detected in step a) is a Hoogsteen base pair between guanine and guanine, a cytosine-cytosine base pair, or a Watson-Crick base pair.

[0092] In one embodiment, the processing of the program of the present invention includes, when it is determined in step a) that a cytosine-cytosine base pair has been detected, further performing principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid to identify the i-motif.

[0093] In one embodiment, the principal component analysis and clustering analysis in step b) of the processing of the program of the present invention are performed using principal component models and clusters created from reference data of circular dichroism spectra of G-quadruplex structure-forming nucleic acids whose G-quadruplex structure topologies are known.

[0094] In one embodiment, when it is determined that a cytosine-cytosine base pair is detected in step a) of the processing of the program of the present invention, principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid are performed using principal component models and clusters created from reference data of circular dichroism spectra of i-motif-forming nucleic acids in which the topology of the i-motif is known.

[0095] In one embodiment, when it is determined in step e) of the processing of the program of the present invention that the test nucleic acid forms two or more types of nucleic acid higher-order structures including a G-quadruplex structure, or when it is determined in step a) of the processing of the program that a cytosine-cytosine base pair or multiple types of base pairs have been detected and the sample solution is determined to be a mixture of two or more types of test nucleic acids, the processing of the program of the present invention may include calculating an estimated value of the mixing ratio of the two or more types of test nucleic acids in the sample solution based on the results of principal component analysis and clustering analysis of the circular dichroism spectra.

[0096] The processing in the program of the present invention is similar to the method of the present invention described above.

[0097] In one embodiment, the program of the present invention may be a program for performing information processing including determining the higher-order structure of a nucleic acid using a decision tree for determining the structure of a nucleic acid, such as the decision tree shown in FIG.

[0098] In one embodiment, the program of the present invention may be operable on a nucleic acid higher-order structure analysis device described below or any information processing device. In one embodiment, the program of the present invention may be operable on the web (online) or offline. The program of the present invention may be operable on a web server.

[0099] The present invention also provides a nucleic acid higher-order structure analysis device for carrying out / executing the above-mentioned method and program of the present invention.

[0100] The nucleic acid higher-order structure analysis device of the present invention includes, for example, a data receiving unit and a processing unit (arithmetic processing unit). The processing unit in the nucleic acid higher-order structure analysis device of the present invention includes a proton nuclear magnetic resonance spectrum analysis processing unit, a G-quadruplex structure topology discrimination processing unit, a circular dichroism temperature change measurement data evaluation processing unit, an ultraviolet temperature change measurement data evaluation processing unit, and a higher-order structure determination processing unit. The nucleic acid higher-order structure analysis device of the present invention may further include a display unit and a display command unit.

[0101] The nucleic acid higher-order structure analysis apparatus of the present invention preferably comprises a data receiving unit that receives proton nuclear magnetic resonance spectrum data of a test nucleic acid in a sample solution, a data receiving unit that receives circular dichroism spectrum data of the nucleic acid, a data receiving unit that receives circular dichroism temperature change measurement data of the nucleic acid, and a data receiving unit that receives ultraviolet temperature change measurement data of the nucleic acid. The nucleic acid higher-order structure analysis apparatus of the present invention may further comprise a data receiving unit that receives result data generated by the nuclear magnetic resonance spectrum analysis processing unit, the G-quadruplex structure topology discrimination processing unit, the circular dichroism temperature change measurement data evaluation processing unit, and / or the ultraviolet temperature change measurement data evaluation processing unit.

[0102] The data receiving unit in the nucleic acid higher-order structure analysis apparatus of the present invention can receive nuclear magnetic resonance spectral data, circular dichroism spectral data, circular dichroism temperature change measurement data, and ultraviolet temperature change measurement data of the nucleic acid. The data receiving unit in the nucleic acid higher-order structure analysis apparatus of the present invention may be connected to measurement devices for nuclear magnetic resonance measurement, circular dichroism spectroscopy measurement, circular dichroism temperature change measurement, and ultraviolet temperature change measurement, respectively, installed inside or outside the apparatus, and configured to receive the nuclear magnetic resonance spectral data, circular dichroism spectral data, circular dichroism temperature change measurement data, and ultraviolet temperature change measurement data. The data receiving unit in the nucleic acid higher-order structure analysis apparatus of the present invention may be connected to a data storage unit installed inside or outside the apparatus, and configured to receive the nuclear magnetic resonance spectral data, circular dichroism spectral data, circular dichroism temperature change spectral data, and ultraviolet temperature change measurement data. The data receiving unit in the nucleic acid higher-order structure analysis device of the present invention may also be connected to a data storage unit inside or outside the device and configured to receive result data generated by the nuclear magnetic resonance spectrum analysis processing unit, the G-quadruplex structure topology discrimination processing unit, the circular dichroism temperature change measurement data evaluation processing unit, and / or the ultraviolet temperature change measurement data evaluation processing unit. The nucleic acid higher-order structure analysis device of the present invention may have a data input unit. In the nucleic acid higher-order structure analysis device of the present invention, the data receiving unit may be connected to the data input unit and receive data input from the data input unit. The data received by the data receiving unit in the nucleic acid higher-order structure analysis device of the present invention may be stored in a data storage unit (such as a storage medium) installed inside the device.

[0103] The proton nuclear magnetic resonance spectrum data received by the data receiving unit is read from the data receiving unit or the data storage unit by the nuclear magnetic resonance spectrum analysis processing unit. The proton nuclear magnetic resonance spectrum data received by the data receiving unit in the nucleic acid higher-order structure analysis device of the present invention is subjected to arithmetic processing by the nuclear magnetic resonance spectrum analysis processing unit provided in the nucleic acid higher-order structure analysis device of the present invention.

[0104] The nuclear magnetic resonance spectrum analysis processor performs arithmetic processing on the read proton nuclear magnetic resonance spectrum data and detects imino proton peaks within a chemical shift value range of 10.0 to 16.0 ppm in the proton nuclear magnetic resonance spectrum. The nuclear magnetic resonance spectrum analysis processor detects the distribution of chemical shift values ​​of imino proton peaks in the proton nuclear magnetic resonance spectrum. In a preferred embodiment, the nuclear magnetic resonance spectrum analysis processor further calculates the signal-to-noise ratio of the imino proton peaks in the proton nuclear magnetic resonance spectrum. The nuclear magnetic resonance spectrum analysis processor stores the obtained data in a data storage unit as needed. The nuclear magnetic resonance spectrum analysis processor may determine the presence of imino proton peaks based on the signal-to-noise ratio. The nuclear magnetic resonance spectrum analysis processor determines the type of base pair detected in the nucleic acid based on the distribution of chemical shift values ​​of the detected peaks and the predetermined criteria described above. The types of base pairs detected by the determination performed by the nuclear magnetic resonance spectrum analysis processing unit include Hoogsteen base pairs between guanine and guanine, cytosine-cytosine base pairs, and Watson-Crick base pairs.

[0105] The circular dichroism spectral data received by the data receiving unit is read from the data receiving unit or the data storage unit by the G-quadruplex topology discrimination processor. The circular dichroism spectral data received by the data receiving unit in the nucleic acid higher-order structure analysis device of the present invention is arithmetically processed by the G-quadruplex topology discrimination processor included in the nucleic acid higher-order structure analysis device of the present invention.

[0106] The G-quadruplex topology discrimination processor performs arithmetic processing on the read circular dichroism spectrum data and performs principal component analysis and clustering analysis of the circular dichroism spectrum. The G-quadruplex topology discrimination processor preferably performs principal component analysis and clustering analysis of the test nucleic acid in the sample solution using principal component models and clusters created from reference data (e.g., reference library data) of circular dichroism spectra of G-quadruplex-forming nucleic acids with known G-quadruplex topology. The G-quadruplex topology discrimination processor may also create principal component models and clusters (learning models) from reference data (e.g., reference library data) of circular dichroism spectra of G-quadruplex-forming nucleic acids with known G-quadruplex topology. Reference data (e.g., reference library data) of circular dichroism spectra of nucleic acids known to have G-quadruplex topology or i-motifs, as well as principal component models and clusters created therefrom, may be stored in a data storage unit inside or outside the nucleic acid higher-order structure analysis device of the present invention. The G-quadruplex topology discrimination processor may use the principal component models and clusters read from the data storage unit to perform principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid. In the principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid, the G-quadruplex topology discrimination processor may calculate principal component scores of the circular dichroism spectrum of the nucleic acid, perform clustering, and generate a discrimination result of the G-quadruplex topology of the test nucleic acid using principal component models and clusters created from reference data (e.g., reference library data) of the circular dichroism spectrum of a G-quadruplex-forming nucleic acid whose G-quadruplex topology is known. The reference data and the data of the principal component models and clusters created therefrom may be input to or referenced by the G-quadruplex topology discrimination processor from a data storage unit inside or outside the nucleic acid higher-order structure analysis device of the present invention.The G-quadruplex topology discrimination processor can discriminate the topology of G4 structures in the test nucleic acid, whether parallel, hybrid, or antiparallel. The G-quadruplex topology discrimination processor stores data of principal component analysis and clustering analysis of circular dichroism spectra in a data storage unit as needed.

[0107] Furthermore, the G-quadruplex topology discrimination processor can calculate an estimated value of the mixture ratio of two or more test nucleic acids in the sample solution based on the results of principal component analysis and clustering analysis of the circular dichroism spectra. The circular dichroism spectral data, or the results of principal component analysis and clustering analysis of the circular dichroism spectra obtained from the circular dichroism spectral data, can be read from the data storage unit by the G-quadruplex topology discrimination processor. The G-quadruplex topology discrimination processor can calculate an estimated value of the mixture ratio of two or more test nucleic acids in the sample solution by arithmetic processing of the circular dichroism spectral data, or the results of principal component analysis and clustering analysis of the circular dichroism spectra obtained from the circular dichroism spectral data. Alternatively, the nucleic acid higher-order structure analysis device of the present invention may be provided with a separate circular dichroism spectral data processor. In this case, instead of the G-quadruplex topology discrimination processor, the circular dichroism spectral data processor can calculate an estimated value of the mixture ratio of two or more test nucleic acids in the sample solution based on the results of principal component analysis and clustering analysis of the circular dichroism spectra. The circular dichroism spectral data, or the results of principal component analysis and clustering analysis of the circular dichroism spectra obtained from the circular dichroism spectral data, are read by the circular dichroism spectral data processor from the data storage unit or as output data from the G-quadruplex topology discrimination processor. The circular dichroism spectral data processor can arithmetically process the circular dichroism spectral data, or the results of principal component analysis and clustering analysis of the circular dichroism spectra obtained from the circular dichroism spectral data, to calculate an estimated value of the mixture ratio of two or more test nucleic acids in the sample solution. In the present invention, the computational process for calculating an estimated value of the mixing ratio of two or more types of test nucleic acids in a sample solution may include, for example, referring to reference data of other nucleic acids and PCA score plot data created from circular dichroism spectral data of nucleic acid samples having multiple known mixing ratios (e.g., mixing ratios of 0:10 to 10:0), and performing multiple regression analysis based on the PCA score plot data.

[0108] The nucleic acid higher-order structure analysis device of the present invention may include an i-motif discrimination processor. When the nuclear magnetic resonance spectrum analysis processor determines that a cytosine-cytosine base pair has been detected, the i-motif discrimination processor further performs principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid to discriminate the i-motif. The i-motif discrimination processor performs arithmetic processing on the read circular dichroism spectrum data and performs principal component analysis and clustering analysis of the circular dichroism spectrum. The i-motif discrimination processor preferably performs principal component analysis and clustering analysis of the test nucleic acid in the sample solution using principal component models and clusters created from reference data (e.g., reference library data) of the circular dichroism spectrum of i-motif-forming nucleic acids known to form an i-motif. The i-motif discrimination processor may also create principal component models and clusters (learning models) from reference data (e.g., reference library data) of the circular dichroism spectrum of i-motif-forming nucleic acids known to form an i-motif. The i-motif discrimination processor may store reference data (e.g., reference library data) of circular dichroism spectra of i-motif-forming nucleic acids known to form an i-motif, as well as principal component models and clusters created therefrom, in a data storage unit inside or outside the nucleic acid higher-order structure analysis device of the present invention. The i-motif discrimination processor may perform principal component analysis and clustering analysis of the circular dichroism spectra of the test nucleic acids using the principal component models and clusters read from such a data storage unit. In the principal component analysis and clustering analysis of the circular dichroism spectra of the test nucleic acids, the i-motif discrimination processor may calculate principal component scores of the circular dichroism spectra of the nucleic acids, perform clustering, and generate discrimination results for the i-motifs of the test nucleic acids, using principal component models and clusters created from reference data (e.g., reference library data) of circular dichroism spectra of i-motif-forming nucleic acids known to form an i-motif.The reference data and the data of the principal component models and clusters created therefrom may be input to or referenced by the i-motif discrimination processor from a data storage unit inside or outside the nucleic acid higher-order structure analysis device of the present invention. The i-motif discrimination processor stores the data of the principal component analysis and clustering analysis of the circular dichroism spectra in a data storage unit as needed.

[0109] The circular dichroism temperature change measurement data received by the data receiving unit is read from the data receiving unit or the data storage unit by a circular dichroism temperature change measurement data evaluation processing unit. The circular dichroism temperature change measurement data received by the data receiving unit in the nucleic acid higher-order structure analysis device of the present invention is arithmetically processed by a circular dichroism temperature change measurement data evaluation processing unit included in the nucleic acid higher-order structure analysis device of the present invention.

[0110] The circular dichroism temperature change measurement data evaluation processor performs arithmetic processing on the read circular dichroism temperature change measurement data and performs curve fitting using a thermodynamic equilibrium model for the circular dichroism temperature change measurement data of the nucleic acid. The circular dichroism temperature change measurement data evaluation processor may further perform processing to set the thermodynamic equilibrium model. The thermodynamic equilibrium model may be recorded in a data storage unit inside or outside the nucleic acid higher-order structure analysis device of the present invention. The circular dichroism temperature change measurement data evaluation processor performs curve fitting using a thermodynamic equilibrium model, for example, a thermodynamic equilibrium model read from the data storage unit. The circular dichroism temperature change measurement data evaluation processor detects a thermodynamic equilibrium model that fits in the curve fitting, and preferably evaluates the fit to a state model that shows equilibrium in three or more states. The circular dichroism temperature change measurement data evaluation processor stores the fitted circular dichroism temperature change measurement data and / or the fitting results in a data storage unit as necessary.

[0111] The ultraviolet temperature change measurement data evaluation processing unit performs arithmetic processing on the read ultraviolet temperature change measurement data and performs curve fitting using a thermodynamic equilibrium model for the ultraviolet temperature change measurement data of the nucleic acid. The ultraviolet temperature change measurement data evaluation processing unit may further perform processing to set the thermodynamic equilibrium model. The thermodynamic equilibrium model may be recorded in a data storage unit inside or outside the nucleic acid higher-order structure analysis device of the present invention. The ultraviolet temperature change measurement data evaluation processing unit performs curve fitting using a thermodynamic equilibrium model, for example, a thermodynamic equilibrium model read from the data storage unit. The ultraviolet temperature change measurement data evaluation processing unit detects a thermodynamic equilibrium model that fits in the curve fitting, and preferably evaluates the fit to a state model that shows equilibrium of two or more states. The ultraviolet temperature change measurement data evaluation processing unit stores the fitted ultraviolet temperature change measurement data and / or the fitting results in a data storage unit as necessary.

[0112] The nucleic acid higher-order structure analysis device of the present invention may further include a higher-order structure determination processor. The higher-order structure determination processor determines the higher-order structure of the test nucleic acid based on result data generated by the nuclear magnetic resonance spectrum analysis processor, the G-quadruplex structure topology discrimination processor, the i-motif discrimination processor, the circular dichroism temperature change measurement data evaluation processor, and the ultraviolet temperature change measurement data evaluation processor. The nucleic acid higher-order structure analysis device of the present invention may include a data storage unit that stores the higher-order structure determination data of the test nucleic acid generated by the higher-order structure determination processor and / or a data recording unit that records the data.

[0113] The nucleic acid higher-order structure analysis device of the present invention comprises: a circular dichroism temperature change measurement unit that performs circular dichroism temperature change measurement on a test nucleic acid at a single wavelength in the range of 220 to 320 nm; an ultraviolet temperature change measurement unit that measures ultraviolet temperature changes of a test nucleic acid at a single wavelength in the range of 290 to 300 nm; may further comprise:

[0114] The nucleic acid higher-order structure analysis device of the present invention may further include a nuclear magnetic resonance spectrum analysis processing unit, a G-quadruplex structure topology discrimination processing unit, an i-motif discrimination processing unit, a circular dichroism temperature change measurement data evaluation processing unit, and an ultraviolet temperature change measurement data evaluation processing unit, and optionally a display unit that displays result data generated by the circular dichroism spectrum data processing unit or higher-order structure determination data of the test nucleic acid generated by the higher-order structure determination processing unit.

[0115] In one embodiment, a data receiving unit for receiving proton nuclear magnetic resonance spectrum data, circular dichroism spectrum data, circular dichroism temperature change measurement data, and ultraviolet temperature change measurement data of the test nucleic acid in the sample solution; a nuclear magnetic resonance spectrum analysis processor that detects imino proton peaks within a chemical shift value range of 10.0 to 16.0 ppm in the proton nuclear magnetic resonance spectrum of the test nucleic acid, and determines the type of base pair detected in the nucleic acid based on the distribution of chemical shift values ​​of the detected peaks; a G-quadruplex topology determination processor that, when it is determined that a Hoogsteen base pair between guanine and guanine has been detected in the nuclear magnetic resonance spectrum analysis processor, determines whether the topology of the G-quadruplex belongs to a parallel type, a hybrid type, or an antiparallel type by performing principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid; and a circular dichroism temperature change measurement data evaluation processing unit that performs curve fitting using a thermodynamic equilibrium model on circular dichroism temperature change measurement data of the test nucleic acid at a single wavelength in the range of 220 to 320 nm, and evaluates the conformity to a state model that shows equilibrium in three or more states; an ultraviolet temperature change measurement data evaluation processing unit that performs curve fitting using a thermodynamic equilibrium model on ultraviolet temperature change measurement data of the test nucleic acid at a single wavelength within a range of 290 to 300 nm, and evaluates the conformity to a state model that shows equilibrium in two or more states; a higher-order structure determination processor that determines the higher-order structure of the test nucleic acid based on result data generated by the nuclear magnetic resonance spectrum analysis processor, the G-quadruplex structure topology discrimination processor, the circular dichroism temperature change measurement data evaluation processor, and the ultraviolet temperature change measurement data evaluation processor; and In one embodiment, the nucleic acid higher-order structure analyzer may include a G-quadruplex topology discrimination processor that calculates an estimated value of the mixing ratio of two or more test nucleic acids in a sample solution based on the results of principal component analysis and clustering analysis of circular dichroism spectra, as necessary, or may include a circular dichroism spectrum data processor that calculates an estimated value of the mixing ratio of two or more test nucleic acids in a sample solution based on the results of principal component analysis and clustering analysis of circular dichroism spectra.

[0116] The nucleic acid higher-order structure analysis device of the present invention is configured so as to be able to carry out the above-described method of the present invention. [Example]

[0117] The present invention will be described in more detail below using examples, although the technical scope of the present invention is not limited to these examples.

[0118] [Example 1] Analysis of the higher-order structure of oligonucleotides of IGA3 and c-MYC promoter sequences in solvent The insulin aptamer IGA3 has been reported to form a parallel G4 structure (e.g., Yoshida W., et al., (2009) Biosens. Bioelectron., 24, 1116-1120). The c-MYC promoter sequence has also been reported to form a parallel G4 structure (e.g., Phan AT, et al., (2004) J. Am. Chem. Soc., 126, 8710-8716). To investigate the G4 structure of IGA3, an oligonucleotide DNA (5'-GGTGGTGGGGGGGGTTGGTAGGGTGTCTTC-3') consisting of the base sequence shown in SEQ ID NO: 1, and the c-MYC promoter sequence (5'-TGAGGGTGGGTAGGGTGGGTAA-3') consisting of the base sequence shown in SEQ ID NO: 2, circular dichroism (CD) spectroscopy was performed under different solvent conditions (water, 0.01x PBS, 0.05x PBS, 0.2x PBS, 0.5x PBS, and 1x PBS, and potassium phosphate buffer for the c-MYC promoter sequence). Furthermore, to investigate the involvement of the guanine-rich sequence of IGA3 in the formation of higher-order structures, we used an oligonucleotide in which guanine was replaced with thymine, T-IGA3 (5'-TTTTTTTTTTTTTTTTTTTATTTTTTCTTC-3'; SEQ ID NO: 3).

[0119] CD spectra were measured at 25°C using a J-820 circular dichroism spectrometer (JASCO Corporation). Spectra were measured using a quartz cuvette with a cell path length of 1 mm. The scan rate was 50 nm / min, the data acquisition interval was 0.5 nm, and the response time was 4 seconds. The average of four scans was shown as the resulting spectrum. CD spectra were measured in this example using oligonucleotides (IGA3, T-IGA3, and c-MYC promoter sequence) at a concentration of 200 μg / mL under each solvent condition.

[0120] Principal component analysis (PCA) of the CD spectra was performed to identify the higher-order structure (topology) of the IGA3 and c-MYC promoter sequences. A PCA model was created using a library of 30 reference CD spectra of G4 structures with known G4 topologies (Table 1; reference data).

[0121] [Table 1]

[0122] In Table 1, Nos. 1 to 22 are previously reported CD spectra, and Nos. 23 to 30 are CD spectra obtained by the inventors of the present application.

[0123] Hierarchical clustering (HCA) was performed using five principal components, and a PCA score plot was created (Figure 2). The principal component scores obtained by PCA of the CD spectra were plotted on the PCA score plot created by PCA / HCA analysis of the reference data.

[0124] Furthermore, for thermodynamic evaluation, circular dichroism (CD) temperature shift measurements and ultraviolet (UV) temperature shift measurements were performed using a J-820 circular dichroism spectrometer (JASCO Corporation). Using 200 μg / mL of oligonucleotides in each solvent condition, CD values ​​at 262 nm for CD temperature shift measurements and absorbance at 295 nm for UV temperature shift measurements were measured using a quartz cuvette with a 1 mm optical path length. Data were acquired in 0.1°C increments over the range of 4 to 95°C (IGA3) or 10 to 90°C (c-MYC promoter sequence) with a temperature gradient of 1°C / min, and the response time was 1 second. The cations (Na + , K. + The effect of ) was evaluated by adding various concentrations of NaCl and / or KCl to PBS.

[0125] The temperature change data obtained by CD temperature change measurement and UV temperature change measurement were subjected to curve fitting using a 2-state model (see below [1]), a 3-state model (see below [2]), and a 4-state model (see below [3]), respectively, to calculate the thermodynamic parameters (Tm, ΔH, ΔS). The search parameters were Tm and ΔH (see below [4] and [5]).

[0126]

number

[0127]

number

[0128]

number

[0129]

number

[0130]

number

[0131] In the above formula, θ obs is the measured CD value (CD temperature change measurement) or absorbance (UV temperature change measurement).

[0132]

number

[0133]

number

[0134] We also performed NMR (nuclear magnetic resonance) measurements using IGA3, T-IGA3, and c-MYC promoter oligonucleotides, as well as dGMP. All NMR experiments were performed using an Avance III HD 500 MHz nuclear magnetic resonance spectrometer (Bruker BioSpin) equipped with a CryoProbe Prodigy. Each oligonucleotide (200 μM) and dGMP (14.4 mM) were prepared in various concentrations in PBS, using 95% / 5% HO / DO or DO. 1 H-NMR measurements were performed. For HD exchange studies, the oligonucleotides and dGMP were dried at room temperature under a nitrogen atmosphere and dissolved in the same volume of DO as before drying. NMR spectra were analyzed using TopSpin software version 3.6.2 (Bruker BioSpin).

[0135] As shown in Figure 3, in 1x PBS, IGA3 exhibited a positive peak near 262 nm and a negative peak near 242 nm. With decreasing PBS concentration, the intensity of the positive and negative peaks at 262 and 242 nm decreased, with the positive peak top shifting to 258 nm in water. In contrast, in 0.05x to 1x PBS, the c-MYC promoter sequence exhibited a positive peak near 265 nm and a negative peak near 244 nm. In 0.01x PBS, the intensity of the positive and negative peaks at 265 nm and 244 nm decreased, with the positive peak top shifting to 257 nm in water. Evaluation based solely on CD spectra indicated that IGA3 formed a parallel G4 structure similar to the c-MYC promoter sequence.

[0136] However, PCA (principal component analysis) showed that the c-MYC promoter sequence formed a parallel G4 structure in PBS at various concentrations, whereas IGA3 did not (Figure 4). Furthermore, measurement of absorbance at 295 nm (derived from the G4 structure) in 1x PBS (UV temperature change measurement) showed a decrease in absorbance between 55 and 75°C for c-MYC, whereas no change in absorbance was observed with temperature for IGA3 (Figure 5). These results indicate that IGA3 in PBS does not have a parallel G4 structure.

[0137] In PCA, the first principal component was the anti-anti glycosidic bond conformation, and the second principal component was the syn-anti glycosidic bond conformation, with a guanine-guanine base step. The CD spectrum of the anti-anti glycosidic bond conformation showed a positive peak at 260 nm and a negative peak at 240 nm, which was similar to the CD spectrum of IGA3 in PBS at various concentrations.

[0138] Therefore, it was shown that IGA3 in PBS does not have a parallel G4 structure, but forms a structure (non-G4 structure) with many anti-anti glycosidic bonds similar to the parallel G4 structure.

[0139] The G4 structure is known to dynamically change its conformation (three-dimensional structure) depending on environmental conditions, especially the coordinated cations. + and / or K +To evaluate the effect of NaCl and / or KCl on the CD spectra of IGA3 in 0.01x PBS, we added NaCl and / or KCl at various concentrations up to a concentration equivalent to that of 1x PBS and measured the CD spectra (Figure 6). When NaCl or KCl alone was added (Figures 6A and 6B), the intensity of the positive peak near 260 nm and the negative peak near 240 nm increased with each cation concentration, but the peak intensities were not comparable to those of 1x PBS. On the other hand, when NaCl and KCl were added simultaneously (Figure 6C), the intensity of the positive peak near 260 nm and the negative peak near 240 nm increased depending on the cation concentration. The CD spectra of IGA3 in 0.01x PBS with 160 mM NaCl and 4.5 mM KCl, equivalent to those of 1x PBS, showed the same intensity as those of IGA3 in 1x PBS (Figure 6D). The cation Na + and K. + It was shown that when coexisting with IGA3, it is strongly involved in the formation of higher-order structures.

[0140] The imino protons from the G4 structure in H2O (corresponding to the Hoogsteen base pair between guanine and guanine) 1H-NMR analysis revealed chemical shifts between 10 and 12 ppm. The results of NMR analysis of the imino protons of each oligonucleotide and dGMP are shown in Figure 7. A sharp imino proton peak was detected in the 10-12 ppm range for the c-MYC promoter sequence in 1x PBS, whereas IGA3 in 1x PBS and water exhibited broad imino proton peaks (Figure 7A). CD spectra and NMR analysis of the imino protons for the c-MYC promoter sequence in 100 mM potassium phosphate buffer (pH 7.0) indicated that the c-MYC promoter sequence forms a single parallel G4 structure. The PCA also indicated that the c-MYC promoter sequence in 100 mM potassium phosphate buffer (pH 7.0) forms a parallel G4 structure (Figure 4). The NMR spectra of the c-MYC promoter sequence in 1x PBS and 100 mM potassium phosphate buffer (pH 7.0) showed similar imino proton peaks, further confirming that the c-MYC promoter sequence in 1x PBS forms a parallel G4 structure. The imino proton peaks were not detected in dGMP and T-IGA3 (Figure 7), indicating that the broad NMR imino proton peak of IGA3 originates from the structure formed by the guanine in IGA3. The NH group of free dGMP is 1 Although it is not detected by H-NMR, it has been reported that when incorporated into a G4 structure, the NH group is detected as an imino proton signal (Wang KB, et al., (2020) J. Am. Chem. Soc., 142, 5204-5211). Therefore, the broad imino proton peak detected in IGA3 was thought to be derived from the imino proton of guanine, which has a slower exchange rate. Since the imino proton formed and stabilized the G4 structure exhibits a slower exchange rate with the solvent proton, the solvent was changed to DO. 1H-NMR was further evaluated (Figure 7B). In DO, the c-MYC promoter sequence and IGA3 in 1x PBS showed residual peaks, whereas the peaks of IGA3 in water almost disappeared. These CD and NMR analyses indicated that IGA3 forms a specific structure distinct from the G4 structure, and that this structure is stabilized by cations.

[0141] To evaluate the possibility of a structure other than G4 formed by IGA3, we performed CD temperature shift measurements of the positive peak at 262 nm observed in the CD spectra (Figures 8 and 9). As shown in Figure 8, the CD temperature shift spectrum (262 nm) of the c-MYC promoter sequence exhibited a sigmoid curve, with the inflection point shifting to higher temperatures with each PBS concentration. In contrast, the CD temperature shift spectrum (262 nm) of IGA3 did not exhibit a regular change with PBS concentration, and multiple inflection points were observed (Figure 9). Curve fitting of the CD temperature shift spectrum using a thermodynamic equilibrium model indicated that the c-MYC promoter sequence corresponded to a two-state model at all PBS concentrations (Figure 8), whereas IGA3 corresponded to a three-state or four-state model, indicating the formation of an intermediate (Figure 9). Table 2 lists the thermodynamic parameters calculated from these models. It has been reported that the thermodynamic parameters of the parallel G4 structure formation of the c-MYC promoter sequence calculated from differential scanning calorimetry (DSC) indicate a two-state model (Jagannath J., Klaus W. (2020) Chem. Eur. J., 26, 17242-17251). The CD temperature shift spectroscopy of the c-MYC promoter sequence in 1x PBS confirmed the previously reported DSC results (Tm: 73.5°C, ΔH: -50.5 kcal mol -1 , ΔS: -145.7 cal mol -1 K -1 , in 10 mM potassium phosphate buffer), which is thought to reflect the formation of a G4 structure. On the other hand, IGA3 did not show the thermodynamic parameters that would be indicative of the formation of a G4 structure.

[0142] [Table 2]

[0143] Electrophoresis was performed to examine the structures formed by IGA3, T-IGA3, and the c-MYC promoter sequence in various concentrations of PBS. The results are shown in Figure 10. For IGA3, a monomer band (30 bp) and a band of approximately 180 bp were detected in each concentration of PBS. Bands exceeding 500 bp were also detected in 0.05x PBS to 1x PBS. On the other hand, for the c-MYC promoter sequence and T-IGA3, only the monomer band was detected in 0.01x PBS and 1x PBS. The approximately 180 bp band and the band exceeding 500 bp indicate the presence of guanine-related polymers in IGA3. Based on the above results, a decision tree for the higher-order structure of nucleic acids (Figure 1) was constructed.

[0144] [Example 2] Determination of the higher-order structure of IGA3 based on a decision tree 1) NMR The structure of insulin aptamer IGA3, which is an oligonucleotide (5'-GGTGGTGGGGGGGGTTGGTAGGGTGTCTTC-3') consisting of the base sequence shown in SEQ ID NO: 1, was determined using a decision tree. IGA3 was dissolved in 1x phosphate buffered saline (PBS) [10 mM NaHPO, 1.76 mM KHPO, 137 mM NaCl, 2.68 mM KCl, pH 7.4] at a concentration of 200 μM, transferred to a 3 mm diameter NMR sample tube, and used as a measurement sample. 1 H-NMR measurement was carried out under the following conditions: Equipment: AVANCE III HD 500 MHz nuclear magnetic resonance spectrometer (Bruker BioSpin) Observation frequency: 500.1 MHz Observation kernel: 1 H nucleus Solvent: H2O : D2O = 95 : 5 Temperature: 25℃ Solvent elimination: ES (Excitation sculpting) method (4.70 ppm) Accumulation count: 512 times

[0145] The NMR spectra were analyzed using TopSpin software version 3.6.2 (Bruker BioSpin). The signal ratio (s) was calculated by dividing the signal in the 10.0-12.0 ppm range by (s) and the signal in the -1.0--2.0 ppm range by (n). The calculated s / n value for IGA3 was 13.2, which exceeded the s / n of 10 used as the criterion for the decision tree (Figure 1). Therefore, it was determined that imino protons were detected in the 10.0-12.0 ppm range. Since no peaks were detected outside the 10.0-12.0 ppm range, it was determined that guanine-guanine Hoogsteen base pairing was detected in IGA3. Based on this result, IGA3 was then subjected to a CD PCA / HCA assay.

[0146] 2) CD PCA / HCA assay For CD evaluation, IGA3 was dissolved in 1x PBS (10 mM NaHPO, 1.76 mM KHPO, 137 mM NaCl, 2.68 mM KCl, pH 7.4) at a concentration of 200 μg / mL and transferred to a quartz cuvette as a measurement sample. CD spectrum measurement was performed under the following measurement conditions. Equipment: Circular dichroism spectrometer J-820 (JASCO Corporation) Observation range: 200-340 nm Cell path length: 1 mm Temperature: 25℃ Response: 4 seconds Data capture interval: 0.5 nm Scanning speed: 50 nm / min Number of times accumulated: 4 times

[0147] To determine the higher-order structure of IGA3, PCA / HCA analysis was performed on the obtained CD spectrum under the following conditions. Reference: CD spectrum data of 30 G4 structures with known G4 topology (Table 1) Wavelength range: 220-320 nm Vertical axis processing: Spectral area normalization Preprocessing: Centered averaging Number of principal components: 5 Classification: Hierarchical Clustering Confidence interval: 95%

[0148] The principal component scores obtained by principal component analysis (PCA) of the CD spectrum of IGA3 were plotted on a PCA score plot created by performing PCA and hierarchical clustering analysis (HCA) using a library of 30 reference CD spectra of G4 structures with known G4 topologies (Table 1; reference data).

[0149] As a result, the principal component scores of the CD spectrum of IGA3 were not plotted in any of the parallel, antiparallel, or hybrid clusters (confidence interval 95%) in the PCA score plot of the reference data, as in Figure 4. Based on this judgment result, CD temperature change measurements were then performed on IGA3 according to the decision tree in Figure 1.

[0150] 3) CD temperature change measurement IGA3 was dissolved at a concentration of 200 μg / mL in 1x PBS [10 mM Na2HPO4, 1.76 mM KH2PO4, 137 mM NaCl, 2.68 mM KCl, pH 7.4] and transferred to a quartz cuvette for measurement. The sample was stored in the instrument at 4°C for at least 5 minutes, and temperature change spectroscopy was initiated once the sample temperature had cooled to 4°C. CD temperature change measurements were performed under the following conditions. Equipment: Circular dichroism spectrometer J-820 (JASCO Corporation) Measurement wavelength: 262 nm Cell path length: 1 mm Evaluation temperature: 4℃~95℃ Response: 1 second Data capture interval: 0.1℃ Temperature gradient: 1°C / min

[0151] The data obtained by the CD temperature change measurement were subjected to curve fitting using the 2-state model ([1] above), the 3-state model ([2] above), and the 4-state model ([3] above). The search parameters were Tm and ΔH ([4] and [5] above).

[0152] As a result, only the fitting of the 4-state model converged. Based on this judgment result, we next performed UV temperature change measurements on IGA3 according to the decision tree in Figure 1.

[0153] 4) UV temperature change measurement IGA3 was dissolved in 1x PBS [10 mM NaHPO, 1.76 mM KHPO, 137 mM NaCl, 2.68 mM KCl, pH 7.4] at a concentration of 200 μg / mL and transferred to a quartz cuvette for measurement. The sample was stored in the instrument at 4°C for at least 5 minutes, and temperature change spectroscopy was initiated once the sample temperature had cooled to 4°C. UV temperature change measurements were performed under the following conditions. Equipment: Circular dichroism spectrometer J-820 (JASCO Corporation) Measurement wavelength: 295 nm (absorbance) (wavelength derived from G4 structure) Cell path length: 1 mm Evaluation temperature: 4℃~95℃ Response: 1 second Data capture interval: 0.1℃ Temperature gradient: 1°C / min

[0154] Curve fitting was performed on the data obtained by UV temperature change measurement using the 2-state model ([1] above), the 3-state model ([2] above), and the 4-state model ([3] above). The search parameters were Tm and ΔH ([4] and [5] above). The results are shown in Figure 11A.

[0155] As a result, fitting by all models did not converge to the data obtained from UV temperature change measurements of IGA3. Therefore, the higher-order structure of IGA3 in 1x PBS was determined to correspond to the "non-G4 structure" in the decision tree. In other words, it was shown that IGA3 does not form a G4 structure. It was thought that the IGA3 in the sample of this example had a nucleic acid structure different from the G4 structure. This result was consistent with the results of Example 1. For comparison, UV temperature change measurements were performed on the c-MYC promoter sequence in the same manner as for IGA3, and curve fitting was performed, which showed compatibility with the 2-state model (Figure 11B). The structure determination was completed according to the decision tree (Figure 1).

[0156] [Example 3] Determination of the higher-order structure of the c-MYC promoter sequence based on a decision tree 1) NMR The structure of the c-MYC promoter sequence, which is an oligonucleotide (5'-TGAGGGTGGGTAGGGTGGGTAA-3') consisting of the base sequence shown in SEQ ID NO: 2, was determined.

[0157] The c-MYC promoter sequence was dissolved in 1x PBS [10 mM NaHPO, 1.76 mM KHPO, 137 mM NaCl, 2.68 mM KCl, pH 7.4] at a concentration of 200 μM and transferred to a 3 mm diameter NMR sample tube to be used as the measurement sample. 1 H-NMR measurement was carried out under the following conditions: Equipment: AVANCE III HD 500 MHz nuclear magnetic resonance spectrometer (Bruker BioSpin) Observation frequency: 500.1 MHz Observation kernel: 1 H nucleus Solvent: H2O : D2O = 95 : 5 Temperature: 25℃ Solvent elimination: ES (Excitation sculpting) method (4.70 ppm) Accumulation count: 512 times

[0158] The NMR spectra were analyzed using TopSpin software version 3.6.2 (Bruker BioSpin). The signal ratio (s / n) was calculated by dividing the signal between 10.0 and 12.0 ppm by (s) and the signal between -1.0 and -2.0 ppm by (n). The calculated s / n value for the c-MYC promoter sequence was 95.2, which exceeded the s / n of 10, the criterion used in the decision tree (Figure 1). Therefore, imino protons were detected. Since no peaks were detected outside the 10.0 to 12.0 ppm range, it was determined that guanine-guanine Hoogsteen base pairing was detected in the c-MYC promoter sequence. Based on this result, the c-MYC promoter sequence was then subjected to CD PCA / HCA assay.

[0159] 2) CD PCA / HCA assay For circular dichroism (CD) evaluation, the c-MYC promoter sequence was dissolved at a concentration of 200 μg / mL in 1x PBS (10 mM NaHPO, 1.76 mM KHPO, 137 mM NaCl, 2.68 mM KCl, pH 7.4) and transferred to a quartz cuvette as the measurement sample. CD spectra were measured under the following conditions: Equipment: Circular dichroism spectrometer J-820 (JASCO Corporation) Observation range: 200-340 nm Cell path length: 1 mm Temperature: 25℃ Response: 4 seconds Data capture interval: 0.5 nm Scanning speed: 50 nm / min Number of times accumulated: 4 times

[0160] To determine the higher-order structure of the c-MYC promoter sequence, PCA / HCA analysis was performed on the obtained CD spectrum under the following conditions. Reference: CD spectrum data of 30 G4 structures with known G4 topology (Table 1) Wavelength range: 220-320 nm Vertical axis processing: Spectral area normalization Preprocessing: Centered averaging Number of principal components: 5 Classification: Hierarchical Clustering Confidence interval: 95%

[0161] The principal component scores obtained by PCA of the CD spectrum of the c-MYC promoter sequence were plotted on a PCA score plot created from a library of 30 reference CD spectra of G4 structures with known G4 topologies (Table 1) in the previous example.

[0162] As a result, the principal component scores of the CD spectrum of the c-MYC promoter sequence were plotted in a parallel cluster (confidence interval 95%) in the PCA score plot created from the reference data. Therefore, the higher-order structure of the c-MYC promoter sequence was determined to be "parallel" in the decision tree (Figure 1). This result was consistent with the results of Example 1. Since the CD PCA / HCA analysis resulted in a "parallel" determination, the c-MYC promoter sequence was determined to have a parallel G4 structure according to the decision tree (Figure 1), and the structure determination was completed.

[0163] [Example 4] Determination of the higher-order structure of an oligonucleotide mixture based on a decision tree - 1 To determine the higher-order structure of the oligonucleotide mixture, a mixture of ParaG4 (5'-GGGGCGGGGCGGGGCGGGGT-3'; SEQ ID NO: 4), an oligonucleotide DNA known to have a parallel G4 structure, and HybrG4 (5'-TAGGGTTAGGGTTAGGGTTAGGGTT-3'; SEQ ID NO: 5), an oligonucleotide DNA known to have a hybrid G4 structure, was used.

[0164] ParaG4 and HybrG4 were dissolved in the following solvents at the following concentrations, heated at 95°C for 10 minutes, and then cooled to 25°C over 30 minutes for annealing. Concentration: 20 μM ParaG4 solvent: 20 mM potassium phosphate, 80 mM KCl (solvent pH 6.8) HybrG4 solvent: 20 mM potassium phosphate, 70 mM KCl (solvent pH 7.0)

[0165] Measurement samples were prepared by mixing the ParaG4 and HybrG4 mixtures prepared as described above at room temperature in volume ratios (concentration ratios were also the same) of ParaG4:HybrG4 = 10:0, 9:1, 8:2, 7:3, 6:4, 5:5, 4:6, 3:7, 2:8, 1:9, and 0:10. These samples were transferred to quartz cells, and CD spectra were measured under the following conditions. Equipment: Circular dichroism spectrometer J-1500 (JASCO Corporation) Observation range: 220-320 nm Cell path length: 1 mm Temperature: room temperature Bandwidth: 1 nm Response: 4 seconds Data capture interval: 0.1 nm Scanning speed: 50 nm / min Number of times accumulated: 4 times

[0166] To determine the higher-order structure of the oligonucleotide, PCA / HCA analysis was performed on the obtained CD spectrum under the following conditions. Reference: CD spectrum data of 30 G4 structures with known G4 topology (Table 1) Wavelength range: 220-320 nm Vertical axis processing: Spectral area normalization Preprocessing: Centered averaging Number of principal components: 5 Classification: Hierarchical Clustering Confidence interval: 95%

[0167] The principal component scores obtained by PCA of the CD spectra were plotted on a PCA score plot created from a library of 30 reference CD spectra (Table 1) of G4 structures with known G4 topologies in the above examples.

[0168] The results are shown in Figure 12. PCA / HCA analysis of mixture samples in which the ParaG4 and HybrG4 ratio was linearly changed (ParaG4:HybrG4 ratio = 0:10 to 10:0) showed that the plot of the ParaG4 and HybrG4 mixture shifted from a parallel cluster to a hybrid cluster (a transition to an equilibrium state) while maintaining linearity as the mixture ratio changed. These results can be used to estimate the mixture ratio of ParaG4 and HybrG4 in any mixture sample by principal component regression.

[0169] [Example 5] Determination of the higher-order structure of an oligonucleotide mixture based on a decision tree - 2 To determine the higher-order structure of the oligonucleotide mixture, a mixture of ParaG4 (5'-GGGGCGGGGCGGGGCGGGGT-3'; SEQ ID NO: 4), an oligonucleotide DNA known to have a parallel G4 structure, and HybrG4 (5'-TAGGGTTAGGGTTAGGGTTAGGGTT-3'; SEQ ID NO: 5), an oligonucleotide DNA known to have a hybrid G4 structure, was used.

[0170] 1) NMR ParaG4 and HybrG4 were dissolved in the following solvents at the following concentrations, heated at 95°C for 10 minutes, and then cooled to room temperature (approximately 25°C) for annealing. Concentration: 200 μM ParaG4 solvent: 20 mM potassium phosphate, 80 mM KCl (solvent pH 7.0) HybrG4 solvent: 20 mM potassium phosphate, 70 mM KCl (solvent pH 7.0)

[0171] The ParaG4 and HybrG4 prepared as described above were mixed at a volume ratio of ParaG4:HybrG4 = 1:1 (the same applies to the concentration ratio) to prepare a measurement sample, a mixture of 100 μM ParaG4 and HybrG4 oligonucleotides. This measurement sample was analyzed under the following measurement conditions: 1 H-NMR measurements were carried out. Equipment: AVANCE III HD 500 MHz nuclear magnetic resonance spectrometer (Bruker BioSpin) Observation frequency: 500.1 MHz Observation kernel: 1 H nucleus Solvent: H2O : D2O = 95 : 5 Temperature: 25℃ Solvent elimination: ES (Excitation sculpting) method (4.70 ppm) Number of times accumulated: 1024

[0172] The NMR spectra obtained were analyzed using TopSpin software version 3.6.2 (Bruker BioSpin) (Figure 13). The signal ratio (s) was calculated by dividing the signal in the 10.0-12.0 ppm range by (s) and the signal ratio (n) by the signal in the -1.0--2.0 ppm range. The calculated s / n value for this oligonucleotide mixture was 129.4, which exceeded the s / n of 10 used as the criterion for the decision tree (Figure 1). Therefore, it was determined that imino protons were detected in the 10.0-12.0 ppm range. Since no peaks were detected outside the 10.0-12.0 ppm range, it was determined that guanine-guanine Hoogsteen base pair formation was detected in this oligonucleotide mixture. Based on this result, this oligonucleotide mixture was then subjected to a CD PCA / HCA assay.

[0173] 2) CD PCA / HCA assay ParaG4 and HybrG4 were dissolved in the following solvents at the following concentrations, heated at 95°C for 10 minutes, and then cooled to room temperature (approximately 25°C) for annealing. Concentration: 20 μM ParaG4 solvent: 20 mM potassium phosphate, 80 mM KCl (solvent pH 7.0) HybrG4 solvent: 20 mM potassium phosphate, 70 mM KCl (solvent pH 7.0)

[0174] The ParaG4 and HybrG4 oligonucleotides prepared as described above were mixed at a volume ratio (concentration ratio) of 1:1 to prepare a 10 μM mixture of ParaG4 and HybrG4 oligonucleotides. CD spectra were measured for this sample under the following conditions: Equipment: Circular dichroism spectrometer J-1500 (JASCO Corporation) Observation range: 220-320 nm Cell path length: 1 mm Temperature: 25℃ Response: 4 seconds Data capture interval: 0.5 nm Scanning speed: 50 nm / min Number of times accumulated: 4 times

[0175] To determine the higher-order structure of this oligonucleotide mixture, PCA / HCA analysis was performed on the obtained CD spectrum under the following conditions. Reference: CD spectrum data of 30 G4 structures with known G4 topology (Table 1) Wavelength range: 220-320 nm Vertical axis processing: Spectral area normalization Preprocessing: Centered averaging Number of principal components: 5 Classification: Hierarchical Clustering Confidence interval: 95%

[0176] Principal component scores obtained by principal component analysis (PCA) of the CD spectrum of this oligonucleotide mixture were plotted on a PCA score plot created by performing principal component analysis (PCA) and hierarchical clustering analysis (HCA) using a library of 30 reference CD spectra of G4 structures with known G4 topologies (Table 1; reference data) (Figure 14).

[0177] As a result, the principal component scores of the CD spectrum of this oligonucleotide mixture were not plotted in any of the parallel, antiparallel, or hybrid clusters (confidence interval 95%) in the PCA score plot of the reference data. Based on this judgment result, CD temperature change measurements were then performed on this oligonucleotide mixture according to the decision tree in Figure 1.

[0178] 3) CD temperature change measurement After the CD spectrum measurement in 2) above in this example, the measurement sample was stored in the device at 4°C for 5 minutes or more, and the CD temperature change spectrum measurement was started after the sample temperature had dropped to 4°C. The CD temperature change measurement was carried out under the following conditions. Equipment: Circular dichroism spectrometer J-1500 (JASCO Corporation) Measurement wavelength: 262 nm Cell path length: 1 mm Evaluation temperature: 4℃~95℃ Response: 1 second Data capture interval: 0.2℃ Temperature gradient: 2°C / min

[0179] The data obtained by the CD temperature change measurement were subjected to curve fitting using the 2-state model ([1] above), the 3-state model ([2] above), and the 4-state model ([3] above). The search parameters were Tm and ΔH ([4] and [5] above).

[0180] As a result, only the fitting of the 3-state model converged (FIG. 15). Based on this judgment result, UV temperature change measurement was then performed on this oligonucleotide mixture according to the decision tree in FIG.

[0181] 4) UV temperature change measurement After the CD spectrum measurement in 2) above in this example, the measurement sample was stored in the device at 4°C for 5 minutes or more, and the UV temperature change spectrum measurement was started after the sample temperature had dropped to 4°C. The UV temperature change measurement was carried out under the following conditions. Equipment: Circular dichroism spectrometer J-1500 (JASCO Corporation) Measurement wavelength: 295 nm (absorbance) (wavelength derived from G4 structure) Cell path length: 1 mm Evaluation temperature: 4℃~95℃ Response: 1 second Data capture interval: 0.2℃ Temperature gradient: 2°C / min

[0182] Curve fitting was performed on the data obtained by UV temperature change measurement using the 2-state model ([1] above), the 3-state model ([2] above), and the 4-state model ([3] above). The search parameters were Tm and ΔH ([4] and [5] above).

[0183] As a result, only the fitting of the 3-state model converged (Figure 16). Therefore, the higher-order structure of this oligonucleotide mixture was determined to correspond to the "nucleic acid mixture containing G4" in the decision tree. In other words, the oligonucleotide mixture explicitly containing two types of G4 structures was correctly determined to be a "nucleic acid mixture containing G4." The structure determination was completed according to the decision tree (Figure 1).

[0184] 5) Estimation of mixture ratio The PCA score plot obtained by the CD PCA / HCA assay in 2) above in this Example and the PCA score plots obtained in Example 4 for mixture ratios of 0:10 to 10:0 were subjected to multiple regression analysis based on [8] (multiple regression C = Xβ) to estimate the mixture ratio of this oligonucleotide mixture. As a result, the content ratio (concentration ratio) of HybrG4 was calculated to be 0.42. Because this oligonucleotide mixture was prepared with an explicit ParaG4:HybrG4 = 1:1 ratio, the value calculated by multiple regression analysis was approximately consistent with the actual HybrG4 content and was therefore reasonable. Thus, it was demonstrated that the present invention allows the mixture ratio of a mixture of multiple nucleic acid molecules to be estimated within a generally reasonable range.

Claims

1. A method for analyzing the higher-order structure of a test nucleic acid in a sample solution, comprising: a) detecting imino proton peaks within a chemical shift value range of 10.0 to 16.0 ppm in the proton nuclear magnetic resonance spectrum of the test nucleic acid, and determining the type of base pair detected in the test nucleic acid based on the distribution of chemical shift values ​​of the detected peaks; b) when it is determined in step a) that a guanine-guanine Hoogsteen base pair has been detected, determining whether the topology of the G-quadruplex structure belongs to a parallel type, a hybrid type, or an antiparallel type by performing principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid; c) if the topology of the G-quadruplex structure in step b) is found to be parallel, hybrid, or antiparallel, determining that the test nucleic acid forms a G-quadruplex structure having that topology; and if the topology of the G-quadruplex structure in step b) is found to be neither parallel, hybrid, nor antiparallel, performing circular dichroism temperature shift measurement of the test nucleic acid at a single wavelength in the range of 220 to 320 nm; d) performing curve fitting using a thermodynamic equilibrium model on the circular dichroism temperature change measurement data acquired in step c) and evaluating the fit to a state model showing equilibrium of three or more states; e) if step d) indicates a fit to a state model showing equilibrium of three or more states, performing ultraviolet temperature change measurements on the test nucleic acid at a single wavelength in the range of 290 to 300 nm; f) performing curve fitting using a thermodynamic equilibrium model on the ultraviolet temperature change measurement data acquired in step e) and evaluating the fit to a state model showing equilibrium of two or more states; g) determining that the test nucleic acid forms two or more nucleic acid higher-order structures including a G-quadruplex structure if step f) shows a match to a state model showing equilibrium in two or more states, and determining that the test nucleic acid does not form a G-quadruplex structure if step f) shows no match to any state model showing equilibrium in two or more states; A method comprising:

2. 2. The method according to claim 1, wherein the type of base pair detected in step a) is a Hoogsteen base pair between guanine and guanine, a cytosine-cytosine base pair, or a Watson-Crick base pair.

3. The method according to claim 1, wherein when it is determined that a cytosine-cytosine base pair is detected in step a), the i-motif is identified by further performing principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid.

4. The method according to any one of claims 1 to 3, further comprising measuring the proton nuclear magnetic resonance spectrum and / or the circular dichroism spectrum of the test nucleic acid.

5. 2. The method of claim 1, wherein the circular dichroism spectrum is measured at wavelengths in the range of 220 to 320 nm inclusive.

6. The method of claim 1, wherein the principal component analysis and clustering analysis in step b) are performed using principal component models and clusters created from reference data of circular dichroism spectra of G-quadruplex structure-forming nucleic acids whose G-quadruplex structure topologies are known.

7. 2. The method according to claim 1, wherein the circular dichroism temperature change measurement is performed at a single wavelength within the range of 260 to 265 nm.

8. 2. The method of claim 1, wherein the ultraviolet temperature change measurement is performed at a single wavelength of 295 nm.

9. The method of claim 1 , wherein the sample solution comprises a cation.

10. The method of claim 1 , wherein the sample solution contains two or more types of the test nucleic acids.

11. The method according to claim 10, further comprising calculating an estimated value of a mixing ratio of two or more types of the test nucleic acids in the sample solution based on the results of principal component analysis and clustering analysis of the circular dichroism spectra.

12. The method of claim 1 , wherein the sample solution further contains a substance that interacts or is likely to interact with the test nucleic acid.

13. a data receiving unit for receiving proton nuclear magnetic resonance spectrum data, circular dichroism spectrum data, circular dichroism temperature change measurement data, and ultraviolet temperature change measurement data of the test nucleic acid in the sample solution; a nuclear magnetic resonance spectrum analysis processor that detects imino proton peaks within a chemical shift value range of 10.0 to 16.0 ppm in the proton nuclear magnetic resonance spectrum of the test nucleic acid, and determines the type of base pair detected in the test nucleic acid based on the distribution of chemical shift values ​​of the detected peaks; a G-quadruplex topology determination processor that, when it is determined that a guanine-guanine Hoogsteen base pair has been detected in the nuclear magnetic resonance spectrum analysis processor, determines whether the topology of the G-quadruplex belongs to a parallel type, a hybrid type, or an antiparallel type by performing principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid; and a circular dichroism temperature change measurement data evaluation processing unit that performs curve fitting using a thermodynamic equilibrium model on circular dichroism temperature change measurement data of the test nucleic acid at a single wavelength in the range of 220 to 320 nm, and evaluates the conformity to a state model showing equilibrium of three or more states; an ultraviolet temperature change measurement data evaluation processing unit that performs curve fitting using a thermodynamic equilibrium model on ultraviolet temperature change measurement data of the test nucleic acid at a single wavelength in the range of 290 to 300 nm, and evaluates the conformity to a state model that shows equilibrium in two or more states; a higher-order structure determination processor that determines the higher-order structure of the test nucleic acid based on result data generated by the nuclear magnetic resonance spectrum analysis processor, the G-quadruplex structure topology discrimination processor, the circular dichroism temperature change measurement data evaluation processor, and the ultraviolet temperature change measurement data evaluation processor; and A nucleic acid higher-order structure analysis device comprising:

14. 14. The apparatus according to claim 13, wherein the type of base pair detected by the nuclear magnetic resonance spectrum analysis processing unit is a Hoogsteen base pair between guanine and guanine, a cytosine-cytosine base pair, or a Watson-Crick base pair.

15. The apparatus of claim 13, further comprising an i-motif discrimination processing unit that, when it is determined that a cytosine-cytosine base pair has been detected in the nuclear magnetic resonance spectrum analysis processing unit, discriminates the i-motif by further performing principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid.

16. The apparatus according to claim 13, wherein the G-quadruplex topology discrimination processor performs principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid using principal component models and clusters created from reference data of circular dichroism spectra of G-quadruplex-forming nucleic acids whose G-quadruplex topologies are known.

17. a circular dichroism temperature change measurement unit that performs circular dichroism temperature change measurement on the test nucleic acid at a single wavelength in the range of 220 to 320 nm; an ultraviolet temperature change measurement unit for measuring ultraviolet temperature changes of the test nucleic acid at a single wavelength in the range of 290 to 300 nm; The apparatus of claim 13 further comprising:

18. 14. The apparatus according to claim 13, wherein the G-quadruplex topology discrimination processor further calculates an estimated value of a mixing ratio of the two or more types of test nucleic acids in the sample solution based on results of principal component analysis and clustering analysis of the circular dichroism spectra.

19. A program for analyzing the higher-order structure of a test nucleic acid in a sample solution, comprising: a) detecting imino proton peaks within a chemical shift value range of 10.0 to 16.0 ppm in the proton nuclear magnetic resonance spectrum of the test nucleic acid, and determining the type of base pair detected in the test nucleic acid based on the distribution of chemical shift values ​​of the detected peaks; b) when it is determined in step a) that a guanine-guanine Hoogsteen base pair has been detected, determining whether the topology of the G-quadruplex structure belongs to a parallel type, a hybrid type, or an antiparallel type by performing principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid; c) determining that the test nucleic acid forms a G-quadruplex structure having a topology that is parallel, hybrid, or antiparallel in step b), and evaluating the suitability of the test nucleic acid to a state model showing equilibrium in three or more states by performing curve fitting using a thermodynamic equilibrium model on data obtained by measuring the temperature change of the G-quadruplex structure at a single wavelength in the range of 220 to 320 nm by circular dichroism in step b) if the topology of the G-quadruplex structure does not belong to any of parallel, hybrid, or antiparallel. d) if step c) indicates a fit to a state model showing equilibrium in three or more states, performing curve fitting using a thermodynamic equilibrium model on the ultraviolet temperature change measurement data of the test nucleic acid at a single wavelength in the range of 290 to 300 nm, and evaluating the fit to a state model showing equilibrium in two or more states; e) determining that the test nucleic acid forms two or more nucleic acid higher-order structures including a G-quadruplex structure if step d) shows a match to a state model showing equilibrium in two or more states, and determining that the test nucleic acid does not form a G-quadruplex structure if step d) does not show a match to any state model showing equilibrium in two or more states; A program that causes a computer to execute a process including the steps of:

20. 20. The program according to claim 19, wherein the type of base pair detected in step a) is a Hoogsteen base pair between guanine and guanine, a cytosine-cytosine base pair, or a Watson-Crick base pair.

21. The program according to claim 19, wherein, when it is determined in step a) that a cytosine-cytosine base pair has been detected, the processing further comprises performing principal component analysis and clustering analysis of the circular dichroism spectrum of the test nucleic acid to identify an i-motif.

22. 20. The program according to claim 19, wherein the principal component analysis and clustering analysis in step b) are performed using principal component models and clusters created from reference data of circular dichroism spectra of G-quadruplex structure-forming nucleic acids whose G-quadruplex structure topology is known.

23. 22. The program according to claim 19, wherein the processing further comprises calculating an estimated value of a mixing ratio of two or more types of the test nucleic acids in the sample solution based on results of principal component analysis and clustering analysis of the circular dichroism spectra.