Structure estimation method, structure estimation device, and program

The method uses a recurrence plot and transformed Jacquard distances to enhance the accuracy of three-dimensional molecular structure reconstruction by applying multidimensional scaling, addressing the limitations of existing probabilistic methods.

WO2026094489A1PCT designated stage Publication Date: 2026-05-07UNIV OF TSUKUBA
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
UNIV OF TSUKUBA
Filing Date
2025-09-26
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Current methods for reconstructing the three-dimensional structures of molecules such as DNA and RNA are limited by the use of probabilistic analytical methods and lack accuracy in utilizing binary information about spatial proximity.

Method used

A method and device that utilize a recurrence plot to generate a distance matrix, transforming Jacquard distances into transformed distances using cubic equations, and apply multidimensional scaling to reconstruct the three-dimensional structure of molecules.

Benefits of technology

The method achieves higher accuracy in reconstructing the three-dimensional structure of molecules by transforming Jacquard distances with overlapping three-dimensional spheres, resulting in improved structural reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025034100_07052026_PF_FP_ABST
    Figure JP2025034100_07052026_PF_FP_ABST
Patent Text Reader

Abstract

A structure estimation method (S1) comprising: a step (S12) for generating a recurrence plot from information about the structures of spatially adjacent moieties of a molecule by representing the primary sequence of the molecule on the time axis; a step (S13) for generating a distance matrix that provides the distance between any two vertices of the recurrence plot; a step (S14) for reconstructing time-series data by using multidimensional scaling with respect to the distance matrix; and a step (S15) for reconstructing the three-dimensional structure of the molecule from the reconstructed time-series data, wherein the distance matrix has, as a component, a post-transformation distance obtained by transforming a Jaccard distance related to the recurrence plot, and the post-transformation distance is obtained as a solution or an approximate solution of an equation in which said post-transformation distance is defined as a variable.
Need to check novelty before this filing date? Find Prior Art

Description

Structural estimation method, structural estimation apparatus, and program

[0001] The present invention relates to a structural estimation method, a structural estimation apparatus, and a program.

[0002] DNA and RNA are macromolecules consisting of nucleotides linked together in double or single strands. Proteins are macromolecules consisting of amino acids linked together in chains.

[0003] Generally, DNA, RNA, and proteins form three-dimensional structures within cells through interactions between nucleic acids, between proteins, or between nucleic acids and proteins. The three-dimensional structures of DNA, RNA, and proteins are thought to be closely related to their functions and are considered important information for understanding biological functions and diseases.

[0004] To analyze the three-dimensional structure of chromosomes, an experimental technique called chromosome conformation capture (3C assay), which detects spatially adjacent sequences, is used. Furthermore, methods such as 4C, 5C, and Hi-C have been developed to comprehensively analyze spatially adjacent nucleic acid fragment information using microarrays and next-generation sequencers. To reconstruct the three-dimensional structure from this data measuring the spatial distance between chromosomes within a cell, information on whether the distance between subsequences in three-dimensional space is close or not is often used. However, currently, it is common to use probabilistic analytical methods or auxiliary information regarding the distance between subsequences to reconstruct the three-dimensional structure (see Non-Patent Documents 1 and 2 below).

[0005] Three-dimensional structural analysis of proteins and nucleic acids utilizes methods such as X-ray crystallography and cryo-electron microscopy, as well as nuclear magnetic resonance (NMR) spectroscopy, which detects interactions between atomic nuclei, atomic bonding, and electronic states as spectra.

[0006] WO2009 / 084524WO2010 / 010675

[0007] Marc A. Marti-Renom and Leonid A. Mirny, Bridging the resolution gap in structural modeling of 3D genome organization, PLoS Computational Biology 7, e1002125(2011).Annick Lesne, Julien Riposo, Paul Roger, Axel Cournac , and Julien Mozziconacci, 3D genome reconstruction from chromosomal contacts, Nature Method 11, 1141-1143 (2014).J. - P. Eckmann, S. O. Kamphorst, and D. Ruelle, Recurrence plots of dynamical systems, Europhysics Letters 4, 973-977 (1987). N. Marwan, M. C. Romano, M. Thiel, and J. Kurths, Recurrence. plots for the analysis of complex systems, Physics Reports 438, 237-329 (2007). Yoshito Hirata, Shunsuke Horai, and Kazuyuki Aihara, Reproduction of distancemtrices and original time series from recurrence plots and their applications, European Physical Journal Special Topics 164, 13-22 (2008) Yoshito Hirata, Shunsuke Horai, and Kazuyuki Aihara, Random number generation mechanism following a normal distribution, Special Topics Application No. 2009-548035, PCT / JP2008 / 073389, Patent No. 4947476. (American patent, Patent No.(US8, 438, 202B2, Date of Patent: May 7, 2013) Yoshito Hirata, Kazuyuki Aihara, Method and apparatus for simultaneous reconstruction of multiple external forces acting on a single system, PCT / JP2009 / 003355. Yoshito Hirata, Motomasa Komuro, Shunsuke Horai, and Kazuyuki Aihara, Faithfulness of recurrence plots:A Mathematical proof, International Journal of Bifurcation and Chaos in press. volume 25, art. no. 155168 (2015) E. W. Dijkstra, A note on two problems in connexion with graphs, Numerische Mathematic 1, 269-271 (1959). J. C. Gower, Some distance properties of latent root and vector methods used in multivariate analysis, Biometrika, 53, 325-338 (1966). Masaaki Tanio, Yoshito Hirata, and Hideyuki Suzuki, Reconstruction of driving forces through recurrence plots, Physics Letters A 373, 2031-2040(2009).Zhijun Duan, Mirela Andronescu, Kevin Schutz, Sean Mcllwain, Yoo Jung Kim, Choli Lee, Hay Shendure, Stanley Fields, C. Anthony Blau, and William S.Noble, A three-dimensional model of the yeast genome, Nature 465, 363-367 (2010)Y. Hirata, A. Oda, K. Ohta and K. Aihara, Three-dimensional reconstruction of single-cell chromosome structure using recurrence plots, Sci. Rep. 6, 34982 (2016) L. Tan, D. Xing, N. Daley and X. S. Xie, Three-dimensional genomic structures of single sensory neurons in mouse visual and olfactory systems, Nature Structural & Molecular Biology 26, 297-307 (2019) Y. Hirata, Y. Kitanishi, H. Sugishita and Y. Gotoh, Fast reconstruction of an original continuous series from a recurrence plot, Chaos 31, 121101 (2021).

[0008] Therefore, the present invention proposes a technique for reconstructing the three-dimensional structure of molecular data, including nucleic acids and proteins, using a mathematical analysis method that utilizes binary information, such as whether the distance between subsequences and subregions is close in three-dimensional space (see Non-Patent Documents 3 and 4 above).

[0009] In view of the above circumstances, the present invention aims to provide a three-dimensional structure reconstruction technique for molecular data using the concept of spatial proximity, which reconstructs the three-dimensional structure of various molecules such as DNA and RNA from binary information indicating whether or not the partial sequences of those molecules are close together.

[0010] To solve the above problems, a structural estimation method according to one aspect of the present invention includes the steps of: generating a recurrence plot from information on the structure of spatially adjacent parts of a molecule by treating the linear arrangement of the molecule as a time axis; generating a distance matrix that gives the distance between any two vertices of the recurrence plot; reconstructing time series data using multidimensional scaling on the distance matrix; and reconstructing the three-dimensional structure of the molecule from the reconstructed time series data, wherein the distance matrix is ​​a matrix whose components are transformed distances obtained by transforming the Jacquard distances relating to the recurrence plot, and the transformed distances are obtained as the solution or approximate solution of an equation with the transformed distances as variables.

[0011] To solve the above problems, a structure estimation device according to one aspect of the present invention comprises: an acquisition unit that acquires information on the structure of spatially adjacent parts of a molecule; a first generation unit that generates a recurrence plot from the information on the structure of spatially adjacent parts of the molecule by treating the linear arrangement of the molecule as a time axis; a second generation unit that generates a distance matrix that gives the distance between any two vertices of the recurrence plot; a first reconstruction unit that reconstructs time series data using multidimensional scaling on the distance matrix; and a second reconstruction unit that reconstructs the three-dimensional structure of the molecule from the reconstructed time series data. The distance matrix is ​​a matrix whose components are transformed distances obtained by transforming the Jacquard distances relating to the recurrence plot, and the transformed distances are obtained as a solution or approximate solution to an equation with the transformed distances as variables.

[0012] Each aspect of the present invention may be implemented by a computer, in which case a program that implements the structural estimation device by a computer by operating the computer as each part (software element) of the structural estimation device, and a computer-readable recording medium on which the program is recorded, also fall within the scope of the present invention.

[0013] According to one aspect of the present invention, it is possible to provide a technique for reconstructing the three-dimensional structure of molecular data using the concept of spatial proximity, which reconstructs the three-dimensional structure of various molecules such as DNA and RNA from binary information indicating whether or not the partial sequences of those molecules are close together.

[0014] This is a flowchart showing the flow of the structural estimation method according to Embodiment 1 of the present invention. This is a diagram for explaining the structural estimation method according to Embodiment 1 of the present invention. This is a block diagram showing the configuration of the structural estimation device according to Embodiment 1 of the present invention. This is a block diagram showing the configuration of the structural estimation device according to Embodiment 2 of the present invention. This is a diagram for explaining the processing by the structural estimation device according to Embodiment 2 of the present invention. This is a diagram for explaining the processing by the structural estimation device according to Embodiment 2 of the present invention. This is a diagram for explaining the processing by the structural estimation device according to Embodiment 2 of the present invention. This is a flowchart showing the flow of the structural estimation method according to Embodiment 2 of the present invention. This is a diagram showing the results according to an embodiment of the present invention. This is a diagram showing the results according to an embodiment of the present invention. This is a diagram showing the results according to an embodiment of the present invention. This is a diagram showing the results according to an embodiment of the present invention. This is a diagram showing the results according to an embodiment of the present invention.

[0015] [Embodiment 1] An embodiment of the present invention will be described in detail below. Figure 1 is a flowchart showing the flow of the structure estimation method S1 according to this embodiment. The structure estimation method S1 can also be described as a method for reconstructing the three-dimensional structure of molecular data using the concept of spatial proximity. In this reconstruction method S1, as an example, the three-dimensional structure of the molecule is reconstructed by calculating the distance in three-dimensional space between any two molecular regions or sequences within the molecule, using binary information on whether or not they are spatially close to each other.

[0016] (Step S11) First, in step S11, information regarding the structure of spatially adjacent regions of a molecule (sometimes simply called data regarding the molecular structure) is obtained. Here, the molecule may be a biomolecule or a non-biomolecule. Furthermore, an example of information regarding the structure of spatially adjacent regions of the molecule is Hi-C (High-throughput chromosome conformation capture) data of a biomolecule, but this is not limited to this embodiment.

[0017] (Step S12) Next, in step S12, a recurrence plot is generated from information about the structure of spatially adjacent regions of the molecule obtained in step S11. As an example, a recurrence plot is generated from the Hi-C data of the biomolecule obtained in step S11. Here, a recurrence plot is originally a plot for visualizing time-series data. The term "recurrence plot" is not limited to this embodiment and may also be expressed as a regression plot or contact map, etc. A recurrence plot is a plot or map that uses a regression or regression-like concept and may be expressed by other terms. Therefore, the process in this step may be described as a process of generating a recurrence plot from information about the structure of spatially adjacent regions of the molecule by treating the primary sequence of the molecule as a time axis. As an example, a recurrence plot can be generated by comparing states (components of a molecule) corresponding to two time points on a two-dimensional plane where the horizontal and vertical axes are the same time axis. If the distance between the states is small, a point is plotted at the corresponding location; otherwise, no point is plotted.

[0018] The recurrence plot can also be expressed as follows: That is, if the recurrence plot is expressed as R(i,j) using the indices i and j that represent the vertical and horizontal axes above, then R(i,j) is R(i,j) = 1, if ||x i - x j|| < ε R(i,j) = 0, otherwise it is given by ε R(i,j) = 0, where x i and x j As an example, the indices i and j represent the positions of the molecular components, respectively. Furthermore, ε represents an infinitesimal distance.

[0019] Generating a recurrence plot corresponds to converting continuous time-series data into binary matrix information indicating proximity. Therefore, a recurrence plot does not include absolute value information in the time-series data. Nevertheless, it is known that the general shape of the time-series data can be reconstructed from a recurrence plot. In particular, when the points are uniformly distributed, a metric space equivalent to the original metric space is reconstructed (see Non-Patent Documents 5-8 above). A specific example of a recurrence plot will be described in more detail in Embodiment 2, which will be discussed later.

[0020] (Step S13) Next, in step S13, a distance matrix is ​​generated that gives the distance between any two vertices of the recurrence plot generated in step S12. The distance matrix generation process in this step can be achieved, for example, by the following series of processes (steps S131 to S133).

[0021] (Step S131) ​​First, a graph is generated from the recurrence plot. The process in this step may also be described as the process of treating the recurrence plot as a graph. Here, the graph includes multiple nodes (vertices) and one or more links (branches) connecting the nodes. In the graph related to this step, the time points of the recurrence plot are vertices, and if points are plotted at locations corresponding to two time points, the graph is generated by connecting the corresponding vertices with branches. Then, the following distance (also called the Jaccard distance in this specification) w(i,j) is assigned to each branch. Here, G i G is the set of time points corresponding to the points plotted on the i-th row of the recurrence plot, jis a set of time points corresponding to the points hit on the recurrence plot of the j-th row, and |A| represents the number of elements in set A.

[0022] (Step S132) Subsequently, the distance after conversion is calculated by converting the distance (Jaccard distance) assigned to each branch in Step S131. Here, as an example, the distance after conversion is obtained as the solution or approximate solution of an equation with the distance after conversion as a variable. More specifically, as an example, the distance after conversion is obtained as the solution or approximate solution of a cubic equation that includes the Jaccard distance as a coefficient and has the distance after conversion as a variable.

[0023] The conversion in this step may also be expressed as a conversion obtained by approximating the Jaccard distance with two overlapping three-dimensional spheres. More specifically, in this step, the distance after conversion is obtained by solving a cubic equation that includes the Jaccard distance w(i, j) as a coefficient By solving, the distance y after conversion can be calculated exactly or approximately as a function of the Jaccard distance w(i, j).

[0024] The above conversion in this step has the following meaning. That is, it corresponds to approximating the local distance (Jaccard distance w(i, j)) using the volumes of two overlapping three-dimensional spheres. More specifically, the Jaccard distance w(i, j) is approximated as and when y = r / d is set, the above formula A2 is obtained. The distance (corrected distance) y converted in this way is used as a new local distance, and the following processing is continued. Although the case of using a cubic equation is described in the above example, this does not limit the present embodiment, and a configuration using an equation of other degrees may also be used.

[0025] (Step S133) In step S133, the shortest path between any two vertices on the graph to which the converted distance y generated in step S132 is assigned is obtained. This can be executed, for example, by using Dijkstra's algorithm (see the above Non-Patent Document 9). Then, a distance matrix in which the distances between any two vertices are given is obtained. In other words, a distance matrix having as components the distances between any two vertices (the shortest path using the converted distance y) is generated. In this way, each process related to step S13 of generating the distance matrix is executed. The distance matrix generated using the Jaccard distance w(i, j) is also referred to as the first distance matrix DM1, and the distance matrix generated using the converted distance y is also referred to as the second distance matrix DM2.

[0026] (Step S14) Subsequently, the time-series data is reconstructed using the multidimensional scaling method for the distance matrix generated in step S13. As an example, the time-series data is reconstructed from the distance matrix generated in step S13 by using the method described in the above Non-Patent Document 10.

[0027] (Step S15) Subsequently, the three-dimensional structure of the molecule is reconstructed from the time-series data reconstructed in step S14. As an example, by regarding the primary sequence of a polymer as the time axis, the method related to the recurrence plot can be directly applied to nucleic acid sequences and amino acid sequences as well. That is, the three-dimensional structures of DNA, RNA, and proteins can be obtained (reconstructed) by using the top three components when reconstructed using the multidimensional scaling method (see the above Non-Patent Document 10). More specifically, by using the method of the above Non-Patent Document 11 for the information that consecutive sequences are spatially close and thickening the diagonal line in the center of the corresponding recurrence plot to be as thick as a line with a width of 3 or more, a method that can always reconstruct the three-dimensional structure can be constructed.

[0028] Figure 2 shows the correspondence between the method for reconstructing the original time series from the recurrence plot and the method for reconstructing the three-dimensional structure of the biomolecule from Hi-C data (structure estimation method S1). In the lower part of Figure 2, the symbol RP represents the recurrence plot generated in step S12 described above, and the symbol 3DD represents the three-dimensional structure of the biomolecule reconstructed in step S15 described above.

[0029] (Effects of the structure estimation method S1) As described above, the structure estimation method S1 according to this embodiment includes: - A step (S12) of generating a recurrence plot from information on the structure of spatially adjacent parts of the molecule by treating the linear arrangement of the molecule as a time axis; - A step (S13) of generating a distance matrix that gives the distance between any two vertices of the recurrence plot; - A step (S14) of reconstructing time series data using multidimensional scaling on the distance matrix; and - A step (S15) of reconstructing the three-dimensional structure of the molecule from the reconstructed time series data. The distance matrix is ​​a matrix whose components are the transformed distance (y = y(w(i,j))) obtained by transforming the Jacquard distance relating to the recurrence plot, and the transformed distance (y = y(w(i,j))) is obtained as a solution or approximate solution of an equation with the transformed distance as a variable.

[0030] According to the structure estimation method S1 configured in this way, the transformed distance y is calculated as a solution or approximate solution to an equation in which the transformed distance is a variable, and the three-dimensional structure of the molecule is reconstructed using the transformed distance y. Therefore, the three-dimensional structure of the molecular data can be reconstructed with high accuracy. More specifically, the structure estimation method S1 configured as described above can reconstruct the three-dimensional structure with higher accuracy compared to the case where the three-dimensional structure of the molecule is reconstructed using Jacquard distance without performing a distance transformation.

[0031] Furthermore, in the structure estimation method S1 described above, as stated above, the transformation obtained by approximating the Jacquard distance with two overlapping three-dimensional spheres is used as the distance transformation. More specifically, in this step, the transformed distance y is obtained by solving a cubic equation (equation A2) in which the Jacquard distance w(i,j) is a coefficient. According to the inventor's findings, by approximating the Jacquard distance with two overlapping three-dimensional spheres in this way, the three-dimensional structure of the molecule can be reconstructed with greater accuracy.

[0032] (Configuration of the structural estimation device 1) Next, the structural estimation device 1 according to this embodiment will be described. The structural estimation device 1 is a device for executing each step of the structural estimation method S1 described above. Figure 3 is a block diagram showing the configuration of the structural estimation device 1. As shown in Figure 3, the structural estimation device 1 includes a control unit 10, a storage unit 20, a communication unit 30, and an input / output unit 40.

[0033] (Communication Unit 30) The communication unit 30 communicates with one or more external devices of the structural estimation device 1. The communication unit 30 transmits data supplied from the control unit 10 to the external device, and supplies data received from the external device to the control unit 10. For example, the communication unit 30 supplies input data INPUT acquired from the external device to the control unit 10, and transmits output data OUTPUT derived by the control unit 10 to the external device.

[0034] (Input / Output Unit 40) The input / output unit 40 is configured to include, for example, at least one of the following input / output devices: a keyboard, mouse, display, printer, touch panel, etc. Alternatively, the input / output unit 40 may be configured to have input / output devices such as a keyboard, mouse, display, printer, touch panel, etc. connected to it. In this configuration, the input / output unit 40 receives various types of information from the connected input device to the structure estimation device 1. The input / output unit 40 also outputs various types of information to the connected output device under the control of the control unit 10A. An interface such as USB (Universal Serial Bus) can be used as the input / output unit 40.

[0035] (Storage Unit 20) The storage unit 20 stores various data referenced by the control unit 10, and various data generated by the control unit 10. As an example, the storage unit 20 stores: - Molecular structure data (HCD) - Recurrence plot (RP) - First distance matrix DM1 - Second distance matrix DM2 - Time series data TSD - Three-dimensional structure data 3DD.

[0036] (Control Unit 10) As shown in Figure 3, the control unit 10 includes a first generation unit 12, a second generation unit 13, a first reconstruction unit 14, and a second reconstruction unit 15.

[0037] (Acquisition unit 11) The acquisition unit 11 acquires information about the structure of spatially adjacent parts of the molecule (molecular structure data HCD). In other words, the acquisition unit 11 performs the processing of step S11 described above.

[0038] (First generation unit 12) The first generation unit 12 generates a recurrence plot (RP) from information about the structure of spatially adjacent regions of the molecule by treating the primary arrangement of the molecule as a time axis. In other words, the first generation unit 12 performs the process of step S12 described above. The specific process performed by the first generation unit 12 has been explained in step S12, so a redundant explanation will be omitted.

[0039] (Second generation unit 13) The second generation unit 13 generates a distance matrix that gives the distance between any two vertices of the recurrence plot. In other words, the second generation unit 13 executes the process of step S13 described above. The specific process that the second generation unit 13 executes has been explained in step S13, so a redundant explanation will be omitted.

[0040] (First reconstruction unit 14) The first reconstruction unit 14 reconstructs the time series data using multidimensional scaling on the distance matrix generated by the second generation unit 13. In other words, the first reconstruction unit 14 performs the processing of step S14 described above.

[0041] (Second reconstruction unit 15) The second reconstruction unit 15 reconstructs the three-dimensional structure of the molecule from the time-series data reconstructed by the first reconstruction unit 14. In other words, the second reconstruction unit 15 performs the process of step S15 described above.

[0042] The structure estimation device 1 configured in this manner produces the same effects as the structure estimation method S1 described above.

[0043] [Embodiment 2] Another embodiment of the present invention will be described below. For the sake of convenience of explanation, components having the same function as those described in the above embodiment will be denoted by the same reference numerals, and redundant explanations may be omitted.

[0044] (Structural Estimation Device 1A) The structural estimation device 1A according to this embodiment will now be described. The structural estimation device 1A is, for example, a device for executing each step of the structural estimation method S1 described above, but is not limited to this, and performs various processes described in this embodiment. Figure 4 is a block diagram showing the configuration of the structural estimation device 1A.

[0045] As shown in Figure 4, the structure estimation device 1A comprises a control unit 10A, a storage unit 20A, a communication unit 30, and an input / output unit 40. The communication unit 30 and the input / output unit 40 have the same configuration as in Embodiment 1, so their description is omitted.

[0046] In addition to the various types of information stored by the storage unit 20 according to Embodiment 1, the storage unit 20A also stores: • First structural data (SD1) • Second structural data (SD2).

[0047] Furthermore, in addition to the various configurations provided by the control unit 10 according to Embodiment 1, the control unit 10A includes a third reconstruction unit 16 and a structure data output unit 17. Also, in the control unit 10A, the first generation unit 12, the second generation unit 13, the first reconstruction unit 14, the second reconstruction unit 15, and the third reconstruction unit 16 may be collectively referred to simply as the reconstruction unit 18. The structure estimation device 1A may also be described as a device that estimates the structure of an object having a one-dimensional sequence. Examples of such objects include chromosomes, which are formed by a series of gene loci and are intricately folded. In the following description, the case in which the structure estimation device 1A estimates the specific structure of a chromosome will be used as an example. However, this is not limited to this embodiment.

[0048] (Acquisition Unit 11) The acquisition unit 11 acquires proximity data. Here, proximity data is data relating to whether an element constituting a one-dimensional array, which is included in an object having a one-dimensional array, is in close proximity to other elements of the one-dimensional array, to a distance of a predetermined distance or less. This proximity data corresponds to the information regarding the structure of spatially adjacent regions (molecular structure data HCD) described in Embodiment 1. The element referred to here is a division of the one-dimensional array of the object into at least two parts. For example, if the object is a chromosome, the element becomes a gene locus. The proximity data is, for example, data generated by chromosome conformation capture. The acquisition unit 11 may be described as having a configuration that acquires information regarding the structure of spatially adjacent regions of a molecule, as in Embodiment 1. Here, the molecule may be a biomolecule or a non-biomolecule. The acquisition unit 11 may be described as having a configuration that acquires Hi-C data of a biomolecule, as in Embodiment 1.

[0049] (First generation unit 12, second generation unit 13, first reconstruction unit 14, second reconstruction unit 15) The first generation unit 12, second generation unit 13, first reconstruction unit 14, and second reconstruction unit 15 reconstruct the three-dimensional structure of the molecule from information (proximity data) regarding the structure of spatially adjacent parts of the molecule, as described in Embodiment 1. Specific processing performed by each of these units will be omitted as appropriate to avoid repetition. In this embodiment, the term "proximity data" may include data such as: • Recurrence plots (recurrence plot RP generated by the first generation unit 12) generated using data generated by the chromosome conformation capture method, • Graphs (graphs generated by the second generation unit 13), • Adjacent lists, or • Distance matrices (distance matrices generated by the second generation unit 13). Here, as in Embodiment 1, the term "recurrence plot" is not limited to this embodiment and may also be expressed as a regression plot or contact map. A recurrence plot is a plot or map that uses a regression or regression-like concept, and may be expressed in other terms. An adjacency list is generated by the second generation unit 13, for example, by listing the vertices (nodes) and line segments (edges) included in the above graph. In this embodiment, the three-dimensional structure data reconstructed by the second reconstruction unit 15 is also referred to as the rough structure data.

[0050] Figure 5 is a diagram showing an example of the relationship between an index for a one-dimensional sequence and the signal before reconstruction according to this embodiment, and is an example of the proximity data (Hi-C data) described above. The horizontal axis in the upper, middle, and lower panels of Figure 5 represents an index for the one-dimensional sequence. This index is, for example, a number assigned to each element when the object of structure estimation is a chromosome. The upper panel of Figure 5 shows the signal for the first element (first component) before reconstruction by the structure estimation device 1A. The middle panel of Figure 5 shows the signal for the second element (second component) before reconstruction by the structure estimation device 1A. The lower panel of Figure 5 shows the signal for the third element (third component) before reconstruction by the structure estimation device 1A. These three signals represent values ​​that have a positive correlation with the distance between each element to which an index is assigned and a specific element.

[0051] Figure 6 shows an example of a recurrence plot RP generated by the first generation unit 12 according to this embodiment, by referring to the proximity data (information on the structure of spatially adjacent parts, Hi-C data) shown in Figure 5. The horizontal and vertical axes in Figure 6 both represent indices related to the one-dimensional array. Furthermore, the recurrence plot RP shown in Figure 6 displays a point when an element to which an index shown on the horizontal axis is assigned and an element to which an index shown on the vertical axis is assigned are close to each other at a distance of less than or equal to a predetermined distance. The first generation unit 12 according to this embodiment may also be described as generating a recurrence plot RP from information on the structure of spatially adjacent parts of the molecule by treating the primary array of the molecule as a time axis, similar to Embodiment 1. Note that a recurrence plot is sometimes referred to as a contact map.

[0052] Figure 7 shows an example of a graph (more specifically, an undirected weighted graph) generated by the second generation unit 13 according to this embodiment by referring to the recurrence plot RP. The white circles in Figure 7 represent each element to which an index is assigned. The line segments in Figure 7 indicate that the distance between the two elements located at both ends of the line segment is less than or equal to the predetermined distance described above. In addition, the line segments in Figure 7 are assigned a distance D between the two elements located at both ends of the line segment. This distance D is calculated by the second generation unit 13 as an example.

[0053] As an example, the distance D is defined by the following formula (B1) which includes the total number of elements M included in the object, index i, and index j. M (i, j) may also be used. Here, the denominator on the right-hand side of equation (B1) represents the number of elements in the union of the elements whose points are displayed in column i and the elements whose points are displayed in column j of the recurrence plot shown in Figure 6. The numerator on the right-hand side of equation (B1) represents the number of elements in the set of elements that are not common to the elements whose points are displayed in column i and the elements whose points are displayed in column j of the recurrence plot shown in Figure 6. The distance calculated by equation (B1) is equal to the value obtained by subtracting the Jaccard coefficient from 1. The distance calculated by equation (B1) is sometimes called the Jaccard distance. Note that instead of such a distance, a value related to the Dice coefficient or Simpson coefficient may be used. Also, the distance calculated by equation (B1) corresponds to the Jaccard distance w(i,j) described in Embodiment 1. For this reason, in this embodiment as well, the distance d M (i, j) is sometimes written as the distance w(i, j). In other words, equation (B1) is a different notation for equation (A1) in Embodiment 1.

[0054] Incidentally, the specific calculation of the above distance D may be performed as follows. More specifically, the second generation unit 13 may be configured to calculate the distance after conversion by converting the above Jaccard distance. Here, the distance after the conversion is obtained, for example, as a solution or an approximate solution of an equation having the distance after the conversion as a variable. More specifically, the distance after the conversion is obtained, for example, as a solution or an approximate solution of a cubic equation including the Jaccard distance as a coefficient and having the distance after the conversion as a variable.

[0055] Further, the above conversion by the second generation unit 13 may be expressed as a conversion obtained by approximating the Jaccard distance with two overlapping three-dimensional spheres. More specifically, in this step, the distance after the conversion is a cubic equation including the Jaccard distance w(i, j) as a coefficient By solving, the distance y after the conversion may be accurately or approximately calculated as a function of the Jaccard distance w(i, j). Incidentally, as described above, instead of the Jaccard distance, a distance related to the Dice coefficient or the Simpson coefficient may be used. Even in the case of using the distance expressed by the Dice coefficient or the Simpson coefficient in this way, similar to the Jaccard distance, the sum of the recurrence plots RP may be approximated by two overlapping three-dimensional spheres.

[0056] The above conversion by the second generation unit 13 has the following meaning. That is, it corresponds to approximating the local distance (Jaccard distance w(i, j)) using the volumes of two overlapping three-dimensional spheres. More specifically, the Jaccard distance w(i, j) is approximated using the volumes of two overlapping three-dimensional spheres as and when y = r / d is set, the above formula B2 is obtained. The distance (corrected distance) y converted in this way may be used as the above distance D (new local distance). Incidentally, in the above example, the case of using a cubic equation has been described, but this does not limit the present embodiment, and a configuration using an equation of another degree may be used.

[0057] Then, the second generation unit 13 is the distance calculated as described above (d MA distance matrix is ​​generated by determining the length of the shortest path on the graph between each vertex based on (i, j), w(i, j), or y). In the same manner as in Embodiment 1, the Jacquard distance w(i, j) (or d M The distance matrix generated using (i, j) is also called the first distance matrix DM1, and the distance matrix generated using the transformed distance y is also called the second distance matrix DM2.

[0058] The first reconstruction unit 14 then applies multidimensional scaling to the distance matrix generated by the second generation unit 13 to reconstruct the time series data. The second reconstruction unit 15 then reconstructs the three-dimensional structure (3DD) of the subject from the time series data reconstructed by the first reconstruction unit 14. More specifically, the first reconstruction unit 14 and the second reconstruction unit 15 use multidimensional scaling on the aforementioned distance matrix and select a predetermined number of components corresponding to the largest eigenvalues ​​(for example, the top three) to obtain a rough three-dimensional structure (3D structure 3DD) of the subject.

[0059] Figure 8 shows an example of the relationship between an index for a one-dimensional sequence and the reconstructed signal according to this embodiment. The horizontal axis in the upper, middle, and lower panels of Figure 8 represents the index for the one-dimensional sequence. Here, in order to facilitate comparison between Figure 5 and Figure 8, the reconstruction results were rotated / inverted and translated so that the three elements in Figure 5 and the three elements in Figure 8 could be associated, and then Figure 8 was created. This index is, for example, a number assigned to each element when the object whose structure is to be estimated is a chromosome. The vertical axis in the upper panel of Figure 8 represents the signal corresponding to the signal shown in the upper panel of Figure 5. The vertical axis in the middle panel of Figure 8 represents the signal corresponding to the signal shown in the middle panel of Figure 5. The vertical axis in the lower panel of Figure 8 represents the signal corresponding to the signal shown in the lower panel of Figure 5.

[0060] The signals shown in the upper part of Figure 8, for example, are assigned indices and represent values ​​that have a positive correlation with the distance between the first element and a specific element related to the upper part of Figure 5. The line graph shown in the upper part of Figure 8 exhibits similar behavior to the line graph shown in the upper part of Figure 5. Similarly, the signals shown in the middle part of Figure 8 exhibit similar behavior to the line graph shown in the middle part of Figure 5. The signals shown in the lower part of Figure 8, for example, are assigned indices and represent values ​​that have a positive correlation with the distance between the first element and a specific element related to the lower part of Figure 5. The line graph shown in the lower part of Figure 8 exhibits similar behavior to the line graph shown in the lower part of Figure 5.

[0061] The process of generating the signal shown in Figure 8, starting from a recurrence plot, graph, or distance matrix, is also simply called reconstruction, and is performed, for example, by the reconstruction unit 18. Furthermore, reconstruction can be performed not only in the three-dimensional case described with reference to Figures 3 to 6, but also in cases of two dimensions or less, or four dimensions or more.

[0062] As described above, the three-dimensional structure data 3DD generated by the second reconstruction unit 15 is also referred to as rough structure data in this embodiment. The rough structure data represents the target structure at the lowest spatial resolution among the structure data representing the target structure. This resolution is expressed by the following formula (B4), which includes, for example, the number of times the resolution is increased L and the total number of points in space N at the spatial resolution shown by the second structure data from which the spatial resolution is finally generated. Here, in equation (B4), the number K represents the number of points in the spatial resolution of the initially generated structural data. Furthermore, equation (B4) shows that by doubling the number of points each time reconstruction is performed, the number of points in the spatial resolution of the generated structural data becomes N after increasing the resolution L times.

[0063] (Third reconstruction unit 16) The third reconstruction unit 16 generates first structural data that shows the target structure with a higher spatial resolution than the rough structural data, using neighboring points in the target structure shown by the rough structural data. In this case, it is preferable that the third reconstruction unit 16 generates second structural data that shows the target structure with a spatial resolution greater than one times and less than or equal to two times the spatial resolution of the target structure shown by the first structural data.

[0064] The third reconstruction unit 16 then generates second structure data, which shows the target structure with a higher spatial resolution than the first structure data, using neighboring points in the target structure shown by the first structure data. In this case, it is preferable that the third reconstruction unit 16 generates second structure data that shows the target structure with a spatial resolution greater than one times and less than or equal to two times the highest spatial resolution of the target structure shown by the already generated second structure data.

[0065] Furthermore, the third reconstruction unit 16 may generate new second structure data that shows the target structure with a higher spatial resolution than the already generated second structure data, using neighboring points of the target structure shown with the highest spatial resolution among the target structures shown by the already generated second structure data. In this case, it is preferable that the third reconstruction unit 16 generates second structure data that shows the target structure with a spatial resolution greater than one times and less than or equal to two times the highest spatial resolution among the target structures shown by the already generated second structure data.

[0066] When the third reconstruction unit 16 generates the first structural data or the second structural data, it calculates the distance D (in other words, d) using the formula (B1) described above. K (c, j)) may be used. Alternatively, when the third reconstruction unit 16 generates the first structural data or the second structural data, d may be used as w(i, j) in the above-described formula (B2). K The transformed distance y = y(d) calculated using (c, j) K (c, j) may be used as the distance D.

[0067] Here, the distance D is the distance when the spatial resolution of the target structure, as indicated by the already generated second structure data, is the highest spatial resolution. Furthermore, this distance D is the distance when the spatial resolution of the target structure, as indicated by the first structure data, is the nearest point M. c,l+1 This corresponds to the nearest point M when this distance D is the spatial resolution of the target structure indicated by the already generated second structure data. c,l+1 It supports this.

[0068] Then, the third reconstruction unit 16 calculates a weighted average a for each index c, which is expressed by the following formula (B5). l Calculate (c). In the following explanation, λ = 1, but this is not limited to this embodiment. Equation (B5) is given by distance d K (c, m i,l|l+1 (c) and the weighted average a calculated when generating the previous second structure data. l+1 (m i,l+1 (c)) Nearest point m i,l+1 This is the formula for calculating the weighted average with respect to (c). Also, the nearest point m included in formula (B5) i,l|l+1 (c) shows the spatial resolution of the target structure as shown by the second structure data generated in this reconstruction, with respect to the nearest point m. i,l+1 This is a neighboring point corresponding to (c). Note that in equation B5, the distance d K (c, m i,l|l+1 (c)) Instead, distance d K The distance after transformation based on y = y(d K (c, m i,l|l+1 (c)) may also be used. More specifically, distance d K (c, m i,l|l+1 (c)) may be replaced with w(i, j) and the transformed distance y derived according to equation B2 may be used.

[0069] First, {a L Let (c) be the set of points in the first structure data defined from the rough structure data. Then, equation (B5), or equation (B5), the distance d KUsing the result of replacing with the transformed distance y, the set of points in the second structure data {a L-1 (c)} is obtained. Thereafter, it is preferable that the third reconstruction unit 16 repeats the process of generating second structure data until the spatial resolution of the target structure indicated by the second structure data exceeds a predetermined resolution. That is, {a l+1 Let (c) be the set representing each point of the first structure data, then equation (B5), or equation (B5) in which the distance d K Using the transformed distance y, we represent each point in the second structure data with the set {a l The calculation to obtain (c) is repeated by decreasing l by 1 from L-1 until it becomes 0. The predetermined resolution referred to here is, for example, 1 kilobase (kb) to 2 kilobases. 1 kilobase is a unit of spatial resolution whose smallest unit is the size of 1000 bases that make up the deoxyribonucleic acid contained in the chromosome. For example, the third reconstruction unit 16 repeats the process of generating second structure data until the index l goes from L-1 to 0, thereby ultimately generating second structure data that shows the target structure with a spatial resolution exceeding the predetermined resolution.

[0070] Furthermore, the aforementioned neighbor points are defined by a matrix obtained by replacing the diagonal elements of the recurrence plot RP and the elements in rows or columns up to a predetermined number of times from the diagonal elements with 1. For example, neighbor points are defined by a matrix obtained by replacing the diagonal elements (i, i) (i=1, 2, ..., M), the elements one column away from the diagonal elements (i, i+1), and the elements one row away from the diagonal elements (i+1, i) with 1.

[0071] (Structural data output unit 17) The structural data output unit 17 outputs the second structural data. For example, the structural data output unit 17 displays the contents indicated by the second structural data on the display provided by the input / output unit 40.

[0072] (Processing flow by structural estimation device 1A) Next, the processing flow by structural estimation device 1A will be explained with reference to Figure 9. Figure 9 is a flowchart showing the processing flow (structural estimation method S1A) by structural estimation device 1A.

[0073] (Step S101) In step S101, the acquisition unit 11 acquires proximity data (Hi-C data) of the target (molecule).

[0074] (Step S102) Next, in step S102, the first generation unit 12, the second generation unit 13, the first reconstruction unit 14, and the second reconstruction unit 15 in the reconstruction unit 18 reconstruct the proximity data acquired in step S101 to generate rough structure data that shows the target structure.

[0075] (Step S103) Next, in step S103, the third reconstruction unit 16 generates first structural data using neighboring points in the target structure indicated by the rough structural data.

[0076] (Step S104) Next, in step S104, the third reconstruction unit 16 generates second structure data using neighboring points in the target structure indicated by the first structure data or the second structure data. Specifically, in the first step S104, the third reconstruction unit 16 generates second structure data using neighboring points in the target structure indicated by the first structure data, and in subsequent steps S104, it generates second structure data using neighboring points in the target structure indicated by the second structure data.

[0077] (Step S105) Next, in step S105, the structural data output unit 17 outputs the second structural data.

[0078] (Effects of the structure estimation device 1A) As described above, the structure estimation device 1A according to this embodiment includes: an acquisition unit 11 that acquires information on the structure of spatially adjacent parts of a molecule; a first generation unit 12 that generates a recurrence plot from the information on the structure of spatially adjacent parts of a molecule by treating the linear arrangement of the molecule as a time axis; a second generation unit 13 that generates a distance matrix that gives the distance between any two vertices of the recurrence plot; a first reconstruction unit 14 that reconstructs time series data using multidimensional scaling on the distance matrix; and a second reconstruction unit 15 that reconstructs the three-dimensional structure of the molecule from the reconstructed time series data. The distance matrix is ​​a matrix whose components are the transformed distance (y = y(w(i,j))) obtained by transforming the Jacquard distance relating to the recurrence plot, and the transformed distance (y = y(w(i,j))) is obtained as the solution or approximate solution of an equation with the transformed distance as the variable.

[0079] With the structure estimation device 1A configured in this way, the transformed distance y is calculated as a solution or approximate solution to an equation in which the transformed distance is a variable, and the three-dimensional structure of the molecule can be reconstructed with high accuracy by using the transformed distance y to reconstruct the three-dimensional structure of the molecule.

[0080] Furthermore, in the structure estimation method S1 described above, as stated above, the transformation obtained by approximating the sum of the recurrence plots with two overlapping three-dimensional spheres is used as the distance transformation. More specifically, in this step, the transformed distance y is obtained by solving a cubic equation (equation B2) in which the Jacquard distance w(i,j) is a coefficient. According to the inventor's findings, by approximating the sum of the recurrence plots with two overlapping three-dimensional spheres in this way, the three-dimensional structure of the molecule can be reconstructed with greater accuracy.

[0081] Furthermore, as described above, the structure estimation device 1A according to this embodiment includes a third reconstruction unit 16, which performs the following processes: - Generates first structure data showing the target structure with a higher spatial resolution than the rough structure data, using neighboring points in the structure of the target (molecule) shown by the rough structure data; and - Generates second structure data showing the target structure with a higher spatial resolution than the first structure data, using neighboring points in the structure of the target (molecule) shown by the first structure data. As a result, the structure estimation device 1A can calculate the structure of chromosomes, etc., with a higher spatial resolution while reducing the computational load for calculating the structure of chromosomes, etc.

[0082] Furthermore, as described above, in the structure estimation device 1A according to this embodiment, in at least one of the processes for generating the first structure data and the process for generating the second structure data, the Jacquard distance (d) with respect to the recurrence plot is calculated. K The transformed distance (y) obtained by transforming (, w), which includes the Jacquard distance as a coefficient and is obtained as the solution to a cubic equation with the transformed distance as the variable, may be used. This makes it possible to reconstruct the three-dimensional structure of the molecule (for example, a chromosome) with even greater accuracy.

[0083] Furthermore, in the process of generating the second structure data, the third reconstruction unit 16 may use neighboring points of the molecular structure shown with the highest spatial resolution among the molecular structures shown by the already generated second structure data to generate second structure data showing the molecular structure with a higher spatial resolution than the already generated second structure data. As a result, the structure estimation device 1A can calculate the structure of chromosomes, etc., with an even higher spatial resolution.

[0084] Furthermore, the third reconstruction unit 16 may repeat the process of generating the second structure data until the spatial resolution of the molecular structure shown by the second structure data exceeds a predetermined resolution. This allows the structure estimation device 1A to calculate the structure of chromosomes, etc., at a desired resolution.

[0085] Furthermore, the structure estimation device 1A may generate second structure data that shows the target structure at a spatial resolution greater than one times but less than or equal to two times the highest spatial resolution of the target structure shown by the already generated second structure data. Alternatively, the structure estimation device 1A may generate second structure data that shows the target structure at a spatial resolution greater than one times but less than or equal to two times the spatial resolution of the target structure shown by the first structure data. This allows the structure estimation device 1A to avoid a situation where points after reconstruction coincide with points before reconstruction.

[0086] In this embodiment, the example of the structure estimation device 1A estimating the structure of a chromosome was used for explanation, but the invention is not limited to this. The structure estimation device 1A may also estimate the structure of, for example, a polymer forming a resin, or a protein having a complexly folded string-like structure.

[0087] [Example 1] Below, an example of the structural estimation device 1 according to Embodiment 1 will be described. Figure 10 is a graph showing an example of the structural estimation device 1 according to Embodiment 1. In this example, the structural estimation device 1 was applied to a toy model of Brownian motion. More specifically, it was applied to Brownian motion with a length of 2000.

[0088] Furthermore, in this embodiment, the true three-dimensional structure data (neighboring information) was defined using the information below the 20th percentile of all the distances between every pair of vertices involved in the Brownian motion. Then, 99.9% of all vertex pairs were rejected, and the three-dimensional structure data of the Brownian motion was reconstructed using only the remaining 0.1% of the data (i.e., used as input data). In this embodiment, the same simulation was performed 20 times.

[0089] In Graph A of Figure 10, "Original" shows the 3D correlation coefficient between the 3D structure data reconstructed using the Jacquard distance w(i,j) before transformation as the distance D component of the aforementioned distance matrix (in other words, the 3D structure data reconstructed using the method of Hirata et al. (2016) (Non-Patent Literature 13)) and the true 3D structure data. Also in Graph A of Figure 10, "Proposed" shows the 3D correlation coefficient between the 3D structure data reconstructed using the transformed distance y = y(w(i,j)) as the distance D component of the aforementioned distance matrix and the true 3D structure data.

[0090] Furthermore, Graph B in Figure 10 shows the difference between the 3D correlation coefficient of the 3D structure data reconstructed using the transformed distance y = y(w(i,j)) as distance D and the 3D correlation coefficient of the 3D structure data reconstructed using the original Jacquard distance w(i,j) as distance D: (3D correlation coefficient for proposed) - (that for original).

[0091] As can be seen from Figure 10, by using the transformed distance y = y(w(i,j)) as the distance D, the three-dimensional structure of the object can be reconstructed with greater accuracy compared to when using the original distance w(i,j).

[0092] The 3D correlation coefficient mentioned above is an index used to compare one 3D object with another. In calculating this index, the two 3D objects have the same number of vertices that correspond one-to-one with each other. First, a distance matrix is ​​calculated for the set of vertices of each object. Next, the correlation coefficient between these two distance matrices is calculated to obtain the 3D correlation coefficient. The 3D correlation coefficient can take values ​​from -1 to 1, and the closer the two objects are to 1, the more similar they are.

[0093] [Example 2] Below, an example using the structural estimation device 1A according to Embodiment 2 will be described. Figure 11 is a graph showing an example using the structural estimation device 1A according to Embodiment 2. In this example, the structural estimation device 1A was applied to retinal cells from Tan et al. (2019) (Non-Patent Literature 14). More specifically, it was applied to 56 retinal cells from Tan et al. (2019) (Non-Patent Literature 14).

[0094] In this example, to eliminate the possibility of adverse effects from sex chromosomes, the analysis was performed after removing the sex chromosomes. Furthermore, in this example, the rough structure data was reconstructed at a resolution of 1024kb, and then reconstructed at a resolution of 256kb through two high-resolution processes (generation of first and second structure data).

[0095] In Graph A of Figure 11, "Original" represents the MaxICM (maximum integrated correctness measure) of the 3D structure data reconstructed by the reconstruction unit 18 using the Jacquard distance w(i,j) before transformation as distance D (in other words, the 3D structure data reconstructed using the method of Hirata et al. (2021) (Non-Patent Literature 15)). In Graph A of Figure 11, "Proposed" represents the MaxICM of the 3D structure data reconstructed by the reconstruction unit 18 using the transformed distance y = y(w(i,j)) as distance D.

[0096] Furthermore, Graph B in Figure 11 shows the difference between the MaxICM of the 3D structure data reconstructed using the transformed distance y = y(w(i,j)) as distance D, and the MaxICM of the 3D structure data reconstructed using the original Jacquard distance w(i,j) as distance D: (MaxICM for proposed) - (MaxICM for original).

[0097] As can be seen from Figure 11, by using the transformed distance y = y(w(i,j)) as the distance D, the three-dimensional structure of the object can be reconstructed with higher consistency compared to when using the original distance w(i,j).

[0098] In calculating the MaxICM (maximum integrated correctness measure) as described above, the ratios of "phased pair," "half-phased pair," and "unphased pair" are first defined. These ratios are defined so that they fall within the threshold for reconstruction. Here, "phased" refers to the case where both chromosome fragments contain a single nucleotide polymorphism, "half-phased" refers to the case where only one chromosome fragment contains the single nucleotide polymorphism, and "unphased" refers to the case where neither chromosome fragment contains the single nucleotide polymorphism. Then, the average of these three ratios is taken. Furthermore, the ratio of points further from the center of the chromosome than the threshold is also used.

[0099] The ICM (integrated correctness measure) is given by the product of (the mean of the three ratios of the pairs within a threshold distance) and (the ratio of points which are farther than the threshold distance from the center of the chromosome). The MaxICM (maximum integrated correctness measure) is obtained by taking the maximum value of the ICM that exceeds the threshold distance. The MaxICM can take values ​​from -1 to 1, and the closer the reconstruction is to perfect, the closer the value is to 1.

[0100] [Example 3] Next, another example of the structural estimation device 1A according to Embodiment 2 will be described. Figures 12 to 14 are graphs showing an example of the structural estimation device 1A according to Embodiment 2. In this example, the structural estimation device 1A was applied to retinal cells and olfactory sensory neuron (OSN) cells.

[0101] In Figures 12 and 13, "Original" for Retina indicates the difference between the LAD density (density of lamina associated domains: density of the portion corresponding to the nuclear membrane) at the outer edge of the Retina, as shown by the 3D structural data reconstructed using the Jacquard distance w(i,j) before conversion as distance D in processing by the reconstruction unit 18, and the LAD density inside the Retina: (LAD density for periphery) - (LAD density for interior) (Figure 12) (NPCLAD density at periphery) - (NPCLAD density at interior) (Figure 13). Also, in Figures 12 and 13, "Proposed" for Retina indicates the difference between the LAD density at the outer edge of the Retina, as shown by the 3D structural data reconstructed using the converted distance y = y(w(i,j)) as distance D in processing by the reconstruction unit 18, and the LAD density inside the Retina. The same applies to "Original" and "Proposed" for OSN in Figures 12 and 13. Figure 12 is a graph where the portion within the detection limit of single-cell Hi-C from the center of the cell nucleus is considered the inside of the cell nucleus, and Figure 13 is a graph where the portion of the chromosome fragment whose distance from the center of the cell nucleus is smaller than the average distance of the chromosome fragment from the center of the cell nucleus is considered the inside of the cell nucleus.

[0102] As can be seen from Figures 12 and 13, by using the transformed distance y = y(w(i,j)), the inverse structure (inside-out structure) in the retina can be reconstructed more appropriately. On the other hand, Figure 14 shows a graph of the case in this embodiment where three retinal nerve rods and three MOE (Main Olfactory Epithelium) nerve OSNs were selected, and the reconstruction unit 18 was used to reconstruct a rough three-dimensional structure at 2048kb initially, followed by 10 calculations to improve the accuracy of the three-dimensional structure, and finally to reconstruct the three-dimensional structure of the chromosome at a resolution of 2kb. More specifically, the upper part of Figure 14 is a graph showing the Maximum Integrated Correctness Measure at each resolution. Furthermore, the lower panel of Figure 14 is a graph showing the difference between the LAD density (density of lamina associated domains: density of the region corresponding to the nuclear membrane) at the outer edge of the retina and the LAD density inside the retina, as shown by the 3D structural data, at each resolution: (NPCLAD density at periphery) - (NPCLAD density at interior). In both the upper and lower panels of Figure 14, the points move from right to left as the resolution of the 3D structure increases. The graph in the upper panel of Figure 14 shows how the Maximum Integrated Correctness Measure, which measures the accuracy of the 3D structure of chromosomes, improves with increasing resolution. Furthermore, the lower graph in Figure 14, similar to Figure 13, shows that in MOE OSN neurons, the NPCLAD is biased to the outside of the cell nucleus, as is typical in neurons. However, in retina neuronal rods, the part that would normally be on the outside of the cell nucleus in neurons is biased to the inside of the cell nucleus, creating a shape that allows light from the outside to be perceived (inside-out structure). This is more clearly shown with higher resolution.

[0103] [Example of implementation using software] The functions of the structural estimation devices 1 and 1A (hereinafter referred to as "devices") can be realized by a program that causes a computer to function as the device, and by a program that causes a computer to function as each control block of the device (especially each part included in the control units 10 and 10A).

[0104] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., memory) as hardware for executing the program. By executing the program using this control device and storage device, the functions described in each of the embodiments are realized.

[0105] The above program may be recorded on one or more computer-readable recording media, not temporary ones. These recording media may or may not be provided by the above device. In the latter case, the program may be supplied to the above device via any wired or wireless transmission medium.

[0106] Furthermore, some or all of the functions of each of the above control blocks can also be realized by logic circuits. For example, an integrated circuit in which logic circuits functioning as each of the above control blocks are formed is also included in the scope of the present invention. In addition, it is also possible to realize the functions of each of the above control blocks by, for example, a quantum computer.

[0107] (Summary) This specification includes at least the following components:

[0108] (Configuration A1) A structure estimation method comprising the steps of: generating a recurrence plot from information on the structure of spatially adjacent parts of a molecule by treating the linear arrangement of the molecule as a time axis; generating a distance matrix that gives the distance between any two vertices of the recurrence plot; reconstructing time series data using multidimensional scaling on the distance matrix; and reconstructing the three-dimensional structure of the molecule from the reconstructed time series data, wherein the distance matrix is ​​a matrix whose components are transformed distances obtained by transforming the Jacquard distances relating to the recurrence plot, and the transformed distances are obtained as the solution or approximate solution of an equation with the transformed distances as variables.

[0109] (Configuration A2) The structure estimation method according to Configuration A1, wherein the transformation is obtained by approximating the Jacquard distance with two overlapping three-dimensional spheres.

[0110] (Configuration A3) The distance y after the transformation is a cubic equation (w(i,j)-2)y, which includes the Jacquard distance w(i,j) as a coefficient. 3 A structural estimation method described in configuration A1 or A2, which is a solution or approximate solution to + (24 - 12w(i,j))y - 16w(i,j) = 0.

[0111] (Configuration A4) A structure estimation method according to any one of Configurations A1 to A3, further comprising the steps of: using the reconstructed three-dimensional structure as rough structure data, generating first structure data that shows the structure of the molecule with a higher spatial resolution than the rough structure data, using neighboring points in the structure of the molecule shown by the rough structure data; and generating second structure data that shows the structure of the molecule with a higher spatial resolution than the first structure data, using neighboring points in the structure of the molecule shown by the first structure data.

[0112] (Configuration A5) The structure estimation method according to Configuration A4, wherein in at least one of the steps of generating the first structure data and generating the second structure data, the transformed distance obtained by transforming the Jacquard distance relating to the recurrence plot is used, and the transformed distance is obtained as the solution to a cubic equation in which the Jacquard distance is included as a coefficient and the transformed distance is a variable.

[0113] (Configuration A6) The structure estimation method according to Configuration A5 or A6, wherein in the step of generating the second structure data, neighboring points of the molecular structure shown with the highest spatial resolution among the molecular structures shown by the already generated second structure data are used to generate the second structure data showing the molecular structure with a higher spatial resolution than the already generated second structure data.

[0114] (Configuration A7) The structure estimation method according to Configuration A6, wherein the step of generating the second structure data is repeated until the spatial resolution of the molecular structure shown by the second structure data exceeds a predetermined resolution.

[0115] (Configuration A8) A structure estimation device comprising: an acquisition unit that acquires information on the structure of spatially adjacent parts of a molecule; a first generation unit that generates a recurrence plot from the information on the structure of spatially adjacent parts of a molecule by treating the linear arrangement of the molecule as a time axis; a second generation unit that generates a distance matrix that gives the distance between any two vertices of the recurrence plot; a first reconstruction unit that reconstructs time series data using multidimensional scaling on the distance matrix; and a second reconstruction unit that reconstructs the three-dimensional structure of the molecule from the reconstructed time series data, wherein the distance matrix is ​​a matrix whose components are transformed distances obtained by transforming the Jacquard distances relating to the recurrence plot, and the transformed distances are obtained as the solution or approximate solution of an equation with the transformed distances as variables.

[0116] (Configuration A9) A program that causes a computer to function as a structure estimation device, wherein the computer functions as: an acquisition unit that acquires information on the structure of spatially adjacent parts of a molecule; a first generation unit that generates a recurrence plot from the information on the structure of spatially adjacent parts of a molecule by treating the linear arrangement of the molecule as a time axis; a second generation unit that generates a distance matrix that gives the distance between any two vertices of the recurrence plot; a first reconstruction unit that reconstructs time series data using multidimensional scaling on the distance matrix; and a second reconstruction unit that reconstructs the three-dimensional structure of the molecule from the reconstructed time series data, wherein the distance matrix is ​​a matrix whose components are transformed distances obtained by transforming the Jacquard distances relating to the recurrence plot, and the transformed distances are obtained as a solution or approximate solution of an equation with the transformed distances as variables.

[0117] The present invention is not limited to the embodiments described above, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.

[0118] 1, 1A...Structural estimation device 11...Acquisition unit 12...First generation unit 13...Second generation unit 14...First reconstruction unit 15...Second reconstruction unit 16...Third reconstruction unit 17...Structural data output unit

Claims

1. A structure estimation method comprising the steps of: generating a recurrence plot from information on the structure of spatially adjacent parts of a molecule by treating the linear arrangement of the molecule as a time axis; generating a distance matrix that gives the distance between any two vertices of the recurrence plot; reconstructing time series data using multidimensional scaling on the distance matrix; and reconstructing the three-dimensional structure of the molecule from the reconstructed time series data, wherein the distance matrix is ​​a matrix whose components are transformed distances obtained by transforming the Jacquard distances relating to the recurrence plot, and the transformed distances are obtained as the solution or approximate solution of an equation with the transformed distances as variables.

2. The structural estimation method according to claim 1, wherein the transformation is obtained by approximating the Jacquard distance with two overlapping three-dimensional spheres.

3. The distance y after the transformation is a cubic equation (w(i,j)-2)y, which includes the Jacquard distance w(i,j) as a coefficient. 3 The structural estimation method according to claim 2, which is a solution or approximate solution to + (24 - 12w(i,j))y - 16w(i,j) = 0.

4. A method for estimating a structure according to any one of claims 1 to 3, further comprising the steps of: using the reconstructed three-dimensional structure as rough structure data, generating first structure data that shows the structure of the molecule with a higher spatial resolution than the rough structure data, using neighboring points in the structure of the molecule shown by the rough structure data; and generating second structure data that shows the structure of the molecule with a higher spatial resolution than the first structure data, using neighboring points in the structure of the molecule shown by the first structure data.

5. The structure estimation method according to claim 4, wherein in at least one of the steps of generating the first structure data and generating the second structure data, a transformed distance obtained by transforming the Jacquard distance relating to the recurrence plot is used, the transformed distance obtained as the solution to a cubic equation in which the Jacquard distance is included as a coefficient and the transformed distance is a variable.

6. The structure estimation method according to claim 4, wherein in the step of generating the second structure data, neighboring points of the molecular structure shown with the highest spatial resolution among the molecular structures shown by the already generated second structure data are used to generate the second structure data showing the molecular structure with a higher spatial resolution than the already generated second structure data.

7. The method for estimating a structure according to claim 6, wherein the step of generating the second structure data is repeated until the spatial resolution of the molecular structure shown by the second structure data exceeds a predetermined resolution.

8. A structure estimation device comprising: an acquisition unit that acquires information on the structure of spatially adjacent parts of a molecule; a first generation unit that generates a recurrence plot from the information on the structure of spatially adjacent parts of the molecule by treating the linear arrangement of the molecule as a time axis; a second generation unit that generates a distance matrix that gives the distance between any two vertices of the recurrence plot; a first reconstruction unit that reconstructs time series data using multidimensional scaling on the distance matrix; and a second reconstruction unit that reconstructs the three-dimensional structure of the molecule from the reconstructed time series data, wherein the distance matrix is ​​a matrix whose components are transformed distances obtained by transforming the Jacquard distances relating to the recurrence plot, and the transformed distances are obtained as a solution or approximate solution of an equation with the transformed distances as variables.

9. A program that causes a computer to function as a structure estimation device, wherein the computer functions as: an acquisition unit that acquires information on the structure of spatially adjacent parts of a molecule; a first generation unit that generates a recurrence plot from the information on the structure of spatially adjacent parts of a molecule by treating the linear arrangement of the molecule as a time axis; a second generation unit that generates a distance matrix that gives the distance between any two vertices of the recurrence plot; a first reconstruction unit that reconstructs time series data using multidimensional scaling on the distance matrix; and a second reconstruction unit that reconstructs the three-dimensional structure of the molecule from the reconstructed time series data, wherein the distance matrix is ​​a matrix whose components are transformed distances obtained by transforming the Jacquard distances relating to the recurrence plot, and the transformed distances are obtained as a solution or approximate solution to an equation with the transformed distances as variables.

Citation Information

Patent Citations

  • Network abnormal traffic refined detection method

    CN117527446A

  • Reconstruction method of three-dimensional structure of biomolecular data using concept of spatial closeness

    JP2017142633A

  • Adaptive Co-Distillation Model

    JP2023513613A