Accuracy prediction method and information processing apparatus

The accuracy prediction program uses regression equations to systematically select high-accuracy division candidates for DMET, enhancing the precision and efficiency of potential energy calculations in quantum chemical computations.

JP2026011009APending Publication Date: 2026-01-23FUJITSU LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024111239
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-10
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

The accuracy of potential energy calculations in quantum chemical calculations varies significantly based on the method of subset division in DMET, making it impractical to search for the best division pattern among countless possibilities.

Method used

An accuracy prediction program generates a regression equation using metrics from known division patterns to predict the accuracy of potential energy calculations for large molecules, allowing for high-accuracy division candidate selection.

Benefits of technology

This approach enables accurate prediction of division candidates, improving the accuracy of potential energy calculations and reducing computational time and load by narrowing down candidates systematically.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026011009000001_ABST
    Figure 2026011009000001_ABST
Patent Text Reader

Abstract

To predict a division candidate with high calculation accuracy when obtaining potential energy.SOLUTION: The information processing apparatus generates a regression equation for predicting calculation accuracy of the potential energy of the first molecule by using each of a plurality of dividing patterns including a plurality of subsets each including one or more atoms included in the first molecule. The information processing apparatus applies the regression equation to a plurality of dividing candidate patterns including a plurality of subsets each including one or more atoms included in the second molecule, and predicts calculation accuracy of the potential energy of the second molecule when each of the plurality of dividing candidates is used.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an accuracy prediction program, an accuracy prediction method, and an information processing device. [Background technology]

[0002] The properties of a molecule can be understood by determining its energy. For example, the stable state of the molecular structure can be determined from the ground state energy of the molecule, and the unstable state of the molecular structure can be determined from the excited energy. Understanding such molecular properties is useful for drug discovery and the discovery of new materials, so quantum chemical calculations are very important. Examples of quantum chemical calculations include the classical algorithm CCSD(T) (Coupled-Cluster Singles-and-Doubles(-and-Triple)) and the quantum algorithm VQE (Variational Quantum Eigensolver), which is designed to run on quantum computers.

[0003] The computational complexity of CCSD(T) is O(n 7 ) is required, but on current computers, 1 From 10 2 For VQE, a similar amount of calculation is expected when using a simulator, and even when using a Noisy Intermediate-Scale Quantum Computer (NISQ), the amount of calculation is expected to increase in polynomial time as the number of orbitals increases. For these reasons, at present, it is not realistic to apply an algorithm to the entire large molecule and calculate the potential energy.

[0004] On the other hand, a method called DMET (Density Matrix Embedding Theory) is known, which divides the atoms contained in a molecule into several subsets, calculates the energies of these subsets individually, and then combines them to calculate the overall potential energy. For example, when calculating the energy of an alanine molecule, DMET divides the atoms contained in alanine into subsets, calculates the energy of each subset after abstracting the interactions with other subsets, and then combines them, thereby reducing the computational effort (problem size). As mentioned above, the algorithm for calculating potential energy requires a very large amount of computation, so using DMET can significantly reduce the computational time. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] US Patent Application Publication No. 2018 / 0096085 Summary of the Invention [Problem to be solved by the invention]

[0006] However, in the above-described subset division technique, the accuracy of the energy calculated for each subset varies depending on the division method, and the accuracy of the potential energy calculation may deteriorate.

[0007] For example, there are countless patterns for dividing a molecule, even if constraints are placed on the problem scale of each subset (e.g., the number of orbitals), and the accuracy of the final energy calculation varies depending on the division method, so it is not realistic to randomly search for the best division from countless candidates.

[0008] In one aspect, an object of the present invention is to provide an accuracy prediction program, an accuracy prediction method, and an information processing device that can predict division candidates with high calculation accuracy when calculating potential energy. [Means for solving the problem]

[0009] In a first proposal, the accuracy prediction program causes a computer to execute the following process: generate a regression equation for predicting the calculation accuracy of the potential energy of a first molecule using a plurality of division patterns each including a plurality of subsets each including one or more of each atom contained in the first molecule; apply the regression equation to a plurality of candidate division patterns each including a plurality of subsets each including one or more of each atom contained in a second molecule; and predict the calculation accuracy of the potential energy of the second molecule when each of the plurality of candidate division patterns is used. [Effects of the Invention]

[0010] According to one embodiment, it is possible to predict division candidates with high calculation accuracy when calculating potential energy. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a diagram illustrating an information processing apparatus according to a first embodiment. [Figure 2] FIG. 2 is a diagram illustrating the molecular division pattern. [Figure 3] FIG. 3 is a functional block diagram of the information processing apparatus according to the first embodiment. [Figure 4] FIG. 4 is a diagram illustrating a data structure used in the first embodiment. [Figure 5] FIG. 5 is a diagram illustrating a data structure used in the first embodiment. [Figure 6] FIG. 6 is a diagram illustrating a data structure used in the first embodiment. [Figure 7] FIG. 7 is a flowchart showing the flow of the process of deriving the regression equation. [Figure 8] FIG. 8 is a flowchart showing the flow of the division candidate presentation process. [Figure 9] FIG. 9 is a diagram illustrating a specific example of a division pattern. [Figure 10] FIG. 10 is a diagram for explaining a specific example of energy accuracy calculation. [Figure 11] FIG. 11 is a diagram for explaining a specific example of regression analysis. [Figure 12] FIG. 12 is a diagram for explaining a specific example of division candidates for a molecule to be calculated. [Figure 13] FIG. 13 is a diagram illustrating a specific example of calculation of estimated energy of a division candidate. [Figure 14] FIG. 14 is a diagram illustrating a specific example of ranking division candidates. [Figure 15] FIG. 15 is a diagram illustrating an example of a screen displaying division candidates. [Figure 16] FIG. 16 is a diagram illustrating an example of a hardware configuration. DETAILED DESCRIPTION OF THE INVENTION

[0012] The following describes in detail embodiments of the accuracy prediction program, accuracy prediction method, and information processing device disclosed herein with reference to the accompanying drawings. Note that the present invention is not limited to these embodiments. Furthermore, the embodiments can be combined as appropriate within a consistent range. [Example]

[0013] (Description of information processing device) Fig. 1 is a diagram illustrating an information processing device 10 according to Example 1. The information processing device 10 shown in Fig. 1 is an example of a computer that divides a group of atoms included in a molecule into several subsets using a theory called DMET, calculates the energy of each subset individually, and then combines the subsets to calculate the potential energy (hereinafter, may be simply referred to as "energy") of the whole (molecule).

[0014] Although DMET is expected to significantly reduce the amount of calculation required for the potential energy of molecules, the accuracy of the final calculated potential energy varies depending on the method of division. Since there are many patterns for dividing molecules, even when restrictions such as the number of orbitals are imposed, it is necessary to explore what kind of division is best.

[0015] Figure 2 is a diagram illustrating molecular division patterns. Figure 2 shows three division patterns (a), (b), and (c) for alanine. Calculating the potential energy using each of the countless possible division patterns would take a significant amount of time and is not realistic. On the other hand, it is also possible to randomly narrow down the countless possible division patterns to the three division patterns (a), (b), and (c) and calculate the accuracy of the potential energy using the narrowed-down division patterns. However, since there are no criteria for narrowing down the candidates, the accuracy of the potential energy may decrease as a result of narrowing down, and therefore, the random narrowing-down method is not a realistic method.

[0016] Therefore, the information processing device 10 according to the first embodiment generates a regression equation for predicting the calculation accuracy of the potential energy of the first molecule using each of a plurality of division patterns including a plurality of subsets each including one or more of each atom included in the first molecule. Subsequently, the information processing device 10 applies the regression equation to a plurality of candidate division patterns (hereinafter sometimes referred to as division candidates) each including a plurality of subsets each including one or more of each atom included in the second molecule, and predicts the calculation accuracy of the potential energy of the second molecule when each of the plurality of candidate division patterns is used.

[0017] That is, when calculating the energy of a large molecule using DMET, the information processing device 10 uses a molecule (first molecule) large enough to calculate the entire energy, and derives a regression equation that predicts calculation accuracy based on the number of orbitals, the number of electrons, etc. of each subset. Then, the information processing device 10 divides the large molecule (second molecule) for which energy is to be calculated so that each subset fits within a calculable size, generating multiple division candidates. Thereafter, the information processing device 10 applies the regression equation for calculating accuracy prediction to each of the multiple division candidates generated, ranks them, and presents them to the user.

[0018] 1, the information processing device 10 divides molecules of a calculable size into division patterns 1 to n (n is a natural number). Next, for division patterns 1 to n, the information processing device 10 collects metrics 1 to n, which are used as evaluation indexes for the model in regression analysis and are also used as measurement standards, and calculates a regression equation using these metrics.

[0019] The information processing device 10 then generates division candidates 1 through n by dividing the target molecule for which energy is to be calculated under the same constraints as when dividing a molecule of a calculable size. The information processing device 10 then collects metrics used in generating a regression equation (regression analysis) for each of the division candidates 1 through n, and applies the collected metrics to the regression equation to predict the accuracy of the energy calculation. The information processing device 10 then calculates the potential energy of the target molecule by executing DMET or the like using, for example, the division candidate 2 with the highest accuracy.

[0020] In this way, the information processing device 10 can predict division candidates with high calculation accuracy when calculating potential energy by applying the regression equation generated using molecules for which accurate potential energy can be calculated to molecules of large size that are the object of calculation.

[0021] (Functional configuration) 3 is a functional block diagram illustrating a functional configuration of the information processing device 10 according to Example 1. As shown in FIG.

[0022] The communication unit 11 is a processing unit that controls communication with other devices, and is realized by, for example, a communication interface. For example, the communication unit 11 receives input of a calculation target molecule for which energy is to be calculated from a user terminal. The communication unit 11 can also transmit various information calculated by the control unit 20 to the user terminal.

[0023] The display unit 12 is a processing unit that displays and outputs various types of information, and is realized by, for example, a display, a touch panel, etc. For example, the display unit 12 displays and outputs various types of information received by the communication unit 11 and various types of information calculated by the control unit 20.

[0024] The storage unit 13 is a processing unit that stores various data and programs executed by the control unit 20, and is realized by, for example, a memory, a hard disk, etc. For example, the storage unit 13 stores a data structure DB 14 that includes data used by the control unit 20 for various processes.

[0025] Here, we will explain the data structures of various information stored in the data structure DB 14. Figures 4, 5, and 6 are diagrams for explaining the data structures used in Example 1. As shown in Figure 4, the data structure DB 14 includes the data structures of an atom list, interatomic bond information, orbital number restrictions, and orbital number table.

[0026] Specifically, the "atom list" is a list of atoms contained in a molecule, expressed as (id, atom type). For example, (0, 'O') indicates that the atom with id=0 is "O." The "atomic bond information" is information about bonds between atoms contained in a molecule, expressed as a symmetric matrix indicating that when the value of number j in row number i is n, there are n multiple bonds from the i-th atom to the j-th atom. The "orbital number limit" is a threshold (upper limit) for the number of orbitals expressed as an integer data type (int), and is set to, for example, 8. The "orbital number table" defines the number of orbitals for each atom, expressed as "atom type → number of orbitals." For example, in the case of "H → 2," it is defined that the number of orbitals for hydrogen "H" is "2."

[0027] Next, as shown in FIG. 5, the data structure DB14 includes data structures of "metric values ​​and accuracy" for the division patterns obtained by dividing the molecules that can be calculated, and "metric values" for the division candidates obtained by dividing the molecules to be calculated.

[0028] "Metric Value and Accuracy" is information that associates the following: accuracy, maximum number of orbitals, minimum number of orbitals, orbital number variance, sum of squares of electron number differences, number of subsets, and bus orbital energy. Here, "accuracy" is the difference between the accurate energy value calculated by CCSD(T) and the energy value calculated by DMET or other methods using the subsets included in the division pattern. "Maximum number of orbitals" is the maximum number of orbitals in the subsets included in the division pattern, including spin. "Minimum number of orbitals" is the minimum number of orbitals in the subsets included in the division pattern, including spin. "Orbital number variance" is the variance of the number of orbitals in each subset included in the division pattern. "Sum of squares of electron number differences" is the sum of squares of the difference between the total number of active electrons and the number of active atoms in each subset. "Number of subsets" is the number of subsets included in the division pattern. "Bus orbital energy" is the energy of the bus orbitals, which represent the interactions between each subset in DMET. In the example of Figure 5, a certain division pattern of a molecule that can be calculated is shown to have an energy precision of 0.1148, a maximum number of orbitals of 16, a minimum number of orbitals of 10, an orbital number dispersion of 6.0, a sum of squares of the difference in the number of electrons of 1302.0, a number of subsets of 4, and a bath orbital energy of 0.43.

[0029] The "metric value" is information about the division candidates obtained by dividing the molecule to be calculated, and includes "maximum number of orbitals, minimum number of orbitals, orbital number dispersion, sum of squares of electron number differences, number of subsets, and bath orbital energy." Each piece of information is the same as described above, so detailed explanations will be omitted. The example in Figure 6 shows that a division candidate for the molecule to be calculated has "maximum number of orbitals of 22, minimum number of orbitals of 8, orbital number dispersion of 8.0, sum of squares of electron number differences of 2302.0, number of subsets of 7, and bath orbital energy of 0.32."

[0030] As shown in FIG. 6, the data structure DB 14 includes data structures for "accuracy estimation regression formula," "subset division candidates," and "subset division candidates and their scores."

[0031] The "accuracy estimation regression equation" is a regression equation for calculating the energy accuracy generated by regression analysis by the control unit 20, and is "f (metric value)." For example, the estimated energy accuracy is expressed as a linear combination of "c0 × sum of squares of electron number difference + c1 × maximum number of orbitals + c2 × minimum number of orbitals + c3 × orbital number variance + c4 × bus orbital energy + c5 × number of subsets + constant term." Note that cX is a coefficient for each metric (X ranges from 0 to 5 in this example).

[0032] A "subset division candidate" is a division candidate obtained by dividing the molecule to be calculated, and is expressed as "(id,id···)". For example, "(0),(1,2)···" indicates that the molecule is divided into a "subset of only atom "O"" with id "0", and a "subset including atom "O" with id "1" and atom "C" with id "2".

[0033] "Subset division candidates and their scores" are rankings based on the accuracy of the energy predicted using a regression equation for the above subset division candidates. Since it is the difference (energy accuracy) between the correct energy calculated by CCSD(T) for the subsets included in each division candidate, the smaller the value, the higher the score. In the example of Figure 6, it is shown that a score of "1" was calculated for the division candidate "(0),(1,2)...".

[0034] The control unit 20 is a processing unit that controls the entire information processing device 10, and is realized by, for example, a processor. The control unit 20 has a regression equation generation unit 30, an inference unit 40, and an energy calculation unit 50. The regression equation generation unit 30, the inference unit 40, and the energy calculation unit 50 are realized by, for example, electronic circuits included in the processor or processes executed by the processor.

[0035] The regression equation generating unit 30 has a dividing unit 31 and a deriving unit 32, and is a processing unit that generates a regression equation for predicting the accuracy of the potential energy using a first molecule of a calculable size.

[0036] The dividing unit 31 is a processing unit that divides each atom included in the first molecule into a plurality of patterns including a plurality of subsets formed by bonding each atom. Specifically, the dividing unit 31 divides the first molecule into a plurality of patterns so that the total number of orbitals, which is the sum of the numbers of orbitals of each atom included in the subset, falls within the number of orbitals that can be calculated and specified by a user or the like. For example, the dividing unit 31 can build a subset from atoms that are bonded to only one other atom by a breadth-first search, and then use this as a basis to move some atoms between subsets to generate a plurality of candidates.

[0037] The dividing unit 31 can also generate multiple division patterns from the pattern (specific pattern) with the largest number of orbitals that is equal to or less than the upper limit. For example, the dividing unit 31 generates multiple division patterns from the specific pattern by moving atoms of each subset included in the specific pattern to other subsets within a range in which the total number of orbitals is equal to or less than the upper limit. At this time, for example, a constraint can be set that limits the movement of one atom of each subset.

[0038] The derivation unit 32 is a processing unit that derives a regression equation that predicts calculation accuracy based on the number of orbitals, the number of electrons, etc. of each subset, using a first molecule that is large enough to determine the overall energy.

[0039] Specifically, the derivation unit 32 performs the following regression analysis for each of the multiple patterns (division patterns) to derive a regression equation. For example, the derivation unit 32 calculates the accurate potential energy of the first molecule using CCSD(T) or the like, and calculates the estimated potential energy, which is the potential energy calculated from each pattern, using DMET or the like. Then, the derivation unit 32 calculates the difference (energy accuracy) between the accurate potential energy of the first molecule and each estimated potential energy, and calculates each metric value shown in FIG. 5. Thereafter, the derivation unit 32 performs regression analysis using the energy accuracy and each metric value, and generates a regression equation for estimating the energy accuracy.

[0040] As mentioned above, the regression equation is expressed as a coefficient and a constant term for each metric. Normalization is performed when deriving the regression equation, and an inverse transformation is performed when predicting the energy accuracy.

[0041] The inference unit 40 has a division unit 41 and a presentation unit 42, and is a processing unit that uses the regression equation generated by the regression equation generation unit 30 to infer division candidates for calculating the energy of the second molecule that is the object of calculation.

[0042] The division unit 41 is a processing unit that generates a plurality of division candidates including a plurality of subsets obtained by bonding each atom included in the second molecule. Specifically, the division unit 41 generates a plurality of division candidates including a plurality of subsets from the second molecule using a method similar to the method used by the division unit 31 when deriving the regression formula. That is, the division unit 41 generates a plurality of division candidates from the second molecule using a breadth search so that the number of division candidates is equal to or less than the upper limit value used when generating the plurality of division patterns for the first molecule.

[0043] The presentation unit 42 is a processing unit that applies a regression equation to a plurality of division candidates, each of which includes a plurality of subsets, generated from the second molecule, and executes a prediction of the calculation accuracy of the potential energy of the second molecule when each of the plurality of division candidates is used. The presentation unit 42 is also a processing unit that outputs information that associates each of the plurality of division candidates with the prediction result of the calculation accuracy of the potential energy of the second molecule when each of the plurality of division candidates is used.

[0044] The energy calculation unit 50 is a processing unit that calculates the potential energy of the second molecule. Specifically, the energy calculation unit 50 calculates the potential energy of the second molecule by DMET using the division candidate with the highest calculation accuracy inferred (predicted) by the presentation unit 42. For example, the energy calculation unit 50 calculates the energy of each of multiple subsets included in the division candidate with the highest calculation accuracy, and combines the energies of each of the multiple subsets to calculate the potential energy of the second molecule. Note that a quantum simulator or the like can also be used for the combination calculation.

[0045] (Flow of the regression equation derivation process) 7 is a flowchart showing the process of deriving a regression equation. As shown in Fig. 7, the regression equation generating unit 30 lists molecules of a calculable scale (S101), and performs the following process loop for the listed molecules (S102 to S106).

[0046] Specifically, the regression equation generation unit 30 generates candidates that can divide the molecule within a range in which the number of orbitals in each subset is equal to or less than a limit (upper limit value) (S103), calculates metric values ​​of the generated candidates (S104), and calculates the energy of the entire molecule using a high-precision algorithm (e.g., CCSD(T)) (S105). Note that steps S103 to S105 may be executed in parallel for each molecule.

[0047] When the loop processing ends (S106), the regression equation generating unit 30 derives a regression equation from the collected accuracy and metric values ​​(S107).

[0048] (Flow of division candidate presentation process) Fig. 8 is a flowchart showing the flow of the division candidate presentation process. As shown in Fig. 8, the inference unit 40 generates division candidates that can be divided from the target molecule within a range in which the number of orbitals in each subset is equal to or less than a limit (S201).

[0049] Next, the inference unit 40 calculates the metric values ​​of the generated division candidates (S202), and applies the metric values ​​to a regression equation to calculate the prediction accuracy (S203). After that, the inference unit 40 sorts the generated division candidates by the calculated metric values, and displays them in a ranked order (S204).

[0050] (Example) Next, a specific example of generating a regression equation using the first molecule and calculating the potential energy of the second molecule will be described with reference to FIGS. 9 to 15. Here, alanine (C3H7N02) is used as an example of a calculable first molecule, and heptanoic acid (C7H 14 O2) is used.

[0051] (Derivation of regression formula: Generation of division pattern) First, the information processing device 10 decomposes alanine into multiple division patterns, with the number of orbitals that can be calculated set as the upper limit of the total number of orbitals in each subset. FIG. 9 is a diagram illustrating a specific example of a division pattern. As shown in FIG. 9, the information processing device 10 generates a pattern in which each atom contained in alanine (C3H7N02) is divided into multiple subsets by a breadth search that takes into account the connections in the molecular structure, under the constraint that the number of orbitals is limited to "8" or less. Here, a breadth-first search is used, starting from an atom that is connected to only one other atom.

[0052] For example, the information processing device 10 generates a pattern "O, O, N, C, C, C, H, H, H, H, H, H, H" with each atom as a subset, and adopts this pattern as the search result (ID = 0) because the maximum number of orbitals for this pattern is "N = 5, O = 5, C = 5", which is less than the orbital number limit value (8).

[0053] Next, the information processing device 10 generates a pattern "O,O,N,CH,C,C,H,H,H,H,H,H" from the pattern "ID=0" by bonding "C" and adjacent "H" according to the molecular structure, and since the maximum number of orbitals in this pattern is "CH=5+1=6", which is less than the orbital number limit value (8), it is adopted as the search result (ID=1).

[0054] Furthermore, the information processing device 10 generates a pattern "O, O, N, CH, CH, C, H, H, H, H, H" from the pattern "ID=1" by bonding "C" and adjacent "H" according to the molecular structure, and since the maximum number of orbitals for this pattern is "CH=5+1=6", which is less than the orbital number limit value (8), it is adopted as the search result (ID=2).

[0055] Next, the information processing device 10 generates a pattern "O, O, N, CH, CHH, C, H, H, H, H" from the pattern "ID=2" by combining "CH" and adjacent "H" according to the molecular structure, and since the maximum number of orbitals for this pattern is "CHH=5+1+1=7", which is less than the orbital number limit value (8), it is adopted as the search result (ID=3).

[0056] Furthermore, the information processing device 10 generates a pattern "O, O, N, CH, CHHH, C, H, H, H" from the pattern "ID=3" by combining "CHH" and adjacent "H" according to the molecular structure, and since the maximum number of orbitals in this pattern is "CH=5+1+1+1=8", which is less than the orbital number limit value (8), it is adopted as the search result (ID=4).

[0057] Next, the information processing device 10 generates a pattern "O, O, NH, CH, CHHH, C, H, H" from the pattern "ID=4" by bonding "N" and adjacent "H" according to the molecular structure, and since the maximum number of orbitals for this pattern is "CHHH=5+1+1+1=8", which is less than the orbital number limit value (8), it is adopted as the search result (ID=5).

[0058] Furthermore, the information processing device 10 generates a pattern "O, O, NHH, CH, CHHH, C, H" from the pattern "ID=5" by bonding "NH" and adjacent "H" according to the molecular structure, and since the maximum number of orbitals for this pattern is "CHHH=5+1+1+1=8" which is less than the orbital number limit value (8), it is adopted as the search result (ID=6).

[0059] Furthermore, the information processing device 10 generates a pattern "OH, O, NHH, CH, CHHH, C" from the pattern "ID=6" by bonding "O" and adjacent "H" according to the molecular structure, and since the maximum number of orbitals for this pattern is "CHHH=5+1+1+1=8" which is less than the orbital number limit value (8), it is adopted as the search result (ID=7).

[0060] Thereafter, the information processing device 10 ends the division because, when adjacent atoms are combined from the pattern "OH, O, NHH, CH, CHHH, C" with ID=7, for example, "CO" or "CH-CHHH", and the maximum number of orbitals exceeds the orbital number limit value (8). As a result, the information processing device 10 generates 8 patterns with ID=0 to 7.

[0061] (Derivation of regression formula: Calculation of energy accuracy) Next, the information processing device 10 further generates a plurality of candidates from the obtained subsets. Here, one atom is moved from each subset, and if the number of orbitals exceeds the limit, the subset is further divided.

[0062] 10 is a diagram illustrating a specific example of energy accuracy calculation. As shown in FIG. 10, the information processing device 10 adopts the pattern with the largest number of orbitals from the patterns generated in FIG. 9 as the division pattern. In this example, the information processing device 10 selects the pattern "OH, O, NHH, CH, CHHH, C" with ID=7, which has the largest number of orbitals and the smallest number of divided subsets. Note that the smaller the number of divided subsets, the easier it is to generate further division patterns.

[0063] Then, the information processing device 10 generates a plurality of split patterns by performing the width search described in Fig. 9 on the pattern "OH, O, NHH, CH, CHHH, C" with ID=7. In the example of Fig. 10, the information processing device 10 generates 11 split patterns such as "H, O, NHH, CH, CHHH, O, C" and "OH, NHH, CH, CHHH, O, C" from the pattern "OH, O, NHH, CH, CHHH, C" with ID=7.

[0064] Next, the information processing device 10 collects metric values ​​(energy precision, sum of squares of electron number differences, maximum number of orbitals, minimum number of orbitals, orbital number variance, bus orbital energy, and number of subsets) for each of the 11 division patterns. Here, the division pattern "OH, O, NHH, CH, CHHH, C" will be used as an example for explanation.

[0065] For example, the information processing device 10 calculates the accurate potential energy of alanine using CCMT(T). Furthermore, the information processing device 10 calculates the energy of each subset "OH", "O", "NHH", "CH", "CHHH", and "C" using DMET or a quantum algorithm, and combines them to calculate an estimated potential energy obtained from the division pattern. The information processing device 10 then calculates the difference between the accurate potential energy and the estimated potential energy as "energy accuracy: 0.8736".

[0066] The information processing device 10 calculates the sum of squares of the difference between the total number of active electrons and the number of active atoms in each subset, and sets the result as "sum of squares of electron number difference: 3807.333." The information processing device 10 doubles the maximum number of orbitals in the division pattern, "CHHH=8," to "16" in consideration of spin, and sets the "minimum number of orbitals" to "10," doubles the minimum number of orbitals in the division pattern, "C=5," in consideration of spin. The information processing device 10 calculates a dispersion value based on the number of orbitals including spin, and sets "orbital number dispersion: 4.555556." The information processing device 10 also calculates the energy of the bus orbital using DMET and sets "bus orbital energy: 3.29E-15." The information processing device 10 also sets the number of subsets in the division pattern "OH, O, NHH, CH, CHHH, C" to "6."

[0067] Using the above-described method, the information processing device 10 collects metric values ​​(energy precision, sum of squares of electron number differences, maximum number of orbitals, minimum number of orbitals, orbital number variance, bus orbital energy, and number of subsets) for each of the 11 division patterns generated from the pattern "OH, O, NHH, CH, CHHH, C" with ID=7.

[0068] (Derivation of regression equation: regression analysis) Next, the information processing device 10 generates a regression equation using the metric value and energy accuracy of each division candidate obtained in Figure 10. For example, the information processing device 10 generates a regression equation that calculates the energy accuracy from the metric value by performing regression analysis with the "energy accuracy" of each division candidate as the objective variable and each metric value as the explanatory variable. In other words, the "energy accuracy" value calculated by the regression equation is information that indicates the difference from the accurate potential energy of alanine using CCMT(T), so the smaller the value, the better the accuracy.

[0069] FIG. 11 is a diagram illustrating a specific example of regression analysis. As shown in FIG. 11, the information processing device 10 generates a regression equation expressed as a linear combination of "c0 × sum of squares of difference in electron number + c1 × maximum number of orbitals + c2 × minimum number of orbitals + c3 × orbital number variance + c4 × bus orbital energy + c5 × number of subsets + constant term" as a regression for calculating the estimated energy precision. Note that the numerical values ​​corresponding to each metric shown in FIG. 11 are coefficients such as c0 and c1 and constant terms. For example, when the maximum number of orbitals is "3.65842533 × 10 -3 " corresponds to the coefficient "c1".

[0070] (Derivation of molecular potential energy: Generation of division candidates) When the generation of the regression equation is completed, the information processing device 10 calculates the heptanoic acid (CH 14 O2), a division pattern is generated by dividing the data using the same method as when generating the regression equation.

[0071] 12 is a diagram illustrating a specific example of division candidates for a molecule to be calculated. As shown in FIG. 12, the information processing device 10 generates patterns in which the molecule is divided into a plurality of subsets by a breadth search that takes into account the connections in the molecular structure so that the orbital number limit of "8" is met. Then, the information processing device 10 identifies the division pattern "OH, O, CHH, CHH, CHH, CHH, CHH, CHHH, C" with ID=6, which has the largest number of orbitals.

[0072] Thereafter, for the division pattern "OH, O, CHH, CHH, CHH, CHH, CHH, CHHH, C" with ID=6, the information processing device 10 moves one atom from each subset to generate division candidates 1 to 6 in which the number of orbitals does not exceed the limit. Then, the information processing device 10 calculates metric values ​​for each of the division candidates 1 to 6 using a method similar to the method described in FIG. 10. As a result, the information processing device 10 can collect metric values ​​for each division candidate.

[0073] (Derivation of molecular potential energy: Calculation of estimated energy) Next, the information processing device 10 calculates the energy accuracy for each of the division candidates 1 to 6. FIG. 13 is a diagram for explaining a specific example of the calculation of the estimated energy of the division candidates. As shown in FIG. 13, the information processing device 10 calculates the metric values ​​of the division candidate 1 (sum of squares of the difference in the number of electrons: 11005.33, maximum number of orbitals: 16, minimum number of orbitals: 10, orbital number variance: 3.654321, bus orbital energy: -1.5×10 -13 , number of subsets: 9) is multiplied by the corresponding coefficients c0 to c5 of the regression equation (the sum of squares of the difference in the number of electrons (= c0), the maximum number of orbitals (= c1), the minimum number of orbitals (= c2), the orbital number dispersion (= c3), the bus orbital energy (= c4), and the number of subsets (= c5)). The energy precision is calculated as 0.20077769, which is the sum of each multiplication value and a constant term.

[0074] As mentioned above, the energy accuracy calculated here is a value that indicates the difference from the accurate potential energy, so the smaller the value, the better the accuracy. Note that Fig. 13 shows an example of division candidate 1 in Fig. 12, but the same process is performed for division candidates 2 to 6.

[0075] (Derivation of molecular potential energy: ranking) Next, the information processing device 10 calculates the "energy accuracy" for each of the division candidates 1 to 6 using the method described with reference to FIG. 13, and ranks them in order of best "energy accuracy."

[0076] Fig. 14 is a diagram illustrating a specific example of ranking division candidates. As shown in Fig. 14, the information processing device 10 calculates the estimated energy accuracy for each of division candidate 1 to division candidate 6 described in Fig. 12 using the method described in Fig. 13. For example, the information processing device 10 calculates "estimated energy accuracy: 0.20077769" for division candidate 1, "estimated energy accuracy: 0.20045908" for division candidate 2, and "estimated energy accuracy: 0.20077769" for division candidate 3. Similarly, the information processing device 10 calculates "estimated energy accuracy: 0.20046347" for division candidate 4, "estimated energy accuracy: 0.20046347" for division candidate 5, and "estimated energy accuracy: 0.20046347" for division candidate 6.

[0077] Then, the information processing device 10 performs ranking in descending order of estimated energy accuracy. Specifically, the information processing device 10 assigns a higher rank to smaller estimated energy accuracy values, and when the estimated energy accuracy values ​​are the same, performs ranking under specified conditions such as assigning a higher rank to a smaller number of subsets, a higher rank to a number earlier in the candidate order, or a higher rank to a larger sum of squares of bus orbital energy or electron number difference. In the example of Fig. 14, the information processing device 10 performs ranking in the order of division candidate 2, division candidate 4, division candidate 5, division candidate 6, division candidate 1, and division candidate 3. Note that the smaller the value, the higher the rank.

[0078] (Derivation of molecular potential energy: Presentation) Finally, the information processing device 10 presents the information ranked in FIG. 14 to the user by outputting it to the display unit 12 or transmitting it to the user terminal.

[0079] FIG. 15 is a diagram illustrating an example of a screen displaying division candidates. As shown in FIG. 15, the information processing device 10 outputs a screen displaying the ranking of each division candidate obtained in FIG. 14. For example, this screen includes a "regression formula sample molecule" indicating a molecule (e.g., alanine) selected for generating the regression formula, a "calculation target molecule" indicating a molecule (e.g., heptanoic acid) selected as a target for calculating potential energy, and a "division candidate list" indicating the division candidate number, subset information of the division candidate, estimated energy accuracy, and rank. Note that the information in the division candidate list uses information obtained in the process leading up to the ranking, as shown in FIG. 14. Furthermore, the information displayed on the screen is merely an example; it is sufficient that at least a division candidate with rank 1 is displayed; other information can be changed as desired.

[0080] Thereafter, the information processing device 10 calculates the potential energy of heptanoic acid using the division candidate with rank 1 or a division candidate selected by the user from the division candidate list.

[0081] (effect) As described above, the information processing device 10 derives a regression equation for estimating the accuracy of potential energy using molecules large enough to calculate the total energy, and calculates the estimated accuracy of energy for division candidates of large molecules for which the energy is to be calculated. As a result, the information processing device 10 can predict division candidates with high calculation accuracy when calculating the potential energy.

[0082] Furthermore, the information processing device 10 derives a regression equation, which is the result of analyzing, by regression analysis, the relationship between the estimated energy accuracy, which is the difference between the energy actually calculated by CCMT(T) or the like and the estimated potential energy, and the metric value. The information processing device 10 uses such a regression equation to determine the energy accuracy of the division candidates for the target molecule, and therefore can improve the accuracy of the potential energy finally obtained compared to when division candidates are generated randomly or when division candidates specified by the user are used.

[0083] Furthermore, the information processing device 10 can prevent an infinite number of division candidates from being generated by imposing constraints (for example, the number of orbitals) when generating division candidates, thereby preventing the series of processing steps from generating a regression equation to calculating the final energy from taking too long. Furthermore, as a result of being able to prevent the information processing device 10 from taking too long, the processing load on the processor of the information processing device 10 can be reduced, and processing speed can be increased.

[0084] Furthermore, the information processing device 10 imposes the same constraints when creating division candidates, both when deriving the regression equation and when calculating the estimated energy accuracy, thereby calculating the estimated energy accuracy under the same conditions as the regression equation, thereby enabling the estimated energy accuracy to be calculated with high accuracy. [Example]

[0085] Although the embodiments of the present invention have been described above, the present invention may be embodied in various different forms other than the above-described embodiments.

[0086] (Numbers, etc.) The numerical values ​​and division methods used in the above embodiments are merely examples and can be changed as desired. Furthermore, the values ​​are not necessarily accurate and are merely examples. Furthermore, the process flow described in each flowchart can be changed as appropriate within a consistent range.

[0087] (metric) In the above embodiment, an example was described in which the metrics used were "maximum number of orbitals, minimum number of orbitals, orbital number dispersion, sum of squares of electron number differences, subset number, and bus orbital energy." However, it is not necessary to use all of these metrics; at least one or a combination of two or more of them can also be used.

[0088] (splitting method) In the above embodiment, the information processing device 10 performs two-stage division. However, the information processing device 10 can also derive a regression formula and calculate an estimated energy accuracy with one stage of division. For example, the information processing device 10 selects one pattern from among patterns including subsets when deriving the regression formula and when calculating the estimated energy accuracy, and generates a further division pattern from the selected pattern. The information processing device 10 then further divides the division pattern to derive a regression formula and calculate an estimated energy accuracy. However, the present invention is not limited to this. In other words, the information processing device 10 can also derive a regression formula by calculating metric values, etc., at the stage of FIG. 9 instead of FIG. 10.

[0089] (system) The information including the processing procedures, control procedures, specific names, various data and parameters shown in the above documents and drawings may be changed arbitrarily unless otherwise specified.

[0090] Furthermore, the specific form of distribution or integration of the components of each device is not limited to that shown in the figure. For example, the regression equation generation unit 30 and the inference unit 40 may be integrated. That is, all or some of the components may be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc. Furthermore, all or any part of the processing functions of each device may be realized by a CPU and a program analyzed and executed by the CPU, or may be realized as hardware using wired logic.

[0091] Furthermore, all or any part of the processing functions performed by each device may be realized by a CPU and a program analyzed and executed by the CPU, or may be realized as hardware using wired logic.

[0092] (Hardware) Fig. 16 is a diagram illustrating an example of a hardware configuration. As shown in Fig. 16, an information processing device 10 includes a communication device 10a, a hard disk drive (HDD) 10b, a memory 10c, and a processor 10d. The components shown in Fig. 16 are connected to each other via a bus or the like.

[0093] The communication device 10a is a network interface card or the like, and communicates with other devices. The HDD 10b stores programs and DBs that operate the functions shown in FIG.

[0094] The processor 10d reads out a program that executes the same processes as the respective processing units shown in FIG. 3 from the HDD 10b or the like and loads it into the memory 10c, thereby operating a process that executes the respective functions described in FIG. 3 or the like. For example, this process executes the same functions as the respective processing units of the information processing device 10. Specifically, the processor 10d reads out a program that has the same functions as the regression equation generation unit 30, the inference unit 40, the energy calculation unit 50, etc. from the HDD 10b or the like. Then, the processor 10d executes a process that executes the same processes as the regression equation generation unit 30, the inference unit 40, the energy calculation unit 50, etc.

[0095] In this way, the information processing device 10 operates as an information processing device that executes an energy calculation method by reading and executing a program. The information processing device 10 can also realize functions similar to those of the above-described embodiment by reading the program from a recording medium using a medium reading device and executing the read program. Note that the program in this other embodiment is not limited to being executed by the information processing device 10. For example, the above-described embodiment may also be applied in the same way to a case where another computer or server executes the program, or a case where these execute the program in cooperation with each other.

[0096] This program may be distributed via a network such as the Internet. Alternatively, this program may be recorded on a computer-readable recording medium such as a hard disk, a flexible disk (FD), a CD-ROM, a magneto-optical disk (MO), or a digital versatile disk (DVD), and may be read out from the recording medium and executed by a computer. [Explanation of symbols]

[0097] 10. Information processing equipment 11 Communications Department 12 Display section 13 Storage section 14 Data Structure DB 20 Control Unit 30 Regression equation generation section 31 Split part 32 Derivation part 40 Reasoning part 41 Split part 42 Presentation part 50 Energy calculation unit

Claims

1. On the computer, generating a regression equation for predicting the calculation accuracy of the potential energy of the first molecule using a plurality of division patterns each including a plurality of subsets each including one or more atoms included in the first molecule; applying the regression equation to a plurality of division candidate patterns including a plurality of subsets each including one or more atoms included in the second molecule, and predicting the calculation accuracy of the potential energy of the second molecule when each of the plurality of division candidate patterns is used; An accuracy prediction program that executes a process.

2. The generating process includes: generating the regression equation for each of the plurality of division patterns by linear combination using at least one of the number of orbitals in each subset included in the division pattern, the number of electrons in each subset, the number of subsets included in the pattern, and energy of a bus orbital expressing an interaction between the subsets; 2. The accuracy prediction program according to claim 1.

3. The generating process includes: When the number of orbitals is used in the regression equation, the regression equation is generated for each of the plurality of division patterns by the linear combination using the maximum number of orbitals and the minimum number of orbitals in each subset included in the division pattern and a variance value of the number of orbitals in each subset.

3. The accuracy prediction program according to claim 2.

4. The generating process includes: generating the plurality of division patterns from the first molecule such that a total number of orbitals obtained by adding up the numbers of orbitals of each atom included in the subset is equal to or less than an upper limit value; Identifying a specific pattern that has the largest total number of trajectories from among the plurality of division patterns; generating a plurality of candidate division patterns from the specific pattern by moving atoms of each subset included in the specific pattern to another subset within a range in which the total number of orbitals is equal to or less than the upper limit value; generating the regression equation using each of the plurality of candidate division patterns; 2. The accuracy prediction program according to claim 1.

5. The process of performing the prediction includes: generating the plurality of candidate division patterns from the second molecule so that the plurality of candidate division patterns is equal to or less than the upper limit value used when generating the plurality of division patterns of the first molecule; 5. The accuracy prediction program according to claim 4.

6. calculating the energy of each of a plurality of subsets included in the division candidate pattern with the highest calculation accuracy among the calculation accuracy of the predicted potential energy of the second molecule; calculating a potential energy of the second molecule by combining the energies of each of the plurality of subsets; 2. The accuracy prediction program according to claim 1, further comprising causing the computer to execute a process.

7. The process of performing the prediction includes: outputting information associating each of the plurality of candidate division patterns with a predicted result of calculation accuracy of the potential energy of the second molecule when each of the plurality of candidate division patterns is used; 2. The accuracy prediction program according to claim 1.

8. The computer generating a regression equation for predicting the calculation accuracy of the potential energy of the first molecule using a plurality of division patterns each including a plurality of subsets each including one or more atoms included in the first molecule; applying the regression equation to a plurality of division candidate patterns including a plurality of subsets each including one or more atoms included in the second molecule, and predicting the calculation accuracy of the potential energy of the second molecule when each of the plurality of division candidate patterns is used; The accuracy prediction method is characterized by performing a process.

9. generating a regression equation for predicting the calculation accuracy of the potential energy of the first molecule using a plurality of division patterns each including a plurality of subsets each including one or more atoms included in the first molecule; applying the regression equation to a plurality of division candidate patterns including a plurality of subsets each including one or more atoms included in the second molecule, and predicting the calculation accuracy of the potential energy of the second molecule when each of the plurality of division candidate patterns is used; An information processing device comprising a control unit.

Citation Information

Patent Citations

  • Simulating quantum systems with quantum computation

    US20180096085A1