Information processing device, information processing method, and information processing program

JP7920939B2Active Publication Date: 2026-09-15TOYOTA JIDOSHA KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023012269
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-01-30
Publication Date
2026-09-15
Estimated Expiration
2043-01-30

Smart Images

  • Figure 0007920939000001
    Figure 0007920939000001
  • Figure 0007920939000002
    Figure 0007920939000002
  • Figure 0007920939000003
    Figure 0007920939000003
Patent Text Reader

Abstract

To improve possibility that an unknown molecule satisfying a performance condition can be discovered compared with a case where it is collated with a database.SOLUTION: A server 10 creates an atomic coordinate image representing an atomic coordinate in a molecule, performs Fourier transformation of the created atomic coordinate image to generate power spectral data, performs principal component analysis for the generated power spectral data, derives a principal component vector representing a base vector of the power spectral data and a principal component score representing an amount, in which the principal component vector is included, from the power spectral data, derives an index value representing a correlation degree between the derived principal component score and performance of the molecule, identifies a principal component vector that is correlated with performance of the molecule on the basis of the derived index value, and outputs principal component power spectral data corresponding to the identified principal component vector.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program. [Background Art]

[0002] For example, Patent Literature 1 describes a method for searching for a novel substance. This method includes the steps of: performing learning on a substance model modeled based on known substances; inputting target physical properties into a learning result to determine at least one candidate substance; and determining a novel substance from the at least one candidate substance. [Prior Art Document] [Patent Document]

[0003] [Patent Document 1] Japanese Unexamined Patent Publication No. 2017-091526 [Summary of the Invention] [Problem to be Solved by the Invention]

[0004] According to the technique described in the above Patent Document 1, the relationship between structural information and property information of known substances is learned by machine learning, and at least one candidate substance is determined by inputting target physical properties into the obtained trained model.

[0005] However, in the above technique, when searching for molecules constituting a substance, known techniques such as ECFP (Extended Connectivity Circular Fingerprints) and RDKit are used to investigate the relationship between molecular feature amounts and performance, and molecules that may satisfy performance conditions are collated with a database. For this reason, it is impossible to search for molecules other than those stored in the database.

[0006] This invention has been made in consideration of the above facts, and aims to provide an information processing device, an information processing method, and an information processing program that can increase the probability of finding unknown molecules that meet performance requirements compared to matching a database. [Means for solving the problem]

[0007] To achieve the above objective, the information processing apparatus described in claim 1 includes: a creation unit that creates an atomic coordinate image representing the atomic coordinates in a molecule; a spectrum generation unit that performs a Fourier transform on the created atomic coordinate image to generate power spectrum data; and a principal component derivation unit that performs principal component analysis on the generated power spectrum data and derives from the power spectrum data principal component vectors representing the basis vectors of the power spectrum data and principal component scores representing the quantities in which the principal component vectors are included. A pre-trained model, which has been machine-learned using training data that correlates principal component scores with molecular performance, The derived principal component score The trained model will input the following and respond to Molecular performance Outputs the principal component score and the performance of the molecule. The system includes an index value derivation unit that derives an index value representing the degree of correlation with the nutrient, an identification unit that identifies a principal component vector correlated with the performance of the nutrient based on the derived index value, and an output unit that outputs principal component power spectrum data, which is power spectrum data corresponding to the identified principal component vector.

[0008] According to the invention described in claim 1, the likelihood of finding an unknown molecule that satisfies the performance requirements can be increased compared to matching a database. Furthermore, the index values ​​can be derived with high accuracy using the pre-trained model.

[0011] Furthermore, claims 2 The information processing device described in claim 1 In the information processing device described above, the creation unit creates a two-dimensional or three-dimensional atomic coordinate image from a molecular data file that stores information representing the structure of the molecule.

[0012] Claim 2 According to the invention described, it is possible to provide two-dimensional or three-dimensional arrangement conditions for atoms.

[0013] Furthermore, claims 3 The information processing device described in claim 1 or Claim 2 The information processing device described above further comprises a map generation unit that identifies a plurality of principal component vectors and generates two-dimensional map data by projecting the principal component scores corresponding to each of the identified plurality of principal component vectors onto a two-dimensional plane as a plurality of plot points.

[0014] Claim 3 According to the invention described above, the correspondence between principal components that have a high correlation with the performance of molecules can be represented as a two-dimensional map.

[0015] Furthermore, claims 4 The information processing device described in claim 3 In the information processing device described above, the output unit causes the display unit to display the atomic coordinate images corresponding to the plot points of the two-dimensional map data together with the two-dimensional map data.

[0016] Claim 4 According to the invention described herein, a user can interpret an atomic coordinate image corresponding to a plotted point.

[0017] Furthermore, in order to achieve the above objective, 5 The information processing method described herein involves an information processing device creating an atomic coordinate image representing the atomic coordinates in a molecule, performing a Fourier transform on the created atomic coordinate image to generate power spectral data, performing principal component analysis on the generated power spectral data, and deriving from the power spectral data principal component vectors representing the basis vectors of the power spectral data, and principal component scores representing the quantities in which the principal component vectors are included. A pre-trained model, which has been machine-learned using training data that correlates principal component scores with molecular performance, The derived principal component score The trained model will input the following and respond to Molecular performance The output is the principal component score and the performance of the molecule. An index value representing the degree of correlation with the nutrient is derived, and based on the derived index value, a principal component vector correlated with the performance of the nutrient is identified, and principal component power spectrum data, which is power spectrum data corresponding to the identified principal component vector, is output.

[0018] Claim 5 According to the invention described in [claim 1], similar to claim 1, the possibility of finding an unknown molecule that satisfies performance conditions can be increased compared to a case of collating a database. Furthermore, the index values ​​can be derived with high accuracy using the pre-trained model.

[0019] Further, in order to achieve the above object, the claim 6 The information processing program described in causes a computer to execute: creating an atomic coordinate image representing atomic coordinates in a molecule; performing Fourier transform on the created atomic coordinate image to generate power spectrum data; performing principal component analysis on the generated power spectrum data; deriving, from the power spectrum data, a principal component vector representing a basis vector of the power spectrum data and a principal component score representing an amount of the principal component vector contained in the power spectrum data; A pre-trained model, which has been machine-learned using training data that correlates principal component scores with molecular performance, deriving an index value representing a degree of correlation between the derived principal component score The trained model will input the following and respond to and performance of the molecule Outputs the principal component score and the performance of the molecule. identifying a principal component vector that correlates with the performance of the molecule based on the derived index value; and outputting principal component power spectrum data, which is power spectrum data corresponding to the identified principal component vector.

[0020] Claim 6 According to the invention described in [claim 1], similar to claim 1, the possibility of finding an unknown molecule that satisfies performance conditions can be increased compared to a case of collating a database. Furthermore, the index values ​​can be derived with high accuracy using the pre-trained model. Effects of the Invention

[0021] As described above, according to the present invention, an effect that the possibility of finding an unknown molecule that satisfies performance conditions can be increased compared to a case of collating with a database is obtained. Brief Description of the Drawings

[0022] [Figure 1] It is a block diagram showing an example of the configuration of an information processing system according to an embodiment. Note: Corrected the typo "であえる" to the standard "である" for accurate translation [Figure 2] This block diagram shows an example of the functional configuration of a server according to the embodiment. [Figure 3] (A) is a diagram showing an example of a molecular data file according to the embodiment. (B) is a diagram showing an example of an atomic coordinate image according to the embodiment. [Figure 4] This figure illustrates the power spectral conversion of an atomic coordinate image according to the embodiment. [Figure 5] This figure illustrates the feature quantification of waveform data according to the embodiment. [Figure 6] (A) is a graph showing an example of plotted points where the ground truth and predicted values ​​for the performance of a molecule output from the trained model according to the embodiment are plotted. (B) and (C) are figures showing an example of index values ​​derived for each principal component vector using the performance of the molecule obtained by the trained model. (D) is a graph showing an example of principal component power spectral data corresponding to the principal component vector. [Figure 7] (A) is a graph showing an example of power spectral data before changing the principal component scores of principal components PC1 and PC3. (B) is a graph showing an example of power spectral data after changing the principal component scores of principal components PC1 and PC3. [Figure 8] This figure shows an example of two-dimensional map data and atomic coordinate images according to the embodiment. [Figure 9] This figure illustrates the prediction using mol file structure information according to the embodiment. [Figure 10] This flowchart shows an example of the processing flow by the information processing program according to the embodiment. [Modes for carrying out the invention]

[0023] Hereinafter, an example of an embodiment for carrying out the technology of this disclosure will be described in detail with reference to the drawings. Components and processes that perform the same operation, action, or function are given the same reference numerals throughout the drawings, and redundant explanations may be omitted as appropriate. Each drawing is only a schematic representation to the extent that the technology of this disclosure can be fully understood. Therefore, the technology of this disclosure is not limited to the illustrated examples. Furthermore, in this embodiment, explanations of configurations not directly related to the present invention or well-known configurations may be omitted.

[0024] Figure 1 is a block diagram showing an example of the configuration of the information processing system 100 according to this embodiment.

[0025] As shown in Figure 1, the information processing system 100 according to this embodiment includes a server 10 and a user terminal 30. The server 10 is an example of an information processing device. The server 10 and the user terminal 30 are connected via a network N so as to be able to communicate with each other.

[0026] Server 10 comprises a CPU (Central Processing Unit) 11, ROM (Read Only Memory) 12, RAM (Random Access Memory) 13, input / output interface (I / O) 14, storage unit 15, display unit 16, operation unit 17, and communication unit 18. Server 10 is configured, for example, as a general-purpose computer device.

[0027] The CPU 11, ROM 12, RAM 13, and I / O 14 are connected to each other via a bus. The I / O 14 is connected to various functional units, including a storage unit 15, a display unit 16, an operation unit 17, and a communication unit 18. These functional units are capable of communicating with the CPU 11 via the I / O 14.

[0028] The control unit is comprised of a CPU 11, ROM 12, RAM 13, and I / O 14. The control unit may be configured as a sub-control unit that controls some of the operations of the server 10, or as part of a main control unit that controls the overall operation of the server 10. Some or all of the blocks in the control unit may use integrated circuits or IC chipsets, such as LSIs (Large Scale Integrations). Individual circuits may be used for each of the above blocks, or circuits that integrate some or all of them may be used. The above blocks may be provided as a single unit, or some of the blocks may be provided separately. Furthermore, parts of each of the above blocks may be provided separately. For the integration of the control unit, dedicated circuits or general-purpose processors may be used, not just LSIs.

[0029] For example, the storage unit 15 can be an HDD (Hard Disk Drive), an SSD (Solid State Drive), or flash memory. The information processing program 15A according to this embodiment is stored in the storage unit 15. This information processing program 15A may also be stored in the ROM 12.

[0030] The information processing program 15A may, for example, be pre-installed on the server 10. The information processing program 15A may also be implemented by storing it on a non-volatile storage medium or distributing it via the network N and installing it on the server 10 as appropriate. Examples of non-volatile storage mediums include CD-ROMs (Compact Disc Read Only Memory), magneto-optical disks, HDDs, DVD-ROMs (Digital Versatile Disc Read Only Memory), flash memory, and memory cards.

[0031] The display unit 16 may include, for example, a liquid crystal display (LCD), an organic EL (electroluminescence) display, etc. The display unit 16 may also have an integrated touch panel. The operation unit 17 is equipped with, for example, a keyboard, mouse, or other device for operation input. The display unit 16 and the operation unit 17 receive various instructions from the user of the server 10. The display unit 16 displays various information such as the results of processing performed in response to the instructions received from the user, and notifications regarding the processing.

[0032] The communication unit 18 is connected to a network N, such as the Internet, LAN (Local Area Network), or WAN (Wide Area Network), and is capable of communicating with the user terminal 30 via the network N.

[0033] The user terminal 30 is operated by the user. Functionally, the user terminal 30 comprises a control unit 31 and a display unit 32, as shown in Figure 1.

[0034] The control unit 31 controls the operation of the user terminal 30. The display unit 32 displays various information in accordance with the control by the control unit 31.

[0035] The server 10 of the information processing system 100 according to this embodiment creates an atomic coordinate image representing the atomic coordinates in a molecule, generates power spectrum data by performing a Fourier transform on the atomic coordinate image, and derives principal component vectors and principal component scores, which will be described later, by performing principal component analysis on the power spectrum data. The server 10 of the information processing system 100 then uses a trained model to derive an index value representing the degree of correlation between the principal component score and molecular performance, identifies principal component vectors with a relatively high correlation to molecular performance based on the derived index value, and outputs principal component power spectrum data, which is the power spectrum data corresponding to the identified principal component vectors. This makes it possible to interpret the shape (power spectrum) of the principal components and clarify the requirements for the molecular structure. In other words, instead of matching a database, it is possible to increase the possibility of finding unknown molecules (unknown structures) that satisfy the performance conditions by providing atomic arrangement conditions using atomic coordinate images.

[0036] Specifically, the CPU 11 of the server 10 according to this embodiment functions as the various parts shown in Figure 2 by writing the information processing program 15A stored in the ROM 12 or memory unit 15 to the RAM 13 and executing it.

[0037] Figure 2 is a block diagram showing an example of the functional configuration of the server 10 according to this embodiment.

[0038] As shown in Figure 2, the CPU 11 of the server 10 according to this embodiment functions as an acquisition unit 11A, a creation unit 11B, a spectrum generation unit 11C, a principal component derivation unit 11D, an index value derivation unit 11E, a specific unit 11F, an output unit 11G, and a map generation unit 11H.

[0039] The acquisition unit 11A acquires molecular data files from the user terminal 30. A molecular data file is a file that stores information representing the structure of a molecule, and a mol file is used as an example.

[0040] The creation unit 11B creates an atomic coordinate image representing the atomic coordinates in the molecule. Specifically, the creation unit 11B creates a two-dimensional or three-dimensional atomic coordinate image from the molecular data file acquired by the acquisition unit 11A. The molecular data file contains information indicating the structural arrangement of the atoms that make up the molecule. The atomic coordinate image is created, for example, based on the structural arrangement of atoms obtained from the molecular data file.

[0041] Figure 3(A) shows an example of a molecular data file according to this embodiment. Figure 3(B) shows an example of an atomic coordinate image according to this embodiment.

[0042] The molecular data file shown in Figure 3(A) is, as an example, a mole file, corresponding to N molecules in the sample size. The atomic coordinate image shown in Figure 3(B) corresponds to one image for each molecule in the molecular data file, with the white dots in the image representing atoms.

[0043] The spectrum generation unit 11C generates power spectrum data by performing a Fourier transform on the atomic coordinate image created by the creation unit 11B.

[0044] Figure 4 is a diagram illustrating the power spectral conversion of the atomic coordinate image according to this embodiment.

[0045] As shown in Figure 4, a Fourier transform is performed on the atomic coordinate image to generate 2D power spectral data. In this Fourier transform, the wave represents a single 2D wavenumber space. In the case of a 3D atomic coordinate image, a 3D Fourier transform is performed to generate 3D power spectral data. Next, the 2D power spectral data is integrated circumferentially to generate 1D power spectral data. This circumferential integration takes into account the superposition of waves and their intensity (pixel values). In the case of 3D power spectral data, the integration is performed circumferentially, similar to the 2D case, to generate 1D power spectral data. In power spectral generation, the features of the image are represented as "waves" by performing a Fourier transform on the atomic coordinate image. This allows us to obtain periodic structure information derived from the "size of particles," "shape of particles," and "arrangement of particles" in the image.

[0046] The principal component derivation unit 11D performs principal component analysis (PCA), a type of dimensionality reduction method, on the power spectral data generated by the spectrum generation unit 11C, and derives principal component vectors and principal component scores from the power spectral data. The principal component vectors represent the basis vectors of the power spectral data. Each principal component vector has one of the spectral values ​​of the principal components as a component. The principal component score is a feature of the power spectral data, and is a coefficient that represents the amount to which the principal component vectors are included, that is, how much of the principal component vector components are included.

[0047] Figure 5 is a diagram illustrating the feature quantification of waveform data according to this embodiment.

[0048] As shown in Figure 5, principal component analysis (PCA) is performed on the waveform data to derive principal component vectors and principal component scores. In the case of X-ray diffraction, for example, the waveform data is represented by the diffraction angle (2θ / degree) on the horizontal axis and the intensity on the vertical axis. In the example in Figure 5, 10 principal component vectors are derived. Hereafter, the 10 principal components will be represented as PC1 to PC10. The principal component vectors shown in Figure 5 are the principal component vector of the first principal component PC1, the principal component vector of the second principal component PC2, and the principal component vector of the third principal component PC3.

[0049] On the other hand, the principal component scores shown in Figure 5 are derived for each of the principal components PC1 to PC10 for each waveform data. In other words, the principal component scores indicate the amount in which each of the principal component vectors PC1 to PC10 is included for each of the multiple waveform data.

[0050] The index value derivation unit 11E derives an index value that represents the degree of correlation between the principal component scores derived by the principal component derivation unit 11D and the performance of the molecule. The index value derivation unit 11E may derive the index value using, for example, a trained model 15B stored in the memory unit 15. The trained model 15B is a model that has been pre-trained to take principal component scores as input and output the performance of the molecule. Specifically, multiple training data sets corresponding to principal component scores and the performance of the molecule may be prepared, a trained model may be generated based on the multiple training data sets, and the index value may be derived using that trained model. In this case, a known machine learning model is used for the trained model 15B. The trained model 15B is generated, for example, by training a machine learning model using a deep learning algorithm. In this embodiment, the "index value" is a value that shows a positive correlation and a negative correlation with respect to the performance of the molecule. In a positive correlation, the "index value" is higher the degree to which it contributes to the nutrient's performance, and in a negative correlation, the index value is lower the degree to which it contributes to the nutrient's performance. Furthermore, "nutrient performance" represents the properties or capabilities that the nutrient possesses.

[0051] Figure 6(A) is a graph showing an example of plotted points representing the ground truth and predicted values ​​for the molecular performance output from the trained model 15B according to this embodiment. The horizontal axis represents the ground truth (true), and the vertical axis represents the predicted value (predict).

[0052] As shown in Figure 6(A), machine learning of the trained model 15B uses, for example, training data (train) and test data (test). Training data (train) is data used to train the machine learning model and is used to update the weights in deep learning. Test data (test) is data used to judge how well the trained model performs.

[0053] Figures 6(B) and 6(C) show examples of index values ​​derived for each principal component vector using the molecular performance obtained by the trained model 15B.

[0054] In the graph shown in Figure 6(B), "wb_pc_01" and "wb_pc_03" represent "PC1" and "PC3," respectively, and "coefficient" represents the index value. In the table shown in Figure 6(C), "wb_pc_01" to "wb_pc_05" represent "PC1" to "PC5," respectively. For simplicity, "PC6" to "PC10" are omitted here. As mentioned above, the index values ​​have both positive and negative correlations, and are expressed in a range of "-1.00" to "1.00" as an example. In other words, the higher the positive correlation, the closer the value is to "1.00," and the higher the negative correlation, the closer it is to "-11.00." For example, the index value for "PC1" is "-0.75," and the index value for "PC3" is "-0.59," both of which show a high negative correlation. When the indicator values ​​for "PC1" and "PC3" in Figure 6(C) are plotted on a graph, the result is the graph shown in Figure 6(B).

[0055] The identification unit 11F identifies principal component vectors based on the index values ​​derived by the index value derivation unit 11E. Specifically, the identification unit 11F identifies principal component vectors whose absolute values ​​of the index values ​​derived by the index value derivation unit 11E are greater than or equal to a threshold. The threshold is set to an appropriate value based on experiments or past knowledge. Specifically, for the index values ​​shown in Figure 6(C), for example, "0.5" is set. In this case, the absolute values ​​of the index values ​​of "PC1" and "PC3" ("0.75" and "0.59", respectively) are greater than or equal to the threshold, so among principal components PC1 to PC10, principal components PC1 and PC3 are identified as principal component vectors that contribute to the performance of the molecule to a high degree.

[0056] The output unit 11G outputs principal component power spectrum data, which is power spectrum data corresponding to the principal component vectors identified by the identification unit 11F. To distinguish it from the molecular power spectrum data described above, the power spectrum data corresponding to the principal component vectors is referred to as principal component power spectrum data. This principal component power spectrum data is obtained by the principal component analysis described above. The output unit 11G outputs the principal component power spectrum data to, for example, the display unit 32 of the user terminal 30.

[0057] Figure 6(D) is a graph showing an example of principal component power spectral data corresponding to principal component vectors. Here, as an example, the principal component power spectral data corresponding to the principal component vectors "PC1" to "PC3" are shown. In the graph of principal component power spectral data, the horizontal axis represents, for example, frequency, and the vertical axis represents spectral value.

[0058] The output unit 11G displays the principal component power spectrum data of "PC1" and "PC3," which have been identified as principal components that contribute significantly to the performance of the molecule, on the display unit 32 of the user terminal 30. This allows the user to interpret the shape of the principal components that contribute significantly to the performance of the molecule.

[0059] Figure 7(A) is a graph showing an example of power spectral data before changing the principal component scores of principal components PC1 and PC3. Figure 7(B) is a graph showing an example of power spectral data after changing the principal component scores of principal components PC1 and PC3. In Figures 7(A) and 7(B), the solid line represents the measured power spectral data, and the dotted line represents the reconstructed data obtained using the principal component vectors and principal component scores.

[0060] As shown in Figures 7(A) and 7(B), when the principal component scores of principal components PC1 and PC3, which contribute significantly to the performance of the molecule, are reduced, the power spectral data of the entire molecule, including principal components PC1 to PC10, also changes. In this case, molecules with atomic coordinates corresponding to the change in spectral shape can be interpreted as high-performance molecules.

[0061] Furthermore, it is desirable that the principal component scores, which are displayed as bars in Figures 7(A) and 7(B), be adjustable on the user terminal 30. For example, by changing the principal component scores as shown by arrows D1 and D2 in Figure 7(B), the waveform of the power spectrum data may be changed as shown by arrows E1 to E4. In this case, since the waveform of the power spectrum data changes in accordance with the changes and adjustments made to the principal component scores displayed as bars, the user can understand the meaning of the principal component scores.

[0062] On the other hand, the map generation unit 11H generates two-dimensional map data by projecting the principal component scores corresponding to each of the multiple principal component vectors identified by the identification unit 11F onto a two-dimensional plane as multiple plot points. In this case, the output unit 11G may display atomic coordinate images corresponding to the plot points of the two-dimensional map data on the display unit 32 of the user terminal 30, along with the two-dimensional map data.

[0063] Figure 8 shows an example of two-dimensional map data and atomic coordinate images according to this embodiment.

[0064] The two-dimensional map data shown in Figure 8 is data obtained by projecting the principal component scores corresponding to each of the principal components PC1 and PC3 identified by the identification unit 11F onto a two-dimensional plane as multiple plot points. The horizontal axis represents principal component PC1, and the vertical axis represents principal component PC3.

[0065] In Figure 8, the coordinate values ​​of a single plotted point in the 2D map data correspond to the principal component scores of principal components PC1 and PC3. As described above, the waveform of the power spectral data changes according to the principal component scores of principal components PC1 and PC3. In other words, since the power spectral data is determined according to the plotted point, an atomic coordinate image can be obtained by transforming this power spectral data. The transformation from power spectral data to atomic coordinate image can be performed, for example, by inverse Fourier transform or by using machine learning. This allows, for example, if a molecule with desired performance is assumed by the user, an atomic coordinate image of the unknown molecule that is expected to have that performance can be obtained without actually imaging the unknown molecule.

[0066] Figure 9 is a diagram illustrating the prediction using mol file structure information according to this embodiment.

[0067] As shown in Figure 9, in the case of prediction using mol file structure information according to this embodiment, molecules can be represented with fewer features compared to prediction using ECFP in the comparative example. Therefore, the computational load during molecule search is reduced.

[0068] Next, the operation of the server 10 according to this embodiment will be explained with reference to Figure 10.

[0069] Figure 10 is a flowchart showing an example of the processing flow by the information processing program 15A according to this embodiment.

[0070] First, when the server 10 is instructed to execute the molecular search process, the CPU 11 starts the information processing program 15A and executes the following processes.

[0071] In step S101 of Figure 10, the CPU 11 obtains a molecular data file from the user terminal 30, for example, the one shown in Figure 3(A) above.

[0072] In step S102, the CPU 11 creates an atomic coordinate image, as shown in Figure 3(B) above, as an example, from the molecular data file acquired in step S101.

[0073] In step S103, the CPU 11 performs a Fourier transform on the atomic coordinate image created in step S102, as shown in Figure 4 above, as an example, to generate power spectral data.

[0074] In step S104, the CPU 11 performs principal component analysis on the power spectral data generated in step S103, as shown in Figure 5 above, as an example, and derives principal component vectors and principal component scores from the power spectral data.

[0075] In step S105, the CPU 11, as an example, uses the trained model 15B to derive an index value that represents the degree of correlation between the principal component score derived in step S104 and the performance of the numerator.

[0076] In step S106, the CPU 11 identifies principal component vectors whose absolute value of the index value derived in step S105 is greater than or equal to a threshold, as shown in Figure 6(C) above, as an example.

[0077] In step S107, the CPU 11 outputs the principal component power spectrum data corresponding to the principal component vector identified in step S106 to the display unit 32 of the user terminal 30, as shown in Figure 6(D) above, as an example, and terminates the series of processes by the information processing program 15A.

[0078] As described above, this embodiment increases the likelihood of finding unknown molecules that meet the performance requirements compared to matching against a database.

[0079] Furthermore, compared to predictions using ECFP, molecules can be represented with fewer features. This reduces the computational load during molecule search.

[0080] In the above embodiment, the term "processor" refers to a processor in a broad sense, and includes general-purpose processors (e.g., CPU: Central Processing Unit, etc.) and dedicated processors (e.g., GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic device, etc.).

[0081] Furthermore, the operation of the processor in the above embodiment may not be performed by a single processor, but may be performed by multiple processors located in physically separate locations working together. Also, the order of the processor's operations is not limited to the order described in the above embodiment, but may be changed as appropriate.

[0082] The above describes an information processing device based on an embodiment. The embodiment may take the form of a program that causes a computer to execute the functions of the information processing device. The embodiment may also take the form of a non-temporary storage medium that is readable by a computer and stores these programs.

[0083] Furthermore, the configuration of the information processing device described in the above embodiment is merely an example, and may be modified as needed without departing from the main purpose.

[0084] Furthermore, the program processing flow described in the above embodiment is just one example, and unnecessary steps may be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0085] Furthermore, although the above embodiment describes a case in which the process according to the embodiment is realized by a software configuration using a computer by executing a program, the embodiment is not limited to this. The embodiment may also be realized by a hardware configuration or a combination of a hardware configuration and a software configuration. [Explanation of symbols]

[0086] 10. Server (Information Processing Device) 11A Acquisition Department 11B Creation Section 11C Spectrum Generation Unit 11D Principal component derivation part 11E Index Value Derivation Section 11F Specific section 11G output section 11H Map Generation Unit 15A Information Processing Program 15B Pre-trained model

Claims

1. A creation unit that creates an atomic coordinate image representing the atomic coordinates within a molecule, A spectrum generation unit generates power spectrum data by performing a Fourier transform on the created atomic coordinate image, A principal component derivation unit performs principal component analysis on the generated power spectrum data and derives principal component vectors representing the basis vectors of the power spectrum data and principal component scores representing the quantities in which the principal component vectors are included from the power spectrum data. An index value derivation unit inputs the derived principal component scores into a pre-trained model that has been machine-learned using training data to which principal component scores and molecular performance are associated, outputs the performance of the corresponding molecule from the trained model, and derives an index value that represents the degree of correlation between the principal component scores and the performance of the molecule. A specification unit that identifies principal component vectors correlated with the performance of the molecule based on the derived index values, An output unit that outputs principal component power spectrum data, which is power spectrum data corresponding to the identified principal component vector, Equipped with an information processing device.

2. The creation unit creates a two-dimensional or three-dimensional atomic coordinate image from a molecular data file that stores information representing the structure of the molecule. The information processing apparatus according to claim 1.

3. The specified unit identifies a plurality of principal component vectors, The system further includes a map generation unit that generates two-dimensional map data by projecting the principal component scores corresponding to each of the identified principal component vectors onto a two-dimensional plane as a plurality of plot points. The information processing apparatus according to claim 1.

4. The output unit displays the atomic coordinate images corresponding to the plot points of the two-dimensional map data on the display unit, along with the two-dimensional map data. The information processing apparatus according to claim 3.

5. Information processing device, Create an atomic coordinate image representing the atomic coordinates within a molecule. The created atomic coordinate image is Fourier transformed to generate power spectral data. Principal component analysis is performed on the generated power spectral data to derive principal component vectors representing the basis vectors of the power spectral data and principal component scores representing the quantities in which the principal component vectors are included. The derived principal component scores are input into a pre-trained model that has been machine-learned using training data in which principal component scores and molecular performance are associated. The trained model outputs the performance of the corresponding molecule, and an index value representing the degree of correlation between the principal component scores and the performance of the molecule is derived. Based on the derived index values, the principal component vectors that correlate with the performance of the molecule are identified. Outputs principal component power spectral data, which is power spectral data corresponding to the identified principal component vector. Information processing methods.

6. Create an atomic coordinate image representing the atomic coordinates within a molecule. The created atomic coordinate image is Fourier transformed to generate power spectral data. Principal component analysis is performed on the generated power spectrum data, and the power spectrum From the Tor data, the principal component vectors representing the basis vectors of the power spectrum data and the principal component scores representing the quantities in which the principal component vectors are included are derived. The derived principal component scores are input into a pre-trained model that has been machine-learned using training data in which principal component scores and molecular performance are associated. The trained model outputs the performance of the corresponding molecule, and an index value representing the degree of correlation between the principal component scores and the performance of the molecule is derived. Based on the derived index values, the principal component vectors that correlate with the performance of the molecule are identified. The process of outputting principal component power spectrum data, which is power spectrum data corresponding to the identified principal component vector, An information processing program designed to be executed by a computer.

Citation Information

Patent Citations

  • Information processing program, information processing device, and information processing method

    CN115132292A

  • Method and device for searching for new material

    JP2017091526A

  • Characteristic prediction system, characteristic prediction method and characteristic prediction program

    JP2022167397A

  • Data processing device and inference method

    JP2023008857A