Compound sequence selection program, compound sequence selection method, and information processing device

JPWO2024224580A5Pending Publication Date: 2026-01-28
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025516434
Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2025-11-10
Publication Date
2026-01-28

AI Technical Summary

Technical Problem

Conventional stepwise screening methods in drug discovery often miss target hit compounds with binding affinity and functionality for target molecules during intermediate rounds, leading to their omission in the final selection.

Method used

A compound sequence selection program and information processing device that calculates the frequency of appearance of partial structures across multiple screening rounds, selecting motifs whose frequency changes to satisfy predetermined conditions, thereby preventing omissions by identifying and retaining compound sequence data with promising motifs from intermediate rounds.

Benefits of technology

The solution effectively prevents the omission of compounds with binding affinity and functionality by selecting compound sequence data that includes motifs with changing frequency patterns, ensuring their inclusion in the final round and improving the reliability of compound selection in drug discovery.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A compound sequence selection program according an embodiment of the present invention causes a computer to perform a calculation process, an extraction process, and a selection process when screening compound sequence data of a plurality of compounds in a plurality of stages and selecting, from among the compound sequence data, compound sequence data for a compound that satisfies a prescribed condition with respect to a target molecule. The calculation process is for calculating, for each of a plurality of types of substructures in the plurality of compounds, the frequency of appearance thereof in compound sequence data extracted in each of the plurality of stages. The extraction process is for extracting, from among the plurality of types of substructures, a substructure for which changes in the calculated frequency of appearance thereof in the plurality of stages satisfy a prescribed condition. The selection process is for selecting, from among the compound sequence data, compound sequence data that includes the extracted substructure.
Need to check novelty before this filing date? Find Prior Art

Description

Compound sequence selection program, compound sequence selection method, and information processing device

[0001] An embodiment of the present invention relates to a compound sequence selection program, a compound sequence selection method, and an information processing device.

[0002] Conventionally, in drug discovery and the like, compounds are selected that have binding affinity to target molecules such as specific proteins or genes, and have functional molecular sequences that inhibit or activate the function of the target molecules.

[0003] For the selection of such compounds, there are conventional techniques such as phage display, aptamer, and mRNA display, which involve stepwise screening from a sequence library of a large number of compounds (approximately 10,000) such as peptides, to select compounds that have binding affinity and functionality for the target molecule.

[0004] International Publication No. 2015 / 076355

[0005] However, the above-mentioned conventional techniques have the problem that the desired hit compound (a compound with binding affinity and functionality for a target molecule) may be missed during the stepwise screening and may not remain until the final round.

[0006] In one aspect, an object of the present invention is to provide a compound sequence selection program, a compound sequence selection method, and an information processing device that can prevent compounds from being overlooked in stepwise screening.

[0007] In one embodiment, the compound sequence selection program causes a computer to execute a calculation process, a selection process, and a selection process when screening compound sequence data of a plurality of compounds in multiple stages to select compound sequence data of a compound that satisfies predetermined conditions for a target molecule from the compound sequence data. The calculation process calculates the occurrence frequency in the compound sequence data extracted at each of the multiple stages for each of multiple types of partial structures in the plurality of compounds. The selection process selects, from the multiple types of partial structures, a partial structure whose calculated occurrence frequency over the multiple stages satisfies predetermined conditions. The selection process selects compound sequence data including the selected partial structure from the compound sequence data.

[0008] According to one embodiment, it is possible to prevent compounds from being overlooked in stepwise screening.

[0009] FIG. 1 is an explanatory diagram illustrating an overview of stepwise screening according to an embodiment. FIG. 2 is a block diagram illustrating an example of the functional configuration of an information processing device according to an embodiment. FIG. 3 is a flowchart illustrating an example of operation of the information processing device according to an embodiment. FIG. 4 is an explanatory diagram illustrating an overview of motif development. FIG. 5 is an explanatory diagram illustrating an overview of motif analysis. FIG. 6 is a diagram showing a graph of the frequency of appearance of each motif in each round. FIG. 7 is a flowchart illustrating an example of operation of the information processing device according to an embodiment. FIG. 8 is an explanatory diagram illustrating motif merging. FIG. 9 is an explanatory diagram illustrating an example of a tree diagram. FIG. 10 is an explanatory diagram illustrating an example of a computer configuration.

[0010] Hereinafter, a compound sequence selection program, a compound sequence selection method, and an information processing device according to the embodiments will be described with reference to the drawings. Components having the same functions in the embodiments will be assigned the same reference numerals, and duplicated descriptions will be omitted. Note that the compound sequence selection program, the compound sequence selection method, and the information processing device described in the following embodiments are merely examples and do not limit the embodiments. Furthermore, the following embodiments may be combined as appropriate within a consistent range.

[0011] 1 is an explanatory diagram illustrating an overview of stepwise screening according to an embodiment. As shown in FIG. 1, the information processing device according to the embodiment performs screening in multiple stages (rounds Rd1 to Rd7) on a sequence library containing compound sequence data of multiple compounds, and selects compound sequence data of compounds that satisfy predetermined conditions (such as having binding affinity or functionality) for a target molecule.

[0012] Here, the screening performed by the information processing device according to the embodiment can be performed by a phage display method, an aptamer method, an mRNA display method, or the like. The screening may also be performed by automatic cleaning using a computer-controlled robot, or by computer screening in which a sequence library of compounds is created on a computer and binding with a target molecule is simulated. In the above screening, whether a compound satisfies predetermined conditions (such as binding affinity or functionality) for the target molecule is determined by whether a value indicating chemical binding strength or binding inhibition is equal to or greater than a predetermined threshold.

[0013] In the case of drug discovery, target molecules correspond to specific proteins (polypeptides) or genes (amino acid sequences) that cause disease, and sequence libraries contain multiple compounds (e.g., peptides) that are candidates for drug discovery.

[0014] In conventional stepwise screening, compound sequence data of compounds that may have binding affinity and functionality for the target molecule may not be extracted in the intermediate rounds Rd1 to Rd6, and may not remain in the screening results 42a up to the final round Rd7.

[0015] Therefore, the information processing device according to the embodiment focuses on multiple types of partial structures (motifs) in multiple compounds in a sequence library and calculates the frequency of occurrence of each motif in compound sequence data extracted in each of multiple rounds. Here, motifs are molecular sequences that constitute parts of various compounds, and appear one or more times within a compound. For example, motifs are sequences of bases, amino acids, etc. that are associated with functions such as binding to or inhibiting binding with a specific molecular sequence.

[0016] Next, the information processing device according to the embodiment selects, from among the multiple types of motifs, motifs whose calculated transitions in occurrence frequency over multiple rounds satisfy a predetermined condition. For example, the information processing device according to the embodiment selects motifs whose transitions in occurrence have decreased in a specific round. Next, the information processing device according to the embodiment selects compound sequence data including the selected motif from a sequence library.

[0017] Specifically, the information processing device according to the embodiment also sets compound sequence data including motifs whose occurrence trends have decreased in rounds Rd4 to Rd5 as the screening result 42b. This allows the information processing device according to the embodiment to select compound sequence data including motifs screened in intermediate rounds Rd1 to Rd6, thereby preventing compounds from being overlooked in stepwise screening.

[0018] 2 is a block diagram showing an example of the functional configuration of an information processing apparatus according to an embodiment. As shown in FIG. 2, the information processing apparatus 1 includes a communication unit 10, an input unit 20, a display unit 30, a storage unit 40, and a control unit 50.

[0019] The communication unit 10 executes data communication with an external device or the like via a network. For example, the communication unit 10 receives a sequence library 41 under the control of the control unit 50 and stores it in the storage unit 40. The communication unit 10 may also receive screening results 42 of each round of automatic cleaning under the control of the control unit 50 and store them in the storage unit 40. The input unit 20 accepts operations from a user. The display unit 30 displays the processing results of the control unit 50. For example, the display unit 30 displays and outputs a selected sequence 44 indicating compound sequence data of a compound selected from the sequence library 41 under the control of the control unit 50.

[0020] The storage unit 40 has a sequence library 41, screening results 42, motif strings 43, and selected sequences 44. For example, the storage unit 40 is realized by a memory or the like.

[0021] The sequence library 41 is library data including the structures (sequences) of a plurality of compounds (e.g., peptides). Specifically, in the sequence library 41, the sequences of bases and amino acids of each compound are represented by the abbreviations a (adenine), c (cytosine), g (guanine), ... Y (tyrosine), I (isoleucine), W (tryptophan), ...

[0022] The screening results 42 are data showing the results of screening in each round from the sequence library 41. For example, the screening results 42 show the number of times each compound in the sequence library 41 was extracted in each round.

[0023] The motif column 43 is data showing each of the multiple types of motifs as a column. Specifically, the motif column 43 shows each of the multiple types of motifs as a column, and each of the compounds in the sequence library 41 as a row, and indicates whether each compound in the sequence library 41 contains each of the multiple types of motifs with 1 (includes) / 0 (does not contain). Each of the multiple types of motifs is set in advance by a user or the like.

[0024] Furthermore, the notation of multiple types of motifs in the motif string 43 indicates the molecular sequence of the corresponding bases, amino acids, etc., such as [7, 2]W→S→C. Here, W→S→C is an abbreviation for the order of the bases, amino acids, etc., indicating that they are arranged in the order of W (tryptophan), S (serine), and C (cytosine). Furthermore, [7, 2] indicates the number of optional molecules (bases, amino acids, etc.) that can be placed between the molecular sequences arranged in W→S→C (W→S, S→C). For example, [7, 2]W→S→C indicates that seven optional molecules can be placed between W→S and two optional molecules can be placed between S→C.

[0025] The selected sequence 44 is compound sequence data of compounds selected from the sequence library 41. The selected sequence 44 includes compound sequence data selected from the sequence library 41 based on the screening results 42a up to the final round Rd7, as well as compound sequence data including motifs (e.g., screened motifs) whose transition in appearance frequency over multiple rounds satisfies predetermined conditions.

[0026] The control unit 50 has a screening processing unit 51, a motif development unit 52, a motif analysis unit 53, a sequence selection unit 54, and an output unit 55. For example, the control unit 50 is realized by a processor.

[0027] The screening processing unit 51 is a processing unit that performs screening in each round by a phage display method, an aptamer method, an mRNA display method, or the like based on the target molecule for the sequence library 41. The screening processing unit 51 stores the results of screening in each round in the memory unit 40 as screening results 42.

[0028] The motif expansion unit 52 is a processing unit that expands multiple types of motifs included in the motif string 43 and counts the presence or absence of appearance of each motif in the compound sequence data extracted in each of the multiple rounds based on the screening results 42. That is, the motif expansion unit 52 calculates the appearance frequency of each motif in the compound sequence data extracted in each of the multiple rounds based on the screening results 42. The motif expansion unit 52 stores the appearance frequency calculated for each motif in each of the multiple rounds in the motif string 43.

[0029] The motif analysis unit 53 analyzes the transition of the occurrence frequency over multiple rounds for each of the multiple types of motifs included in the motif sequence 43, based on the occurrence frequency calculated in each of the multiple rounds by the motif development unit 52. Then, the motif analysis unit 53 selects motifs whose transition of the occurrence frequency over the analyzed multiple rounds satisfies a predetermined condition.

[0030] Specifically, the motif analysis unit 53 uses a machine learning model based on Wide Learning (registered trademark: hereinafter referred to as WL) to statistically examine the changes in the frequency of occurrence of each of multiple types of motifs and combinations of motifs in each round.

[0031] For example, the motif analysis unit 53 hypothesizes each of a plurality of types of motifs and combinations of motifs. As an example, the motif analysis unit 53 hypothesizes each of motifs such as [7,2]W→S→C, [0,7]V→W→S..., and combinations of motifs such as [7,2]W→S→C^[0,7]V→W→S....

[0032] Next, the motif analysis unit 53 trains a machine learning model that outputs whether a hypothesis is correct (appears) or incorrect (does not appear) for each of the multiple motifs in each of the multiple rounds, based on the occurrence frequency calculated by the motif development unit 52 for each of the multiple motifs. Next, the motif analysis unit 53 compares the outputs of the machine learning model in successive rounds to statistically examine the transition in the occurrence frequency of each of the multiple motifs and motif combinations in each round. For example, the motif analysis unit 53 calculates statistical values ​​(such as chi-square values) for each of the multiple motifs and motif combinations in each round.

[0033] Next, the motif analysis unit 53 selects motifs whose appearance frequency in each round satisfies a predetermined condition (e.g., decrease from the previous round, increase from the previous round, etc.) for each of the multiple types of motifs and combinations of motifs in each round. For example, the motif analysis unit 53 selects motifs whose appearance in each round is unique based on statistical values ​​(e.g., chi-square values) for each of the multiple types of motifs and combinations of motifs in each round.

[0034] Based on the motif selected by the motif analysis unit 53, the sequence selection unit 54 selects compound sequence data including the selected motif from the compound sequence data in the sequence library 41. The sequence selection unit 54 stores the selected compound sequence data as selected sequences 44 in the storage unit 40 together with compound sequence data included in the screening results 42 remaining until the final round.

[0035] The output unit 55 is a processing unit that outputs the compound sequence data included in the selected sequence 44 to the display unit 30 or the like. When outputting the compound sequence data included in the selected sequence 44, the output unit 55 may output the compound sequence data selected from the sequence library 41 based on the motif selected by the motif analysis unit 53, and the compound sequence data included in the screening result 42 that remained until the final round, in a distinguished manner. This allows the user to easily distinguish between compound sequence data that remained until the final round in stepwise screening and compound sequence data that includes a motif screened in an intermediate round.

[0036] 3 is a flowchart showing an example of the operation of the information processing apparatus according to the embodiment. As shown in FIG. 3, when the process starts, the control unit 50 reads and acquires the sequence library 41 from the storage unit 40 (S1).

[0037] Next, the screening processing unit 51 performs a stepwise screening process on the sequence library 41 based on the target molecule (S2). Specifically, the screening processing unit 51 sets the number of rounds (K) to 1 (initial value) (S21).

[0038] Next, the screening processing unit 51 performs K rounds of screening (S22), increments K (S23), and then determines whether K<N (final round) is true (S24). If K<N is true (S24: Yes), the screening processing unit 51 returns to S22. If K<N is not true (S24: No), the screening processing unit 51 proceeds to S4.

[0039] In the stepwise screening process, the control unit 50 selects compound sequence data from the sequence library 41 that includes motifs whose transition in appearance frequency over multiple rounds up to the final round satisfies a predetermined condition (S3).

[0040] Specifically, the motif development unit 52 acquires the sequences (screening results 42) obtained from each round (S31). Next, the motif development unit 52 develops multiple types of motifs included in the motif string 43, and calculates the frequency of occurrence of each motif in the compound sequence data extracted in each of the multiple rounds based on the screening results 42 (S32).

[0041] 4 is an explanatory diagram illustrating an overview of motif development. As shown in Fig. 4, the motif development unit 52 compares the compound sequence data included in the screening results 42 of each round with each of the multiple motifs listed in the motif string 43. In this comparison, the motif development unit 52 counts the presence or absence of a motif by assigning a value of 1 if the compound sequence data contains a corresponding motif, and 0 if not.

[0042] As shown in the illustrated example, when comparing the first row "YIWDTGTFYLSRTCGSGSGS" of the compound sequence data included in the screening result 42 with the motif [7,2]W→S→C, there is an "S" seven positions after the "W" in the compound sequence data, and a "C" two positions after the "S" in the compound sequence data. Therefore, the motif expansion section 52 sets the value of motif [7,2]W→S→C to 1, indicating a match. Note that when the same motif appears multiple times in the compound sequence data, the motif expansion section 52 may set the value to the number of times it appears, rather than 1, indicating a match.

[0043] Returning to Figure 3, after S32, the motif analysis unit 53 learns each round using the above-mentioned WL based on the occurrence frequency calculated by the motif development unit 52 for each of the multiple types of motifs included in the motif sequence 43 in each of the multiple rounds (S33).

[0044] Next, the motif analysis unit 53 analyzes the transition in the frequency of occurrence of each of the multiple motifs and combinations of motifs in each round by comparing the output of the machine learning model in successive rounds based on the machine learning model in each round learned by WL. Next, based on the analysis results, the motif analysis unit 53 selects motifs whose transition in the frequency of occurrence of each of the multiple motifs in each round satisfies a predetermined condition (e.g., decrease from the previous round, increase from the previous round, etc.) (S34).

[0045] 5 is an explanatory diagram illustrating an overview of motif analysis. As shown in FIG. 5, the motif analysis unit 53 can classify multiple types of motifs into classes CL1 and CL2 in each round (Rd1 to Rd6) up to the final round (e.g., Rd7) by using the machine learning model for each round learned in WL. Specifically, class CL1 is a class of motifs that do not appear (motifs that do not bind to the target molecule), and class CL2 is a class of motifs that appear (motifs that bind to the target molecule).

[0046] The motif analysis unit 53 compares the outputs of such machine learning models in successive rounds to statistically analyze the transitions in the frequency of occurrence of each of the multiple types of motifs and combinations of motifs in each round. For example, the motif analysis unit 53 calculates statistical values ​​(such as chi-squared values) for each of the multiple types of motifs and combinations of motifs in each round. Based on the calculated statistical values ​​(such as chi-squared values), the motif analysis unit 53 selects motifs whose transitions in frequency of occurrence satisfy a predetermined condition (e.g., decrease from the previous round, increase from the previous round, unique occurrence, etc.).

[0047] By performing such statistical analysis, for example, by comparing classes CL1 and CL2 based on the machine learning models of rounds Rd5 and Rd6, the motif analysis unit 53 can obtain motifs that are abundant in round Rd5 but are rare in round Rd6. That is, the motif analysis unit 53 can obtain motifs that are abundant in compound sequence data that failed to bind in round Rd6 in the sequence library 41.

[0048] Returning to FIG. 3, after S34, the sequence selection unit 54 selects a sequence (compound-containing sequence data including the motif) corresponding to the selected motif from the compound sequence data in the sequence library 41 based on the motif selected by the motif analysis unit 53 (S35).

[0049] Next, the sequence selection unit 54 uses the selected motifs to cluster the motifs based on the tendency of changes in appearance frequency over multiple rounds (S36).

[0050] 6 is a diagram showing a graph of the frequency of occurrence of each motif in each round. As shown in FIG. 6, in the information processing device 1, a graph G1 of the frequency of occurrence of each of multiple types of motifs in each round can be obtained through analysis by the motif analysis unit 53. Based on this graph G1, the sequence selection unit 54 clusters motifs that have similar trends in the transition of their frequency of occurrence over multiple rounds.

[0051] For example, as shown by the oval portion of graph G1, the trends in the transition of the frequency of appearance over multiple rounds of motif 1 and motif 2 are similar. Therefore, the sequence selection unit 54 clusters motif 1 and motif 2.

[0052] 3, the sequence selection unit 54 selects, for each cluster obtained by clustering, a representative sequence from each motif in the cluster from the sequence library 41 (S37), and returns the process to S22. Specifically, the sequence selection unit 54 selects, from the sequence library 41, compound sequence data including each motif in the cluster.

[0053] After S2, the control unit 50 outputs the compound sequence data selected in S3 together with the compound sequence data included in the screening results 42 remaining until the final round (S4).

[0054] Specifically, the sequence selection unit 54 selects compound sequence data included in the screening results 42 remaining until the final round (S41), and defines these together with the compound sequence data selected in S3 as selected sequences 44. Next, the output unit 55 outputs the compound sequence data included in the selected sequences 44 to the display unit 30 or the like (S42), and the process ends.

[0055] In the above embodiment, screening for one target (target molecule) was exemplified, but the present invention may also be applied to a dual binder or the like that screens for sequences that have the property of being able to bind in common to multiple types of targets.

[0056] FIG. 7 is a flowchart showing an example of the operation of the information processing device 1 according to the embodiment, and shows an example of the operation of screening by Dual Binder when there are two targets (target molecules) A ​​and B.

[0057] As shown in FIG. 7, in the case of Dual Binder, the screening processing unit 51 reads and acquires the sequence library 41 for target A from the storage unit 40 (S1a), and performs stepwise screening processing on the sequence library 41 (S2a).

[0058] Similarly, the screening processing unit 51 reads and acquires the sequence library 41 for the target B from the storage unit 40 (S1b), and performs a stepwise screening process on the sequence library 41 (S2b).

[0059] Next, the control unit 50 selects compound sequence data from the sequence library 41 that includes motifs whose appearance frequency over multiple rounds up to the final round satisfies predetermined conditions in the stepwise screening process of targets A and B (S5).

[0060] Specifically, the motif development unit 52 develops multiple types of motifs contained in the motif string 43 (S51), and calculates the frequency of occurrence of each motif in the compound sequence data extracted in each of the multiple rounds based on the screening results 42 of targets A and B (S52).

[0061] Next, the motif analysis unit 53 uses the above-mentioned WL to perform motif analysis for each of the targets A and B (S52). Next, based on the results of the motif analysis for each of the targets A and B, the motif analysis unit 53 selects motifs that have common trends in the frequency of appearance of multiple types of motifs in the rounds (S53).

[0062] Next, the sequence selection unit 54 selects a sequence (compound-containing sequence data including the motif) corresponding to the selected motif from the compound sequence data in the sequence library 41 based on the motif selected by the motif analysis unit 53 (S54).

[0063] After S5, the output unit 55 outputs the compound sequence data selected in S5 together with the compound sequence data included in the screening results 42 remaining until the final round (S6). Here, the output unit 55 identifies co-occurring motifs in the compound sequence data selected in S5. Next, the control unit 50 creates and outputs a tree diagram for sequentially identifying motifs other than the co-occurring motif included in the selected compound sequence data based on the co-occurring motifs.

[0064] Specifically, the output unit 55 identifies co-occurring motifs in the compound sequence data selected in S5, and then merges the identified motifs (S61).

[0065] 8 is an explanatory diagram illustrating motif merging. As shown in FIG. 8, the output unit 55 identifies motifs that are common to a plurality of compound sequences and that appear most frequently as co-occurring motifs based on the occurrence frequencies shown in the motif string 43 for each compound sequence included in the screening result 42. In the illustrated example, the output unit 55 identifies the three-residue motifs [1,10] Y→W→C, [7,2] W→S→C, and [2,0] T→L→S, which correspond to sequences 1, 2, 3, and 4 (highest number of occurrences), as co-occurring motifs.

[0066] Next, the output unit 55 merges the identified motifs so that the sequence configuration is consistent. Specifically, the output unit 55 arranges the three-residue motifs [1,10] Y→W→C, [7,2] W→S→C, and [2,0] T→L→S in parallel so that the sequence configuration is consistent, and merges them to obtain a putative motif of "Y-W---T--LS--C."

[0067] Returning to FIG. 7, the output unit 55 creates a tree diagram for sequentially identifying motifs other than the co-occurring motifs contained in the selected compound sequence data, based on the estimated motif obtained by merging the co-occurring motifs (S62).

[0068] 9 is an explanatory diagram illustrating an example of a tree diagram. In FIG. 9, sequences 1, 2, and 3 in the selected compound sequence data are assumed to contain no motif other than the putative motif obtained by merging co-occurring motifs. Sequence 4 in the selected compound sequence data is assumed to contain motif A in addition to the putative motif. Sequence 5 in the selected compound sequence data is assumed to contain motif B in addition to the putative motif.

[0069] 9, the output unit 55 determines that sequences 1, 2, and 3, which contain a putative motif obtained by merging co-occurring motifs, are relevant (functional) among the selected compound sequence data. Here, for sequence 4, which contains motif A other than the putative motif obtained by merging co-occurring motifs, motif A can be identified by carrying out a functionality test. Furthermore, for sequence 5, which contains motif B other than the putative motif obtained by merging co-occurring motifs, motif B can be identified by carrying out a functionality test.

[0070] Therefore, the output unit 55 creates a branched tree diagram Tr so that motif A can be identified if functionality is present in the functionality test of sequence 4, and motif B can be identified if functionality is present in the functionality test of sequence 5.

[0071] Returning to FIG. 7, the output unit 55 outputs the created tree diagram Tr together with the compound sequence data (S63), and the process ends.

[0072] As described above, when the information processing device 1 selects compound sequence data of a compound that satisfies predetermined conditions for a target molecule (target) from the sequence library 41 by screening the sequence library 41 in multiple rounds (stages), the control unit 50 executes a calculation process, a selection process, and a selection process. In the calculation process, the occurrence frequency in the compound sequence data extracted in each of the multiple rounds is calculated for each of multiple types of motifs (substructures) in multiple compounds. In the selection process, a motif is selected from the multiple types of motifs, the change in the calculated occurrence frequency over the multiple rounds satisfies predetermined conditions. In the selection process, compound sequence data including the selected motif is selected from the sequence library 41.

[0073] This allows the information processing device 1 to select compound sequence data containing motifs whose appearance frequency over multiple rounds satisfies a predetermined condition, and to select compound sequence data according to the transition of the selection tendency of compounds through stepwise screening. Therefore, the information processing device 1 can prevent compounds from being overlooked in the selection through stepwise screening.

[0074] Furthermore, the information processing device 1 selects motifs whose frequency of occurrence has decreased in a specific round among multiple rounds. This allows the information processing device 1 to select compound sequence data including motifs whose frequency of occurrence tends to decrease due to oversight in multiple rounds. Therefore, the information processing device 1 can more reliably prevent compounds from being overlooked in the selection of compounds through stepwise screening.

[0075] Furthermore, the information processing device 1 clusters the selected motifs based on their transitional tendency over multiple rounds, and selects, for each cluster, compound sequence data containing the motif in the cluster. This allows the information processing device 1 to select compound sequence data containing the clustered motif in accordance with the transition of the clustered motif in stepwise screening.

[0076] Furthermore, the information processing device 1 calculates the frequency of occurrence of each of multiple target molecules (targets A and B) in compound sequence data extracted in each of multiple rounds of screening. The information processing device 1 selects, for each of the multiple target molecules, a motif whose calculated transition in occurrence frequency over multiple rounds satisfies a predetermined condition. This allows the information processing device 1 to select, for each of the multiple target molecules, compound sequence data including a motif whose transition in occurrence frequency over multiple rounds satisfies a predetermined condition. Therefore, the information processing device 1 can prevent compound sequence data with properties that can bind to multiple types of target molecules from being overlooked through stepwise screening.

[0077] Furthermore, the information processing device 1 identifies co-occurring motifs in the selected compound sequence data, and outputs a tree diagram Tr for sequentially identifying motifs other than the co-occurring motifs contained in the selected compound sequence data based on the co-occurring motifs. This allows the information processing device 1 to easily identify motifs other than the co-occurring motifs that have binding and functionality.

[0078] Note that the components of each device shown in the figure do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.

[0079] Furthermore, the various processing functions of the screening processing unit 51, motif development unit 52, motif analysis unit 53, sequence selection unit 54, and output unit 55 performed by the control unit 50 of the information processing device 1 may be executed in whole or in part on a CPU (or a microcomputer such as an MPU or MCU (Micro Controller Unit)). Needless to say, the various processing functions may be executed in whole or in part on a program analyzed and executed by the CPU (or a microcomputer such as an MPU or MCU), or on hardware using wired logic. Furthermore, the various processing functions performed by the information processing device 1 may be executed by multiple computers working together using cloud computing.

[0080] The various processes described in the above embodiments can be realized by executing a program prepared in advance on a computer. Therefore, an example of a computer configuration (hardware) that executes a program having the same functions as those of the above embodiments will be described below. Fig. 10 is an explanatory diagram illustrating an example of the computer configuration.

[0081] 10 , computer 200 has a CPU 201 that executes various types of arithmetic processing, an input device 202 that accepts data input, a monitor 203, and a speaker 204. Computer 200 also has a medium reading device 205 that reads programs and the like from a storage medium, an interface device 206 for connecting with various devices, and a communication device 207 for connecting with external devices via wired or wireless communication. Computer 200 also has a RAM 208 that temporarily stores various types of information, and a hard disk drive 209. Each unit (201 to 209) within computer 200 is connected to a bus 210.

[0082] The hard disk drive 209 stores a program 211 for executing various processes in the functional configuration (e.g., the screening processing unit 51, the motif development unit 52, the motif analysis unit 53, the sequence selection unit 54, and the output unit 55) described in the above embodiment. The hard disk drive 209 also stores various data 212 referenced by the program 211. The input device 202, for example, accepts input of operation information from an operator. The monitor 203, for example, displays various screens operated by the operator. The interface device 206 is connected to, for example, a printing device. The communication device 207 is connected to a communication network such as a LAN (Local Area Network) and exchanges various information with external devices via the communication network.

[0083] The CPU 201 reads out the program 211 stored in the hard disk drive 209, expands it in the RAM 208, and executes it to perform various processes related to the above-mentioned functional configuration (e.g., the screening processing unit 51, the motif development unit 52, the motif analysis unit 53, the sequence selection unit 54, and the output unit 55). The program 211 does not have to be stored in the hard disk drive 209. For example, the program 211 stored in a storage medium readable by the computer 200 may be read out and executed. Examples of the storage medium readable by the computer 200 include portable storage media such as CD-ROMs, DVD disks, and USB (Universal Serial Bus) memories, semiconductor memories such as flash memories, and hard disk drives. The program 211 may also be stored in a device connected to a public line, the Internet, a LAN, or the like, and the computer 200 may read out and execute the program 211 from the device.

[0084] DESCRIPTION OF SYMBOLS 1...Information processing device 10...Communication unit 20...Input unit 30...Display unit 40...Storage unit 41...Sequence library 42, 42a, 42b...Screening results 43...Motif string 44...Selected sequence 50...Control unit 51...Screening processing unit 52...Motif development unit 53...Motif analysis unit 54...Sequence selection unit 55...Output unit 200...Computer 201...CPU 202...Input device 203...Monitor 204...Speaker 205...Media reading device 206...Interface device 207...Communication device 208...RAM 209...Hard disk device 210...Bus 211...Program 212...Various data A, B...Target CL1, CL2...Class G1...Graph Rd1 to Rd7...Round Tr...Tree diagram

Claims

1. When screening compound sequence data of a plurality of compounds in a plurality of stages to select compound sequence data of a compound that satisfies a predetermined condition for a target molecule from the compound sequence data, calculating an appearance frequency in the compound sequence data extracted in each of the plurality of stages for each of a plurality of types of partial structures in the plurality of compounds; selecting a substructure from the plurality of types of substructures, the substructure having the calculated transition of the occurrence frequency over the plurality of stages satisfying a predetermined condition; selecting compound sequence data including the selected partial structure from the compound sequence data; A compound sequence selection program characterized by causing a computer to execute processing.

2. the selecting process selects the partial structure whose appearance frequency has decreased at a specific stage among the plurality of stages; 2. The compound sequence selection program according to claim 1.

3. the selecting process includes clustering the selected partial structures based on the transition tendency, and selecting, for each cluster, compound sequence data including the partial structure in the cluster; 2. The compound sequence selection program according to claim 1.

4. the calculating step calculates an appearance frequency of each of the plurality of target molecules in compound sequence data extracted at each of the plurality of stages of the screening; the selecting process selects, for each of the plurality of target molecules, a partial structure in which the calculated transition of the occurrence frequency over the plurality of stages satisfies a predetermined condition; 2. The compound sequence selection program according to claim 1.

5. further causing the computer to execute a process of identifying co-occurring partial structures in the selected compound sequence data, and outputting a tree diagram for sequentially identifying partial structures other than the co-occurring partial structures contained in the selected compound sequence data based on the co-occurring partial structures; 2. The compound sequence selection program according to claim 1.

6. When screening compound sequence data of a plurality of compounds in a plurality of stages to select compound sequence data of a compound that satisfies a predetermined condition for a target molecule from the compound sequence data, calculating an appearance frequency in the compound sequence data extracted in each of the plurality of stages for each of a plurality of types of partial structures in the plurality of compounds; selecting a substructure from the plurality of types of substructures, the substructure having the calculated transition of the occurrence frequency over the plurality of stages satisfying a predetermined condition; selecting compound sequence data including the selected partial structure from the compound sequence data; A compound sequence selection method characterized in that the processing is carried out by a computer.

7. When screening compound sequence data of a plurality of compounds in a plurality of stages to select compound sequence data of a compound that satisfies a predetermined condition for a target molecule from the compound sequence data, calculating an appearance frequency in the compound sequence data extracted in each of the plurality of stages for each of a plurality of types of partial structures in the plurality of compounds; selecting a substructure from the plurality of types of substructures, the substructure having the calculated transition of the occurrence frequency over the plurality of stages satisfying a predetermined condition; selecting compound sequence data including the selected partial structure from the compound sequence data; An information processing device comprising: a control unit that executes processing.