Method for calculating prediction score used for forecasting onset of specific infectious disease, generation method, calculation apparatus, and computer program

The proposed calculation method addresses the time constraint of short-read analysis by identifying high-reliability bacterial species and calculating prediction scores quickly, enhancing the timeliness and accuracy of infectious disease predictions.

JP2025086131APending Publication Date: 2025-06-06FUKUOKA UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023199976
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Short-read analysis methods for predicting the onset of infectious diseases require a significant amount of time (two to three days) to produce results, making it impractical for timely predictions.

Method used

A calculation method that involves obtaining gene sequences from a subject's sample, randomly selecting gene sequences to form sequence groups, identifying bacterial species using a pre-generated database, determining high-reliability bacterial species, and calculating a prediction score using species-specific correction values.

Benefits of technology

This method enables rapid calculation of prediction scores for infectious disease onset, improving the timeliness and accuracy of predictions compared to traditional short-read analysis methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025086131000001_ABST
    Figure 2025086131000001_ABST
Patent Text Reader

Abstract

To easily and accurately predict the onset of chorioamnionitis.SOLUTION: Provided is a method for calculating a prediction score used to predict the onset of a specific infectious disease in a subject, the method comprising: obtaining sequencing results as sequence information, including a plurality of genetic sequences obtained using samples collected from the subject; randomly selecting a predetermined number of genetic sequences included in the sequence information; generating a plurality of sets of sequence groups; identifying, for each of the generated sequence groups, a strain contained in the sequence group by referring to a pre-generated strain database that associates genetic sequences with strains; determining at least one strain included in all of the sequence groups as a high-confidence strain likely to influence the infectious disease; and calculating a prediction score using a strain-specific correction value predetermined for the strain determined as a high-confidence strain.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a calculation method, generation method, calculation device, and computer program for calculating a prediction score used to predict the onset of a specific infectious disease. [Background technology]

[0002] The results of metagenomic analysis may be used to predict the onset of a particular infectious disease. The accuracy of infectious disease onset prediction using metagenomic analysis depends on the nature of the metagenomic analysis.

[0003] For example, metagenomic analysis widely uses short-read analysis, which reads short fragments of about 300 bp in length, and allows the analysis of samples from a large number of subjects (e.g., 300 or more) at once. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] International Publication No. 2021 / 112236 Summary of the Invention [Problem to be solved by the invention]

[0005] However, with short-read analysis, analysis must be started after samples from a certain number of subjects have been collected, and it takes two to three days from the start of the analysis to obtaining results, meaning that there is an issue with this method in that results cannot be obtained in a short period of time after samples are taken from the subjects.

[0006] In view of the above, the present invention relates to a calculation method, a generation method, a calculation device, and a computer program for calculating a prediction score used to predict the onset of a specific infectious disease. [Means for solving the problem]

[0007] The calculation method according to the present disclosure is a calculation method for calculating a prediction score used to predict the onset of a specific infectious disease in a subject, which includes obtaining a sequence result including a plurality of gene sequences obtained using a sample collected from the subject as sequence information, randomly selecting a predetermined number of gene sequences included in the sequence information, generating a plurality of sequence groups, referring to a pre-generated bacterial species database that associates gene sequences with bacterial species, identifying the bacterial species included in each of the generated sequence groups, determining at least one or more bacterial species included in all of the sequence groups as high-reliability bacterial species that are likely to affect the infectious disease, and calculating a prediction score using a value that is predetermined for the bacterial species determined as the high-reliability bacterial species. Effect of the Invention

[0008] According to the calculation method, calculation device, and computer program disclosed herein, it is possible to calculate a prediction score used to predict the onset of a specific infectious disease. [Brief description of the drawings]

[0009] [Figure 1] FIG. 2 is a block diagram showing a configuration of a calculation device. [Diagram 2] FIG. 1 is a schematic diagram showing an overview of determination of high-confidence bacterial species. [Diagram 3] FIG. 1 is a schematic diagram illustrating calculation of a prediction score using CAM and NCAM values. [Figure 4] 13 is a flowchart illustrating a calculation process executed by the prediction device. [Diagram 5] 5 is a flowchart illustrating a calculation process executed by the prediction device following FIG. 4. [Figure 6] FIG. 2 is a block diagram showing a configuration of a generating device. [Figure 7] FIG. 2 is a schematic diagram illustrating learning data used in the generating device. [Figure 8] FIG. 1 is a schematic diagram illustrating the execution of a random forest and Volta analysis using training data. [Figure 9]FIG. 1 is a schematic diagram illustrating the determination of predicted related bacterial species. [Figure 10] FIG. 1 is a diagram illustrating the determination of specific correction values ​​for each bacterial species. [Figure 11] FIG. 1 is a schematic diagram illustrating the determination of positively or negatively associated predicted species in each subgroup. [Figure 12] FIG. 13 is a schematic diagram for explaining the determination of a positively associated predicted species or a negatively associated predicted species in the entire learning data. [Figure 13] 11 is a flowchart illustrating a process executed by the generating device. [Figure 14] FIG. 1 is a schematic diagram illustrating learning data used in an experiment. [Figure 15] This is a graph showing the results of a comparison of different read numbers using the results of long read analysis. [Figure 16] 13 is a graph showing the results of a comparison between the use of a specific correction value and the absence of the use of the specific correction value. [Figure 17] This is a graph showing the results of a comparison of different read numbers using the results of long read analysis. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] The calculation method, generation method, calculation device, and computer program according to the present disclosure will be described below with reference to the drawings. The calculation method, calculation device, and computer program according to the present disclosure calculate a prediction score used to predict the degree of possibility of a subject developing a specific infectious disease, such as chorioamnionitis. The calculation method and calculation device calculate the prediction score using an analysis result including a gene sequence obtained from a sample of the subject. Specifically, the gene sequence obtained from the sample used here may be the gene sequence of a bacterium contained in the sample of the subject. For example, the calculated prediction score is used to determine that the higher the value, the higher the possibility of developing a specific infectious disease.

[0011] In the following description, a "subject" is a person who is the subject of prediction of onset of a specific infectious disease. Specifically, chorioamnionitis is used as an example of a specific infectious disease. Therefore, the following description will be given using an example of prediction of the likelihood of onset of chorioamnionitis. However, the present disclosure is not limited to chorioamnionitis, and the calculation method, generation method, calculation device, and computer program can be applied to any disease associated with bacteria.

[0012] The "sample" refers to a subject's secretions, etc., including resident bacterial flora. Specifically, the sample may be vaginal secretions, feces, skin, amniotic fluid, saliva, etc. By using vaginal secretions, etc. as a sample, non-invasive prediction is possible.

[0013] [Calculation device] A calculation device 1 according to the embodiment will be described with reference to Fig. 1. The calculation device 1 is realized by an information processing device including an arithmetic circuit 11, a storage device 12, an input device 13, an output device 14, and a communication device 15.

[0014] The arithmetic circuit 11 is a controller that controls the entire calculation device 1. The arithmetic circuit 11 can execute various processes related to the calculation of the predicted score by executing the calculation program P1 stored in the storage device 12. The arithmetic circuit 11 may be various processors such as a CPU, an MPU, a GPU, an FPGA, a DSP, an ASIC, or a dedicated hardware circuit.

[0015] The storage device 12 is a storage medium for recording various information. The storage device 12 is realized, for example, by a RAM, a ROM, a flash memory, an SSD (Solid State Drive), a hard disk drive, or other storage devices, or by an appropriate combination of these. The storage device 12 stores a calculation program P1, which is a computer program executed by the arithmetic circuit 11, as well as a bacterial species database 121, a related predicted bacterial species database 122, correction value data 123, and the like, which are used in processes such as calculation of a prediction score.

[0016] The bacterial species database 121 includes data that associates the type of bacterial species with a gene sequence that indicates the bacterial species. For this bacterial species database 121, a generally available database as an open source can be used.

[0017] The associated predicted species database 122 includes positively associated predicted species (CAM-associated predicted species) that are bacterial species associated with infectious diseases, specifically, positive results for chorioamnionitis, and negatively associated predicted species (NCAM-associated predicted species) that are bacterial species associated with negative results for infectious diseases. Specifically, it includes a distinction as to whether each species is a positively associated predicted species or a negatively associated predicted species. For example, this associated predicted species database 122 is generated by a generating device 2 described below using the analysis results of samples obtained from subjects known to be positive or negative for infectious diseases.

[0018] The correction value data 123 is data including a species-specific correction value used to calculate a prediction score predetermined for each species. For example, the correction value data 123 is generated by a generating device 2 described later using the analysis results of a sample obtained from a subject who is known to be positive or negative for an infectious disease.

[0019] The input device 13 is an input means such as an operation button, a keyboard, a mouse, a touch panel, a microphone, etc., which is used to input a request for calculation of a prediction score and data. The output device 14 is an output means such as a display, a speaker, etc., which is used to output the calculation result of the prediction score and data.

[0020] The communication device 15 is a communication means for enabling data communication with an external device (not shown). The data communication may be performed wirelessly and / or according to known wired communication standards. For example, wired data communication is performed by using a communication controller of a semiconductor integrated circuit that operates in compliance with the Ethernet (registered trademark) standard and / or the USB (registered trademark) standard as the communication device 15. Wireless data communication is performed by using a communication controller of a semiconductor integrated circuit that operates in compliance with the IEEE802.11 standard for LAN (Local Area Network) and / or the fourth / fifth generation mobile communication system, so-called 4G / 5G, for mobile communication, as the communication device 15.

[0021] The arithmetic circuit 11 acquires a sequence result including gene sequences of a plurality of bacteria obtained using a sample collected from the subject as sequence information of the subject. Specifically, the sequence result may be a group of data of base sequences acquired in advance from a sample collected from a pregnant woman who is the subject. The arithmetic circuit 11 stores the acquired sequence result in the storage device 12. The sequence result of the gene sequence used here may be a result obtained using a long-read method. In the long-read analysis, for example, a base sequence exceeding 100 Kbp can be read. In this example, the sequence result is a long-read sequence result obtained using a sequencer of Oxford Nanopore Technologies, Inc. Compared to the short-read analysis, the long-read analysis can obtain a result within one day (several hours). In addition, since the cost of the long-read analysis is lower than that of the short-read analysis, the long-read analysis can be performed alone. On the other hand, one of the problems with the long-read analysis is that the error rate is relatively high compared to the short-read analysis, but the arithmetic circuit 21 of the present disclosure can obtain a highly accurate prediction score by selecting a highly reliable bacterial species and calculating a prediction score.

[0022] The arithmetic circuit 11 randomly selects a predetermined number of gene sequences included in the sequence information of the sequencing result, and generates a plurality of sequence groups. In the example shown in FIG. 2, the arithmetic circuit 11 generates 100 sequence groups from the target sequencing result. Each sequence group includes a predetermined number of randomly selected reads, for example, 5,000 reads. The arithmetic circuit 11 stores data indicating the generated plurality of sequence groups in the storage device 12. The arithmetic circuit 11 can use, for example, stratified random sampling to generate the sequence groups.

[0023] The arithmetic circuit 11 refers to a bacterial species database 121 that associates previously generated gene sequences with the types of bacterial species, and identifies the bacterial species contained in each generated sequence group. In the example shown in Fig. 2, bacterial species such as Burkholderia plantarii, Paenibacillus peoriae, and Lactobacillus casei are identified from sequence group 1. In the example shown in Fig. 1, the arithmetic circuit 11 reads out and uses the bacterial species database 121 stored in the storage device 12, but is not limited to this. For example, a bacterial species database stored in an external storage device (not shown) may be read out and used.

[0024] The arithmetic circuit 11 determines the bacterial species contained in all sequence groups as high-reliability bacterial species that are likely to affect infectious diseases. In the present disclosure, a "high-reliability bacterial species" is a bacterial species used to predict whether or not a subject is likely to develop chorioamnionitis. In the example of FIG. 2, the arithmetic circuit 11 determines the bacterial species identified in all 100 sequence groups as high-reliability bacterial species. Specifically, in the example shown in FIG. 2, bacterial species such as Candidatus Solibacter usitatus and Lactobacillus casei are determined as high-reliability bacterial species.

[0025] The arithmetic circuit 11 calculates the predicted score using a species-specific correction value (correction value data 123) that is predetermined for the species determined to be a highly reliable species. The calculation of the predicted score by the arithmetic circuit 11 will be specifically described below.

[0026] The arithmetic circuit 11 refers to the associated predicted species database 122 and determines whether each of the high-reliability species corresponds to either a positive associated predicted species or a negative associated predicted species. The arithmetic circuit 11 also determines, as a CAM value, the sum of species-specific correction values ​​(correction value data 123) that are predetermined for one or more species determined as a positive associated predicted species. The arithmetic circuit 11 also determines, as an NCAM value, the sum of species-specific correction values ​​(correction value data 123) that are predetermined for one or more species determined as a negative associated predicted species. The arithmetic circuit 11 then calculates a prediction score using the difference between the CAM value and the NCAM value.

[0027] Specifically, for each of the high-reliability bacterial species, when it is a positive-associated predicted bacterial species, the calculation circuit 11 adds the value (species-specific correction value) obtained for each species as shown in FIG. 3 to obtain a CAM value. Also, when the high-reliability bacterial species is a negative-associated predicted bacterial species, the calculation circuit 11 adds the value (species-specific correction value) obtained for each species as shown in FIG. 3 to obtain an NCAM value. The specific correction value of each species used here is stored in the storage device as correction value data 123. This correction value data 123 is generated, for example, by the generating device 2 described later using FIG. 6, as described later using FIG. 10. Furthermore, the calculation circuit 11 calculates a value obtained by subtracting the NCAM value from the CAM value as a prediction score, as shown in FIG. 3.

[0028] 《Calculation process of predicted score》 4 and 5, the process of calculating a predicted score in the calculation device 1 will be described. First, as shown in Fig. 4, the arithmetic circuit 11 acquires a target sequence result (S001).

[0029] Thereafter, the arithmetic circuit 11 generates a predetermined number of sequence groups from the sequence result acquired in step S001 (S002). In the example of the present disclosure, the arithmetic circuit 11 generates 100 sequence groups including 5000 randomly selected sequences as described above with reference to Fig. 2. The arithmetic circuit 11 also stores sequence group data indicating each of the generated sequence groups in the storage device 12.

[0030] Next, the arithmetic circuit 11 selects one sequence group from the plurality of sequence groups generated in step S002 (S003).

[0031] Next, the arithmetic circuit 11 performs a quality check on the sequence group selected in step S003 (S004). The arithmetic circuit 11 removes low-quality sequences, such as so-called barcode sequences and adapter sequences. The arithmetic circuit 11 performs the quality check by using, for example, a general-purpose program.

[0032] The arithmetic circuit 11 also refers to the bacterial species database 121 stored in the storage device 12, and identifies the bacterial species included in the sequence group selected in step S13 (S005). In the example of the present disclosure, the bacterial species are identified for all of sequence groups 1 to 100.

[0033] The arithmetic circuit 11 repeats the processes of steps S003 to S006 until the identification of the bacterial species is completed for all sequence groups. When the identification of the bacterial species is completed for all sequence groups (YES in S006), the arithmetic circuit 11 determines the bacterial species identified in a sequence group with a predetermined high reliability as a high reliability bacterial species (S007). For example, the bacterial species identified in all (100%) sequence groups may be determined as a high reliability bacterial species. Also, for example, the bacterial species identified in 98% sequence groups may be determined as a high reliability bacterial species. In the example of the present disclosure, the bacterial species identified in 100 sequence groups is determined as a high reliability bacterial species.

[0034] Next, as shown in the flowchart of FIG. 5, the arithmetic circuit 11 refers to the associated predicted species database 122 to determine, for each high-confidence species, whether it is a positive associated predicted species that influences positive results, or a negative associated predicted species that influences negative results (S008).

[0035] Furthermore, the arithmetic circuit 11 adds a predetermined species-specific correction value to each positively associated predicted species to calculate a CAM value (S009).

[0036] Furthermore, the arithmetic circuit 11 adds a predetermined species-specific correction value to each negative-associated predicted species to calculate the NCAM value (S010).

[0037] Then, the arithmetic circuit 11 subtracts the sum of the species-specific corrected values ​​(NCAM value) of each negatively associated bacterial species calculated in step S010 from the sum of the species-specific corrected values ​​(CAM value) of each positively associated bacterial species calculated in step S010 to calculate a prediction score (S011).

[0038] In this manner, the calculation device 1 of the present disclosure can use the sequencing results to calculate a prediction score used to predict the onset of a target infectious disease. In this case, the calculation device 1 of the present disclosure selects a highly reliable bacterial species, calculates the CAM value and the NCAM value using a specific correction value determined for each bacterial species, and calculates a prediction score. This makes it possible to obtain a highly accurate prediction score even when using the sequencing results of the long-read analysis of the present disclosure, which is generally considered to have a high error rate.

[0039] <Generation device> The generating device 2 that generates the related predicted bacterial species database 122 and the correction value data 123 used in the calculation device 1 will be described with reference to FIG. 6. Note that, here, an example in which the related predicted bacterial species database 122 and the correction value data 123 are generated in the generating device 2 will be described, but the related predicted bacterial species database 122 may be generated by other information processing devices such as the calculation device 1. The generating device 2 is realized by an information processing device including a calculation circuit 21, a storage device 22, an input device 23, an output device 24, and a communication device 25. Specifically, the generating program P2 stored in the storage device 22 is read and executed, thereby performing each process as the generating device. The calculation circuit 21, the storage device 22, the input device 23, the output device 24, and the communication device 25 are respectively realized by the same specific means as the calculation circuit 11, the storage device 12, the input device 13, the output device 14, and the communication device 15 described above with reference to FIG. 1.

[0040] The arithmetic circuit 21 acquires learning data 221 including a pair of a label indicating whether a specific infectious disease of a plurality of subjects is positive or negative, and a bacterial species identified from a sequence result including a plurality of gene sequences obtained using a sample collected from each subject. For example, the arithmetic circuit 21 reads out and uses the learning data 221 stored in advance in the storage device 22 as shown in FIG. 6. Alternatively, the arithmetic circuit 21 may access an external storage device (not shown) to read out and use the learning data 221. As shown in FIG. 7, the learning data 221 is labeled with positive or negative obtained as a result of a diagnosis of a plurality of subjects in the past, and is associated with one or more bacterial species identified in the sequence result of a sample collected from each subject. Note that the number of positive subjects and the number of negative subjects do not need to be the same in the learning data 221.

[0041] The arithmetic circuit 21 determines each bacterial species associated with all subjects in the learning data 221 as a highly reliable bacterial species that is a bacterial species that is highly likely to affect infectious diseases. If the learning data 211 includes bacterial species analyzed from samples of 100 subjects, the bacterial species analyzed from the samples of all 100 subjects are determined to be highly reliable bacterial species that are highly likely to affect infectious diseases.

[0042] The arithmetic circuit 21 generates a plurality of subgroups randomly selected from both a group of bacterial species labeled with a positive label and a group of a plurality of bacterial species labeled with a negative label in the learning data 221. For example, a stratified random sampling method is used to generate the subgroups. As shown in FIG. 7, each subgroup includes a set of bacterial species identified with the labels of a plurality of positive subjects and a set of bacterial species identified with the labels of a plurality of negative subjects. Note that one label and bacterial species set may be included in a plurality of subgroups. For example, even if there are 56 sets of labels and bacterial species, 100 groups may be generated with different combinations. Here, the number of positive label and bacterial species sets included in each subgroup is the same. Also, the number of negative label and bacterial species sets included in each subgroup is the same. On the other hand, the number of positive label and bacterial species sets included in each subgroup may be different from the number of negative label and bacterial species sets.

[0043] The arithmetic circuit 21 determines the predicted bacterial species that are predicted to have influenced the positive or negative judgment of the infectious disease using a machine learning algorithm that determines the importance of the determined high reliability bacterial species when dividing each subgroup into two groups, positive or negative. As shown in FIG. 8, the machine learning algorithm used here is an algorithm that determines the positive specificity and negative specificity of each bacterial species using Random Forest and Boruta analysis. As shown in FIG. 8, the arithmetic circuit 21 determines the values ​​of meanImp, MedianImp, minImp, and maxImp for each bacterial species using Random Forest. In addition, the arithmetic circuit 21 determines the value of normHits using Boruta analysis. Furthermore, the arithmetic circuit 21 obtains the results Confirmed when it is predicted that each positive or negative judgment has been influenced, Tentative when it is unclear whether it has influenced the positive or negative judgment, and Rejected when it is predicted that it has not influenced the positive or negative judgment, according to the values ​​of meanImp, MedianImp, minImp, maxImp, and normHits, using Boruta analysis. The arithmetic circuit 21 determines the bacterial species for which a result of Confirmed or Tentative has been obtained as the predicted bacterial species. Therefore, in the example shown in Fig. 8, the arithmetic circuit 21 determines Dehalobacter sp. DCA, Bacillus infantis, Burkholderia ambifaria, Dehalobacter sp. CF, etc. as the predicted bacterial species of the subgroup 100. In addition, as shown in Fig. 9, the arithmetic circuit 21 determines the bacterial species determined as the predicted bacterial species in any of the subgroups as the related predicted bacterial species.

[0044] The arithmetic circuit 21 can obtain a species-specific correction value used in the calculation of the prediction score for each species determined as a related predicted species, according to the number of times the species was determined as a predicted species in each subgroup. In the example shown in FIG. 10, the arithmetic circuit 11 multiplies the number of times the species was determined as a predicted species by a predetermined number, 0.01, to obtain a species-specific correction value. In FIG. 10, each species is marked with a circle when it is determined as a predicted species for each subgroup, and marked with a cross when it is not determined. Specifically, Burkholderia stabilis is marked with a circle because it was determined as a predicted species in subgroup 1, marked with a circle because it was determined as a predicted species in subgroup 2, and marked with a circle because it was determined as a predicted species in subgroup 100. This Burkholderia stabilis was determined as a predicted species in 17 groups in all of subgroups 1 to 100. Therefore, the arithmetic circuit 21 obtained a species-specific correction value of 0.17 for Burkholderia stabilis. In this way, the arithmetic circuit 21 obtains a species-specific correction value for each predicted related species by using the number of groups determined as the predicted species. For all predicted related species, the arithmetic circuit 21 generates correction value data 123 by associating the predicted related species with the species-specific correction value. In the example shown in FIG. 10, the dashed line portion is associated with the correction value data 123. Furthermore, the arithmetic circuit 21 stores the generated correction value data 123 in the storage device 12.

[0045] The arithmetic circuit 21 classifies each predicted species as a positive-associated predicted species when the proportion of the predicted species associated with a positive label in each subgroup is high, and as a negative-associated predicted species when the proportion of the predicted species associated with a negative label is high. FIG. 11 shows a method for classifying the predicted species in subgroup 1 into a positive-associated predicted species and a negative-associated predicted species. Each of subgroups 1 to 100 includes pairs of labels and species obtained from samples of 14 positive subjects and pairs of labels and species obtained from samples of 13 negative subjects. In the example of FIG. 11, the number of predicted species is 16. FIG. 11 associates each predicted bacterial species in subgroup 1 with the number of bacteria associated with positive labels in subgroup 1 (CAM number of people), the ratio of the number of CAMs to the number of positive labels in subgroup 1 (CAM ratio), the number of bacteria associated with negative labels in subgroup 1 (NCAM number of people), the ratio of the number of NCAMs to the number of negative labels in subgroup 1 (NCAM ratio), and the classification result (classification result) of whether the bacteria is a positively associated predicted bacterial species or a negatively associated predicted bacterial species based on the CAM ratio and NCAM ratio. For example, it shows that Ureaplasma parvum is identified from samples of eight positive subjects and four negative subjects. It also shows that in subgroup 1, 57.14% (8 / 14) of the total positive subjects (14) contain Ureaplasma parvum. It also shows that in subgroup 1, 30.77% (4 / 13) of the total negative subjects (13) contain Ureaplasma parvum. In this case, since the proportion associated with the positive label is high (dashed line), Ureaplasma parvum is classified as a "positively associated predicted species." In Figure 11, "positively associated predicted species" is written as "CAM." For example, it shows that Lactobacillus helveticus was identified from samples of eight positive subjects and 11 negative subjects. And, in subgroup 1, 57.14% (8 / 14) of the total positive subjects (14) contained Lactobacillus helveticus.In addition, in subgroup 1, 84.62% (11 / 13) of the total negative subjects (13 people) contain Lactobacillus helveticus. In this case, since the proportion associated with the negative label is high (dashed line), the classification result is that Lactobacillus helveticus is classified as a "negatively associated predicted species." In addition, in FIG. 11, "negatively associated predicted species" is written as "NCAM." In this way, all predicted species in all subgroups from subgroups 1 to 100 are classified as either positively associated predicted species (CAM) or negatively associated predicted species (NCAM).

[0046] The arithmetic circuit 21 generates an associated predicted species database 122 indicating whether each species is a positive associated predicted species or a negative associated predicted species in the entire learning data based on the classification result for each subgroup. Specifically, the arithmetic circuit 21 uses the classification result of whether each associated predicted species is classified as a positive associated predicted species or a negative associated predicted species in each subgroup to determine whether the species is a positive associated predicted species or a negative associated predicted species in the entire learning data 211. More specifically, if the number of species classified as positive associated predicted species is greater than the number of species classified as negative associated predicted species, the arithmetic circuit 21 determines that the species is a positive associated predicted species in the entire learning data 211. On the other hand, if the number of species classified as negative associated predicted species is greater than the number of species classified as positive associated predicted species, the arithmetic circuit 21 determines that the species is a negative associated predicted species in the entire learning data 211.

[0047] As shown in FIG. 12, the arithmetic circuit 21 obtains the number of groups (CAM group number) that are positively associated predicted species (CAM) and the number of groups (NCAM group number) that are negatively associated predicted species (NCAM) for each associated predicted species. For example, in the example of FIG. 12, the number of CAM groups for Burkholderia stabilis is "4" and the number of NCAM groups is "96". Therefore, the arithmetic circuit 21 judges Burkholderia stabilis as a negatively associated predicted species (NCAM). On the other hand, the number of CAM groups for Dialister pneumosintes is "90" and the number of NCAM groups is "9". Therefore, the arithmetic circuit 21 judges Dialister pneumosintes as a positively associated predicted species (CAM). Note that associated predicted species are not necessarily included in all subgroups. Therefore, for example, as shown in FIG. 12, even if there are 100 subgroups, the total number of CAM groups and the number of NCAM groups is not necessarily 100. Furthermore, when the arithmetic circuit 21 has determined that all associated predicted bacterial species are either positively associated predicted bacterial species or negatively associated predicted bacterial species, it stores an associated predicted bacterial species database 122 generated by associating the species with the determination results (broken line portion) in the storage device 22. The calculation device 1 uses the associated predicted bacterial species database 122 thus generated to calculate the prediction score.

[0048] <<Generation Process>> 13, a generation process for generating various data used in the calculation process, which is executed by the generation device 2, will be described. In the pre-processing, the predicted bacterial species is identified, and the associated predicted bacterial species database 122 and the correction value data 123 are generated. First, the arithmetic circuit 21 acquires the learning data 221 (S101).

[0049] Next, the arithmetic circuit 21 determines the bacterial species contained in the identification results of all subjects as high-reliability bacterial species of the learning data 221 (S102).

[0050] After that, the arithmetic circuit 21 generates a predetermined number of subgroups from the learning data 221 acquired in step S101 (S103). In the example of the present disclosure, the arithmetic circuit 21 generates 100 subgroups by randomly selecting a predetermined set.

[0051] Next, the arithmetic circuit 21 selects one subgroup from the multiple subgroups generated in step S103 (S104).

[0052] Next, the arithmetic circuit 21 executes the random forest and Volta analysis on the subgroup selected in step S104 (S105), thereby obtaining the results of Confirmed, Tentative, and Rejected for each bacterial species included in the subgroup.

[0053] Furthermore, the arithmetic circuit 21 determines the bacterial species judged as Confirmed and Tentative as predicted bacterial species using the results obtained in step S105 (S106).

[0054] The arithmetic circuit 21 repeats the processes of steps S104 to S107 until the determination of the predicted bacterial species is completed for all subgroups. The processes of steps S104 to S107 described above are processes related to the specification of the predicted bacterial species.

[0055] When the determination of the predicted species has been completed for all subgroups (YES in S107), the calculation circuit 21 determines the positively associated predicted species and the negatively associated predicted species for each predicted species using the ratio of labels associated with each subgroup (S108).

[0056] Next, the arithmetic circuit 21 uses the result obtained in step S108 to generate an associated predicted species database 122, which is a list of whether each predicted species is a CAM-associated predicted species or an NCAM-associated predicted species (S109). At this time, the arithmetic circuit 21 stores the generated associated predicted species database 122 in the storage device 22.

[0057] Furthermore, the arithmetic circuit 21 of each prediction device 1 calculates a specific correction value for each predicted bacterial species according to the number of groups in which each bacterial species is included (S110).

[0058] 《Comparison results》 An example of learning data used in machine learning will be described with reference to FIG. 14. Specifically, the learning data includes a plurality of pairs of a label indicating whether the subject is negative or positive for chorioamnionitis and the type of bacterial species obtained as the analysis result of the sample. The total number of samples of the learning data used this time is 81. In addition, 37 of the samples are analysis results of negative subject samples, and 44 are analysis results of positive subject samples. Here, a set of 56 subjects, which is 70% of the total number of samples, is used as a training set, and a set of 25 subjects, which is 30%, is used as a test set. The training set shown in FIG. 14 corresponds to the learning data 211 described above with reference to FIG. 7. In the training set, 30 is the analysis result of a negative subject, and 26 is the analysis result of a positive subject. In addition, in the test set, 14 is the analysis result of a negative subject, and 11 is the analysis result of a negative subject. During training, 13 negative and 14 positive subject samples were selected using stratified random sampling to generate 100 subgroups, as shown in Figure 14.

[0059] FIG. 15 shows a comparison of the prediction accuracy (AUC value) of the training set and the test set obtained by determining high-confidence bacterial species using 100 subgroups obtained as shown in FIG. 14. In FIG. 15, the prediction accuracy is compared when the number of reads is 3000, 4000, 5000, 6000, and 8000 randomly extracted, and when the number of reads is all reads (1 hour). In the example shown in FIG. 15, the number of reads is about the same for the training set and the test set when the number of reads is 5000, so a highly accurate prediction score was obtained. Note that the data for the training set and the test set shown in FIG. 15 were generated using the results of long-read analysis.

[0060] FIG. 16 shows a comparison of prediction accuracy between prediction scores obtained without using the species-specific correction value and prediction scores obtained using the species-specific correction value described above in this disclosure. In FIG. 16, prediction accuracy is compared when the number of reads is 3000, 5000, and 8000 randomly extracted. From the example shown in FIG. 16, it can be seen that the prediction results were stabilized by using the species-specific correction value. In particular, when the number of reads was 8000, the prediction scores obtained without using the species-specific correction value were significantly different between the training set and the test set, but by using the species-specific correction value, it can be seen that the scores changed to the same level. Note that the training set and test set shown in FIG. 16 were also generated using the results of long read analysis, similar to the training set and test set shown in FIG. 15.

[0061] FIG. 17 shows an example of prediction accuracy when using a training set and a test set obtained using the results of short read analysis. In the example shown in FIG. 17, a part of the results of short read analysis is used as a training set and the rest is used as a test set, as shown in FIG. 14. Also, as shown in FIG. 14, 100 subgroups are randomly generated from the training set to determine high-reliability bacterial species, and the prediction accuracy (AUC value) of the training set and the test set are compared. In FIG. 17, the prediction accuracy is compared between the case where the number of reads is randomly extracted as 3000, 4000, 5000, 6000, and 8000, and the case where the number of reads is all reads (2 days). In the example shown in FIG. 17, in all cases, the prediction accuracy is significantly different between the training set and the test set, and a highly accurate prediction score was not obtained.

[0062] As described above, by using the calculation device 1 of the present disclosure, [Industrial Applicability]

[0063] The present disclosure is useful for predicting the onset of certain infectious diseases. [Explanation of symbols]

[0064] 1. Calculation device 11 Arithmetic circuit 12 Storage device 13 Input Devices 14 Output Devices 15. Communications Equipment

Claims

1. A method for calculating a prediction score used to predict the onset of a specific infectious disease in a subject, comprising: Obtaining a sequence result including a plurality of gene sequences obtained using a sample collected from the subject as sequence information; Randomly selecting a predetermined number of the gene sequences included in the sequence information to generate a plurality of sequence groups; Identifying the species of bacteria contained in each of the generated sequence groups by referring to a previously generated bacterial species database that associates gene sequences with bacterial species; determining at least one or more bacterial species contained in all of the sequence groups as high-confidence bacterial species that are likely to affect the infectious disease; The predicted score is calculated using a species-specific correction value that is predetermined for the species determined as the high-reliability species. Calculation method.

2. In calculating the prediction score, With reference to a related predicted species database including positively associated predicted species determined to be associated with a positive determination of the onset of the infectious disease and negatively associated predicted species determined to be associated with a negative determination of the onset of the infectious disease, it is determined whether each of the high reliability species corresponds to either a positively associated predicted species or a negatively associated predicted species, the sum of species-specific correction values ​​predetermined for one or more species determined to be the positively associated predicted species is set as a CAM value, the sum of species-specific correction values ​​predetermined for one or more species determined to be the negatively associated predicted species is set as a corrected NCAM value, and the difference therebetween is calculated as the prediction score. The calculation method according to claim 1 .

3. The gene sequence results were obtained using a long-read method. The calculation method according to claim 1 .

4. The infection is chorioamnionitis The calculation method according to claim 1 .

5. A method for generating the predicted related species database used in the calculation method of claim 2, comprising the steps of: Obtaining learning data including pairs of labels indicating whether a particular infectious disease is positive or negative for a plurality of subjects and bacterial species obtained using the sequencing results of samples collected from each subject; The bacterial species obtained from the sequencing results of all subjects are determined as high-confidence bacterial species that are likely to affect the infectious disease; generating a plurality of subgroups each including a predetermined number of species randomly selected from both a group of a plurality of species labeled as positive and a group of a plurality of species labeled as negative in the learning data; For each of the subgroups, a predicted bacterial species that is predicted to have influenced the positive or negative judgment regarding the infectious disease is determined from the determined high-confidence bacterial species using a machine learning algorithm; For each of the predicted species, if the proportion associated with a positive label in each of the subgroups is high, the predicted species is classified as a positive-associated predicted species, and if the proportion associated with a negative label is high, the predicted species is classified as a negative-associated predicted species; Based on the result of the classification, the database of associated predicted species is generated, which indicates whether each species is a positively associated predicted species or a negatively associated predicted species. Generation method.

6. A method for calculating a species-specific correction value used in the calculation method of claim 2, comprising the steps of: For each of the predicted bacterial species determined by the generation method of claim 5, the species-specific correction value is calculated according to the number of times the predicted bacterial species is determined as the predicted bacterial species in each of the subgroups. Calculation method.

7. The species-specific correction value is a value obtained by multiplying the number of times by a predetermined number. The calculation method according to claim 6.

8. The subgroups are generated from a predetermined number of label and strain pairs selected using a stratified random sampling method. The method of claim 7.

9. A calculation device including an arithmetic circuit and calculating a prediction score used to predict onset of a specific infectious disease in a subject, The arithmetic circuit includes: Obtaining a sequence result including a plurality of gene sequences obtained using a sample collected from the subject as sequence information; Randomly selecting a predetermined number of the gene sequences included in the sequence information to generate a plurality of sequence groups; Identifying the species of bacteria contained in each of the generated sequence groups by referring to a species database that associates gene sequences with types of bacterial species; determining at least one or more bacterial species contained in all of the sequence groups as high-confidence bacterial species that are likely to affect the infectious disease; The predicted score is calculated using a species-specific correction value that is predetermined for the species determined as the high-reliability species. Calculation device.

10. A computer program product causing a computer to carry out the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Method for predicting onset of chorioamnionitis

    WO2021112236A1