Information processing system, information processing method, and program

The information processing system enhances health advice accuracy by clustering bacterial data based on similarity and transition frequency, addressing the variability of bacterial flora patterns and daily fluctuations.

JP7742103B2Active Publication Date: 2025-09-19CYKINSO INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021122746
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-07-27
Publication Date
2025-09-19
Estimated Expiration
2041-07-27

AI Technical Summary

Technical Problem

Conventional techniques fail to accurately understand a subject's health condition based on the abundance of a single human resident bacterial species, as bacterial flora patterns vary significantly among individuals and fluctuate daily, leading to ineffective lifestyle advice.

Method used

An information processing system that clusters human resident bacterial data into multiple clusters based on both bacterial flora similarity and transition frequency, using algorithms like kmeans++ and Fast Greedy, to generate personalized lifestyle advice.

Benefits of technology

Improves the accuracy of deriving health improvement advice by considering individual bacterial flora patterns and transition frequencies, providing tailored advice that is effective for each subject.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007742103000001
    Figure 0007742103000001
  • Figure 0007742103000002
    Figure 0007742103000002
  • Figure 0007742103000003
    Figure 0007742103000003
Patent Text Reader

Abstract

To provide an information processing system, an information processing method and a program for increasing the accuracy of deriving improvement advice for improving the health state of users on the basis of the state of human resident bacteria.SOLUTION: In an information processing system, an information processing device includes a clustering unit for clustering the unit data of a plurality of subjects into a plurality of clusters on the basis of the frequency of transition of the respective unit data of the plurality of subjects, which are obtained by time-series information of each of the plurality of subjects, while time-series information has preliminary been acquired for each of the plurality of subjects, in which two or more of unit data that includes at least information regarding two or more human microbial species are arranged in a time direction.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system, an information processing method, and a program. [Background technology]

[0002] BACKGROUND ART Conventionally, in order to understand the health condition of a subject, the normal human bacteria of the subject have been examined. For example, Patent Document 1 proposes a technology that examines the intestinal bacteria of a subject and suggests dietary improvements based on a predetermined correlation between the intestinal bacteria and the diet. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-165716 Summary of the Invention [Problem to be solved by the invention]

[0004] However, conventional techniques, including the technique described in Patent Document 1, alone do not necessarily enable accurate understanding of the health condition of a subject. That is, for example, since a large number and variety of human resident bacterial species function as a whole in the human body, it is difficult to accurately grasp the health condition of a subject based solely on the abundance of a single human resident bacterial species, such as healthy human resident bacterial species and unwell human resident bacterial species. In other words, it is necessary to consider the multiple human resident bacterial species present in the subject's body and the pattern (combination) of the number of each human resident bacterial species. Furthermore, in many cases, advice on lifestyle improvement has been provided based on the assumption that there are certain bacterial species that everyone should increase or decrease in order to achieve better health. However, the bacterial flora patterns of multiple subjects are different, and while each subject's bacterial flora pattern fluctuates on a daily basis, the extent to which this can change also differs from subject to subject. Therefore, advice aimed at transitioning to a universally accepted ideal bacterial flora pattern may not be as effective for each subject. The present invention has been made in light of these circumstances, and aims to improve the accuracy of understanding a subject's health condition based on the species of bacteria normally present in humans and of providing advice on improving lifestyle habits.

[0005] The present invention aims to improve the accuracy of deriving improvement advice that can improve a user's health condition based on the state of human resident bacteria. [Means for solving the problem]

[0006] In order to achieve the above object, an information processing system according to one aspect of the present invention comprises: In a state where time-series information consisting of two or more unit data items arranged in a time direction, each of which includes information on two or more human microbial species, is previously acquired for each of a plurality of subjects, a clustering means for clustering the unit data of the plurality of subjects into a plurality of clusters based on the frequency of transition of each of the unit data of the plurality of subjects obtained from the time-series information of each of the plurality of subjects; An information processing system comprising:

[0007] An information processing method and a program corresponding to the information processing system according to one aspect of the present invention are also provided as an information processing method and a program according to one aspect of the present invention. [Effects of the Invention]

[0008] According to the present invention, it is possible to provide a technology that improves the accuracy of deriving improvement advice that can improve the health condition of a subject based on the state of human resident bacteria. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a block diagram showing a hardware configuration of an information processing device according to an embodiment of an information processing system of the present invention. [Figure 2] 2 is a functional block diagram showing an example of a functional configuration of the information processing device of FIG. 1; [Figure 3] 3 is a schematic diagram illustrating an example of the flow of a learning process executed by the information processing device of FIG. 2. FIG. [Figure 4] FIG. 3 is a diagram showing an example of the result of bacterial flora similarity clustering performed by an information processing device having the functional configuration of FIG. 2. [Figure 5] 3 is a diagram showing an example of a result of time-series clustering performed by an information processing device having the functional configuration of FIG. 2. FIG. [Figure 6] 3 is a flowchart illustrating an example of the flow of an evaluation process executed by the information processing device of FIG. 2. [Figure 7] FIG. 7 is a diagram showing an example of an evaluation result including improvement advice generated by the improvement advice generation process of FIG. 6. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. FIG. 1 is a block diagram showing the hardware configuration of an information processing device according to an embodiment of the information processing system of the present invention.

[0011] The information processing device 1 includes a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, a bus 14, an input / output interface 15, an input unit 16, an output unit 17, a memory unit 18, a communication unit 19, and a drive 20. The CPU 11 executes various processes according to a program recorded in the ROM 12 or a program loaded from the storage unit 18 into the RAM 13 . The RAM 13 also stores data and the like necessary for the CPU 11 to execute various processes.

[0012] The CPU 11, ROM 12, and RAM 13 are connected to one another via a bus 14. An input / output interface 15 is also connected to this bus 14. An input unit 16, an output unit 17, a memory unit 18, a communication unit 19, and a drive 20 are connected to the input / output interface 15. The input unit 16 is composed of a keyboard, a mouse, etc., and inputs various information in response to user instructions. The output unit 17 is composed of a display, a speaker, etc., and outputs images and sounds. The storage unit 18 is configured with a hard disk or the like, and stores various types of information data.

[0013] The communication unit 19 controls communication with other terminals (not shown) via the network, and removable media 31, which may be a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, is appropriately attached to the drive 20. Programs read from the removable media 31 by the drive 20 are installed in the storage unit 18 as needed. The removable media 31 can also store various data stored in the storage unit 18 in the same way as the storage unit 18.

[0014] FIG. 2 is a functional block diagram illustrating an example of a functional configuration of the information processing apparatus of FIG. In this embodiment, the information processing device 1 can execute two processes: a learning process and an evaluation process. As will be described in more detail later, the learning process refers to a series of processes for learning clustering using human resident bacterial data and other data as learning data, taking into consideration the similarity of the human resident bacterial data and the transition frequency of the human resident bacterial data. In addition, the learning process calculates improvement advice within such clusters. In addition, the evaluation process refers to a series of processes that use the model obtained as a result of the learning process to output the human resident bacterial data of the user being evaluated as evaluation data, such as the cluster to which the human resident bacterial data of the user being evaluated belongs, and improvement advice for the user being evaluated. The concept of transition frequency of human resident bacteria data will be described in detail later.

[0015] In the CPU 11 of the information processing device 1, a learning unit 111 functions during learning processing, and an evaluation unit 112 functions during evaluation processing. In addition, a learning DB 200, a bacterial flora similarity model DB 300, a transition frequency model DB 400, and an improvement advice DB 500 are provided in one area of ​​the storage unit 18.

[0016] In the learning unit 111, a learning data management unit 121, a clustering unit 122, and an improvement advice calculation unit 123 function.

[0017] The learning data management unit 121 acquires learning data used in the learning process, processes it appropriately, and stores and manages it in the learning DB 200. In the learning data management unit 121, a learning user information acquisition unit 131, a time-series information collection unit 132, a preprocessing unit 133, and a learning data integration management unit 134 function.

[0018] The learning user information acquiring unit 131 acquires, as learning user information, human resident bacteria data and questionnaire data for each of N subjects (N is an integer value of 2 or more).

[0019] The time-series information collecting unit 132 collects time-series information on the human commensal bacteria data and questionnaire data for each of the N subjects.

[0020] Here, the human resident bacterial data refers to data that includes at least information on two or more human resident bacterial species. For example, data showing the amount of each of multiple human resident bacterial species collected from feces is an example of human resident bacterial data.

[0021] Here, "human resident bacteria" refers mainly to bacteria that are constantly present in the human body, not just transiently. For example, so-called intestinal bacteria are an example of human resident bacteria. Furthermore, a specific species of human resident bacteria obtained as a result of a specific classification is called a "human resident bacterial species." The amount of each species of human commensal bacteria present may be "0," i.e., the total absence of that species of human commensal bacteria or the total absence of all species of human commensal bacteria.

[0022] Here, in this embodiment, the human resident bacterial data refers to bacterial flora sample data obtained at each time from N subjects who are the subject of the learning process or users who are the subject of the evaluation process. That is, human resident bacterial data (bacterial flora sample data) is data containing information about human resident bacteria for one time, obtained by one biological sample collection (sampling) and its analysis. Specifically, for example, the human resident bacterial data includes the species and lineages of a plurality of human resident bacterial species, their ratios, and the like. The learning user information acquiring unit 131 acquires the human resident bacteria data for each of the N subjects as unit data each time a biological sample is collected (sampling) once.

[0023] The questionnaire data includes at least information about the health status and lifestyle habits of each of the N subjects. For example, questions about the age, sex, account information, BMI (Body Mass Index), health status, bowel movements, diet, lifestyle habits, mental state, and body shape are asked to each of the N subjects who provided the human resident bacteria data, and the results (answers) are collected as the questionnaire data. In this embodiment, every time a biological sample is collected (sampling) once for a predetermined subject, the learning user information acquiring unit 131 acquires questionnaire data for the predetermined subject.

[0024] Here, each of the N subjects is assigned an individual ID. The IDs of the N subjects are also associated with the human resident bacteria data and the questionnaire data. In other words, the human resident bacteria data and questionnaire data with the same ID are data on the same subject.

[0025] For each of the N subjects, such biological sampling is performed two or more times at intervals.

[0026] Therefore, the time series information collection unit 132 acquires, for each of the N subjects, information consisting of two or more pieces of human resident bacterial data (unit data) arranged in the time direction, each time a biological sample is collected (sampling), acquired by the learning user information acquisition unit 131, as time series information of the human resident bacterial data. Specifically, for example, if biological samples are collected and analyzed once a month for a specific subject, information consisting of the human resident bacterial data (unit data) for each collection arranged in the same number of times as the number of times continued in the time direction will be obtained as time-series information on the human resident bacterial data for the specific subject.

[0027] In addition, the time series information collection unit 132 acquires, for each of the N subjects, information consisting of two or more pieces of questionnaire data acquired by the learning user information acquisition unit 131 arranged in the time direction each time a biological sample is collected (sampling), as time series information of the questionnaire data.

[0028] The preprocessing unit 133 performs preprocessing on the time-series information of the human resident bacterial data and questionnaire data collected for each of the N subjects by the time-series information collecting unit 132, converting the data into a predetermined format and numerical values ​​so that the data can be used for the clustering and calculations described below.

[0029] The learning data integration management unit 134 stores and manages the time series information of the human resident bacterial data and questionnaire data for a specified subject that have been preprocessed by the preprocessing unit 133 as learning data for the specified subject, as learning user data for N subjects in the learning DB 200. That is, the learning data integration management unit 134 manages the human resident bacteria data, the questionnaire data, and their time-series information as integrated learning data by associating them with each other.

[0030] The clustering unit 122 clusters the human resident bacteria data included in the training data of N subjects stored in the training DB 200. Specifically, the human resident bacterial data (unit data) of N subjects is clustered into multiple clusters based on the frequency of transitions of each of the human resident bacterial data (unit data) of N subjects obtained from the time series information of each of the N subjects. Specifically, in the clustering section 122, a bacterial flora similarity clustering section 141 and a transition frequency clustering section 142 function.

[0031] Here, the meanings of terms that are the premise of the clustering process will be briefly explained. "Clustering" refers to generating a subset that satisfies a predetermined condition from a set to be analyzed, and "cluster" refers to the subset generated by clustering.

[0032] The clusters resulting from the clustering may overlap with each other, and there may be analysis subjects that do not belong to any of the clusters. However, in the following description, it is assumed that the clusters resulting from the clustering in this embodiment do not overlap with each other.

[0033] The bacterial flora similarity clustering unit 141 clusters the human resident bacterial data (unit data) of N subjects into M clusters (M is an integer value of 2 or greater independent of N) using a predetermined algorithm based on the similarity of the bacterial flora itself. The bacterial flora similarity clustering unit 141 then stores and manages information about the clusters clustered from the viewpoint of the similarity of the bacterial flora themselves as a bacterial flora similarity clustering model MOD1 in the bacterial flora similarity model DB300.

[0034] Although details will be described later using FIG. 3, the bacterial flora similarity clustering unit 141 performs clustering using the following algorithm as a predetermined algorithm from the viewpoint of the similarity of the bacterial flora itself.

[0035] That is, the bacterial flora similarity clustering unit 141 performs clustering based on the similarity between each of the multiple pieces of human resident bacterial data (unit data) for each of the N subjects. For example, the bacterial flora similarity clustering unit 141 performs clustering so that the bacterial flora of the human resident bacteria (human resident bacterial data for each time) that are determined to have high similarity in terms of the species, lineages, ratios, etc. thereof that make up the bacterial flora belong to the same cluster. In this way, the bacterial flora similarity clustering unit 141 performs clustering from the perspective of the similarity of the bacterial flora itself (hereinafter referred to as "microbial flora similarity") as a predetermined algorithm.

[0036] The bacterial flora similarity between multiple pieces of human resident bacterial data (unit data) can be defined as follows: For example, the smaller the distance index (e.g., Euclidean distance) between the unit data in the space where the unit data can exist, the higher the bacterial flora similarity. Here, the Euclidean distance takes into account the proportion of the constituent species and their abundance. In this way, the bacterial flora similarity clustering section 141 can cluster unit data whose bacterial flora similarity is equal to or greater than a predetermined value (distance index is equal to or less than a predetermined value) into small clusters.

[0037] Here, the clusters resulting from the clustering by the above-described bacterial flora similarity clustering unit 141 are called "small clusters." That is, "small clusters" are clusters resulting from clustering performed from the perspective of bacterial flora similarity.

[0038] The transition frequency clustering unit 142 clusters the small clusters (results of clustering by bacterial flora similarity) clustered by the bacterial flora similarity clustering unit 141 into L clusters (L is an integer value between 1 and N) using a predetermined algorithm based on the transition frequency of the bacterial flora. Then, the transition frequency clustering unit 142 stores and manages information about the clusters clustered in terms of the transition frequency of the bacterial flora as a transition frequency clustering model MOD2 in the transition frequency model DB 400.

[0039] Although details will be described later with reference to FIG. 3, the transition frequency clustering unit 142 performs clustering by adopting the following algorithm as the predetermined algorithm.

[0040] That is, the transition frequency clustering unit 142 performs clustering based on the frequency (number or number) of transitions between the human commensal bacteria data in the time-series information collected for each of the N subjects. For example, the transition frequency clustering unit 142 performs clustering based on which small clusters the human resident bacteria of each of the N subjects (human resident bacteria data for each time) can transition between and which small clusters have a high frequency of transition, so that those determined to have a high frequency belong to the same cluster. In this way, the transition frequency clustering unit 142 performs clustering from the perspective of the frequency of mutual transitions between small clusters (hereinafter referred to as "transition frequency") as a predetermined algorithm from the perspective of the transition frequency of the bacterial flora.

[0041] Here, the clusters resulting from the clustering performed by the transition frequency clustering unit 142 are referred to as "large clusters." That is, the "large clusters" are clusters resulting from clustering performed from the perspective of bacterial flora similarity.

[0042] As described above, the transition frequency clustering unit 142 can cluster M small clusters into L large clusters.

[0043] Here, a specific method of clustering from the viewpoint of transition frequency, which is performed by the transition frequency clustering unit 142, will be outlined. That is, the transition frequency clustering unit 142 first generates a weighted undirected graph based on the time-series information of each of the N subjects, using the magnitude of the transition frequency between small clusters as the weight between the small clusters. Next, the transition frequency clustering unit 142 clusters the small clusters into large clusters based on the weighted undirected graph.

[0044] Then, the transition frequency clustering unit 142 associates the unit data that belonged to each small cluster as belonging to the large cluster into which the data has been grouped. In this way, the transition frequency clustering unit 142 clusters the data (unit data) of the multiple human indigenous bacterial species for each of the N subjects into large clusters.

[0045] To summarize the above, the transition frequency clustering model MOD2 of this embodiment can be said to be a model that can identify the large cluster to which unit data belongs based on multiple human resident bacterial data (unit data) for each of N subjects.

[0046] As a result, the transition frequency clustering unit 142 can cluster small clusters into large clusters as follows. That is, for example, although not shown, suppose that a first subject has transitioned (mutually transitioned) from the first small cluster to the second small cluster to the third small cluster according to the time-series information of the human resident bacterial species data (unit data). Similarly, suppose that a second subject has transitioned (mutually transitioned) from the first small cluster to the second small cluster to the fourth small cluster. Furthermore, suppose that a third subject has transitioned from the fifth small cluster to the sixth small cluster to the seventh small cluster.

[0047] In this case, for example, the transition frequency clustering unit 142 can cluster the data into two large clusters: a first large cluster consisting of the first to fourth small clusters, and a second large cluster consisting of the fifth to seventh small clusters. That is, in this example, the transition frequency clustering unit 142 can perform clustering on the condition that the frequency of transitions between the small cluster group consisting of the first to fourth small clusters and the small cluster group consisting of the fifth to seventh small clusters is low (0 times in this example).

[0048] As an example where the condition for the frequency of transition is different from that described above, the transition frequency clustering unit 142 can perform clustering on the condition that the frequency of transitions (mutual transitions) between the first small cluster and the second small cluster is high (in this example, two times).

[0049] In this way, the transition frequency clustering unit 142 uses a predetermined algorithm to cluster the small clusters to which the human resident bacteria data (unit data) belong, based on the transition frequency, that is, from which small cluster to which small cluster the transition can occur and with which frequency. Furthermore, the transition frequency clustering unit 142 clusters M small clusters into L large clusters based on the time-series information of the human commensal bacteria data (unit data) of each of the N subjects.

[0050] To summarize the above, the clustering unit 122 is composed of a bacterial flora similarity clustering unit 141 and a transition frequency clustering unit 142, and can cluster multiple bacterial flora sample data (unit data) for each of N subjects based on both the bacterial flora similarity and the transition frequency. During the evaluation process, the transition frequency clustering model MOD2 is used to identify the large cluster to which the user's unit data belongs based on the user's human resident bacteria data (unit data). This makes it possible to identify other unit data and small clusters to which the user's unit data can transition.

[0051] Next, we will explain the benefits of taking transition frequencies into consideration when clustering data on human resident bacteria (unit data). First, conventional clustering and questionnaire aggregation will be described using the first to fourth unit data (human resident bacteria data from one biological sample collection (sampling)) as an example of questionnaire aggregation. Here, it is assumed that the first unit of data and the third unit of data are data obtained from a first subject, and the second unit of data and the fourth unit of data are data obtained from a second subject. The unit data includes information such as the species and strains that make up the bacterial flora of each of a plurality of human resident bacterial species, as well as their amounts and ratios.

[0052] For example, in conventional clustering of human resident bacterial species, the state of a subject's bacterial flora has been grasped simply based on whether the human resident bacterial data (unit data) from a single biological sample collection (sampling) of a subject is similar to other biological sample collections (samples).Furthermore, in conventional questionnaire aggregation, the questionnaires have been aggregated simply by correlating them with the human resident bacterial data without distinguishing between subjects.

[0053] Specifically, for example, in a space where unit data may exist, suppose that the first and second unit data are arranged in proximity (close in terms of distance) compared to other unit data. Similarly, suppose that the third and fourth unit data are arranged in proximity (close in terms of distance) compared to other unit data. Furthermore, suppose that the set of the first and second unit data and the set of the third and fourth unit data are not in proximity (close in terms of distance). Conventionally, when such unit data are clustered, the microbiota similarity is clustered based on a distance index (e.g., Euclidean distance) in the space where the unit data may exist. As a result, a first cluster including the first and second unit data and a second cluster including the third and fourth unit data are generated as clustering results.

[0054] As a result, for example, by comparing the lifestyle habits of subjects in a first cluster who answered "in good health" in a questionnaire with those of subjects who answered "in poor health," lifestyle advice that could improve the health of subjects belonging to that cluster was derived. Specifically, for example, if there are many subjects in "good health" near the first unit of data and many subjects in "in poor health" near the second unit of data, this lifestyle advice was derived to change the bacterial flora from the state near the second unit of data to the state near the first unit of data. This was also the case within the second cluster.

[0055] In contrast, this service performs clustering based on transition frequency. In other words, it can be considered that the first unit data and the third unit data were data obtained from the first subject, and the second unit data and the fourth unit data were data obtained from the second subject. That is, it is considered that there is an easy mutual transition between the first and third unit data, and between the second and fourth unit data, and that there is an insufficient mutual transition between the first and second unit data, and between the third and fourth unit data. This can be understood from the fact that the distance between bacterial flora states was calculated by treating the abundance of each bacterial species as an equal numerical value, and was defined without taking into account information on whether the abundance of each bacterial species is likely to change or not, or whether it is easy to transition between bacterial flora states.

[0056] In other words, if "there is easy mutual transition between the first and third unit data, and between the second and fourth unit data, but there is no easy mutual transition between the first and second unit data, and between the third and fourth unit data," then even if it is assumed based on the data, it is actually difficult to "change the bacterial flora from a state near the second unit data to a state near the first unit data," and such lifestyle advice is considered to be ineffective. In this case, it is considered possible to derive actually effective lifestyle advice by linking the ranges where there is easy mutual transition as clusters and identifying "clusters including the first and third unit data" and "clusters including the second and fourth unit data," respectively.

[0057] In other words, rather than simply considering the condition that the bacterial flora between unit data is highly similar, the accuracy of deriving lifestyle advice can be improved by considering whether transition is possible (or has actually occurred), in other words, by using clusters that take transition frequency into consideration. Conversely, it can be said that as a result of performing conventional clustering, improvement advice is given to transition to unit data that cannot be transitioned from the unit data of the first subject, and as a result, appropriate improvement advice is not given.

[0058] In the above example, the method of clustering taking into account the bacterial flora similarity and the method of clustering taking into account the transition frequency were explained using an example of clustering unit data, but the same applies to clustering from small clusters.

[0059] That is, conventionally, for example, 100,000 unit data sets have been clustered into 10 clusters simply by taking into consideration the similarity of the bacterial flora. However, in this embodiment, first, 100,000 unit data are clustered into 1,000 (M) small clusters in consideration of the bacterial flora similarity. Then, the 1,000 (M) small clusters are clustered into 10 (L) large clusters. As a result, a transition frequency clustering model MOD2 is generated for determining to which of the 10 (L) large clusters the 100,000 unit data belong.

[0060] In this manner, in this embodiment, unit data is first clustered into small clusters and then into large clusters. As a result, for example, unit data that is virtually unchanged (in an extreme example, multiple samples taken simultaneously from the same subject) is managed as one small cluster. As a result, more useful data is obtained as data on transition frequency (number of times) rather than the transition frequency between unit data. For example, simply considering the transition frequency for unit data may result in data such as "the number of transitions is either 1 or 0." However, with this service, unit data is first clustered as small clusters. This results in data with a range of transitions (for example, the number of transitions may be between 0 and 10), making it easier to process when clustering into large clusters.

[0061] As described above, in this embodiment, clustering is performed based on both the similarity of bacterial flora and transition frequency, based on information such as the species and strains that make up the bacterial flora of each of a plurality of human resident bacterial species, as well as their amounts and ratios. Because transition frequency is taken into consideration in the clustering of human resident bacterial data (unit data), this service can contribute to the generation of more accurate improvement advice.

[0062] In this manner, in this embodiment, the bacterial species and strains that constitute the bacterial flora of human resident bacteria, their ratios, etc. are classified into M large clusters based on the bacterial flora similarity and transition frequency. As a result, the improvement advice calculation unit 123 and the improvement advice generation unit 153, which will be described later, provide improvement advice according to the large clusters.

[0063] The improvement advice calculation unit 123 calculates improvement advice based on the large clusters clustered by the transition frequency clustering unit 142, and the time-series information of the human commensal bacteria data and questionnaire data of the subject. Improvement advice is intended to improve the subject's health condition and is advice that indicates the direction in which the subject should change their lifestyle habits (such as menu planning, daily rhythm, exercise habits, and intake of supplements). The improvement advice calculation unit 123 stores the calculated improvement advice in the improvement advice DB 500 and manages it.

[0064] The improvement advice calculation unit 123 can also generate a "microflora state score" as part of the improvement advice. Here, the "microflora state score" is a score that represents the expected health state of each unit data (or its surroundings, i.e., unit data included in a small cluster, for example). That is, the microbiota state score in this embodiment is a value that differs among a plurality of unit data included in a certain large cluster. Then, lifestyle habits and the like that will increase the microbiota state score within that large cluster, that is, that will result in a microbiota state of unit data with a high microbiota state score, are generated as improvement advice in this embodiment.

[0065] Specifically, for example, the improvement advice calculation unit 123 calculates improvement advice by cross-tabulating the health status and lifestyle habits based on the questionnaire data for each large cluster clustered by the transition frequency clustering unit 142. That is, by cross-tabulating the health status and lifestyle habits contained in the questionnaire data for each large cluster, it is possible to calculate, for subjects belonging to a certain large cluster, what lifestyle habits tend to lead to better or worse health status.

[0066] Conventionally, the condition of a subject has been grasped based on the similarity of the flora of the subject's resident human bacterial data (unit data) at a certain time, and improvement advice has been calculated. In other words, conventionally, such nearly identical small clusters have been clustered solely based on the similarity of the bacterial flora and treated as the same cluster. As a result, the composition of the preferred bacterial flora within a range that can be easily transitioned for one subject may be different from the composition of the preferred bacterial flora within a range that can be easily transitioned for another subject, and therefore the lifestyle advice that can improve the health status of one subject may also be different from the lifestyle advice that can improve the health status of another subject. Furthermore, lifestyle advice that was effective in improving the health of one subject may not necessarily improve the health of another subject.

[0067] In contrast to this, in this embodiment, the condition of a certain subject is grasped based on the flora similarity and transition frequency of human resident bacterial data (unit data) of a certain subject at a certain time, and improvement advice is calculated. As a result, if, for example, the unit data of the first to third subjects at a certain time have a high degree of bacterial flora similarity between the first and second subjects and have similar human resident bacterial species but belong to different small clusters, and the unit data of the third subject have a low degree of bacterial flora similarity between the first and second subjects, and the small clusters are combined into large clusters based on the transition frequency, and the unit data of the first and third subjects belong to the same large cluster and the unit data of the second subject belongs to a different large cluster, different improvement advice can be calculated for the first and second subjects. In this example, the first subject can be given appropriate advice to bring him closer to the third subject who is in the same large cluster and in better health. That is, in this embodiment, large clusters are extracted based on the bacterial flora similarity and transition frequency, and improvement advice is calculated, thereby providing more accurate improvement advice to each subject. Although the above example has been described using the first to third subjects, the same applies to other users to be evaluated.

[0068] Furthermore, to add, even if certain fluctuations occur in the bacterial flora on a regular basis, the large clusters are a range of attributes that are likely to remain unchanged as characteristics that the person has to a certain extent fixedly, and can be said to be a kind of "constitution." This "constitution" (large cluster) may change to another "constitution" (large cluster) when a particularly large fluctuation occurs. However, in this embodiment, the level of the range that is thought to remain the same "constitution" (large cluster) with the degree of daily change is identified, and improvement advice is provided.

[0069] Next, an example of the flow of the evaluation process will be described. As described above, the evaluation unit 112 functions during the evaluation process.

[0070] In the evaluation unit 112, an evaluation user information acquisition unit 151, a transition frequency evaluation unit 152, and an improvement advice generation unit 153 function.

[0071] The evaluation user information acquisition unit 151 acquires the human resident bacteria data (unit data) of the user to be evaluated as the user data of the user to be evaluated. The user to be evaluated may be a person different from the above-mentioned N subjects, or may be any one of the N subjects.

[0072] The transition frequency evaluation unit 152 uses the trained transition frequency clustering model MOD2 stored in the transition frequency model DB 400 to evaluate to which large cluster the human indigenous bacteria data (unit data) of the user to be evaluated belongs.

[0073] The improvement advice generation unit 153 generates improvement advice for the user to be evaluated, based on the large cluster to which the user belongs as a result of evaluation by the transition frequency evaluation unit 152, out of the improvement advice for each of the M large clusters calculated by the improvement advice calculation unit 123.

[0074] That is, for example, the improvement advice generation unit 153 generates improvement advice for the user to be evaluated, such as lifestyle habits, to transition the user to a unit data (or small cluster) that is considered to have a better health condition within the large cluster to which the user is considered to belong. At this time, the improvement advice generation unit 153 can refer to the improvement advice calculated by the improvement advice calculation unit 123 from the improvement advice DB 500.

[0075] Specifically, for example, a case will be described in which, in a certain large cluster, there is a correlation tendency that the more unit data there is of a certain human resident bacterial species, the more likely it is that the more responses to a questionnaire indicating good health are made. In this case, the improvement advice calculation unit 123 calculates improvement advice for "users in the large cluster who have a high level of that human resident bacterial species," such as "lifestyle habits, etc. for reducing that human resident bacterial species." As a result, the improvement advice generation unit 153 generates improvement advice for "lifestyle habits, etc. for reducing that human resident bacterial species" for users of unit data who belong to that large cluster and have a high level of that human resident bacterial species.

[0076] The improvement advice generation unit 153 may generate improvement advice that includes information other than that corresponding to the large cluster to which the unit data to be evaluated belongs. That is, for example, the improvement advice may include improvement advice for transitioning from the large cluster to another large cluster that is likely to transition but occurs at a low frequency and that has a correlation tendency with more responses indicating better health status. In this way, the improvement advice generation unit 153 can also generate improvement advice that includes continuous or multiple policies.

[0077] As described above, in this embodiment, it is determined to which of the L large clusters the unit data of the user to be evaluated belongs, which are clustered using the time-series data of N subjects in the learning process. Then, improvement advice for the user to be evaluated is generated based on the large cluster to which the unit data of the user to be evaluated belongs. As a result, improvement advice is presented to the user to be evaluated that takes into consideration whether or not the user's own unit data can actually transition. In other words, the improvement advice is actually useful to the user to be evaluated.

[0078] FIG. 3 is a schematic diagram illustrating an example of the flow of the learning process executed by the information processing device of FIG. First, with reference to FIG. 3, the human resident bacteria data and questionnaire data for each of N subjects (N is an integer value of 2 or more) that are acquired by the learning data management unit 121 and managed after being processed as appropriate will be described. FIG. 3 illustrates N subjects U1 to Un. The following describes processing for a specific subject Uk (k is any integer value between 1 and N). Note that, hereinafter, a symbol with a subscript k indicates a numerical value or the like related to the subject Uk.

[0079] The subject Uk undergoes biological sample collection (sampling) Lk times (Lk is an integer value equal to or greater than 1) at predetermined time intervals. Fig. 3 shows the human resident bacteria data Bk-1 to Bk-Lk for the subject Uk as unit data acquired by the learning user information acquisition unit 131 each time a biological sample is collected (sampling) from the subject Uk. Fig. 3 also shows the questionnaire data Ak-1 to Ak-Lk acquired corresponding to each of the human resident bacteria data Bk-1 to Bk-Lk.

[0080] In the following, when there is no need to distinguish between the questionnaire data Ak-1 to Ak-Lk of subject Uk, they will be collectively referred to as "questionnaire data Ak." When referring to "questionnaire data Ak," the human resident bacteria data Bk to Bk-Lk will be referred to as "human resident bacteria data Bk." Furthermore, when there is no need to distinguish between the subjects U1 to Un, they will be collectively referred to as "subject U." When referring to "subject U," "questionnaire data A1 to An" will be collectively referred to as "questionnaire data A." Furthermore, "human resident bacteria data B1 to Bn" will be collectively referred to as "human resident bacteria data B."

[0081] In step ST11, the bacterial flora similarity clustering unit 141 clusters a plurality of unit data (human resident bacteria data B for each time) of each of the N subjects into M small clusters. The bacterial flora similarity clustering unit 141 stores and manages a model for such clustering as a bacterial flora similarity clustering model MOD1 in the bacterial flora similarity model DB 300. As a result, each of the human resident bacteria data B of the N subjects U is associated with a small cluster to which it belongs.

[0082] The algorithm used for clustering in the processing of step ST11 can be a method that handles cluster center point coordinates and in which all samples belong to one of the clusters. Specifically, for example, in this embodiment, the "kmeans++" method is used. Other methods, such as "kmeans," satisfy the above conditions. It is preferable that the number of clusters P is a sufficiently large number appropriate for the total number of samples. In this embodiment, a value of approximately 100 to 300 (i.e., the above-mentioned M = approximately 100 to 300) is used. Furthermore, the general Euclidean distance is used to define the distance for clustering.

[0083] Next, in step ST12, the transition frequency clustering unit 142 clusters the M small clusters resulting from the processing of step ST11 into L large clusters (L is an integer value between 1 and N) using a predetermined algorithm that uses time-series information consisting of the human resident bacteria data Bk (unit data) of N subjects Uk. That is, the small clusters are clustered using time-series information of Lk sessions of human resident bacteria data Bk-1 to Bk-Lk, based on the concept of transition frequency. As described above, the transition frequency clustering unit 142 stores and manages a model for such clustering in the transition frequency model DB 400 as a transition frequency clustering model MOD2. At this time, as described above in the description of the transition frequency clustering unit 142, the human resident bacteria data Bk (unit data) of N subjects Uk are associated with the large clusters.

[0084] First, the transition frequency clustering unit 142 sequentially processes each of the subjects U1 to Un as a (predetermined) subject Uk, and for all pairs of human resident bacterial data B from consecutive test times for the subject Uk (for example, for a subject Uk who has taken the test three times (Lk is 3), there are two pairs: "human resident bacterial data Bk-1 from the first test and human resident bacterial data Bk-2 from the second test" and "human resident bacterial data Bk-2 from the second test and human resident bacterial data Bk-3 from the third test"), identifies the pairs of small clusters to which the unit data from each test belong, and records a link between the two small clusters.

[0085] Next, the transition frequency clustering unit 142 counts the number of links connecting pairs of small clusters and sets the counted number as the weight of the link connecting the pair of small clusters. In this embodiment, links whose entry and exit are the same small cluster, i.e., links indicating that the small cluster to which the pair belongs did not change between consecutive test sessions, are deleted. Based on the weights of the links connecting pairs of small clusters obtained in this way, an undirected weighted graph (hereinafter referred to as a "weighted undirected graph") is generated, in which small clusters are regarded as nodes.

[0086] Here, the transition frequency clustering unit 142 of this embodiment excludes nodes as follows: That is, among the nodes constituting the generated weighted undirected graph, nodes that are not connected to other nodes, i.e., nodes that are not connected by links, are excluded. In this embodiment, nodes that are not connected to the most connected part of the graph (the pair of small clusters with the largest link weight) are also excluded. For example, when the first to fourth small clusters (nodes) are connected in a single line, and the weight of the link between the first and second small clusters is the largest, the fourth small cluster that is not connected to either the first or second small cluster is excluded. Thus, in this embodiment, nodes that are not connected to the most connected part of the graph (the pair of small clusters with the largest link weight) are removed, and the removed nodes are recorded for later use. When the maximally connected portion does not occupy the majority of the entire graph to be analyzed, it is not necessary to exclude nodes that are not connected to the maximally connected portion.

[0087] The transition frequency clustering unit 142 of this embodiment clusters small clusters into large clusters based on the weighted undirected graph obtained as described above, and on the time series information (information composed of unit data of two or more occurrences arranged in the time direction) of the human resident bacteria data B (unit data) of each of the N subjects. Specifically, for example, the transition frequency clustering unit 142 of this embodiment performs clustering according to the "Fast Greedy" algorithm based on the weights of the links (corresponding to the number of transition frequencies between small clusters) between the nodes (small clusters) that make up the generated weighted undirected graph. Note that the algorithm is not limited to this, and any other algorithm that can incorporate a weighting factor into clustering an undirected graph may be used. The number of clusters used in clustering in the transition frequency clustering unit 142 is an appropriate number depending on the purpose of using the clusters. In this embodiment, the number of clusters used is between 5 and 10. The clusters created in this way are called network (NW) clusters because they are in the form of a network (NW) connecting nodes.

[0088] As a result, each of the human resident bacteria data B (unit data) for the N subjects is associated with a small cluster, and further associated with the above-mentioned NW cluster. More specifically, if the small cluster to which a given human resident bacteria data B (unit data) belongs belongs to one of the NW clusters generated in the processing of step ST12, the NW cluster is designated as the belonging destination of the human resident bacteria data B. If the node (small cluster) to which the human resident bacteria data B belongs is to be excluded in the processing of step ST13, the node (small cluster) that is the closest in distance from the coordinates of the human resident bacteria data B among the center point coordinates of the nodes (small clusters) that were not excluded is identified (the distance may be defined as Euclidean distance), and the node (small cluster) is designated as the belonging destination of the human resident bacteria data B.

[0089] Regarding the human resident bacteria data B belonging to a node (small cluster) excluded from the NW cluster, in addition to using the nearby small cluster as described above, it is also possible to treat it as an "exception sample that did not fit into any NW cluster."

[0090] Next, in step ST13, the improvement advice calculation unit 123 calculates improvement advice by cross-tabulating the health status and lifestyle habits based on the questionnaire data for each large cluster clustered by the microbiota similarity clustering unit 141 and the transition frequency clustering unit 142. That is, for each large cluster, the improvement advice calculation unit 123 cross-tabulates the health conditions and lifestyle habits included in the questionnaire data corresponding to the unit data belonging to that large cluster. Specifically, for example, the improvement advice calculation unit 123 calculates, by statistical processing, "what kind of lifestyle habits will result in what kind of health conditions for subjects belonging to that large cluster?"

[0091] An example of the flow of the learning process executed by the information processing device in FIG. 2 has been described above with reference to FIG. An example of the results of bacterial flora similarity clustering will be described below with reference to FIG.

[0092] FIG. 4 is a diagram showing an example of the result of bacterial flora similarity clustering performed by an information processing device having the functional configuration of FIG. FIG. 4 shows each of the small clusters resulting from clustering all of the multiple human resident bacterial data B (unit data) for each of the N subjects U1 to Un, who are the set of subjects to be analyzed, by the bacterial flora similarity clustering unit 141, arranged on a two-dimensional graph. The horizontal axis Ax11 and the vertical axis Ax12 in FIG. 4 are the first and second principal components, respectively, in the principal coordinate analysis. In FIG. 4, each of the small clusters (for example, small cluster SCL1) is shown positioned at the center point (for example, (Ax11, Ax12)=(0.17, 0.33)) of the human commensal bacterial data B belonging to the small cluster.

[0093] Each point on the graph in Fig. 4, i.e., a small cluster (for example, small cluster SCL1), is colored into seven clusters (for example, clusters CL2 and CL7 shown in Fig. 4) as shown in the legend. Specifically, for example, the small cluster SCL in Fig. 4 belongs to cluster CL7.

[0094] Here, the seven clusters shown in Fig. 4 are clusters based only on information about the microbiota similarity. In other words, in the example of Fig. 4, the state of subject U is clustered based only on the similarity (microbiota similarity) of the human resident bacterial data B (unit data) in a single biological sample collection (sampling) of subject U.

[0095] Note that there are no clear boundaries between the seven color-coded clusters in Figure 4 because the graph in Figure 4 is the result of principal coordinate analysis, and the actual human resident bacterial data B is multidimensional data with three or more variables, and the clustering is based on this multidimensional human resident bacterial data B.

[0096] Conventionally, human resident bacterial data Bk obtained as a result of a single biological sample collection (sampling) of a subject Uk has been classified into clusters based solely on the similarity of the bacterial flora, such as the seven clusters shown in Figure 4. Improvement advice has then been given according to the cluster to which the subject belongs. That is, first, questionnaire data is compiled for each cluster (e.g., clusters CL2 and CL7). Then, based on the results of the compilation of the questionnaire data, it is determined which of the clusters (e.g., clusters CL2 and CL7) has a better health condition. Then, when human resident bacteria data Bk is obtained as a result of a single biological sample collection (sampling) of a certain subject Uk, the cluster to which the human resident bacteria data Bk belongs (e.g., cluster CL7) is identified. As a result, improvement advice is given, such as what lifestyle improvements should be made in order to move from cluster CL7 to cluster CL2.

[0097] Above, we have explained, using Figure 4, an example of clustering data on human resident bacteria into small clusters, and an example of conventional clustering based only on bacterial flora similarity. Next, an example of time-series clustering in this embodiment will be described with reference to FIG.

[0098] FIG. 5 is a diagram showing an example of a result of time-series clustering performed by an information processing device having the functional configuration of FIG. Figure 5 shows the nodes (small clusters) and their links (transition frequencies) resulting from clustering the time series information (information consisting of two or more occurrences of human resident bacterial data (unit data) arranged in the time direction) of each of the N subjects being analyzed, arranged on a two-dimensional graph. The horizontal axis Ax21 and the vertical axis Ax22 in FIG. 5 are the first and second principal components, respectively, in the principal coordinate analysis.

[0099] In Figure 5, for pairs of small clusters that have a high transition frequency (for example, the pair of small clusters SCL2 and SCL3), a link (for example, link L1) is shown indicating that the small cluster to which the pair belongs has changed between consecutive test sessions. Furthermore, link L2 for the pair of small clusters SCL3 and SCL4 is shown with a thicker line compared to link L1 in the example above. This represents the weight of each of links L1 and L2. In other words, the thicker the line, the greater the weight of the link, indicating that the number of mutual transitions corresponding to that link has occurred (higher transition frequency).

[0100] That is, each of the small clusters resulting from clustering all of the multiple human resident bacterial data B (unit data) for each of the N subjects U1 to Un by the bacterial flora similarity clustering unit 141 is shown as each point in Figure 5.

[0101] 5. The weight of the link between the small clusters (nodes), which is the magnitude of the transition frequency between the small clusters used by the transition frequency clustering unit 142 as explained in FIG. 3 and the like, is shown as a straight line in FIG.

[0102] Each point in the graph in Fig. 5, i.e., each small cluster (e.g., small cluster SCL, etc.), is colored into seven clusters as shown in the legend. That is, the color of the small clusters in the graph in Fig. 5 indicates the clusters resulting from clustering the small clusters into large clusters based on the transition frequency. That is, the colors of the clusters in FIG. 5 indicate the results of clustering based on the concept of transition frequency using time-series information.

[0103] Even if the same small clusters are clustered, the clusters based on the bacterial flora similarity shown in FIG. 4 and the clusters based on the time-series information shown in FIG. 5 usually yield different clustering results.

[0104] Next, the flow of the evaluation process executed by the information processing device 1 having the functional configuration of FIG. 2 will be described with reference to FIG. FIG. 6 is a flowchart illustrating an example of the flow of the evaluation process executed by the information processing device of FIG.

[0105] The evaluation process is started when evaluation information of the user to be evaluated is acquired at any timing required by the user to be evaluated, and the following steps S21 to S23 are executed.

[0106] In step S21, the evaluation user information acquisition unit 151 acquires human resident bacteria data (unit data) of a user different from the N subjects as user data to be evaluated.

[0107] In step S22, the transition frequency evaluation unit 152 uses the trained transition frequency clustering model MOD2 stored in the transition frequency model DB 400 to evaluate to which of the L large clusters the unit data of the user to be evaluated belongs.

[0108] In step S23, the improvement advice generation unit 153 generates improvement advice for the user to be evaluated, based on the large cluster to which the user belongs as a result of evaluation by the transition frequency evaluation unit 152, out of the improvement advice for each of the M large clusters calculated by the improvement advice calculation unit 123.

[0109] In this way, once improvement advice for the user to be evaluated is generated, the advice generation process ends.

[0110] Next, an example of changes in the evaluation results for a user who is the subject of evaluation will be described with reference to FIG.

[0111] FIG. 7 is a diagram showing an example of an evaluation result including the improvement advice generated by the improvement advice generating process of FIG.

[0112] FIG. 7(A) is an example of an evaluation result including improvement advice based only on the bacterial flora similarity, without using time-series information or large clusters in this embodiment. FIG. 7(B) is an example of an evaluation result using time-series information and large clusters in this embodiment, that is, including improvement advice based on the degree of belonging to a large cluster based on the bacterial flora similarity and time-series information.

[0113] An example of improvement advice in this embodiment will be described below by comparing FIGS. 7(A) and (B).

[0114] The leftmost column of the table shown in FIG. 7(A) shows, from top to bottom, the item names of "bacterial flora cluster type," "bacterial flora state score," "improvement advice," and "skin concerns." That is, each column to the right of that shows the details of the evaluation results for each item of "bacterial flora cluster type," "bacterial flora condition score," "improvement advice," and "skin concerns" for each test.

[0115] For example, in the first to third tests in the example of Figure 7(A), the "bacterial flora cluster type" changes from "I" to "VI" to "I." At this time, the "bacterial flora state score" also changes from "D" to "B" to "D." As described above, the example in Fig. 7(A) is an example of an evaluation result including improvement advice based only on the bacterial flora similarity, without using the time-series information or large clusters in this embodiment. In other words, the "microbial flora cluster type" in Fig. 7(A) corresponds to the cluster based only on the bacterial flora similarity information shown in different colors in the explanation of Fig. 4.

[0116] The content of the items "bacterial flora cluster type," "bacterial flora state score," "improvement advice," and "skin concerns" is the same in the first and third tests. Furthermore, there is a one-to-one correspondence between the "bacterial flora cluster type" and the "bacterial flora state score." This is because, as explained in FIG. 4, questionnaire data is compiled for each cluster based only on the bacterial flora similarity. Then, based on the results of compiling the questionnaire data, it is determined which of the multiple clusters represents the better health condition. That is, in the example of Figure 7(A), the bacterial flora cluster type "I" is defined as having a relatively poor bacterial flora condition score, resulting in a bacterial flora condition score of "D." Also, the bacterial flora cluster type "VI" is defined as having a relatively good bacterial flora condition score of "B."

[0117] When human resident bacterial data Bk is obtained as a result of a single biological sample collection (sampling) of a certain subject Uk, the cluster to which the human resident bacterial data Bk belongs (for example, bacterial flora cluster type "I") is identified. As a result, improvement advice is given, such as what lifestyle improvements should be made in order to belong to bacterial flora cluster type "I" to bacterial flora cluster type "VI." This is because the contents of the "bacterial flora status score" and "improvement advice" items were based on the bacterial flora cluster type.

[0118] In other words, the bacterial flora cluster type in conventional evaluation results was obtained by clustering data on human resident bacteria based only on the degree of similarity. That is, in the example of Figure 7(A), the evaluation was such that "if the bacterial flora state is state I (if it is similar to state I), the bacterial flora state is bad (the bacterial flora state score is low)" and "if the bacterial flora state is state VI (if it is similar to state VI), the bacterial flora state is good (the bacterial flora state score is high)." Furthermore, the advice given to improve the condition was dependent on the cluster (bacterial flora state) that was clustered based solely on the degree of bacterial flora similarity. Specifically, for example, the advice was evaluated as "If the bacterial flora state is state I, you should eat green and yellow vegetables and seafood," and "If the bacterial flora state is VI, you should eat soybeans."

[0119] As a result, although advice for improvement was given in the second test, by the third test, the skin had returned to the state it was in at the first test. Furthermore, the "skin concerns" also recurred (and they were very concerned). This means that for some subjects, the advice for improvement was inappropriate, and their skin condition worsened rather than improved, or the previous improvement was ignored, or the previous improvement was merely a facade.

[0120] In this way, in examples of evaluation results that do not utilize the concept of transition frequency used in this service, the microbiota cluster type is determined solely based on the microbiota similarity, and the score for the microbiota status also depends solely on the microbiota cluster type, which may not necessarily reflect the subject's constitution or situation.

[0121] In contrast, for example, in the first to third tests in the example of Figure 7(B), the "bacterial flora cluster type" remains unchanged from "2." Also, at this time, the "bacterial flora state score" changes from "D," to "B," to "B." As described above, the example in Fig. 7(B) is an example of an evaluation result based on large clusters using the concept of transition frequency in this embodiment. That is, the "microbial flora cluster type" in Fig. 7(B) indicates a cluster based on the bacterial flora similarity and transition frequency in this embodiment.

[0122] The "bacterial flora cluster type" is the same for the first to third tests. However, the "bacterial flora status score" item changes from "D" to "B." This corresponds to the fact that, although the human resident bacteria data B (unit data) for each test belong to the same large cluster, each test is assigned a different bacterial flora status score within the same large cluster.

[0123] Furthermore, the "bacterial flora cluster type" and "bacterial flora status score" are the same in the second and third tests. At this time, the "improvement advice" has changed from "intake of green and yellow vegetables and intake of mixed grain rice" to "intake of green and yellow vegetables." In other words, even if the human resident bacteria data B (unit data) for that test belongs to the same cluster as the previous large cluster and the bacterial flora status score (health status) has not changed, different improvement advice may be generated. In other words, the content of the "improvement advice" item is based on the "bacterial flora cluster type" and "bacterial flora state score" which are based on large clusters using the concept of transition frequency.

[0124] That is, in this embodiment, a large cluster or some other similar cluster can be adopted as the "microbial flora cluster type." Therefore, the bacterial flora tends to transition within the same "microbial flora cluster type." Therefore, it is more appropriate to treat a user who follows improvement advice as having "transitioned to a better state within the same bacterial flora cluster type" as a result of the improvement, rather than "transitioning to a different bacterial flora cluster type." In other words, even if the subjects are of the same "bacterial flora cluster type," it is preferable to vary the strength (quantity) of the advice for improvement depending on whether the subject already has a good bacterial flora or not.

[0125] In contrast, conventional methods are premised on evaluation in units of flora cluster type (which cluster a user belongs to based solely on flora similarity), such that flora cluster type "VI" is a good flora and flora cluster type "I" is a bad flora. Therefore, the improvement advice for transitioning from flora cluster type "I" to "VI" will remain the same as long as the user to whom the improvement advice is to be provided belongs to flora cluster type "I." In this way, in this embodiment, more appropriate improvement advice is provided without relying solely on the bacterial flora similarity and without being limited to the bacterial flora cluster type.

[0126] In this way, this service uses not only the microbial flora similarity but also time-series information to cluster based on the concept of transition frequency, making it possible to associate the above-mentioned unit data with large clusters, and to provide more detailed improvement advice aimed at achieving a better microbial flora state for each unit data belonging to that large cluster.

[0127] Here, in conventional clustering based only on bacterial flora similarity, for example, the clustering results (the bacterial flora cluster type in Figure 7(A)) may differ between multiple tests due to the influence of bacterial species normally present in humans, which tend to fluctuate between multiple tests. In this embodiment, clustering is performed using the concept of transition frequency, and therefore large clusters are generated that take into account the influence of bacterial species that tend to fluctuate indigenous to humans. In other words, this embodiment can provide improvement advice that suppresses the influence of bacterial species that tend to fluctuate between multiple tests. In this way, this embodiment is capable of performing clustering by taking into account additional aspects (elements) that were lacking in the information on the distance between clusters as the conventional bacterial flora similarity.

[0128] Furthermore, academically, there are two theories: "a theory that it is better for the bacterial flora to change" and "a theory that it is better for the bacterial flora to be stable." The large cluster using the concept of transition frequency evaluated in this embodiment is a concept that can influence these theories.

[0129] The above has described an embodiment of an information processing system to which the present invention is applied. However, the present invention may also be applied to, for example, the following embodiment.

[0130] For example, in the above-described embodiment, the learning data management unit 121 acquires and manages human resident bacterial data B (human resident bacterial data Bk-1 to Bk-Lk for subject Uk) for each of N subjects U1 to Un, which are composed of two or more sets of human resident bacterial data B arranged in two or more time directions, but this is not particularly limited to this. In other words, it is sufficient to obtain in advance unit data (e.g., human resident bacterial data B and questionnaire data A in Figure 3) containing at least information on two or more human resident bacterial species for each of N subjects (N is an integer value of 2 or greater).

[0131] For example, the bacterial flora similarity clustering unit 141 clusters the human resident bacterial data (unit data) of each of the N subjects into small clusters using the algorithm described in Figures 2 and 3, but is not limited to this. For example, an algorithm other than the algorithms described in FIGS. 2 and 3 may be employed. That is, for example, the bacterial flora similarity clustering unit 141 is sufficient to cluster the unit data of each of the N subjects into clusters (small clusters in the above example) smaller than the final clustering (large clusters in the above example) using the bacterial flora similarity.

[0132] For example, the transition frequency clustering unit 142 clusters the time series information (information consisting of two or more occurrences of unit data arranged in the time direction) of the human resident bacterial data B (unit data) of each of the N subjects using the algorithm described in Figures 3 and 3, but this is not particularly limited to this. For example, an algorithm other than the algorithm described in FIG. 3 may be adopted. That is, it is sufficient for the transition frequency clustering unit 142 to perform clustering using the concept of transition frequency based on the time-series information of each of the N subjects.

[0133] Furthermore, for example, the evaluation process is performed on user data of a user different from the data of the N subjects in the learning process, but the present invention is not particularly limited to this. That is, each time data of the user to be evaluated is obtained, a learning process may be performed, and the bacterial flora cluster type calculated in the learning process and improvement advice may be provided to the user to be evaluated.

[0134] Furthermore, for example, in the above-described embodiment, the service targets normal human bacteria, but is not limited to this. That is, the service can target a variety of microbiota. Specifically, for example, the human microbiota can be adopted as the target of the service. That is, the target can be not only bacteria that are constantly present in the human body, such as normal human bacteria, but also the microbiota of microorganisms (bacteria, fungi, algae, protozoa, etc.), including foreign microorganisms. Note that the target microbiota may also include non-living organisms such as viruses. Furthermore, the target is not limited to the microbial flora in a specific part of the human body (for example, the digestive tract such as the intestines), but may be the microbial flora in various parts of the human body (the digestive tract, skin, oral cavity and nasal cavity, or a combination thereof, etc.). Furthermore, the bacterial flora or microorganisms are not limited to those in humans, but may be those in a specific subject, such as plants or environmental microorganisms in soil, ocean, or air. In this case, it becomes possible to classify predetermined objects such as plants, soil, oceans, and the atmosphere based not only on similarity but also on transition frequency.

[0135] Furthermore, for example, the above-described series of processes can be executed by hardware or software. In other words, the functional configuration of FIG. 2 is merely an example and is not particularly limited. That is, it is sufficient for the information processing system to have the function of executing the above-described series of processes as a whole, and the functional blocks and databases used to realize these functions are not particularly limited to the example of FIG. 2. Furthermore, the locations of the functional blocks and databases are not particularly limited to those of FIG. 2 and may be arbitrary. For example, the functional blocks and databases of the information processing device 1 may be transferred to another information processing device (not shown). That is, for example, the information processing device may be divided into an information processing device that performs learning processing and an information processing device that performs evaluation processing. Furthermore, one functional block may be configured as a single piece of hardware, a single piece of software, or a combination thereof.

[0136] Furthermore, for example, when a series of processes is executed by software, the programs that make up the software are installed into a computer or the like from a network or a recording medium. The computer may be a computer built on dedicated hardware. The computer may also be a computer capable of executing various functions by installing various programs, such as a server, a general-purpose smartphone, or a personal computer.

[0137] Furthermore, for example, the recording medium containing such a program may be configured not only as a removable medium (not shown) that is distributed separately from the device itself in order to provide the program to the user (the provider of this service or the user to be evaluated), but also as a recording medium that is provided to the user in a state where it is pre-installed in the device itself.

[0138] In this specification, the steps of describing a program to be recorded on a recording medium include not only processes that are performed chronologically in accordance with the order, but also processes that are not necessarily performed chronologically but are performed in parallel or individually. In addition, in this specification, the term "system" refers to an overall device that is made up of a plurality of devices, a plurality of means, etc. That is, in the above-described embodiment, the information processing system was configured from the information processing device 1 of FIG. 1 having the multiple functional blocks (embodiments of multiple means) of FIG. 2, but this is not particularly limited to this, and for example, the multiple functional blocks (embodiments of multiple means) of FIG. 2 may be distributed and arranged in any of multiple devices connected via a network, etc., to configure the information processing system.

[0139] In other words, the information processing system to which the present invention is applied can take various forms having the following configurations.

[0140] That is, an information processing system to which the present invention is applied (for example, the information processing device 1 shown in FIG. 2 etc.) For each of a plurality of subjects, time-series information (e.g., human resident bacteria data Bk-1 to Bk-Lk in FIG. 3) consisting of two or more unit data (e.g., human resident bacteria data Bk-1 in FIG. 3) arranged in the time direction, each unit data including at least information on two or more human microbial species is previously acquired, A clustering means (e.g., the clustering unit 122 in FIG. 2) that clusters the unit data of the plurality of subjects into a plurality of clusters based on the frequency of transitions of each of the unit data of the plurality of subjects (e.g., the transition frequency in the specification) obtained from the time-series information of each of the plurality of subjects. Equipped with.

[0141] This allows the transition frequency to be taken into account when clustering the human resident bacterial data (unit data) of a certain subject at a certain time.The transition frequency of the subject or user can be taken into account by using the clustering results. Previously, the condition of a subject was understood and improvement advice was calculated based on the similarity of the bacterial flora of the subject's resident human bacterial data (unit data) at a certain point in time, but this method provides more appropriate clustering and makes it possible to provide more useful improvement advice.

[0142] Furthermore, the clustering means a first clustering means (for example, the bacterial flora similarity clustering unit 141 in FIG. 2) that clusters the unit data of the plurality of subjects into M first clusters (for example, small clusters in the specification) (M is an integer value of 3 or more); a second clustering means (e.g., a transition frequency clustering unit 142) for clustering the M first clusters having a predetermined relationship into L second clusters (e.g., large clusters in the specification) (L is a positive integer value less than M); may include:

[0143] As a result, the human indigenous bacteria data (unit data) is first clustered into the first cluster, and then clustered into the second cluster. When clustering with consideration of transition frequency, two-stage clustering is performed, which allows the statistics of the transition frequency data to be adjusted appropriately, enabling clustering that more appropriately takes transition frequency into consideration.

[0144] Furthermore, the first clustering means clustering the human microbial species based on their similarity; The second clustering means A second cluster can be generated by setting the relationship between the frequency of transitions of the unit data of the plurality of subjects, obtained from the time-series information of each of the plurality of subjects, as the predetermined relationship.

[0145] First, unit data including at least information on human resident bacteria is clustered into a first cluster. Then, based on the relationship between the transition frequencies of the unit data of each of the plurality of subjects obtained from the time-series information of the plurality of subjects, the first cluster is further clustered into L second clusters. In this way, learning is performed to cluster unit data belonging to a group that satisfies the condition that each of the plurality of subjects has a predetermined relationship into L second clusters. That is, unlike conventional clustering, data is not simply clustered based on the similarity of each piece of data, but is clustered based on a predetermined relationship different from that based on time-series information. As a result, when the unit data of each of a plurality of subjects are grouped into M clusters, not only the unit data of each subject at that time point but also the predetermined relationship is taken into consideration. For example, clustering can be performed based on whether or not there is mutual transition between unit data as the predetermined relationship. In other words, the second cluster will be a more accurate cluster in terms of reflecting the predetermined relationship.

[0146] As a result, the second cluster is generated from the viewpoint of whether a transition between the unit data constituting the cluster can occur or is likely to occur, which is a predetermined relationship separate from the similarity. As a result, the second cluster becomes a cluster with higher accuracy in terms of whether or not the transition is likely to occur.

[0147] For each of the unit data of the plurality of subjects, data including at least data on health status and lifestyle habits is acquired in advance as questionnaire data, an improvement advice calculation means (for example, the improvement advice calculation unit 123 in FIG. 2 ) that tally the questionnaire data and calculates improvement advice within each of the L second clusters based on the tallying results; The sensor may further include:

[0148] As a result, improvement advice is calculated based on the M clusters that more accurately reflect the physical constitution of the subject. In other words, improvement advice that is more accurately tailored to the physical constitution of the subject is calculated. As a result, the accuracy of advice for improvement based on the subject's normal human bacterial species is improved.

[0149] Also, An acquisition means (for example, the evaluation user information acquisition unit 151 in FIG. 2) for acquiring the unit data of the user to be evaluated as user data; a second cluster determination means (for example, the transition frequency evaluation unit 152 in FIG. 2) for determining, among the L second clusters, a second cluster to which the user data belongs as a cluster to which the user data belongs; The system may further include an improvement advice generation means (e.g., improvement advice generation unit 153 in Figure 2) that generates improvement advice that includes at least the improvement advice calculated by the improvement advice calculation means that is within the range of the belonging cluster determined by the second cluster determination means.

[0150] As a result, for users other than the N subjects, it is calculated which of the M clusters has a high degree of belonging, i.e., which cluster accurately reflects the user's constitution. Then, improvement advice tailored to the cluster with a high degree of belonging, i.e., which cluster accurately reflects the user's constitution, is generated. As a result, more accurate improvement advice is provided to the user to be evaluated. [Explanation of symbols]

[0151] 1 Information processing device, 11 CPU, 18 Storage unit, 20 Drive, 31 Removable media, 111 Learning unit, 112 Evaluation unit, 121 Learning data management unit, 122 Clustering unit, 123 Improvement advice calculation unit, 131 User information acquisition unit for learning, 132 Time series information collection unit, 133 Preprocessing unit, 134 Learning data integration management unit, 141 Microbial flora similarity clustering unit, 142 Transition frequency clustering unit, 151 User information acquisition unit for evaluation, 152 Transition frequency evaluation unit, 153 Improvement advice generation unit, 200 Learning DB, 300 Microbial flora similarity model DB, 400 Transition frequency model DB, 500 Improvement advice DB

Claims

1. In a state where time-series information consisting of two or more unit data items each containing information on two or more human microbial species is previously acquired for each of a plurality of subjects, a first clustering means for clustering the unit data of each of the plurality of subjects into M first clusters (M is an integer value of 3 or more) based on the similarity of the human microbial species; a second clustering means for clustering the M first clusters into L second clusters (L is a positive integer value less than M) based on the frequency of transition of each of the unit data of the plurality of subjects, which is obtained from the time-series information of each of the plurality of subjects; an improvement advice generating means for aggregating the questionnaire data, with data including at least data on health conditions and lifestyle habits of each of the unit data of the plurality of subjects being acquired in advance as questionnaire data, and generating improvement advice within the range of each of the L second clusters based on the aggregation results; An information processing system comprising:

2. The first clustering means clustering the human microbial species based on their similarity; The second clustering means generating the second cluster based on a relationship between the frequency of transitions of the unit data of the plurality of subjects, the relationship being obtained from the time-series information of each of the plurality of subjects; The information processing system according to claim 1 .

3. The frequency of the transition of each of the unit data of the plurality of subjects obtained from the time series information of each of the plurality of subjects is The number of transitions from the first cluster to which the unit data belongs to another first cluster. The information processing system according to claim 2 .

4. An acquisition means for acquiring the unit data of the user to be evaluated as user data; a second cluster determination means for determining, among the L second clusters, a second cluster to which the user data belongs as a cluster to which the user data belongs; a second improvement advice generating means for generating improvement advice that includes at least the improvement advice within the range of the belonging cluster determined by the second cluster determining means, among the improvement advice generated by the improvement advice generating means; The information processing system according to claim 1 , further comprising:

5. An information processing method executed by an information processing system, In a state where time-series information consisting of two or more unit data items each containing information on two or more human microbial species is previously acquired for each of a plurality of subjects, a first clustering step of clustering the unit data of each of the plurality of subjects into M first clusters (M is an integer value of 3 or more) based on the similarity of the human microbial species; a second clustering step of clustering the M first clusters into L second clusters (L is a positive integer value less than M) based on the frequency of transition of each of the unit data of the plurality of subjects, which is obtained from the time-series information of each of the plurality of subjects; an improvement advice generating step of aggregating the questionnaire data, with data including at least data on health conditions and lifestyle habits of each of the unit data of the plurality of subjects being acquired in advance as questionnaire data, and generating improvement advice within the range of each of the L second clusters based on the aggregation result; An information processing method including:

6. A computer, In a state where time-series information consisting of two or more unit data items each containing information on two or more human microbial species is previously acquired for each of a plurality of subjects, a first clustering step of clustering the unit data of each of the plurality of subjects into M first clusters (M is an integer value of 3 or more) based on the similarity of the human microbial species; a second clustering step of clustering the M first clusters into L second clusters (L is a positive integer value less than M) based on the frequency of transition of each of the unit data of the plurality of subjects, which is obtained from the time-series information of each of the plurality of subjects; an improvement advice generating step of aggregating the questionnaire data, with data including at least data on health conditions and lifestyle habits of each of the unit data of the plurality of subjects being acquired in advance as questionnaire data, and generating improvement advice within the range of each of the L second clusters based on the aggregation result; A program that executes control processing including:

Citation Information

Patent Citations

  • Meal support system based on intestinal resident bacterial analysis information

    JP2012165716A

  • Prediction model construction device

    JP2016173728A

  • Etiological analysis device and disease prediction device

    JP2018124702A

  • Information processing device, information processing method, and computer program

    JP2021086556A