Prediction method and program
A machine learning model utilizing lifestyle data predicts Bifidobacterium abundance in the gut by training on dietary and behavioral factors, improving the accuracy of intestinal bacteria presence estimation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- MORINAGA MILK IND CO LTD
- Filing Date
- 2024-10-24
- Publication Date
- 2026-05-12
AI Technical Summary
Existing methods for predicting the presence state of intestinal bacteria, such as those involving the genus Blautia, lack accuracy and effectiveness in utilizing relevant lifestyle and dietary factors.
A prediction method using a machine learning model trained on lifestyle habits, including food and beverage consumption, standing time, and electronic device use before bed, to estimate the abundance of Bifidobacterium bacteria in the human gut, utilizing a database to extract subject groups with high or low proportions and inputting user-specific data to predict Bifidobacterium occupancy rates.
Accurately predicts the abundance of Bifidobacterium bacteria by leveraging key lifestyle factors, enhancing the precision of intestinal flora state prediction.
Smart Images

Figure 2026076720000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method for predicting the presence state of intestinal bacteria.
Background Art
[0002] There is a technology for predicting the presence state of intestinal bacteria in the human intestine. Regarding this, for example, in Patent Document 1, based on data showing the relationship between an index value obtained from stool information regarding a subject's stool and the ratio of the number of bacteria of the genus Blautia in the intestinal flora, a method for predicting the ratio of the number of bacteria of the genus Blautia in the intestinal flora of the subject is disclosed.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] An object of the present disclosure is to predict the abundance of the genus Bifidobacterium bacteria.
Means for Solving the Problems
[0005] One aspect of the present disclosure is a prediction method for estimating the presence state of the genus Bifidobacterium bacteria in the human intestine, A first step involves using a database containing information on the lifestyle habits of the subjects, which includes at least one of the following: (1) the proportion of Bifidobacterium bacteria in the gut microbiota of multiple subjects, (2) information on the amount of food and beverages consumed by the subjects over a predetermined period, and (3) attribute information of the subjects, which includes at least one of the following: (i) the age and body fat percentage of the subjects, (ii) information representing the amount of time the subjects stand over a predetermined period, and (iii) information indicating whether or not electronic devices are used at a predetermined time before going to bed, to extract from the subjects a group of subjects with a relatively high proportion of Bifidobacterium bacteria in their gut microbiota and a group of subjects with a relatively low proportion of Bifidobacterium bacteria in their gut microbiota, and (1) information on the amount of food and beverages consumed, and (2) the extracted group of subjects. The computer performs the following steps: a second step of generating a machine learning model in which the explanatory variables are one or more items selected from the attribute information of the subject, information representing the amount of time spent standing during a predetermined period, and information indicating whether or not electronic devices are used during a predetermined time before going to bed, and the dependent variable is a value relating to the occupancy rate of Bifidobacterium bacteria in the gut microbiota of the extracted subject; and a third step of inputting information including (1) information regarding the intake of predetermined foods and beverages, and (2) information including one or more items selected from the attribute information of the target user, information representing the amount of time spent standing during a predetermined period, and information indicating whether or not electronic devices are used during a predetermined time before going to bed, into the machine learning model to obtain a value relating to the occupancy rate of Bifidobacterium bacteria in the gut microbiota of the target user. This is a prediction method.
[0006] Other embodiments include an information processing device for performing the above-mentioned prediction method, a program for causing a computer to perform the above-mentioned prediction method, or a computer-readable storage medium that non-temporarily stores the program. [Effects of the Invention]
[0007] According to this disclosure, it is possible to predict the abundance of bacteria of the genus Bifidobacterium. [Brief explanation of the drawing]
[0008] [Figure 1] A diagram showing an overview of the processing performed by a server device that implements the prediction method according to the first embodiment. [Figure 2] A diagram illustrating the components of a server device that performs the prediction method according to the first embodiment. [Figure 3] A flowchart of the model generation process in the prediction method according to the first embodiment. [Figure 4] A diagram showing an example of a screen displaying questions related to the amount of food and beverages consumed. [Figure 5] This diagram shows an example of a screen displaying questions related to the amount of time spent standing per day. [Figure 6] This diagram shows an example of a screen displaying questions regarding the use of electronic devices before going to bed. [Figure 7] A flowchart illustrating the process of predicting the occupancy rate of Bifidobacterium bacteria using a generated model in a prediction method according to the first embodiment. [Figure 8] A diagram showing an example of a screen displaying the determination result according to the first embodiment. [Figure 9] A flowchart of the process for generating a model by learning in the prediction method according to the second embodiment. [Figure 10] A figure showing experimental results in the generation of machine learning models. [Modes for carrying out the invention]
[0009] (overview) In recent years, with the advancement of machine learning technology, it is sometimes possible to predict the state of a biological sample being examined using machine learning techniques.
[0010] For example, a predictive model can be generated to predict the state of a person's gut flora based on the correlation between the individual's food and beverage intake, or their responses to questions about their lifestyle, and the state of their gut flora. Then, by inputting the responses of the person being tested to this predictive model, it is possible to predict the state of their gut flora.
[0011] However, in the conventional technology described above, there is room for improvement in the items used to train the predictive model. Therefore, in the prediction method according to the present invention, items that can more accurately predict the state of the intestinal flora of the organism being examined are adopted as items used to train the predictive model.
[0012] A prediction method relating to one aspect of this disclosure is: A predictive method for estimating the presence of Bifidobacterium bacteria in the human gut, comprising: (1) the occupancy rate of Bifidobacterium bacteria in the gut microbiota of multiple subjects; (2) information on the amount of food and beverages consumed by the subjects over a predetermined period; and (3) any of the following: (i) attribute information of the subjects, including at least one of the subjects' age and body fat percentage; (ii) information representing the amount of time the subjects spend standing over a predetermined period; and (iii) information indicating whether or not electronic devices are used at a predetermined time before going to bed. A first step involves using a database containing information on the lifestyle habits of the subjects, including but not limited to, to extract from the subjects a group of subjects with a relatively high proportion of Bifidobacterium bacteria in their gut microbiota and a group of subjects with a relatively low proportion of Bifidobacterium bacteria in their gut microbiota, and then, from among (1) information on the intake of the predetermined food and beverages, and (2) the attribute information of the extracted subject group, information representing the time spent standing during the predetermined period, and information indicating whether or not electronic devices were used during the predetermined time before going to bed. A second step of generating a machine learning model that uses one or more selected items and the value related to the occupancy rate of Bifidobacterium bacteria in the gut microbiota of the extracted subject as explanatory variables and the objective variable, and a third step of inputting information including (1) information on the intake of the predetermined food or drink and (2) one or more selected items from the attribute information of the user to be predicted, the information representing the time spent standing during the predetermined period, and the information indicating the presence or absence of use of an electronic device at a predetermined time before bedtime into the machine learning model to obtain a value related to the occupancy rate of Bifidobacterium bacteria in the gut microbiota of the target user are executed by a computer.
[0013] The control unit acquires information from a database including the occupancy rate of Bifidobacterium bacteria in the gut microbiota of a plurality of subjects, the attribute information of the subjects, and the information on the lifestyle habits of the subjects. Then, based on the above information, the control unit extracts a group of subjects with a relatively high occupancy rate of Bifidobacterium bacteria in the gut microbiota and a group of subjects with a relatively low occupancy rate of Bifidobacterium bacteria in the gut microbiota.
[0014] The attribute information includes, for example, either the age or the body fat percentage of the subject. The attribute information may include the gender of the subject, or the weight, etc.
[0015] The information on the lifestyle habits of the subject is, for example, the information on the intake of a predetermined food or drink during a predetermined period of the subject, the information representing the time spent standing during a predetermined period of the subject, and the information indicating the presence or absence of use of an electronic device at a predetermined time before bedtime. Further, the information may include the frequency of exercise (resistance exercise or aerobic exercise) of the subject, the sleep time, or the bedtime, etc.
[0016] Also, for example, the control unit extracts a group of subjects in the upper predetermined percentage and a group of subjects in the lower predetermined percentage based on the occupancy rate of Bifidobacterium bacteria in the gut microbiota of a plurality of subjects. The control unit may also extract a group of subjects in the middle predetermined percentage.
[0017] Then, based on the contribution rate of the subject's attribute information or information regarding the subject's lifestyle habits to the prediction of the value of the occupancy rate of Bifidobacterium bacteria in the gut microbiota of a plurality of subjects, the control unit generates a machine learning model for predicting the value of the occupancy rate of Bifidobacterium bacteria in the gut microbiota of the user to be predicted from the attribute information of the user to be predicted or information regarding the lifestyle habits of the user to be predicted.
[0018] Subsequently, the control unit inputs information including one or more items selected from information regarding the intake amount of a predetermined food or drink, the attribute information of the user to be predicted, information representing the time spent standing in a predetermined period, and information indicating the presence or absence of use of an electronic device at a predetermined time before going to bed into the generated machine learning model, thereby obtaining the value of the occupancy rate of Bifidobacterium bacteria in the gut microbiota of the target user.
[0019] The value regarding the occupancy rate of Bifidobacterium bacteria in the gut microbiota may be represented as a percentage or by a symbol (or word) indicating the amount divided into several levels. Further, the machine learning model may output the absolute amount of Bifidobacterium bacteria in the gut microbiota of the target user.
[0020] According to such a configuration, the prediction method according to the present disclosure can predict the abundance of Bifidobacterium bacteria.
[0021] Also, the predetermined food or drink may be saccharides or dairy products.
[0022] Thereby, the prediction method according to the present disclosure can predict the value of the occupancy rate of Bifidobacterium bacteria in the gut microbiota using the factors found to have a high contribution rate to the prediction of the value of the occupancy rate of Bifidobacterium bacteria.
[0023] Furthermore, the group of subjects with a relatively high proportion of Bifidobacterium bacteria in their gut microbiota may be the group of subjects whose proportion is in the top one-third of the group of subjects, and the group of subjects with a relatively low proportion of Bifidobacterium bacteria in their gut microbiota may be the group of subjects whose proportion is in the bottom one-third of the group of subjects.
[0024] The following describes specific embodiments of this disclosure with reference to the drawings. Unless otherwise specified, the hardware configurations, module configurations, functional configurations, etc., described in each embodiment are not intended to limit the technical scope of the disclosure to those configurations alone.
[0025] (First embodiment) [Overview of the processes performed by the system] An overview of the processing of the prediction method according to the first embodiment will be described with reference to Figure 1. Figure 1 is a diagram showing an overview of the processing performed by a server device 100 that executes the prediction method according to the first embodiment. An information processing device according to one aspect of this disclosure is implemented as a server device 100. In this embodiment, the server device 100 generates a machine learning model that predicts the occupancy rate of Bifidobacterium bacteria in the gut microbiota of a user to be predicted, and uses the machine learning model to predict a value related to the occupancy rate of Bifidobacterium bacteria in the gut microbiota of a user to be predicted.
[0026] First, let's explain the model generation phase. First, the server device 100 obtains information from multiple subjects, including one or more items selected from the following: the proportion of Bifidobacterium bacteria in the subject's gut microbiota, information on the amount of food and beverages consumed, attribute information such as the subject's body fat percentage or age, information on the subject's standing time during a predetermined period, and information on whether or not the subject uses electronic devices during a predetermined time before going to sleep.
[0027] Next, the server device 100 extracts from the multiple subjects for which the above information was obtained a group of subjects with a high prevalence of Bifidobacterium bacteria and a group of subjects with a low prevalence of Bifidobacterium bacteria.
[0028] The server device 100 then generates a machine learning model in which the proportion of Bifidobacterium bacteria in the subject's gut microbiota is the dependent variable, with explanatory variables being information on the amount of food and beverages consumed, attribute information representing the subject's body fat percentage or age, information representing the subject's standing time during a predetermined period, and information indicating whether the subject uses electronic devices during a predetermined time before going to sleep. The server device 100 then generates a machine learning model that outputs the proportion of Bifidobacterium bacteria as, for example, one of two signs (high or low).
[0029] Next, we will describe the prediction phase for the proportion of Bifidobacterium bacteria in the gut microbiota of the target users.
[0030] First, the server device 100 includes one or more items selected from the following: information on the amount of food and beverages consumed by the subject during a predetermined period; attribute information representing the subject's body fat percentage or age, etc., for the user being predicted; information representing the amount of time the subject spends standing during a predetermined period; and information indicating whether or not the subject uses electronic devices during a predetermined time before going to sleep. To obtain the necessary information.
[0031] Next, the server device 100 obtains a value regarding the proportion of Bifidobacterium bacteria in the target user as input to the machine learning model generated in the generation phase, using the four pieces of information mentioned above. For example, the server device 100 obtains from the machine learning model whether the proportion of Bifidobacterium bacteria in the target user's gut microbiota is high or low.
[0032] As described above, the server device 100 generates a machine learning model in which the percentage of Bifidobacterium bacteria in the gut microbiota of each of the subjects is the objective variable, using explanatory variables such as attribute information of multiple subjects, information on the amount of food and beverages consumed, information on the amount of time spent standing, and information on whether or not electronic devices are used before going to bed. The server device 100 then obtains the percentage of Bifidobacterium bacteria in the gut microbiota of the target user as input to the generated machine learning model, using the information corresponding to the explanatory variables of the target user. As a result, the prediction method executed by the server device 100 can predict the amount of Bifidobacterium bacteria present in the gut of the target user.
[0033] [Server device configuration] Next, the hardware and software configuration of the system including the server device 100 will be described. Figure 2 is a diagram illustrating the components of the server device 100 that performs the prediction method according to the first embodiment.
[0034] The server device 100 can be configured as a computer having a processor (CPU, GPU, etc.), main memory (RAM, ROM, etc.), auxiliary storage (EPROM, hard disk drive, removable media, etc.), and communication circuits. The auxiliary storage contains an operating system (OS), various programs, various tables, etc., and by executing the programs stored therein, various functions (software modules) that match a predetermined purpose, as described later, can be realized. However, some or all of the functions may be realized as hardware modules by hardware circuits such as ASICs and FPGAs.
[0035] The server device 100 is configured to include a control unit 110, a storage unit 120, and a communication unit 130.
[0036] The control unit 110 is a computing unit that realizes various functions of the server device 100 by executing a predetermined program. The control unit 110 can be realized by a hardware processor such as a CPU. The control unit 110 also utilizes RAM, ROM (Read It may be configured to include (Only Memory), cache memory, etc.
[0037] In this embodiment, the control unit 110 of the server device 100 is configured with four software modules: an acquisition unit 111, an extraction unit 112, a generation unit 113, and a calculation unit 114. Each software module may be implemented by the control unit 110 (CPU, etc.) executing a program stored in the storage unit 120. The information processing performed by the software modules is synonymous with the information processing performed by the control unit 110 (CPU, etc.).
[0038] The acquisition unit 111 acquires, for each of several subjects, the occupancy rate of Bifidobacterium bacteria in the gut microbiota, attribute information, information on the amount of food and beverages consumed during a predetermined period, information on the amount of time spent standing during a predetermined period, and information on whether or not electronic devices were used before going to sleep. For example, the acquisition unit 111 communicates with the server device 100. The above information may be obtained from the input section of a terminal that can perform this operation.
[0039] The extraction unit 112 extracts from among multiple subjects whose information is acquired by the acquisition unit 111 a group of subjects with a high proportion of Bifidobacterium bacteria and a group of subjects whose proportion of Bifidobacterium bacteria is low. For example, the extraction unit 112 may extract from among multiple subjects whose information is acquired by the acquisition unit 111 a group of subjects with the top 30% and the bottom 30% proportion of Bifidobacterium bacteria in their gut microbiota.
[0040] The generation unit 113 generates a machine learning model in which at least two of the following are explanatory variables: attribute information representing the body fat percentage or age of the user to be predicted; information regarding the amount of food and beverages consumed by the user to be predicted over a predetermined period; information representing the amount of time the user to be predicted spends standing over a predetermined period; and information indicating whether or not the user uses electronic devices during a predetermined time before going to bed. The target variable is the proportion of Bifidobacterium bacteria in the gut microbiota of the user to be predicted.
[0041] The calculation unit 114 inputs information on the amount of a predetermined food or beverage consumed during a predetermined period, along with one of the following, into a machine learning model generated by the generation unit 113: attribute information of the target user, information on the amount of time spent standing during a predetermined period, and information on whether or not electronic devices were used before going to bed. The calculation unit 114 then obtains the percentage of Bifidobacterium bacteria present in the target user, which is output from the machine learning model.
[0042] The memory unit 120 is a means for storing information and is composed of storage media such as RAM, magnetic disks, and flash memory. The memory unit 12 stores programs executed by the control unit 11, data used by those programs, and so on.
[0043] The communication unit 130 is a wireless communication interface for connecting the server device 100 to an external network. The communication unit 130 is configured to communicate with external terminal devices, etc., via a cellular communication network such as a wireless LAN, 3G, 4G, or 5G.
[0044] Note that the configuration shown in Figure 2 is just one example, and all or part of the illustrated functions may be performed using specially designed circuits. Furthermore, program storage and execution may be performed using combinations of main memory and auxiliary memory other than those shown.
[0045] [Processing in the prediction method] Next, the specific details of the processing in the prediction method according to one embodiment of this disclosure will be described. Figure 3 is a flowchart of the model generation process in the prediction method according to the first embodiment. In Figure 3, the process by which the server device 100 generates a machine learning model for predicting the occupancy rate of Bifidobacterium bacteria in the gut microbiota of the target user will be described based on an operation to start training the machine learning model.
[0046] When an operation is performed on the server device 100 to start training a machine learning model, the server device 100 starts the operation that begins in step S10. Alternatively, the server device 100 may periodically start the operation that begins in step S10.
[0047] First, in step S10, the acquisition unit 111 acquires information for each of the multiple subjects, including the percentage of Bifidobacterium bacteria in the gut microbiota, information on the amount of a predetermined food or beverage consumed during a predetermined period, attribute information, information on the amount of time spent standing during a predetermined period, and information on whether or not electronic devices were used before going to bed. The information to be acquired includes attribute information, information regarding the amount of time spent standing during a specified period, This could include all of the information regarding the use of electronic devices before going to bed, or it could include one or more of the information.
[0048] Figure 4 shows an example of a screen displaying questions regarding the intake of food and beverages. As shown in screen 200 of Figure 4, the acquisition unit 111 may acquire information indicating the amount and frequency of intake of milk or dairy products for each of multiple subjects over a predetermined period. For example, the predetermined period may be one month, and the intake frequency may be a value selected from "did not drink / eat," "less than once a week," "once a week," "2-3 times a week," "4-6 times a week," "daily," or "more than twice a day." Specifically, milk or dairy products may include milk, cheese, yogurt, or ice cream. The recommended intake per serving may be one glass (approximately 200 ml) of milk, one slice (approximately 20 g) of cheese, one cup of yogurt (approximately 120 g), or one cup of ice cream (approximately 80 g). Food and beverages may also contain sugar.
[0049] Figure 5 shows an example of a screen displaying questions about the amount of time spent standing per day. As shown in screen 300 of Figure 5, the acquisition unit 111 may acquire information representing the average amount of time each of multiple subjects spends standing over a predetermined period. For example, the predetermined period may be one day. The amount of time spent standing may also be a value selected from "hardly stood," "less than 1 hour," "1 to 2 hours," "2 to 3 hours," "4 to 6 hours," "7 to 8 hours," or "8 hours or more."
[0050] Figure 6 shows an example of a screen displaying questions regarding the use of electronic devices before going to bed. As shown in screen 400 of Figure 6, the acquisition unit 111 may acquire information on whether each of multiple subjects uses electronic devices before going to bed over a predetermined period. For example, the predetermined period may be one month. The use of electronic devices may also be a value selected from "0 times", "less than once a week", "once a week", "2-3 times a week", "4-6 times a week", and "every day".
[0051] Next, in step S11, the generation unit 113 uses as input data information that includes one or more items selected from information regarding the amount of a predetermined food or beverage consumed during a predetermined period, attribute information of multiple subjects, information regarding the amount of time spent standing during a predetermined period, and information regarding the use of electronic devices before going to bed, and generates training data for supervised learning, with the proportion of Bifidobacterium bacteria in the gut microbiota of the subjects as the ground truth data. The generation unit 113 may also set as input data only one, two, or three of the following: information regarding the amount of a predetermined food or beverage consumed during a predetermined period, attribute information of multiple subjects, information regarding the amount of time spent standing during a predetermined period, or information regarding the use of electronic devices before going to bed. In this case, the generation unit 113 may always include age included in the attribute information and information regarding the amount of a predetermined food or beverage consumed during a predetermined period. The predetermined food or beverage may be milk or dairy products.
[0052] Next, in step S12, the generation unit 113 generates a machine learning model that learns the relationship between the input data and the proportion of Bifidobacterium bacteria in the gut microbiota, using training data that includes information corresponding to the input data in step S11 for multiple subjects and the proportion of Bifidobacterium bacteria in the gut microbiota linked to the information corresponding to each of the input data for multiple subjects.
[0053] The generation unit 113 repeats the process in steps S10 to S12 a predetermined number of times or for a predetermined period of time.
[0054] Next, in step S13, the generation unit 113 generates information regarding the amount of food and beverages consumed by the target user over a predetermined period, and the target user's body fat percentage or age, etc. The system outputs a machine learning model that uses one or more items selected from attribute information, information representing the amount of time the target user spends standing over a specified period, and information indicating whether the target user uses electronic devices during a specified time before going to bed, as explanatory variables, and the proportion of Bifidobacterium bacteria in the target user's gut microbiota as the dependent variable.
[0055] Next, we will describe the process by which the server device 100 uses the machine learning model generated in Figure 3 to predict the occupancy rate of Bifidobacterium bacteria in the gut microbiota of the target user. Figure 7 is a flowchart of the process in the prediction method according to the first embodiment in which the occupancy rate of Bifidobacterium bacteria is predicted by the generated model.
[0056] The server device 100 starts processing when an operation is performed to initiate a prediction using the machine learning model generated in Figure 3 for the user to be predicted.
[0057] In step S20, the calculation unit 114 acquires information including information on the amount of a predetermined food or beverage consumed during a predetermined period, and one or more items selected from the attribute information of the target user, information on the amount of time spent standing during a predetermined period, and information on whether or not electronic devices were used before going to bed. For example, the calculation unit 114 may acquire the above information input from the input unit of a terminal that can communicate with the server device 100. If the explanatory variables of the machine learning model generated in step S13 of Figure 3 are not all of the above information, but only some of it, the generation unit 113 may acquire only the corresponding information.
[0058] Next, in step S21, the calculation unit 114 inputs the information acquired in step S20 into the machine learning model generated in step S13 of Figure 3.
[0059] Next, in step S22, the calculation unit 114 obtains the percentage of Bifidobacterium bacteria in the gut microbiota of the user being predicted from the machine learning model. For example, the calculation unit 114 may obtain the percentage of Bifidobacterium bacteria in the gut microbiota of the user being predicted from the machine learning model, expressed as a percentage.
[0060] Figure 8 shows an example of a screen displaying the judgment results. As shown in screen 500a of Figure 8(a), if the occupancy rate of Bifidobacterium bacteria in the gut microbiota of the user being predicted, as acquired by the calculation unit 114, is higher than a predetermined threshold, a message indicating that the occupancy rate is high may be displayed on the display unit associated with the server device 100. For example, the message "Your gut tends to have a high concentration of Bifidobacteria" may be displayed on the display unit of a terminal (smartphone or tablet, etc.) where a predetermined application has been downloaded.
[0061] Furthermore, as shown in screen 500b of Figure 8(b), if the occupancy rate of Bifidobacterium bacteria in the gut microbiota of the target user, as obtained by the calculation unit 114, is lower than a predetermined threshold, a message indicating that the occupancy rate is low may be displayed on the display unit associated with the server device 100. For example, the message "You tend to have a low amount of Bifidobacterium in your gut" may be displayed on the display unit of a terminal (smartphone or tablet, etc.) where the predetermined application has been downloaded. Also, if the occupancy rate of Bifidobacterium bacteria in the gut microbiota of the target user, as obtained by the calculation unit 114, is low, lifestyle advice may be displayed on the terminal where the predetermined application has been downloaded, along with the above message. For example, the message "We recommend lactulose and yogurt containing Bifidobacterium" may be displayed on the terminal where the predetermined application has been downloaded, along with the above message.
[0062] As described above, the prediction method in this embodiment uses data from multiple subjects The machine learning model is generated using the attribute information of the target user, information on the amount of food and beverages consumed, information on the amount of time spent standing, and information on whether or not electronic devices are used before going to bed as explanatory variables, and the proportion of Bifidobacterium bacteria in the gut microbiota of each target user as the dependent variable.The prediction method in this embodiment then uses the generated machine learning model to obtain the proportion of Bifidobacterium bacteria in the gut microbiota of the target user.As a result, the prediction method in this embodiment can predict the amount of Bifidobacterium bacteria present in the target user.
[0063] (Second Embodiment) In the first embodiment, the server device 100 generated a machine learning model that obtains the proportion of Bifidobacterium bacteria in the gut microbiota of a target user based on the contribution rate of each piece of information acquired by the acquisition unit 111 to the prediction of the proportion of Bifidobacterium bacteria in the subject's gut microbiota. In other words, this machine learning model can be described as a regression model.
[0064] However, the server device 100 can also generate a machine learning model that obtains the proportion of Bifidobacterium bacteria in the gut microbiota of the target user by classifying each piece of information acquired by the acquisition unit 111 according to a predetermined range of values for the proportion of Bifidobacterium bacteria in the gut microbiota associated with that information, using a classification model. This classification model can be obtained as a result of representing the proportion of Bifidobacterium bacteria in the gut microbiota of the target user with a symbol that has a predetermined range of values. In the second embodiment, the server device 100 extracts a group of subjects with a high proportion of Bifidobacterium bacteria in their gut microbiota and a group of subjects with a low proportion of Bifidobacterium bacteria in their gut microbiota, and generates a classification model in which the abundance or absence of Bifidobacterium bacteria in the gut microbiota of the target user is the objective variable, using each piece of information acquired by the acquisition unit 111 as an explanatory variable. Figure 9 is a flowchart of the process of generating a model by learning in the prediction method according to the second embodiment. The process described in Figure 9 is executed instead of the process described in Figure 3.
[0065] When an operation is performed on the server device 100 to start training a machine learning model, the server device 100 starts the operation that begins in step S30. Alternatively, the server device 100 may periodically start the operation that begins in step S30.
[0066] First, in step S30, the acquisition unit 111 acquires information for each of the multiple subjects, including the percentage of Bifidobacterium bacteria in the gut microbiota, information on the amount of food and beverages consumed over a predetermined period, attribute information, information on the amount of time spent standing over a predetermined period, and information on whether or not electronic devices were used before going to bed.
[0067] Next, in step S31, the generating unit 113 determines for each of the multiple subjects whether the proportion of Bifidobacterium bacteria in each subject's gut microbiota is in the top one-third of the multiple subjects. This step results in a positive determination if the generating unit 113 determines that the proportion of Bifidobacterium bacteria in each subject's gut microbiota is in the top one-third of the multiple subjects.
[0068] If the result in this step is positive, the process proceeds to step S32.
[0069] If the result in this step is negative, the process proceeds to step S33.
[0070] If the process transitions to step S32, the generation unit 113 determines the target in step S31 and The subjects who exhibited these characteristics were selected as a group with a high proportion of Bifidobacterium bacteria in their gut microbiota.
[0071] If the process proceeds to step S33, the generation unit 113 determines for each of the multiple subjects whether the proportion of Bifidobacterium bacteria in each subject's gut microbiota is in the lower third of the multiple subjects. This step is affirmative if the generation unit 113 determines that the proportion of Bifidobacterium bacteria in each subject's gut microbiota is in the lower third of the multiple subjects.
[0072] If the result in this step is positive, the process proceeds to step S34.
[0073] If the result in this step is negative, the process proceeds to step S35.
[0074] If the process proceeds to step S34, the generation unit 113 extracts the subjects who were determined in step S33 as a group of subjects with a low proportion of Bifidobacterium bacteria in their gut microbiota.
[0075] If the process proceeds to step S35, the generation unit 113 extracts the subjects who were determined in step S33 as a group of subjects with an intermediate proportion of Bifidobacterium bacteria in their gut microbiota.
[0076] Next, in step S36, the generation unit 113 takes as input data information that includes one or more items selected from the following: information on the amount of a predetermined food and beverage consumed by each of several subjects over a predetermined period, attribute information, information on the amount of time spent standing over a predetermined period, and information on whether or not electronic devices were used before going to bed, and generates training data for supervised learning, with the value of the occupancy rate of Bifidobacterium bacteria in the subject's gut microbiota as the ground truth data. The value of the occupancy rate of Bifidobacterium bacteria indicates whether the subject belongs to a group with a relatively high occupancy rate (upper group) or a group with a relatively low occupancy rate (lower group). The generation unit 113 may also set only one, two, or three of the following as input data: information on the amount of a specified food or beverage consumed during a specified period, along with attribute information of multiple subjects, information on the amount of time spent standing during a specified period, or information on whether or not electronic devices were used before going to bed. In this case, the generation unit 113 may always include the age included in the attribute information and information on the amount of a specified food or beverage consumed during a specified period. The specified food or beverage may be milk or dairy products.
[0077] Next, in step S37, the generation unit 113 generates a machine learning model that learns the relationship between the input data and the group to which the subject belongs, using training data that includes information corresponding to the input data in step S36 for the user to be predicted, and a value indicating the group to which the user to be predicted belongs.
[0078] The generation unit 113 repeats steps S30 to S37 a predetermined number of times or for a predetermined period of time.
[0079] Next, in step S38, the generation unit 113 outputs a machine learning model in which the explanatory variables are one or more items selected from among information on the amount of food and beverages consumed by the target user over a predetermined period, attribute information such as the target user's body fat percentage or age, information on the amount of time the target user spends standing over a predetermined period, and information on whether or not the target user uses electronic devices during a predetermined time before going to bed, and the dependent variable is a value indicating whether the proportion of Bifidobacterium bacteria in the target user's gut microbiota is in the upper or lower group among multiple subjects.
[0080] The machine learning model output in step S38 is used to predict the value of the proportion of Bifidobacterium bacteria in the gut microbiota of the target user, similar to the process described in Figure 7 in the first embodiment.
[0081] As described above, the prediction method in this embodiment uses data from multiple subjects to generate a machine learning model in which the explanatory variables are information on the intake of predetermined foods and beverages, attribute information of the target user, information on standing time, or information on whether or not electronic devices are used before going to bed, and the dependent variable is a value indicating whether the proportion of Bifidobacterium bacteria in the gut microbiota of the target user falls into the upper or lower group among the multiple subjects. Then, the prediction method in this embodiment uses the generated machine learning model to obtain a value indicating whether the proportion of Bifidobacterium bacteria in the gut microbiota of the target user falls into the upper or lower group among the multiple subjects. As a result, the prediction method in this embodiment can efficiently predict the amount of Bifidobacterium bacteria present in the target user.
[0082] Here, we will explain the experimental results regarding the number of explanatory variables used by machine learning models.
[0083] The number of explanatory variables used by machine learning models can increase in an attempt to improve the accuracy of the model. On the other hand, a large number of input items in the prediction process can be cumbersome for the user. Therefore, the inventors conducted experiments to find a pattern that reduces the number of explanatory variables while ensuring the prediction accuracy of the machine learning model.
[0084] Figure 10 shows the experimental results in the generation of machine learning models. Numbers 1 to 5 in Figure 10 show multiple machine learning models obtained by changing the information used as explanatory variables, with the target variable being the percentage of Bifidobacterium bacteria in the gut microbiota of the target user.
[0085] For example, No. 1 represents a machine learning model that uses "age," "milk and dairy product intake," and "body fat percentage" as explanatory variables. AUC is a value that represents the classification accuracy of the machine learning model (Area Under the Curve). The closer the AUC value is to 1, the higher the prediction accuracy.
[0086] As shown in Figure 10, the Area Contribution (AUC) of machine learning models No. 1 to No. 4, which include milk and dairy product intake as an explanatory variable, is closer to 1 than that of No. 5, which does not include milk and dairy product intake as an explanatory variable. Since an AUC closer to 1 indicates higher classification accuracy, it can be concluded that including "milk and dairy product intake" as an explanatory variable is effective for accurately predicting the target variable.
[0087] The experimental results confirmed that including all the explanatory variables used in No. 2 helps improve the accuracy of the machine learning model. However, increasing the number of items makes it cumbersome for the user. From the results of this experiment, it was found that even if the number of explanatory variables is limited to a small number, as in No. 1, including milk and dairy product intake as an explanatory variable ensures a certain level of accuracy in the machine learning model. On the other hand, if the number of explanatory variables is reduced too much, the accuracy of the prediction decreases. Since a certain level of accuracy can be obtained with an AUC of around 0.7, the results of this experiment showed that selecting "age," "milk and dairy product intake," and "body fat percentage" as explanatory variables allows for a balance between prediction accuracy and user convenience.
[0088] Furthermore, from the results of No. 4, we did not include "age" as an explanatory variable, but instead considered "time spent standing," It has also been suggested that an AUC close to 0.7 can be obtained by including three explanatory variables: "use of electronic devices before bedtime" and "amount of milk and dairy products consumed." In other words, it has been found that it is not always necessary to include "age" as an explanatory variable in order to ensure predictive accuracy.
[0089] (Other variations) The embodiments described above are merely examples, and this disclosure may be modified as appropriate without departing from its essence. For example, the processes and means described in this disclosure can be freely combined and implemented as long as no technical inconsistencies arise.
[0090] The present disclosure can also be realized by supplying a computer program implementing the functions described in the embodiments above to a computer, and having one or more processors in the computer read and execute the program. Such a computer program may be provided to the computer by a non-temporary computer-readable storage medium that can be connected to the computer's system bus, or it may be provided to the computer via a network. Non-temporary computer-readable storage mediums include, for example, any type of disk such as magnetic disks (floppy disks, hard disk drives (HDDs), etc.), optical disks (CD-ROMs, DVDs, Blu-ray discs, etc.), read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic cards, flash memory, optical cards, and any type of medium suitable for storing electronic instructions. [Explanation of Symbols]
[0091] 100... Server device 110... Control Unit 111...Acquisition part 112...Extraction part 113...Generation section 114...Calculation Department 120...Storage section 130... Communications Department
Claims
1. A predictive method for estimating the presence of Bifidobacterium bacteria in the human gut, A first step involves using a database containing information on the lifestyle habits of several subjects, which includes at least one of the following: (1) the proportion of Bifidobacterium bacteria in the gut microbiota of multiple subjects; (2) information on the amount of food and beverages consumed by the subjects over a predetermined period; and (3) any of the following: (i) attribute information of the subjects, including at least the age and body fat percentage of the subjects; (ii) information representing the amount of time the subjects spend standing over a predetermined period; and (iii) information indicating whether or not electronic devices are used at a predetermined time before going to bed, to extract from the subjects a group of subjects with a relatively high proportion of Bifidobacterium bacteria in their gut microbiota and a group of subjects with a relatively low proportion of Bifidobacterium bacteria in their gut microbiota. The second step involves generating a machine learning model that uses (1) information on the amount of predetermined food and beverages consumed, and (2) one or more items selected from the attribute information of the extracted subject group, information representing the time spent standing during the predetermined period, and information indicating whether or not electronic devices were used during the predetermined time before going to bed, as explanatory variables, and the value relating to the occupancy rate of Bifidobacterium bacteria in the gut microbiota of the extracted subjects as the dependent variable. A computer performs a third step in which it inputs information including (1) information on the amount of predetermined food and beverages consumed, and (2) information including one or more items selected from the attribute information of the user to be predicted, information representing the amount of time spent standing during the predetermined period, and information indicating whether or not electronic devices are used during the predetermined time before going to bed, to the machine learning model to obtain a value relating to the occupancy rate of Bifidobacterium bacteria in the gut microbiota of the target user. Prediction method.
2. The aforementioned specified food and beverage is sugar or dairy product. The prediction method according to claim 1.
3. The group of subjects with a relatively high proportion of Bifidobacterium bacteria in their gut microbiota is the group of subjects whose proportion is in the top one-third of the group of subjects, and the group of subjects with a relatively low proportion of Bifidobacterium bacteria in their gut microbiota is the group of subjects whose proportion is in the bottom one-third of the group of subjects. The prediction method according to claim 1.
4. A program for causing a computer to execute the prediction method described in any one of claims 1 to 3.