Nutrient predictions based on food purchase data
A machine-learning model using datasets like FIES and Euromonitor predicts nutrient availability from food purchase data, addressing inaccuracies in traditional nutrition assessment methods by providing reliable and standardized nutrient information.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SOCIETE DES PRODUITS NESTLE SA
- Filing Date
- 2025-11-25
- Publication Date
- 2026-06-04
AI Technical Summary
Traditional methods for assessing nutrition, such as dietary surveys and Global Dietary Databases, are cumbersome, unreliable, and lack accuracy due to self-misreporting biases, variability in food compositions, and inconsistent data quality, making it difficult to provide reliable and standardized nutrient information.
A machine-learning model trained on datasets like the Philippines Family Income Expenditure Survey and Euromonitor data is used to predict nutrient availability from food purchase data, incorporating feature mapping and optimization techniques to ensure accuracy and consistency across regions.
The model provides cost-effective, real-time nutrient predictions with high accuracy, enabling comprehensive insights into nutrient intake patterns and dietary diversity, benefiting individuals and stakeholders with actionable nutritional insights.
Smart Images

Figure EP2025084124_04062026_PF_FP_ABST
Abstract
Description
[0001] NUTRIENT PREDICTIONS BASED ON FOOD PURCHASE DATA
[0002]
[0001] Traditional ways of assessing the nutrition of a subject, e.g., an individual, or a group of subjects, e.g., a household, can be cumbersome and unreliable. For example, dietary or nutritional surveys are sometimes used, but can be prone to self-misreporting biases, variability in food compositions, inaccuracies in portion size estimates, limited capture of irregular eating, high costs, and / or lack of availability due to regions, e.g., in low and middle-income countries. Food balance sheets are sometimes used, but such systems can also vary in data quality and availability, lack granularity, focus on food commodities, and often overlook available dietary diversity and / or consumption patterns. Even the use of Global Dietary Databases (GDDs), which aim to provide accurate global dietary intake estimates, can be limited in effectiveness due to data quality and coverage, challenges in standardization, data accessibility issues, and failure to accurately capture cultural variations in some instances. Thus, it would be beneficial to provide systems and methods of more accurately providing nutrition information to subjects in a simpler way that may be less labor intensive and does not rely on manual mapping nutrition based on dietary or nutritional surveys, food balance sheets, or the like.
[0003] BRIEF DESCRIPTION OF THE DRAWINGS
[0004]
[0002] FIG. 1 schematically illustrates a flow diagram of an example machinelearning training and testing system in accordance with the present disclosure;
[0005]
[0003] FIG. 2 schematically illustrates a flow diagram of an example machinelearning model prediction and validation system in accordance with the present disclosure; and
[0006]
[0004] FIG. 3 schematically illustrates a food purchase data collection and analysis system in accordance with the present disclosure. DETAILED DESCRIPTION
[0007]
[0005] In accordance with examples of the present disclosure, nutrient prediction systems and methods that utilize food purchase data can be an effective way of predicting nutrient values in a manner that is cost-effective, fast, and reliable. More specifically, nutrient predictions can be derived from food purchase data of a subject (or group of subjects, e.g.,. households) using a machine-learning model or algorithm that may have been trained, optimized / fine-tuned, e.g., model fitting, tested / validated, etc., to provide subjects with real-time or near-real time information regarding nutrient availability. In some examples, these systems and methods can be adapted to be more specific with respect to the region, e.g., country, where the subject(s) reside. Furthermore, these systems and methods can also provide information or data that support various insights generated from the machine-learning predictions, adding additional value to other interested parties, such as food manufacturers, researchers, government administrative agencies related to food health, cafeterias, etc.
[0008]
[0006] In more specific detail, the development of the nutrient prediction systems described herein can assist with understanding nutrient availability from food baskets, regardless of the type, including both fresh foods and pre-packaged foods. The use of artificial intelligence or machine-learning models for nutrient predictions utilizing food purchase data can provide for the implementation of more comprehensive tools usable by a subject(s) for purposes of analysis of nutrient availability across an average or typical food basket, and such tools can be more accurate as they may relate to differences in various regions or countries. For example, by utilizing expenditure records or sales data that may be available or is made available by the subject(s), it is possible to gain insight into actual food choices and consumption patterns of the subject(s), e.g., individual or household subjects. Furthermore, using the advanced algorithms and data analysis techniques described herein, the nutrient contribution of various foods can be provided to the subject(s) and / or other interested parties. In some examples, a more consistent and standardized analysis of nutrient availability may be realized with the systems and methods described herein, overcoming challenges that may exist related to inconsistencies that may otherwise be problematic, e.g., regional or country-specific choices, feature names inconsistently used from various datasets, etc.
[0009]
[0007] In accordance with this, a method of training and testing a machine-learning model to predict nutrient availability can include dividing a dataset into a training data portion and a testing data portion. The dataset can include a plurality of data points individually representing food purchase data of a household and a plurality of features individually representing a food category. In additional detail, and under the control of at least one processor, the method can include using the training data portion for feature selection and model fitting to generate a machine-learning model for testing and testing the machine-learning model using the testing data portion. In some examples, the dataset can be a modified dataset from multiple source datasets, wherein a first source dataset provides purchase data for fresh and packaged foods for a region or country and a second source dataset includes a correlation between food purchase data and nutrient availability. In some examples, the method can include conducting feature mapping (or variable mapping) to correlate feature names of the first dataset and the second dataset for consistency. As an example, the first source dataset can include the Philippines Family Income Expenditure Survey (FIES) dataset from a single country (the Philippines) and the second source dataset can include the Euromonitor dataset, which is available for multiple countries. In this example and others, a modified dataset that may be used can include a first source dataset, e.g., FIES dataset, modified to include the feature names from a second dataset, e.g., Euromonitor dataset. In some examples, the machine-learning model can be validated in its ability to predict nutrient availability using food purchase data by evaluating at least a portion of one country’s data from the second source dataset to compare against truth data. A prediction of nutrient availability that performs at least about 70% as well as the truth data can be considered to be an acceptable prediction, for example.
[0010]
[0008] In another example, a method of predicting nutrient availability based on food purchase data, under the control of at least one processor, can include collecting food purchase data of a subject and predicting nutrient availability based on the food purchase data using a machine-learning model. The food purchase data can include information regarding food items purchased and location of the food purchase, which are used in making the prediction. In some examples, collecting the food purchase data may include collecting the food purchase data electronically from a purchase data collection device at or after a point of sale by the subject. In other examples, the method can include transmitting the food purchase data from the purchase data collection device over a network to an analysis server. In other examples, the machine-learning model can analyze the food purchase data at the analysis server to predict the nutrient availability. In other examples, the method can include displaying the nutrient availability on a client device. The client device can provide access to a computer interface that includes a plurality of functions selected from a location selection function, a nutrient prediction function, a nutrient adequacy function, a food group comparison function, or a combination thereof.
[0011]
[0009] In another example, a food purchase data collection and analysis system, under the control of at least one processor, can include a purchase data collection device, an analysis server, and a network. The food data collection device can collect food purchase data at a point of sale of a food basket including information regarding food items purchased and location of the food purchase. The analysis server can include a machine-learning model to predict nutrient availability based on the food purchase data, including data related to both the food items purchased and the location of the food purchase. The network can be included for transmitting the food purchase data from the purchase data collection device to the analysis server for analysis using the machine-learning model. In some examples, the system can include an electronic payment terminal data linked to the purchase data collection device, a client device connected or connectable with the analysis server over the network to display the nutrient availability, or a combination thereof. In some additional examples, the client device can provide access to a computer interface that includes a plurality of functions selected from a location selection function, a nutrient prediction function, a nutrient adequacy function, a food group comparison function, or a combination thereof.
[0012]
[0010] It is noted that when discussing examples related to the various systems and methods herein, such discussions can be considered applicable to one another whether or not they are explicitly discussed in the context of that example. Thus, for example, when discussing a “food basket” in the context of a food purchase data collection and analysis system, such disclosure is also relevant to and directly supported in the context of the various methods, and vice versa.
[0013] [Oil] Furthermore, terms used herein will have their ordinary meaning in the relevant technical field unless specified otherwise. In some instances, there are terms defined more specifically throughout the specification, with a few more general terms included at the end of the specification.
[0014]
[0012] As used herein, the term “food” as used herein can include both foods and drinks, and may also include both fresh foods (and drinks) as well as pre-packaged foods (and drinks), unless the context is more specifically identified herein.
[0015]
[0013] The term “subject” or “subjects” or “subject(s)” refers to individual subjects, e.g., an individual user, or groups of subjects, e.g., a household of users or other groups sharing a common kitchen (such as a restaurant or cafeteria of users where nutrient availability would be the same), that may use the food purchase data collection and analysis systems and methods described herein, which may also be referred to herein as the “nutrient prediction system(s).” Thus, reference to a singular “subject” also includes examples where a group of subjects may be sharing food in a common household or other common preparative kitchen.
[0016]
[0014] The term “interested party” or “interested parties” refers to any private organization(s) or interest, e.g., food manufacturers, cafeterias, hospitals, food distributors, food banks, charities, non-government agencies (NGAs), researchers, etc., and / or public or government entities, e.g., public schools, government agencies, military, public shelters, government researchers, etc. that may benefit from information collected from the food purchase data collection and analysis systems and methods of the present disclosure, including information generated by use by any number of subjects.
[0017] Training and Testing Machine-learning Models
[0018]
[0015] Machine-learning (ML) uses statistical algorithms to learn from data and generalize unseen data and provide a modern technique of making targeted predictions. In accordance with the present disclosure, nutrient intake predictions can be made using food purchase data collected which includes information regarding individual food items purchased and regional location of food acquisition. However, for a machine-learning model to make these types of predictions, the model(s) can first be trained using relevant data that correlates food items or classes of foods with nutrient value. Furthermore, as many regions or countries have different food availability, and thus different nutrient availability, a machinelearning model can be likewise trained to consider regional food choices and / or preferences when making these predictions based on nutrient availability.
[0019]
[0016] In accordance with some examples herein, by leveraging expenditure records or sales data of food, or food baskets purchased by a subject (defined as including an individual, a group of individuals, e.g., a household or shared kitchen / restaurant), it becomes possible to gain insights into the actual food choices and consumption patterns of the subject. Without adequate data for training and testing the appropriate machine-learning algorithm(s), it can be difficult to account for the regional, e.g., country, differences related to properly correlating individual foods and / or classes of foods regionally purchased with nutritional content. Another challenge includes accounting for different formats of expenditure data in individual regions, e.g., making it difficult to conduct consistent and standardized analysis.
[0020]
[0017] Thus, in accordance with examples of the present disclosure, it can be convenient to use public datasets that correlate foods with nutrient value and nutrient availability. In training (and ultimately testing / validating) the machine-learning models of the present disclosure, two data sources were used, though it is understood that other data sources can also be used to train and / or test machine-learning models. In accordance with examples herein, the datasets selected for use were the Philippines Family Income Expenditure Survey (FIES) and Euromonitor. Essentially, by utilizing advanced algorithms and data analysis techniques, these datasets (or others) can be used to train and validate / test machine-learning models to analyse the nutrient contributions from various foods or food categories, and with enough training data used in developing the artificial intelligence platform, can further be implemented in a more accurate manner across different regions, e.g., countries, thus allowing for a more comprehensive understanding of the nutritional composition of population diets and providing valuable insights on the existing nutritional gaps, which would be of interest to various stakeholders, such as food manufacturers, health organizations, governments, or the like. Furthermore, this type of information regarding nutrients in specific foods or classes of food, e.g., food groups, pre-packaged vs. fresh foods, etc., can be valuable to individuals or households interested in enhancing their nutritional intake.
[0021]
[0018] The FIES is a Household Consumption and Expenditure Survey implemented by the Philippines government, and part of its value is that it provides reasonable predictors for 26 essential nutrients based on 132,926 data points (with each data point representing an individual household) and 247 features (representing different food categories). This dataset can be used, for example, as a good source of food supply data from a specific country, namely the Philippines. To illustrate, the FIES includes data for both fresh foods and a limited number of packaged foods, and also provides expenditure and volume of purchase data (though it does not provide consumption data). The FIES provides valuable data related to predictors for 26 essential nutrients. Euromonitor, on the other hand, provides volume of food purchase data for many countries, and furthermore, has modeled purchase data for some countries. This dataset includes both fresh foods and a much wider variety of packaged foods. The Euromonitor dataset includes 11 data points per country, with each data point corresponding to an individual year ranging from 2012 to 2022, and also contains 387 features representing different food categories. Either of these datasets can be individually validated prior to their use in the training and / or testing of the artificial intelligence, to potentially ensure that the predictions will be reasonably accurate. For example, the dataset(s) can be validated using other surveys for comparison. As one example, the Euromonitor dataset can be validated using a single year or multiple years against commercial company datasets, e.g., the National Nutrition Survey (NNS), of the same year(s). More specifically, a dataset validation was carried out for the year 2018 comparing the Euromonitor data with the NNS data for foods such as vitamin C-rich fruits; dried beans, nuts, and seeds; fat and oils; eggs; starchy roots and tubers; milk and milk products; meat and meat products; fish and fish products; and vegetables. In this comparison of databases, the liner regression value was confirmed to be r2=0.9449, indicating a strong correlation between two different datasets. Thus, the Euromonitor data was found to be a reliable source of data which may be extrapolated out to indicate good accuracy across various regions / countries.
[0022]
[0019] As two different datasets are described herein by way of example, e.g., FIES and Euromonitor, it is notable that the FIES dataset is particularly useful for its food supply data and predictors of 26 essential nutrients, and the Euromonitor dataset is particularly useful for training and testing the machine-learning algorithm(s) as described hereinafter. With this in mind, there may be issues related to establishing a consistency between two datasets when used for training and testing artificial intelligence, e.g., feature lists may be inconsistent. To promote better feature consistency between multiple datasets, a feature mapping process can be conducted. The feature mapping may involve matching feature names in a first dataset, e.g., the FIES dataset, with the corresponding feature names in a second dataset, e.g., the Euromonitor dataset. This can be carried out manually and / or by electronically comparing regular expressions to align feature names so that they correspond properly across datasets. For example, a mapping table may be generated that associates the features from the first dataset, e.g., the FIES feature name, with the corresponding features from the second dataset, e.g., the Euromonitor feature (or variable) names. In the present example, such a table can provide a reference for assigning the FIES purchase expenditure numbers (based on correlated features) to the appropriate Euromonitor features. Upon completing the feature mapping process, a new dataset can be generated that combines the numerical data from the FIES dataset with the feature (or variable) names from the Euromonitor dataset. The result of this process provides combined dataset including 132,926 data points, 77 features, and 26 response features in this example.
[0023]
[0020] In further detail, in addition to the feature mapping to correlate the use of multiple datasets, a correlation study of nutrient value as it relates to food purchases can be carried out to validate expectations that food purchases or food purchase patterns have a high correlation with nutrient value. In accordance with this, an exploratory data analysis (EDA) can be conducted to understand and gain insight into how food purchases correlate with nutrient availability to a subject(s). In conducting an EDA using protein (fresh and packaged meats) and polyunsaturated fats (edible oils and fats), it was validated that there is a reasonable linear relationship between food expenditure / purchase and nutrient availability.
[0024]
[0021] With the combined dataset in place (via feature mapping) and having validated a suitable linear correlation between food purchases and nutrient availability, the information is in place to select the most predictive features and machine-learning model(s) for generating predictions with a high confidence of accuracy. For example, any of a number of methodologies can be used to identify the most predictive features selected for use, such as Recursive Feature Elimination (RFE), Random Forest, and / or LASSO, to name a few. This is sometimes referred to as “feature optimization.” It is noted, however, that optimization does not infer or require that the most optimal set of features be selected, but rather that the features selected provide an acceptable enough result having a high enough degree of accuracy to be useful. For example, feature optimization combined with selection of an acceptable machine-learning model can result in predictions within reasonable range that are at least about 70% accurate, at least about 80% accurate, or at least about 90% accurate, on average across all food purchases considered.
[0025]
[0022] The subset of features that are selected using feature optimization can be evaluated against multiple machine-learning models to determine which model performed better with those specific features. For example, the most relevant features can be selected for each nutrient, e.g., using RFE or other feature optimization process. By selecting features individually based on the specific nutrient, the machine-learning model can evaluate the most valuable predictors in correlating food purchase with nutrient availability. Additional fine- tuning of the machine-learning model can include the validation of the various model parameters. Example validation processes that can be used include 5-fold cross-validation, leave-one-out, boostrap resampling, and / or stratified cross-validation, for example. Further evaluation can be carried out using mean square error (MSE), and multiple machine-learning models can be used to determine which model provides better, more predictively accurate results. The equation for MSE is shown in Formula 1, by way of example, as follows:
[0026] Formula 1 where n=number of data points, Yi=observed values, and Yi=predicted values.
[0027]
[0023] Table 1 below illustrates an example where four nutrients, namely protein, fiber, iron, and vitamin B3, were evaluated using both the LASSO model and the XGBoost model for comparative purposes, and the MSE value for each nutrient and each model was determined as follows: Table 1
[0028]
[0024] Based on the comparison of the results and the lower MSE values obtained for the XGBoost model compared to the LASSO model for these four nutrients (as well as the remaining 22 nutrients provided by the FIES, totaling 26 nutrients), the XGBoost model can be selected for use as being more accurate for predicting nutrient availability / intake, as it outperformed the LASSO model using the dataset as described above. This conclusion using this dataset is strengthened by the fact that all 26 nutrients generated lower MSE values in the XGBoost model compared to the LASSO model. Once a machine-learning model is selected for use, any available data that may be reliable can be used to train the machine-learning model selected (and / or other models for comparison). In examples herein, both the FIES datasets and the Euromonitor datasets can be used, particularly after the data properly correlated to operate together. In further detail, in some examples, the Euromonitor data can be used to train and test the machine-learning model.
[0029]
[0025] FIG. 1 illustrates a flow diagram for the machine-learning training and testing system 100 which, as shown in this example, can include multiple development phases such as feature mapping 110 to generate a modified dataset 120, splitting dataset for training data 130 and testing data 140 for training vs. testing, feature selecting 150, model fitting / fine tuning parameters 160 resulting in multiple tuned models 170 for testing, and tuned model comparison 180 resulting in a final machine-learning model 190 selected for making nutrient predictions based on food purchase data. As an initial note, training and testing the machine-learning algorithm in accordance with FIG. 1 (or other similar artificial intelligence build) can be carried out for each individual nutrient for more accurate predictive results. Thus, if there are 26 nutrients being evaluated, each nutrient may be built based on its own set of features, its own fine tuning, and / or its own machine-learning model, for example. In the instant case, it was found that one machine-learning model, e.g., XGBoost, outperformed other models across all 26 nutrients of the FIES dataset, so the use of multiple machine-learning algorithms on a nutrient-by-nutrient basis was not a consideration due to the results provided by the XGBoost algorithm. With that stated, it may be the case that with other sets of training data, features, etc., similar results could be achieved by using multiple machine-learning models.
[0030]
[0026] In further detail regarding the training and testing of machine-learning models or algorithms 100, as shown, feature mapping 110 can be used to merge or otherwise utilize data from multiple sources of data. As mentioned, feature mapping may involve matching feature names in a first dataset, e.g., the FIES dataset, with the corresponding feature names in a second dataset, e.g., the Euromonitor dataset, as described previously. Thus, in this instance, the FIES dataset was used for both training and testing, but the dataset was modified to be consistent with the Euromonitor dataset feature naming convention so that the Euromonitor dataset could be used effectively as the dataset for making predictions. Thus, upon completing the feature mapping (if more than one dataset is used), a modified dataset 120 can be generated for training and testing that includes all of the FIES datapoints, but which is modified to use feature naming conventions found in the Euromonitor dataset. Once trained and tested, with the consistent naming convention, the Euromonitor dataset can then be seamlessly used as the main dataset for making predictions on a country by country basis, as the feature names will align properly. In this specific example, the FIES data used in this example included 132,926 data points and 247 features, and the Euromonitor data included 10 data points and 387 features. The modified FIES data used for training and testing included 132,926 data points, 77 features, and 26 response features.
[0031]
[0027] The modified dataset in this example is then split into multiple portions of data. For example, a first portion of the dataset can be used as training data 130 to train a machine-learning model(s), and a second portion of the dataset can be used as testing data 140 to test or validate the predictive accuracy of the model(s). In the example shown, about 80% of the modified dataset is shown as being used for training the model, and about 20% of the modified dataset (the portion not used for training) is shown as being used for validating and / or testing the predictive accuracy of the model, though other percentagebased splits of the modified dataset can be used for training and testing, e.g., 60% to 90% of modified data used for training and 10% to 40% of modified data used for testing. As described previously, exploratory data analysis (EDA) (not shown in FIG. 1) can be conducted to understand and gain insight into how food purchases correlate with nutrient availability. In this instance, the correlation between purchases and nutrient availability was found to be reasonably linear in its correlation. Conducting the EDA also provides information regarding the selection of features for providing and improving the predictive capabilities of the machine-learning models that are tested and ultimately assists with selecting the most predictive model for use. This can be done for each of the nutrients for which the model is designed to predict nutrient availability.
[0032]
[0028] With the modified dataset in place (via feature mapping against the Euromonitor dataset), and in some instances having validated a suitable linear correlation between food purchases and nutrient availability, the information is in place for feature selection 150 that may be more suitable for use with various machine-learning models that would be suitable for predicting nutrient availability based on food purchase data with high confidence of accuracy. In machine-learning, a “feature” is a characteristic or attribute of a dataset that can be used to train a model. Finding and / or selecting features for use in the context of a machine-learning model that provides the data that leads to better results is often selected and weighted heavier for use. Other features that may be redundant or reduce the predictive value of a machine-learning model can remain unused or given less weight. Thus, the selection of appropriate features against the backdrop of a specific machine-learning model(s) can lead to improved predictive capabilities. Typically, features relate to the inputs entered into a machine-learning algorithm that contribute significantly to the accuracy and performance of a given model. In some examples, features may be selected using any of a number of methodologies, including recursive feature elimination (RFE), random forest, LASSO, for example. A few example features that may be selected in accordance with the present disclosure include iron intake-related food and Vitamin C-related food profiles. Notably, features can be selected generally across multiple nutrients, but for better accuracy, each specific nutrient may include its own list of relevant features for use.
[0033]
[0029] In some examples, model fitting / fine tuning parameters 160 of machinelearning models can be carried out to further enhance the predictive accuracy of the model. For example, the model can undergo validation of the various model parameters. Training and / or validating one or more of the machine-learning algorithms can include or be based on a linearly fitted model utilizing a stochastic gradient decent algorithm, a tree-based pipeline optimization, and a k-fold cross-validation, or the like. In one specific example, the model fitting can be carried out using 5 -fold cross-validation.
[0034]
[0030] With the modified FIES data assimilated and the features selected and finetuned, multiple tuned models 170 can be chosen for additional testing. Various methods can be used for tuned model comparison 180, e.g., 20% of untouched modified dataset can be used for comparison purposes. Example model comparison approaches include using methods such as mean square error (MSE), mean absolute error (MAE), or root mean square error (RMSE). An example of using MSE for comparing machine-learning models is shown by way of example at Formula 1 above. Upon comparing the machine-learning models, a final machine-learning model 190 can be selected, which can be any of a number of machinelearning models, including XGBoost, LASSO, LightGBM, decision tree (DT), artificial neural network (ANN), long short-term memory (LSTM), etc.
[0035] Machine-learning Model Predictions and Validation
[0036]
[0031] Once a good machine-learning model has been selected that can predict results with reasonable accuracy based on the training data and the testing data, the model can be used to make predictions using new data entered during use of the model, such as data corresponding to food basket purchases and / or food purchase patterns of a subject(s), e.g., individuals, groups sharing a kitchen or food resources, etc. Referring now to FIG. 2 by way of example, a nutrient prediction validation system 200 is shown that utilizes the machine-learning training and testing system 100 as illustrated by way of example in FIG. 1. Regarding nutrient prediction, food purchase data 210 can be inputted into the final machine-learning model selected as a result of the machine-learning training and testing, e.g., XGBoost with appropriately selected features that have been fine-tuned, etc. Using the inputted food purchase data and the final machine-learning model, nutrient prediction 220 can be made using the Euromonitor dataset on a regional or country-by-country basis, for example. Prediction validation 230 based on the Euromonitor datasets can be carried out using truth data, e.g., scientific literature from 1 to a few countries. For example, by comparing the nutrient prediction generated by the machine-learning model with truth data, the accuracy of the predictions can be verified and quantified. An example acceptable prediction can be one that aligns with the truth data by greater than about 70% in accuracy.
[0037]
[0032] With the machine-learning model, loaded with purchase data resulting in nutrient prediction, and furthermore validated against truth data, the nutrient prediction system and method may be ready for food purchase data collection and analysis 300 to predict nutrient availability via use of the machine-learning model, which model may have been validated by truth data. Essentially, the machine-learning model to be used by a subject may be loaded on a client device or may be loaded on an analysis server across a network that is accessed by the subject, for example. Interaction with the machine-learning model by subjects, for example, can occur at a client device via a computer application, an internet website or gateway, a computer dashboard, etc. As a note, data collected based on the use of the machine-learning models as described herein may also be available to other interested parties for viewing via a similarly configured client device(s). Thus, interested parties, such as private organization(s) and / or government entities, as described previously, may also benefit from information collected from the food purchase data collection and analysis systems and methods of the present disclosure. For example, a subject may benefit directly from the nutrient availability predictions generated by the machine-learning model described herein, but other interested parties may also benefit from the information generated by the subject, as well as many other subjects that may also be using the food purchase data collection and analysis systems described herein.
[0038]
[0033] Regarding the client device and interface therewith by the subject or other interested party, the example of a computer dashboard is considered by way of example. In this example, the dashboard can provide the subject and / or any other interested parties interfacing with the dashboard with multiple functions. Example functions can include regional (or country) selection(s), nutrient prediction(s), nutrient adequacy, and / or food group comparison(s), to name a few. A regional (or country) selection function, for example, can allow a subject to select specific regions or contraries of interest from a dropdown menu or interactive map. This feature can allow for a more focused analysis of nutrient intake patterns in different regions / countries, enabling the subject or other interested party to compare the results across regions. A nutrient prediction function, as another example, may display the predicted nutrient intake values for the selected region / country. This nutrient prediction function can provide visualizations, such as line graphs or bar charts to showcase the trends and variations in nutrient intake over time. In some examples, subjects can explore different nutrients individually or compare multiple nutrients side by side. In another example, a nutrient adequacy function can be included to assess the adequacy of nutrient intake based on recommended dietary guidelines or reference values by comparing the predicted nutrient intake values with the corresponding recommended intake levels, highlighting any potential inadequacies or excesses. This information can help subject(s) or other interested parties to identify areas where interventions or adjustments to dietary patterns may be helpful. A food group comparison function can also be included, which may allow users to explore the contribution of different food groups to nutrient intake in the selected country. This function can present visualizations, such as pie charts or stacked bar graphs, etc., to illustrate the proportion of nutrients derived from various food groups. This insight may not only benefit the subject(s) using the food purchase data collection and analysis systems / methods, but can also guide interested parties, such as policymakers, companies in the food industry, etc., in developing strategies to promote balanced and diverse diets on a regional or country-by-country basis. As mentioned, these types of visualizations, summaries, interactive features, or the like that may be viewable on a client device, such as a dashboard, computer application, website, etc., can make it easier for a subject(s) to explore and understand the nutrient intake patterns derived from the machinelearning model, and / or can facilitate data-driven decision-making by providing the subject(s) with actionable insights into the nutritional landscape of their country or countries where they may be visiting. Furthermore, other interested parties may also use this data from the subject(s) along with many other subject(s) using these systems and methods to make data- based decisions regarding nutritional health of larger populations on a region or country-bycountry basis.
[0039] Food Purchase Data Collection and Analysis for Nutrient Predictions using Machine-learning Model
[0040]
[0034] In accordance with the present disclosure, an example food purchase data collection and analysis system 300 is shown that can be used with the nutrient prediction systems and methods described herein. As a note, this is merely one example of how food purchase data can be collected from a subject(s), e.g., an individual, a household, a group using a common kitchen such as a restaurant or cafeteria, etc., for subsequent utilization by the subject, group of subjects, or various stakeholders, e.g., food manufacturers, health organizations, governments, etc. With that stated, it is understood that there are many other arrangements that can be utilized with the food purchase data collection and analysis systems and methods herein.
[0041]
[0035] As an example, a food purchase data collection and analysis system 300 can include an analysis server system 310 (which can be a single server, part of a larger network of servers, cloud computing system, etc.), a client device 330, and a network 320 connecting the client device(s) to the analysis server system. For example, the analysis of a food basket 360 or multiple food baskets purchased by a user may be provided to the analysis server system by the subject, or may be provided by purchase data collected as a result of the food basket purchases. Though a network is shown connecting the client device(s) with the analysis server systems, in some examples, the systems can be implemented locally on the client device(s) and / or can be implemented locally by connecting to an analysis server via a non-network link (wired or wireless). The term “non-network” does not infer that the client device(s) are not also connected to the server analysis system via the network 320, but merely that there may be a local non-network wireless or wired connection that can more directly connect one or more of the client devices with the analysis server system. Examples of non-network links or connections that can be used include more local communication connections via wired connections, Bluetooth, WIFI, etc.
[0042]
[0036] The collection of purchase data from customers or subjects can occur at the point of purchase of the food basket 360 via purchase data collection device 350 upon payment using an electronic payment terminal 340, which is sometimes referred to as a point of sale (POS) terminal, a credit card machine, a car reader, a PIN pad, an EFTPOS terminal (or PDQ terminal), etc. The electronic payment terminal, of whatever type, can interface with the subject in some manner, such as via a payment card, smartphone, biometrics, etc., to authorize electronic funds transfers. Regardless of the specific technology, the electronic payment terminal can capture information related to payment from the subject for payment, and typically utilizes a network 320 connection to access payment authorization. Notably, the network shown in FIG. 3 schematically illustrates a single network, but it is noted that the network that is used to transmit data to the merchant services provider (or bank) for authorization may be a different network than is used for operation of the nutrient prediction systems of the present disclosure. For example, many electronic payment terminals transmit data over a cellular network, via Bluetooth or Wi-Fi connection, or in some instances a satellite network, though some systems may still communicate over standard telephone lines, Ethernet connections, or the like.
[0043]
[0037] In accordance with examples herein, the electronic payment terminal 340 can be linked to the purchase data collection device 350 (also connected via a network 320) so that the purchasing of a food basket 360 can become associated with the specific food items purchased via a data link 370. For example, a purchase data collection device may be in the form of a system that can be used to scan or otherwise record individual food items selected for purchase (which is typically used to generate electronic data associated with the purchased foods, and may be used to generate a consumer receipt, e.g., paper receipt, electronic receipt, etc.). The total for payment may be calculated by the purchase data collection device, and the data link can provide that total for payment to the electronic payment terminal for collecting payment. However, the purchase data collection device that has captured the food purchase data can be stored or uploaded over a network to the analysis server 310 where the food purchase data can be analyzed and processed using the artificial intelligence or the machine-learning model (shown at 300 of FIG. 2) that may have been trained and tested (shown at 100 of FIG. 2) and subjected to the nutrient prediction validation (shown at 200 of FIG. 2) in accordance with the present disclosure. The processed data analyzed by the machine-learning model can then be delivered to the client device 330 with results, via a computer dashboard, a computer application, a website, a computer gateway, or the like.
[0044]
[0038] In further detail regarding the client device 330 that can be used, the food purchase data collection and analysis systems 300 can independently include a single computing device or can include multiple computing devices, a cluster of computing devices, or the like. Thus, the use of the term “client device” (in the singular) refers to either a single client device or multiple client devices, regardless of how they are associated with one another or the network. The client device can thus include one or more physical processors communicatively coupled to one or more memory devices, input / output devices, or the like.
[0039] As used herein, the term “processor” may be referred to as a central processing unit (CPU) or other similar terminology. In accordance with this, any of the processors in use in any of the devices that may be included as part of the nutrient prediction system, including as part of the analysis server, the client device, the electronic payment terminal, or the purchase data collection device, for example. Each may include one or more devices capable of executing instructions and encoding arithmetic, logical, and / or I / O operations. In one illustrative example, a processor may implement a Von Neumann architectural model and may include an arithmetic logic unit (ALU), a control unit, and a plurality of registers. In some aspects, a processor may be a single core processor that is typically capable of executing one instruction at a time (or process a single pipeline of instructions) and / or a multi-core processor that may simultaneously execute multiple instructions. In some examples, a processor may be implemented as a single integrated circuit, two or more integrated circuits, and / or may be a component of a multi-chip module in which individual microprocessor dies are included in a single integrated circuit package and hence share a single socket.
[0045]
[0040] The term “memory” or “memory device” as used herein can refer to a volatile or non-volatile memory device, such as RAM, ROM, EEPROM, or any other device capable of storing data, e.g., purchase data, and / or carrying instructions that can be executed by the processor. Input / output devices can include a network device, e.g., a network adapter or any other component that connects a computer to a network 320, a peripheral component interconnect (PCI) device, storage devices, disk drives, etc. In some instances, there may be other useful components as well, such as connected sound or video adaptors, photo / video cameras, printer devices, keyboards, displays, etc. In several aspects, a computing device provides an interface, such as an API or web service, which provides some or all of the purchase data to other computing devices for further processing. Access to the interface can be open and / or secured using any of a variety of techniques, such as by using client authorization keys, as appropriate to the requirements of specific applications of the disclosure.
[0046]
[0041] In further detail regarding the network 320, this can be established to include a LAN (local area network), a WAN (wide area network), a telephone network, e.g., Public Switched Telephone Network (PSTN), a Session Initiation Protocol (SIP) network, a wireless network, a point-to-point network, a star network, a token ring network, a hub network, wireless networks (including protocols such as EDGE, 3G, 4G LIE, Wi-Fi, 5G, WiMAX, or the like), the internet, or the like. A variety of authorization and authentication techniques, such as username / password, Open Authorization (OAuth), Kerberos, SecurelD, digital certificates, or more, may be used to secure the communications. It will be appreciated that the network connections shown in the example food purchase data collection system or method 300 is merely illustrative, and thus, any other known electronic communication setups or methodologies of establishing one or more communication links between the various components shown, in examples where each of these components are indeed used, may be implemented.
[0047] Example Embodiments
[0048]
[0042] In accordance with the disclosure herein, the following examples are illustrative of several embodiments of the present technology.
[0049] 1. A method of training and testing a machine-learning model to predict nutrient availability, comprising: dividing a dataset into a training data portion and a testing data portion, wherein the dataset includes: a plurality of data points individually representing food purchase data of a household, and a plurality of features individually representing a food category; and under the control of at least one processor: using the training data portion for feature selection and model fitting to generate a machine-learning model for testing, and testing the machine-learning model using the testing data portion.
[0050] 2. The method of example 1, wherein the dataset is a modified dataset from multiple source datasets, wherein: a first source dataset provides purchase data for fresh and packaged foods for a region or country, and a second source dataset includes a correlation between food purchase data and nutrient availability, wherein the method includes conducting feature mapping to correlate feature names of the first dataset and the second dataset for consistency.
[0051] 3. The method of example 2, wherein training and testing includes using the modified dataset, wherein the modified dataset includes the first source dataset modified to include the feature names from the second dataset.
[0052] 4. The method of one of examples 2 or 3, wherein the first source dataset includes a Philippines Family Income Expenditure Survey (FIES) dataset from a single country, and wherein the second source dataset includes a Euromonitor dataset for multiple countries.
[0053] 5. The method of one of examples 1 to 4, further comprising validating the machinelearning model to predict nutrient availability using food purchase data by evaluating at least a portion of one country’s data from the second source dataset to compare against truth data, wherein a prediction of nutrient availability that performs at least about 70% as well as the truth data is considered to be an acceptable prediction.
[0054] 6. The method of one of examples 1 to 5, wherein feature selection is carried out by recursive feature elimination (RFE), random forest, LASSO, or a combination thereof and wherein the model fitting is carried out by k-fold cross-validation, leave-one-out, bootstrap resampling, or a combination thereof.
[0055] 7. The method of one of examples 1 to 6, wherein the machine-learning model that is generated is selected over other machine-learning models by using mean square error (MSE) analysis, mean absolute error (MAE) analysis, root mean square error (RMSE) analysis, or a combination thereof.
[0056] 8. The method of one of examples 1 to 7, wherein the training portion includes from about 60% to about 90% of the dataset and the testing portion includes from about 10% to about 40% of the dataset.
[0057] 9. The method of one of examples 1 to 8, wherein the machine-learning model includes XGBoost.
[0058] 10. A method of predicting nutrient availability based on food purchase data, under the control of at least one processor, comprising: collecting food purchase data of a subject, wherein the food purchase data includes information regarding food items purchased and location of the food purchase; and predicting nutrient availability based on the food purchase data using a machinelearning model that includes data related to both the food items purchased and the location of the food purchase.
[0059] 11. The method of example 10, wherein collecting the food purchase data includes collecting the food purchase data electronically from a purchase data collection device at or after a point of sale by the subject.
[0060] 12. The method of example 11, further comprising transmitting the food purchase data from the purchase data collection device over a network to an analysis server.
[0061] 13. The method of example 12, wherein the machine-learning model analyzes the food purchase data at the analysis server to predict the nutrient availability.
[0062] 14. The method of one of examples 10-13, further comprising displaying the nutrient availability on a client device.
[0063] 15. The method of example 14, wherein the client device provides access to a computer interface that includes a plurality of functions selected from a location selection function, a nutrient prediction function, a nutrient adequacy function, a food group comparison function, or a combination thereof.
[0064] 16. A food purchase data collection and analysis system, under the control of at least one processor, comprising: a purchase data collection device to collect food purchase data at a point of sale of a food basket, wherein the food purchase data includes information regarding food items purchased and location of the food purchase; an analysis server including a machine-learning model predict nutrient availability based on the food purchase data, including data related to both the food items purchased and the location of the food purchase; and a network for transmitting the food purchase data from the purchase data collection device to the analysis server for analysis using the machine-learning model.
[0065] 17. The system of example 16, further comprising an electronic payment terminal data linked to the purchase data collection device.
[0066] 18. The system of one of examples 16 or 17, further comprising a client device connected or connectable with the analysis server over the network to display the nutrient availability. 19. The system of example 18, wherein the client device provides access to a computer interface that includes a plurality of functions selected from a location selection function, a nutrient prediction function, a nutrient adequacy function, a food group comparison function, or a combination thereof.
[0067]
[0043] It is noted that although the nutrient prediction systems and methods described herein include various example details, methods may be performed by various software and / or hardware configurations, including various processing logic that may include hardware (circuitry, dedicated logic, etc.), software, or a combination of both. For example, methods may be implemented and executed as instructed on a machine, where the instructions are included on at least one computer readable medium or one non-transitory machine-readable storage medium.
[0068]
[0044] It will be appreciated that all of the disclosed methods and procedures described herein can be implemented using one or more computer programs, components, applications, program modules, etc. These components may be provided as a series of computer instructions on any conventional computer readable medium or machine-readable storage medium, including volatile or non-volatile memory, such as RAM, ROM, flash memory, magnetic or optical disks, optical memory, or other storage media. The instructions may be provided as software or firmware and / or may be implemented in whole or in part in hardware components such as ASICs, FPGAs, DSPs, or any other similar devices. The instructions may be configured to be executed by one or more processors which, when executing the series of computer instructions, performs or facilitates the performance of all or part of the disclosed methods and procedures. As will be appreciated, the functionality of the program modules may be combined or distributed as desired in various aspects of the disclosure.
Claims
CLAIMSWhat Is Claimed Is:
1. A method of training and testing a machine-learning model to predict nutrient availability, comprising: dividing a dataset into a training data portion and a testing data portion, wherein the dataset includes: a plurality of data points individually representing food purchase data of a household, and a plurality of features individually representing a food category; and under the control of at least one processor: using the training data portion for feature selection and model fitting to generate a machine-learning model for testing, and testing the machine-learning model using the testing data portion.
2. The method of claim 1, wherein the dataset is a modified dataset from multiple source datasets, wherein: a first source dataset provides purchase data for fresh and packaged foods for a region or country, and a second source dataset includes a correlation between food purchase data and nutrient availability, wherein the method includes conducting feature mapping to correlate feature names of the first dataset and the second dataset for consistency.
3. The method of claim 2, wherein training and testing includes using the modified dataset, wherein the modified dataset includes the first source dataset modified to include the feature names from the second dataset.
234. The method of claim 2, wherein the first source dataset includes a Philippines Family Income Expenditure Survey (FIES) dataset from a single country, and wherein the second source dataset includes a Euromonitor dataset for multiple countries.
5. The method of claim 1, further comprising validating the machine-learning model to predict nutrient availability using food purchase data by evaluating at least a portion of one country’s data from the second source dataset to compare against truth data, wherein a prediction of nutrient availability that performs at least about 70% as well as the truth data is considered to be an acceptable prediction.
6. The method of claim 1, wherein feature selection is carried out by recursive feature elimination (RFE), random forest, LASSO, or a combination thereof and wherein the model fitting is carried out by k-fold cross-validation, leave-one-out, bootstrap resampling, or a combination thereof.
7. The method of claim 1, wherein the machine-learning model that is generated is selected over other machine-learning models by using mean square error (MSE) analysis, mean absolute error (MAE) analysis, root mean square error (RMSE) analysis, or a combination thereof.
8. The method of claim 1, wherein the training portion includes from about 60% to about 90% of the dataset and the testing portion includes from about 10% to about 40% of the dataset.
9. The method of claim 1, wherein the machine-learning model includes XGBoost.
10. A method of predicting nutrient availability based on food purchase data, under the control of at least one processor, comprising: collecting food purchase data of a subject, wherein the food purchase data includes information regarding food items purchased and location of the food purchase; andpredicting nutrient availability based on the food purchase data using a machinelearning model that includes data related to both the food items purchased and the location of the food purchase.
11. The method of claim 10, wherein collecting the food purchase data includes collecting the food purchase data electronically from a purchase data collection device at or after a point of sale by the subject.
12. The method of claim 11, further comprising transmitting the food purchase data from the purchase data collection device over a network to an analysis server.
13. The method of claim 12, wherein the machine-learning model analyzes the food purchase data at the analysis server to predict the nutrient availability.
14. The method of claim 10, further comprising displaying the nutrient availability on a client device.
15. The method of claim 14, wherein the client device provides access to a computer interface that includes a plurality of functions selected from a location selection function, a nutrient prediction function, a nutrient adequacy function, a food group comparison function, or a combination thereof.
16. A food purchase data collection and analysis system, under the control of at least one processor, comprising: a purchase data collection device to collect food purchase data at a point of sale of a food basket, wherein the food purchase data includes information regarding food items purchased and location of the food purchase; an analysis server including a machine-learning model predict nutrient availability based on the food purchase data, including data related to both the food items purchased and the location of the food purchase; anda network for transmitting the food purchase data from the purchase data collection device to the analysis server for analysis using the machine-learning model.
17. The system of claim 16, further comprising an electronic payment terminal data linked to the purchase data collection device.
18. The system of claim 16, further comprising a client device connected or connectable with the analysis server over the network to display the nutrient availability.
19. The system of claim 18, wherein the client device provides access to a computer interface that includes a plurality of functions selected from a location selection function, a nutrient prediction function, a nutrient adequacy function, a food group comparison function, or a combination thereof.26