Personalized feeding recommendation assistant and method by using image recognition technology thereof
The method uses image recognition and user-specific training to enhance the accuracy and personalization of feeding recommendations, addressing the limitations of existing systems by integrating user habits and real-time feedback for improved dietary guidance.
Patent Information
- Application Number
- PCT/CN2024/105929
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-17
- Publication Date
- 2026-01-22
AI Technical Summary
Existing dietary recommendation systems lack personalization and accuracy, particularly for individuals with unique needs, and face challenges in handling noisy or incomplete data, ethical considerations, and the complexity of recognizing handwritten data.
A method using image recognition technology to extract textual information from visual data, train models on user habits, and generate personalized feeding recommendations based on scientific logic, incorporating user-specific data and real-time feedback for continuous improvement.
Enhances the accuracy and personalization of feeding recommendations, ensuring they align with individual needs and developmental stages, while reducing human error and improving reliability through dynamic learning and adaptability.
Smart Images

Figure CN2024105929_22012026_PF_FP_ABST
Abstract
Description
Personalized feeding recommendation assistant and method by using image recognition technology thereofField of the invention
[0001] The present invention relates to a personalized feeding recommendation assistant, especially a personalized feeding recommendation method using image recognition technology, a system and an apparatus thereof.Background
[0002] In the healthcare field, for example, to care-receivers like people (e.g., an infant or a senior person) , advances in technology allow helping caregivers with the task of tracking the developments of the care-receivers and quickly detecting certain anomalies.
[0003] Traditional dietary guidelines are often generic and fail to consider an individual's specific context, such as their meal habits, activity levels, emotions, and growth patterns. In the other words, existing methods lack personalization and often require manual input from caregivers. The challenge lies in providing personalized dietary recommendations that cater to an individual's unique needs, preferences, and lifestyle.
[0004] Therefore, there is a need for providing an automated system that considers individual characteristics and provides precise feeding guidance.
[0005] Currently, there are several applications utilizing optical character recognition (OCR) technologies for recommendation tools. Respective OCR models in these tools are trained by the general dataset which has achieved preliminary results in the recognition of feeding content. However, due to the particularity and complexity of feeding content, the general model is still insufficient in the recognition of feeding content, especially for the recognition of handwritten data.
[0006] Therefore, there is a need for improving the recognition effect of the OCR model for achieving more accurate recommendations for individuals.
[0007] Furthermore, balancing personalized recommendations with ethical and / or nutritional guidelines is recommended. For instance, avoiding recommendations that promote unhealthy behaviors or eating disorders. Therefore, there is a need for ensuring transparency about how recommendations are generated and avoiding biases.
[0008] Additionally, collecting data from the individuals can be challenging, meaning that recorded data can be noisy, incomplete, or inconsistent. Variability exists in how users input information (e.g., subjective meal descriptions, varying sleep tracking accuracy) . Therefore, there is also need for developing robust algorithms that handle missing or erroneous data.
[0009] In the case of the dietary recommendation is prepared for an infant or individuals for higher care cases such as elderly or incapacitated, the above mentioned challenges can be more complex.
[0010] For example, infants undergo rapid growth and development. Their nutritional needs change significantly during the first two years of life. Therefore, designing personalized recommendations that adapt to these developmental stages is recommended.
[0011] Since the infant has not enough cognitive skills to show their responses for feeding, understanding an infant's hunger cues and satiety signals is important. Responsive feeding promotes healthy eating behaviors for the infants.
[0012] Furthermore, regular growth assessments for infants are critical. Tracking weight, length, and head circumference of the infant ensures adequate nutrition. Therefore, recommendations prepared for infants should emphasize also growth monitoring.
[0013] In summary, personalized feeding recommendations for infants require a holistic approach, considering developmental stages, individual needs, and caregiver education. Therefore, there is a need for providing an automated system by addressing these challenges, thereby optimizing infant or higher care individuals nutrition and promote healthy growth.Summary of the invention
[0014] The disclosure relates to providing personalized feeding recommendations based on a user′s recording habits related to dietary and activity habits of a subject, preferably a mammal subject, and more preferably human subject. Especially, the mammal is selected from those needs care or intends to improve either dietary or activity situation.
[0015] The disclosure provides solutions for obtaining visual image data representing data related to the dietary and activity habits of the subject, extracting textual information from the visual image data using optical character recognition, training models to determine the user′s recording habits data, processing input data based on the determined habits data, analyzing the input data based on scientific logic data, and generating personalized feeding recommendations using trained models.
[0016] There is disclosed a method for providing personalized feeding recommendations based on a trained model of a user′s recording habits related to a dietary and an activity of a subject, comprising the steps of:
[0017] obtaining visual image data representing data related to the dietary and the activity of the subject by using a camera of a computing device, wherein the data comprises at least one of feeding data about type, quantity, and frequency of food consumed by the subject, personal data about gender, age, weight, and length of the subject, and activity data about interaction, emotion, sleep, excreta, and growth scale of the subject, wherein the computing device is configured to run an application with a user interface, wherein the user interface gives directions for scanning the visual image data to a user of the computing device,
[0018] employing a trained optical character recognition, OCR, model to scan and extract textual information from the obtained visual image data, thereby converting the data related to the dietary and the activity of the subject into a machine-readable data,
[0019] training a user's recording habits model to determine the user's recording habits data related to the dietary and activity of the subject based on the machine-readable data,
[0020] processing an input data based on the determined user's recording habits data and the machine-readable data to obtained processed input data,
[0021] training a dietary habits model to determine the subject's dietary habits data based on the processed input data,
[0022] training an activity habits model to determine the subject's activity habits data based on the processed input data,
[0023] analysing the processed input data based on scientific logic data to obtain analysed input data, wherein the scientific logic data includes information on at least one of principles of subject nutrition, developmental stages related to the activity of the subject, and dietary guidelines related to the dietary of the subject,
[0024] training a feeding recommendation model using at least one of the analysed input data, the subject's dietary habits data, and the subject's activity habits data,
[0025] generating a personalized feeding recommendation for the subject by using the trained feeding recommendation model.
[0026] The integration of a trained OCR model enhances the accuracy and efficiency of data extraction from visual image data, reducing the likelihood of human error in manual data entry and improving the reliability of the personalized feeding recommendations.
[0027] The use of a user′s recording habits model to process input data ensures that the feeding recommendations are tailored not only to the subject′s dietary and activity data but also to the user′s unique recording patterns, providing a more customized and relevant user experience.
[0028] By incorporating scientific logic data in the analysis of input data, the method ensures that the personalized feeding recommendations are grounded in established nutritional principles and guidelines, thereby enhancing the credibility and effectiveness of the dietary advice provided.
[0029] The method provides the advantage of personalized feeding recommendations based on a trained model of a user′s recording habits, allowing for customized dietary and activity recommendations tailored to the individual′s needs.
[0030] By employing a trained optical character recognition (OCR) model to extract textual information from visual image data, the method ensures accurate and efficient conversion of data related to the dietary and activity of the subject into machine-readable data.
[0031] By utilizing scientific logic data, including principles of subject nutrition, developmental stages, and dietary guidelines, to analyze the processed input data, the method provides reliable and evidence-based recommendations. The method further incorporates a user interface on a computing device, providing clear directions for scanning visual image data and facilitating user interaction and engagement with the application. This allows for continuous improvement and updating of the trained models and user′s recording habits data through retraining based on user input, ensuring the recommendations remain up-to-date and relevant. User input may include but not limited to textual information and visual information etc.
[0032] In an embodiment, the method further comprises obtaining the visual image data comprising:
[0033] after capturing, using the camera of the computing device, the visual image data,
[0034] transmitting the visual image data to a server.
[0035] Transmitting visual image data to a server allows for the leveraging of more powerful computational resources and storage capabilities than those available on a computing device, potentially improving the speed and accuracy of data processing.
[0036] In an embodiment, the method further comprises machine readable data comprising the quantity of food consumed in terms of at least one of units of weight, units of volume, pieces, and cup.
[0037] Including quantitative measures such as weight, volume, pieces, and cups in the machine-readable data allows for precise tracking of food consumption, which can lead to more accurate assessments of dietary intake and better-informed feeding recommendations.
[0038] The ability to quantify food consumption in various units provides flexibility in accommodating different types of foods and user preferences, enhancing the adaptability of the method to diverse dietary habits, cultural practices and efficiently manage nutrition.
[0039] In an embodiment, the method further comprises generating a personalized feeding recommendation of the subject based on the trained model comprising:
[0040] generating a report related to the analysed input data, wherein the report comprises information associated with the determined user's recording habits data.
[0041] Generating a report that includes information on the user′s recording habits data provides valuable feedback to the user, which can be used to improve the accuracy of future recordings and adherence to the personalized feeding recommendations.
[0042] The report serves as a tangible record of the analyzed input data and the user′s dietary and activity patterns, which can be used for tracking progress over time and making adjustments to the feeding recommendations as needed.
[0043] In an embodiment, the method further comprises:
[0044] obtaining a correlation model between the determined subject's activity habits data and the determined subject's dietary habits data, based on the analysed input data.
[0045] Obtaining a correlation model between the subject′s activity habits data and dietary habits data allows for the identification of patterns and relationships that may not be immediately apparent, leading to more nuanced and effective feeding recommendations.
[0046] The correlation model can help in predicting the impact of changes in activity levels on dietary needs, enabling proactive adjustments to feeding recommendations to maintain or improve the subject′s overall health and well-being.
[0047] In an embodiment, the method further comprises training a feeding recommendation model using the analysed input data comprising:
[0048] training a feeding recommendation model using the analysed input data and the obtained correlation model.
[0049] This may enhance the accuracy of the feeding recommendation model by incorporating user-specific data, leading to more tailored and effective dietary suggestions.
[0050] It may also facilitate continuous improvement of the recommendation system through iterative learning, ensuring the model remains relevant and adapts to changing dietary patterns or user needs.
[0051] Further, it may increase user trust and engagement with the system by providing recommendations that are perceived as more credible due to the personalized approach based on analyzed input data.
[0052] In an embodiment, the method further comprises:
[0053] transmitting the personalized feeding recommendations and the generated report to the user through a user interface.
[0054] This may improve user convenience by delivering personalized feeding recommendations and reports directly to the user′s preferred interface, promoting regular interaction and adherence to suggested dietary plans.
[0055] It may also enhance the user experience by providing actionable insights in an accessible format, which may encourage positive behavioral changes regarding nutrition.
[0056] Further, it may streamline the communication process between the system and the user, potentially increasing the efficiency of the dietary management process.
[0057] In an embodiment, the method further comprises:
[0058] receiving a user input on the transmitted personalized feeding recommendations and report, through the user interface.
[0059] This may enable the system to capture real-time feedback from users, allowing for immediate adjustments to recommendations and enhancing user satisfaction.
[0060] This may further provide valuable data on user preferences and behaviors, which can be used to refine the recommendation algorithms by retraining the trained models and improve future recommendations.
[0061] Furthermore, this may encourage user interaction and engagement with the system, fostering a collaborative approach to dietary management and increasing the likelihood of sustained use.
[0062] In an embodiment, the method further comprises:
[0063] retraining the trained user's recording habits based on the received user input, in order to update the user's recording habits data.
[0064] This allows the system to adapt to changes in user behavior over time, maintaining the relevance and effectiveness of the feeding recommendations. In that way, a dynamic learning environment can be created, wherein the system evolves based on user feedback, leading to a more personalized and user-centric approach.
[0065] Therefore, the likelihood of user disengagement by ensuring that the system′s recommendations remain aligned with the user′s evolving recording habits and preferences.
[0066] In an embodiment, the method further comprises scientific logic data being predetermined data stored in a memory of the computing device.
[0067] This provides a reliable foundation for the recommendation system by utilizing scientifically validated data, which can enhance the credibility and accuracy of the recommendations.
[0068] Furthermore, this may reduce the dependency on real-time data collection by leveraging pre-existing scientific logic data, thereby speeding up the recommendation process.
[0069] It may also ensure consistency in the recommendations provided to users by referencing a standardized dataset, which can be particularly beneficial for regulatory compliance and quality control.
[0070] In an embodiment, the method further comprises scientific logic data being simultaneously retrieved from at least one resource through the server, while analysing the processed input data based on scientific logic data to obtain analysed input data.
[0071] This may enhance the accuracy of the feeding recommendation model by incorporating real-time scientific logic data, ensuring that the recommendations are based on the latest research and findings, thereby providing a dynamic and adaptive learning environment for the model by allowing continuous data integration from various resources.
[0072] This may further facilitate a comprehensive approach to feeding recommendations by leveraging a diverse set of data sources, potentially improving user trust and adherence to the suggested guidelines.
[0073] In an embodiment, the method further comprises:
[0074] retraining the trained feeding recommendation model based on at least one of the received user inputs, the updated user's recording habits data, and the scientific logic data.
[0075] This may enable the feeding recommendation model to evolve and improve over time by incorporating new user inputs and updated habit data, leading to a system that better reflects the changing needs and preferences of the user.
[0076] It may increase the model′s predictive performance and reliability by continuously refining its parameters, which can result in more effective and tailored feeding recommendations for the subject.
[0077] It may promote a user-centric design by allowing the model to adapt to individual feedback and habit changes, thereby enhancing user satisfaction and engagement with the feeding recommendation system.
[0078] In an embodiment, the method further comprises generating a personalized feeding recommendation of the subject based on the trained model comprising:
[0079] generating a personalized activity recommendation based on the determined the subject's activity habits data.
[0080] This may deliver a holistic approach to health and well-being by not only providing feeding recommendations but also suggesting personalized activities, which can contribute to the overall development and health of the subject.
[0081] It may encourage a balanced lifestyle by integrating activity habits into the recommendation process, potentially leading to better health outcomes and adherence to a healthy routine.
[0082] Furthermore, this may tailor recommendations to the unique activity patterns of the subject, which can increase the effectiveness of the feeding recommendations by aligning them with the subject′s lifestyle and energy expenditure.
[0083] In an embodiment, the subject needs care.
[0084] In an embodiment, the subject is selected from mammal subject.
[0085] In an embodiment, the mammal is up to 3 years old, preferably up to 2 years old, and more preferably up to 1 years old.
[0086] In an embodiment, the subject is an infant.
[0087] As such, this addresses the specific nutritional needs and developmental considerations of infants, which are critical during the early stages of life and can have long-term impacts on health and growth.
[0088] Apart from providing a specialized solution for a vulnerable population group, such a method can assist caregivers in making informed decisions about infant feeding practices. It may enhance the precision of feeding recommendations by focusing on the unique physiological and metabolic requirements of infants, thereby supporting proper nutrition and development.
[0089] In an embodiment, the method further comprises dietary habits data comprising at least one of breastfeeding data and nutritional composition data, and wherein the activity habits data comprises at least one of sleep time period, awake time period, and excreta data.
[0090] This may incorporate comprehensive dietary and activity data, including breastfeeding and nutritional composition, which can lead to more accurate and personalized feeding recommendations tailored to the individual needs of the infant. It may enable the identification of correlations between dietary habits and activity patterns, such as sleep and awake times, which can inform more effective feeding schedules and nutritional plans.
[0091] Furthermore, early detection and intervention may be facilitated for potential dietary or activity-related issues by monitoring excreta data, contributing to the overall health and well-being of the infant.
[0092] There is disclosed a non-transitory computer readable medium for providing personalized feeding recommendations based on a trained model of a user′s recording habits related to a dietary and an activity of a subject, the non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to perform the method according to any one to
[0093] By utilizing a trained model, the medium provides dynamic and adaptive suggestions that evolve with the user′s changing dietary patterns and physical activities, ensuring that the recommendations remain relevant and personalized over time.
[0094] The implementation of such a medium can lead to increased user engagement and adherence to dietary recommendations, as the personalized nature of the advice may be more appealing and easier to integrate into the user′s lifestyle.
[0095] There is disclosed a system for providing personalized feeding recommendations based on a trained model of a user′s recording habits related to a dietary and an activity of a subject, the system comprising: one or more processors; one or more modules, and one or more memories storing instructions that, when executed by the one or more processors, cause the system to perform operations the method according to any one to
[0096] The system′s architecture, comprising one or more processors and modules, allows for scalable and efficient processing of complex datasets, enabling real-time generation of personalized feeding recommendations.
[0097] Integration of multiple modules within the system facilitates a comprehensive analysis of user data, encompassing both dietary intake and physical activity, which can lead to more accurate and holistic nutritional guidance.Brief description of the drawings
[0098] The present invention will be discussed in more detail below, with reference to the attached drawings, in which:
[0099] Fig. 1 shows a schematic diagram of environment of personalized feeding recommendation assistant computing system according to an embodiment of the present invention.
[0100] Fig. 2 shows a schematic illustration of exemplary display screen of an electronic device, wherein the display screen includes visual image data representing user's recording data captured by a camera of the electronic device according to an embodiment of the present invention.
[0101] Figs. 3a and 3b show schematic illustrations of exemplary display screens of an electronic device with each figure showing machine readable data obtained by using the system of Fig. 1.
[0102] Figs. 4a, 4b, 5a, and 5b show schematic illustrations of exemplary display regions in the display screen of Fig. 2, and its conversion to the machine-readable data.
[0103] Fig. 6 shows a schematic illustration of an electronic device according to an embodiment of the present invention.
[0104] Fig. 7 shows a schematic llustration of an electronic device and a server according to an embodiment of the present invention.Description of embodiments
[0105] Embodiments of the present disclosure will be described herein below with reference to the accompanying drawings. However, the embodiments of the present disclosure are not limited to the specific embodiments and should be construed as including all modifications, changes, equivalent devices and methods, and / or alternative embodiments of the present disclosure.
[0106] The terms “have, ” “may have, ” “include, ” and “may include” as used herein indicate the presence of corresponding features (for example, elements such as numerical values, functions, operations, or parts) , and do not preclude the presence of additional features.
[0107] The terms “A or B, ” “at least one of A or / and B, ” or “one or more of A or / and B” as used herein include all possible combinations of items enumerated with them. For example, “A or B, ” “at least one of A and B, ” or “at least one of A or B” means (1) including at least one A, (2) including at least one B, or (3) including both at least one A and at least one B.
[0108] The terms such as “first” and “second” as used herein may modify various elements regardless of an order and / or importance of the corresponding elements, and do not limit the corresponding elements. These terms may be used for the purpose of distinguishing one element from another element. For example, a first printing form and a second printing form may indicate different printing forms regardless of the order or importance. For example, a first element may be referred to as a second element without departing from the scope the present invention, and similarly, a second element may be referred to as a first element.
[0109] It will be understood that, when an element (for example, a first element) is “ (operatively or communicatively) coupled with / to” or “connected to” another element (for example, a second element) , the element may be directly coupled with / to another element, and there may be an intervening element (for example, a third element) between the element and another element. To the contrary, it will be understood that, when an element (for example, a first element) is “directly coupled with / to” or “directly connected to” another element (for example, a second element) , there is no intervening element (for example, a third element) between the element and another element.
[0110] The expression “configured to (or set to) ” as used herein may be used interchangeably with “suitable for, ” “having the capacity to, ” “designed to, ” “adapted to, ” “made to, ” or “capable of” according to a context. The term “configured to (set to) ” does not necessarily mean “specifically designed to” in a hardware level. Instead, the expression “apparatus configured to. . . ” may mean that the apparatus is “capable of. . . ” along with other devices or parts in a certain context.
[0111] The terms used in describing the various embodiments of the present disclosure are for the purpose of describing particular embodiments and are not intended to limit the present disclosure. As used herein, the singular forms are intended to include the plural forms as well, unless the context clearly indicates otherwise. All of the terms used herein including technical or scientific terms have the same meanings as those generally understood by an ordinary skilled person in the related art unless they are defined otherwise. The terms defined in a generally used dictionary should be interpreted as having the same or similar meanings as the contextual meanings of the relevant technology and should not be interpreted as having ideal or exaggerated meanings unless they are clearly defined herein. According to circumstances, even the terms defined in this disclosure should not be interpreted as excluding the embodiments of the present disclosure.
[0112] As used herein, the term OCR model may refer generally to optical character recognition software, that is a computer program product that is able to convert a digital image or video representation including handwritten or typed or printed writing to machine readable writing in e.g. ASCII or Unicode encoding. Such computer program products may also include intelligent character recognition (ICR) and / or intelligent word recognition (IWR) . It may include general machine learning models that are able to process image or video data and distill written or typed text from it, such as large language models (LLMs) with vision capabilities. In an example definition, an OCR model is able to convert a pixel-based representation of written or typed data into a ASCII or Unicode or other representation of the same written or typed data.
[0113] As used herein, the term "system" refers to a collection of one or more processors and memory.
[0114] The term "trained model" as used herein refers to a computer-based model that is capable of providing personalized feeding recommendations based on the user's recording habits.
[0115] A “users recording habits model” refers to a model, such as a machine learning model, that is trained to determine a user's recording habits data related to the dietary and / or to the activity of the subject based on data entered by the user. It may be trained to demarcate, distinguish, or segment different portions of (handwritten) data in visual image data and label these as different data such as data relating to dietary habits or data relating to activity habits.
[0116] A "dietary habits model" designates a model, such as a machine learning model, that is trained to determine the dietary habits of a subject based on the machine-readable data, to be used for training a feeding recommendation model.
[0117] An "activity habits model" designates a model, such as a machine learning model, that is trained to determine the subject's activity habits data based on the processed input data, to be used for training a feeding recommendation model.
[0118] As used herein, the term "computing device" refers to any device that is configured to run an application with a user interface. It may be a mobile phone, or a general computing device. It may be a tablet, personal digital assistant, smart watch, general wearable computing device, or the like.
[0119] The term "user interface" as used herein refers to an interface between a human user or operator and one or more devices that enables communication between the user and the device (s) .
[0120] A "determined user's recording habits data" designates the user's recording habits data related to the dietary and activity of the subject. A "machine readable data" designates a data that can be processed by a machine, such as a computer, to obtain a result.
[0121] As used herein, the term "frequency of food" refers to the amount, type, quantity, and frequency of food consumed by a subject. As used herein, the term "quantity of food" refers to the amount of food consumed by a subject. A "unit of weight" designates a unit of weight of food, such as a gram, a milligram, a milliliter, a liter, a kilogram, a kilogram, a gram, a gram, a milligram, a milliliter, a liter, a kilogram, a kilogram. A "unit of volume" designates a unit of volume of food consumed by the subject. As used herein, the term "growth scale" refers to a growth scale that measures the growth rate of a subject.
[0122] As used herein, the term "personalized activity recommendation" refers to a feeding recommendation of a subject based on at least one recorded activity habits data (e.g., type, quantity, and frequency of food consumed by the subject) or personal data about gender, age, weight, and length of the subject.
[0123] Fig. 1 shows a schematic diagram of environment of personalized feeding recommendation assistant computing system 100 according to an embodiment of the present invention. Said illustrative system maybe embedded as a number of machine-readable components, such as instructions, modules, data structures and / or other components which maybe implemented as computer hardware, firmware, software, or a combination thereof. For ease of description, the personalized feeding recommendation assistant computing system 100 may be referred as “image recognition assistant” or “image recognition system” or by similar terminology.
[0124] The personalized feeding recommendation assistant computing system 100 includes image recognition module 120, image library 140, processor 160, analysing tool 130, recommendation module 170, and a number of training models 150. In an exemplary embodiment, all of the components of computing system 100 may be embodied on a single computing device, such as a smart phone, tablet computer, or wearable computing device, in other implementations.
[0125] The image recognition module 120 may execute a number of algorithms for identifying features in a visual image data representing user's recording data, which can be used to identify data shown in the visual image in order to generate machine readable data 121. The visual image 110 may contain handwritten data related to the dietary and the activity of the subject, logged by the user. The visual image may be obtained by using a camera of a computing device (not shown in Fig. 1) , and may comprise at least one of feeding data about type, quantity, and frequency of food consumed by the subject, personal data about gender, age, weight, and length of the subject, and activity data about interaction, emotion, sleep, excreta, and growth scale of the subject.
[0126] The visual image 110 may include at least one of the video information, image information, three-dimensional information and thermal image information. In the whole document, an image may be used as an example of the visual image, but it should be understood that the image is only an example and other forms of visual information are included in the present invention as well. For example, an image or a video may be captured (either displayed or not displayed, may be stored in the memory or not) , e.g., by a camera, thermal Imager, fetched from a memory, received from an external device via a telecommunication unit, or via other means, in order to obtain the visual image 110.
[0127] In an exemplary embodiment, the computing device may run an application with a user interface, wherein the user interface gives directions to the image recognition module 120 for scanning the visual image to the user of the computing device. For ease of description, the computing device may be referred as “electronic device” or “computing device” or by similar terminology.
[0128] The image recognition module 120 outputs or generates the machine-readable data after the scanning of the visual image 110. The machine-readable data 121 literally indicates the handwritten data included the user's recording data. As shown in Fig. 1, the machine-readable data is generated algorithmically by the image recognition module 120 in the computing system 100, by scanning the visual image 110.
[0129] In exemplary embodiments, to generate the machine-readable data 121, the illustrative image recognition module 120 may access one or more feature detection algorithms and / or one or more text identification modules. The text identification modules may be embodied as software, hardware, or a combination thereof. The text identification modules may be developed or trained by correlating suitable large sample of the handwritten data logged by the user, as mentioned above, where the sample of said user-generated handwritten data is stored in the image library 140.
[0130] While generating the machine-readable data 121, the image recognition module 120 may estimate food volume, or portion size, of each handwritten data logged by the user in the visual image 110. For example, in a custom set of forms, the user will record or log the amount of food consumed in the morning / afternoon / evening in a certain unit, such as ingesting eight glasses of water, several fruits, a few slices of bread, the grams of solid food the baby ingested. Accordingly, the machine-readable data 121 may comprise the quantity of food consumed in terms of at least one of units of weight, units of volume, pieces, slices and cup.
[0131] In some embodiments, while determining the quantity of the food consumed by the subject, the image recognition module may access mapping table or knowledge base via external databases, in order to define the features related to these quantities in the visual image, more accurately.
[0132] Once the image recognition module 120 has identified the text present in the textual information data including handwritten data or visual information data and outputted the machine-readable data 121, training models 150 of the computing system 100 may analyse the machine-readable data in the recording habits model 151.
[0133] The recording habits model 151 is trained or developed based on the machine-readable data 121, for determine the user's recording habits data related to the dietary and activity of the subject. Exemplarily, the recording habits model 151 may generate one or more output for defining the determined user's recording habits data and may display said data detected in the visual image 110.
[0134] In some embodiments, the recording habits model 151 may use the similar formatting of the handwritten data logged by the user, for example, table, classification, or etc. Accordingly, the recording habits model 151 may further demarcate different portions of the handwritten data in the visual image 110 that have been identified as different data such as dietary habits or activity habits.
[0135] After the recording habits model 151 generated the user's recording habits data 155, it may be directed to processor 160 together with the machine-readable data 121, for processing an input data 161 based on the determined user's recording habits data and the machine-readable data.
[0136] In an exemplary embodiment, once the input data is generated or determined by the processor, dietary habits model 152 and activity habits model 153 may be trained or developed based on said processed input data 161, in order to determine dietary habits data 156, and activity habits data 157 to be used for training feeding recommendation model 154.
[0137] The user maybe the subject. Although, no user input is needed in the above-mentioned steps, the user may also provide input on any of the above mentioned output data, in order to improve the accuracy of the training models.
[0138] In the case of the subject is an infant which shows low cognitive skills, the user provides handwritten data in the visual image on behalf of the infant. Since, the infants undergo rapid growth and development, their nutritional needs also change significantly during the first two years of life.
[0139] Therefore, when the dietary habits model 152 and the activity habits model 153 utilizes the processed input data associated with the infant, the determined dietary habits data may comprise at least one of breastfeeding data and nutritional composition data, and the determined activity habits data comprises at least one of sleep time period, awake time period, and excreta data. In that way, the computing system 100 can provide a holistic view of infant nutrition, considering not only the nutritional content of foods but also developmental stages, feeding behaviors, and individual variability.
[0140] In some embodiments, parallel to the operations in the dietary habits model 152 and the activity habits model 153, or independently from said operations, the processed input data 161 may be directed to the analysing tool 130, in order to analyse the data included in the processed input data based on scientific logic data (not shown in Fig. 1) which can be retrieved from at least one of memory of the computing device, Internet, or server that the portable device is connected. Once the processed input data 161 is analysed by the analysing tool 130, the analysed input data 131 is directed to the feeding recommendation model 154.
[0141] In particularly, the analysing tool 130 compares the processed input data 161 with up-todate scientific logic data that includes scientific information on at least one of principles of subject nutrition, developmental stages related to the activity of the subject, and dietary guidelines related to the dietary of the subject. In that way, the feeding recommendation model 154 can be trained dynamically for providing / generating more accurate recommendations. This dynamic learning approach ensures accurate and up-to-date recommendations.
[0142] In some embodiments, the feeding recommendation model 154 may include one or more additional models for obtaining a correlation model between the determined subject's activity habits data and the determined subject's dietary habits data, based on the analysed input data. Accordingly, the feeding recommendation model may be further trained based on the analysed input data 131 and the obtained correlation model.
[0143] Once the personalized feeding recommendations are determined or generated by the feeding recommendation model 154, those recommendations can be output / display by the recommendation module 170 to the user, through the user interface. The personalized feeding recommendation can be output / displayed by way of graphics or comparison tables based on the scientific logic data. Said recommendations may further include food items, portion sizes, and meal timings by ensuring dietary diversity for balanced nutrient intake for the subject.
[0144] The feeding recommendation module 170 may further generate a report related to the analysed input data 131. Said report may comprise information associated with the determined user's recording habits data.
[0145] The feeding recommendation module 170 may further output one or more information associated with the displayed personalized feeding recommendations, to one or more training models included in the training models 150.
[0146] Once the personalized feeding recommendations are displayed to the user, the user may further provide input on the displayed personalized feeding recommendations and reports, through the user interface. The feeding recommendation module 170 may output said user input to one or more models included in the training model 150, in order to retrain said models. For example, the feeding recommendation module 170 may output said user input to the recording habits model for retraining the trained user's recording habits based on the received user input, in order to update the user's recording habits data.
[0147] Fig. 2 shows a schematic illustration of exemplary display screen of an electronic device, wherein the display screen displays the visual image 110 of user's recording data captured by a camera of the electronic device. As shown, the display screen displays an illustrative output of the visual image 110 including the handwritten data 111 logged by the user.
[0148] The camera of the electronic device may scan the handwritten data 111 in the visual image for the image recognition module 120 to determine the machine-readable data 121. The handwritten data may include a data logged in a predefined form. The image recognition module 120 may identify regions in the visual image in order to group the handwritten data 111 while converting it to the machine-readable data 121.
[0149] While identifying those regions in the handwritten data 111, the image recognition module 120 may access to the external databases in order to accurately place the fields of the predetermined field in the machine-readable data 121. As shown exemplarily in Fig. 2, the image recognition module 120 may identify at least two regions 111a and 111b in the handwritten data 111 included in the visual image 110, in order to determine related format for the machine-readable data 121.
[0150] The image recognition module 120 employs a trained OCR model to scan and extract textual information from the obtained visual image, thereby converting the data in the visual image into the machine-readable data 121.
[0151] Example models that can be employed for detecting the features in the handwritten data 111 are available in the OpenCV, Keras, and TensorFlow frameworks. More recently, the following models have been introduced LayoutLM, FOTS, and STAR-Net. In addition, use may be made of Large Language Models (LLMs) with vision capabilities, such as (but not limited to) OpenAI GPT-4.0, OpenAI GPT-4 with Vision, LLaVA-1.5, BakLLaVA, Google Gemini, Google PaliGemma, and Anthropic Claude 3.
[0152] Such AI models may provide jointly model interactions between text and layout information across scanned document images, by combining 2D location information, image embeddings, and text for pretraining the model. Such models are fine-tuned based on open data set (mostly collected by the user) , and may utilize an OCR engine (such as Tesseract) for obtaining bounding box information associated with each feature (token) in the handwritten data 1 11. These bounding boxes may represent 2D positions of the features (tokens) in the handwritten data 1 1 1. In this way, the pretrained model may jointly learn the handwritten text in the handwritten data 111, and the layout, by improving accuracy in document level learning model.
[0153] More information about the models can be found in the following papers:
[0154] Xu, Y., Li, M., Cui L., Huang, S., Wei, F., Zhou M. (2019) . “LayoutLM: Pre-training of Text and Layout for Document Image Understanding” . This paper introduces the LayoutLM.
[0155] Liu, X., Liang, D., Yan, S., Chen, D., Qiao, Y., Yan, J., (2018) . “FOTS: Fast Oriented Text Spotting with a Unified Network” . This paper introduces the FOTS.
[0156] Mao, J., Chang, Y., Yin, X., Nie, B., (2023) . “Star-Net: Improving Single Image Desnowing Model With More Efficient Connection and Diverse Feature Interaction” . This paper introduces the STAR-Net.
[0157] Farooq Alvi (2024) . “Deep Learning for Computer Vision: Models &Real World Applications” . This paper discusses the model of OpenCV.
[0158] According to some embodiments, in order to further improve the recognition effect of the model, the trained OCR model of the image recognition module 120, 90 feeding tables annotated in the dataset in order to obtain more than 5, 000 image-text samples and said OCR model is fine-tuned the training on this specialized dataset. The model needle trained on a large-scale generic dataset may be able to learn general features and may be asked again to learn on the dataset of feeding content to accelerate convergence and improve performance. During the fine-tuning process of the trained OCR model, the low-level feature layers of some pre-trained models are frozen, and only the high-level feature layers and output layers are trained to avoid overfitting. In that way, the trained OCR model improves the accuracy and efficiency of text detection and recognition through a number of improvements and optimizations to the detection and recognition model. Whether it is a Chinese scenario, an English digital scenario, or a multi-language scenario, said OCR model demonstrates strong performance advantages and provides a more reliable and efficient recommendation.
[0159] In an exemplary embodiment, in the fine-tuning process of the OCR model, a variety of model fine-tuning strategies may be employed, including:
[0160] ● Phased fine-tuning: Initial fine-tuning on a larger, related dataset followed by fine-tuning on the target dataset. This allows you to make better use of the features of an existing dataset while optimizing the model′s performance on specific tasks.
[0161] ● Layer-by-layer fine-tuning: Thawing the model layer by layer, starting at the top (near the output layer) and gradually training to the low layer (near the input layer) , can alleviate the instability in model training.
[0162] In some embodiments, in order to improve model performance and recognition of the trained OCR model, in addition to above the following strategies may be used:
[0163] ● Image enhancement: Rotate, scale, translate, add noise, and colour jitter to augment the training data to improve the robustness of the model to different handwriting and formula deformations.
[0164] ● Post-processing and correction: The NLP algorithm is used for text post-processing, which can correct the recognition results according to the contextual information and improve the accuracy of handwriting and formulas.
[0165] Fig. 3A shows a schematic illustration of the display screen of an electronic device. The display screen as shown, includes the machine-readable data 121 obtained by using the system 100 of Fig. 1. For example, once the user interface 180 of the electronic device scans the visual image 110 for training the image recognition module 120 (for example as shown in Fig. 2) , the image recognition module 120 determines (outputs) the machine-readable data 121.
[0166] As exemplarily shown in Fig. 3A, the user interface may display different regions e.g. 181, 182, for defining different data in the machine readable data 121, which is associated with the handwritten data of the visual image 110 (For example the handwritten data 111 of Fig. 2) . In order to identify these regions 181, 182, and accordingly the layout in the displayed machine-readable data 121, the image recognition module 120 may be developed by collecting information about the surrounding situation in which the visual image was captured. The image recognition module 120 may collect the information by looking at meta data associated with the visual image 110 (e.g., date, time, location) . Said information may include geographic location, the time, and the date at which the image 110 was captured. In that way, said information may be utilized for training the AI models mentioned above (in connection with Fig. 2) in order to identify the handwritten text in the virtual image 110.
[0167] Fig. 3B shows a schematic illustration of the display screen of an electronic device according to an embodiment. The display screen of the user interface 180 displays the processed input data 161 that is determined by the processor of the system of Fig. 1, based on the machine-readable data 121.
[0168] For example, once the training models 150 in Fig. 1 determine the dietary and activity habits data (156 and 157) of the user, the system 100 may direct those data to the processor. The processor of the system 100 may utilize said habit data of the user together with the machine-readable data of Fig. 3A, in order process an input data for training the feeding recommendation model 154.
[0169] As exemplarily shown in Fig. 3B, the processed input data 161 may be displayed to the user through the user interface 181, either before or simultaneously directed to the feeding recommendation model 154 of the system 100 of Fig. 1. Not necessarily, but alternatively, the user may provide input on the displayed input data 161, and the system 100 may output said user input in order to re-train one or more training models 150 in the system 100. Thereby, the system allows the training models to adapt or learn continuously and real time, based on user feedback.
[0170] Figs. 4A and 4B shows a schematic illustration of the region 111a the scanned visual image 110 in Fig. 2 and it is correspondence layout 121a in the machine-readable data 121 which is determined by the image recognition module 120.
[0171] As shown in Fig. 4A, the region 111a in the handwritten data 111 logged by the user includes Chinese characters written by the user in the predefined form. The AI model used in the image recognition module 120 may segment individual characters from the text using edge detection, contours, and contour filtering. After said segmentation, the AI model may label each character in the handwritten data, as shown in the region 121a of the machine-readable data 121.
[0172] However, said labelling may include errors or misidentifications, like the misidentified character 122 in the machine-readable data 121, in Fig. 4A. Since the image recognition module 120 and the training models interacts continuously in the system 100 of Fig. 1, such a misidentification can be determined by the recording habits model 151, before requiring any user input in between. For example, once the machine-readable data including the character 122 output to the recording habits model 151, said model may be already pretrained about the recording habits of the user related to said data by requesting open data set, and therefore, can identify the misidentified character 122 by the image recognition module 120.
[0173] As shown in Fig. 4B, the AI model of the image recognition module 120 may learn from the recording habits data 155 output from the recording habits model 151 of the system 100 of Fig. 1, and correctly identify the character 123 in the region 121a of the machine-readable data 121, even during the scanning of the visual image via the camera of the electronic device.
[0174] Figs. 5A and 5B shows a schematic illustration of the region 111b the scanned visual image 110 in Fig. 2 and it is correspondence layout 121b in the machine-readable data 121 which is determined by the image recognition module 120.
[0175] As shown in Fig. 5A, the region 111a in the handwritten data 111 logged by the user includes some feeding quantity indications written by the user in the predefined form. The AI model used in the image recognition module 120 may segment individual characters from the text using edge detection, contours, and contour filtering. After said segmentation, the AI model may label each character in the handwritten data, as shown in the region 121a of the machine-readable data 121. In this approach, the image recognition module 120 may use image enhancement technology and NLP algorithm to improve recognition accuracy.
[0176] Similar to Fig. 4A, as shown in Fig. 5A the machine-readable data 121 may also label some characters wrongly, e.g. character 124 in Fig. 5A. However, due to the continuous learning of the training models 150 between the AI models of the image recognition module 120, said character may be correctly identified and / or labelled by the image recognition module 120, even during the scanning of the visual image 110, as shown in Fig. 5B.
[0177] Figure 5 schematically shows the electronic device 10 according to an embodiment of the present invention. The electronic device 10 has a display unit 501, which may be touch screen suitable for displaying information and handling user input. The device 10 further has a camera 502 for recording the visual information of the excreta, a processor 503 for processing recorded visual information, a memory 504 for storing the visual information, program data, the AI engine, and the like, and a communication unit 505 for communication with other devices over wired or wireless connections. In an embodiment the processor 503 is programmed to process recorded images using the AI engine and to generally implement the processes as described in this application.
[0178] Figure 6 schematically shows the electronic device 10 and a server 100 according to an embodiment of the present invention. The electronic device 10 and server 100 can communicate over a wired or wireless link. In an embodiment, the device 10 sends recorded visual information to the server. The server has a processor, a memory, and a communication unit. The server 100 can be programmed to process received visual image using the AI engine and to send back the results to the device 10. In addition, the server may store the obtained results and / or the received images and / or any intermediate calculation results. The server may be further arranged to implement the afore mentioned trained machine learning method.
[0179] Various other embodiments of the invention will be apparent to the skilled person when having read the above disclosure in connection with the drawings, all of which are within the scope of the invention and accompanying claims.
Claims
1.A method for providing personalized feeding recommendations based on a trained model of a user′s recording habits related to a dietary and / or an activity of a subject, comprising:obtaining visual image data representing data related to the dietary and the activity of the subject by using a camera of a computing device, wherein said data comprises at least one of feeding data about type, quantity, and frequency of food consumed by the subject, personal data about gender, age, weight, and length of the subject, and activity data about interaction, emotion, sleep, excreta, and growth scale of the subject, wherein the computing device is configured to run an application with a user interface, wherein the user interface gives directions for scanning the visual image data to a user of the computing device,employing a trained optical character recognition, OCR, model to scan and extract textual information from the obtained visual image data, thereby converting the data related to the dietary and the activity of the subject into a machine-readable data,training a user’s recording habits model to determine the user’s recording habits data related to the dietary and activity of the subject based on the machine-readable data,processing an input data based on the determined user’s recording habits data and the machine-readable data to obtained processed input data,training a dietary habits model to determine the subject’s dietary habits data based on the processed input data, and / or,training an activity habits model to determine the subject’s activity habits data based on the processed input data,analysing the processed input data based on scientific logic data to obtain analysed input data, wherein the scientific logic data includes information on at least one of principles of subject nutrition, developmental stages related to the activity of the subject, and dietary guidelines related to the dietary of the subject,training a feeding recommendation model using at least one of the analysed input data, the subject’s dietary habits data, and the subject’s activity habits data,generating a personalized feeding recommendation for the subject by using the trained feeding recommendation model.2.The method according to claim 1, wherein the obtaining the visual image data comprises:after capturing, using the camera of the computing device, the visual image data,transmitting the visual image data to a server.3.The method according to any one of the preceding claims, wherein the machine-readable data comprises the quantity of food consumed in terms of at least one of units of weight, units of volume, pieces, and cup.4.The method according to any one of the preceding claims, wherein generating a personalized feeding recommendation of the subject based on the trained model comprises:generating a report related to the analysed input data, wherein the report comprises information associated with the determined user’s recording habits data.5.The method according to any one of the preceding claims further comprising:obtaining a correlation model between the determined subject’s activity habits data and the determined subject’s dietary habits data, based on the analysed input data.6.The method according to claim 5, wherein the training a feeding recommendation model using the analysed input data comprises:training a feeding recommendation model using the analysed input data and the obtained correlation model.7.The method according to any one of the preceding claims further comprising:transmitting the personalized feeding recommendations and the generated report to the user through a user interface.8.The method according to claim 7 further comprising:receiving a user input on the transmitted personalized feeding recommendations and report, through the user interface.9.The method according to claim 8 further comprising:retraining the trained user’s recording habits based on the received user input, in order to update the user’s recording habits data.10.The method according to any one of the preceding claims, wherein the scientific logic data is predetermined data stored in a memory of the computing device.11.The method according to any one of claims 1 to 9, wherein the scientific logic data is simultaneously retrieved from at least one resource through the server while analysing the processed input data based on scientific logic data to obtain analysed input data.12.The method according to any one of the preceding claims further comprising:retraining the trained feeding recommendation model based on at least one of the received user inputs, the updated user’s recording habits data, and the scientific logic data.13.The method according to any one of the preceding claims, wherein the generating a personalized feeding recommendation of the subject based on the trained model comprises:generating a personalized activity recommendation based on the determined the subject’s activity habits data.14.The method according to any one of the preceding claims, wherein the subject needs care.15.The method according to any one of the preceding claims, wherein the subject is selected from mammal subject.16.The method according to claim 15, wherein the mammal is up to 3 years old, preferably up to 2 years old, and more preferably up to 1 years old.17.The method according to any one of the preceding claims, wherein the subject is an infant.18.The method according to claim 17, wherein the dietary habits data comprises at least one of breastfeeding data and nutritional composition data, and wherein the activity habits data comprises at least one of sleep time period, awake time period, and excreta data.19.A non-transitory computer readable medium for providing personalized feeding recommendations based on a trained model of a user′s recording habits related to a dietary and an activity of a subject, the non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 18.20.A system (100) for providing personalized feeding recommendations based on a trained model of a user′s recording habits related to a dietary and an activity of a subject, the system comprising: one or more processors; one or more modules, and one or more memories storing instructions that, when executed by the one or more processors, cause the system to perform operations the method according to any one of claims 1 to 18.
Citation Information
Patent Citations
Automatic monitor and management system for intake of infant formula milk
CN102968563A
Breast-feeding management method, device and equipment in infant lactation period and storage medium
CN116092637A
Method for preparing personalized nutrition recommendations for infants
CN116391232A
New dietary survey method based on food image recognition technology
CN117936030A
Systems and methods for food analysis, personalized recommendations, and health management
US20190290172A1