System

A system using image recognition and AI determines food safety for babies and pets by analyzing images on a terminal, addressing the challenge of inconsistent information and ensuring quick, reliable dietary management.

JP2026019219APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024120628
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

The challenge of determining the safety of food for babies and pets in households, where various food ingredients are available, is complicated by the inconsistency of information from different websites, leading to a high risk of inappropriate feeding and the need for quick, reliable decision-making.

Method used

A system that allows users to take images of food using a terminal, analyze the images with an image recognition algorithm, and use a chat-based AI model to determine food safety, providing results on the terminal.

Benefits of technology

Enables users to easily and quickly identify safe and unsafe foods for babies and pets by leveraging image recognition and AI for accurate dietary management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019219000001_ABST
    Figure 2026019219000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system, comprising: means for a user to capture an image of food using a camera of a device and send the image to a server; means for the server to analyze the received image and identify a type of the food using an image-recognition algorithm; means for the server to collect information about the identified food using a chat-type AI model and determine whether it is safe for a baby or a pet to consume the food; and means for the server to send a determination result to the user's device and display the result on the device.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] When managing the diet of babies and pets (dogs and cats), the question "Is it safe to eat this?" often arises. With the wide variety of food ingredients available in the home, there is a high risk of accidentally feeding something inappropriate to babies or pets. Furthermore, different websites provide different information, making it difficult to obtain reliable information and requiring quick decisions. To solve this problem, a system that allows easy access to accurate and comprehensive information is needed. [Means for solving the problem]

[0005] To solve this problem, the present invention provides the following means: a means for a user to take an image of food using a camera on a terminal and send the image to a server; a means for the server to analyze the received image and identify the type of food using an image recognition algorithm; and a means for the server to collect information about the identified food using a chat-based AI model and determine whether the food is safe for babies and pets to consume. The system further includes a means for the server to send the determination result to the user's terminal and display the result on the terminal; a means for providing different information depending on whether the determination result relates to a baby or a pet; a means for the image recognition algorithm to use a deep learning model; and a means for the chat-based AI model to use natural language processing. This realizes a system that allows users at home to easily determine what is and is not edible for babies and pets.

[0006] "User" refers to an individual who wishes to use the system to verify food safety.

[0007] "Device" refers to a smartphone, tablet, or other internet-enabled electronic device used by a User.

[0008] "Camera" refers to a device built into the device that is used to take pictures of food.

[0009] "Image" refers to photographic data of food taken with a camera.

[0010] "Server" refers to a central processing unit that analyzes images and searches for information.

[0011] "Image recognition algorithm" refers to a set of computational steps used to analyze a received image and identify its content.

[0012] "Food type" refers to a particular category or specific name of food identified by an image recognition algorithm.

[0013] A "chat-type AI model" refers to an artificial intelligence system that uses natural language processing technology to provide comprehensive information in response to questions.

[0014] "Determining" refers to the process of determining whether a particular food is safe for babies or pets to consume based on the information collected.

[0015] A "deep learning model" is a type of artificial intelligence algorithm that uses a multi-layer neural network to analyze and learn from complex data.

[0016] "Natural language processing" refers to technology that allows computers to understand, generate, and manipulate human language. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] This invention relates to a system for supporting dietary management for babies and pets (dogs and cats). This system works by having the user take a picture of food using the camera on their device and send it to a server. The server analyzes the received image and identifies the type of food using an image recognition algorithm. Next, it uses a chat-based AI model to collect safety information about the identified food and, based on that information, determines whether the food is safe for babies and pets to consume. The result of the determination is sent to the user's device and displayed on the device.

[0039] Program processing explanation

[0040] 1. Image capture and transmission

[0041] User: Activates the camera on a device such as a smartphone or tablet and takes a picture of food. For example, the user takes a picture of an apple.

[0042] Device: The captured image is temporarily saved in local storage. After that, the network connection is confirmed and the captured image is sent to the server.

[0043] 2. Image Recognition

[0044] Server: Receives image data sent from the device and temporarily stores it.

[0045] Server: Analyzes the received image using an image recognition algorithm to identify the type of food in the image. The algorithm uses a deep learning model to learn the characteristics of the ingredients and classify them accordingly.

[0046] Server: Get a specific food type, such as "apple."

[0047] 3. Information Search

[0048] Server: Generates queries to the chat-based AI model based on the identified food types, for example, formulating specific questions such as "Is it safe for babies to eat apples?"

[0049] Server: Sends the generated queries to the chat-based AI model and collects comprehensive information from databases and the Internet.

[0050] Server: Analyzes the information returned as a response from the AI ​​model and determines the safety of food based on that information. For example, it may determine that "it is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."

[0051] 4. Results display

[0052] Server: Formats the data so that the obtained information and judgment results are presented to the user in an easy-to-understand manner.

[0053] Server: Sends the judgment result to the user's terminal.

[0054] Terminal: Analyzes the received results and displays them on the user interface. For example, it displays advice such as "Apples are suitable for babies, but it is recommended to cut them into small pieces or grate them."

[0055] Specific examples

[0056] Example 1: An apple for a baby

[0057] User: Take a photo of an apple, save it on the device, and send the image to the server.

[0058] Server: Identifies an apple using an image recognition algorithm. Sends a query to the chat-based AI model asking, "Is it safe for babies to eat apples?"

[0059] Server: Collects information from the AI ​​model and determines whether it is safe for babies to eat apples. Sends the result to the user's device.

[0060] Device: Displays the result: "Apples are suitable for babies, but it is recommended that they be cut into small pieces or grated."

[0061] Example 2: Chocolate for dogs

[0062] User: Take a photo of the chocolate and save it on the device. Send the image to the server.

[0063] Server: Identifies chocolate using an image recognition algorithm. Sends a query to the chat-based AI model asking, "Is it safe for dogs to eat chocolate?"

[0064] Server: Collects information from the AI ​​model and determines whether chocolate is harmful to dogs. Sends the result to the user's device.

[0065] Device: Displays the result: "Chocolate is harmful to dogs and should not be given to them."

[0066] This allows you to quickly and easily determine what foods your baby or pets can and cannot eat in your home.

[0067] The processing flow will be explained below.

[0068] Step 1:

[0069] The user activates the device's camera and takes an image of food, for example, an apple.

[0070] Step 2:

[0071] The device temporarily saves the captured image in local storage, then checks for network connectivity and sends the image data to the server's API endpoint.

[0072] Step 3:

[0073] The server receives the image data sent from the terminal, stores it temporarily, and returns a reception confirmation response to the terminal.

[0074] Step 4:

[0075] The server analyzes the received image data using an image recognition algorithm. Specifically, the image is input into a deep learning model for image recognition to identify the type of food.

[0076] Step 5:

[0077] The server obtains the identified food type (e.g., "apple") based on the results of the image recognition algorithm.

[0078] Step 6:

[0079] The server generates a query to the chat-based AI model based on the identified food type, for example, creating a specific question such as "Is it safe for babies to eat apples?"

[0080] Step 7:

[0081] The server generates queries and sends them to a chat-based AI model, which collects comprehensive information from databases and the Internet.

[0082] Step 8:

[0083] The server receives the response from the chat-based AI model and analyzes its content, obtaining, for example, information that "it is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."

[0084] Step 9:

[0085] The server determines the safety of the food based on the information obtained, and then formats the determination results and the information that supports them into a data format.

[0086] Step 10:

[0087] The server then sends the formatted result to the user's device. Once the transmission is confirmed, the server records a log and prepares for the next process.

[0088] Step 11:

[0089] The terminal receives the result data from the server, checks the integrity of the data, and returns a reception confirmation response to the server.

[0090] Step 12:

[0091] The device analyzes the received data and displays it on the user interface, for example, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them."

[0092] Through these steps, users can easily and quickly determine what foods babies and pets in the home can and cannot eat.

[0093] Example 1

[0094] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0095] In modern society, dietary management for babies and animals is an important issue, and there is a need for a system that can quickly and accurately obtain information for providing appropriate food. However, manually collecting and assessing information requires time and effort, and there is a risk of making decisions based on incorrect information. Conventional systems have had difficulty efficiently providing information needed to accurately determine food safety. The present invention aims to solve these problems and provide a system that supports dietary management for babies and animals.

[0096] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0097] In this invention, the server includes a means for a user to take an image of a food using a camera on a terminal and send the image to the server, a means for the server to analyze the received image and identify the type of food using an image recognition algorithm, and a means for the server to collect information about the identified food using a generative AI model and determine whether the food is safe for babies or animals to consume, thereby enabling users to easily identify foods suitable for babies and animals and confirm their safety.

[0098] A "user" is a person who uses the system to take pictures of food and transmits the pictures to the server.

[0099] A "terminal" refers to a computer device used by a user, such as a smartphone or tablet, that has a camera function and a network connection function.

[0100] A "camera" is an image capturing device built into a terminal.

[0101] "Food" refers to all food and drink consumed by babies and animals.

[0102] A "server" is a computer system that processes data submitted by users and performs specific functions.

[0103] An "image recognition algorithm" is a computational method for analyzing a received image and identifying objects within the image.

[0104] A "deep learning model" is a type of machine learning that uses a multi-layered neural network to learn patterns from large amounts of data and is used to analyze images, audio, and other data.

[0105] A "generative AI model" is an algorithm that generates a response in natural language in response to an input prompt.

[0106] A "prompt sentence" is a sentence entered to ask a generative AI model for specific information.

[0107] "Baby" refers to an infant or young child, especially a child within the first year of life.

[0108] "Animals" refers to pets kept at home, specifically mammals such as dogs and cats.

[0109] "Ingestion" refers to the act of taking food into the mouth and digesting and absorbing it.

[0110] "Displaying on the terminal" refers to making the results visible to the user in the form of text or graphics on the terminal screen.

[0111] The present invention relates to a system that supports the dietary management of babies and animals, and operates by allowing a user to take an image of food using a camera on a terminal and send the image to a server. The system operates as follows.

[0112] First, a user uses a device such as a smartphone or tablet to activate the camera and take an image of food. For example, the user takes a photo of an apple. The device temporarily saves the captured image in local storage and then checks for an Internet connection. Once it confirms that a network connection has been established, it sends the captured image to the server.

[0113] The server receives the image data sent from the device and temporarily stores it. It then uses an image recognition algorithm to analyze the received image. This algorithm uses deep learning models such as TensorFlow and PyTorch to learn the characteristics of food and classify it based on that. The server then identifies the type of food in the image and obtains the identified food type, such as "apple."

[0114] Next, the server generates a query to a generative AI model based on the identified food type. Specifically, it creates a question such as, "Is it safe for babies to eat apples?" The generative AI model used is OpenAI's GPT-4. The server sends the generated query to the generative AI model, which collects comprehensive information from databases and the internet.

[0115] The server analyzes the information returned as a response from the generative AI model and uses it to determine the safety of food. For example, it may determine that "apples are safe for babies to eat, but it is recommended that they be cut into small pieces or grated." The result is then formatted and converted into a user-friendly format.

[0116] Finally, the server sends the result of the judgment to the user's device, which then analyzes the result and displays it on the user interface. For example, it might say, "Apples are suitable for babies, but it is recommended that you cut them into small pieces or grate them."

[0117] Specific examples

[0118] Apples for babies

[0119] A user takes a photo of an apple, saves it on their device, and sends it to the server.

[0120] The server uses an image recognition algorithm to identify the apple and sends the generative AI model a query: "Is it safe for a baby to eat the apple?"

[0121] The server collects information from the AI ​​model, determines whether it is safe for babies to eat apples, and sends the result to the user's device.

[0122] The device will display the result: "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them."

[0123] Chocolate for animals

[0124] The user takes a photo of the chocolate, saves it on the device, and sends it to the server.

[0125] The server uses an image recognition algorithm to identify the chocolate and sends the generative AI model a query: "Is chocolate safe for dogs to eat?"

[0126] The server collects information from the AI ​​model, determines whether chocolate is harmful to dogs, and sends the result to the user's device.

[0127] The device will display the result: "Chocolate is harmful to dogs and should not be given to them."

[0128] This makes it possible to quickly and easily determine what foods babies and animals in the home can and cannot eat.

[0129] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0130] Step 1:

[0131] A user activates the camera on a device such as a smartphone or tablet and takes a picture of food. The input is the food to be photographed (e.g., an apple), and the output is an image file saved on the device (e.g., IMG_20230101.jpg). Specifically, the user opens the camera app, focuses on the apple, and presses the shutter button.

[0132] Step 2:

[0133] The device temporarily saves the captured image in local storage. It then checks the Internet connection. The input is the captured image file, and the output is the temporarily saved image file and the result of checking the network connection status. Specifically, it saves the image file in a specific folder and checks the Wi-Fi and mobile data connection status.

[0134] Step 3:

[0135] Once the device has confirmed a network connection, it uses the REST API to send image data to the server. The input is the temporarily saved image file and the network connection status, and the output is the image data sent to the server. Specifically, it creates an HTTP POST request, encodes the image file, and sends it to the specified URL on the server.

[0136] Step 4:

[0137] The server receives image data sent from the terminal and temporarily stores it. The input is the image data sent from the terminal, and the output is the image data temporarily stored in the server. Specifically, the received image file is stored in the server's temporary directory (e.g., / tmp).

[0138] Step 5:

[0139] The server analyzes the received image using an image recognition algorithm. The input is the temporarily stored image data, and the output is the identified type of food (e.g., apple). Specifically, the image file is input into a deep learning model (e.g., TensorFlow), which extracts and classifies the food's features.

[0140] Step 6:

[0141] The server generates a query to the generative AI model based on the identified food type. The input is the identified food type (e.g., apple), and the output is a generated prompt (e.g., "Is it safe for babies to eat apples?"). Specifically, the server uses the identified food type to create a question-style prompt.

[0142] Step 7:

[0143] The server sends the generated query to the generative AI model and collects information from databases and the Internet. The input is the generated prompt, and the output is the information returned by the generative AI model (e.g., "It is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."). Specific operations include creating an API request and sending the prompt to the generative AI model.

[0144] Step 8:

[0145] The server analyzes the response from the generative AI model and determines the safety of the food. The input is the information returned from the generative AI model, and the output is the safety judgment result (e.g., "safe"). Specifically, it extracts information about safety from the response text and determines the judgment result.

[0146] Step 9:

[0147] The server formats the resulting judgment results in a data format that is easy to understand for the user. The input is the safety judgment result, and the output is the formatted data (e.g., "Apples are suitable for babies, but it is recommended that they be cut into small pieces or grated"). Specific operations include converting the data into a format such as JSON.

[0148] Step 10:

[0149] The server sends the result of the judgment to the user's device. The input is the formatted data, and the output is the data received by the user's device. The specific operation is to send the data as an API response.

[0150] Step 11:

[0151] The device analyzes the received judgment result and displays it on the user interface. The input is the received judgment result data, and the output is the judgment result displayed to the user (e.g., "Apples are suitable for babies, but it is recommended that they be cut into small pieces or grated"). Specific actions include displaying a pop-up message or notification.

[0152] (Application example 1)

[0153] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0154] In recent years, the number of users of online food delivery services has increased, and there is a growing need to confirm in advance whether the food being ordered is safe, especially for households with babies or pets. However, current food delivery systems lack a function to quickly and easily determine whether the ordered food is suitable for babies or pets. This requires users to spend time researching various information themselves, and the reliability of that information cannot be guaranteed. Therefore, the present invention aims to provide a food delivery system that quickly and reliably determines the safety of food and provides it to users.

[0155] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0156] In this invention, the server includes: means for a user to take an image of food using a camera on a terminal and send the image to the server; means for the server to analyze the received image and identify the type of food using an image recognition algorithm; means for a food delivery application for a user to take an image of the food to be ordered and use the image to obtain safety information about the ingredients; means for the server to collect information about the identified food using a chat-based generative AI model and determine whether the food is safe for babies and pets to consume; and means for the server to send the determination result to the user's terminal and display the result on the terminal. This allows a user to use the food delivery application to check the effects of the food to be ordered on babies and pets and ensure its safety.

[0157] "Terminal" means a portable electronic device for inputting, processing, and outputting information.

[0158] A "camera" is a device for taking still images and videos, and is built into a terminal.

[0159] A "server" is a computer system that stores, processes, and distributes information over a network.

[0160] An "image recognition algorithm" is a computational method for analyzing image data and identifying specific objects or features.

[0161] "Food type" is information that indicates the specific classification or name of the food.

[0162] A "chat-type generative AI model" is an artificial intelligence model that answers questions and generates information based on text data.

[0163] A "food delivery application" is software that allows users to order and have food delivered online.

[0164] This invention is a system applied to a food delivery application that allows users to check the safety of food they plan to order. It consists of a terminal used by the user, a server, an image recognition algorithm, a generative AI model, and a food delivery application.

[0165] System configuration

[0166] 1. Terminal

[0167] The user uses a device such as a smartphone or tablet. The device has a built-in camera and is used to take pictures of food. The captured image data is temporarily stored on the device and then sent to the server.

[0168] 2. Server

[0169] The server has multiple functions. It receives and stores image data, analyzes it using an image recognition algorithm to identify the type of food, and uses a generative AI model to collect safety information about the identified food. It then sends the collected information to the user's device and displays it in an appropriate format.

[0170] 3. Image Recognition Algorithm

[0171] The server uses TensorFlow and PyTorch as image recognition algorithms to automatically identify the type of food from the captured image. By using a deep learning model, food can be identified with high accuracy.

[0172] 4. Generative AI Models

[0173] The generative AI model uses a chat-type AI model such as GPT-4. Once the type of food is identified, specific prompts such as "Is it safe for babies to eat XX?" or "Is it safe for dogs to eat XX?" are sent to the generative AI model to obtain relevant safety information.

[0174] 5. Food delivery applications

[0175] The food delivery application provides an interface for users to check the safety of the food they plan to order. The safety information sent from the server is displayed on the user's device, allowing them to easily check whether the food is suitable for babies and pets before ordering.

[0176] Processing flow

[0177] 1. Image capture

[0178] The user uses the smartphone camera to take a photo of the food they plan to order.

[0179] 2. Sending images

[0180] The device sends the captured image to the server, which receives and stores the image data.

[0181] 3. Image Analysis

[0182] The server uses image recognition algorithms to analyze the captured image and identify the type of food, such as "pizza," "spaghetti," or "carbonara."

[0183] 4. Safety assessment

[0184] Based on the type of food identified, prompts such as "Is it safe for babies to eat XX?" or "Is it safe for dogs to eat XX?" are sent to the generative AI model, and safety information is collected from the generative AI model.

[0185] 5. Display results

[0186] The server analyzes the collected data and sends it to the device in an easy-to-understand format. The results are displayed on the user's device, allowing them to see the effects on their baby or pet.

[0187] Specific examples

[0188] Example 1: Pizza for babies

[0189] Prompt: "Is it safe for babies to eat pizza?"

[0190] The user takes a photo of the pizza and sends it to the system.

[0191] The server uses an image recognition algorithm to identify the pizza and send a prompt to the generative AI model.

[0192] Based on information from the AI ​​model, a decision is made as to whether the pizza is suitable.

[0193] Example 2: Carbonara for dogs

[0194] Prompt: "Is it safe for dogs to eat carbonara?"

[0195] The user takes a photo of the carbonara and sends it to the system.

[0196] The server uses an image recognition algorithm to identify carbonara and send a prompt to the generative AI model.

[0197] Based on information from the AI ​​model, the results will show whether carbonara is harmful to dogs.

[0198] In this way, users can quickly and easily check whether the food they plan to order is suitable for babies or pets.

[0199] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0200] Step 1:

[0201] Input: The user takes a photo of the food they plan to order.

[0202] How it works: The user activates the device's camera, takes a picture of the food they plan to order (e.g., pizza, carbonara), and temporarily saves it to the device's local storage.

[0203] Step 2:

[0204] Input: Captured image data.

[0205] Operation: The device checks for network connectivity and sends the stored image data to the server.

[0206] Output: Image data is sent to the server.

[0207] Step 3:

[0208] Input: Image data received by the server.

[0209] How it works: The server temporarily stores the image data it receives, then applies image recognition algorithms (using TensorFlow or PyTorch) to identify the type of food in the image.

[0210] Output: The type of food is identified (e.g. "pizza" or "carbonara").

[0211] Step 4:

[0212] Input: The type of food identified.

[0213] How it works: The server generates queries for the generative AI model based on the identified food types, creating specific questions like "Is it safe for babies to eat pizza?" or "Is it safe for dogs to eat carbonara?"

[0214] Output: The generated query.

[0215] Step 5:

[0216] Input: The generated query.

[0217] How it works: The server sends the generated queries to the generative AI model, which collects comprehensive information from the internet and databases. The generative AI model then generates answers to the questions and provides safety information.

[0218] Output: Response (safety information) from the generative AI model.

[0219] Step 6:

[0220] Input: The response from the generative AI model.

[0221] How it works: The server analyzes the collected information and formats it in a way that is easy for the user to understand. For example, it will give a judgment result such as "Pizza is suitable for babies, but it is recommended to cut it into small pieces."

[0222] Output: The result of the test for the user.

[0223] Step 7:

[0224] Input: Judgment result.

[0225] Operation: The server sends the result of the judgment to the user's device. The device analyzes the received result and displays it on the interface of the food delivery application.

[0226] Output: The result displayed on the device (e.g., "Pizza is suitable for babies, but it is recommended to cut it into small pieces").

[0227] In this way, users can easily check whether the food they plan to order is suitable for babies or pets.

[0228] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0229] This invention relates to a system that supports dietary management for babies and pets (dogs and cats), and in particular, it combines an emotion engine that recognizes the user's emotions. This system provides a process for quickly determining the safety of food.

[0230] Program processing explanation

[0231] 1. Image capture and transmission

[0232] User: Activates the camera on a device such as a smartphone or tablet and takes a picture of food. For example, the user takes a picture of an apple.

[0233] Device: The captured image is temporarily saved in local storage, and after checking the network connection, the image data is sent to the server's API endpoint.

[0234] 2. Image Recognition

[0235] Server: Receives image data sent from the terminal, temporarily stores it, and returns a reception confirmation response to the terminal.

[0236] Server: Analyzes the received image using an image recognition algorithm to identify the type of food in the image. The algorithm uses a deep learning model to learn the characteristics of the ingredients and classify them accordingly.

[0237] Server: Get a specific food type, such as "apple."

[0238] 3. Information Search

[0239] Server: Generates a query to the chat-based AI model based on the identified food type, for example, formulating a specific question such as "Is it safe for babies to eat apples?"

[0240] Server: Sends the generated queries to the chat-based AI model and collects comprehensive information from databases and the Internet.

[0241] Server: Receives the response from the AI ​​model and analyzes its content. For example, it obtains information such as "It is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."

[0242] 4. Emotion recognition

[0243] Device: When a user takes a picture, the emotion engine is activated and analyzes the user's facial expressions and voice. This analysis is performed in real time to understand the user's emotional state.

[0244] Terminal: Transmits the acquired user emotional state data to the server.

[0245] 5. Judgment and result display

[0246] Server: Determines the safety of food based on the obtained information and the user's emotional state. For example, if the result is that it is safe for babies or pets, or if the user's emotional state is deemed unsafe, it provides additional advice or a warning.

[0247] Server: Formats the judgment results and the information that forms the basis for them into a data format.

[0248] Server: The formatted result is sent to the user's device. Once the transmission is confirmed, the server records a log and prepares for the next process.

[0249] Terminal: Receives the result data from the server, checks the integrity of the data, and returns a reception confirmation response to the server.

[0250] Terminal: Analyzes the received data and displays it on the user interface. For example, it might display "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them," and if the user's emotions are anxious, it might display additional advice such as "If you are worried, start with a small amount and monitor the baby's reaction."

[0251] Specific examples

[0252] Example 1: Applying apples and emotion recognition to babies

[0253] User: Take a photo of an apple, save it on the device, and send the image to the server.

[0254] Server: Identifies an apple using an image recognition algorithm. Sends a query to the chat-based AI model asking, "Is it safe for babies to eat apples?"

[0255] Server: Collects information from the AI ​​model and determines whether it is safe for babies to eat apples.

[0256] Device: When the user takes a photo, facial recognition and voice analysis are used to determine the emotional state as anxiety.

[0257] Server: The judgement result was accompanied by additional advice: "If you are concerned, start with a small amount and monitor the condition."

[0258] Device: The result is displayed as follows: "Apples are suitable for babies, but it is recommended that they be cut into small pieces or grated. If you are concerned, start with a small amount and monitor the baby's condition."

[0259] Example 2: Applying chocolate and emotion recognition to dogs

[0260] User: Take a photo of the chocolate and save it on the device. Send the image to the server.

[0261] Server: Identifies chocolate using an image recognition algorithm. Sends a query to the chat-based AI model asking, "Is it safe for dogs to eat chocolate?"

[0262] Server: Collects information from the AI ​​model and determines that chocolate is harmful to dogs.

[0263] Device: When the user takes a photo, facial recognition and voice analysis are used to determine the emotional state as surprise.

[0264] Server: The verdict included a warning that "chocolate is extremely dangerous to dogs and should never be given to them."

[0265] Device: Displays the result: "Chocolate is harmful to dogs and should never be given to them," with an additional warning: "Move out of reach immediately."

[0266] This allows the system to recognize the user's emotions and provide appropriate advice and warnings based on those emotions, making it easy and safe to manage the diet of babies and pets at home.

[0267] The processing flow will be explained below.

[0268] Step 1:

[0269] The user activates the camera on the device and takes an image of food, for example, an apple.

[0270] Step 2:

[0271] The device temporarily saves the captured image in local storage, then checks for network connectivity and sends the image data to the server's API endpoint.

[0272] Step 3:

[0273] The server receives the image data sent from the terminal, stores it temporarily, and returns a reception confirmation response to the terminal.

[0274] Step 4:

[0275] The server analyzes the received image data using an image recognition algorithm that identifies food characteristics and identifies a particular food type.

[0276] Step 5:

[0277] The server obtains the identified food type based on the results of the image recognition algorithm, for example, identifying "apple."

[0278] Step 6:

[0279] The server generates a query to the chat-based AI model based on the type of food identified, for example, "Is it safe for babies to eat apples?"

[0280] Step 7:

[0281] The server generates queries and sends them to a chat-based AI model to collect information about food safety.

[0282] Step 8:

[0283] The server receives the response from the chat-based AI model and analyzes its content, obtaining, for example, information that "it is safe for babies to eat apples, and it is recommended that they be cut into small pieces or grated."

[0284] Step 9:

[0285] The device analyzes the user's facial expressions and voice in real time to identify the user's emotional state. For example, if the user looks anxious or speaks in an anxious voice, it will be determined that the user is anxious.

[0286] Step 10:

[0287] The device transmits the user's emotional state data to the server, which then incorporates this data into the decision-making process.

[0288] Step 11:

[0289] The server determines the safety of food based on the information obtained and the user's emotional state. For example, if the user seems anxious, it may provide additional advice or warnings.

[0290] Step 12:

[0291] The server formats the information, including the judgment result and any additional advice, into a data format and prepares it for presentation to the user.

[0292] Step 13:

[0293] The server then sends the formatted result to the user's device. Once the transmission is confirmed, the server records a log and prepares for the next process.

[0294] Step 14:

[0295] The terminal receives the result data from the server, checks the integrity of the data, and returns a reception confirmation response to the server.

[0296] Step 15:

[0297] The device analyzes the received data and displays it on the user interface. For example, it might say, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them. If you're worried, start with a small amount and monitor your baby's condition."

[0298] This allows the system to recognize the user's emotions and provide appropriate advice and warnings based on those emotions, making it easy and safe to manage the diet of babies and pets at home.

[0299] Example 2

[0300] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0301] Currently, there is a lack of systems that provide appropriate information quickly and accurately for managing the diet of babies and pets. Furthermore, there are no systems that can provide advice or warnings that take into account the user's emotions. This leaves users with insufficient support to provide food to their babies and pets with peace of mind. Furthermore, in existing systems, information collection to determine food safety is often done manually, which is a very time-consuming process.

[0302] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0303] In this invention, the server includes: a means for a user to take an image of food using a camera on the terminal and send the image to the server; a means for the terminal to analyze the user's facial expressions and voice and acquire emotion data; a means for the server to analyze the received image and identify the type of food using an image recognition algorithm; a means for the server to collect information about the identified food using a generative AI model and determine whether the food is safe for babies or pets to ingest; and a means for the server to send additional advice or warnings based on the determination results and emotion data to the user's terminal and display the results on the terminal. This allows the user to provide quick and appropriate information when feeding babies or pets, ensuring safety, and providing advice or warnings according to the user's emotions.

[0304] "Terminal" refers to a smartphone, tablet, or other portable information processing device used by a user.

[0305] "Server" refers to a computer system that receives, stores, analyzes, searches for information, and transmits results of image data.

[0306] "User" refers to a person who uses the system to manage the diet of a baby or pet.

[0307] "Image recognition algorithm" refers to a computational method for analyzing received image data and identifying the type of food depicted in the image.

[0308] A "deep learning model" refers to an artificial intelligence technology that learns and classifies the characteristics of ingredients as part of an image recognition algorithm.

[0309] "Generative AI model" refers to a system that uses chat-based artificial intelligence to collect and generate information about food.

[0310] A "prompt" refers to a query or question that is input to a generative AI model.

[0311] "Emotion engine" refers to technology that analyzes a user's facial expressions and voice to identify their emotional state.

[0312] "Determination result" refers to data containing the evaluation results of the server's evaluation of food safety.

[0313] "Emotion data" refers to data that indicates the user's emotional state obtained through facial recognition or voice analysis.

[0314] "Additional advice or warning" refers to supplementary advice or a message urging caution that is provided to the user based on the judgment result.

[0315] This invention relates to a system that supports dietary management for babies and pets, and in particular, it combines an emotion engine that recognizes the user's emotions. This system provides a process for quickly determining the safety of food.

[0316] A user takes a picture of food using a device such as a smartphone or tablet. For example, the user takes a picture of an apple. The device temporarily saves the captured image in local storage, and after confirming network connectivity, sends the image data to the server's API endpoint.

[0317] The server receives the image data sent from the device and temporarily stores it. It then returns a receipt confirmation response to the device. It then uses an image recognition algorithm to analyze the received image and identify the type of food in the image. For example, it uses a deep learning framework such as TensorFlow or PyTorch. The algorithm learns the characteristics of the ingredients contained in the image and classifies them based on that. The server then obtains the identified type of food, such as "apple."

[0318] The server then generates a query to a generative AI model (such as OpenAI's ChatGPT) based on the identified food type, for example, creating a specific question such as "Is it safe for babies to eat apples?" The generated query is a prompt sentence of the form:

[0319] "Is it safe for babies to eat apples?"

[0320] The server sends this prompt to the generative AI model, collects comprehensive information from the internet and databases, receives the response from the generative AI model, and analyzes its content. For example, the server obtains the information that "It is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."

[0321] When a user takes a picture, the device runs an emotion engine (such as the Affectiva SDK) to analyze the user's facial expressions and voice. This analysis is performed in real time to understand the user's emotional state. The device then transmits the acquired data on the user's emotional state to the server.

[0322] The server determines the safety of the food based on the obtained information and the user's emotional state. For example, even if the result indicates that the food is safe for babies or pets, if the user's emotions are deemed to be uneasy, the server will add additional advice or warnings. The server then formats the result and the information that forms the basis of the result into a data format and sends the formatted result to the user's device. Once the transmission is confirmed, the server records the log and prepares for the next process.

[0323] The device receives the result data sent from the server and checks the integrity of the data. It then returns a reception confirmation response to the server. Finally, the device analyzes the received data and displays it on the user interface. For example, it might display, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them. If you are concerned, start with a small amount and monitor the baby's condition," and if the user's emotions are uneasy, it might display additional advice such as, "If you are concerned, start with a small amount and monitor the baby's condition."

[0324] This system recognizes the user's emotions and provides appropriate advice and warnings based on those emotions, making it easy and safe to manage the diet of babies and pets at home.

[0325] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0326] Step 1:

[0327] A user takes a picture of food using the device's camera. Specifically, the user launches the camera app, places food (e.g., an apple) in the center of the screen, and presses the shutter button. The input is the camera operation and the captured food image, and the output is an image file saved in local storage.

[0328] Step 2:

[0329] The device temporarily saves the image captured in local storage and checks the network connection. Specifically, it saves the image file in the device's local storage and then checks whether the WiFi or mobile data connection is enabled. The input is the image file captured in step 1 and the network status, and the output is the state ready to be sent.

[0330] Step 3:

[0331] The device sends image data to the server's API endpoint. Specifically, the device sends the image file to the server as an HTTP POST request. The input is the saved image file and the server's API endpoint, and the output is the completion status of the transmission to the server.

[0332] Step 4:

[0333] The server receives image data sent from the terminal and temporarily stores it. Specifically, the server extracts image data from the HTTP request it receives and stores it in a temporary directory on the server. The input is the image data sent from the terminal, and the output is the temporarily stored image file and a response confirming receipt.

[0334] Step 5:

[0335] The server uses an image recognition algorithm to analyze the received image and identify the type of food. Specifically, the server analyzes the image using a deep learning model such as TensorFlow or PyTorch to identify the type of food (e.g., "apple"). The input is a temporarily saved image file, and the output is information about the identified type of food (e.g., "apple").

[0336] Step 6:

[0337] The server generates a query to the generative AI model based on the identified food type. Specifically, the server generates a prompt sentence, "Is it safe for babies to eat apples?" and sends it to the generative AI model. The input is food type information (e.g., "apple"), and the output is the generated query. An example of a prompt sentence: "Is it safe for babies to eat apples?"

[0338] Step 7:

[0339] The server sends the generated query to the generative AI model and collects information from databases and the Internet. Specifically, the server sends a prompt to the generative AI model as an HTTP request and receives food safety information in response. The input is the generated query, and the output is the received safety information.

[0340] Step 8:

[0341] The server analyzes the response from the AI ​​model and determines the safety of the food. Specifically, it analyzes the received safety information and obtains a judgment result such as "It is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated." The input is the received safety information, and the output is the judgment result.

[0342] Step 9:

[0343] When a user takes a picture, the device uses an emotion engine to analyze the user's facial expressions and voice to identify their emotional state. Specifically, the device's front camera and microphone are used to collect facial expressions and voice in real time, which are then analyzed by the emotion engine. The input is the user's facial expression and voice data, and the output is the identified emotional state data (e.g., "anxiety").

[0344] Step 10:

[0345] The emotional state data acquired by the device is sent to the server. Specifically, the emotional state data is converted into JSON format and sent to the server as an HTTP request. The input is the emotional state data, and the output is the status of completion of transmission to the server.

[0346] Step 11:

[0347] The server determines the safety of the food based on the information and emotional state obtained, and generates additional advice or a warning. Specifically, it adds additional advice to the judgment result, such as "If you are concerned, start with a small amount and monitor the situation." The input is the judgment result and emotional state data, and the output is the final judgment result and additional advice.

[0348] Step 12:

[0349] The server sends the final judgment result and additional advice to the user's device. Specifically, the server formats the data in JSON format and sends it to the user's device as an HTTP request. The input is the final judgment result and additional advice, and the output is the status of completion of transmission to the device.

[0350] Step 13:

[0351] The terminal receives the result data from the server and checks the integrity of the data. Specifically, it checks the received data using a checksum or other method and returns a response confirming receipt to the server. The input is the result data sent from the server, and the output is the response confirming receipt.

[0352] Step 14:

[0353] The device analyzes the result data and displays it on the user interface. Specifically, it displays the results and advice in the application's UI component, informing the user, for example, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them. If you are concerned, start with a small amount and monitor the baby's condition." The input is the analyzed result data, and the output is the display on the user interface.

[0354] (Application example 2)

[0355] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0356] Avoiding the risk of minors and animals accidentally ingesting harmful foods is an important issue in home dietary management. Users often have concerns or doubts about food safety, and appropriate information and advice that take these feelings into account is needed. However, existing systems rarely offer the functionality to integrate emotion recognition and food safety assessment in real time. Therefore, a system that allows users to manage their diet with peace of mind is needed.

[0357] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for a user to take an image of a food using a camera of a terminal and send the image to an information processing device; means for the information processing device to analyze the received image and identify the type of food using an image recognition algorithm; means for the information processing device to collect information about the identified food using an interactive AI model and determine whether the food is safe for minors or animals to consume; means for the information processing device to send the determination result to the user's terminal and display the result on the terminal; means for the terminal to analyze emotions from the user's facial expressions and voice and send the emotions to the information processing device; and means for the information processing device to add additional advice or warnings based on the obtained emotion data. This makes it possible to determine the safety of food in real time and provide appropriate information and advice according to the user's emotions.

[0358] A "terminal" is a portable information and communication device operated by a user, and is equipped with a camera and a display.

[0359] An "information processing device" is a computer system that analyzes received data and processes it using a specific algorithm.

[0360] An "image recognition algorithm" is a computational method for analyzing image data to identify specific objects or features.

[0361] An "interactive AI model" is a program that uses artificial intelligence to answer questions and search for information in natural language.

[0362] A "minor" is a child or young person who is not legally recognized as an adult.

[0363] "Animals" are living creatures such as dogs and cats kept as pets in homes.

[0364] "Emotion data" is information that indicates the emotional state of the user analyzed from facial expressions and voice.

[0365] "Additional advice and warnings" are supplemental information provided to help users provide food safely and with peace of mind.

[0366] System Program Overview

[0367] The system that realizes this application example consists of the following major components:

[0368] 1. Hardware

[0369] Terminal: A portable information and communication device (smartphone, tablet, etc.) operated by a user, equipped with a camera and display.

[0370] Information processing device: A server system that analyzes received data and processes it using specific algorithms. It is desirable for this server to be equipped with a high-performance CPU and GPU.

[0371] 2. Software

[0372] Image recognition algorithm: Built using deep learning libraries such as TensorFlow.

[0373] Conversational AI model: An engine used for natural language processing and question answering.

[0374] Emotion recognition software: Programs for facial expression recognition and speech analysis (OpenCV and other facial expression analysis tools).

[0375] Program processing overview

[0376] 1. Image capture and transmission

[0377] The user takes a photo of the food using the device's camera.

[0378] The terminal transmits the captured image to the information processing device.

[0379] 2. Image Recognition

[0380] The server analyzes the received images and uses image recognition algorithms to identify the type of food.

[0381] It uses a deep learning model to analyze the features of food in an image and identify its type.

[0382] 3. Information Search

[0383] The server sends information about the identified food as a query to the conversational AI model.

[0384] The AI ​​model collects the necessary information from databases and the internet to determine whether the food is safe for minors and animals.

[0385] 4. Emotion recognition

[0386] When a user takes a picture, the device analyzes emotions in real time from facial expressions and voice.

[0387] The analyzed emotion data is transmitted to an information processing device.

[0388] 5. Judgment and result display

[0389] The server generates additional advice and warnings based on the food safety assessment results and emotion data, and sends them to the user's device.

[0390] The terminal displays the received results on a user interface.

[0391] Specific examples

[0392] Example 1: Judging whether an apple is suitable for babies

[0393] 1. A user takes a photo of an apple and asks the camera, "Are apples safe for babies?"

[0394] 2. The device sends the emotional data analyzed as "anxiety" to the server.

[0395] 3. The server generates advice such as, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them. If you are concerned, start with a small amount and monitor their condition." and provides it to the user.

[0396] Prompt Sentence Examples

[0397] Take a photo of the food you are about to give your baby. Then ask, "Is an apple safe for my baby?" Then, enter your emotion (e.g., "anxious" or "surprised").

[0398] Specific examples of the technologies used

[0399] Image recognition algorithm: TensorFlow

[0400] Conversational AI model: Natural language processing engine

[0401] Emotion recognition software: OpenCV

[0402] These technologies allow users to confidently verify food safety and provide appropriate diets for minors and animals.

[0403] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0404] Step 1:

[0405] The user takes a picture of food using the device's camera and temporarily saves the captured image data in local storage. This is done using a camera application or the device's camera function. The input is the captured image data, and the output is an image file saved in the device's local storage.

[0406] Step 2:

[0407] The device checks the network connection and sends the captured image data to the information processing device. The transmission is performed using HTTP or other communication protocols. The input is an image file saved in local storage, and the output is the image data sent to the server.

[0408] Step 3:

[0409] The server temporarily stores the received image data and begins analysis using an image recognition algorithm. The type of food is identified using a deep learning library such as TensorFlow. The input is the image data stored on the server, and the output is the type of food identified through analysis.

[0410] Step 4:

[0411] The server sends a query to the conversational AI model based on the identified food type to collect safety information about the food. For example, a query such as "Is it safe for babies to eat apples?" is generated and the AI ​​model searches for information. The input is the identified food type information, and the output is information about the food obtained from the AI ​​model.

[0412] Step 5:

[0413] The device analyzes emotions from the user's facial expressions and voice when taking an image. It uses OpenCV and other facial expression analysis tools to obtain the user's emotional data in real time. The input is the user's facial expression and voice data, and the output is the emotional data obtained through the analysis.

[0414] Step 6:

[0415] The device sends the analyzed emotional data to the server using a communication method such as the HTTP protocol. The input is the emotional data analyzed on the device, and the output is the emotional data sent to the server.

[0416] Step 7:

[0417] The server generates additional advice and warnings based on food safety information and emotion data. Based on the obtained data, it generates appropriate messages and adds advice to alleviate the user's anxiety. The input is food safety information and emotion data, and the output is the generated advice or warning message.

[0418] Step 8:

[0419] The server generates advice and warning messages and sends them to the user's terminal. The input is the generated message and the output is the message sent to the user's terminal.

[0420] Step 9:

[0421] The terminal analyzes the received result data and displays it on the user interface. Based on this information, the user can decide whether to provide food to minors or animals. The input is the received message data, and the output is the information displayed on the user interface.

[0422] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0423] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0424] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0425] [Second embodiment]

[0426] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0427] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0428] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0429] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0430] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0431] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0432] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0433] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0434] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0435] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0436] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0437] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0438] This invention relates to a system for supporting dietary management for babies and pets (dogs and cats). This system works by having the user take a picture of food using the camera on their device and send it to a server. The server analyzes the received image and identifies the type of food using an image recognition algorithm. Next, it uses a chat-based AI model to collect safety information about the identified food and, based on that information, determines whether the food is safe for babies and pets to consume. The result of the determination is sent to the user's device and displayed on the device.

[0439] Program processing explanation

[0440] 1. Image capture and transmission

[0441] User: Activates the camera on a device such as a smartphone or tablet and takes a picture of food. For example, the user takes a picture of an apple.

[0442] Device: The captured image is temporarily saved in local storage. After that, the network connection is confirmed and the captured image is sent to the server.

[0443] 2. Image Recognition

[0444] Server: Receives image data sent from the device and temporarily stores it.

[0445] Server: Analyzes the received image using an image recognition algorithm to identify the type of food in the image. The algorithm uses a deep learning model to learn the characteristics of the ingredients and classify them accordingly.

[0446] Server: Get a specific food type, such as "apple."

[0447] 3. Information Search

[0448] Server: Generates queries to the chat-based AI model based on the identified food types, for example, formulating specific questions such as "Is it safe for babies to eat apples?"

[0449] Server: Sends the generated queries to the chat-based AI model and collects comprehensive information from databases and the Internet.

[0450] Server: Analyzes the information returned as a response from the AI ​​model and determines the safety of food based on that information. For example, it may determine that "it is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."

[0451] 4. Results display

[0452] Server: Formats the data so that the obtained information and judgment results are presented to the user in an easy-to-understand manner.

[0453] Server: Sends the judgment result to the user's terminal.

[0454] Terminal: Analyzes the received results and displays them on the user interface. For example, it displays advice such as "Apples are suitable for babies, but it is recommended to cut them into small pieces or grate them."

[0455] Specific examples

[0456] Example 1: An apple for a baby

[0457] User: Take a photo of an apple, save it on the device, and send the image to the server.

[0458] Server: Identifies an apple using an image recognition algorithm. Sends a query to the chat-based AI model asking, "Is it safe for babies to eat apples?"

[0459] Server: Collects information from the AI ​​model and determines whether it is safe for babies to eat apples. Sends the result to the user's device.

[0460] Device: Displays the result: "Apples are suitable for babies, but it is recommended that they be cut into small pieces or grated."

[0461] Example 2: Chocolate for dogs

[0462] User: Take a photo of the chocolate and save it on the device. Send the image to the server.

[0463] Server: Identifies chocolate using an image recognition algorithm. Sends a query to the chat-based AI model asking, "Is it safe for dogs to eat chocolate?"

[0464] Server: Collects information from the AI ​​model and determines whether chocolate is harmful to dogs. Sends the result to the user's device.

[0465] Device: Displays the result: "Chocolate is harmful to dogs and should not be given to them."

[0466] This allows you to quickly and easily determine what foods your baby or pets can and cannot eat in your home.

[0467] The processing flow will be explained below.

[0468] Step 1:

[0469] The user activates the device's camera and takes an image of food, for example, an apple.

[0470] Step 2:

[0471] The device temporarily saves the captured image in local storage, then checks for network connectivity and sends the image data to the server's API endpoint.

[0472] Step 3:

[0473] The server receives the image data sent from the terminal, stores it temporarily, and returns a reception confirmation response to the terminal.

[0474] Step 4:

[0475] The server analyzes the received image data using an image recognition algorithm. Specifically, the image is input into a deep learning model for image recognition to identify the type of food.

[0476] Step 5:

[0477] The server obtains the identified food type (e.g., "apple") based on the results of the image recognition algorithm.

[0478] Step 6:

[0479] The server generates a query to the chat-based AI model based on the identified food type, for example, creating a specific question such as "Is it safe for babies to eat apples?"

[0480] Step 7:

[0481] The server generates queries and sends them to a chat-based AI model, which collects comprehensive information from databases and the Internet.

[0482] Step 8:

[0483] The server receives the response from the chat-based AI model and analyzes its content, obtaining, for example, information that "it is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."

[0484] Step 9:

[0485] The server determines the safety of the food based on the information obtained, and then formats the determination results and the information that supports them into a data format.

[0486] Step 10:

[0487] The server then sends the formatted result to the user's device. Once the transmission is confirmed, the server records a log and prepares for the next process.

[0488] Step 11:

[0489] The terminal receives the result data from the server, checks the integrity of the data, and returns a reception confirmation response to the server.

[0490] Step 12:

[0491] The device analyzes the received data and displays it on the user interface, for example, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them."

[0492] Through these steps, users can easily and quickly determine what foods babies and pets in the home can and cannot eat.

[0493] Example 1

[0494] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0495] In modern society, dietary management for babies and animals is an important issue, and there is a need for a system that can quickly and accurately obtain information for providing appropriate food. However, manually collecting and assessing information requires time and effort, and there is a risk of making decisions based on incorrect information. Conventional systems have had difficulty efficiently providing information needed to accurately determine food safety. The present invention aims to solve these problems and provide a system that supports dietary management for babies and animals.

[0496] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0497] In this invention, the server includes a means for a user to take an image of a food using a camera on a terminal and send the image to the server, a means for the server to analyze the received image and identify the type of food using an image recognition algorithm, and a means for the server to collect information about the identified food using a generative AI model and determine whether the food is safe for babies or animals to consume, thereby enabling users to easily identify foods suitable for babies and animals and confirm their safety.

[0498] A "user" is a person who uses the system to take pictures of food and transmits the pictures to the server.

[0499] A "terminal" refers to a computer device used by a user, such as a smartphone or tablet, that has a camera function and a network connection function.

[0500] A "camera" is an image capturing device built into a terminal.

[0501] "Food" refers to all food and drink consumed by babies and animals.

[0502] A "server" is a computer system that processes data submitted by users and performs specific functions.

[0503] An "image recognition algorithm" is a computational method for analyzing a received image and identifying objects within the image.

[0504] A "deep learning model" is a type of machine learning that uses a multi-layered neural network to learn patterns from large amounts of data and is used to analyze images, audio, and other data.

[0505] A "generative AI model" is an algorithm that generates a response in natural language in response to an input prompt.

[0506] A "prompt sentence" is a sentence entered to ask a generative AI model for specific information.

[0507] "Baby" refers to an infant or young child, especially a child within the first year of life.

[0508] "Animals" refers to pets kept at home, specifically mammals such as dogs and cats.

[0509] "Ingestion" refers to the act of taking food into the mouth and digesting and absorbing it.

[0510] "Displaying on the terminal" refers to making the results visible to the user in the form of text or graphics on the terminal screen.

[0511] The present invention relates to a system that supports the dietary management of babies and animals, and operates by allowing a user to take an image of food using a camera on a terminal and send the image to a server. The system operates as follows.

[0512] First, a user uses a device such as a smartphone or tablet to activate the camera and take an image of food. For example, the user takes a photo of an apple. The device temporarily saves the captured image in local storage and then checks for an Internet connection. Once it confirms that a network connection has been established, it sends the captured image to the server.

[0513] The server receives the image data sent from the device and temporarily stores it. It then uses an image recognition algorithm to analyze the received image. This algorithm uses deep learning models such as TensorFlow and PyTorch to learn the characteristics of food and classify it based on that. The server then identifies the type of food in the image and obtains the identified food type, such as "apple."

[0514] Next, the server generates a query to a generative AI model based on the identified food type. Specifically, it creates a question such as, "Is it safe for babies to eat apples?" The generative AI model used is OpenAI's GPT-4. The server sends the generated query to the generative AI model, which collects comprehensive information from databases and the internet.

[0515] The server analyzes the information returned as a response from the generative AI model and uses it to determine the safety of food. For example, it may determine that "apples are safe for babies to eat, but it is recommended that they be cut into small pieces or grated." The result is then formatted and converted into a user-friendly format.

[0516] Finally, the server sends the result of the judgment to the user's device, which then analyzes the result and displays it on the user interface. For example, it might say, "Apples are suitable for babies, but it is recommended that you cut them into small pieces or grate them."

[0517] Specific examples

[0518] Apples for babies

[0519] A user takes a photo of an apple, saves it on their device, and sends it to the server.

[0520] The server uses an image recognition algorithm to identify the apple and sends the generative AI model a query: "Is it safe for a baby to eat the apple?"

[0521] The server collects information from the AI ​​model, determines whether it is safe for babies to eat apples, and sends the result to the user's device.

[0522] The device will display the result: "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them."

[0523] Chocolate for animals

[0524] The user takes a photo of the chocolate, saves it on the device, and sends it to the server.

[0525] The server uses an image recognition algorithm to identify the chocolate and sends the generative AI model a query: "Is chocolate safe for dogs to eat?"

[0526] The server collects information from the AI ​​model, determines whether chocolate is harmful to dogs, and sends the result to the user's device.

[0527] The device will display the result: "Chocolate is harmful to dogs and should not be given to them."

[0528] This makes it possible to quickly and easily determine what foods babies and animals in the home can and cannot eat.

[0529] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0530] Step 1:

[0531] A user activates the camera on a device such as a smartphone or tablet and takes a picture of food. The input is the food to be photographed (e.g., an apple), and the output is an image file saved on the device (e.g., IMG_20230101.jpg). Specifically, the user opens the camera app, focuses on the apple, and presses the shutter button.

[0532] Step 2:

[0533] The device temporarily saves the captured image in local storage. It then checks the Internet connection. The input is the captured image file, and the output is the temporarily saved image file and the result of checking the network connection status. Specifically, it saves the image file in a specific folder and checks the Wi-Fi and mobile data connection status.

[0534] Step 3:

[0535] Once the device has confirmed a network connection, it uses the REST API to send image data to the server. The input is the temporarily saved image file and the network connection status, and the output is the image data sent to the server. Specifically, it creates an HTTP POST request, encodes the image file, and sends it to the specified URL on the server.

[0536] Step 4:

[0537] The server receives image data sent from the terminal and temporarily stores it. The input is the image data sent from the terminal, and the output is the image data temporarily stored in the server. Specifically, the received image file is stored in the server's temporary directory (e.g., / tmp).

[0538] Step 5:

[0539] The server analyzes the received image using an image recognition algorithm. The input is the temporarily stored image data, and the output is the identified type of food (e.g., apple). Specifically, the image file is input into a deep learning model (e.g., TensorFlow), which extracts and classifies the food's features.

[0540] Step 6:

[0541] The server generates a query to the generative AI model based on the identified food type. The input is the identified food type (e.g., apple), and the output is a generated prompt (e.g., "Is it safe for babies to eat apples?"). Specifically, the server uses the identified food type to create a question-style prompt.

[0542] Step 7:

[0543] The server sends the generated query to the generative AI model and collects information from databases and the Internet. The input is the generated prompt, and the output is the information returned by the generative AI model (e.g., "It is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."). Specific operations include creating an API request and sending the prompt to the generative AI model.

[0544] Step 8:

[0545] The server analyzes the response from the generative AI model and determines the safety of the food. The input is the information returned from the generative AI model, and the output is the safety judgment result (e.g., "safe"). Specifically, it extracts information about safety from the response text and determines the judgment result.

[0546] Step 9:

[0547] The server formats the resulting judgment results in a data format that is easy to understand for the user. The input is the safety judgment result, and the output is the formatted data (e.g., "Apples are suitable for babies, but it is recommended that they be cut into small pieces or grated"). Specific operations include converting the data into a format such as JSON.

[0548] Step 10:

[0549] The server sends the result of the judgment to the user's device. The input is the formatted data, and the output is the data received by the user's device. The specific operation is to send the data as an API response.

[0550] Step 11:

[0551] The device analyzes the received judgment result and displays it on the user interface. The input is the received judgment result data, and the output is the judgment result displayed to the user (e.g., "Apples are suitable for babies, but it is recommended that they be cut into small pieces or grated"). Specific actions include displaying a pop-up message or notification.

[0552] (Application example 1)

[0553] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0554] In recent years, the number of users of online food delivery services has increased, and there is a growing need to confirm in advance whether the food being ordered is safe, especially for households with babies or pets. However, current food delivery systems lack a function to quickly and easily determine whether the ordered food is suitable for babies or pets. This requires users to spend time researching various information themselves, and the reliability of that information cannot be guaranteed. Therefore, the present invention aims to provide a food delivery system that quickly and reliably determines the safety of food and provides it to users.

[0555] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0556] In this invention, the server includes: means for a user to take an image of food using a camera on a terminal and send the image to the server; means for the server to analyze the received image and identify the type of food using an image recognition algorithm; means for a food delivery application for a user to take an image of the food to be ordered and use the image to obtain safety information about the ingredients; means for the server to collect information about the identified food using a chat-based generative AI model and determine whether the food is safe for babies and pets to consume; and means for the server to send the determination result to the user's terminal and display the result on the terminal. This allows a user to use the food delivery application to check the effects of the food to be ordered on babies and pets and ensure its safety.

[0557] "Terminal" means a portable electronic device for inputting, processing, and outputting information.

[0558] A "camera" is a device for taking still images and videos, and is built into a terminal.

[0559] A "server" is a computer system that stores, processes, and distributes information over a network.

[0560] An "image recognition algorithm" is a computational method for analyzing image data and identifying specific objects or features.

[0561] "Food type" is information that indicates the specific classification or name of the food.

[0562] A "chat-type generative AI model" is an artificial intelligence model that answers questions and generates information based on text data.

[0563] A "food delivery application" is software that allows users to order and have food delivered online.

[0564] This invention is a system applied to a food delivery application that allows users to check the safety of food they plan to order. It consists of a terminal used by the user, a server, an image recognition algorithm, a generative AI model, and a food delivery application.

[0565] System configuration

[0566] 1. Terminal

[0567] The user uses a device such as a smartphone or tablet. The device has a built-in camera and is used to take pictures of food. The captured image data is temporarily stored on the device and then sent to the server.

[0568] 2. Server

[0569] The server has multiple functions. It receives and stores image data, analyzes it using an image recognition algorithm to identify the type of food, and uses a generative AI model to collect safety information about the identified food. It then sends the collected information to the user's device and displays it in an appropriate format.

[0570] 3. Image Recognition Algorithm

[0571] The server uses TensorFlow and PyTorch as image recognition algorithms to automatically identify the type of food from the captured image. By using a deep learning model, food can be identified with high accuracy.

[0572] 4. Generative AI Models

[0573] The generative AI model uses a chat-type AI model such as GPT-4. Once the type of food is identified, specific prompts such as "Is it safe for babies to eat XX?" or "Is it safe for dogs to eat XX?" are sent to the generative AI model to obtain relevant safety information.

[0574] 5. Food delivery applications

[0575] The food delivery application provides an interface for users to check the safety of the food they plan to order. The safety information sent from the server is displayed on the user's device, allowing them to easily check whether the food is suitable for babies and pets before ordering.

[0576] Processing flow

[0577] 1. Image capture

[0578] The user uses the smartphone camera to take a photo of the food they plan to order.

[0579] 2. Sending images

[0580] The device sends the captured image to the server, which receives and stores the image data.

[0581] 3. Image Analysis

[0582] The server uses image recognition algorithms to analyze the captured image and identify the type of food, such as "pizza," "spaghetti," or "carbonara."

[0583] 4. Safety assessment

[0584] Based on the type of food identified, prompts such as "Is it safe for babies to eat XX?" or "Is it safe for dogs to eat XX?" are sent to the generative AI model, and safety information is collected from the generative AI model.

[0585] 5. Display results

[0586] The server analyzes the collected data and sends it to the device in an easy-to-understand format. The results are displayed on the user's device, allowing them to see the effects on their baby or pet.

[0587] Specific examples

[0588] Example 1: Pizza for babies

[0589] Prompt: "Is it safe for babies to eat pizza?"

[0590] The user takes a photo of the pizza and sends it to the system.

[0591] The server uses an image recognition algorithm to identify the pizza and send a prompt to the generative AI model.

[0592] Based on information from the AI ​​model, a decision is made as to whether the pizza is suitable.

[0593] Example 2: Carbonara for dogs

[0594] Prompt: "Is it safe for dogs to eat carbonara?"

[0595] The user takes a photo of the carbonara and sends it to the system.

[0596] The server uses an image recognition algorithm to identify carbonara and send a prompt to the generative AI model.

[0597] Based on information from the AI ​​model, the results will show whether carbonara is harmful to dogs.

[0598] In this way, users can quickly and easily check whether the food they plan to order is suitable for babies or pets.

[0599] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0600] Step 1:

[0601] Input: The user takes a photo of the food they plan to order.

[0602] How it works: The user activates the device's camera, takes a picture of the food they plan to order (e.g., pizza, carbonara), and temporarily saves it to the device's local storage.

[0603] Step 2:

[0604] Input: Captured image data.

[0605] Operation: The device checks for network connectivity and sends the stored image data to the server.

[0606] Output: Image data is sent to the server.

[0607] Step 3:

[0608] Input: Image data received by the server.

[0609] How it works: The server temporarily stores the image data it receives, then applies image recognition algorithms (using TensorFlow or PyTorch) to identify the type of food in the image.

[0610] Output: The type of food is identified (e.g. "pizza" or "carbonara").

[0611] Step 4:

[0612] Input: The type of food identified.

[0613] How it works: The server generates queries for the generative AI model based on the identified food types, creating specific questions like "Is it safe for babies to eat pizza?" or "Is it safe for dogs to eat carbonara?"

[0614] Output: The generated query.

[0615] Step 5:

[0616] Input: The generated query.

[0617] How it works: The server sends the generated queries to the generative AI model, which collects comprehensive information from the internet and databases. The generative AI model then generates answers to the questions and provides safety information.

[0618] Output: Response (safety information) from the generative AI model.

[0619] Step 6:

[0620] Input: The response from the generative AI model.

[0621] How it works: The server analyzes the collected information and formats it in a way that is easy for the user to understand. For example, it will give a judgment result such as "Pizza is suitable for babies, but it is recommended to cut it into small pieces."

[0622] Output: The result of the test for the user.

[0623] Step 7:

[0624] Input: Judgment result.

[0625] Operation: The server sends the result of the judgment to the user's device. The device analyzes the received result and displays it on the interface of the food delivery application.

[0626] Output: The result displayed on the device (e.g., "Pizza is suitable for babies, but it is recommended to cut it into small pieces").

[0627] In this way, users can easily check whether the food they plan to order is suitable for babies or pets.

[0628] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0629] This invention relates to a system that supports dietary management for babies and pets (dogs and cats), and in particular, it combines an emotion engine that recognizes the user's emotions. This system provides a process for quickly determining the safety of food.

[0630] Program processing explanation

[0631] 1. Image capture and transmission

[0632] User: Activates the camera on a device such as a smartphone or tablet and takes a picture of food. For example, the user takes a picture of an apple.

[0633] Device: The captured image is temporarily saved in local storage, and after checking the network connection, the image data is sent to the server's API endpoint.

[0634] 2. Image Recognition

[0635] Server: Receives image data sent from the terminal, temporarily stores it, and returns a reception confirmation response to the terminal.

[0636] Server: Analyzes the received image using an image recognition algorithm to identify the type of food in the image. The algorithm uses a deep learning model to learn the characteristics of the ingredients and classify them accordingly.

[0637] Server: Get a specific food type, such as "apple."

[0638] 3. Information Search

[0639] Server: Generates a query to the chat-based AI model based on the identified food type, for example, formulating a specific question such as "Is it safe for babies to eat apples?"

[0640] Server: Sends the generated queries to the chat-based AI model and collects comprehensive information from databases and the Internet.

[0641] Server: Receives the response from the AI ​​model and analyzes its content. For example, it obtains information such as "It is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."

[0642] 4. Emotion recognition

[0643] Device: When a user takes a picture, the emotion engine is activated and analyzes the user's facial expressions and voice. This analysis is performed in real time to understand the user's emotional state.

[0644] Terminal: Transmits the acquired user emotional state data to the server.

[0645] 5. Judgment and result display

[0646] Server: Determines the safety of food based on the obtained information and the user's emotional state. For example, if the result is that it is safe for babies or pets, or if the user's emotional state is deemed unsafe, it provides additional advice or a warning.

[0647] Server: Formats the judgment results and the information that forms the basis for them into a data format.

[0648] Server: The formatted result is sent to the user's device. Once the transmission is confirmed, the server records a log and prepares for the next process.

[0649] Terminal: Receives the result data from the server, checks the integrity of the data, and returns a reception confirmation response to the server.

[0650] Terminal: Analyzes the received data and displays it on the user interface. For example, it might display "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them," and if the user's emotions are anxious, it might display additional advice such as "If you are worried, start with a small amount and monitor the baby's reaction."

[0651] Specific examples

[0652] Example 1: Applying apples and emotion recognition to babies

[0653] User: Take a photo of an apple, save it on the device, and send the image to the server.

[0654] Server: Identifies an apple using an image recognition algorithm. Sends a query to the chat-based AI model asking, "Is it safe for babies to eat apples?"

[0655] Server: Collects information from the AI ​​model and determines whether it is safe for babies to eat apples.

[0656] Device: When the user takes a photo, facial recognition and voice analysis are used to determine the emotional state as anxiety.

[0657] Server: The judgement result was accompanied by additional advice: "If you are concerned, start with a small amount and monitor the condition."

[0658] Device: The result is displayed as follows: "Apples are suitable for babies, but it is recommended that they be cut into small pieces or grated. If you are concerned, start with a small amount and monitor the baby's condition."

[0659] Example 2: Applying chocolate and emotion recognition to dogs

[0660] User: Take a photo of the chocolate and save it on the device. Send the image to the server.

[0661] Server: Identifies chocolate using an image recognition algorithm. Sends a query to the chat-based AI model asking, "Is it safe for dogs to eat chocolate?"

[0662] Server: Collects information from the AI ​​model and determines that chocolate is harmful to dogs.

[0663] Device: When the user takes a photo, facial recognition and voice analysis are used to determine the emotional state as surprise.

[0664] Server: The verdict included a warning that "chocolate is extremely dangerous to dogs and should never be given to them."

[0665] Device: Displays the result: "Chocolate is harmful to dogs and should never be given to them," with an additional warning: "Move out of reach immediately."

[0666] This allows the system to recognize the user's emotions and provide appropriate advice and warnings based on those emotions, making it easy and safe to manage the diet of babies and pets at home.

[0667] The processing flow will be explained below.

[0668] Step 1:

[0669] The user activates the camera on the device and takes an image of food, for example, an apple.

[0670] Step 2:

[0671] The device temporarily saves the captured image in local storage, then checks for network connectivity and sends the image data to the server's API endpoint.

[0672] Step 3:

[0673] The server receives the image data sent from the terminal, stores it temporarily, and returns a reception confirmation response to the terminal.

[0674] Step 4:

[0675] The server analyzes the received image data using an image recognition algorithm that identifies food characteristics and identifies a particular food type.

[0676] Step 5:

[0677] The server obtains the identified food type based on the results of the image recognition algorithm, for example, identifying "apple."

[0678] Step 6:

[0679] The server generates a query to the chat-based AI model based on the type of food identified, for example, "Is it safe for babies to eat apples?"

[0680] Step 7:

[0681] The server generates queries and sends them to a chat-based AI model to collect information about food safety.

[0682] Step 8:

[0683] The server receives the response from the chat-based AI model and analyzes its content, obtaining, for example, information that "it is safe for babies to eat apples, and it is recommended that they be cut into small pieces or grated."

[0684] Step 9:

[0685] The device analyzes the user's facial expressions and voice in real time to identify the user's emotional state. For example, if the user looks anxious or speaks in an anxious voice, it will be determined that the user is anxious.

[0686] Step 10:

[0687] The device transmits the user's emotional state data to the server, which then incorporates this data into the decision-making process.

[0688] Step 11:

[0689] The server determines the safety of food based on the information obtained and the user's emotional state. For example, if the user seems anxious, it may provide additional advice or warnings.

[0690] Step 12:

[0691] The server formats the information, including the judgment result and any additional advice, into a data format and prepares it for presentation to the user.

[0692] Step 13:

[0693] The server then sends the formatted result to the user's device. Once the transmission is confirmed, the server records a log and prepares for the next process.

[0694] Step 14:

[0695] The terminal receives the result data from the server, checks the integrity of the data, and returns a reception confirmation response to the server.

[0696] Step 15:

[0697] The device analyzes the received data and displays it on the user interface. For example, it might say, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them. If you're worried, start with a small amount and monitor your baby's condition."

[0698] This allows the system to recognize the user's emotions and provide appropriate advice and warnings based on those emotions, making it easy and safe to manage the diet of babies and pets at home.

[0699] Example 2

[0700] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0701] Currently, there is a lack of systems that provide appropriate information quickly and accurately for managing the diet of babies and pets. Furthermore, there are no systems that can provide advice or warnings that take into account the user's emotions. This leaves users with insufficient support to provide food to their babies and pets with peace of mind. Furthermore, in existing systems, information collection to determine food safety is often done manually, which is a very time-consuming process.

[0702] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0703] In this invention, the server includes: a means for a user to take an image of food using a camera on the terminal and send the image to the server; a means for the terminal to analyze the user's facial expressions and voice and acquire emotion data; a means for the server to analyze the received image and identify the type of food using an image recognition algorithm; a means for the server to collect information about the identified food using a generative AI model and determine whether the food is safe for babies or pets to ingest; and a means for the server to send additional advice or warnings based on the determination results and emotion data to the user's terminal and display the results on the terminal. This allows the user to provide quick and appropriate information when feeding babies or pets, ensuring safety, and providing advice or warnings according to the user's emotions.

[0704] "Terminal" refers to a smartphone, tablet, or other portable information processing device used by a user.

[0705] "Server" refers to a computer system that receives, stores, analyzes, searches for information, and transmits results of image data.

[0706] "User" refers to a person who uses the system to manage the diet of a baby or pet.

[0707] "Image recognition algorithm" refers to a computational method for analyzing received image data and identifying the type of food depicted in the image.

[0708] A "deep learning model" refers to an artificial intelligence technology that learns and classifies the characteristics of ingredients as part of an image recognition algorithm.

[0709] "Generative AI model" refers to a system that uses chat-based artificial intelligence to collect and generate information about food.

[0710] A "prompt" refers to a query or question that is input to a generative AI model.

[0711] "Emotion engine" refers to technology that analyzes a user's facial expressions and voice to identify their emotional state.

[0712] "Determination result" refers to data containing the evaluation results of the server's evaluation of food safety.

[0713] "Emotion data" refers to data that indicates the user's emotional state obtained through facial recognition or voice analysis.

[0714] "Additional advice or warning" refers to supplementary advice or a message urging caution that is provided to the user based on the judgment result.

[0715] This invention relates to a system that supports dietary management for babies and pets, and in particular, it combines an emotion engine that recognizes the user's emotions. This system provides a process for quickly determining the safety of food.

[0716] A user takes a picture of food using a device such as a smartphone or tablet. For example, the user takes a picture of an apple. The device temporarily saves the captured image in local storage, and after confirming network connectivity, sends the image data to the server's API endpoint.

[0717] The server receives the image data sent from the device and temporarily stores it. It then returns a receipt confirmation response to the device. It then uses an image recognition algorithm to analyze the received image and identify the type of food in the image. For example, it uses a deep learning framework such as TensorFlow or PyTorch. The algorithm learns the characteristics of the ingredients contained in the image and classifies them based on that. The server then obtains the identified type of food, such as "apple."

[0718] The server then generates a query to a generative AI model (such as OpenAI's ChatGPT) based on the identified food type, for example, creating a specific question such as "Is it safe for babies to eat apples?" The generated query is a prompt sentence of the form:

[0719] "Is it safe for babies to eat apples?"

[0720] The server sends this prompt to the generative AI model, collects comprehensive information from the internet and databases, receives the response from the generative AI model, and analyzes its content. For example, the server obtains the information that "It is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."

[0721] When a user takes a picture, the device runs an emotion engine (such as the Affectiva SDK) to analyze the user's facial expressions and voice. This analysis is performed in real time to understand the user's emotional state. The device then transmits the acquired data on the user's emotional state to the server.

[0722] The server determines the safety of the food based on the obtained information and the user's emotional state. For example, even if the result indicates that the food is safe for babies or pets, if the user's emotions are deemed to be uneasy, the server will add additional advice or warnings. The server then formats the result and the information that forms the basis of the result into a data format and sends the formatted result to the user's device. Once the transmission is confirmed, the server records the log and prepares for the next process.

[0723] The device receives the result data sent from the server and checks the integrity of the data. It then returns a reception confirmation response to the server. Finally, the device analyzes the received data and displays it on the user interface. For example, it might display, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them. If you are concerned, start with a small amount and monitor the baby's condition," and if the user's emotions are uneasy, it might display additional advice such as, "If you are concerned, start with a small amount and monitor the baby's condition."

[0724] This system recognizes the user's emotions and provides appropriate advice and warnings based on those emotions, making it easy and safe to manage the diet of babies and pets at home.

[0725] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0726] Step 1:

[0727] A user takes a picture of food using the device's camera. Specifically, the user launches the camera app, places food (e.g., an apple) in the center of the screen, and presses the shutter button. The input is the camera operation and the captured food image, and the output is an image file saved in local storage.

[0728] Step 2:

[0729] The device temporarily saves the image captured in local storage and checks the network connection. Specifically, it saves the image file in the device's local storage and then checks whether the WiFi or mobile data connection is enabled. The input is the image file captured in step 1 and the network status, and the output is the state ready to be sent.

[0730] Step 3:

[0731] The device sends image data to the server's API endpoint. Specifically, the device sends the image file to the server as an HTTP POST request. The input is the saved image file and the server's API endpoint, and the output is the completion status of the transmission to the server.

[0732] Step 4:

[0733] The server receives image data sent from the terminal and temporarily stores it. Specifically, the server extracts image data from the HTTP request it receives and stores it in a temporary directory on the server. The input is the image data sent from the terminal, and the output is the temporarily stored image file and a response confirming receipt.

[0734] Step 5:

[0735] The server uses an image recognition algorithm to analyze the received image and identify the type of food. Specifically, the server analyzes the image using a deep learning model such as TensorFlow or PyTorch to identify the type of food (e.g., "apple"). The input is a temporarily saved image file, and the output is information about the identified type of food (e.g., "apple").

[0736] Step 6:

[0737] The server generates a query to the generative AI model based on the identified food type. Specifically, the server generates a prompt sentence, "Is it safe for babies to eat apples?" and sends it to the generative AI model. The input is food type information (e.g., "apple"), and the output is the generated query. An example of a prompt sentence: "Is it safe for babies to eat apples?"

[0738] Step 7:

[0739] The server sends the generated query to the generative AI model and collects information from databases and the Internet. Specifically, the server sends a prompt to the generative AI model as an HTTP request and receives food safety information in response. The input is the generated query, and the output is the received safety information.

[0740] Step 8:

[0741] The server analyzes the response from the AI ​​model and determines the safety of the food. Specifically, it analyzes the received safety information and obtains a judgment result such as "It is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated." The input is the received safety information, and the output is the judgment result.

[0742] Step 9:

[0743] When a user takes a picture, the device uses an emotion engine to analyze the user's facial expressions and voice to identify their emotional state. Specifically, the device's front camera and microphone are used to collect facial expressions and voice in real time, which are then analyzed by the emotion engine. The input is the user's facial expression and voice data, and the output is the identified emotional state data (e.g., "anxiety").

[0744] Step 10:

[0745] The emotional state data acquired by the device is sent to the server. Specifically, the emotional state data is converted into JSON format and sent to the server as an HTTP request. The input is the emotional state data, and the output is the status of completion of transmission to the server.

[0746] Step 11:

[0747] The server determines the safety of the food based on the information and emotional state obtained, and generates additional advice or a warning. Specifically, it adds additional advice to the judgment result, such as "If you are concerned, start with a small amount and monitor the situation." The input is the judgment result and emotional state data, and the output is the final judgment result and additional advice.

[0748] Step 12:

[0749] The server sends the final judgment result and additional advice to the user's device. Specifically, the server formats the data in JSON format and sends it to the user's device as an HTTP request. The input is the final judgment result and additional advice, and the output is the status of completion of transmission to the device.

[0750] Step 13:

[0751] The terminal receives the result data from the server and checks the integrity of the data. Specifically, it checks the received data using a checksum or other method and returns a response confirming receipt to the server. The input is the result data sent from the server, and the output is the response confirming receipt.

[0752] Step 14:

[0753] The device analyzes the result data and displays it on the user interface. Specifically, it displays the results and advice in the application's UI component, informing the user, for example, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them. If you are concerned, start with a small amount and monitor the baby's condition." The input is the analyzed result data, and the output is the display on the user interface.

[0754] (Application example 2)

[0755] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0756] Avoiding the risk of minors and animals accidentally ingesting harmful foods is an important issue in home dietary management. Users often have concerns or doubts about food safety, and appropriate information and advice that take these feelings into account is needed. However, existing systems rarely offer the functionality to integrate emotion recognition and food safety assessment in real time. Therefore, a system that allows users to manage their diet with peace of mind is needed.

[0757] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for a user to take an image of a food using a camera of a terminal and send the image to an information processing device; means for the information processing device to analyze the received image and identify the type of food using an image recognition algorithm; means for the information processing device to collect information about the identified food using an interactive AI model and determine whether the food is safe for minors or animals to consume; means for the information processing device to send the determination result to the user's terminal and display the result on the terminal; means for the terminal to analyze emotions from the user's facial expressions and voice and send the emotions to the information processing device; and means for the information processing device to add additional advice or warnings based on the obtained emotion data. This makes it possible to determine the safety of food in real time and provide appropriate information and advice according to the user's emotions.

[0758] A "terminal" is a portable information and communication device operated by a user, and is equipped with a camera and a display.

[0759] An "information processing device" is a computer system that analyzes received data and processes it using a specific algorithm.

[0760] An "image recognition algorithm" is a computational method for analyzing image data to identify specific objects or features.

[0761] An "interactive AI model" is a program that uses artificial intelligence to answer questions and search for information in natural language.

[0762] A "minor" is a child or young person who is not legally recognized as an adult.

[0763] "Animals" are living creatures such as dogs and cats kept as pets in homes.

[0764] "Emotion data" is information that indicates the emotional state of the user analyzed from facial expressions and voice.

[0765] "Additional advice and warnings" are supplemental information provided to help users provide food safely and with peace of mind.

[0766] System Program Overview

[0767] The system that realizes this application example consists of the following major components:

[0768] 1. Hardware

[0769] Terminal: A portable information and communication device (smartphone, tablet, etc.) operated by a user, equipped with a camera and display.

[0770] Information processing device: A server system that analyzes received data and processes it using specific algorithms. It is desirable for this server to be equipped with a high-performance CPU and GPU.

[0771] 2. Software

[0772] Image recognition algorithm: Built using deep learning libraries such as TensorFlow.

[0773] Conversational AI model: An engine used for natural language processing and question answering.

[0774] Emotion recognition software: Programs for facial expression recognition and speech analysis (OpenCV and other facial expression analysis tools).

[0775] Program processing overview

[0776] 1. Image capture and transmission

[0777] The user takes a photo of the food using the device's camera.

[0778] The terminal transmits the captured image to the information processing device.

[0779] 2. Image Recognition

[0780] The server analyzes the received images and uses image recognition algorithms to identify the type of food.

[0781] It uses a deep learning model to analyze the features of food in an image and identify its type.

[0782] 3. Information Search

[0783] The server sends information about the identified food as a query to the conversational AI model.

[0784] The AI ​​model collects the necessary information from databases and the internet to determine whether the food is safe for minors and animals.

[0785] 4. Emotion recognition

[0786] When a user takes a picture, the device analyzes emotions in real time from facial expressions and voice.

[0787] The analyzed emotion data is transmitted to an information processing device.

[0788] 5. Judgment and result display

[0789] The server generates additional advice and warnings based on the food safety assessment results and emotion data, and sends them to the user's device.

[0790] The terminal displays the received results on a user interface.

[0791] Specific examples

[0792] Example 1: Judging whether an apple is suitable for babies

[0793] 1. A user takes a photo of an apple and asks the camera, "Are apples safe for babies?"

[0794] 2. The device sends the emotional data analyzed as "anxiety" to the server.

[0795] 3. The server generates advice such as, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them. If you are concerned, start with a small amount and monitor their condition." and provides it to the user.

[0796] Prompt Sentence Examples

[0797] Take a photo of the food you are about to give your baby. Then ask, "Is an apple safe for my baby?" Then, enter your emotion (e.g., "anxious" or "surprised").

[0798] Specific examples of the technologies used

[0799] Image recognition algorithm: TensorFlow

[0800] Conversational AI model: Natural language processing engine

[0801] Emotion recognition software: OpenCV

[0802] These technologies allow users to confidently verify food safety and provide appropriate diets for minors and animals.

[0803] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0804] Step 1:

[0805] The user takes a picture of food using the device's camera and temporarily saves the captured image data in local storage. This is done using a camera application or the device's camera function. The input is the captured image data, and the output is an image file saved in the device's local storage.

[0806] Step 2:

[0807] The device checks the network connection and sends the captured image data to the information processing device. The transmission is performed using HTTP or other communication protocols. The input is an image file saved in local storage, and the output is the image data sent to the server.

[0808] Step 3:

[0809] The server temporarily stores the received image data and begins analysis using an image recognition algorithm. The type of food is identified using a deep learning library such as TensorFlow. The input is the image data stored on the server, and the output is the type of food identified through analysis.

[0810] Step 4:

[0811] The server sends a query to the conversational AI model based on the identified food type to collect safety information about the food. For example, a query such as "Is it safe for babies to eat apples?" is generated and the AI ​​model searches for information. The input is the identified food type information, and the output is information about the food obtained from the AI ​​model.

[0812] Step 5:

[0813] The device analyzes emotions from the user's facial expressions and voice when taking an image. It uses OpenCV and other facial expression analysis tools to obtain the user's emotional data in real time. The input is the user's facial expression and voice data, and the output is the emotional data obtained through the analysis.

[0814] Step 6:

[0815] The device sends the analyzed emotional data to the server using a communication method such as the HTTP protocol. The input is the emotional data analyzed on the device, and the output is the emotional data sent to the server.

[0816] Step 7:

[0817] The server generates additional advice and warnings based on food safety information and emotion data. Based on the obtained data, it generates appropriate messages and adds advice to alleviate the user's anxiety. The input is food safety information and emotion data, and the output is the generated advice or warning message.

[0818] Step 8:

[0819] The server generates advice and warning messages and sends them to the user's terminal. The input is the generated message and the output is the message sent to the user's terminal.

[0820] Step 9:

[0821] The terminal analyzes the received result data and displays it on the user interface. Based on this information, the user can decide whether to provide food to minors or animals. The input is the received message data, and the output is the information displayed on the user interface.

[0822] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0823] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0824] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0825] [Third embodiment]

[0826] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0827] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0828] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0829] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0830] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0831] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0832] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0833] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0834] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0835] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0836] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0837] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0838] This invention relates to a system for supporting dietary management for babies and pets (dogs and cats). This system works by having the user take a picture of food using the camera on their device and send it to a server. The server analyzes the received image and identifies the type of food using an image recognition algorithm. Next, it uses a chat-based AI model to collect safety information about the identified food and, based on that information, determines whether the food is safe for babies and pets to consume. The result of the determination is sent to the user's device and displayed on the device.

[0839] Program processing explanation

[0840] 1. Image capture and transmission

[0841] User: Activates the camera on a device such as a smartphone or tablet and takes a picture of food. For example, the user takes a picture of an apple.

[0842] Device: The captured image is temporarily saved in local storage. After that, the network connection is confirmed and the captured image is sent to the server.

[0843] 2. Image Recognition

[0844] Server: Receives image data sent from the device and temporarily stores it.

[0845] Server: Analyzes the received image using an image recognition algorithm to identify the type of food in the image. The algorithm uses a deep learning model to learn the characteristics of the ingredients and classify them accordingly.

[0846] Server: Get a specific food type, such as "apple."

[0847] 3. Information Search

[0848] Server: Generates queries to the chat-based AI model based on the identified food types, for example, formulating specific questions such as "Is it safe for babies to eat apples?"

[0849] Server: Sends the generated queries to the chat-based AI model and collects comprehensive information from databases and the Internet.

[0850] Server: Analyzes the information returned as a response from the AI ​​model and determines the safety of food based on that information. For example, it may determine that "it is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."

[0851] 4. Results display

[0852] Server: Formats the data so that the obtained information and judgment results are presented to the user in an easy-to-understand manner.

[0853] Server: Sends the judgment result to the user's terminal.

[0854] Terminal: Analyzes the received results and displays them on the user interface. For example, it displays advice such as "Apples are suitable for babies, but it is recommended to cut them into small pieces or grate them."

[0855] Specific examples

[0856] Example 1: An apple for a baby

[0857] User: Take a photo of an apple, save it on the device, and send the image to the server.

[0858] Server: Identifies an apple using an image recognition algorithm. Sends a query to the chat-based AI model asking, "Is it safe for babies to eat apples?"

[0859] Server: Collects information from the AI ​​model and determines whether it is safe for babies to eat apples. Sends the result to the user's device.

[0860] Device: Displays the result: "Apples are suitable for babies, but it is recommended that they be cut into small pieces or grated."

[0861] Example 2: Chocolate for dogs

[0862] User: Take a photo of the chocolate and save it on the device. Send the image to the server.

[0863] Server: Identifies chocolate using an image recognition algorithm. Sends a query to the chat-based AI model asking, "Is it safe for dogs to eat chocolate?"

[0864] Server: Collects information from the AI ​​model and determines whether chocolate is harmful to dogs. Sends the result to the user's device.

[0865] Device: Displays the result: "Chocolate is harmful to dogs and should not be given to them."

[0866] This allows you to quickly and easily determine what foods your baby or pets can and cannot eat in your home.

[0867] The processing flow will be explained below.

[0868] Step 1:

[0869] The user activates the device's camera and takes an image of food, for example, an apple.

[0870] Step 2:

[0871] The device temporarily saves the captured image in local storage, then checks for network connectivity and sends the image data to the server's API endpoint.

[0872] Step 3:

[0873] The server receives the image data sent from the terminal, stores it temporarily, and returns a reception confirmation response to the terminal.

[0874] Step 4:

[0875] The server analyzes the received image data using an image recognition algorithm. Specifically, the image is input into a deep learning model for image recognition to identify the type of food.

[0876] Step 5:

[0877] The server obtains the identified food type (e.g., "apple") based on the results of the image recognition algorithm.

[0878] Step 6:

[0879] The server generates a query to the chat-based AI model based on the identified food type, for example, creating a specific question such as "Is it safe for babies to eat apples?"

[0880] Step 7:

[0881] The server generates queries and sends them to a chat-based AI model, which collects comprehensive information from databases and the Internet.

[0882] Step 8:

[0883] The server receives the response from the chat-based AI model and analyzes its content, obtaining, for example, information that "it is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."

[0884] Step 9:

[0885] The server determines the safety of the food based on the information obtained, and then formats the determination results and the information that supports them into a data format.

[0886] Step 10:

[0887] The server then sends the formatted result to the user's device. Once the transmission is confirmed, the server records a log and prepares for the next process.

[0888] Step 11:

[0889] The terminal receives the result data from the server, checks the integrity of the data, and returns a reception confirmation response to the server.

[0890] Step 12:

[0891] The device analyzes the received data and displays it on the user interface, for example, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them."

[0892] Through these steps, users can easily and quickly determine what foods babies and pets in the home can and cannot eat.

[0893] Example 1

[0894] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0895] In modern society, dietary management for babies and animals is an important issue, and there is a need for a system that can quickly and accurately obtain information for providing appropriate food. However, manually collecting and assessing information requires time and effort, and there is a risk of making decisions based on incorrect information. Conventional systems have had difficulty efficiently providing information needed to accurately determine food safety. The present invention aims to solve these problems and provide a system that supports dietary management for babies and animals.

[0896] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0897] In this invention, the server includes a means for a user to take an image of a food using a camera on a terminal and send the image to the server, a means for the server to analyze the received image and identify the type of food using an image recognition algorithm, and a means for the server to collect information about the identified food using a generative AI model and determine whether the food is safe for babies or animals to consume, thereby enabling users to easily identify foods suitable for babies and animals and confirm their safety.

[0898] A "user" is a person who uses the system to take pictures of food and transmits the pictures to the server.

[0899] A "terminal" refers to a computer device used by a user, such as a smartphone or tablet, that has a camera function and a network connection function.

[0900] A "camera" is an image capturing device built into a terminal.

[0901] "Food" refers to all food and drink consumed by babies and animals.

[0902] A "server" is a computer system that processes data submitted by users and performs specific functions.

[0903] An "image recognition algorithm" is a computational method for analyzing a received image and identifying objects within the image.

[0904] A "deep learning model" is a type of machine learning that uses a multi-layered neural network to learn patterns from large amounts of data and is used to analyze images, audio, and other data.

[0905] A "generative AI model" is an algorithm that generates a response in natural language in response to an input prompt.

[0906] A "prompt sentence" is a sentence entered to ask a generative AI model for specific information.

[0907] "Baby" refers to an infant or young child, especially a child within the first year of life.

[0908] "Animals" refers to pets kept at home, specifically mammals such as dogs and cats.

[0909] "Ingestion" refers to the act of taking food into the mouth and digesting and absorbing it.

[0910] "Displaying on the terminal" refers to making the results visible to the user in the form of text or graphics on the terminal screen.

[0911] The present invention relates to a system that supports the dietary management of babies and animals, and operates by allowing a user to take an image of food using a camera on a terminal and send the image to a server. The system operates as follows.

[0912] First, a user uses a device such as a smartphone or tablet to activate the camera and take an image of food. For example, the user takes a photo of an apple. The device temporarily saves the captured image in local storage and then checks for an Internet connection. Once it confirms that a network connection has been established, it sends the captured image to the server.

[0913] The server receives the image data sent from the device and temporarily stores it. It then uses an image recognition algorithm to analyze the received image. This algorithm uses deep learning models such as TensorFlow and PyTorch to learn the characteristics of food and classify it based on that. The server then identifies the type of food in the image and obtains the identified food type, such as "apple."

[0914] Next, the server generates a query to a generative AI model based on the identified food type. Specifically, it creates a question such as, "Is it safe for babies to eat apples?" The generative AI model used is OpenAI's GPT-4. The server sends the generated query to the generative AI model, which collects comprehensive information from databases and the internet.

[0915] The server analyzes the information returned as a response from the generative AI model and uses it to determine the safety of food. For example, it may determine that "apples are safe for babies to eat, but it is recommended that they be cut into small pieces or grated." The result is then formatted and converted into a user-friendly format.

[0916] Finally, the server sends the result of the judgment to the user's device, which then analyzes the result and displays it on the user interface. For example, it might say, "Apples are suitable for babies, but it is recommended that you cut them into small pieces or grate them."

[0917] Specific examples

[0918] Apples for babies

[0919] A user takes a photo of an apple, saves it on their device, and sends it to the server.

[0920] The server uses an image recognition algorithm to identify the apple and sends the generative AI model a query: "Is it safe for a baby to eat the apple?"

[0921] The server collects information from the AI ​​model, determines whether it is safe for babies to eat apples, and sends the result to the user's device.

[0922] The device will display the result: "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them."

[0923] Chocolate for animals

[0924] The user takes a photo of the chocolate, saves it on the device, and sends it to the server.

[0925] The server uses an image recognition algorithm to identify the chocolate and sends the generative AI model a query: "Is chocolate safe for dogs to eat?"

[0926] The server collects information from the AI ​​model, determines whether chocolate is harmful to dogs, and sends the result to the user's device.

[0927] The device will display the result: "Chocolate is harmful to dogs and should not be given to them."

[0928] This makes it possible to quickly and easily determine what foods babies and animals in the home can and cannot eat.

[0929] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0930] Step 1:

[0931] A user activates the camera on a device such as a smartphone or tablet and takes a picture of food. The input is the food to be photographed (e.g., an apple), and the output is an image file saved on the device (e.g., IMG_20230101.jpg). Specifically, the user opens the camera app, focuses on the apple, and presses the shutter button.

[0932] Step 2:

[0933] The device temporarily saves the captured image in local storage. It then checks the Internet connection. The input is the captured image file, and the output is the temporarily saved image file and the result of checking the network connection status. Specifically, it saves the image file in a specific folder and checks the Wi-Fi and mobile data connection status.

[0934] Step 3:

[0935] Once the device has confirmed a network connection, it uses the REST API to send image data to the server. The input is the temporarily saved image file and the network connection status, and the output is the image data sent to the server. Specifically, it creates an HTTP POST request, encodes the image file, and sends it to the specified URL on the server.

[0936] Step 4:

[0937] The server receives image data sent from the terminal and temporarily stores it. The input is the image data sent from the terminal, and the output is the image data temporarily stored in the server. Specifically, the received image file is stored in the server's temporary directory (e.g., / tmp).

[0938] Step 5:

[0939] The server analyzes the received image using an image recognition algorithm. The input is the temporarily stored image data, and the output is the identified type of food (e.g., apple). Specifically, the image file is input into a deep learning model (e.g., TensorFlow), which extracts and classifies the food's features.

[0940] Step 6:

[0941] The server generates a query to the generative AI model based on the identified food type. The input is the identified food type (e.g., apple), and the output is a generated prompt (e.g., "Is it safe for babies to eat apples?"). Specifically, the server uses the identified food type to create a question-style prompt.

[0942] Step 7:

[0943] The server sends the generated query to the generative AI model and collects information from databases and the Internet. The input is the generated prompt, and the output is the information returned by the generative AI model (e.g., "It is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."). Specific operations include creating an API request and sending the prompt to the generative AI model.

[0944] Step 8:

[0945] The server analyzes the response from the generative AI model and determines the safety of the food. The input is the information returned from the generative AI model, and the output is the safety judgment result (e.g., "safe"). Specifically, it extracts information about safety from the response text and determines the judgment result.

[0946] Step 9:

[0947] The server formats the resulting judgment results in a data format that is easy to understand for the user. The input is the safety judgment result, and the output is the formatted data (e.g., "Apples are suitable for babies, but it is recommended that they be cut into small pieces or grated"). Specific operations include converting the data into a format such as JSON.

[0948] Step 10:

[0949] The server sends the result of the judgment to the user's device. The input is the formatted data, and the output is the data received by the user's device. The specific operation is to send the data as an API response.

[0950] Step 11:

[0951] The device analyzes the received judgment result and displays it on the user interface. The input is the received judgment result data, and the output is the judgment result displayed to the user (e.g., "Apples are suitable for babies, but it is recommended that they be cut into small pieces or grated"). Specific actions include displaying a pop-up message or notification.

[0952] (Application example 1)

[0953] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0954] In recent years, the number of users of online food delivery services has increased, and there is a growing need to confirm in advance whether the food being ordered is safe, especially for households with babies or pets. However, current food delivery systems lack a function to quickly and easily determine whether the ordered food is suitable for babies or pets. This requires users to spend time researching various information themselves, and the reliability of that information cannot be guaranteed. Therefore, the present invention aims to provide a food delivery system that quickly and reliably determines the safety of food and provides it to users.

[0955] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0956] In this invention, the server includes: means for a user to take an image of food using a camera on a terminal and send the image to the server; means for the server to analyze the received image and identify the type of food using an image recognition algorithm; means for a food delivery application for a user to take an image of the food to be ordered and use the image to obtain safety information about the ingredients; means for the server to collect information about the identified food using a chat-based generative AI model and determine whether the food is safe for babies and pets to consume; and means for the server to send the determination result to the user's terminal and display the result on the terminal. This allows a user to use the food delivery application to check the effects of the food to be ordered on babies and pets and ensure its safety.

[0957] "Terminal" means a portable electronic device for inputting, processing, and outputting information.

[0958] A "camera" is a device for taking still images and videos, and is built into a terminal.

[0959] A "server" is a computer system that stores, processes, and distributes information over a network.

[0960] An "image recognition algorithm" is a computational method for analyzing image data and identifying specific objects or features.

[0961] "Food type" is information that indicates the specific classification or name of the food.

[0962] A "chat-type generative AI model" is an artificial intelligence model that answers questions and generates information based on text data.

[0963] A "food delivery application" is software that allows users to order and have food delivered online.

[0964] This invention is a system applied to a food delivery application that allows users to check the safety of food they plan to order. It consists of a terminal used by the user, a server, an image recognition algorithm, a generative AI model, and a food delivery application.

[0965] System configuration

[0966] 1. Terminal

[0967] The user uses a device such as a smartphone or tablet. The device has a built-in camera and is used to take pictures of food. The captured image data is temporarily stored on the device and then sent to the server.

[0968] 2. Server

[0969] The server has multiple functions. It receives and stores image data, analyzes it using an image recognition algorithm to identify the type of food, and uses a generative AI model to collect safety information about the identified food. It then sends the collected information to the user's device and displays it in an appropriate format.

[0970] 3. Image Recognition Algorithm

[0971] The server uses TensorFlow and PyTorch as image recognition algorithms to automatically identify the type of food from the captured image. By using a deep learning model, food can be identified with high accuracy.

[0972] 4. Generative AI Models

[0973] The generative AI model uses a chat-type AI model such as GPT-4. Once the type of food is identified, specific prompts such as "Is it safe for babies to eat XX?" or "Is it safe for dogs to eat XX?" are sent to the generative AI model to obtain relevant safety information.

[0974] 5. Food delivery applications

[0975] The food delivery application provides an interface for users to check the safety of the food they plan to order. The safety information sent from the server is displayed on the user's device, allowing them to easily check whether the food is suitable for babies and pets before ordering.

[0976] Processing flow

[0977] 1. Image capture

[0978] The user uses the smartphone camera to take a photo of the food they plan to order.

[0979] 2. Sending images

[0980] The device sends the captured image to the server, which receives and stores the image data.

[0981] 3. Image Analysis

[0982] The server uses image recognition algorithms to analyze the captured image and identify the type of food, such as "pizza," "spaghetti," or "carbonara."

[0983] 4. Safety assessment

[0984] Based on the type of food identified, prompts such as "Is it safe for babies to eat XX?" or "Is it safe for dogs to eat XX?" are sent to the generative AI model, and safety information is collected from the generative AI model.

[0985] 5. Display results

[0986] The server analyzes the collected data and sends it to the device in an easy-to-understand format. The results are displayed on the user's device, allowing them to see the effects on their baby or pet.

[0987] Specific examples

[0988] Example 1: Pizza for babies

[0989] Prompt: "Is it safe for babies to eat pizza?"

[0990] The user takes a photo of the pizza and sends it to the system.

[0991] The server uses an image recognition algorithm to identify the pizza and send a prompt to the generative AI model.

[0992] Based on information from the AI ​​model, a decision is made as to whether the pizza is suitable.

[0993] Example 2: Carbonara for dogs

[0994] Prompt: "Is it safe for dogs to eat carbonara?"

[0995] The user takes a photo of the carbonara and sends it to the system.

[0996] The server uses an image recognition algorithm to identify carbonara and send a prompt to the generative AI model.

[0997] Based on information from the AI ​​model, the results will show whether carbonara is harmful to dogs.

[0998] In this way, users can quickly and easily check whether the food they plan to order is suitable for babies or pets.

[0999] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1000] Step 1:

[1001] Input: The user takes a photo of the food they plan to order.

[1002] How it works: The user activates the device's camera, takes a picture of the food they plan to order (e.g., pizza, carbonara), and temporarily saves it to the device's local storage.

[1003] Step 2:

[1004] Input: Captured image data.

[1005] Operation: The device checks for network connectivity and sends the stored image data to the server.

[1006] Output: Image data is sent to the server.

[1007] Step 3:

[1008] Input: Image data received by the server.

[1009] How it works: The server temporarily stores the image data it receives, then applies image recognition algorithms (using TensorFlow or PyTorch) to identify the type of food in the image.

[1010] Output: The type of food is identified (e.g. "pizza" or "carbonara").

[1011] Step 4:

[1012] Input: The type of food identified.

[1013] How it works: The server generates queries for the generative AI model based on the identified food types, creating specific questions like "Is it safe for babies to eat pizza?" or "Is it safe for dogs to eat carbonara?"

[1014] Output: The generated query.

[1015] Step 5:

[1016] Input: The generated query.

[1017] How it works: The server sends the generated queries to the generative AI model, which collects comprehensive information from the internet and databases. The generative AI model then generates answers to the questions and provides safety information.

[1018] Output: Response (safety information) from the generative AI model.

[1019] Step 6:

[1020] Input: The response from the generative AI model.

[1021] How it works: The server analyzes the collected information and formats it in a way that is easy for the user to understand. For example, it will give a judgment result such as "Pizza is suitable for babies, but it is recommended to cut it into small pieces."

[1022] Output: The result of the test for the user.

[1023] Step 7:

[1024] Input: Judgment result.

[1025] Operation: The server sends the result of the judgment to the user's device. The device analyzes the received result and displays it on the interface of the food delivery application.

[1026] Output: The result displayed on the device (e.g., "Pizza is suitable for babies, but it is recommended to cut it into small pieces").

[1027] In this way, users can easily check whether the food they plan to order is suitable for babies or pets.

[1028] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1029] This invention relates to a system that supports dietary management for babies and pets (dogs and cats), and in particular, it combines an emotion engine that recognizes the user's emotions. This system provides a process for quickly determining the safety of food.

[1030] Program processing explanation

[1031] 1. Image capture and transmission

[1032] User: Activates the camera on a device such as a smartphone or tablet and takes a picture of food. For example, the user takes a picture of an apple.

[1033] Device: The captured image is temporarily saved in local storage, and after checking the network connection, the image data is sent to the server's API endpoint.

[1034] 2. Image Recognition

[1035] Server: Receives image data sent from the terminal, temporarily stores it, and returns a reception confirmation response to the terminal.

[1036] Server: Analyzes the received image using an image recognition algorithm to identify the type of food in the image. The algorithm uses a deep learning model to learn the characteristics of the ingredients and classify them accordingly.

[1037] Server: Get a specific food type, such as "apple."

[1038] 3. Information Search

[1039] Server: Generates a query to the chat-based AI model based on the identified food type, for example, formulating a specific question such as "Is it safe for babies to eat apples?"

[1040] Server: Sends the generated queries to the chat-based AI model and collects comprehensive information from databases and the Internet.

[1041] Server: Receives the response from the AI ​​model and analyzes its content. For example, it obtains information such as "It is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."

[1042] 4. Emotion recognition

[1043] Device: When a user takes a picture, the emotion engine is activated and analyzes the user's facial expressions and voice. This analysis is performed in real time to understand the user's emotional state.

[1044] Terminal: Transmits the acquired user emotional state data to the server.

[1045] 5. Judgment and result display

[1046] Server: Determines the safety of food based on the obtained information and the user's emotional state. For example, if the result is that it is safe for babies or pets, or if the user's emotional state is deemed unsafe, it provides additional advice or a warning.

[1047] Server: Formats the judgment results and the information that forms the basis for them into a data format.

[1048] Server: The formatted result is sent to the user's device. Once the transmission is confirmed, the server records a log and prepares for the next process.

[1049] Terminal: Receives the result data from the server, checks the integrity of the data, and returns a reception confirmation response to the server.

[1050] Terminal: Analyzes the received data and displays it on the user interface. For example, it might display "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them," and if the user's emotions are anxious, it might display additional advice such as "If you are worried, start with a small amount and monitor the baby's reaction."

[1051] Specific examples

[1052] Example 1: Applying apples and emotion recognition to babies

[1053] User: Take a photo of an apple, save it on the device, and send the image to the server.

[1054] Server: Identifies an apple using an image recognition algorithm. Sends a query to the chat-based AI model asking, "Is it safe for babies to eat apples?"

[1055] Server: Collects information from the AI ​​model and determines whether it is safe for babies to eat apples.

[1056] Device: When the user takes a photo, facial recognition and voice analysis are used to determine the emotional state as anxiety.

[1057] Server: The judgement result was accompanied by additional advice: "If you are concerned, start with a small amount and monitor the condition."

[1058] Device: The result is displayed as follows: "Apples are suitable for babies, but it is recommended that they be cut into small pieces or grated. If you are concerned, start with a small amount and monitor the baby's condition."

[1059] Example 2: Applying chocolate and emotion recognition to dogs

[1060] User: Take a photo of the chocolate and save it on the device. Send the image to the server.

[1061] Server: Identifies chocolate using an image recognition algorithm. Sends a query to the chat-based AI model asking, "Is it safe for dogs to eat chocolate?"

[1062] Server: Collects information from the AI ​​model and determines that chocolate is harmful to dogs.

[1063] Device: When the user takes a photo, facial recognition and voice analysis are used to determine the emotional state as surprise.

[1064] Server: The verdict included a warning that "chocolate is extremely dangerous to dogs and should never be given to them."

[1065] Device: Displays the result: "Chocolate is harmful to dogs and should never be given to them," with an additional warning: "Move out of reach immediately."

[1066] This allows the system to recognize the user's emotions and provide appropriate advice and warnings based on those emotions, making it easy and safe to manage the diet of babies and pets at home.

[1067] The processing flow will be explained below.

[1068] Step 1:

[1069] The user activates the camera on the device and takes an image of food, for example, an apple.

[1070] Step 2:

[1071] The device temporarily saves the captured image in local storage, then checks for network connectivity and sends the image data to the server's API endpoint.

[1072] Step 3:

[1073] The server receives the image data sent from the terminal, stores it temporarily, and returns a reception confirmation response to the terminal.

[1074] Step 4:

[1075] The server analyzes the received image data using an image recognition algorithm that identifies food characteristics and identifies a particular food type.

[1076] Step 5:

[1077] The server obtains the identified food type based on the results of the image recognition algorithm, for example, identifying "apple."

[1078] Step 6:

[1079] The server generates a query to the chat-based AI model based on the type of food identified, for example, "Is it safe for babies to eat apples?"

[1080] Step 7:

[1081] The server generates queries and sends them to a chat-based AI model to collect information about food safety.

[1082] Step 8:

[1083] The server receives the response from the chat-based AI model and analyzes its content, obtaining, for example, information that "it is safe for babies to eat apples, and it is recommended that they be cut into small pieces or grated."

[1084] Step 9:

[1085] The device analyzes the user's facial expressions and voice in real time to identify the user's emotional state. For example, if the user looks anxious or speaks in an anxious voice, it will be determined that the user is anxious.

[1086] Step 10:

[1087] The device transmits the user's emotional state data to the server, which then incorporates this data into the decision-making process.

[1088] Step 11:

[1089] The server determines the safety of food based on the information obtained and the user's emotional state. For example, if the user seems anxious, it may provide additional advice or warnings.

[1090] Step 12:

[1091] The server formats the information, including the judgment result and any additional advice, into a data format and prepares it for presentation to the user.

[1092] Step 13:

[1093] The server then sends the formatted result to the user's device. Once the transmission is confirmed, the server records a log and prepares for the next process.

[1094] Step 14:

[1095] The terminal receives the result data from the server, checks the integrity of the data, and returns a reception confirmation response to the server.

[1096] Step 15:

[1097] The device analyzes the received data and displays it on the user interface. For example, it might say, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them. If you're worried, start with a small amount and monitor your baby's condition."

[1098] This allows the system to recognize the user's emotions and provide appropriate advice and warnings based on those emotions, making it easy and safe to manage the diet of babies and pets at home.

[1099] Example 2

[1100] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1101] Currently, there is a lack of systems that provide appropriate information quickly and accurately for managing the diet of babies and pets. Furthermore, there are no systems that can provide advice or warnings that take into account the user's emotions. This leaves users with insufficient support to provide food to their babies and pets with peace of mind. Furthermore, in existing systems, information collection to determine food safety is often done manually, which is a very time-consuming process.

[1102] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1103] In this invention, the server includes: a means for a user to take an image of food using a camera on the terminal and send the image to the server; a means for the terminal to analyze the user's facial expressions and voice and acquire emotion data; a means for the server to analyze the received image and identify the type of food using an image recognition algorithm; a means for the server to collect information about the identified food using a generative AI model and determine whether the food is safe for babies or pets to ingest; and a means for the server to send additional advice or warnings based on the determination results and emotion data to the user's terminal and display the results on the terminal. This allows the user to provide quick and appropriate information when feeding babies or pets, ensuring safety, and providing advice or warnings according to the user's emotions.

[1104] "Terminal" refers to a smartphone, tablet, or other portable information processing device used by a user.

[1105] "Server" refers to a computer system that receives, stores, analyzes, searches for information, and transmits results of image data.

[1106] "User" refers to a person who uses the system to manage the diet of a baby or pet.

[1107] "Image recognition algorithm" refers to a computational method for analyzing received image data and identifying the type of food depicted in the image.

[1108] A "deep learning model" refers to an artificial intelligence technology that learns and classifies the characteristics of ingredients as part of an image recognition algorithm.

[1109] "Generative AI model" refers to a system that uses chat-based artificial intelligence to collect and generate information about food.

[1110] A "prompt" refers to a query or question that is input to a generative AI model.

[1111] "Emotion engine" refers to technology that analyzes a user's facial expressions and voice to identify their emotional state.

[1112] "Determination result" refers to data containing the evaluation results of the server's evaluation of food safety.

[1113] "Emotion data" refers to data that indicates the user's emotional state obtained through facial recognition or voice analysis.

[1114] "Additional advice or warning" refers to supplementary advice or a message urging caution that is provided to the user based on the judgment result.

[1115] This invention relates to a system that supports dietary management for babies and pets, and in particular, it combines an emotion engine that recognizes the user's emotions. This system provides a process for quickly determining the safety of food.

[1116] A user takes a picture of food using a device such as a smartphone or tablet. For example, the user takes a picture of an apple. The device temporarily saves the captured image in local storage, and after confirming network connectivity, sends the image data to the server's API endpoint.

[1117] The server receives the image data sent from the device and temporarily stores it. It then returns a receipt confirmation response to the device. It then uses an image recognition algorithm to analyze the received image and identify the type of food in the image. For example, it uses a deep learning framework such as TensorFlow or PyTorch. The algorithm learns the characteristics of the ingredients contained in the image and classifies them based on that. The server then obtains the identified type of food, such as "apple."

[1118] The server then generates a query to a generative AI model (such as OpenAI's ChatGPT) based on the identified food type, for example, creating a specific question such as "Is it safe for babies to eat apples?" The generated query is a prompt sentence of the form:

[1119] "Is it safe for babies to eat apples?"

[1120] The server sends this prompt to the generative AI model, collects comprehensive information from the internet and databases, receives the response from the generative AI model, and analyzes its content. For example, the server obtains the information that "It is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."

[1121] When a user takes a picture, the device runs an emotion engine (such as the Affectiva SDK) to analyze the user's facial expressions and voice. This analysis is performed in real time to understand the user's emotional state. The device then transmits the acquired data on the user's emotional state to the server.

[1122] The server determines the safety of the food based on the obtained information and the user's emotional state. For example, even if the result indicates that the food is safe for babies or pets, if the user's emotions are deemed to be uneasy, the server will add additional advice or warnings. The server then formats the result and the information that forms the basis of the result into a data format and sends the formatted result to the user's device. Once the transmission is confirmed, the server records the log and prepares for the next process.

[1123] The device receives the result data sent from the server and checks the integrity of the data. It then returns a reception confirmation response to the server. Finally, the device analyzes the received data and displays it on the user interface. For example, it might display, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them. If you are concerned, start with a small amount and monitor the baby's condition," and if the user's emotions are uneasy, it might display additional advice such as, "If you are concerned, start with a small amount and monitor the baby's condition."

[1124] This system recognizes the user's emotions and provides appropriate advice and warnings based on those emotions, making it easy and safe to manage the diet of babies and pets at home.

[1125] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1126] Step 1:

[1127] A user takes a picture of food using the device's camera. Specifically, the user launches the camera app, places food (e.g., an apple) in the center of the screen, and presses the shutter button. The input is the camera operation and the captured food image, and the output is an image file saved in local storage.

[1128] Step 2:

[1129] The device temporarily saves the image captured in local storage and checks the network connection. Specifically, it saves the image file in the device's local storage and then checks whether the WiFi or mobile data connection is enabled. The input is the image file captured in step 1 and the network status, and the output is the state ready to be sent.

[1130] Step 3:

[1131] The device sends image data to the server's API endpoint. Specifically, the device sends the image file to the server as an HTTP POST request. The input is the saved image file and the server's API endpoint, and the output is the completion status of the transmission to the server.

[1132] Step 4:

[1133] The server receives image data sent from the terminal and temporarily stores it. Specifically, the server extracts image data from the HTTP request it receives and stores it in a temporary directory on the server. The input is the image data sent from the terminal, and the output is the temporarily stored image file and a response confirming receipt.

[1134] Step 5:

[1135] The server uses an image recognition algorithm to analyze the received image and identify the type of food. Specifically, the server analyzes the image using a deep learning model such as TensorFlow or PyTorch to identify the type of food (e.g., "apple"). The input is a temporarily saved image file, and the output is information about the identified type of food (e.g., "apple").

[1136] Step 6:

[1137] The server generates a query to the generative AI model based on the identified food type. Specifically, the server generates a prompt sentence, "Is it safe for babies to eat apples?" and sends it to the generative AI model. The input is food type information (e.g., "apple"), and the output is the generated query. An example of a prompt sentence: "Is it safe for babies to eat apples?"

[1138] Step 7:

[1139] The server sends the generated query to the generative AI model and collects information from databases and the Internet. Specifically, the server sends a prompt to the generative AI model as an HTTP request and receives food safety information in response. The input is the generated query, and the output is the received safety information.

[1140] Step 8:

[1141] The server analyzes the response from the AI ​​model and determines the safety of the food. Specifically, it analyzes the received safety information and obtains a judgment result such as "It is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated." The input is the received safety information, and the output is the judgment result.

[1142] Step 9:

[1143] When a user takes a picture, the device uses an emotion engine to analyze the user's facial expressions and voice to identify their emotional state. Specifically, the device's front camera and microphone are used to collect facial expressions and voice in real time, which are then analyzed by the emotion engine. The input is the user's facial expression and voice data, and the output is the identified emotional state data (e.g., "anxiety").

[1144] Step 10:

[1145] The emotional state data acquired by the device is sent to the server. Specifically, the emotional state data is converted into JSON format and sent to the server as an HTTP request. The input is the emotional state data, and the output is the status of completion of transmission to the server.

[1146] Step 11:

[1147] The server determines the safety of the food based on the information and emotional state obtained, and generates additional advice or a warning. Specifically, it adds additional advice to the judgment result, such as "If you are concerned, start with a small amount and monitor the situation." The input is the judgment result and emotional state data, and the output is the final judgment result and additional advice.

[1148] Step 12:

[1149] The server sends the final judgment result and additional advice to the user's device. Specifically, the server formats the data in JSON format and sends it to the user's device as an HTTP request. The input is the final judgment result and additional advice, and the output is the status of completion of transmission to the device.

[1150] Step 13:

[1151] The terminal receives the result data from the server and checks the integrity of the data. Specifically, it checks the received data using a checksum or other method and returns a response confirming receipt to the server. The input is the result data sent from the server, and the output is the response confirming receipt.

[1152] Step 14:

[1153] The device analyzes the result data and displays it on the user interface. Specifically, it displays the results and advice in the application's UI component, informing the user, for example, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them. If you are concerned, start with a small amount and monitor the baby's condition." The input is the analyzed result data, and the output is the display on the user interface.

[1154] (Application example 2)

[1155] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1156] Avoiding the risk of minors and animals accidentally ingesting harmful foods is an important issue in home dietary management. Users often have concerns or doubts about food safety, and appropriate information and advice that take these feelings into account is needed. However, existing systems rarely offer the functionality to integrate emotion recognition and food safety assessment in real time. Therefore, a system that allows users to manage their diet with peace of mind is needed.

[1157] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for a user to take an image of a food using a camera of a terminal and send the image to an information processing device; means for the information processing device to analyze the received image and identify the type of food using an image recognition algorithm; means for the information processing device to collect information about the identified food using an interactive AI model and determine whether the food is safe for minors or animals to consume; means for the information processing device to send the determination result to the user's terminal and display the result on the terminal; means for the terminal to analyze emotions from the user's facial expressions and voice and send the emotions to the information processing device; and means for the information processing device to add additional advice or warnings based on the obtained emotion data. This makes it possible to determine the safety of food in real time and provide appropriate information and advice according to the user's emotions.

[1158] A "terminal" is a portable information and communication device operated by a user, and is equipped with a camera and a display.

[1159] An "information processing device" is a computer system that analyzes received data and processes it using a specific algorithm.

[1160] An "image recognition algorithm" is a computational method for analyzing image data to identify specific objects or features.

[1161] An "interactive AI model" is a program that uses artificial intelligence to answer questions and search for information in natural language.

[1162] A "minor" is a child or young person who is not legally recognized as an adult.

[1163] "Animals" are living creatures such as dogs and cats kept as pets in homes.

[1164] "Emotion data" is information that indicates the emotional state of the user analyzed from facial expressions and voice.

[1165] "Additional advice and warnings" are supplemental information provided to help users provide food safely and with peace of mind.

[1166] System Program Overview

[1167] The system that realizes this application example consists of the following major components:

[1168] 1. Hardware

[1169] Terminal: A portable information and communication device (smartphone, tablet, etc.) operated by a user, equipped with a camera and display.

[1170] Information processing device: A server system that analyzes received data and processes it using specific algorithms. It is desirable for this server to be equipped with a high-performance CPU and GPU.

[1171] 2. Software

[1172] Image recognition algorithm: Built using deep learning libraries such as TensorFlow.

[1173] Conversational AI model: An engine used for natural language processing and question answering.

[1174] Emotion recognition software: Programs for facial expression recognition and speech analysis (OpenCV and other facial expression analysis tools).

[1175] Program processing overview

[1176] 1. Image capture and transmission

[1177] The user takes a photo of the food using the device's camera.

[1178] The terminal transmits the captured image to the information processing device.

[1179] 2. Image Recognition

[1180] The server analyzes the received images and uses image recognition algorithms to identify the type of food.

[1181] It uses a deep learning model to analyze the features of food in an image and identify its type.

[1182] 3. Information Search

[1183] The server sends information about the identified food as a query to the conversational AI model.

[1184] The AI ​​model collects the necessary information from databases and the internet to determine whether the food is safe for minors and animals.

[1185] 4. Emotion recognition

[1186] When a user takes a picture, the device analyzes emotions in real time from facial expressions and voice.

[1187] The analyzed emotion data is transmitted to an information processing device.

[1188] 5. Judgment and result display

[1189] The server generates additional advice and warnings based on the food safety assessment results and emotion data, and sends them to the user's device.

[1190] The terminal displays the received results on a user interface.

[1191] Specific examples

[1192] Example 1: Judging whether an apple is suitable for babies

[1193] 1. A user takes a photo of an apple and asks the camera, "Are apples safe for babies?"

[1194] 2. The device sends the emotional data analyzed as "anxiety" to the server.

[1195] 3. The server generates advice such as, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them. If you are concerned, start with a small amount and monitor their condition." and provides it to the user.

[1196] Prompt Sentence Examples

[1197] Take a photo of the food you are about to give your baby. Then ask, "Is an apple safe for my baby?" Then, enter your emotion (e.g., "anxious" or "surprised").

[1198] Specific examples of the technologies used

[1199] Image recognition algorithm: TensorFlow

[1200] Conversational AI model: Natural language processing engine

[1201] Emotion recognition software: OpenCV

[1202] These technologies allow users to confidently verify food safety and provide appropriate diets for minors and animals.

[1203] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1204] Step 1:

[1205] The user takes a picture of food using the device's camera and temporarily saves the captured image data in local storage. This is done using a camera application or the device's camera function. The input is the captured image data, and the output is an image file saved in the device's local storage.

[1206] Step 2:

[1207] The device checks the network connection and sends the captured image data to the information processing device. The transmission is performed using HTTP or other communication protocols. The input is an image file saved in local storage, and the output is the image data sent to the server.

[1208] Step 3:

[1209] The server temporarily stores the received image data and begins analysis using an image recognition algorithm. The type of food is identified using a deep learning library such as TensorFlow. The input is the image data stored on the server, and the output is the type of food identified through analysis.

[1210] Step 4:

[1211] The server sends a query to the conversational AI model based on the identified food type to collect safety information about the food. For example, a query such as "Is it safe for babies to eat apples?" is generated and the AI ​​model searches for information. The input is the identified food type information, and the output is information about the food obtained from the AI ​​model.

[1212] Step 5:

[1213] The device analyzes emotions from the user's facial expressions and voice when taking an image. It uses OpenCV and other facial expression analysis tools to obtain the user's emotional data in real time. The input is the user's facial expression and voice data, and the output is the emotional data obtained through the analysis.

[1214] Step 6:

[1215] The device sends the analyzed emotional data to the server using a communication method such as the HTTP protocol. The input is the emotional data analyzed on the device, and the output is the emotional data sent to the server.

[1216] Step 7:

[1217] The server generates additional advice and warnings based on food safety information and emotion data. Based on the obtained data, it generates appropriate messages and adds advice to alleviate the user's anxiety. The input is food safety information and emotion data, and the output is the generated advice or warning message.

[1218] Step 8:

[1219] The server generates advice and warning messages and sends them to the user's terminal. The input is the generated message and the output is the message sent to the user's terminal.

[1220] Step 9:

[1221] The terminal analyzes the received result data and displays it on the user interface. Based on this information, the user can decide whether to provide food to minors or animals. The input is the received message data, and the output is the information displayed on the user interface.

[1222] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1223] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1224] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1225] [Fourth embodiment]

[1226] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1227] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1228] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1229] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1230] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1231] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1232] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1233] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1234] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1235] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1236] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1237] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1238] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1239] This invention relates to a system for supporting dietary management for babies and pets (dogs and cats). This system works by having the user take a picture of food using the camera on their device and send it to a server. The server analyzes the received image and identifies the type of food using an image recognition algorithm. Next, it uses a chat-based AI model to collect safety information about the identified food and, based on that information, determines whether the food is safe for babies and pets to consume. The result of the determination is sent to the user's device and displayed on the device.

[1240] Program processing explanation

[1241] 1. Image capture and transmission

[1242] User: Activates the camera on a device such as a smartphone or tablet and takes a picture of food. For example, the user takes a picture of an apple.

[1243] Device: The captured image is temporarily saved in local storage. After that, the network connection is confirmed and the captured image is sent to the server.

[1244] 2. Image Recognition

[1245] Server: Receives image data sent from the device and temporarily stores it.

[1246] Server: Analyzes the received image using an image recognition algorithm to identify the type of food in the image. The algorithm uses a deep learning model to learn the characteristics of the ingredients and classify them accordingly.

[1247] Server: Get a specific food type, such as "apple."

[1248] 3. Information Search

[1249] Server: Generates queries to the chat-based AI model based on the identified food types, for example, formulating specific questions such as "Is it safe for babies to eat apples?"

[1250] Server: Sends the generated queries to the chat-based AI model and collects comprehensive information from databases and the Internet.

[1251] Server: Analyzes the information returned as a response from the AI ​​model and determines the safety of food based on that information. For example, it may determine that "it is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."

[1252] 4. Results display

[1253] Server: Formats the data so that the obtained information and judgment results are presented to the user in an easy-to-understand manner.

[1254] Server: Sends the judgment result to the user's terminal.

[1255] Terminal: Analyzes the received results and displays them on the user interface. For example, it displays advice such as "Apples are suitable for babies, but it is recommended to cut them into small pieces or grate them."

[1256] Specific examples

[1257] Example 1: An apple for a baby

[1258] User: Take a photo of an apple, save it on the device, and send the image to the server.

[1259] Server: Identifies an apple using an image recognition algorithm. Sends a query to the chat-based AI model asking, "Is it safe for babies to eat apples?"

[1260] Server: Collects information from the AI ​​model and determines whether it is safe for babies to eat apples. Sends the result to the user's device.

[1261] Device: Displays the result: "Apples are suitable for babies, but it is recommended that they be cut into small pieces or grated."

[1262] Example 2: Chocolate for dogs

[1263] User: Take a photo of the chocolate and save it on the device. Send the image to the server.

[1264] Server: Identifies chocolate using an image recognition algorithm. Sends a query to the chat-based AI model asking, "Is it safe for dogs to eat chocolate?"

[1265] Server: Collects information from the AI ​​model and determines whether chocolate is harmful to dogs. Sends the result to the user's device.

[1266] Device: Displays the result: "Chocolate is harmful to dogs and should not be given to them."

[1267] This allows you to quickly and easily determine what foods your baby or pets can and cannot eat in your home.

[1268] The processing flow will be explained below.

[1269] Step 1:

[1270] The user activates the device's camera and takes an image of food, for example, an apple.

[1271] Step 2:

[1272] The device temporarily saves the captured image in local storage, then checks for network connectivity and sends the image data to the server's API endpoint.

[1273] Step 3:

[1274] The server receives the image data sent from the terminal, stores it temporarily, and returns a reception confirmation response to the terminal.

[1275] Step 4:

[1276] The server analyzes the received image data using an image recognition algorithm. Specifically, the image is input into a deep learning model for image recognition to identify the type of food.

[1277] Step 5:

[1278] The server obtains the identified food type (e.g., "apple") based on the results of the image recognition algorithm.

[1279] Step 6:

[1280] The server generates a query to the chat-based AI model based on the identified food type, for example, creating a specific question such as "Is it safe for babies to eat apples?"

[1281] Step 7:

[1282] The server generates queries and sends them to a chat-based AI model, which collects comprehensive information from databases and the Internet.

[1283] Step 8:

[1284] The server receives the response from the chat-based AI model and analyzes its content, obtaining, for example, information that "it is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."

[1285] Step 9:

[1286] The server determines the safety of the food based on the information obtained, and then formats the determination results and the information that supports them into a data format.

[1287] Step 10:

[1288] The server then sends the formatted result to the user's device. Once the transmission is confirmed, the server records a log and prepares for the next process.

[1289] Step 11:

[1290] The terminal receives the result data from the server, checks the integrity of the data, and returns a reception confirmation response to the server.

[1291] Step 12:

[1292] The device analyzes the received data and displays it on the user interface, for example, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them."

[1293] Through these steps, users can easily and quickly determine what foods babies and pets in the home can and cannot eat.

[1294] Example 1

[1295] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1296] In modern society, dietary management for babies and animals is an important issue, and there is a need for a system that can quickly and accurately obtain information for providing appropriate food. However, manually collecting and assessing information requires time and effort, and there is a risk of making decisions based on incorrect information. Conventional systems have had difficulty efficiently providing information needed to accurately determine food safety. The present invention aims to solve these problems and provide a system that supports dietary management for babies and animals.

[1297] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1298] In this invention, the server includes a means for a user to take an image of a food using a camera on a terminal and send the image to the server, a means for the server to analyze the received image and identify the type of food using an image recognition algorithm, and a means for the server to collect information about the identified food using a generative AI model and determine whether the food is safe for babies or animals to consume, thereby enabling users to easily identify foods suitable for babies and animals and confirm their safety.

[1299] A "user" is a person who uses the system to take pictures of food and transmits the pictures to the server.

[1300] A "terminal" refers to a computer device used by a user, such as a smartphone or tablet, that has a camera function and a network connection function.

[1301] A "camera" is an image capturing device built into a terminal.

[1302] "Food" refers to all food and drink consumed by babies and animals.

[1303] A "server" is a computer system that processes data submitted by users and performs specific functions.

[1304] An "image recognition algorithm" is a computational method for analyzing a received image and identifying objects within the image.

[1305] A "deep learning model" is a type of machine learning that uses a multi-layered neural network to learn patterns from large amounts of data and is used to analyze images, audio, and other data.

[1306] A "generative AI model" is an algorithm that generates a response in natural language in response to an input prompt.

[1307] A "prompt sentence" is a sentence entered to ask a generative AI model for specific information.

[1308] "Baby" refers to an infant or young child, especially a child within the first year of life.

[1309] "Animals" refers to pets kept at home, specifically mammals such as dogs and cats.

[1310] "Ingestion" refers to the act of taking food into the mouth and digesting and absorbing it.

[1311] "Displaying on the terminal" refers to making the results visible to the user in the form of text or graphics on the terminal screen.

[1312] The present invention relates to a system that supports the dietary management of babies and animals, and operates by allowing a user to take an image of food using a camera on a terminal and send the image to a server. The system operates as follows.

[1313] First, a user uses a device such as a smartphone or tablet to activate the camera and take an image of food. For example, the user takes a photo of an apple. The device temporarily saves the captured image in local storage and then checks for an Internet connection. Once it confirms that a network connection has been established, it sends the captured image to the server.

[1314] The server receives the image data sent from the device and temporarily stores it. It then uses an image recognition algorithm to analyze the received image. This algorithm uses deep learning models such as TensorFlow and PyTorch to learn the characteristics of food and classify it based on that. The server then identifies the type of food in the image and obtains the identified food type, such as "apple."

[1315] Next, the server generates a query to a generative AI model based on the identified food type. Specifically, it creates a question such as, "Is it safe for babies to eat apples?" The generative AI model used is OpenAI's GPT-4. The server sends the generated query to the generative AI model, which collects comprehensive information from databases and the internet.

[1316] The server analyzes the information returned as a response from the generative AI model and uses it to determine the safety of food. For example, it may determine that "apples are safe for babies to eat, but it is recommended that they be cut into small pieces or grated." The result is then formatted and converted into a user-friendly format.

[1317] Finally, the server sends the result of the judgment to the user's device, which then analyzes the result and displays it on the user interface. For example, it might say, "Apples are suitable for babies, but it is recommended that you cut them into small pieces or grate them."

[1318] Specific examples

[1319] Apples for babies

[1320] A user takes a photo of an apple, saves it on their device, and sends it to the server.

[1321] The server uses an image recognition algorithm to identify the apple and sends the generative AI model a query: "Is it safe for a baby to eat the apple?"

[1322] The server collects information from the AI ​​model, determines whether it is safe for babies to eat apples, and sends the result to the user's device.

[1323] The device will display the result: "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them."

[1324] Chocolate for animals

[1325] The user takes a photo of the chocolate, saves it on the device, and sends it to the server.

[1326] The server uses an image recognition algorithm to identify the chocolate and sends the generative AI model a query: "Is chocolate safe for dogs to eat?"

[1327] The server collects information from the AI ​​model, determines whether chocolate is harmful to dogs, and sends the result to the user's device.

[1328] The device will display the result: "Chocolate is harmful to dogs and should not be given to them."

[1329] This makes it possible to quickly and easily determine what foods babies and animals in the home can and cannot eat.

[1330] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1331] Step 1:

[1332] A user activates the camera on a device such as a smartphone or tablet and takes a picture of food. The input is the food to be photographed (e.g., an apple), and the output is an image file saved on the device (e.g., IMG_20230101.jpg). Specifically, the user opens the camera app, focuses on the apple, and presses the shutter button.

[1333] Step 2:

[1334] The device temporarily saves the captured image in local storage. It then checks the Internet connection. The input is the captured image file, and the output is the temporarily saved image file and the result of checking the network connection status. Specifically, it saves the image file in a specific folder and checks the Wi-Fi and mobile data connection status.

[1335] Step 3:

[1336] Once the device has confirmed a network connection, it uses the REST API to send image data to the server. The input is the temporarily saved image file and the network connection status, and the output is the image data sent to the server. Specifically, it creates an HTTP POST request, encodes the image file, and sends it to the specified URL on the server.

[1337] Step 4:

[1338] The server receives image data sent from the terminal and temporarily stores it. The input is the image data sent from the terminal, and the output is the image data temporarily stored in the server. Specifically, the received image file is stored in the server's temporary directory (e.g., / tmp).

[1339] Step 5:

[1340] The server analyzes the received image using an image recognition algorithm. The input is the temporarily stored image data, and the output is the identified type of food (e.g., apple). Specifically, the image file is input into a deep learning model (e.g., TensorFlow), which extracts and classifies the food's features.

[1341] Step 6:

[1342] The server generates a query to the generative AI model based on the identified food type. The input is the identified food type (e.g., apple), and the output is a generated prompt (e.g., "Is it safe for babies to eat apples?"). Specifically, the server uses the identified food type to create a question-style prompt.

[1343] Step 7:

[1344] The server sends the generated query to the generative AI model and collects information from databases and the Internet. The input is the generated prompt, and the output is the information returned by the generative AI model (e.g., "It is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."). Specific operations include creating an API request and sending the prompt to the generative AI model.

[1345] Step 8:

[1346] The server analyzes the response from the generative AI model and determines the safety of the food. The input is the information returned from the generative AI model, and the output is the safety judgment result (e.g., "safe"). Specifically, it extracts information about safety from the response text and determines the judgment result.

[1347] Step 9:

[1348] The server formats the resulting judgment results in a data format that is easy to understand for the user. The input is the safety judgment result, and the output is the formatted data (e.g., "Apples are suitable for babies, but it is recommended that they be cut into small pieces or grated"). Specific operations include converting the data into a format such as JSON.

[1349] Step 10:

[1350] The server sends the result of the judgment to the user's device. The input is the formatted data, and the output is the data received by the user's device. The specific operation is to send the data as an API response.

[1351] Step 11:

[1352] The device analyzes the received judgment result and displays it on the user interface. The input is the received judgment result data, and the output is the judgment result displayed to the user (e.g., "Apples are suitable for babies, but it is recommended that they be cut into small pieces or grated"). Specific actions include displaying a pop-up message or notification.

[1353] (Application example 1)

[1354] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1355] In recent years, the number of users of online food delivery services has increased, and there is a growing need to confirm in advance whether the food being ordered is safe, especially for households with babies or pets. However, current food delivery systems lack a function to quickly and easily determine whether the ordered food is suitable for babies or pets. This requires users to spend time researching various information themselves, and the reliability of that information cannot be guaranteed. Therefore, the present invention aims to provide a food delivery system that quickly and reliably determines the safety of food and provides it to users.

[1356] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1357] In this invention, the server includes: means for a user to take an image of food using a camera on a terminal and send the image to the server; means for the server to analyze the received image and identify the type of food using an image recognition algorithm; means for a food delivery application for a user to take an image of the food to be ordered and use the image to obtain safety information about the ingredients; means for the server to collect information about the identified food using a chat-based generative AI model and determine whether the food is safe for babies and pets to consume; and means for the server to send the determination result to the user's terminal and display the result on the terminal. This allows a user to use the food delivery application to check the effects of the food to be ordered on babies and pets and ensure its safety.

[1358] "Terminal" means a portable electronic device for inputting, processing, and outputting information.

[1359] A "camera" is a device for taking still images and videos, and is built into a terminal.

[1360] A "server" is a computer system that stores, processes, and distributes information over a network.

[1361] An "image recognition algorithm" is a computational method for analyzing image data and identifying specific objects or features.

[1362] "Food type" is information that indicates the specific classification or name of the food.

[1363] A "chat-type generative AI model" is an artificial intelligence model that answers questions and generates information based on text data.

[1364] A "food delivery application" is software that allows users to order and have food delivered online.

[1365] This invention is a system applied to a food delivery application that allows users to check the safety of food they plan to order. It consists of a terminal used by the user, a server, an image recognition algorithm, a generative AI model, and a food delivery application.

[1366] System configuration

[1367] 1. Terminal

[1368] The user uses a device such as a smartphone or tablet. The device has a built-in camera and is used to take pictures of food. The captured image data is temporarily stored on the device and then sent to the server.

[1369] 2. Server

[1370] The server has multiple functions. It receives and stores image data, analyzes it using an image recognition algorithm to identify the type of food, and uses a generative AI model to collect safety information about the identified food. It then sends the collected information to the user's device and displays it in an appropriate format.

[1371] 3. Image Recognition Algorithm

[1372] The server uses TensorFlow and PyTorch as image recognition algorithms to automatically identify the type of food from the captured image. By using a deep learning model, food can be identified with high accuracy.

[1373] 4. Generative AI Models

[1374] The generative AI model uses a chat-type AI model such as GPT-4. Once the type of food is identified, specific prompts such as "Is it safe for babies to eat XX?" or "Is it safe for dogs to eat XX?" are sent to the generative AI model to obtain relevant safety information.

[1375] 5. Food delivery applications

[1376] The food delivery application provides an interface for users to check the safety of the food they plan to order. The safety information sent from the server is displayed on the user's device, allowing them to easily check whether the food is suitable for babies and pets before ordering.

[1377] Processing flow

[1378] 1. Image capture

[1379] The user uses the smartphone camera to take a photo of the food they plan to order.

[1380] 2. Sending images

[1381] The device sends the captured image to the server, which receives and stores the image data.

[1382] 3. Image Analysis

[1383] The server uses image recognition algorithms to analyze the captured image and identify the type of food, such as "pizza," "spaghetti," or "carbonara."

[1384] 4. Safety assessment

[1385] Based on the type of food identified, prompts such as "Is it safe for babies to eat XX?" or "Is it safe for dogs to eat XX?" are sent to the generative AI model, and safety information is collected from the generative AI model.

[1386] 5. Display results

[1387] The server analyzes the collected data and sends it to the device in an easy-to-understand format. The results are displayed on the user's device, allowing them to see the effects on their baby or pet.

[1388] Specific examples

[1389] Example 1: Pizza for babies

[1390] Prompt: "Is it safe for babies to eat pizza?"

[1391] The user takes a photo of the pizza and sends it to the system.

[1392] The server uses an image recognition algorithm to identify the pizza and send a prompt to the generative AI model.

[1393] Based on information from the AI ​​model, a decision is made as to whether the pizza is suitable.

[1394] Example 2: Carbonara for dogs

[1395] Prompt: "Is it safe for dogs to eat carbonara?"

[1396] The user takes a photo of the carbonara and sends it to the system.

[1397] The server uses an image recognition algorithm to identify carbonara and send a prompt to the generative AI model.

[1398] Based on information from the AI ​​model, the results will show whether carbonara is harmful to dogs.

[1399] In this way, users can quickly and easily check whether the food they plan to order is suitable for babies or pets.

[1400] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1401] Step 1:

[1402] Input: The user takes a photo of the food they plan to order.

[1403] How it works: The user activates the device's camera, takes a picture of the food they plan to order (e.g., pizza, carbonara), and temporarily saves it to the device's local storage.

[1404] Step 2:

[1405] Input: Captured image data.

[1406] Operation: The device checks for network connectivity and sends the stored image data to the server.

[1407] Output: Image data is sent to the server.

[1408] Step 3:

[1409] Input: Image data received by the server.

[1410] How it works: The server temporarily stores the image data it receives, then applies image recognition algorithms (using TensorFlow or PyTorch) to identify the type of food in the image.

[1411] Output: The type of food is identified (e.g. "pizza" or "carbonara").

[1412] Step 4:

[1413] Input: The type of food identified.

[1414] How it works: The server generates queries for the generative AI model based on the identified food types, creating specific questions like "Is it safe for babies to eat pizza?" or "Is it safe for dogs to eat carbonara?"

[1415] Output: The generated query.

[1416] Step 5:

[1417] Input: The generated query.

[1418] How it works: The server sends the generated queries to the generative AI model, which collects comprehensive information from the internet and databases. The generative AI model then generates answers to the questions and provides safety information.

[1419] Output: Response (safety information) from the generative AI model.

[1420] Step 6:

[1421] Input: The response from the generative AI model.

[1422] How it works: The server analyzes the collected information and formats it in a way that is easy for the user to understand. For example, it will give a judgment result such as "Pizza is suitable for babies, but it is recommended to cut it into small pieces."

[1423] Output: The result of the test for the user.

[1424] Step 7:

[1425] Input: Judgment result.

[1426] Operation: The server sends the result of the judgment to the user's device. The device analyzes the received result and displays it on the interface of the food delivery application.

[1427] Output: The result displayed on the device (e.g., "Pizza is suitable for babies, but it is recommended to cut it into small pieces").

[1428] In this way, users can easily check whether the food they plan to order is suitable for babies or pets.

[1429] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1430] This invention relates to a system that supports dietary management for babies and pets (dogs and cats), and in particular, it combines an emotion engine that recognizes the user's emotions. This system provides a process for quickly determining the safety of food.

[1431] Program processing explanation

[1432] 1. Image capture and transmission

[1433] User: Activates the camera on a device such as a smartphone or tablet and takes a picture of food. For example, the user takes a picture of an apple.

[1434] Device: The captured image is temporarily saved in local storage, and after checking the network connection, the image data is sent to the server's API endpoint.

[1435] 2. Image Recognition

[1436] Server: Receives image data sent from the terminal, temporarily stores it, and returns a reception confirmation response to the terminal.

[1437] Server: Analyzes the received image using an image recognition algorithm to identify the type of food in the image. The algorithm uses a deep learning model to learn the characteristics of the ingredients and classify them accordingly.

[1438] Server: Get a specific food type, such as "apple."

[1439] 3. Information Search

[1440] Server: Generates a query to the chat-based AI model based on the identified food type, for example, formulating a specific question such as "Is it safe for babies to eat apples?"

[1441] Server: Sends the generated queries to the chat-based AI model and collects comprehensive information from databases and the Internet.

[1442] Server: Receives the response from the AI ​​model and analyzes its content. For example, it obtains information such as "It is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."

[1443] 4. Emotion recognition

[1444] Device: When a user takes a picture, the emotion engine is activated and analyzes the user's facial expressions and voice. This analysis is performed in real time to understand the user's emotional state.

[1445] Terminal: Transmits the acquired user emotional state data to the server.

[1446] 5. Judgment and result display

[1447] Server: Determines the safety of food based on the obtained information and the user's emotional state. For example, if the result is that it is safe for babies or pets, or if the user's emotional state is deemed unsafe, it provides additional advice or a warning.

[1448] Server: Formats the judgment results and the information that forms the basis for them into a data format.

[1449] Server: The formatted result is sent to the user's device. Once the transmission is confirmed, the server records a log and prepares for the next process.

[1450] Terminal: Receives the result data from the server, checks the integrity of the data, and returns a reception confirmation response to the server.

[1451] Terminal: Analyzes the received data and displays it on the user interface. For example, it might display "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them," and if the user's emotions are anxious, it might display additional advice such as "If you are worried, start with a small amount and monitor the baby's reaction."

[1452] Specific examples

[1453] Example 1: Applying apples and emotion recognition to babies

[1454] User: Take a photo of an apple, save it on the device, and send the image to the server.

[1455] Server: Identifies an apple using an image recognition algorithm. Sends a query to the chat-based AI model asking, "Is it safe for babies to eat apples?"

[1456] Server: Collects information from the AI ​​model and determines whether it is safe for babies to eat apples.

[1457] Device: When the user takes a photo, facial recognition and voice analysis are used to determine the emotional state as anxiety.

[1458] Server: The judgement result was accompanied by additional advice: "If you are concerned, start with a small amount and monitor the condition."

[1459] Device: The result is displayed as follows: "Apples are suitable for babies, but it is recommended that they be cut into small pieces or grated. If you are concerned, start with a small amount and monitor the baby's condition."

[1460] Example 2: Applying chocolate and emotion recognition to dogs

[1461] User: Take a photo of the chocolate and save it on the device. Send the image to the server.

[1462] Server: Identifies chocolate using an image recognition algorithm. Sends a query to the chat-based AI model asking, "Is it safe for dogs to eat chocolate?"

[1463] Server: Collects information from the AI ​​model and determines that chocolate is harmful to dogs.

[1464] Device: When the user takes a photo, facial recognition and voice analysis are used to determine the emotional state as surprise.

[1465] Server: The verdict included a warning that "chocolate is extremely dangerous to dogs and should never be given to them."

[1466] Device: Displays the result: "Chocolate is harmful to dogs and should never be given to them," with an additional warning: "Move out of reach immediately."

[1467] This allows the system to recognize the user's emotions and provide appropriate advice and warnings based on those emotions, making it easy and safe to manage the diet of babies and pets at home.

[1468] The processing flow will be explained below.

[1469] Step 1:

[1470] The user activates the camera on the device and takes an image of food, for example, an apple.

[1471] Step 2:

[1472] The device temporarily saves the captured image in local storage, then checks for network connectivity and sends the image data to the server's API endpoint.

[1473] Step 3:

[1474] The server receives the image data sent from the terminal, stores it temporarily, and returns a reception confirmation response to the terminal.

[1475] Step 4:

[1476] The server analyzes the received image data using an image recognition algorithm that identifies food characteristics and identifies a particular food type.

[1477] Step 5:

[1478] The server obtains the identified food type based on the results of the image recognition algorithm, for example, identifying "apple."

[1479] Step 6:

[1480] The server generates a query to the chat-based AI model based on the type of food identified, for example, "Is it safe for babies to eat apples?"

[1481] Step 7:

[1482] The server generates queries and sends them to a chat-based AI model to collect information about food safety.

[1483] Step 8:

[1484] The server receives the response from the chat-based AI model and analyzes its content, obtaining, for example, information that "it is safe for babies to eat apples, and it is recommended that they be cut into small pieces or grated."

[1485] Step 9:

[1486] The device analyzes the user's facial expressions and voice in real time to identify the user's emotional state. For example, if the user looks anxious or speaks in an anxious voice, it will be determined that the user is anxious.

[1487] Step 10:

[1488] The device transmits the user's emotional state data to the server, which then incorporates this data into the decision-making process.

[1489] Step 11:

[1490] The server determines the safety of food based on the information obtained and the user's emotional state. For example, if the user seems anxious, it may provide additional advice or warnings.

[1491] Step 12:

[1492] The server formats the information, including the judgment result and any additional advice, into a data format and prepares it for presentation to the user.

[1493] Step 13:

[1494] The server then sends the formatted result to the user's device. Once the transmission is confirmed, the server records a log and prepares for the next process.

[1495] Step 14:

[1496] The terminal receives the result data from the server, checks the integrity of the data, and returns a reception confirmation response to the server.

[1497] Step 15:

[1498] The device analyzes the received data and displays it on the user interface. For example, it might say, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them. If you're worried, start with a small amount and monitor your baby's condition."

[1499] This allows the system to recognize the user's emotions and provide appropriate advice and warnings based on those emotions, making it easy and safe to manage the diet of babies and pets at home.

[1500] Example 2

[1501] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1502] Currently, there is a lack of systems that provide appropriate information quickly and accurately for managing the diet of babies and pets. Furthermore, there are no systems that can provide advice or warnings that take into account the user's emotions. This leaves users with insufficient support to provide food to their babies and pets with peace of mind. Furthermore, in existing systems, information collection to determine food safety is often done manually, which is a very time-consuming process.

[1503] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1504] In this invention, the server includes: a means for a user to take an image of food using a camera on the terminal and send the image to the server; a means for the terminal to analyze the user's facial expressions and voice and acquire emotion data; a means for the server to analyze the received image and identify the type of food using an image recognition algorithm; a means for the server to collect information about the identified food using a generative AI model and determine whether the food is safe for babies or pets to ingest; and a means for the server to send additional advice or warnings based on the determination results and emotion data to the user's terminal and display the results on the terminal. This allows the user to provide quick and appropriate information when feeding babies or pets, ensuring safety, and providing advice or warnings according to the user's emotions.

[1505] "Terminal" refers to a smartphone, tablet, or other portable information processing device used by a user.

[1506] "Server" refers to a computer system that receives, stores, analyzes, searches for information, and transmits results of image data.

[1507] "User" refers to a person who uses the system to manage the diet of a baby or pet.

[1508] "Image recognition algorithm" refers to a computational method for analyzing received image data and identifying the type of food depicted in the image.

[1509] A "deep learning model" refers to an artificial intelligence technology that learns and classifies the characteristics of ingredients as part of an image recognition algorithm.

[1510] "Generative AI model" refers to a system that uses chat-based artificial intelligence to collect and generate information about food.

[1511] A "prompt" refers to a query or question that is input to a generative AI model.

[1512] "Emotion engine" refers to technology that analyzes a user's facial expressions and voice to identify their emotional state.

[1513] "Determination result" refers to data containing the evaluation results of the server's evaluation of food safety.

[1514] "Emotion data" refers to data that indicates the user's emotional state obtained through facial recognition or voice analysis.

[1515] "Additional advice or warning" refers to supplementary advice or a message urging caution that is provided to the user based on the judgment result.

[1516] This invention relates to a system that supports dietary management for babies and pets, and in particular, it combines an emotion engine that recognizes the user's emotions. This system provides a process for quickly determining the safety of food.

[1517] A user takes a picture of food using a device such as a smartphone or tablet. For example, the user takes a picture of an apple. The device temporarily saves the captured image in local storage, and after confirming network connectivity, sends the image data to the server's API endpoint.

[1518] The server receives the image data sent from the device and temporarily stores it. It then returns a receipt confirmation response to the device. It then uses an image recognition algorithm to analyze the received image and identify the type of food in the image. For example, it uses a deep learning framework such as TensorFlow or PyTorch. The algorithm learns the characteristics of the ingredients contained in the image and classifies them based on that. The server then obtains the identified type of food, such as "apple."

[1519] The server then generates a query to a generative AI model (such as OpenAI's ChatGPT) based on the identified food type, for example, creating a specific question such as "Is it safe for babies to eat apples?" The generated query is a prompt sentence of the form:

[1520] "Is it safe for babies to eat apples?"

[1521] The server sends this prompt to the generative AI model, collects comprehensive information from the internet and databases, receives the response from the generative AI model, and analyzes its content. For example, the server obtains the information that "It is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated."

[1522] When a user takes a picture, the device runs an emotion engine (such as the Affectiva SDK) to analyze the user's facial expressions and voice. This analysis is performed in real time to understand the user's emotional state. The device then transmits the acquired data on the user's emotional state to the server.

[1523] The server determines the safety of the food based on the obtained information and the user's emotional state. For example, even if the result indicates that the food is safe for babies or pets, if the user's emotions are deemed to be uneasy, the server will add additional advice or warnings. The server then formats the result and the information that forms the basis of the result into a data format and sends the formatted result to the user's device. Once the transmission is confirmed, the server records the log and prepares for the next process.

[1524] The device receives the result data sent from the server and checks the integrity of the data. It then returns a reception confirmation response to the server. Finally, the device analyzes the received data and displays it on the user interface. For example, it might display, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them. If you are concerned, start with a small amount and monitor the baby's condition," and if the user's emotions are uneasy, it might display additional advice such as, "If you are concerned, start with a small amount and monitor the baby's condition."

[1525] This system recognizes the user's emotions and provides appropriate advice and warnings based on those emotions, making it easy and safe to manage the diet of babies and pets at home.

[1526] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1527] Step 1:

[1528] A user takes a picture of food using the device's camera. Specifically, the user launches the camera app, places food (e.g., an apple) in the center of the screen, and presses the shutter button. The input is the camera operation and the captured food image, and the output is an image file saved in local storage.

[1529] Step 2:

[1530] The device temporarily saves the image captured in local storage and checks the network connection. Specifically, it saves the image file in the device's local storage and then checks whether the WiFi or mobile data connection is enabled. The input is the image file captured in step 1 and the network status, and the output is the state ready to be sent.

[1531] Step 3:

[1532] The device sends image data to the server's API endpoint. Specifically, the device sends the image file to the server as an HTTP POST request. The input is the saved image file and the server's API endpoint, and the output is the completion status of the transmission to the server.

[1533] Step 4:

[1534] The server receives image data sent from the terminal and temporarily stores it. Specifically, the server extracts image data from the HTTP request it receives and stores it in a temporary directory on the server. The input is the image data sent from the terminal, and the output is the temporarily stored image file and a response confirming receipt.

[1535] Step 5:

[1536] The server uses an image recognition algorithm to analyze the received image and identify the type of food. Specifically, the server analyzes the image using a deep learning model such as TensorFlow or PyTorch to identify the type of food (e.g., "apple"). The input is a temporarily saved image file, and the output is information about the identified type of food (e.g., "apple").

[1537] Step 6:

[1538] The server generates a query to the generative AI model based on the identified food type. Specifically, the server generates a prompt sentence, "Is it safe for babies to eat apples?" and sends it to the generative AI model. The input is food type information (e.g., "apple"), and the output is the generated query. An example of a prompt sentence: "Is it safe for babies to eat apples?"

[1539] Step 7:

[1540] The server sends the generated query to the generative AI model and collects information from databases and the Internet. Specifically, the server sends a prompt to the generative AI model as an HTTP request and receives food safety information in response. The input is the generated query, and the output is the received safety information.

[1541] Step 8:

[1542] The server analyzes the response from the AI ​​model and determines the safety of the food. Specifically, it analyzes the received safety information and obtains a judgment result such as "It is safe for babies to eat apples, but it is recommended that they be cut into small pieces or grated." The input is the received safety information, and the output is the judgment result.

[1543] Step 9:

[1544] When a user takes a picture, the device uses an emotion engine to analyze the user's facial expressions and voice to identify their emotional state. Specifically, the device's front camera and microphone are used to collect facial expressions and voice in real time, which are then analyzed by the emotion engine. The input is the user's facial expression and voice data, and the output is the identified emotional state data (e.g., "anxiety").

[1545] Step 10:

[1546] The emotional state data acquired by the device is sent to the server. Specifically, the emotional state data is converted into JSON format and sent to the server as an HTTP request. The input is the emotional state data, and the output is the status of completion of transmission to the server.

[1547] Step 11:

[1548] The server determines the safety of the food based on the information and emotional state obtained, and generates additional advice or a warning. Specifically, it adds additional advice to the judgment result, such as "If you are concerned, start with a small amount and monitor the situation." The input is the judgment result and emotional state data, and the output is the final judgment result and additional advice.

[1549] Step 12:

[1550] The server sends the final judgment result and additional advice to the user's device. Specifically, the server formats the data in JSON format and sends it to the user's device as an HTTP request. The input is the final judgment result and additional advice, and the output is the status of completion of transmission to the device.

[1551] Step 13:

[1552] The terminal receives the result data from the server and checks the integrity of the data. Specifically, it checks the received data using a checksum or other method and returns a response confirming receipt to the server. The input is the result data sent from the server, and the output is the response confirming receipt.

[1553] Step 14:

[1554] The device analyzes the result data and displays it on the user interface. Specifically, it displays the results and advice in the application's UI component, informing the user, for example, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them. If you are concerned, start with a small amount and monitor the baby's condition." The input is the analyzed result data, and the output is the display on the user interface.

[1555] (Application example 2)

[1556] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1557] Avoiding the risk of minors and animals accidentally ingesting harmful foods is an important issue in home dietary management. Users often have concerns or doubts about food safety, and appropriate information and advice that take these feelings into account is needed. However, existing systems rarely offer the functionality to integrate emotion recognition and food safety assessment in real time. Therefore, a system that allows users to manage their diet with peace of mind is needed.

[1558] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for a user to take an image of a food using a camera of a terminal and send the image to an information processing device; means for the information processing device to analyze the received image and identify the type of food using an image recognition algorithm; means for the information processing device to collect information about the identified food using an interactive AI model and determine whether the food is safe for minors or animals to consume; means for the information processing device to send the determination result to the user's terminal and display the result on the terminal; means for the terminal to analyze emotions from the user's facial expressions and voice and send the emotions to the information processing device; and means for the information processing device to add additional advice or warnings based on the obtained emotion data. This makes it possible to determine the safety of food in real time and provide appropriate information and advice according to the user's emotions.

[1559] A "terminal" is a portable information and communication device operated by a user, and is equipped with a camera and a display.

[1560] An "information processing device" is a computer system that analyzes received data and processes it using a specific algorithm.

[1561] An "image recognition algorithm" is a computational method for analyzing image data to identify specific objects or features.

[1562] An "interactive AI model" is a program that uses artificial intelligence to answer questions and search for information in natural language.

[1563] A "minor" is a child or young person who is not legally recognized as an adult.

[1564] "Animals" are living creatures such as dogs and cats kept as pets in homes.

[1565] "Emotion data" is information that indicates the emotional state of the user analyzed from facial expressions and voice.

[1566] "Additional advice and warnings" are supplemental information provided to help users provide food safely and with peace of mind.

[1567] System Program Overview

[1568] The system that realizes this application example consists of the following major components:

[1569] 1. Hardware

[1570] Terminal: A portable information and communication device (smartphone, tablet, etc.) operated by a user, equipped with a camera and display.

[1571] Information processing device: A server system that analyzes received data and processes it using specific algorithms. It is desirable for this server to be equipped with a high-performance CPU and GPU.

[1572] 2. Software

[1573] Image recognition algorithm: Built using deep learning libraries such as TensorFlow.

[1574] Conversational AI model: An engine used for natural language processing and question answering.

[1575] Emotion recognition software: Programs for facial expression recognition and speech analysis (OpenCV and other facial expression analysis tools).

[1576] Program processing overview

[1577] 1. Image capture and transmission

[1578] The user takes a photo of the food using the device's camera.

[1579] The terminal transmits the captured image to the information processing device.

[1580] 2. Image Recognition

[1581] The server analyzes the received images and uses image recognition algorithms to identify the type of food.

[1582] It uses a deep learning model to analyze the features of food in an image and identify its type.

[1583] 3. Information Search

[1584] The server sends information about the identified food as a query to the conversational AI model.

[1585] The AI ​​model collects the necessary information from databases and the internet to determine whether the food is safe for minors and animals.

[1586] 4. Emotion recognition

[1587] When a user takes a picture, the device analyzes emotions in real time from facial expressions and voice.

[1588] The analyzed emotion data is transmitted to an information processing device.

[1589] 5. Judgment and result display

[1590] The server generates additional advice and warnings based on the food safety assessment results and emotion data, and sends them to the user's device.

[1591] The terminal displays the received results on a user interface.

[1592] Specific examples

[1593] Example 1: Judging whether an apple is suitable for babies

[1594] 1. A user takes a photo of an apple and asks the camera, "Are apples safe for babies?"

[1595] 2. The device sends the emotional data analyzed as "anxiety" to the server.

[1596] 3. The server generates advice such as, "Apples are suitable for babies, but we recommend cutting them into small pieces or grating them. If you are concerned, start with a small amount and monitor their condition." and provides it to the user.

[1597] Prompt Sentence Examples

[1598] Take a photo of the food you are about to give your baby. Then ask, "Is an apple safe for my baby?" Then, enter your emotion (e.g., "anxious" or "surprised").

[1599] Specific examples of the technologies used

[1600] Image recognition algorithm: TensorFlow

[1601] Conversational AI model: Natural language processing engine

[1602] Emotion recognition software: OpenCV

[1603] These technologies allow users to confidently verify food safety and provide appropriate diets for minors and animals.

[1604] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1605] Step 1:

[1606] The user takes a picture of food using the device's camera and temporarily saves the captured image data in local storage. This is done using a camera application or the device's camera function. The input is the captured image data, and the output is an image file saved in the device's local storage.

[1607] Step 2:

[1608] The device checks the network connection and sends the captured image data to the information processing device. The transmission is performed using HTTP or other communication protocols. The input is an image file saved in local storage, and the output is the image data sent to the server.

[1609] Step 3:

[1610] The server temporarily stores the received image data and begins analysis using an image recognition algorithm. The type of food is identified using a deep learning library such as TensorFlow. The input is the image data stored on the server, and the output is the type of food identified through analysis.

[1611] Step 4:

[1612] The server sends a query to the conversational AI model based on the identified food type to collect safety information about the food. For example, a query such as "Is it safe for babies to eat apples?" is generated and the AI ​​model searches for information. The input is the identified food type information, and the output is information about the food obtained from the AI ​​model.

[1613] Step 5:

[1614] The device analyzes emotions from the user's facial expressions and voice when taking an image. It uses OpenCV and other facial expression analysis tools to obtain the user's emotional data in real time. The input is the user's facial expression and voice data, and the output is the emotional data obtained through the analysis.

[1615] Step 6:

[1616] The device sends the analyzed emotional data to the server using a communication method such as the HTTP protocol. The input is the emotional data analyzed on the device, and the output is the emotional data sent to the server.

[1617] Step 7:

[1618] The server generates additional advice and warnings based on food safety information and emotion data. Based on the obtained data, it generates appropriate messages and adds advice to alleviate the user's anxiety. The input is food safety information and emotion data, and the output is the generated advice or warning message.

[1619] Step 8:

[1620] The server generates advice and warning messages and sends them to the user's terminal. The input is the generated message and the output is the message sent to the user's terminal.

[1621] Step 9:

[1622] The terminal analyzes the received result data and displays it on the user interface. Based on this information, the user can decide whether to provide food to minors or animals. The input is the received message data, and the output is the information displayed on the user interface.

[1623] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1624] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1625] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1626] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1627] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1628] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1629] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1630] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1631] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1632] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1633] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1634] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1635] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1636] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1637] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1638] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1639] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1640] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1641] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1642] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1643] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1644] The following is further disclosed regarding the above embodiment.

[1645] (Claim 1)

[1646] A means for a user to take an image of food using a camera of the terminal and transmit the image to a server;

[1647] means for analyzing the image received by the server and identifying the type of food using an image recognition algorithm;

[1648] The server collects information about the identified food using a chat-type AI model and determines whether the food is safe for babies and pets to consume;

[1649] a means for the server to transmit the determination result to a terminal of the user and display the result on the terminal;

[1650] A system including:

[1651] (Claim 2)

[1652] The system according to claim 1, further comprising means for providing different information depending on whether the determination result relates to a baby or a pet.

[1653] (Claim 3)

[1654] 10. The system of claim 1, wherein the image recognition algorithm uses a deep learning model.

[1655] (Claim 4)

[1656] 10. The system of claim 1, wherein the chat AI model uses natural language processing.

[1657] "Example 1"

[1658] (Claim 1)

[1659] A means for a user to take an image of food using a camera of the terminal and transmit the image to a server;

[1660] means for analyzing the image received by the server and identifying the type of food using an image recognition algorithm;

[1661] a means for the server to collect information about the identified food using a generative AI model and determine whether the food is safe for consumption by babies or animals;

[1662] a means for the server to transmit the determination result to a terminal of the user and display the result on the terminal;

[1663] A system including:

[1664] (Claim 2)

[1665] The system according to claim 1, further comprising means for providing different information depending on whether the determination result relates to a baby or an animal.

[1666] (Claim 3)

[1667] 10. The system of claim 1, wherein the image recognition algorithm uses a deep learning model.

[1668] "Application Example 1"

[1669] (Claim 1)

[1670] A means for a user to take an image of food using a camera of the terminal and transmit the image to a server;

[1671] means for analyzing the image received by the server and identifying the type of food using an image recognition algorithm;

[1672] The server collects information about the identified food using a chat-based generative AI model and determines whether the food is safe for babies and pets to consume.

[1673] In a food delivery application, a means for a user to take an image of food to be ordered and obtain safety information of ingredients using the image;

[1674] a means for the server to transmit the determination result to a terminal of the user and display the result on the terminal;

[1675] A system including:

[1676] (Claim 2)

[1677] The system according to claim 1, further comprising means for providing different information depending on whether the determination result relates to a baby or a pet.

[1678] (Claim 3)

[1679] 10. The system of claim 1, wherein the image recognition algorithm uses a deep learning model.

[1680] "Example 2: Combining Emotion Engines"

[1681] (Claim 1)

[1682] A means for a user to take an image of food using a camera of the terminal and transmit the image to a server;

[1683] means for analyzing facial expressions and voice of the user and acquiring emotion data;

[1684] means for analyzing the image received by the server and identifying the type of food using an image recognition algorithm;

[1685] a means for the server to collect information about the identified food using a generative AI model and determine whether the food is safe for babies or pets to consume;

[1686] a means for the server to transmit additional advice or warning based on the judgment result and emotion data to the user's terminal and display the result on the terminal;

[1687] A system including:

[1688] (Claim 2)

[1689] The system according to claim 1, further comprising means for providing different information depending on whether the determination result relates to a baby or a pet.

[1690] (Claim 3)

[1691] 10. The system of claim 1, wherein the image recognition algorithm uses a deep learning model.

[1692] "Application example 2 when combining emotion engines"

[1693] (Claim 1)

[1694] A means for a user to take an image of food using a camera of the terminal and transmit the image to an information processing device;

[1695] means for analyzing the image received by the information processing device and identifying the type of food using an image recognition algorithm;

[1696] a means for collecting information about the identified food using an interactive AI model by the information processing device and determining whether the food is safe for consumption by minors or animals;

[1697] a means for transmitting the determination result to a terminal of a user by the information processing device and displaying the result on the terminal;

[1698] means for the terminal to analyze emotions from facial expressions and voice of the user and transmit the emotions to an information processing device;

[1699] means for adding additional advice or warning based on the emotion data obtained by the information processing device;

[1700] A system including:

[1701] (Claim 2)

[1702] The system of claim 1 , further comprising means for providing different information depending on whether the determination result relates to a minor or an animal.

[1703] (Claim 3)

[1704] 10. The system of claim 1, wherein the image recognition algorithm uses a deep learning model. [Explanation of symbols]

[1705] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for a user to take an image of food using a camera of the terminal and transmit the image to a server; means for analyzing the image received by the server and identifying the type of food using an image recognition algorithm; The server collects information about the identified food using a chat-type AI model and determines whether the food is safe for babies and pets to consume; a means for the server to transmit the determination result to a terminal of the user and display the result on the terminal; A system including:

2. The system according to claim 1 , further comprising means for providing different information depending on whether the determination result relates to a baby or a pet.

3. The system of claim 1 , wherein the image recognition algorithm uses a deep learning model.

4. The system of claim 1 , wherein the chat-based AI model uses natural language processing.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A