Method and apparatus food recognition

KR103013901B1Active Publication Date: 2026-09-02KAKAO ENTERPRISE CORP +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
KR1020220103071
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-08-18
Publication Date
2026-09-02
Estimated Expiration
2042-08-18

Smart Images

  • Figure R1020220103071_ABST
    Figure R1020220103071_ABST
Patent Text Reader

Abstract

The present invention provides a food recognition method comprising the steps of: extracting a plurality of regions of interest containing each food from a food image containing a plurality of foods; obtaining a first feature vector containing overall feature information of the food image from the food image; obtaining a plurality of second feature vectors containing individual feature information of each food from the plurality of regions of interest; generating a plurality of third feature vectors by combining the first feature vector with each of the plurality of second feature vectors; and determining the names of a plurality of foods using the plurality of third feature vectors.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a food recognition method and device. Background Technology

[0003] With the recent increase in interest in maintaining a balanced diet for health management, various diet-related applications are offering meal management services. These applications suggest appropriate exercise and meal plans based on dietary records.

[0004] In addition, the diet management service also provides a diet camera feature that recognizes foods included in images and automatically inputs the name and calories of each food.

[0005] Figure 1 is a configuration diagram of a conventional food recognition device.

[0006] Referring to FIG. 1, a conventional food recognition device is configured to include a food detector (11) and a food classifier (12).

[0007] Here, the food detector (11) extracts multiple regions of interest (2) containing each food from a food image (1) containing multiple foods. For example, the food detector (11) can extract multiple regions of interest (2) containing seaweed, spinach, yogurt, seaweed soup, donggeurangtteok, and barley rice, respectively, from the food image (1).

[0008] Then, the food detector (11) crops each of the multiple regions of interest (2) and inputs them into the food classifier.

[0009] The food classifier (12) receives a cropped multiple interest regions (2) as input and analyzes the features of the multiple interest regions (2) to determine and classify the names of the foods (e.g., seaweed, spinach, yogurt, seaweed soup, meatballs, and barley rice).

[0010] However, the aforementioned conventional food recognition device determines and classifies the name of the food by considering only the characteristics of the individual interest region (2) without considering the relative relationship between multiple foods included in the food image, so it may incorrectly determine and classify the name of the food when the food has a similar shape or color (e.g., milk and makgeolli).

[0011] In addition, conventional food recognition devices have a problem in that they misidentify the food to be recognized as an adjacent food when multiple foods included in a food image are located too close together. The problem to be solved

[0013] In order to solve the problems of the conventional technology described above, the present invention aims to provide a food recognition method capable of relatively accurate food classification when classifying foods that are similar in shape and color.

[0014] In addition, the present invention aims to provide a food recognition method that can prevent the food being recognized from being mistaken for an adjacent food by concentrating the feature vector of the food being recognized to the central part of the region of interest.

[0015] The technical problems to be solved by the present invention are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which the present invention belongs from the description below. means of solving the problem

[0017] To solve the aforementioned problem, the present invention provides a food recognition method implemented by a food recognition device, comprising the steps of: extracting a plurality of regions of interest containing each food from a food image containing shape and color information of a plurality of foods; obtaining a first feature vector containing overall feature information of the food image from the food image; obtaining a plurality of second feature vectors containing individual feature information of each food from the plurality of regions of interest; generating a plurality of third feature vectors by combining the first feature vector with each of the plurality of second feature vectors; and determining what kind of food a plurality of foods is using the plurality of third feature vectors.

[0018] Here, the step of determining the names of multiple foods may be a step of generating correlation information representing the correlation between multiple foods using the total feature information included in the third feature vector, and determining the names of multiple foods using the correlation information and the individual feature information included in the third feature vector.

[0019] Additionally, the step of extracting multiple regions of interest may include the step of inputting a food image into a food detector, the step of the food detector detecting multiple foods in the food image, the step of generating multiple regions of interest such that each of the detected multiple foods is included, and the step of the food detector cropping and outputting the multiple regions of interest.

[0020] In addition, the step of detecting multiple foods may include the step of acquiring location information of multiple regions of interest.

[0021] Additionally, the step of cropping and outputting multiple regions of interest may be a step in which a food detector crops multiple regions of interest based on location information of multiple regions of interest and outputs them to multiple food feature extractors.

[0022] Additionally, the step of acquiring the second feature vector may include applying a focus map extraction function to the second feature vector to focus the second feature vector into a first region where the food to be recognized is located among a plurality of regions of interest.

[0023] Here, the concentration map extraction function may be a function in which multiple pixels are arranged in multiple rows and columns corresponding to the size of multiple regions of interest, and weights are set for multiple pixels.

[0024] In addition, the concentration map extraction function may set the weight of the first region higher than that of the second region other than the first region among the multiple regions of interest.

[0025] In addition, the present invention provides a food recognition device comprising: a food detector that extracts a plurality of regions of interest containing each food from a food image containing shape and color information of a plurality of foods and obtains a first feature vector containing overall feature information of the food image from the food image; a food feature extractor that obtains a plurality of second feature vectors containing individual feature information of each food from a plurality of regions of interest; a feature combiner that generates a plurality of third feature vectors by combining the first feature vectors with the plurality of second feature vectors respectively; and a food classifier that determines what kind of food a plurality of foods is using the plurality of third feature vectors. Effects of the invention

[0027] According to the present invention, when classifying foods with similar shapes and colors, the food is classified using correlations (contextual information) between multiple foods included in a food image, thus having the advantage of enabling relatively accurate food classification.

[0028] In addition, according to the present invention, the feature vector of the food to be recognized is concentrated in the central part of the region of interest, thereby preventing the food to be recognized from being mistaken for an adjacent food.

[0029] The effects obtainable from the present invention are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the description below. Brief explanation of the drawing

[0031] Figure 1 is a configuration diagram of a conventional food recognition device. FIG. 2 is a configuration diagram of a food recognition device according to an embodiment of the present invention. FIG. 3 is a diagram illustrating the effect of the centralization process of a food recognition device according to an embodiment of the present invention. FIG. 4 is a flowchart of a food recognition method according to an embodiment of the present invention. Specific details for implementing the invention

[0032] To fully understand the structure and effects of the present invention, preferred embodiments of the present invention are described with reference to the attached drawings. However, the present invention is not limited to the embodiments disclosed below, but can be implemented in various forms and various modifications can be made. The description of the embodiments is provided merely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention. In the attached drawings, components are depicted with their sizes enlarged compared to their actual size for convenience of explanation, and the proportions of each component may be exaggerated or reduced.

[0033] Terms such as 'first' and 'second' may be used to describe various components, but said components should not be limited by said terms. These terms may be used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, 'first component' may be named 'second component,' and similarly, 'second component' may be named 'first component.' Furthermore, singular expressions include plural expressions unless the context clearly indicates otherwise. Unless otherwise defined, terms used in the embodiments of the present invention may be interpreted in the sense commonly known to those skilled in the art.

[0035] FIG. 2 is a configuration diagram of a food recognition device according to an embodiment of the present invention.

[0036] A food recognition device according to an embodiment of the present invention is a device that recognizes the type or name of food included in a food image, and can be implemented by at least one of a user terminal and a server.

[0037] A user terminal may be a mobile terminal, desktop computer, laptop computer, and tablet computer used by the user, but is not limited thereto, and may be various electronic devices equipped with a camera, display, communication device, memory, etc.

[0038] Here, the meaning that a food recognition device according to an embodiment of the present invention is implemented in a user terminal means that after installing a food recognition application on the user terminal, the type or name of food in a food image is recognized using the installed food recognition application.

[0039] Specifically, when a food image is generated by photographing food using a camera equipped on a user terminal, a food recognition application installed on the user terminal can determine the type or name of the food included in the food image.

[0040] The server is a device that provides food recognition services to a user terminal and may be equipped with a processor, a database, a communication device, etc.

[0041] Here, the meaning that the food recognition device according to an embodiment of the present invention is implemented on a server means that after installing a food recognition processor on the server, the type or name of food in a food image is recognized using the installed food recognition processor.

[0042] Specifically, when a user terminal transmits a food image generated by the user terminal to a server, a food recognition processor installed on the server can determine the type or name of the food included in the received food image and transmit the determined type or name of the food to the user terminal.

[0043] Hereinafter, a food recognition device according to an embodiment of the present invention is described as being implemented in a user terminal as an example, but it is obvious that it can be implemented in a server as described above.

[0044] Referring to FIG. 2, a food recognition device according to an embodiment of the present invention may be configured to include a food detector (110), a food feature extractor (121, 122), a feature combiner (131, 132), and a food classifier (141, 142).

[0045] A camera equipped in a user terminal can photograph food to generate a food image (10), and the generated food image (10) can be stored in memory. Here, the food image (10) may include a plurality of foods. For example, the food image (10) may include a sandwich and milk.

[0046] The food detector (110) can extract multiple regions of interest (21, 22) containing each food from a food image (10) containing shape and color information of multiple foods. In the example described above, the region of interest (21) may contain a sandwich, and the region of interest (22) may contain milk.

[0047] Specifically, the food detector (110) can receive a food image (10) from a camera or memory and detect multiple foods in the food image (10). Meanwhile, the food detector (110) can detect whether an object included in the food image (10) is food, but may not detect the type or name of the food.

[0048] The food detector (110) may have a first artificial neural network generated by learning a plurality of food images and a plurality of regions of interest containing food.

[0049] Here, the first Artificial Neural Network (ANN) is a statistical learning algorithm in machine learning and cognitive science inspired by biological neural networks (specifically the brain within the animal central nervous system). The first Artificial Neural Network refers to a model in general that possesses problem-solving capabilities by having artificial neurons, which form a network through synaptic connections, change the strength of these connections through learning. For example, the first Artificial Neural Network can be generated by training based on a Convolutional Neural Network (CNN).

[0050] The first artificial neural network may be composed of an input layer into which a food image (10) is input, an output layer into which a region of interest (21, 22) is output, and a plurality of hidden layers existing between the input layer and the output layer.

[0051] The food detector (110) can generate multiple regions of interest (21, 22) such that each of the detected foods is included, and can crop each of the multiple regions of interest (21, 22) from the food image (10) and output them.

[0052] Additionally, the food detector (110) can acquire location information of a plurality of regions of interest (21, 22). Here, the location information of the regions of interest (21, 22) may include the center coordinates, width, and height of the regions of interest (21, 22).

[0053] The food detector (110) can crop the multiple regions of interest (21, 22) in the food image (10) based on the location information of the multiple regions of interest (21, 22) and output them to the multiple food feature extractors (121, 122). In the example described above, the food detector (110) can output a first region of interest (21) including a sandwich to the first food feature extractor (121) and output a second region of interest (22) including milk to the second food feature extractor (122).

[0054] The food detector (110) can obtain a first feature vector (30) containing global feature information of the food image (10) from the food image (10). Here, the global feature information may include information such as the shape and color of a plurality of foods included in the food image (10).

[0055] The first feature vector (30) can be used to generate correlation information (context information) between multiple foods included in the food image (10). Here, the correlation information refers to information indicating the correlation between multiple foods that go well together when consumed, and by using this correlation information, it is possible to determine which food each food is included with in the food image (10).

[0056] The food detector (110) can extract and use various features such as edge features, corner features, or LoG (Laplacian of Gaussian) and DoG (Difference of Gaussian) to detect all features in the food image (10), and can also use various existing feature description methods including SIFT (Scale-invariant feature transform), SURF (Speeded Up Robust Features), and HOG (Histogram of Oriented Gradients).

[0057] To this end, the food detector (110) may be provided with a first artificial neural network generated by learning a plurality of food images and the total feature information (30) of a plurality of foods included in the plurality of food images. For example, the first artificial neural network may be generated by learning based on a Convolutional Neural Network (CNN).

[0058] The first artificial neural network may be composed of an input layer into which a food image (10) is input, an output layer into which all feature information of the food image (10) is output, and a plurality of hidden layers existing between the input layer and the output layer.

[0059] The food detector (110) can output a first feature vector (30) containing all feature information of the food image (10) to a feature combiner (131, 132).

[0060] The food feature extractors (121, 122) may be provided in multiple numbers. Here, a plurality of regions of interest (21, 22) cropped by the food detector (110) may each be input to any one of the plurality of food feature extractors (121, 122).

[0061] The food feature extractor (121, 122) may be provided with a plurality of interest regions (21, 22) containing food and a second artificial neural network generated by learning the features of the food included in each of the plurality of interest regions (21, 22). For example, the second artificial neural network may be generated by learning based on a Convolutional Neural Network (CNN).

[0062] The second artificial neural network may be composed of an input layer into which a region of interest (21, 22) is input, an output layer that outputs individual feature information of food included in the region of interest (21, 22), and a plurality of hidden layers existing between the input layer and the output layer.

[0063] The food feature extractor (121, 122) can obtain a plurality of second feature vectors (41, 42) containing local feature information of each food in a plurality of regions of interest (21, 22) using the aforementioned second artificial neural network.

[0064] Here, individual feature information may include information such as the color and shape of each food, and the second feature vector (41, 42) may be used to determine the type or name of the food included in the food image (10). In the example described above, the first food feature extractor (121) may extract features of a sandwich from the first region of interest (21), and the second food feature extractor (122) may extract features of milk from the second region of interest (22).

[0065] The food feature extractor (121, 122) can extract and use various features such as edge features, corner features, LoG (Laplacian of Gaussian), DoG (Difference of Gaussian) to detect individual features of food included in each region of interest (21, 22), and can also use various existing feature description methods including SIFT (Scale-invariant feature transform), SURF (Speeded Up Robust Features), and HOG (Histogram of Oriented Gradients).

[0066] The feature combiner (131, 132) can receive a first feature vector (30) from the food detector (110) and receive a plurality of second feature vectors (41, 42) from the food feature extractor (121, 122).

[0067] The feature combiner (131, 132) can combine the first feature vector (30) with a plurality of second feature vectors (41, 42) respectively to generate a plurality of third feature vectors (51, 52). Here, the third feature vectors (51, 52) may include overall feature information of the food image (10) and individual feature information of each food.

[0068] In the example described above, the first feature combiner (131) combines the first feature vector (30) representing the correlation between the sandwich and the milk with the second feature vector (41) representing the individual features of the sandwich, and the second feature combiner (132) combines the first feature vector (30) representing the correlation between the sandwich and the milk with the second feature vector (42) representing the individual features of the milk.

[0069] The feature combiner (131, 132) can output a plurality of third feature vectors to the food classifier (141, 142).

[0070] The food classifier (141, 142) can determine the types or names of multiple foods using multiple third feature vectors (51, 52) each received from the feature combining unit (131, 132).

[0071] Specifically, the food classifier (141, 142) can generate correlation information indicating correlations between multiple foods using the total feature information included in the third feature vector (51, 52), and can determine the types or names of multiple foods using the correlation information and the individual feature information included in the third feature vector (51, 52).

[0072] For example, the first food classifier (141) can generate correlation information indicating the correlation between sandwiches and milk, and can determine the name of the food as sandwiches using the correlation information and the individual feature information of the sandwiches. Here, the first food classifier (141) can determine the name of the food as sandwiches using only the individual feature information if there are no similar foods similar in shape and color to the sandwiches, but if there are multiple similar foods similar in shape and color to the sandwiches, it can determine the name of the food by using the total feature information to confirm that milk is included together in the food image, and among the multiple similar foods, the sandwich that is correlated with milk can be determined as the name of the food.

[0073] In particular, the second food classifier (142) generates correlation information indicating the correlation between sandwiches and milk, and can determine the name of the food as milk using the correlation information and individual characteristic information of milk.

[0074] In other words, the second food classifier (142) can accurately classify milk and makgeolli that are similar in shape and color by utilizing the correlation between sandwiches and milk. Here, the second food classifier (142) cannot accurately distinguish between milk and makgeolli based solely on individual feature information because there is a similar food, namely makgeolli, that is similar in shape and color to milk. In this case, the second food classifier (142) can determine the name of the food by using the overall feature information to confirm that the sandwich is included together in the food image, thereby determining the milk that is correlated with the sandwich among milk and makgeolli. Accordingly, it is possible to prevent misidentifying milk as makgeolli that is similar in color and shape.

[0075] Thus, the food recognition device according to the embodiment of the present invention classifies food by utilizing the correlation between a plurality of foods included in the food image (10) when classifying foods that are similar in shape and color, thereby enabling relatively accurate food classification.

[0076] The food recognition process of the food recognition device according to an embodiment of the present invention can be expressed by the following mathematical formulas.

[0077] A first feature vector (including all feature information acquired by the food detector (110) ) can be defined by the following mathematical formula 1.

[0078] [Mathematical Formula 1]

[0079]

[0080] Here, is a food image, and is a function for extracting a first feature vector, and means an artificial neural network equipped in a food detector (110).

[0081] According to the above mathematical formula 1, food images ( If you input ), the first feature vector ( ) is extracted.

[0082] A second feature vector (including individual feature information obtained by a food feature extractor (121, 122)) ) can be defined by the following mathematical formula 2.

[0083] [Mathematical Formula 2]

[0084]

[0085] Here, is food image( It refers to the i-th region of interest (where i is an integer greater than or equal to 2) in ), and b i represents the coordinates and size of the i-th region of interest. And, L( ) is a function for extracting the second feature vector, and represents an artificial neural network equipped in the food feature extractor (121, 122).

[0086] b in mathematical equation 2 i It can be defined by the following mathematical formula 3.

[0087] [Mathematical Formula 3]

[0088]

[0089] Here, x i and y i is the center coordinate of the region of interest, and w i and h i represents the width and height of the region of interest.

[0090] According to the above mathematical formulas 2 and 3, if the coordinates and size of the region of interest are input into L( ), the second feature vector ( ) is extracted.

[0091] A third feature vector (f) formed by combining the first feature vector and the second feature vector. i ) can be defined by the following mathematical formula 4.

[0092] [Mathematical Formula 4]

[0093]

[0094] Here, concatenate() is a function for combining the first feature vector and the second feature vector.

[0095] The type or name of food classified by the food classifier (141, 142) (c i ) can be defined by the following mathematical formula 5.

[0096] [Mathematical Formula 5]

[0097]

[0098] Here, Classify() is a function for classifying the types or names of food.

[0099] Combining the above mathematical formulas 1 to 5, the type or name of food classified by the food classifier (141, 142) is c i ) can be simplified by the following mathematical formula 6.

[0100] [Mathematical Formula 6]

[0101]

[0102] Meanwhile, during the process of the food detector (110) extracting the region of interest (21, 22), a part of another food may be included in the extracted region of interest (21, 22). In this case, the food feature extractor (121, 122) extracts even the features of the other food that are partially included in the region of interest (21, 22).

[0103] Accordingly, the food classifier (141, 142) may incorrectly determine and classify the type or name of food included in the area of ​​interest (21, 22).

[0104] To solve this problem, the food feature extraction unit (121, 122) applies a concentration map extraction function to the second feature vector to concentrate the second feature vector into the central part of a plurality of regions of interest, that is, the first region where the food to be recognized is located, thereby preventing the food to be recognized from being mistaken for an adjacent food. This utilizes the principle that the food to be recognized is mainly placed in the central part of the region of interest, and some of the adjacent food is mainly placed in the peripheral part of the region of interest, that is, the second region other than the first region, thereby further highlighting the features of the region of interest while minimizing the influence of the features of the peripheral part of the region of interest.

[0105] Here, the concentration map extraction function is a function in which multiple pixels are arranged in multiple rows and columns corresponding to the size of multiple regions of interest, and weights are set for multiple pixels.

[0106] As described above, when a portion of other food is included in the region of interest (21, 22), since the portion of other food is mainly included in the periphery of the region of interest (21, 22), it is desirable for the concentration map extraction function to set the weight of the central portion (first region) of the multiple regions of interest higher than the periphery (second region) of the multiple regions of interest.

[0107] Such a focal map extraction function can be implemented as one of the hidden layers of an artificial neural network equipped in a food feature extractor (121, 122). Here, the hidden layer in which the focal map extraction function is implemented will be referred to as the focal layer.

[0108] A second feature vector generated through a centralization process of a food recognition device according to an embodiment of the present invention ( ) can be expressed by the following mathematical formula 7.

[0109] [Mathematical Formula 7]

[0110]

[0111] Here, M() is a function for extracting the concentration map.

[0112] FIG. 3 is a diagram illustrating the effect of the centralization process of a food recognition device according to an embodiment of the present invention.

[0113] Referring to FIG. 3, after extracting regions of interest (21, 22, 23) containing each food from a food image (10) containing multiple foods, the second feature vector was concentrated in the central part (first region) and displayed as a heatmap.

[0114] The more red the area of ​​interest (21, 22, 23) is displayed, the better the concentration is.

[0115] FIG. 3 (a) shows a case where the focal layer is provided in the 3rd layer of the 12 hidden layers of an artificial neural network, FIG. 3 (b) shows a case where the focal layer is provided in the 7th layer of the hidden layers of an artificial neural network, FIG. 3 (c) shows a case where the focal layer is provided in the 9th layer of the hidden layers of an artificial neural network, and FIG. 3 (d) shows a case where the focal layer is provided in the 11th layer of the hidden layers of an artificial neural network.

[0116] As shown in FIG. 3, it can be seen that the central part (first region) of the region of interest (21, 22, 23) in FIG. 3(d) is marked in relatively more red, which means that the concentration layer is provided in the latter part of the hidden layers, which has a higher concentration.

[0117] Therefore, it is preferable to place the concentration layer in the latter part of the hidden layers of the artificial neural network equipped in the food feature extractor (121, 122).

[0118] FIG. 4 is a flowchart of a food recognition method according to an embodiment of the present invention.

[0119] Hereinafter, with reference to FIGS. 2 to 4, a food recognition method according to an embodiment of the present invention will be described, but details identical to those described above will be omitted.

[0120] First, a plurality of regions of interest (21, 22) containing each food are extracted from a food image (10) containing a plurality of foods (S10).

[0121] At this time, multiple regions of interest (21, 22) can be generated so that each of the detected multiple foods is included, and each of the multiple regions of interest (21, 22) can be cropped from the food image (10) and output.

[0122] Next, a first feature vector containing all feature information of the food image (10) is obtained from the food image (10) (S20). Here, the first feature vector can be used to generate correlation information (context information) between a plurality of foods included in the food image (10).

[0123] Next, a plurality of second feature vectors containing individual feature information of each food in a plurality of regions of interest (21, 22) are obtained using an artificial neural network (S30).

[0124] Here, individual feature information refers to information such as the color and shape of each food, and the second feature vector can be used to determine the type or name of the food included in the food image (10).

[0125] Next, a plurality of third feature vectors are generated by combining a first feature vector with a plurality of second feature vectors, respectively (S40).

[0126] Next, multiple types or names of multiple foods are determined using multiple third feature vectors (S50).

[0127] Specifically, correlation information representing the correlation between multiple foods can be generated using the total feature information included in the third feature vector, and the types or names of multiple foods can be determined using the correlation information and the individual feature information included in the third feature vector.

[0128] As such, the food recognition method according to the embodiment of the present invention classifies food by utilizing the correlation between a plurality of foods included in the food image (10) when classifying foods that are similar in shape and color, thereby enabling relatively accurate food classification.

[0130] Although specific embodiments have been described in the detailed description of the present invention, it is understood that various modifications are possible within the scope of the invention. Therefore, the scope of the present invention is not limited to the described embodiments and should be defined by the claims set forth below and equivalents thereof. Explanation of the symbols

[0132] 110: Food detector 121, 122: Food characteristic extractor 131, 132: Feature combiner 141, 142: Food classifier

Claims

Claim 1 A method for recognizing food implemented by a food recognition device, comprising: a step of extracting a plurality of regions of interest containing each food from a food image containing shape and color information of a plurality of foods; a step of obtaining a first feature vector containing overall feature information of the food image from the food image; a step of obtaining a plurality of second feature vectors containing individual feature information of each food from the plurality of regions of interest; and a step of determining what kind of food the plurality of foods are using the first feature vector and the plurality of second feature vectors, wherein the step of obtaining the second feature vector includes a step of applying a focus map extraction function to the second feature vector to focus the second feature vector to a first region where the food to be recognized is located among the plurality of regions of interest, and wherein the focus map extraction function is implemented as a focus layer, which is one of the hidden layers of an artificial neural network provided in a food feature extractor that obtains the second feature vector before applying the focus map extraction function, and wherein the focus layer is placed in the latter part of the hidden layers. Claim 2 A food recognition method according to claim 1, further comprising the step of generating a plurality of third feature vectors by combining the plurality of second feature vectors to which the first feature vector and the concentration map extraction function are applied, and the step of determining what kind of food the plurality of foods are is a step of determining what kind of food the plurality of foods are by using the plurality of third vectors. Claim 3 A food recognition method according to claim 2, wherein the step of determining the names of the plurality of foods is to generate correlation information representing the correlation between the plurality of foods using the total feature information included in the third feature vector, and to determine the names of the plurality of foods using the correlation information and the individual feature information included in the third feature vector. Claim 4 A food recognition method according to claim 1, wherein the step of extracting a plurality of regions of interest comprises: a step of detecting a plurality of foods in a food image; and a step of generating a plurality of regions of interest such that each of the detected plurality of foods is included. Claim 5 In claim 4, the step of detecting a plurality of foods includes the step of obtaining location information of the plurality of regions of interest. Claim 6 In claim 5, the step of extracting the plurality of regions of interest is a step of cropping the plurality of regions of interest based on location information of the plurality of regions of interest. Claim 7 delete Claim 8 A food recognition method according to claim 1, wherein the concentration map extraction function is a function in which a plurality of pixels are arranged in a plurality of rows and columns corresponding to the size of the plurality of regions of interest, and weights are set for the plurality of pixels. Claim 9 In claim 8, the concentration map extraction function is a food recognition method in which the first region is set to have a higher weight than the second region other than the first region among the plurality of regions of interest. Claim 10 A food recognition device comprising: a food detector that extracts multiple regions of interest containing each food from a food image containing shape and color information of multiple foods, and obtains a first feature vector containing overall feature information of the food image from the food image; a food feature extractor that obtains multiple second feature vectors containing individual feature information of each food from the multiple regions of interest; and a food classifier that determines what kind of food the multiple foods are using the first feature vector and the multiple second feature vectors, wherein the food feature extractor applies a focus map extraction function to the second feature vector to focus the second feature vector to the region where the food to be recognized is located among the multiple regions of interest, and the focus map extraction function is implemented as a focus layer which is one of the hidden layers of an artificial neural network provided in the food feature extractor, and the focus layer is placed in the latter part of the hidden layers.

Citation Information

Patent Citations

  • System and method for nutritional analysis using food image recognition

    JP2018528545A

  • Information processor, information processing method, and program

    US20130058566A1