Image-based dish recognition device and method

Through the image-based dish recognition method, the image is analyzed and the similarity of dish identifiers is calculated using the prediction model, and the problem of inaccurate dish recognition in the prior art is solved, and the dish recognition with high recall rate and high specificity is achieved, which improves the user experience.

CN113892110BActive Publication Date: 2025-05-09VERSUNI HLDG BV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080025081.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-02
Filing Date
2020-03-20
Publication Date
2025-05-09
Estimated Expiration
2040-03-20

AI Technical Summary

Technical Problem

The existing dietary recording system requires users to input and record consumed dishes in large quantities, and due to the overlap of category definitions and the overlap of training data, machine learning models have confusion and inaccurate classification in dish recognition, affecting the user experience.

Method used

Using an image-based dish recognition method, by obtaining images describing the dishes to be identified, using a prediction model to analyze the images to determine the candidate topic, calculating the similarity between the dish identifier and the candidate topic, selecting the dish identifier with the highest correlation score as the variant dish identifier, and outputting the heart-shaped dish identifier and the variant dish identifier.

Benefits of technology

It realizes high recall and high specificity of dishes, reduces the user input needs, improves the user-friendliness and intuitiveness of the system, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113892110B_ABST
    Figure CN113892110B_ABST
Patent Text Reader

Abstract

A computer-implemented method for performing image-based dish recognition is provided. The method includes: obtaining (202) a first image depicting a dish to be recognized; analyzing (204) the first image using a predictive model to determine a first candidate theme; obtaining (206) a reference set of dish identifiers; for each dish identifier in the reference set, calculating (208) an association score indicating a similarity between the dish represented by the corresponding dish identifier and the first candidate theme; selecting (210) one or more dish identifiers in the reference set having the highest association score as one or more variant dish identifiers of the first candidate theme; and outputting (212) a centroid dish identifier of the first candidate theme and one or more variant dish identifiers of the first candidate theme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image-based dish recognition device and method. Background Art

[0002] There are many food logging systems currently available on the market that can be used to monitor an individual's nutritional intake. These systems typically rely on receiving user input to log the food consumed, a process that requires a significant amount of human interaction and a certain degree of competence and motivation from the user. To achieve the desired level of effectiveness, the user must accurately and persistently log the dishes consumed. Summary of the invention

[0003] As mentioned above, users of food logging systems need to accurately and persistently input and / or log the dishes they consume. Therefore, it is very important to provide a user-friendly and intuitive logging tool to increase the likelihood of the user's ability to stick to the daily habit of inputting and / or logging the food consumed. One of the key factors in achieving user-friendliness and intuitiveness is the ability of the system to correctly respond to the user's intent within a minimum number of user interaction steps.

[0004] To this end, machine learning and deep learning processes can be used to improve the understanding of user intent while reducing the amount of user input and / or recorded dish intervention required. Although machine learning processes (such as processes involving neural networks) can be used to extract patterns that are common in a category and different from other categories from a large amount of labeled training data, for many foods, such as Chinese cooking dishes, there is a lack of clear definitions or specifications associated with the food, which may lead to category specification problems. For example, some dishes may be composed of similar food ingredients and / or have similar appearances but different names, and some dishes with similar names may have significantly different food ingredients. For this reason, it is usually difficult to specify a set of mutually exclusive categories with fine granularity, and it is difficult to collect exemplary data based on such a set. Overlap between category definitions or actual overlap of training data may produce confusion in machine learning and lead to inaccurate classification, which may cause frustration to users in use. However, if the category definitions differ greatly, and only highly representative data are collected to avoid overlap, high specificity and low sensitivity in dish recognition will result - this is undesirable because the system will not be suitable for identifying the rich and diverse dishes in real life. Therefore, it would be advantageous to provide a user-friendly image-based dish recognition method that can accurately recognize user input (i.e., high recall) and is also able to respond to changes in user input (i.e., high specificity).

[0005] In order to better solve one or more of the above-mentioned problems, in the first aspect, a computer-implemented method for image-based dish recognition is provided. The method includes: obtaining a first image depicting a dish to be recognized; analyzing the first image using a prediction model to determine a first candidate theme, wherein the first candidate theme includes a plurality of candidate dish identifiers, each candidate dish identifier is associated with a candidate dish, and one of the candidate dish identifiers is a centroid dish identifier associated with a candidate dish that best represents the first candidate theme; obtaining a reference set of dish identifiers; for each dish identifier in the reference set, calculating an association score indicating the similarity between the dish represented by the corresponding dish identifier and the first candidate theme; selecting one or more dish identifiers in the reference set with the highest association score as one or more variant dish identifiers of the first candidate theme; and outputting the centroid dish identifier of the first candidate theme and one or more variant dish identifiers of the first candidate theme.

[0006] In some embodiments, outputting the centroid dish identifier of the first candidate theme and one or more variant dish identifiers of the first candidate theme includes: displaying the first candidate theme; and upon receiving user input to expand the first candidate theme, displaying the centroid dish identifier of the first candidate theme and one or more variant dish identifiers of the first candidate theme. In these embodiments, the centroid dish identifier of the first candidate theme may be displayed above the one or more variant dish identifiers of the first candidate theme, and the one or more variant dish identifiers of the first candidate theme may be displayed in descending order according to the corresponding relevance scores.

[0007] In some embodiments, multiple candidate dish identifiers in the first candidate theme may represent dishes that are similar to each other.

[0008] In some embodiments, the method may further include: analyzing the first image using the predictive model to determine one or more additional candidate topics. In these embodiments, each additional candidate topic may include a plurality of candidate dish identifiers, each candidate dish identifier is associated with a candidate dish, and one of the candidate dish identifiers is a centroid dish identifier that best represents the corresponding candidate topic. In addition, the method may further include outputting the one or more additional candidate topics.

[0009] In some embodiments, the method may further include performing the following method steps for each of one or more additional candidate topics: for each dish identifier in the reference set, calculating an association score indicating the similarity between the dish represented by the corresponding dish identifier and the corresponding additional candidate topic; selecting one or more dish identifiers in the reference set with the highest association scores as one or more variant dish identifiers of the corresponding additional candidate topic; and outputting the centroid dish identifier of the corresponding additional candidate topic and one or more variant dish identifiers of the corresponding additional candidate topic.

[0010] In some embodiments, the method may further include determining a ranking of the first candidate topic and one or more additional candidate topics using the predictive model. In these embodiments, the ranking may indicate a decreasing similarity between the centroid dish of the corresponding candidate topic and the dish depicted by the obtained first image. The method may further include displaying the first candidate topic and one or more additional candidate topics based on the determined ranking.

[0011] In some embodiments, the method may further include: obtaining a plurality of recipes, wherein each of the plurality of recipes includes: a dish identifier, a plurality of food ingredients, and one or more cooking instructions; selecting a core subset of recipes from the obtained plurality of recipes; calculating a similarity score between recipes in the core subset based on at least one of the following: a similarity between dish identifiers of two recipes, a similarity between food ingredients of the two recipes, and a similarity between cooking instructions of the two recipes; clustering the plurality of recipes into a plurality of reference topics based on the similarity scores of the plurality of recipes; and for each of the plurality of reference topics, selecting a recipe having the highest cosine similarity with the corresponding reference topic as a centroid recipe, wherein the dish identifier of the selected recipe is a centroid dish identifier of the corresponding reference topic. In these embodiments, determining the first candidate topic may include selecting the first candidate topic from the plurality of reference topics.

[0012] In some embodiments, the method may further include determining a popularity score for each of the plurality of recipes, wherein the popularity score indicates popularity or prevalence of the corresponding recipe. In these embodiments, selecting the core subset of recipes may be performed based on the popularity scores of the plurality of recipes.

[0013] In some embodiments, calculating the similarity score between the recipes in the core subset may be performed based on one or more synonyms of at least one of: a dish identifier, a plurality of food ingredients, and cooking instructions of the two recipes.

[0014] In some embodiments, clustering the plurality of recipes into a plurality of reference topics may be performed based on K-means clustering or singular value decomposition.

[0015] In some embodiments, the method may further include determining, for each of the plurality of reference topics, a plurality of keywords based on recipes of the corresponding reference topic. In these embodiments, each of the plurality of keywords is associated with at least one of a cooking technique and a food ingredient;

[0016] In some embodiments, the method may further include: selecting one of a plurality of reference themes; acquiring a second image based on a centroid dish identifier of the selected reference theme and at least one of a plurality of keywords of the selected reference theme; and training a prediction model based on the second image and the selected reference theme.

[0017] In some embodiments, the prediction model may be at least one of: a convolutional neural network, a residual neural network, and a dense neural network.

[0018] In a second aspect, a computer program product is provided, comprising a computer readable medium, wherein the computer readable medium contains computer readable code, which, when executed on a suitable computer or processor, causes the computer or processor to perform the method as described herein.

[0019] In a third aspect, an image-based dish recognition device is provided. The device includes a processor, which is configured to: obtain a first image depicting a dish to be recognized; analyze the first image using a prediction model to determine a first candidate theme, wherein the first candidate theme includes a plurality of candidate dish identifiers, each candidate dish identifier is associated with a candidate dish, and one of the candidate dish identifiers is a centroid dish identifier associated with a candidate dish that best represents the first candidate theme; obtain a reference set of dish identifiers; for each dish identifier in the reference set, calculate an association score indicating the similarity between the dish represented by the corresponding dish identifier and the first candidate theme; select one or more dish identifiers in the reference set with the highest association score as one or more variant dish identifiers of the first candidate theme; and output the centroid dish identifier of the first candidate theme and the one or more variant dish identifiers of the first candidate theme.

[0020] According to the above aspects and embodiments, the limitations of the prior art are solved. In particular, the above aspects and embodiments enable high recall and high specificity in image-based dish recognition. The above embodiments achieve high sensitivity by performing image recognition at a coarse (e.g., topic) granularity, and achieve high specificity by performing a topic variant search based on natural language processing (NLP).

[0021] Therefore, an improved method and apparatus for image-based dish recognition is provided. These and other aspects of the disclosure will become apparent and elucidated with reference to the embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] For a better understanding of the embodiments, and to more clearly show how they may be implemented, reference is now made, by way of example only, to the accompanying drawings, in which:

[0023] Figure 1 is a block diagram of an image-based dish recognition apparatus according to one embodiment;

[0024] Figure 2 An image-based dish recognition method according to one embodiment is shown;

[0025] Figure 3A showing an example of a first candidate topic and a plurality of additional candidate topics outputted in a first state on a display according to one embodiment; and

[0026] Figure 3B The second state is shown on the display. Figure 3A A first candidate topic and multiple additional candidate topics. DETAILED DESCRIPTION

[0027] As described above, an improved device and method of operating the same are provided that solve the existing problems.

[0028] Figure 1 A block diagram of an apparatus 100 according to one embodiment is shown, which can be used to perform image-based dish recognition. Although the operation of the apparatus 100 is described below in the context of performing dish recognition for a single image, it should be understood that the apparatus 100 can perform image-based dish recognition for each of a plurality of images.

[0029] like Figure 1 As shown, the device includes a processor 102 for controlling the operation of the device 100 and can implement the methods described herein. The processor 102 may include one or more processors, processing units, multi-core processors or modules configured or programmed to control the device 100 in the manner described herein. In a specific embodiment, the processor 102 may include multiple software and / or hardware modules, each of which is configured to perform or is used to perform each or multiple steps of the method described herein.

[0030] In brief, the processor 102 is configured to obtain a first image depicting a dish to be identified, and analyze the first image using a predictive model to determine a first candidate theme. The first candidate theme includes a plurality of dish identifiers, each dish identifier is associated with a candidate dish, and one of the candidate dish identifiers is a centroid dish identifier associated with a candidate dish that best represents the first candidate theme.

[0031] The processor 102 is further configured to obtain a reference set of dish identifiers, and for each dish identifier in the reference set, calculate an association score indicating a similarity between the dish represented by the corresponding dish identifier and the first candidate theme. Subsequently, the processor 102 is configured to select one or more dish identifiers in the reference set with the highest association score as one or more variant dish identifiers of the first candidate theme, and output a centroid dish identifier of the first candidate theme and one or more variant dish identifiers of the first candidate theme.

[0032] In some embodiments, the device 100 may also include at least one user interface 104. Alternatively or additionally, at least one user interface 104 may be located outside the device 100 (i.e., separated from or away from the device 100). For example, at least one user interface 104 may be part of another device. The user interface 104 may be used to provide information obtained from the method described herein to the user of the device 100. Alternatively or additionally, the user interface 104 may be configured to receive user input. For example, the user interface 104 may allow the user of the device 100 to manually input instructions, data, or information. In these embodiments, the processor 102 may be configured to obtain user input from one or more user interfaces 104.

[0033] The user interface 104 may be any user interface capable of presenting (or outputting or displaying) information to a user of the device 100. Alternatively or additionally, the user interface 104 may be any user interface that enables a user of the device 100 to provide user input, interact with the device 100, and / or control the device 100. For example, the user interface 104 may include one or more switches, one or more buttons, a keypad, a touch screen or application (e.g., on a tablet or smart phone), a display screen, a graphical user interface (GUI) or other visual rendering component, one or more speakers, one or more microphones or any other audio component, one or more lights, a component that provides tactile feedback (e.g., a vibration function), or any other user interface or combination of user interfaces.

[0034] In some embodiments, the apparatus 100 may include a memory 106. Alternatively or additionally, one or more memories 106 may be located external to the apparatus 100 (i.e., separate from or remote from the apparatus 100). For example, one or more memories 106 may be part of another device. The memory 106 may be configured to store program code that may be executed by the processor 102 to perform the methods described herein. The memory may be used to store information, data, signals, and measurements acquired or created by the processor 102 of the apparatus 100. For example, the memory 106 may be used to store (e.g., in a local file) a reference set of dish identifiers. The processor 102 may be configured to control the memory 106 to store a reference set of dish identifiers.

[0035] In some embodiments, the device 100 may include a communication interface (or circuit) 108 for enabling the device 100 to communicate with any interface, memory, and / or device inside or outside the device 100. The communication interface 108 can communicate with any interface, memory, and / or device wirelessly or via a wired connection. For example, the communication interface 108 can communicate with one or more user interfaces 104 wirelessly or via a wired connection. Similarly, the communication interface 108 can communicate with one or more memories 106 wirelessly or via a wired connection.

[0036] It should be understood that Figure 1 Only components necessary to illustrate one aspect of the apparatus 100 are shown, and in actual implementation, the apparatus 100 may include components alternative to or in addition to those shown.

[0037] Figure 2 A computer-implemented method for image-based dish recognition according to one embodiment is shown. The method shown can generally be executed by the processor 102 of the device 100 or executed under the control of the processor 102.

[0038] See also Figure 2 , at block 202, a first image depicting a dish to be identified is acquired. More specifically, the first image may be acquired by the processor 102 of the device 100. The first image may be acquired from the memory 106 of the device 100, or from an external memory or database. For example, the first image may be acquired from a memory of an imaging device (e.g., a camera) external to the device 100. The user may take a photo of the dish before eating, and the photo may then be transmitted to the processor 102 of the device 100 to perform dish identification.

[0039] return Figure 2, at box 204, the first image obtained at box 202 is analyzed using a prediction model to determine a first candidate topic. More specifically, the first image can be acquired by the processor 102 of the device 100. The first candidate topic includes multiple candidate dish identifiers, each candidate dish identifier is associated with a candidate dish, and one of the candidate dish identifiers is a centroid dish identifier associated with a candidate dish that best represents the first candidate topic. In some embodiments, the multiple candidate dish identifiers in the first candidate topic can represent dishes that are similar to each other. In some embodiments, the prediction model can be at least one of the following: a convolutional neural network, a residual neural network, and a dense neural network.

[0040] As will be described in further detail below, determining the first candidate topic at block 204 may include selecting the first candidate topic from a plurality of reference topics using a predictive model. In more detail, among the plurality of reference topics, the processor 102 may be configured to select a topic having a candidate dish that is most likely the dish depicted in the first image. Additionally or alternatively, the processor 102 may be configured to select a topic having a centroid dish that is most likely the dish depicted in the first image.

[0041] return Figure 2 , at block 206, a reference set of dish identifiers is obtained. More specifically, the reference set of dish identifiers may be obtained by the processor 102 of the apparatus 100. As will be described in more detail below, in some embodiments, the method may further include the step of obtaining a plurality of recipes, wherein each of the plurality of recipes includes: a dish identifier, a plurality of food ingredients, and one or more cooking instructions. In these embodiments, the reference set of dish identifiers may be obtained from a plurality of obtained recipes, i.e., the dish identifiers may be extracted from at least some of the obtained recipes.

[0042] return Figure 2 At block 208, a relevance score is calculated for each dish identifier in the reference set obtained in block 206. More specifically, the relevance score of each dish identifier in the reference set may be calculated by the processor 102 of the apparatus 100. The relevance score of the dish identifier indicates the similarity between the dish represented by the corresponding dish identifier and the first candidate theme.

[0043] In some embodiments, the association score for each dish identifier in the reference set can be calculated based on a cosine similarity metric. In these embodiments, the method may also include determining a vector for each dish identifier in the reference set based on the relative frequency of occurrence of one or more semantic keywords (e.g., the dish identifier itself and its aliases / synonyms, food ingredients, and instructions) in the recipes associated with the corresponding dish identifier. The method may also include determining a vector for the first candidate topic based on the relative frequency of occurrence of one or more semantic keywords in the recipes associated with the candidate dish identifiers in the first candidate topic. Alternatively, the method may include determining a vector for the first candidate topic based on the relative frequency of occurrence of one or more semantic keywords in the recipes associated with the centroid dish identifiers in the first candidate topic.

[0044] Therefore, the similarity between the dish represented by the corresponding dish identifier and the first candidate topic can be calculated based on the close distance between the vector of the corresponding dish identifier in the reference set and the vector of the first candidate topic, or based on the angle between the vector of the corresponding dish identifier in the reference set and the vector of the first candidate topic.

[0045] return Figure 2 At block 210, one or more dish identifiers with the highest relevance score in the reference set are selected as one or more variant dish identifiers of the first candidate theme. More specifically, one or more dish identifiers with the highest relevance score in the reference set may be selected by the processor 102 of the apparatus 100 as one or more variant dish identifiers of the first candidate theme.

[0046] In some embodiments, selecting one or more dish identifiers in the reference set as one or more variant dish identifiers of the first candidate theme can be performed based on at least one of a predetermined threshold of the association score and a predetermined number of dish identifiers to be selected. For example, the processor 102 can be configured to either select only dish identifiers whose association scores exceed the predetermined threshold, or select only four dish identifiers with the highest association scores among all dish identifiers in the reference set, or select only four dish identifiers with the highest association scores among all dish identifiers in the reference set and assume that the association scores of these four dish identifiers all exceed the predetermined threshold.

[0047] return Figure 2 At block 212, the centroid dish identifier of the first candidate theme and one or more variant dish identifiers of the first candidate theme are output. More specifically, the centroid dish identifier and one or more variant dish identifiers of the first candidate theme may be output via the user interface 104 (e.g., display screen) of the device 100 under the control of the processor 102.

[0048] In some embodiments, outputting the centroid dish identifier of the first candidate theme and the one or more variant dish identifiers of the first candidate theme at block 212 includes: displaying the first candidate theme, and upon receiving user input to expand the first candidate theme, displaying the centroid dish identifier of the first candidate theme and the one or more variant dish identifiers of the first candidate theme. In these embodiments, the centroid dish identifier of the first candidate theme may be displayed above the one or more variant dish identifiers of the first candidate theme. Furthermore, in these embodiments, the one or more variant dish identifiers of the first candidate theme may be displayed in descending order of the corresponding relevance scores.

[0049] Although Figure 2 204. Although not shown, in some embodiments, the method may further include analyzing the first image obtained at box 202 using a predictive model to determine one or more additional candidate topics, and outputting the one or more additional candidate topics. Each additional candidate topic may include multiple candidate dish identifiers, each candidate dish identifier being associated with a candidate dish. In addition, for each additional candidate topic, one of the candidate dish identifiers may be a centroid dish identifier that best represents the corresponding candidate topic. In these embodiments, the predictive model used to determine one or more additional candidate topics may be the same as the predictive model used to determine the first candidate topic at box 204. In some embodiments, the processor 102 of the device 100 may be configured to determine a predetermined number of additional candidate topics or a maximum predetermined number of additional candidate topics.

[0050] In some embodiments, outputting each of the one or more additional candidate themes may include outputting a centroid dish identifier of the corresponding additional candidate theme and one or more variant dish identifiers of the corresponding additional candidate theme. In some embodiments, outputting each of the one or more additional candidate themes may include displaying the corresponding additional candidate theme, and upon receiving user input to expand the corresponding additional candidate theme, displaying the centroid dish identifier of the corresponding additional candidate theme and one or more variant dish identifiers of the corresponding additional candidate theme. Figure 3A and Figure 3B An example of a first candidate topic being output together with one or more additional candidate topics is shown in FIG.

[0051] In some embodiments where the method includes analyzing the first image using a predictive model to determine one or more additional candidate topics, the method may further include performing the following steps for each of the one or more additional candidate topics: for each dish identifier in the reference set, calculating an association score indicating a similarity between the dish represented by the corresponding dish identifier and the corresponding additional candidate topic; selecting one or more dish identifiers in the reference set with the highest association score as one or more variant dish identifiers of the corresponding additional candidate topic; and outputting a centroid dish identifier of the corresponding additional candidate topic and one or more variant dish identifiers of the corresponding additional candidate topic. In these embodiments, the association score for each dish identifier in the reference set may be calculated based on a cosine similarity metric.

[0052] Although Figure 2 202. Although not shown, in some embodiments the method may further include determining a ranking of the first candidate topic and one or more additional candidate topics using the predictive model. In these embodiments, the ranking may indicate a decreasing similarity between the centroid dish of the corresponding candidate topic and the dish depicted in the first image obtained at block 202. In addition, in these embodiments, the method may further include displaying the first candidate topic and one or more additional candidate topics based on the determined ranking. Figure 3A and Figure 3B An example of ranking of the determined first candidate topic and one or more additional candidate topics is shown.

[0053] Although Figure 2 106. Although not shown, in some embodiments, the method may further include obtaining a plurality of recipes. Each of the plurality of recipes includes: a dish identifier, a plurality of food ingredients, and one or more cooking instructions. The dish identifier, the plurality of food ingredients, and the one or more cooking instructions may be provided and / or stored (e.g., in the memory 106) in a tagged or untagged format. Examples of tagged formats include Extensible Markup Language (XML) and JavaScript Object Notation (JSON), and examples of untagged formats include delimited plain text, regular expressions, and spreadsheets.

[0054] After the step of obtaining a plurality of recipes, the method may further include selecting a core subset of recipes from the obtained plurality of recipes, and calculating a similarity score between each recipe in the core subset based on at least one of the following: a similarity between dish identifiers of two recipes, a similarity between food ingredients of two recipes, and a similarity between cooking instructions of two recipes. Subsequently, the method may include clustering the plurality of recipes into a plurality of reference topics based on the similarity scores of the plurality of recipes, and selecting a centroid recipe for each of the plurality of reference topics. The dish identifier of the selected recipe is a centroid dish identifier of the corresponding reference topic. At this point, recipe selection may be performed so that the recipe having the highest cosine similarity with the mean center of the corresponding reference topic is selected as the centroid recipe. Alternatively, recipe selection may be performed so that the recipe having the smallest proximity distance with the mean center of the corresponding reference topic is selected as the centroid recipe. In addition, in these embodiments, determining the first candidate topic at block 204 may include selecting the first candidate topic from the plurality of reference topics.

[0055] In these embodiments, calculating the similarity scores between recipes in the core subset can be performed based on the cosine similarity metric. Each of the multiple recipes obtained can be represented by a vector, which is determined based on the relative frequency of occurrence of one or more semantic keywords in the recipe (e.g., the dish identifier of the recipe and its aliases / synonyms, food ingredients, and instructions) (e.g., determined by the processor 102 of the device 100). The similarity score between two recipes in the multiple recipes can be represented by the angle between two vectors corresponding to the two recipes. In these embodiments, the similarity scores between the recipes in the calculated core subset can be adjusted for irrelevant common associations. In addition, in these embodiments, the calculation of the similarity scores can be performed based on at least the similarity of the food ingredients of the two recipes. In this case, the similarity scores between the recipes in the calculated core subset can be adjusted for the proportion of the food ingredients in the total weight of the recipe (vectorized and proportion-weighted term frequency-inverse document frequency metric).

[0056] As an example, refer to Figure 1As described, multiple recipes can be obtained from a recipe database. In this example, the recipe database may include recipes associated with multiple dishes N, and each dish n is associated with a corresponding structured text description consisting of L(n) items. A total number of items (or keywords) M can be extracted from the recipe database. For each dish n in the N dishes, the occurrence count of item m is f(n,m), so the fractional item frequency of item m in dish n is f(n,m) / L(n). Item m also appears in the structured text descriptions of N'(m) / N dishes, where repetitions in the structured text descriptions associated with a single dish are not counted. Therefore, the fractional document frequency of item m is N'(m) / N. Therefore, the relevant term frequency T(n,m) (i.e., "term frequency inverse document frequency (TF-IDF)") can be adjusted from the original fractional term frequency for the document frequency according to the following formula:

[0057]

[0058] Thus, each dish n has a term vector of length M with TF-IDF values.

[0059] The cosine similarity value between the item vectors of any two dishes n1, n2 can be determined using the following formula:

[0060]

[0061] Among them, w m Represents the importance weighting of the terms according to, for example, the proportion of food ingredients in the total weight of a recipe. In this case, the closer the two vectors are to each other, the greater the cosine similarity value.

[0062] In these embodiments, selecting the recipe with the highest cosine similarity to the corresponding reference topic as the centroid recipe can be performed based on at least one of a similarity metric and a proximity distance metric. For example, the recipe with the maximum similarity metric (i.e., the minimum angle between the vector representing the recipe and the vector of the corresponding reference topic) can be selected as the centroid recipe. As another example, the recipe with the minimum proximity distance metric can be selected as the centroid recipe. The proximity distance metric can be represented by a proximity value D. In more detail, the proximity value D(n1, n2) between the term vectors of any two dishes n1, n2 (associated with the corresponding recipes) can be determined using the following formula:

[0063]

[0064] The proximity value is the opposite of the cosine similarity value, that is, the closer the two vectors are to each other, the smaller the proximity value. In addition, in these embodiments, the method may also include determining a popularity score for each recipe in a plurality of recipes. The popularity score of a recipe may indicate the popularity or prevalence of the recipe. For example, the popularity score of a recipe may be determined based on a search result count of the full name of the corresponding dish returned by an Internet search engine or a cooking website or a food / meal order service website. In these embodiments, selecting a core subset of recipes may be performed based on the popularity scores of a plurality of recipes. For example, a core subset of recipes may include recipes having a popularity score above a predetermined threshold. As another example, a core subset of recipes may include a predetermined number (e.g., the 200 with the highest popularity scores) of recipes with the highest popularity scores.

[0065] In some embodiments, the plurality of recipes may be obtained by the processor 102 of the device 100. Additionally, in some embodiments, the plurality of recipes may be obtained from a recipe database. For example, the structured recipe database may be stored in the memory 106 of the device 100, and the processor 102 may obtain the plurality of recipes from the memory 106. The recipe database may be configured such that new recipes may be introduced to continuously expand the recipe database. Additionally, as described above, in these embodiments, obtaining a reference set of dish identifiers at block 206 may include extracting the dish identifier from the plurality of dishes obtained.

[0066] As described above, in some embodiments the method may further include calculating a similarity score between recipes in the core subset. In these embodiments, calculating the similarity score between recipes in the core subset may be performed based on one or more synonyms of at least one of: a dish identifier, a plurality of food ingredients, and cooking instructions of the two recipes. One or more synonyms of at least one of: a dish identifier, a plurality of food ingredients, and cooking instructions of the recipes may be obtained from a synonym dictionary stored in the memory 106 or retrieved by the processor 102. The synonym dictionary may include synonyms for cooking terms, such as food ingredients, characteristic flavors, cooking methods, and cooking styles.

[0067] In some embodiments, in addition to the dish identifier (referred to herein as the "original dish identifier"), each of the plurality of recipes obtained may also include one or more equivalent dish identifiers. The one or more equivalent dish identifiers may be synonyms of the original dish identifier, since a particular dish may have different names depending on different factors such as region. For example, a dish with an original dish identifier of "cheese omelet" may have equivalent dish identifiers such as "cheese omelette" and "omelette du fromage". In these embodiments, calculating the similarity score between the recipes in the core subset may be performed based on the original dish identifier and at least one equivalent dish identifier.

[0068] As described above, in some embodiments the method may further include clustering the plurality of recipes into a plurality of reference topics. In more detail, recipes having low covariance or high similarity (expressed as a similarity score) with each other may be clustered into a plurality of reference topics. In these embodiments, the clustering process may be performed based on K-means clustering or singular value decomposition.

[0069] Although Figure 2 Although not shown, the method may also include, for each of the plurality of reference themes, determining a plurality of keywords based on recipes of the corresponding reference theme. In these embodiments, each of the plurality of keywords may be associated with at least one of a cooking technique and a food ingredient. The plurality of keywords may be extracted from recipes associated with the centroid dishes of the selected reference theme.

[0070] As described above, in some embodiments, the method may further include clustering a plurality of recipes into a plurality of reference themes. In these embodiments, the method may further include selecting one of the plurality of reference themes, acquiring a second image based on a centroid dish identifier of the selected reference theme and at least one of a plurality of keywords of the selected reference theme, and training a prediction model based on the second image and the selected reference theme. In these embodiments, the prediction model may be trained based on ground truth using the selected reference theme as a label. The prediction model may also be further validated based on the second image and the selected reference theme. It should be understood that more than one second image may be acquired based on the centroid dish identifier of the selected reference theme and at least one of a plurality of keywords of the selected reference theme, and training the prediction model may be performed based on more than one of these second images and the selected reference theme.

[0071] In these embodiments, obtaining the second image may include retrieving the second image based on an image search using at least one of the centroid dish identifiers of a plurality of keywords and / or a selected reference topic in an Internet search engine. For example, the second image may be an image that is highly correlated with the centroid dish identifier as the search keyword. In this case, the second image will increase the intra-topic diversity when training the predictive model. In addition, the search may be conducted with appropriate inclusiveness. For example, full-word matching of a single keyword or a combination of different keywords may be used. In more detail, in the case of Chinese text, a keyword may consist of more than one Chinese character without separators. In this case, full-word matching refers to matching of the entire keyword (i.e., each Chinese character in the keyword), rather than just matching of a portion of the keyword (i.e., only some of the Chinese characters in the keyword).

[0072] Although Figure 2 , but the method may also include receiving user input selecting a dish identifier. In some embodiments, the user is able to select from the output centroid dish identifier and one or more variant dish identifiers of the first candidate theme. In addition, in embodiments where a candidate theme and one or more variant dish identifiers are output for one or more additional candidate themes, the user is able to select from these output dish identifiers. The dish identifier selected by the user may be stored (e.g., in the memory 106 of the device 100) for dietary analysis purposes and / or dietary history recording of the user.

[0073] Figure 3A and Figure 3B A first candidate topic and a plurality of additional candidate topics are shown outputted on a display in a first state and a second state, respectively, according to one embodiment.

[0074] Figure 3A and Figure 3B The output shown in Figure 1 104 of the device 100. In more detail, the output can be performed after the following steps: the processor 102 has determined the first candidate theme and multiple additional candidate themes, has calculated the association score of each dish identifier in the reference set with each candidate theme, has selected a variant dish identifier for each of the first candidate theme and multiple additional candidate themes, and has determined the ranking of the first candidate theme and multiple additional candidate themes. In this embodiment, the ranking indicates a decreasing similarity between the centroid dish of the corresponding candidate theme and the dish depicted by the first image.

[0075] like Figure 3AAs shown, in the first state, the first candidate topic 31 ("candidate topic A") is displayed together with three additional candidate topics 32, 33, 34 ("candidate topic B", "candidate topic C" and "candidate topic D") and their determined corresponding rankings. In addition, the candidate topics are displayed in descending order of the ranking of the determined corresponding candidate topics. Figure 3A It can be considered as an example of the initial output stage of candidate topics 31, 32, 33, 34.

[0076] When receiving (e.g., via user interface 104) user input to expand any one of candidate topics 31, 32, 33, 34, the centroid dish identifier of the candidate topic and one or more variant dish identifiers of the candidate topic are also displayed. In this embodiment, the user can select a candidate topic to expand by clicking the inverted triangle icon of the corresponding candidate topic.

[0077] An example of a user selecting a first candidate topic 31 for expansion, thereby initiating a transition from the first state to the second state is shown in Figure 3B . After receiving the specific user input, the centroid dish identifier 310 of the first candidate theme 31 and the first variant dish identifier 311, the second variant dish identifier 312, the third variant dish identifier 313 and the fourth variant dish identifier 314 are displayed below the candidate theme identifier of the first candidate theme 31. More specifically, the centroid dish identifier 310 is displayed above the variant dish identifiers 311, 312, 313, 314 (as the default option for the candidate theme). In addition, in this embodiment, the first to fourth variant dish identifiers 311, 312, 313, 314 of the first candidate theme 31 are displayed in descending order according to their corresponding relevance scores.

[0078] In this way, a two-level format of image-based dish recognition results can be presented to the user. Specifically, the user is provided with an overview of candidate topics depicting a first image of a dish to be recognized in a first state, while, if desired, the user is provided with an option to view centroid dishes and variant dishes of the candidate topics. In this embodiment, the candidate topic can also be "minimized" by clicking on the triangle icon corresponding to the expanded candidate topic, thereby returning to the first state from the second state.

[0079] Therefore, an improved method and apparatus for image-based dish recognition is provided that overcomes existing problems.

[0080] A computer program product comprising a computer readable medium is also provided, wherein the computer readable medium has computer readable code stored therein, and the computer readable code is configured to, when executed on a suitable computer or processor, cause the computer or processor to perform one or more methods described herein. Therefore, it should be understood that the present disclosure is equally applicable to computer programs, in particular computer programs on or in a carrier suitable for putting the embodiments of the present disclosure into practice. The program may be in the form of source code, object code, code intermediate source, and object code such as partially compiled form, or any other form suitable for use in the implementation of the method according to the embodiment described herein.

[0081] It should also be understood that such programs can have many different architectural designs. For example, the program code that implements the functions of the method or system can be subdivided into one or more subroutines. The various different distribution modes of functions between these subroutines will be obvious to those skilled in the art. The subroutines can be stored together in an executable file to form a self-contained program. This executable file can include computer executable instructions, such as processor instructions and / or interpreter instructions (such as Java interpreter instructions). Alternatively, one or more or all subroutines can be stored in at least one external library file and, for example, statically or dynamically linked to the main program during operation. The main program includes at least one call to at least one subroutine. The subroutines can also include function calls to each other.

[0082] One embodiment of a computer program product includes computer executable instructions corresponding to each processing stage of at least one method set forth herein. These instructions may be subdivided into subroutines and / or stored in one or more files that may be statically or dynamically linked. Another embodiment of a computer program product includes computer executable instructions corresponding to each device of at least one system and / or product set forth herein. These instructions may be subdivided into subroutines and / or stored in one or more files that may be statically or dynamically linked.

[0083] The carrier of a computer program may be any entity or device capable of carrying the program. For example, the carrier may include a data storage device, such as a ROM, such as a CDROM or a semiconductor ROM, or a magnetic recording medium, such as a hard disk. In addition, the carrier may be a transmissible carrier, such as an electrical signal or an optical signal, which may be transmitted via an electric cable or an optical cable or by radio or other means. When the program is contained in such a signal, the carrier may be constituted by such a cable or other device or apparatus. Alternatively, the carrier may be an integrated circuit in which the program is embedded, which is suitable for executing the relevant method or for use in the execution of the relevant method.

[0084] In the process of practicing the claimed invention, by studying the drawings, the disclosure and the appended claims, variations of the disclosed embodiments can be understood and implemented by those skilled in the art. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "one" or "an" does not exclude a plurality. A single processor or other unit can meet multiple functions described in the claims. The fact that certain measures are recorded in mutually different dependent claims does not indicate that a combination of these measures cannot be used to obtain advantages. The computer program may be stored / distributed on an appropriate medium, such as an optical storage medium or solid-state medium provided with the hardware or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems. Any figure marks in the claims should not be understood as limiting their scope.

Claims

1. A computer-implemented method for performing image-based dish recognition, the method comprising: Acquiring (202) a first image depicting a dish to be identified; obtaining a plurality of recipes, wherein each of the plurality of recipes comprises: a dish identifier, a plurality of food ingredients, and one or more cooking instructions; clustering the plurality of recipes into a plurality of reference topics; analyzing (204) the first image using a predictive model to determine a first candidate theme, wherein the first candidate theme includes a plurality of candidate dish identifiers, each of the candidate dish identifiers being associated with a candidate dish, one of the candidate dish identifiers being a centroid dish identifier associated with a candidate dish that best represents the first candidate theme, and wherein determining the first candidate theme includes: selecting the first candidate theme from the plurality of reference themes; Obtaining (206) a reference set of dish identifiers; calculating (208) for each of the dish identifiers in the reference set, a relevance score indicating a similarity between the dish represented by the corresponding dish identifier and the first candidate topic based on a proximity distance between a vector representing the corresponding dish identifier in the reference set and a vector representing a first candidate topic, or based on an angle between a vector representing the corresponding dish identifier in the reference set and a vector representing the first candidate topic, wherein the vector is a result of a cluster analysis; selecting (210) one or more dish identifiers in the reference set having the highest relevance scores as one or more variant dish identifiers of the first candidate theme; and The centroid dish identifier of the first candidate theme and one or more variant dish identifiers of the first candidate theme are output (212).

2. The computer-implemented method of claim 1 , wherein outputting ( 212 ) a centroid dish identifier ( 310 ) of the first candidate theme ( 31 ) and one or more variant dish identifiers ( 311 , 312 , 313 , 314 ) of the first candidate theme comprises: Display the first candidate topic; as well as When receiving user input for expanding the first candidate theme, displaying a centroid dish identifier of the first candidate theme and one or more variant dish identifiers of the first candidate theme, The centroid dish identifier of the first candidate theme is displayed above the one or more variant dish identifiers of the first candidate theme, and the one or more variant dish identifiers of the first candidate theme are displayed in descending order according to the corresponding relevance scores. 3 . The computer-implemented method of claim 1 , wherein a plurality of candidate dish identifiers in the first candidate theme represent dishes that are similar to each other.

4. The computer-implemented method of claim 1 or 2, further comprising: analyzing the first image using a predictive model to determine one or more additional candidate topics, wherein each of the additional candidate topics includes a plurality of candidate dish identifiers, each of the candidate dish identifiers is associated with a candidate dish, and one of the candidate dish identifiers is a centroid dish identifier that best represents the corresponding candidate topic; and The one or more additional candidate topics are output.

5. The computer-implemented method of claim 4, further comprising, for each of the one or more additional candidate topics: For each dish identifier in the reference set, calculating a relevance score indicating a similarity between the dish represented by the corresponding dish identifier and the corresponding additional candidate topic; selecting one or more dish identifiers with the highest relevance scores in the reference set as one or more variant dish identifiers of the corresponding additional candidate themes; as well as The centroid dish identifier of the corresponding additional candidate theme and the one or more variant dish identifiers of the corresponding additional candidate theme are output.

6. The computer-implemented method of claim 4, further comprising: determining a ranking of the first candidate topic and the one or more additional candidate topics using the predictive model, wherein the ranking indicates decreasing similarity between the centroid dishes of the respective candidate topics and the dishes depicted in the obtained first image; and Based on the determined ranking, the first candidate topic and the one or more additional candidate topics are displayed.

7. The computer-implemented method of claim 1 or 2, further comprising: Selecting a core subset of recipes from the obtained plurality of recipes; calculating a similarity score between recipes in the core subset based on at least one of: a similarity between the dish identifiers of two recipes, a similarity between the food ingredients of two recipes, and a similarity between the cooking instructions of two recipes; clustering the plurality of recipes into a plurality of reference topics based on similarity scores of the plurality of recipes; as well as For each of the plurality of reference topics, a recipe having the highest cosine similarity with the corresponding reference topic is selected as a centroid recipe, wherein the dish identifier of the selected recipe is the centroid dish identifier of the corresponding reference topic.

8. The computer-implemented method of claim 7, further comprising: determining a popularity score for each recipe in the plurality of recipes, wherein the popularity score indicates popularity or prevalence of the corresponding recipe, Wherein selecting the core subset of recipes is performed based on popularity scores of the plurality of recipes.

9. The computer-implemented method of claim 7, wherein: Calculating the similarity score between the recipes in the core subset is performed based on one or more synonyms of at least one of: cooking instructions, a dish identifier, and a plurality of food ingredients of the two recipes.

10. The computer-implemented method of claim 7, wherein: Clustering the plurality of recipes into a plurality of reference topics is performed based on K-means clustering or singular value decomposition.

11. The computer-implemented method of claim 7, further comprising: For each of the plurality of reference topics, a plurality of keywords are determined based on the recipe of the corresponding reference topic, wherein each of the plurality of keywords is associated with at least one of a cooking technique and a food ingredient.

12. The computer-implemented method of claim 7, further comprising: selecting one of the plurality of reference subjects; acquiring a second image based on the centroid dish identifier of the selected reference theme and at least one of a plurality of keywords of the selected reference theme; as well as The prediction model is trained based on the second image and the selected reference subject.

13. The computer-implemented method of any one of claims 1-2, 5-6 and 8, wherein: The prediction model is at least one of: a convolutional neural network, a residual neural network, and a dense neural network.

14. A computer program product comprising a computer-readable medium having computer-readable code implemented therein, wherein the computer-readable code is configured to, when executed on a suitable computer or processor, cause the computer or processor to perform the method of any one of claims 1 to 13.

15. A device (100) for performing image-based dish recognition, the device comprising a processor (102), the processor being configured to: Acquiring a first image depicting a dish to be identified; A plurality of recipes is obtained, wherein each of the plurality of recipes comprises: a dish identifier, a plurality of food ingredients, and one or more cooking instructions; clustering the plurality of recipes into a plurality of reference topics; analyzing the first image using a predictive model to determine a first candidate theme, wherein the first candidate theme includes a plurality of candidate dish identifiers, each of the candidate dish identifiers being associated with a candidate dish, one of the candidate dish identifiers being a centroid dish identifier associated with a candidate dish that best represents the first candidate theme, and wherein determining the first candidate theme includes: selecting the first candidate theme from the plurality of reference themes; Get a reference collection of dish identifiers; calculating, for each of the dish identifiers in the reference set, a correlation score indicating a similarity between the dish represented by the corresponding dish identifier and the first candidate topic based on a proximity distance between a vector representing the corresponding dish identifier in the reference set and a vector representing the first candidate topic, or based on an angle between a vector representing the corresponding dish identifier in the reference set and a vector representing the first candidate topic, wherein the vector is a result of cluster analysis; selecting one or more dish identifiers with the highest relevance scores in the reference set as one or more variant dish identifiers of the first candidate theme; and The centroid dish identifier of the first candidate theme and the one or more variant dish identifiers of the first candidate theme are output.

Citation Information

Patent Citations

  • Automated Food Recognition and Nutritional Estimation With a Personal Mobile Electronic Device

    US20160063734A1

  • Digital recipe library and network with food image recognition services

    US20180308143A1