A dish identification method and system based on dish sample retention
Through the dish sample retention method, combined with the target detection and measurement learning model, tableware interference is reduced, and a weighted and integrated new dish library is established, which solves the problems of high cost and poor flexibility of dish recognition in the existing technology, and improves the dish recognition rate and user experience.
Patent Information
- Application Number
- CN202310704282.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-14
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-06-14
AI Technical Summary
The existing dish recognition technology has the problems of high cost, poor flexibility and low recognition rate, especially when changing dishes, it is necessary to change tableware or purchase new labels, and deep learning methods are easily disturbed by external environment, resulting in poor user experience.
The method of retaining dishes is adopted to locate the dishes position through the target detection network, and semantic segmentation is performed to obtain the pure dishes area map, and the measurement learning model is used to compare with the canteen dishes total library, and a new dish library is established through a weighted fusion mechanism to reduce tableware interference and improve recognition accuracy; the number of errors is recorded during settlement and the feature library is updated to integrate error characteristics to improve recognition rate.
It significantly improves the dish recognition rate and accuracy, reduces tableware interference, improves user experience, and improves recognition efficiency and accuracy by automatically updating the feature library.
Smart Images

Figure CN116665208B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of dish sample retention and dish recognition, and particularly to a dish recognition method and system based on dish sample retention. Background Art
[0002] The traditional group meal industry is competing to rapidly transform towards intelligent group meals. The previous group meal industry faced great technical barriers, but with the development of science and technology in the new era, they have been broken through one by one. The first and foremost is the dish recognition technology, which has developed vigorously in recent years and has been widely used in the entire group meal industry.
[0003] Currently, dish recognition mainly includes the RFID-based recognition method and the method based on deep learning and image recognition.
[0004] RFID-based dish recognition is a comprehensive catering solution based on the Internet of Things RFID technology and cloud computing technology. By installing a desktop high-frequency reader / writer integrated machine on the dish settlement table, installing RFID chips at the bottom of various bowls and dishes, and placing the bowls and dishes with RFID chips on the settlement table, dish recognition can be achieved. However, it requires the cafeteria to purchase a batch of dinner plates embedded with RFID tags and enter dish information on the tags through a special RFID editor. The RFID tag recognition method has very high accuracy, but each dinner plate needs to be pre-embedded with a tag to mark dish information. When the cafeteria changes dishes, it is necessary to modify the RFID tag information of a large number of dinner plates or purchase a new batch of dinner plates, resulting in relatively high cost consumption. Moreover, the long-term cleaning and high-temperature disinfection of the dinner plates, or the collision between the dinner plates, will reduce the service life of the RFID tags.
[0005] Dish recognition based on deep learning and image recognition is to locate the position coordinate information of each bowl and dish captured by the camera through a deep learning object detection network, and then extract the feature information of each dish through a deep convolutional neural network, and compare it with the pre-established dish feature retrieval library one by one, so as to quickly realize the automatic recognition of dishes. When encountering a new dish, this method only needs to input a picture of the new dish and add its features to the dish feature library, and then the recognition of this type of dish can be quickly realized, which has extremely high flexibility and scalability. Compared with the RFID method, it eliminates the dish production equipment and production process, has low constraints on dining staff, and removes the dependence on tableware. When replacing new dishes, the existing tableware can be directly used, greatly reducing the procurement cost. In addition, the deep learning method can also perform complex recognition such as fruits, beverages, and quantities. However, the dish recognition method based on depth is easily affected by external environmental light interference, reducing the recognition rate. The diversity of cooking techniques and combinations of dishes, as well as the differences in regions, cultures, etc., result in a wide variety of types and a large amount of data in the dish feature library of each restaurant. If the feature comparison and recognition are performed with the large dish library of the restaurant every time a meal is taken, it will greatly affect the dish recognition rate and the user experience is poor. Summary of the invention
[0006] The object of the present invention is to provide a dish recognition method and system based on dish sampling to improve the dish recognition rate, so as to enhance the user experience.
[0007] The dish identification method based on dish sample retention of the present invention comprises:
[0008] Sample dishes:
[0009] Obtain sample dish images, locate the position coordinates of each dish in the sample dish images through the object detection network, semantically segment each dish sub-image according to the position coordinates of each dish, and obtain a pure dish area image without tableware;
[0010] After the pure dish area pictures are extracted with the metric learning model, they are compared with each dish category in the canteen dish database one by one, and classified according to the highest similarity score to obtain the corresponding category of each dish;
[0011] According to the corresponding categories of each dish, the dishes are automatically added to the menu list of the meal, and the menu of the meal and the pure dish area map of each sample category of dishes are obtained;
[0012] Weighted Fusion:
[0013] According to the menu of the meal, the corresponding dish image information is screened from the canteen dish library to obtain the meal library. After feature extraction of the pure dish area map of each sample category dish, it is weightedly fused with the pictures of the meal library to obtain a new dish library.
[0014] Settlement:
[0015] Get the current meal picture, locate the position coordinates of each dish in the current meal picture through the target detection algorithm, cut each dish sub-graph according to the position coordinates of each dish, and obtain the dish sub-graph containing tableware;
[0016] After the features of the dish sub-graph containing tableware are extracted by the metric learning model, they are compared with the current dish graphs in the new dish library one by one. If the dishes corresponding to each category are identified, settlement will be made.
[0017] As a preferred solution of the present invention, after extracting features from the dish sub-graph containing tableware through a metric learning model, feature comparison is performed one by one with the dish pictures of the current meal in the new dish library. If the dishes corresponding to each category cannot be identified, the number of errors of each dish and the features of the erroneous dish pictures are recorded and accumulated, and the features of these erroneous samples are stored to obtain an erroneous picture feature library, which is set to be valid for the current meal and automatically cleared after the meal is over.
[0018] Determine the number of errors in each category of dishes. If it is greater than the preset error threshold, the features of the error category dish images are merged into the new dish library.
[0019] As a preferred solution of the present invention, the calculation method of weighted fusion is as follows, assuming that the dish category A is:
[0020] 1. The dish category A under the current meal database has the following dish images: image1, ..., image N (N is the number of pictures, N<=5), the corresponding features of each picture are: feature (1,1536) ,…,featureN (1,1536) (N is the number of pictures, N<=5, and the dimension of each feature is 1536), then the representative features of this category of dishes are:
[0021] 2. Dish category A under the pure dish area map feature of the retained category dishes, the corresponding dish image is: image tmp ; Representative features: featureT (1,1536) ;
[0022] 3. Find the feature similarity of dish A under the pure dish area map features of the current meal database and the reserved sample category dishes. From 1 and 2 above, we get:
[0023]
[0024] 4. The representative features of dish category A in the new dish library after weighted fusion are:
[0025] Feature (1,1536) =(1-weight)featureA (1,1536) +weight*featureT (1,1536) ; weight is the fusion weight of the feature, when similarity>=0.7, weight=0.5; when 0.5<=similarity<0.7, weight=0.1; when similarity<0.5, weight=0.
[0026] As a preferred solution of the present invention, the canteen dish library is specifically composed of a transformer network-based distillation backbone network, which trains a feature extraction model with learning and generalization capabilities, and a canteen dish library corresponding to the feature retrieval of the canteen dish library is established through the model.
[0027] The dish identification system based on dish sample retention of the present invention is characterized by comprising:
[0028] Dish Sampling Terminal:
[0029] Collect pictures of sampled dishes, locate the position coordinates of each dish in the picture of the sampled dish through the target detection network, and semantically segment each dish sub - picture according to the position coordinates of each dish to obtain pictures of pure dish areas excluding tableware;
[0030] After extracting features from the pictures of pure dish areas through the metric learning model, compare them one by one with each dish category in the general cafeteria dish library, and classify them according to the highest similarity score value to obtain the corresponding categories of each dish;
[0031] Automatically add them to the menu list of the current meal according to the corresponding categories of each dish obtained, obtain the menu of the current meal and the pictures of pure dish areas of each sampled category of dishes and send them to the weighted fusion terminal;
[0032] Weighted Fusion Terminal:
[0033] Receive the menu of the current meal and the pictures of pure dish areas of each sampled category of dishes, screen the corresponding dish picture information from the large cafeteria dish library to obtain the current meal dish library, extract features from the pictures of pure dish areas of each sampled category of dishes, and perform weighted fusion with the pictures in the current meal dish library to obtain a new dish library and send it to the settlement terminal;
[0034] Settlement Terminal:
[0035] Collect pictures of the current meal dishes, locate the position coordinates of each dish in the picture of the current meal dishes through the target detection algorithm, and cut each dish sub - picture according to the position coordinates of each dish to obtain dish sub - pictures containing tableware;
[0036] After extracting features from the dish sub - pictures containing tableware through the metric learning model, compare the features with the current meal dish pictures in the new dish library one by one. If the dishes corresponding to each category are recognized, settlement is carried out.
[0037] As a preferred solution of the present invention, the settlement terminal also extracts features from the dish sub - pictures containing tableware through the metric learning model, compares the features with the current meal dish pictures in the new dish library one by one. If the dishes corresponding to each category cannot be recognized, record and accumulate the error occurrence times of each category of dishes and the features of the pictures of the error - category dishes, store the features of these error samples to obtain an error - picture feature library, set the error - picture feature library as valid for the current meal, and automatically clear it after the meal.
[0038] Judge the error occurrence times of each dish. If it is greater than the preset error occurrence times threshold, send the features of the pictures of the error - category dishes to the weighted fusion terminal for fusion with the new dish library.
[0039] As a preferred solution of the present invention, the weighted fusion calculation of the weighted fusion terminal is as follows. Let the dish category be A:
[0040] 1. For the dish category A under the current meal's dish library, the corresponding dish pictures are: image1, …, image N (where N is the number of pictures, N <= 5), and the corresponding features of each picture are: feature (1,1536) , …, featureN (1,1536) (where N is the number of pictures, N <= 5, and the dimension of each feature feature is 1536), then the representative feature of this category of dishes is:
[0041] 2. For the dish category A under the pure dish area map feature of the sampled dish category, the corresponding dish picture is: image tmp ; the representative feature is: featureT (1,1536) ;
[0042] 3. Calculate the feature similarity similarity of dish A under the current meal's dish library and the pure dish area map feature of the sampled dish category. From the above 1 and 2, we get:
[0043]
[0044] 4. For the dish category A under the new dish library, the representative feature obtained through weighted fusion is:
[0045] Feature (1,1536) =(1 - weight)featureA (1,1536) +weight * featureT (1,1536) ; weig h t is the fusion weight of the feature. When simi l arity >= 0.7, weight = 0.5; when 0.5 <= similarity < 0.7, weight = 0.1; when similarity < 0.5, weight = 0.
[0046] As a preferred solution of the present invention, the total cafeteria dish library at the dish sampling end is specifically a large cafeteria dish library that trains a feature extraction model with learning and generalization capabilities based on the distillation backbone network of the transformer network, and establishes a feature retrieval corresponding to the cafeteria dish library through this model.
[0047] Advantages of the present invention:
[0048] The dish sample retention is responsible for obtaining the recipe of the dishes served in the cafeteria before the meal starts. It accurately condenses the identification from the original cafeteria general warehouse to a smaller range of only the dish library for the current meal, effectively reducing the identification calculation on the settlement side and significantly improving the dish identification rate. Additionally, since there are significant differences between the tableware used for serving the sample dishes and the tableware for the dishes when customers actually dine, to reduce the interference of the tableware and improve the contribution of the sample dishes to the identification accuracy of the settlement dishes, by semantically segmenting each dish sub-image according to the position coordinates of each dish, the dish identification accuracy can be effectively improved. Furthermore, after the sample retention is completed, the recipe for the current meal and the pure dish area map of each sample retention category of dishes are obtained. Since the portion of the sample dishes is small and the tableware used for serving is very different from the tableware of the dishes during the settlement identification of customer dining, by introducing a semantic segmentation model in the dish positioning link during the sample retention, accurately extracting the dish area, eliminating the interference of non-dish items such as bowls and plates, and forming the pure dish area map of each sample retention category of dishes. The recipe for the current meal is used to screen the dish library for the current meal used for settlement identification from the cafeteria general warehouse. Through a weighted fusion mechanism, after the features of the pure dish area map of each sample retention category of dishes are extracted, they are weighted and fused with the pictures in the dish library for the current meal, thus ensuring the accuracy of the settlement identification, improving the dish identification rate, and greatly enhancing the user experience.
[0049] During settlement, it automatically records and accumulates the error occurrence times of each category of dishes and the features of the pictures of the error category dishes, stores the features of these error samples to obtain an error picture feature library, which is valid for the current meal and is automatically cleared after the meal, without affecting the cafeteria's general dish library, thereby further improving the efficiency of settlement identification. Additionally, during settlement identification, by judging the error occurrence times of each category of dishes, if it is greater than the preset error occurrence threshold, it proves that the original dish features of this category of dishes are less representative, resulting in a low identification rate for this category of dishes. At this time, the features of the pictures of the error category dishes are fused with the new dish library, thereby being able to improve the identification accuracy of the frequently erred category of dishes, further improving the dish identification rate, and well enhancing the user experience. Brief Description of the Drawings
[0050] Figure 1 It is a schematic flowchart of a dish identification method based on dish sample retention according to the present invention. Detailed Embodiments
[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the protection scope of the present invention.
[0052] The present invention provides a dish identification method based on dish sample retention, as Figure 1 shown, including:
[0053] Food sample retention:
[0054] S1. Obtain pictures of the retained food samples, locate the position coordinates of each dish in the pictures of the retained food samples through a target detection network, and semantically segment each dish sub-picture according to the position coordinates of each dish to obtain pictures of pure dish areas after removing tableware.
[0055] The person in charge of food sample retention is responsible for obtaining the menu of the dishes served in the cafeteria before the meal. It is accurately concentrated from the original cafeteria general library to a smaller library for dish recognition only for the current meal, effectively reducing the recognition calculation on the settlement side, significantly improving the recognition accuracy, and increasing the dish recognition rate. In addition, since there are great differences between the tableware for serving the retained food samples and the tableware for the dishes when customers actually dine, in order to reduce the interference of the tableware and improve the contribution of the retained food samples to the recognition accuracy of the settlement dishes, each dish sub-picture is semantically segmented according to the position coordinates of each dish, thereby effectively increasing the dish recognition rate.
[0056] S2. After the pictures of the pure dish areas are extracted with features by a metric learning model, they are compared with each dish category in the cafeteria dish general library one by one, and classified according to the highest similarity score value to obtain the corresponding categories of each dish.
[0057] The cafeteria dish general library can train a feature extraction model with learning and generalization capabilities by distilling the backbone network based on the transformer network, and establish a large cafeteria dish library for feature retrieval corresponding to the cafeteria dish library through this model, so as to improve the comparison accuracy and further improve the accuracy and efficiency of dish recognition.
[0058] S3. Automatically add them to the menu list of the current meal according to the corresponding categories of each dish obtained, and obtain the menu of the dishes served in the current meal and the pictures of the pure dish areas of each retained category of dishes.
[0059] Weighted fusion:
[0060] S4. According to the menu of the dishes served in the current meal, screen the corresponding dish picture information from the large cafeteria dish library to obtain the current meal dish library. After extracting the features of the pictures of the pure dish areas of each retained category of dishes, they are weighted and fused with the pictures in the current meal dish library to obtain a new dish library. Among them, the dish data in the current meal dish library is temporarily valid and only valid for the current meal, and is automatically cleared after the meal, and each category of retained dishes has only one picture as the feature representative of each dish in the current meal.
[0061] The calculation method of weighted fusion is as follows. Let the dish category be A:
[0062] 1. For the dish category A under the current meal dish library, the corresponding dish pictures are: image1,..., image N(N is the number of pictures, N<=5), the corresponding features of each picture are: feature (1,1536) ,…,featureN (1,1536) (N is the number of pictures, N<=5, and the dimension of each feature is 1536), then the representative features of this category of dishes are:
[0063] 2. Dish category A under the pure dish area map feature of the retained category dishes, the corresponding dish image is: image tmp ; Representative features: featureT (1,1536) ;
[0064] 3. Find the feature similarity of dish A under the pure dish area map features of the current meal database and the reserved sample category dishes. From 1 and 2 above, we get:
[0065]
[0066] 4. The representative features of dish category A in the new dish library after weighted fusion are:
[0067] Feature (1,1536) =(1-weight)featureA (1,1536) +weight*featureT (1,1536) ; weight is the fusion weight of the feature, when similarity>=0.7, weight=0.5; when 0.5<=similarity<0.7, weight=0.1; when similarity<0.5, weight=0.
[0068] Settlement:
[0069] S5. Obtain a picture of the current meal dish, locate the position coordinates of each dish in the picture of the current meal dish through a target detection algorithm, cut each dish sub-graph according to the position coordinates of each dish, and obtain a dish sub-graph containing tableware.
[0070] S6. After extracting features from the dish sub-graph containing tableware through the metric learning model, the features are compared with the dish graphs of the new dish library one by one. If the dishes corresponding to each category are identified, settlement is performed.
[0071] If the dishes corresponding to each category cannot be identified, the number of errors in each category of dishes and the features of the pictures of the erroneous categories of dishes are recorded and accumulated, and the features of these erroneous samples are stored to obtain a wrong picture feature library. The wrong picture feature library is set to be valid for the current meal and is automatically cleared after the meal is over, without affecting the total library of canteen dishes, thereby further improving the efficiency of settlement recognition.
[0072] Judge the number of error occurrences of each category of dishes. If it is greater than the preset error occurrence threshold, fuse the features of the pictures of the dishes in the error category into the new dish library, which proves that the original dish features of this category of dishes are less representative, resulting in a low recognition rate of this category of dishes. At this time, fuse the features of the pictures of the dishes in the error category with the new dish library, so as to improve the recognition accuracy of the dishes in the frequently wrong category, further improve the dish recognition rate, and greatly enhance the user experience.
[0073] The method for settling the dish recognition by fusing the features of wrong pictures is as follows:
[0074] Suppose there is a category A dish. During the recognition process of the current meal, if the number of recognition errors is counted as M times, then the features of the category A dish under the pure dish area map feature of the sampled dishes are: (M is the number of recognition errors), and the corresponding dish features in the current meal dish library are Feature (1,1536) . If the number of recognition errors of the category A dish in the current meal is greater than the set error occurrence threshold, then when a category A dish appears again in the new order to be recognized, the features Feature (1,1536) of the category A dish in the current meal dish library will be replaced by the fused new features: Feature′ (1,1536) =(1 - weight)Feature (1,1536) +weight*featureE (1,1536) , where weight is the feature fusion weight and weight = 0.5.
[0075] The present invention provides a dish recognition system based on dish sampling, including:
[0076] Dish sampling end: The dish sampling end refers to a device equipped with a weighing module, a camera module, a display module, a self-adhesive printing module, etc. It can verify the login permission through face recognition or input of the identity number, etc., and log in to the dish sampling end.
[0077] By logging in to the dish sample collection terminal to collect pictures of sampled dishes, the position coordinates of each dish in the sampled dish pictures are located by the target detection network. According to the position coordinates of each dish, each dish sub-picture is semantically segmented to obtain pictures of the pure dish areas after removing tableware. After the pictures of the pure dish areas are extracted with features by the metric learning model, they are compared one by one with each dish category in the cafeteria dish library. They are classified according to the highest similarity score value to obtain the corresponding categories of each dish. The cafeteria dish library is a large cafeteria dish library with a feature retrieval corresponding to the cafeteria dish library established by training a feature extraction model with learning and generalization capabilities based on the transformer network distilled backbone network, so as to improve the accuracy of comparison and further improve the accuracy and efficiency of dish recognition. According to the corresponding categories of each dish obtained, they are automatically added to the menu list of the current meal, and the menu of the current meal and the pictures of the pure dish areas of each sampled category of dishes are obtained and sent to the weighted fusion terminal.
[0078] The weighted fusion terminal refers to the background server. The background receives the menu of the current meal and the pictures of the pure dish areas of each sampled category of dishes, and screens the corresponding dish picture information from the large cafeteria dish library to obtain the current meal dish library. After the pictures of the pure dish areas of each sampled category of dishes are extracted with features, they are weighted and fused with the pictures in the current meal dish library to obtain a new dish library and sent to the settlement terminal. Among them, the dish data in the current meal dish library is set to be temporarily valid in the background, only valid for the current meal, and automatically cleared after the meal. And there is only one picture for each sampled category of dishes as the feature representative of each dish in the current meal. The calculation method of the background for weighted fusion is as follows. Let the dish category be A:
[0079] 1. For the dish category A in the current meal dish library, the corresponding dish pictures are: image1,..., image N (N is the number of pictures, N <= 5), and the corresponding features of each picture are: feature (1,1536) ,..., featureN (1,1536) (N is the number of pictures, N <= 5, and the dimension of each feature feature is 1536), then the representative feature of this category of dishes is:
[0080] 2. For the dish category A under the feature of the picture of the pure dish area of the sampled category of dishes, the corresponding dish picture is: image tmp ; the representative feature is: featureT (1,1536) ;
[0081] 3. Calculate the feature similarity similarity of dish A under the current meal dish library and the feature of the picture of the pure dish area of the sampled category of dishes. From the above 1 and 2, we get:
[0082]
[0083] 4. For the dish category A under the new dish library, the representative features obtained through weighted fusion are as follows:
[0084] Feature (1,1536) =(1 - weight)featureA (1,1536) +weight*featureT (1,1536) ; weight is the fusion weight of the features. When similarity >= 0.7, weight = 0.5; when 0.5 <= similarity < 0.7, weight = 0.1; when similarity < 0.5, weight = 0.
[0085] The settlement terminal refers to the settlement counter. The settlement counter calls the camera to collect dish pictures in real time. When the requirements of the dish motion and stillness detection algorithm are met, the pictures of the current dining dishes are collected. Through the object detection algorithm, the position coordinates of each dish in the current dining dish picture are located, and each dish sub - picture is cut according to the position coordinates of each dish to obtain the dish sub - pictures containing tableware; then, after the dish sub - pictures containing tableware are extracted with features by the metric learning model, they are compared with the dish pictures of the current meal in the new dish library one by one. If the dishes corresponding to each category are recognized, settlement is carried out. If the dishes corresponding to each category cannot be recognized, the error occurrence times of each category of dishes and the features of the pictures of the error - category dishes are recorded and accumulated, and the features of these error samples are stored to obtain the wrong - picture feature library. The wrong - picture feature library is set as valid for the current meal and is automatically cleared after the meal; the error occurrence times of each category of dishes are judged. If it is greater than the preset error occurrence threshold, the features of the pictures of the error - category dishes are sent to the weighted fusion terminal for fusion with the new dish library. The method for the settlement counter to identify and fuse the wrong - picture features of dishes is as follows: Assume that for category A dishes, during the recognition process of the current meal, if the number of recognition errors is counted as M times, then the features of category A dishes under the pure dish area picture features of the sampled category of dishes are: (M is the number of recognition errors), and the corresponding dish features in the current - meal dish library are Feature(1, 1536). If the number of recognition errors of category A dishes in the current meal is greater than the set error occurrence threshold, then when category A dishes appear again in the new order, the feature Feature(1, 1536) of category A dishes in the current - meal dish library will be replaced by the new fused feature: Feature (1,1536) =(1 - weight)Feature (1,1536) +weight*featureE (1,1536) , where weight is the feature fusion weight and weight = 0.5.
[0086] Since the same chef may produce the same dish at different times with significant differences due to various external influences, some dishes in the dish library for the current meal used for settlement identification may have a low similarity to the actual dishes served in the cafeteria during the current meal. By automatically recording and accumulating the error counts of each dish and the features of the pictures of the error dishes during settlement, the features of these error samples are stored to obtain an error picture feature library, which is valid for the current meal and is automatically cleared after the meal, without affecting the overall cafeteria dish library, thus further improving the efficiency of settlement identification. In addition, during settlement identification, by judging the error counts of each dish, if it is greater than the preset error count threshold, it proves that the original dish features of this type of dish are less representative, resulting in a low recognition rate for this type of dish. At this time, the features of the pictures of the error dishes are fused with the new dish library, which can improve the recognition accuracy of frequently misidentified dishes, further improve the dish recognition rate, and greatly enhance the user experience.
[0087] As a preferred implementation of the present invention, in the description of this specification, the description referring to terms such as "preferred" means that the specific features, structures, materials, or characteristics described in connection with this embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expression of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0088] The above embodiments are only used to illustrate the detailed solutions of the present invention. The present invention is not limited to the above detailed solutions, that is, it does not mean that the present invention must rely on the above detailed solutions to be implemented. Those skilled in the art should understand that any improvement to the present invention, the equivalent replacement of the raw materials of the products of the present invention, the addition of auxiliary components, and the selection of specific methods, etc., all fall within the protection scope and the disclosure scope of the present invention.
Claims
1. A dish recognition method based on dish sample retention, characterized in that, include: Sample dishes: Obtain sample dish images, locate the position coordinates of each dish in the sample dish images through the object detection network, semantically segment each dish sub-image according to the position coordinates of each dish, and obtain a pure dish area image without tableware; After the pure dish area pictures are extracted with the metric learning model, they are compared with each dish category in the canteen dish database one by one, and classified according to the highest similarity score to obtain the corresponding category of each dish; According to the corresponding categories of each dish, the dishes are automatically added to the menu list of the meal, and the menu of the meal and the pure dish area map of each sample category of dishes are obtained; Weighted Fusion: According to the menu of the meal, the corresponding dish image information is screened from the canteen dish library to obtain the meal library. After feature extraction of the pure dish area map of each sample category dish, it is weightedly fused with the pictures of the meal library to obtain a new dish library. Settlement: Get the current meal picture, locate the position coordinates of each dish in the current meal picture through the target detection algorithm, cut each dish sub-graph according to the position coordinates of each dish, and obtain the dish sub-graph containing tableware; After the features of the dish sub-graph containing tableware are extracted by the metric learning model, they are compared with the current dish graphs in the new dish library one by one. If the dishes corresponding to each category are identified, settlement will be made.
2. The dish recognition method based on dish sample retention according to claim 1, wherein, After extracting features from the dish sub-graph containing tableware through the metric learning model, the features are compared with the dish images of the current meal in the new dish library one by one. If the dishes corresponding to each category cannot be identified, the number of errors for each dish and the features of the erroneous dish images are recorded and accumulated, and the features of these erroneous samples are stored to obtain the erroneous image feature library, which is set to be valid for the current meal and automatically cleared after the meal is over. Determine the number of errors for each dish. If it is greater than the preset error threshold, the features of the error dish image will be integrated into the new dish library.
3. The dish identification method based on dish sample retention according to claim 1 is characterized in that: The calculation method of weighted fusion is as follows, assuming that the dish category is A:
1. For the dish category A under the current meal's dish library, the corresponding dish pictures are: image1, …, image N (N) is the number of pictures, N <= 5), and the corresponding features of each picture are: feature (1,1536) , …, featureN (1,1536) (N is the number of pictures, N <= 5, and the dimension of each feature feature is 1536), then the representative feature of this category of dishes is:
2. For Dish Category A under the characteristics of the pure dish area map of the sample retention category dishes, the corresponding dish picture is: image tmp ; The representative feature is: featureT (1,1536) ; 3. Find the feature similarity of dish A under the pure dish area map features of the current meal database and the reserved sample category dishes. From 1 and 2 above, we get:
4. The representative features of dish category A in the new dish library after weighted fusion are: Feature (1,1536) = (1 - weight) * featureA (1,1536) + weight * featureT (1,1536) ; where weight is the fusion weight of the feature. When similarity >= 0.7, weight = 0.5; when 0.5 <= similarity < 0.7, weight = 0.1; when similarity < 0.5, weight = 0.
4. The method for dish recognition based on dish sample retention according to claim 1, wherein The canteen menu database is specifically composed of a transformer-based network distillation backbone network, which trains a feature extraction model with learning and generalization capabilities. Through this model, a canteen menu database with feature retrieval corresponding to the canteen menu database is established.
5. A dish recognition system based on dish sample retention, characterized in that, include: Sample storage for dishes: Collect sample dishes, locate the coordinates of each dish in the sample dishes through the target detection network, and semantically segment each dish sub-image according to the coordinates of each dish to obtain a pure dish area image without tableware. After the pure dish area pictures are extracted with the metric learning model, they are compared with each dish category in the canteen dish database one by one, and classified according to the highest similarity score to obtain the corresponding category of each dish; According to the corresponding categories of each dish, the dishes are automatically added to the menu list of the meal, and the menu of the meal and the pure dish area map of each sample category are obtained and sent to the weighted fusion end; Weighted fusion end: Receive the recipe of the current meal and the pure dish area map of each sampled dish category, screen the corresponding dish picture information from the large cafeteria dish library to obtain the current meal dish library. After extracting the features of the pure dish area map of each sampled dish category, perform weighted fusion with the pictures in the current meal dish library to obtain a new dish library and send it to the settlement terminal; Settlement terminal: Collect the pictures of the current meal dishes, use the object detection algorithm to locate the position coordinates of each dish in the current meal dish picture, and cut each dish sub-picture according to the position coordinates of each dish to obtain the dish sub-picture containing tableware; After extracting the features of the dish sub-picture containing tableware through the metric learning model, compare the features with the current meal dish pictures in the new dish library one by one. If the dishes corresponding to each category are identified, settlement will be carried out.
6. The dish recognition system based on dish sample retention according to claim 5, wherein, The settlement terminal also extracts the features of the dish sub-picture containing tableware through the metric learning model, and compares the features with the current meal dish pictures in the new dish library one by one. If the dishes corresponding to each category cannot be identified, record and accumulate the error times of each dish and the features of the error dish pictures, store the features of these error samples to obtain an error picture feature library, set the error picture feature library as valid for the current meal, and automatically clear it after the meal ends; Judge the error times of each category of dishes. If it is greater than the preset error times threshold, send the features of the dish pictures of the error category to the weighted fusion terminal to be fused with the new dish library.
7. The dish recognition system based on dish sample retention according to claim 5, wherein The weighted fusion calculation of the weighted fusion terminal is as follows. Let the dish category be A:
1. For the dish category A under the current meal dish library, the corresponding dish pictures are: image1, …, image N (where N is the number of pictures, N <= 5), and the corresponding features of each picture are: feature (1,1536) , …, featureN (1,1536) (where N is the number of pictures, N <= 5, and the dimension of each feature feature is 1536), then the representative feature of this category of dishes is:
2. For dish category A under the feature of the pure dish area map of the sample-keeping category dishes, the corresponding dish picture is: image tmp ; The representative feature is: featureT (1,1536) ; 3. Calculate the feature similarity similarity of dish A under the features of the current meal dish library and the pure dish area map of the sampled dish category. From the above 1 and 2, we get:
4. For dish category A in the new dish library, the representative feature obtained through weighted fusion is: Feature (1,1536) =(1 - weight)featureA (1,1536) + weight*featureT (1,1536) ; weight is the fusion weight of the feature. When similarity >= 0.7, weight = 0.5; when 0.5 <= similarity < 0.7, weight = 0.1; when similarity < 0.5, weight = 0.
8. The dish recognition system based on dish sample retention according to claim 5, characterized in that, The specific cafeteria dish library of the dish sampling terminal is a large cafeteria dish library that trains a feature extraction model with learning and generalization capabilities based on the distillation backbone network of the transformer network, and establishes a feature retrieval corresponding to the cafeteria dish library through this model.
Citation Information
Patent Citations
Open dish identification method based on deep learning target detection and metric learning
CN112115906A
Linkage control method, system and device for food sample reservation
CN112184494A