A method and device for assisting dish recognition based on tableware information

By combining target detection and tableware recognition algorithms, tableware information is used to assist dish recognition, the problems of low recognition accuracy and poor generalization ability caused by high dishes are solved, and efficient and accurate dish recognition and correction are achieved to adapt to the variety of dishes in different scenarios.

CN119107639BActive Publication Date: 2025-07-22GUANGZHOU PAIKEPUSHI INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411132514.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2025-07-22
Estimated Expiration
2044-08-16

AI Technical Summary

Technical Problem

In the prior art, there are problems in the identification of dishes with low recognition accuracy and poor generalization ability due to high similarity of dishes. Especially when different seasons and meal segments change, traditional visual methods are greatly affected by the environment, while deep learning methods are difficult to distinguish dishes with the same cooking methods.

Method used

By combining target detection, dish recognition and tableware recognition algorithms, tableware information is used to assist dish recognition, tableware ID is bound to improve recognition accuracy, tableware shape and color characteristics are used to correct dish categories, and secondary correction is carried out in combination with target detection and tableware recognition algorithm.

Benefits of technology

It improves the accuracy and generalization ability of dish recognition, reduces labor costs, improves recognition speed and robustness, adapts to the diversity of dishes in different scenarios, and ensures the user's dining experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119107639B_ABST
    Figure CN119107639B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for assisting dish recognition based on tableware information. The method includes: using a target detection algorithm to locate the position of the dish in the target dish picture, and using a dish recognition algorithm to extract features of the located dish to obtain the dish category to be entered, calculating the similarity value between the dish category to be entered and the dishes in the dish set, and comparing whether the similarity value is less than the similarity threshold; if so, storing the dish category to be entered into the database; if not, storing the dish category to be entered and the bound tableware ID into the database; when recognizing the dish, determining whether there is a category with a bound tableware ID in the category of the recognition result information; when there is a dish category bound by the tableware ID in the dish set, determining whether the dish category bound by the tableware ID exists in the recognition result information; if so, modifying the dish information category; if not, not modifying; the present invention can assist in the secondary recognition and correction of dishes through the tableware recognition algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dish recognition, and more particularly to a method and device for assisting dish recognition based on tableware information. Background Art

[0002] With the development of modern technology, the settlement methods in group meal scenarios such as original restaurants and canteens have tended to be phased out. This is mainly because the manual settlement method has problems such as high labor cost input and low efficiency, resulting in a poor dining experience for users. With the development of modern vision-related technologies, a dish recognition and settlement system based on vision and deep learning has been introduced. In the current technology field, significant progress has been made in dish recognition technology. These technologies are mainly divided into two categories: namely, the method based on traditional vision and the method based on deep learning. First of all, in the method of traditional vision, the dish recognition technology relies on image processing technology. First, by analyzing feature information such as the color, shape, and texture of the dish as dish information, and then calculating the similarity between feature classes to judge the dish category name; the dish recognition technology based on the deep learning network is to collect data to train the model. This model has a good feature expression ability for different dishes. By calculating the similarity between classes of the features obtained by the model, the category name of the dish can be judged.

[0003] The dish recognition algorithm based on traditional image processing is affected by various objective factors such as the lighting of the on-site environment and the actual color of the dish, and has poor generalization ability for the complex and diverse dishes in different canteens; and due to the changes in the number of dish categories in different seasons and meal periods, its adaptability will also decrease. In addition, factors such as the replacement of chefs that affect the increase in the similarity between dishes will also make the recognition effect worse; while the dish recognition based on deep learning can effectively reduce the influence brought by on-site environmental factors and has a certain degree of generalization for the complex and diverse dishes in the canteen, but it still cannot avoid the fact that different dishes with the same cooking method are difficult to distinguish from the dish images due to their high similarity. Although the correct category exists in the returned TopN (N≥5) when using deep neural network recognition, it can be corrected by the method of manually correcting the category name, but this will undoubtedly increase the labor cost. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method and device for assisting dish recognition based on tableware information. When the similarity between dishes is extremely high, the shape and color of the tableware bound to the dishes are recognized by the tableware recognition algorithm to improve the accuracy of dish recognition, so as to solve the deficiencies in the prior art.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A method for assisting dish recognition based on tableware information, including the following steps:

[0007] S1. When entering dishes, obtain the target dish image to be warehoused;

[0008] Use the target detection algorithm to locate the position of the dish in the target dish image to obtain the coordinate area of the dish, and use the dish recognition algorithm to extract the features of the dish in the coordinate area to obtain the dish category to be entered;

[0009] S2. Read the dish set in the dish general library, and calculate the similarity value between the dish category to be entered and each dish category in the dish set;

[0010] Compare each of the similarity values with each similarity threshold in the preset threshold set to obtain the corresponding comparison results;

[0011] When the similarity value is less than the similarity threshold, the dish category to be entered is warehoused, otherwise execute step S3;

[0012] S3. When the similarity value is greater than the similarity threshold, bind the dish category to be entered with a tableware ID, use the tableware recognition algorithm to extract the features of the tableware ID bound to the dish category to be entered, and warehouse the dish category to be entered and the bound tableware ID;

[0013] S4. When recognizing a dish, use the dish recognition algorithm to match and recognize the features of the dish category to be recognized and each dish category in the dish set, and output the recognition result information TopN;

[0014] Judge whether there is a category with a bound tableware ID in the categories of the recognition result information TopN; if not, directly output the recognition result; if so, use the tableware recognition algorithm to recognize the tableware ID, and judge whether there is a dish category bound to the tableware ID in the dish set;

[0015] When there is no dish category bound to the tableware ID in the dish set, output the dish recognition result;

[0016] When there is a dish category bound to the tableware ID in the dish set, then judge whether the dish category bound to the tableware ID exists in the recognition result information TopN. If so, modify the recognition result information Top1 to the dish information corresponding to the tableware ID, otherwise, do not modify the dish recognition result information TopN.

[0017] Further, in step S1, using the target detection algorithm to locate the position of the dish in the target dish image to obtain the coordinate area of the dish, and using the dish recognition algorithm to extract the features of the dish in the coordinate area to obtain the dish category to be entered, specifically:

[0018] Locate the position information of the dishes in the target dish picture based on the target detection model, obtain the coordinate area of the dishes, and cut the picture within the coordinate area into sub-pictures;

[0019] Extract features from each sub-picture based on the dish recognition model, and determine the dish categories to be entered based on the extracted features.

[0020] Further, in step S2, read the dish set in the dish total library, and calculate the similarity value between the dish categories to be entered and each dish category in the dish set. The calculation method of the similarity value is as follows:

[0021] Read the feature information of N categories in the dish set of the dish total library. The calculation formula of the similarity value is:

[0022]

[0023] In the above formula, features j refers to the feature of the j-th dish category in the total library; refers to the i-th feature value of the j-th category of dishes in the total library; featuresA refers to the feature of the dish category to be entered, and featuresA i refers to the i-th feature value of the dish category to be entered; refers to saving the similarity values between all dish category features in the dish total library and the features of the dishes to be entered; similarityVector is the storage vector of all similarity values, with a dimension of N.

[0024] Further, the target detection model is a trained YoloV8 target detection network model. The specific training process of the YoloV8 target detection network model includes:

[0025] Collect a large number of dish data sets with different lighting conditions and different backgrounds for target box annotation, and divide the annotated dish data sets into a training set and a validation set according to a ratio;

[0026] Take the YoloV8 network as the basic network, and use the training set to train the basic network to obtain the target detection model;

[0027] Use the validation set to verify the target detection model, and select the optimal target detection model.

[0028] Further, the training of the basic network using the training set is specifically:

[0029] Initialize the basic network, and perform data augmentation on the dish pictures in the training set;

[0030] Optimize the comprehensive loss function containing classification and bounding box regression losses in the dish pictures after data augmentation using the SGD optimization algorithm;

[0031] The data augmentation includes: rotation, scaling, and color jitter.

[0032] Furthermore, the dish recognition model is a trained student model, and the specific training process of the student model includes:

[0033] Based on the optimal object detection model, locate the regional positions of the dishes in the dish dataset, cut the pictures of the dishes within the regional positions into sub - pictures, label the dish category names of the sub - pictures, and divide the labeled dish category names of the sub - pictures into a training set and a validation set according to a ratio;

[0034] Construct a Vision Transformer model as the teacher model, and use the training set to train the teacher model to obtain the output result of the teacher model;

[0035] Construct a ResNet backbone network as the student model, use the pre - trained weights of the teacher model to train the student model, and use the training set to train the student model to obtain the output result of the student model;

[0036] Jointly optimize the student model through multiple iterations of classification loss and distillation loss. When the difference between the teacher model and the student model is minimized, save the optimal student model. Specifically:

[0037] Use the classification loss to evaluate the difference between the output result of the student model and the dish category names of the labeled sub - pictures;

[0038] Use the distillation loss to evaluate the difference between the output result of the student model and the output result of the teacher model;

[0039] Extract features from the pictures of the sub - pictures within the regional positions based on the optimal student model to obtain the dish information after feature extraction;

[0040] Calculate the similarity value between the dish information after feature extraction and the dish features in the dish dataset, and select the category with the highest similarity value as the recognition result.

[0041] Furthermore, in step S3, the tableware recognition algorithm is executed using a trained tableware positioning and segmentation model. The training process of the tableware positioning and segmentation model is specifically as follows:

[0042] Obtain a number of tableware pictures under different lighting conditions and backgrounds, and divide the number of tableware pictures into a training set and a validation set according to a ratio;

[0043] Use the training set to train the YOLACT algorithm model to obtain the tableware positioning and segmentation model;

[0044] Verify the tableware positioning and segmentation model based on the validation set, and select the optimal tableware positioning and segmentation model.

[0045] Furthermore, in step S3, use the tableware recognition algorithm to extract the features of the tableware ID bound to the category of the dish to be entered, specifically including:

[0046] Use the optimal tableware positioning and segmentation model to locate the tableware image in the target dish image, and extract the contour of the tableware image.

[0047] Use the Canny algorithm and dilation and erosion methods in traditional image algorithms to optimize the contour of the tableware image extracted by the tableware positioning and segmentation model, and obtain the optimized contour map of the tableware image; and extract the edge feature information of the optimized contour map of the tableware image.

[0048] Match the edge feature information of the contour map of the tableware image with the edge feature information of the tableware already in the library, calculate the gradient cosine similarity value between the two, and store the edge feature information of the contour map of the tableware image at the maximum matching score value at the current angle in the gradient cosine similarity value.

[0049] Compare the maximum matching score value at the current angle with the global maximum matching score value. If the maximum score value at the current angle is greater than the global maximum matching score value, then use the ID corresponding to the tableware image as the recognition result.

[0050] Furthermore, the extraction of the contour of the tableware image is specifically: the shape and color of the tableware image.

[0051] The present invention provides a device for assisting dish recognition based on tableware information, including:

[0052] An acquisition module, which scans the dish to be entered through a camera device to obtain a target dish image;

[0053] A target detection algorithm module, which is used to locate the position of the dish in the target dish image to obtain the coordinate area of the dish;

[0054] A dish recognition algorithm module, which is used to extract the features of the dish in the coordinate area to obtain the category of the dish to be entered;

[0055] A tableware recognition algorithm module, which is used to extract the features of the tableware ID bound to the category of the dish to be entered;

[0056] A classification processing module, which is used to calculate the similarity value between the category of the dish to be entered and each dish category in the dish set;

[0057] Compare each of the similarity values with each similarity threshold in a preset threshold set to obtain corresponding comparison results.

[0058] The present invention also provides an electronic device, including a memory and a processor. The memory is used to store a computer program, and the computer program is configured to be executed by the processor. When the computer program is executed by the processor, a method for assisting dish recognition based on tableware information is implemented.

[0059] According to the specific embodiments provided by the present invention, the present invention has the following technical effects:

[0060] In this application, when a dish is entered, the position of the dish in the target dish picture is located through a target detection algorithm, and the features of the dish are extracted using a dish recognition algorithm to obtain the category of the dish to be entered. Then, it is compared with the categories in the dish set respectively. When the category is not in the dish set, the tableware ID is recognized through a tableware recognition algorithm and bound to the dish to be recognized, and then it is stored in the database. When recognizing, if the dish to be recognized is not in the dish set, the tableware ID is recognized through the tableware recognition algorithm, and then the category of the dish to be recognized is obtained. The method of assisting dish recognition through tableware can adapt to dish recognition in different scenarios, and can effectively improve the recognition accuracy of dishes. And through the tableware recognition algorithm, it can effectively assist in the secondary recognition and correction of dishes. Moreover, without adding other hardware and other costs, it can effectively improve the dish recognition rate and its generalization ability, and can also ensure that its recognition speed is not affected, further improving the recognition ability under different scenarios and dish diversity, effectively enhancing the robustness and accuracy of the algorithm, and being compatible to ensure the user's dining experience in terms of recognition accuracy. Description of the Drawings

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0062] The following further illustrates the method and device for assisting dish recognition based on tableware information of the present invention with reference to the drawings;

[0063] Figure 1 is the overall flowchart in the method for assisting dish recognition based on tableware information provided by the present invention. Detailed Embodiments

[0064] The following further describes in detail the specific embodiments of the present invention with reference to the drawings and embodiments. The following embodiments are used to illustrate the present invention but not to limit the scope of the present invention.

[0065] To better understand the purpose, structure, and function of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.

[0066] The present invention provides a method for assisting dish recognition based on tableware information, including the following steps:

[0067] S1. When entering a dish, use a camera device to scan the dish to be entered to obtain a target dish picture; use a target detection algorithm to locate the position of the dish in the target dish picture to obtain the coordinate area of the dish, and use a dish recognition algorithm to extract the features of the dish in the coordinate area to obtain the dish category to be entered;

[0068] It should be noted that: the position coordinates of the dish target are located as ROI(x, y, w, h) through the trained YoloV8 target detection network algorithm. If the target box cannot be correctly detected, the corresponding dish is boxed manually.

[0069] S2. Read the dish set in the dish library and calculate the similarity value between the dish category to be entered and each dish category in the dish set;

[0070] Compare each of the similarity values with each similarity threshold in the preset threshold set to obtain the corresponding comparison results;

[0071] When the similarity value is less than the similarity threshold, the dish category to be entered is stored in the database, otherwise, execute step S3;

[0072] S3. When the similarity value is greater than the similarity threshold, bind the dish category to be entered with a tableware ID, use a tableware recognition algorithm to extract the features of the tableware ID bound to the dish category to be entered, and store the dish category to be entered and the bound tableware ID in the database;

[0073] S4. When recognizing a dish, use a dish recognition algorithm to match and recognize the features of the dish category to be recognized and each dish category in the dish set, and output the recognition result information TopN;

[0074] Judge whether there is a category with a bound tableware ID in the categories of the recognition result information TopN; if not, directly output the recognition result; if so, use a tableware recognition algorithm to recognize the tableware ID, and judge whether there is a dish category bound to the tableware ID in the dish set;

[0075] It should be noted that: in the integrated settlement, the recognition results of all dishes are integrated, and the corresponding dish categories are output for settlement.

[0076] When the dish category bound by the tableware ID does not exist in the dish set, output the dish recognition result;

[0077] When the dish category bound by the tableware ID exists in the dish set, then determine whether the dish category bound by the tableware ID exists in the TopN of the recognition result information. If it is, modify the Top1 of the recognition result information to the dish information corresponding to the tableware ID. Otherwise, do not modify the dish recognition result information TopN.

[0078] The target detection model is a trained YoloV8 target detection network model. The specific training process of the Yo1oV8 target detection network model includes:

[0079] Collect a large number of dish data sets with different lighting conditions and different backgrounds for target box annotation, and divide the annotated dish data sets into a training set and a validation set according to a ratio;

[0080] Use the YoloV8 network as the basic network, and use the training set to train the basic network to obtain the target detection model;

[0081] Use the validation set to verify the target detection model and select the optimal target detection model.

[0082] The training of the basic network using the training set is specifically as follows: Initialize the basic network and perform data augmentation on the target dish pictures in the training set;

[0083] Use the SGD optimization algorithm to optimize the comprehensive loss function containing classification and bounding box regression losses in the data-augmented target dish pictures;

[0084] The data augmentation includes: rotation, scaling, and color jitter.

[0085] It should be noted that when using the YoloV8 network architecture to train the dish detection model, first load the pre-trained model for initialization to accelerate the convergence speed, and perform data augmentation (such as rotation, scaling, color jitter, etc.) on the pictures in the training data set to improve the robustness of the model. Use the SGD optimization algorithm to optimize the comprehensive loss function containing classification and bounding box regression losses.

[0086] The dish recognition model is a trained student model. The specific training process of the student model includes:

[0087] Based on the optimal target detection model, locate the regional positions of the dishes in the dish data set and cut the pictures of the dishes in the regional positions into sub-pictures, then label the dish category names of the sub-pictures, and divide the labeled dish category names of the sub-pictures into a training set and a validation set according to a ratio;

[0088] Build a Vision Transformer model as the teacher model, and use the training set to train the teacher model to obtain the output results of the teacher model;

[0089] It should be noted that the selected target position of the dish is cut into sub-images as the target images of the dishes to be stored in the library. Use the ResNet backbone network based on Transformer network distillation to train the metric learning model DML that can learn and extract dish features, and extract the features of the dishes in the sub-images through this dish recognition model (DML).

[0090] Build a ResNet backbone network as the student model, use the pre-trained weights of the teacher model to train the student model, and use the training set to train the student model to obtain the output results of the student model;

[0091] It should be noted that: Select the ResNet backbone network as the student model and initialize it with pre-trained weights to speed up the training process;

[0092] Jointly optimize the student model through multiple iterations of the classification loss and the distillation loss. When the difference between the teacher model and the student model is minimized, save the optimal student model. Specifically:

[0093] Use the classification loss to evaluate the difference between the output results of the student model and the dish category names of the labeled sub-images;

[0094] Use the distillation loss to evaluate the difference between the output results of the student model and the output results of the teacher model;

[0095] Extract the features of the sub-image pictures in the area based on the optimal student model to obtain the dish information after feature extraction;

[0096] Calculate the similarity value between the dish information after feature extraction and the dish features in the dish dataset, and select the category with the highest similarity value as the recognition result.

[0097] It should be noted that: Evaluate the performance of the student model on the validation set and perform fine-tuning to ensure that its dish feature extraction ability is optimized. Select the optimal trained feature extraction model as the dish metric learning model DML. The dish features extracted by this model have a dimension of 1X1536, that is, featuresA (1×1536) 。

[0098] The optimal student model is the dish recognition model. Based on the dish recognition model, extract the features of the dishes in the coordinate area to obtain the dish information after feature extraction, and calculate the similarity value between the dish information after feature extraction and the dish features in the dish total library dataset, and select the category with the highest similarity value as the recognition result.

[0099] It should be noted that: if there is no original general library dish category in the similarity value similarityVector that is greater than the set similarity threshold similarThre (the default similarThre = 0.8), the features corresponding to this dish category are directly stored in the database; otherwise, the following steps are executed.

[0100] In step S3, the tableware recognition algorithm is executed using the trained tableware positioning and segmentation model. The training process of the tableware positioning and segmentation model is specifically as follows:

[0101] Obtain a number of tableware pictures under different lighting conditions and backgrounds, and divide the number of tableware pictures into a training set and a validation set according to a ratio.

[0102] Use the training set to train the YOLACT algorithm model to obtain the tableware positioning and segmentation model.

[0103] Based on the validation set, verify the tableware positioning and segmentation model and select the optimal tableware positioning and segmentation model.

[0104] In step S3, using the tableware recognition algorithm to extract the features of the tableware ID bound to the dish category to be entered specifically includes:

[0105] Use the optimal tableware positioning and segmentation model to locate the tableware pictures in the target dish picture and extract the contours of the tableware pictures.

[0106] Use the Canny algorithm and dilation and erosion methods in traditional image algorithms to optimize the contours of the tableware pictures extracted by the tableware positioning and segmentation model to obtain the optimized contour map of the tableware pictures; and extract the edge feature information of the optimized contour map of the tableware pictures.

[0107] Based on the edge feature information of the contour map of the tableware picture, match it with the edge feature information of the already stored tableware, calculate the gradient cosine similarity value between the two, and store the edge feature information of the contour map of the tableware picture at the maximum matching score value at the current angle in the gradient cosine similarity value.

[0108] Compare the maximum matching score value at the current angle with the global maximum matching score value. If the maximum score value at the current angle is greater than the global maximum matching score value, the ID corresponding to the tableware picture is used as the recognition result.

[0109] The tableware picture to be recognized includes: the shape and color of the tableware.

[0110] In step S2, when reading the dish set in the dish general library and calculating the similarity value between the dish category to be entered and each dish category in the dish set, the calculation method of the similarity threshold is as follows:

[0111] The dish recognition algorithm extracts features of the dishes within the coordinate region, and the dimension of the feature extraction is 1×1536;

[0112] Read the feature information of N categories in the dataset of the dish library, and the calculation formula of the similarity threshold is:

[0113]

[0114] In the above formula, Teatures j refers to the feature of the j-th dish category in the library; refers to the i-th feature value of the j-th category of dishes in the library; featuresA refers to the feature of the dish category to be entered, featuresA i refers to the i-th feature value of the dish category to be entered;, refers to saving the similarity values between all dish category features in the dish library and the features of the dish to be entered; similarityVector is the storage vector of all similarity values, with a dimension of N.

[0115] The present invention also provides a device for assisting dish recognition based on tableware information, including:

[0116] An acquisition module that scans the dish to be entered through a camera device to obtain a target dish picture;

[0117] A target detection algorithm module for positioning the position of the dish in the target dish picture to obtain the coordinate region of the dish;

[0118] A dish recognition algorithm module for extracting features of the dish within the coordinate region to obtain the dish category to be entered;

[0119] A tableware recognition algorithm module for extracting the features of the tableware ID bound to the dish category to be entered;

[0120] A classification processing module for calculating the similarity values between the dish category to be entered and each dish category in the dish set;

[0121] Compare each of the similarity values with each similarity threshold in the preset threshold set to obtain the corresponding comparison results.

[0122] The present invention also provides an electronic device, including a memory and a processor, where the memory is used to store a computer program, and the computer program is configured to be executed by the processor. When the computer program is executed by the processor, it implements the method for assisting dish recognition based on tableware information.

[0123] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for assisting dish recognition based on tableware information, characterized in that Including the following steps: S1. When entering dishes, obtain the target dish image to be warehoused; Use the target detection algorithm to locate the position of the dish in the target dish image to obtain the coordinate area of the dish, and use the dish recognition algorithm to extract the features of the dish in the coordinate area to obtain the dish category to be entered; S2. Read the dish set in the dish total library and calculate the similarity value between the dish category to be entered and each dish category in the dish set; Compare each of the similarity values with each similarity threshold in the preset threshold set to obtain the corresponding comparison result; When the similarity value is less than the similarity threshold, the dish category to be entered is warehoused, otherwise, execute step S3; S3. When the similarity value is greater than the similarity threshold, bind the dish category to be entered with a tableware ID, use the tableware recognition algorithm to extract the features of the tableware ID bound to the dish category to be entered, and warehouse the dish category to be entered and the bound tableware ID; S4. When identifying a dish, use the dish recognition algorithm to match and identify the features of the dish category to be identified and each dish category in the dish set, and output the recognition result information TopN; Judge whether there is a category with a bound tableware ID in the categories of the recognition result information TopN; if not, directly output the recognition result; If so, use the tableware recognition algorithm to identify the tableware ID, and judge whether there is a dish category bound to the tableware ID in the dish set; When there is no dish category bound to the tableware ID in the dish set, output the dish recognition result; When there is a dish category bound to the tableware ID in the dish set, judge whether the dish category bound to the tableware ID exists in the recognition result information TopN. If so, modify the recognition result information Top1 to the dish information corresponding to the tableware ID, otherwise, do not modify the dish recognition result information TopN.

2. The method for assisting dish recognition based on tableware information according to claim 1, wherein In step S1, using the target detection algorithm to locate the position of the dish in the target dish image to obtain the coordinate area of the dish, and using the dish recognition algorithm to extract the features of the dish in the coordinate area to obtain the dish category to be entered, specifically: Based on the target detection model, locate the position information of the dish in the target dish image to obtain the coordinate area of the dish, and cut the image in the coordinate area into sub-images; Based on the dish recognition model, extract the features of each sub-image, and determine the dish category to be entered based on the extracted features.

3. The method for assisting dish recognition based on tableware information according to claim 1, wherein In step S2, when reading the dish set in the dish total library and calculating the similarity value between the dish category to be entered and each dish category in the dish set, the calculation method of the similarity value is: Read the feature information of N categories in the dish set in the dish total library, and the calculation formula of the similarity value is: In the above formula, features j refers to the j-th dish category feature in the general database; refers to the i-th feature value of the j-th category of dishes in the general database; featuresA refers to the dish category feature to be entered, featuresA i refers to the i-th feature value of the dish category to be entered; refers to the similarity values of all dish category features in the saved dish general database and the features of the dish to be entered; similarityVector is the storage vector of all similarity values, with a dimension of N.

4. The method for assisting dish recognition based on tableware information according to claim 2, wherein The target detection model is a trained YoloV8 target detection network model, and the specific training process of the YoloV8 target detection network model includes: Collect a large number of dish data sets with different lighting conditions and different backgrounds for target box annotation, and divide the annotated dish data sets into a training set and a validation set according to a ratio; Using the YoloV8 network as the base network, training the base network with the training set to obtain an object detection model; Validating the object detection model with the validation set and selecting the optimal object detection model.

5. The method for assisting dish recognition based on tableware information according to claim 4, characterized in that The training of the base network using the training set is specifically as follows: Initializing the base network and performing data augmentation on the dish pictures in the training set; Using the SGD optimization algorithm to optimize the comprehensive loss function including classification and bounding box regression losses in the data-augmented dish pictures; The data augmentation includes: rotation, scaling, and color jitter.

6. The method for assisting dish recognition based on tableware information according to claim 3, wherein The dish recognition model is the trained student model, and the specific training process of the student model includes: Based on the optimal object detection model, locating the regional positions of the dishes in the dish dataset and cutting the pictures of the dishes within the regional positions into sub-pictures, then annotating the dish category names of the sub-pictures, and dividing the dish category names of the annotated sub-pictures into a training set and a validation set according to a ratio; Constructing a Vision Transformer model as the teacher model, training the teacher model with the training set to obtain the output result of the teacher model; Constructing a ResNet backbone network as the student model, training the student model with the pre-trained weights of the teacher model, and training the student model with the training set to obtain the output result of the student model; Jointly optimizing the student model through multiple iterations of classification loss and distillation loss. When the difference between the teacher model and the student model is minimized, save the optimal student model. Specifically: Using the classification loss to evaluate the difference between the output result of the student model and the dish category names of the marked sub-pictures; Using the distillation loss to evaluate the difference between the output result of the student model and the output result of the teacher model; Extracting the features of the pictures of the sub-pictures within the regional positions based on the optimal student model to obtain the dish information after feature extraction; Calculating the similarity value between the dish information after feature extraction and the dish features in the dish dataset, and selecting the category with the highest similarity value as the recognition result.

7. The method for assisting dish recognition based on tableware information according to claim 1, characterized in that In step S3, the tableware recognition algorithm is executed using the trained tableware positioning and segmentation model. The training process of the tableware positioning and segmentation model is specifically as follows: Obtaining a number of tableware pictures under different lighting conditions and backgrounds, and dividing the number of tableware pictures into a training set and a validation set according to a ratio; Training the YOLACT algorithm model with the training set to obtain a tableware positioning and segmentation model; Validating the tableware positioning and segmentation model based on the validation set and selecting the optimal tableware positioning and segmentation model.

8. The method for assisting dish recognition based on tableware information according to claim 7, wherein In step S3, extracting the features of the tableware ID bound to the dish category to be entered using the tableware recognition algorithm specifically includes: Using the optimal tableware positioning and segmentation model to locate the tableware pictures in the target dish picture and extract the contours of the tableware pictures; Using the Canny algorithm and dilation and erosion methods in traditional image algorithms to optimize the contours of the tableware pictures extracted by the tableware positioning and segmentation model to obtain an optimized contour map of the tableware pictures; and extracting the edge feature information of the optimized contour map of the tableware pictures; Match the edge feature information of the contour map of the tableware picture with the edge feature information of the tableware already stored in the database, calculate the gradient cosine similarity value between the two, and store the edge feature information of the contour map of the tableware picture with the maximum matching score value at the current angle in the gradient cosine similarity value; Compare the maximum matching score value at the current angle with the global maximum matching score value. If the maximum score value at the current angle is greater than the global maximum matching score value, use the ID corresponding to the tableware picture as the recognition result.

9. The method for assisting dish recognition based on tableware information according to claim 7, wherein The extraction of the contour of the tableware picture is specifically: the shape and color of the tableware picture.

10. An apparatus for assisting dish recognition based on tableware information, which is used to execute the method for assisting dish recognition based on tableware information according to any one of claims 1-9, characterized in that, Including: An acquisition module that scans the dish to be entered through a camera device to obtain a target dish picture; A target detection algorithm module for positioning the position of the dish in the target dish picture to obtain the coordinate area of the dish; A dish recognition algorithm module for extracting the features of the dish in the coordinate area to obtain the category of the dish to be entered; A tableware recognition algorithm module for extracting the features of the tableware ID bound to the category of the dish to be entered; A classification processing module for calculating the similarity value between the category of the dish to be entered and each dish category in the dish set; Compare each of the similarity values with each similarity threshold in the preset threshold set to obtain the corresponding comparison result.

Citation Information

Patent Citations

  • Intelligent catering settlement system based on RFID fusion dish image matching identification

    CN110852733A

  • Tableware management method for intelligent catering

    CN112862634A