Artificial intelligence detection method for distinguishing optical discs and leftovers in dinner taking plate
By using high-frame-rate depth cameras and deep learning technology, combined with Fourier transform, the precise distinction between the discs and leftovers on the meal plate is achieved, solving the inaccuracy and instability of traditional technologies when monitoring the food consumption in the meal plate, and improving the accuracy and stability of identification.
Patent Information
- Application Number
- CN202510051218.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-16
AI Technical Summary
Traditional manual observation and primary image recognition technologies have problems of inaccuracy, instability and high cost when monitoring food consumption in the meal tray in the catering industry, and are greatly affected by light and background environment.
High frame rate and high resolution depth cameras are used to collect point cloud data, combined with Fourier transform and deep learning technology, and through point cloud preprocessing, image segmentation and detection, multi-path result fusion, etc., the precise distinction between the discs and leftovers on the meal plate is achieved.
It improves the accuracy and stability of identification, reduces the risk of false alarms and underreports, maintains efficient identification performance in complex catering environments, and reduces dependence on lighting and background environments.
Smart Images

Figure CN120013876A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision, deep learning and three-dimensional data processing technology, and specifically to an artificial intelligence detection method for distinguishing between CDs and leftovers in a meal tray. Background Art
[0002] In the catering industry, accurate and efficient monitoring of food consumption on plates is a crucial task. It is not only related to the rational use of resources, but also can effectively reduce food waste, optimize catering service processes, and improve customer satisfaction and corporate operating efficiency. However, traditionally, this monitoring process mainly relies on manual observation and recording, or uses relatively basic image recognition technology. These methods have exposed many limitations in practical applications;
[0003] Although manual observation is intuitive, it is easily affected by human factors, such as subjective judgment, fatigue, and limited observation time of the observer, which may lead to inaccurate and unstable data. In addition, manual monitoring is costly, especially in large catering venues or fast food chains, where it is difficult to achieve timely and comprehensive monitoring of every plate.
[0004] On the other hand, although early image recognition technologies have achieved automation to a certain extent, their recognition accuracy is often unsatisfactory. These technologies are easily affected by changes in lighting conditions. For example, strong light, weak light, or uneven light may cause image quality to deteriorate, thus affecting the recognition effect. At the same time, complex background environments, such as the diversity of tableware, tabletop textures, and occlusions from other objects, will also bring additional challenges to image recognition and increase the risk of misidentification.
[0005] Therefore, we propose an artificial intelligence detection method for distinguishing between clean plates and leftovers in meal trays. Summary of the invention
[0006] The purpose of the present invention is to provide an artificial intelligence detection method for distinguishing between clean plates and leftovers in a meal tray, thereby solving the problems raised in the background technology.
[0007] To achieve the above-mentioned purpose, the present invention provides the following technical solution: an artificial intelligence detection method for distinguishing between clean plates and leftovers in a meal tray, comprising the following method steps:
[0008] Step 1: Data collection: Use a point cloud collection device to collect point cloud data. The device includes a high frame rate depth camera module and a data processing module. Each frame of point cloud data is immediately transferred to the processing module for pre-processing after collection to ensure real-time performance.
[0009] Step 2: Point cloud preprocessing
[0010] Noise filtering: Calculate the maximum and minimum values of each dimension in each frame of point cloud data, filter out the noise point cloud that exceeds the range, and for each point (x, y, z), determine whether its coordinate value is within the set maximum and minimum range, if not, filter it out;
[0011] Through filtering: limit the point cloud range, retain only valid points, set the range of X, Y, and Z dimensions, and only retain point cloud data within the range;
[0012] Projection generates depth image: Get the maximum value of the valid point cloud in the Z direction, project the point cloud to the XY plane to form a depth image;
[0013] Step 3: Image segmentation and detection
[0014] Path 1: Traditional method processing
[0015] Mean filtering: Perform mean filtering on the depth image to smooth out noise and fill in abnormal points;
[0016] Edge detection: Use the Canny edge detection algorithm to extract the edge contour of the disc;
[0017] Hough circle transform: The center area of the circle is detected by Hough circle transform, and all the plate data on the plate are detected;
[0018] Characteristic differentiation: Distinguish between CDs and leftovers based on their different characteristics;
[0019] Path 2: Deep Learning Semantic Segmentation Detection
[0020] Data annotation: perform CD annotation on the depth image to facilitate subsequent depth model training;
[0021] Semantic segmentation: Input the depth image to the DeeplabV3+ semantic segmentation model, and the model returns the candidate detection results of the CD area;
[0022] Result correction: Use the detection results of the deep learning model to correct the detection results of the traditional method;
[0023] Step 4: Multi-path result fusion
[0024] Calculate IOU: Calculate the intersection over union (IOU) of the disc area detected by the traditional method and the deep learning method;
[0025] Result fusion: Use deep learning methods to correct the non-disc areas detected by traditional methods;
[0026] Step 5: Result processing and output
[0027] Prevent duplicate detection: limit the acquisition rate of the depth camera to prevent repeated acquisition of the same data;
[0028] Feature analysis and identification: After extracting the features of CDs and leftovers according to the above method, analyze the feature data, set a reasonable threshold to distinguish them, and realize CD identification;
[0029] Result saving and report generation: Save the recognition results in the form of image annotations and generate statistical analysis reports to facilitate subsequent processing and use.
[0030] As a preferred implementation of the present invention, the Hough circle transform is specifically:
[0031] Multiple different thresholds are designed for the circular area detected by Hough transform, and multiple circular areas are screened out again. The circular area is feature extracted using Fourier transform, and the features corresponding to each circle are weighted and averaged to obtain the features of the final circular area.
[0032] As a preferred implementation of the present invention, the result fusion in step 4 is specifically as follows:
[0033] Determine the size of the IOU of the two regions and set a threshold σ;
[0034] If the IOU is greater than the threshold σ, it is determined to be an overlapping area, and the semantic segmentation result is preferentially used as the CD detection result.
[0035] As a preferred embodiment of the present invention, preventing duplication detection in step 5 is specifically as follows:
[0036] The radius, position and other statistics of the circular area detected by multiple adjacent data are performed, and a reasonable threshold is designed to determine whether the data on the adjacent frames is duplicate data. If it is duplicate data, the detection is skipped.
[0037] As a preferred embodiment of the present invention, during data acquisition, the depth camera acquires point cloud data and RGB image data sequences of the optical disk and the surrounding environment at a rate of 30 frames per second, and the data format is PLY.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] The present invention utilizes a high frame rate and high resolution depth camera, which can capture the subtle features of food on the plate, including shape, edge, and height information, and effectively excludes invalid data by accurately calculating the maximum and minimum values of each dimension of the point cloud data and applying the straight-through filtering technology, thereby ensuring the accuracy of the recognition result;
[0040] The method of the present invention is not limited by lighting conditions and background environment. The depth camera directly obtains the three-dimensional information of the object, avoiding the recognition error caused by the change of lighting in the traditional image recognition technology. At the same time, by performing pre-processing steps such as smoothing and edge detection on the depth image, the anti-interference ability of the system is further enhanced, so that the stable recognition performance can be maintained even in a complex and changeable catering environment.
[0041] Combining Fourier transform and deep learning technology, the present invention can automatically extract the features of optical disks and leftovers and effectively distinguish them. Fourier transform can capture the unique features of round dinner plates and leftovers in the frequency domain, while the deep learning model learns to recognize complex patterns and textures by training a large amount of sample data. The combination of the two greatly improves the intelligent level of recognition;
[0042] The present invention cleverly combines the feature extraction method based on Fourier transform and the semantic segmentation method based on deep learning. By setting appropriate thresholds, it achieves accurate distinction between optical disks and leftovers. This fusion strategy not only improves the accuracy of recognition, but also reduces the risks of false positives and false negatives, making the system more reliable and efficient in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0044] Figure 1 This is an operation diagram of an artificial intelligence detection method for distinguishing between clean plates and leftovers in a meal tray according to the present invention. DETAILED DESCRIPTION
[0045] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the present invention is further explained below in conjunction with specific implementation methods.
[0046] Example
[0047] An artificial intelligence detection method for distinguishing between clean plates and leftovers in a meal tray is specifically implemented as follows:
[0048] Data collection
[0049] A point cloud data and RGB image data sequence of the CD and its surroundings are collected by a point cloud acquisition device, which includes a high frame rate depth camera module and a data processing module. The depth camera collects point cloud data at a rate of 30 frames per second, and the data is in PLY format. In order to ensure the real-time processing, each frame of point cloud data is immediately transferred to the processing module for preprocessing after collection.
[0050] Point cloud preprocessing
[0051] Perform the following steps for each frame of point cloud data:
[0052] Calculate the maximum and minimum values of each dimension and filter out the noise point cloud:
[0053] x min =min(x i ),x max =max(x i )
[0054] y min =min(y i ),y max =max(y i )
[0055] z min =min(z i ),z max =max(z i )
[0056] where x i ,y i , z i is the coordinate value of each point in the point cloud.
[0057] Through filtering: limit the point cloud range and keep only valid points:
[0058] x min ≤x i ≤x max ,y min ≤y i ≤y max , z min ≤z i ≤z max
[0059] Get the maximum value z of the valid point cloud in the Z direction max , project the point cloud onto the XY plane to form a depth image:
[0060] Depth(x,y)=z max .
[0061] 3. Image segmentation and detection
[0062] The generated depth image is processed through two paths:
[0063] Path 1: Traditional processing
[0064] 1) Perform mean filtering on the depth image, smooth out the noise, and fill in the outliers:
[0065]
[0066] 2) Use Canny edge detection to extract the edge contour of the disc;
[0067] 3) Detect the center area of the circle through Hough circle transform, detect all the plate data on the plate, design multiple different thresholds for the circular area detected by Hough transform, such as 0.5, 0.7, 0.9, etc. (0.5 times the circle radius), screen out multiple circular areas again, use Fourier transform to extract features of the circular area, and perform weighted averaging on the features corresponding to each circle to obtain the features of the final circular area. Finally, distinguish between CDs and leftovers based on their different features.
[0068] Path 2: Deep Learning Semantic Segmentation Detection
[0069] In order to further improve the accuracy of CD detection, deep learning methods are used to assist in improving performance;
[0070] Disc labeling of depth images to facilitate subsequent depth model training;
[0071] The depth image is input to the DeeplabV3+ semantic segmentation model, and the model returns the candidate detection results of the CD area. The results are used to correct the detection results of the traditional method.
[0072] 4. Multi-path result fusion
[0073] Fusion of semantic segmentation results with traditional methods:
[0074] 1) Calculate the IOU (intersection over union) of the traditional method and deep learning method to detect the disc area:
[0075]
[0076] 2) The non-CD area detected by the traditional method is corrected using the deep learning method, the IOU size of the two areas is determined, and a threshold σ is set. If the IOU is greater than the threshold, it is determined to be an overlapping area, and the semantic segmentation result is preferentially used as the CD detection result.
[0077] 5. Result processing and output
[0078] 1) In order to prevent repeated detection of data, a variety of methods are used to process the data; when collecting data, the acquisition rate of the depth camera is limited to prevent repeated acquisition of the same data. In addition, the radius, position and other statistics of the circular areas detected by multiple adjacent data are performed, and a reasonable threshold is designed to determine whether the data on the adjacent frames is duplicate data. If it is duplicate data, the detection is skipped.
[0079] 2) After extracting the features of CDs and leftovers according to the above method, the feature data is analyzed. Since the features of CDs and leftovers are quite different, reasonable thresholds can be set to distinguish them, and CD recognition can be achieved in turn.
[0080] 3) Save the recognition results in the form of image annotations and generate a statistical analysis report for subsequent processing and use.
[0081] In summary, the present invention utilizes a high frame rate and high resolution depth camera, and can capture the subtle features of food on the plate, including shape, edge, and height information, etc., and effectively excludes invalid data by accurately calculating the maximum and minimum values of each dimension of the point cloud data and applying the straight-through filtering technology, thereby ensuring the accuracy of the recognition result;
[0082] The method of the present invention is not limited by lighting conditions and background environment. The depth camera directly obtains the three-dimensional information of the object, avoiding the recognition error caused by the change of lighting in the traditional image recognition technology. At the same time, by performing pre-processing steps such as smoothing and edge detection on the depth image, the anti-interference ability of the system is further enhanced, so that the recognition performance can be kept stable even in a complex and changeable catering environment.
[0083] Combining Fourier transform and deep learning technology, the present invention can automatically extract the features of optical disks and leftovers and effectively distinguish them. Fourier transform can capture the unique features of round dinner plates and leftovers in the frequency domain, while the deep learning model learns to recognize complex patterns and textures by training a large amount of sample data. The combination of the two greatly improves the intelligent level of recognition.
[0084] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic features of the present invention. Therefore, no matter from which point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention.
[0085] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.
Claims
1. An artificial intelligence detection method for distinguishing between clean plates and leftovers in a meal tray, characterized in that: The method comprises the following steps: Step 1: Data collection: Use a point cloud collection device to collect point cloud data. The device includes a high frame rate depth camera module and a data processing module. Each frame of point cloud data is immediately transferred to the processing module for pre-processing after collection to ensure real-time performance. Step 2: Point cloud preprocessing Noise filtering: Calculate the maximum and minimum values of each dimension in each frame of point cloud data, filter out the noise point cloud that exceeds the range, and for each point (x, y, z), determine whether its coordinate value is within the set maximum and minimum range, if not, filter it out; Through filtering: limit the point cloud range, retain only valid points, set the range of X, Y, and Z dimensions, and only retain the point cloud data within the range; Projection generates depth image: Get the maximum value of the valid point cloud in the Z direction, project the point cloud to the XY plane to form a depth image; Step 3: Image segmentation and detection Path 1: Traditional method processing Mean filtering: Perform mean filtering on the depth image to smooth out noise and fill in abnormal points; Edge detection: Use the Canny edge detection algorithm to extract the edge contour of the disc; Hough circle transform: The center area of the circle is detected by Hough circle transform, and all the plate data on the plate are detected; Characteristic differentiation: Distinguish between CDs and leftovers based on their different characteristics; Path 2: Deep Learning Semantic Segmentation Detection Data annotation: perform CD annotation on the depth image to facilitate subsequent depth model training; Semantic segmentation: Input the depth image to the DeeplabV3+ semantic segmentation model, and the model returns the candidate detection results of the CD area; Result correction: Use the detection results of the deep learning model to correct the detection results of the traditional method; Step 4: Multi-path result fusion Calculate IOU: Calculate the intersection over union (IOU) of the disc area detected by the traditional method and the deep learning method; Result fusion: Use deep learning methods to correct the non-disc areas detected by traditional methods; Step 5: Result processing and output Prevent duplicate detection: limit the acquisition rate of the depth camera to prevent repeated acquisition of the same data; Feature analysis and identification: After extracting the features of CDs and leftovers according to the above method, analyze the feature data, set a reasonable threshold to distinguish them, and realize CD identification; Result saving and report generation: Save the recognition results in the form of image annotations and generate statistical analysis reports to facilitate subsequent processing and use.
2. The artificial intelligence detection method for distinguishing between clean plates and leftovers in a meal tray according to claim 1, characterized in that: The Hough circle transform is specifically: Design multiple different thresholds for the circular area detected by Hough transform, screen out multiple circular areas again, use Fourier transform to extract features of the circular area, take weighted average of the features corresponding to each circle, and obtain the features of the final circular area.
3. The artificial intelligence detection method for distinguishing between clean plates and leftovers in a meal tray according to claim 1, characterized in that: The specific results of step 4 are as follows: Determine the size of the IOU of the two regions and set a threshold σ; If the IOU is greater than the threshold σ, it is determined to be an overlapping area, and the semantic segmentation result is preferentially used as the CD detection result.
4. The artificial intelligence detection method for distinguishing between clean plates and leftovers in a meal tray according to claim 1, characterized in that: The specific steps to prevent duplicate detection in step 5 are: The radius, position and other statistics of the circular area detected by multiple adjacent data are performed, and a reasonable threshold is designed to determine whether the data on the adjacent frames is duplicate data. If it is duplicate data, the detection is skipped.
5. The artificial intelligence detection method for distinguishing between clean plates and leftovers in a meal tray according to claim 1, characterized in that: During data collection, the depth camera collects point cloud data and RGB image data sequences of the optical disk and its surrounding environment at a rate of 30 frames per second, and the data format is PLY.