Food waste detection method and system based on image processing

By using high-resolution multispectral imaging and deep learning technology, the problem of fine segmentation and multi-category identification of leftover food in kitchen waste has been solved, achieving efficient and accurate volume and weight estimation and improving the automation level of kitchen waste management.

CN120953607APending Publication Date: 2025-11-14INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS +1

Patent Information

Application Number
CN202511056412.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies for detecting leftover food in kitchen waste suffer from problems such as insufficient segmentation accuracy, inaccurate category identification, and large volume estimation errors. In particular, they are difficult to achieve efficient and accurate automated identification and measurement in complex scenarios.

Method used

By employing high-resolution multispectral imaging combined with multi-feature fusion deep learning, and through multi-angle image acquisition, an improved Retinex algorithm and superpixel segmentation, sparse representation fusion network and stereo vision reconstruction, we can achieve fine segmentation, multi-class recognition and 3D volume estimation of leftover food, and dynamically correct errors caused by occlusion and stacking.

Benefits of technology

It improves the accuracy of segmentation and identification of leftover food in kitchen waste, reduces the error in volume and weight estimation, and enhances the accuracy and efficiency of automated detection, providing reliable data support for the reduction and resource utilization of kitchen waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953607A_ABST
    Figure CN120953607A_ABST
Patent Text Reader

Abstract

The invention discloses a food waste detection method and system based on image processing, belongs to the field of image recognition, and aims to realize efficient and automatic recognition and quantification of kitchen waste. According to the method, residual food images are collected at multiple periods and multiple angles in a kitchen garbage can or a dinner plate recovery area through high-resolution and multi-spectral imaging equipment, preprocessing is carried out in combination with an improved Retinex algorithm and a space self-adaptive denoising technology, and the image quality is improved. Afterwards, fine segmentation of a food area is achieved through a multi-scale super-pixel segmentation and graph segmentation algorithm, and multi-category intelligent recognition is conducted on remaining food through a recognition network fused with multi-modal features. The system further combines stereoscopic vision and Monte Carlo sampling to dynamically and accurately count the volume or weight of various residual foods. The method has the advantages of high adaptability and accurate statistical result, and can provide data support for catering management, resource recovery, nutrition evaluation and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing, and more specifically relates to a method and system for detecting food waste based on image processing. The "food waste detection" mentioned in the title refers to the detection of leftover food in kitchen waste. Background Technology

[0002] With economic development and improved living standards, the catering industry and household food consumption have seen continuous growth, resulting in a large amount of food waste and leftovers. Food waste typically contains a mixture of staple foods, meats, vegetables, and other food categories, and the residue is varied in form and type, with fragments of different sizes, often resulting in mixing, stacking, and obscuring. In the fields of food safety, resource recycling, nutritional analysis, and waste management, achieving automated and precise identification and statistics of leftover food is crucial for promoting the goals of "reduction, resource recovery, and harmless disposal." However, existing methods for detecting and counting leftover food mainly rely on manual experience, which is inefficient, subjective, and lacks effective means of identifying and measuring fragmented, mixed, and heavily obscured food waste.

[0003] In recent years, food recognition technologies based on machine vision and deep learning have been gradually applied to areas such as restaurant inventory management, smart ordering, and dietary management. However, these technologies primarily focus on whole foods or standard plates, and are less adaptable to complex food waste. On the one hand, traditional segmentation methods struggle to handle mixed, stacked, poorly defined, and varied food waste, leading to inaccurate segmentation. On the other hand, single appearance or color features are easily affected by cooking methods and lighting changes, making it difficult to achieve accurate multi-dimensional identification of different types. Furthermore, the estimation of the volume and weight of food waste commonly relies on projected area inference or simple volume models, failing to fully utilize three-dimensional information and easily resulting in significant deviations in real-world scenarios.

[0004] The existing technology, titled "Food Waste Detection Method, Device, and System," application number CN202310146854.X, proposes a method for automatically estimating the quality of food waste by acquiring images of the food to be detected at a food recycling station using an image acquisition device and based on the edible area in the image and a preset correspondence between food category, volume, and mass. This solution enables intelligent and automated identification of food waste behavior, saving labor costs and improving detection efficiency.

[0005] However, analysis of existing technologies reveals the following shortcomings and limitations:

[0006] 1. Limited accuracy in edible area extraction: Existing methods mainly rely on traditional image segmentation or simple region discrimination, lacking the ability to finely segment food mixtures, stacking, occlusion and fragmented residues in complex scenes, making it difficult to accurately identify the specific boundaries and areas of multiple food categories.

[0007] 2. Low accuracy in food category and volume identification: This method mainly relies on the preset static correspondence between category, volume and mass, and does not make full use of multispectral imaging and depth information. As a result, the estimation error is large for foods of different shapes, stacks and deformations in real-world scenarios, and it is difficult to adapt to diverse and complex residual forms.

[0008] 3. Lack of dynamic correction and three-dimensional quantitative capabilities: Existing solutions usually rely only on the area and static relationship of two-dimensional images to infer volume, and do not make sufficient use of three-dimensional spatial information such as height, depth and viewpoint changes, making it difficult to make dynamic corrections and accurate volume / mass estimations for stacked or occluded food.

[0009] 4. Limited feature representation and intelligent recognition capabilities: Traditional methods often rely on manually defined features or shallow classifiers, which have limited recognition accuracy, generalization ability and robustness when faced with food residues of multiple categories and different forms.

[0010] To address the aforementioned shortcomings, this invention proposes a food waste detection method based on high-resolution multispectral imaging and multi-feature fusion deep learning. Through systematic innovations in image acquisition, preprocessing, region segmentation, intelligent recognition, and 3D volume estimation, it overcomes the technical bottlenecks of existing methods in terms of segmentation accuracy, 3D quantification, adaptability to complex scenes, and automated recognition. This method achieves efficient and accurate automatic detection of mixed stacks and multiple types of leftover food, providing strong technical support for applications such as intelligent waste sorting, resource recycling, and nutrition monitoring. Summary of the Invention

[0011] This invention aims to solve the problems of multi-category mixing, fragment stacking, blurred boundaries, and overlapping occlusion that make accurate segmentation and identification difficult in existing methods of identifying and counting leftover food, as well as measurement errors caused by insufficient utilization of three-dimensional information and occlusion in estimating residual volume or weight. It proposes an intelligent method that can automatically and accurately segment and identify multiple categories of leftover food, and combines depth vision and dynamic error correction to achieve high-precision volume and quantity statistics. To achieve the above objectives, this invention employs the following technical solution: The method includes:

[0012] Images of leftover food are collected using high-resolution, multispectral imaging equipment placed in food waste bins or dish recycling areas. Through timed shooting and sensor triggering mechanisms, images of leftover food are automatically collected in multiple time periods, from multiple angles, and in multiple scenarios. The collected images contain timestamps and spatial metadata information.

[0013] The collected images of remaining food are preprocessed by using a fusion of the improved Retinex algorithm and a spatial adaptive denoising method to correct the illumination and suppress noise, thereby enhancing the food texture and edge features in the images.

[0014] The preprocessed image is segmented into regions. Multi-scale boundary-aware superpixel segmentation and fully variational regularized graph cut algorithm are used to achieve fine segmentation of the remaining food region and background. For food stacking or occlusion, shape prior is used to further optimize the segmentation results.

[0015] Food type identification is performed on the segmented food regions. A sparse representation fusion network based on the fusion of texture, color, shape and multispectral features is used, and a transfer learning capsule network is combined to perform multi-category intelligent identification of the remaining food.

[0016] Based on the identified food regions, the quantity, volume, or weight of leftover food in each category is statistically analyzed. Using the area information of the segmented regions and the depth information obtained from stereo vision, the food regions are mapped to three-dimensional space. The volume is inferred through Monte Carlo sampling, and the volume estimation results are optimized by combining stacking and occlusion correction factors. Finally, the weight of leftover food is obtained by combining the known density of each type of food, thus achieving comprehensive detection and quantification of food waste.

[0017] In one embodiment, the high-resolution, multispectral imaging device is equipped with automatic exposure, HDR, and extreme low-light illumination supplementation functions. It can automatically adjust shooting parameters under different lighting conditions to ensure clear images of leftover food in strong backlight, low light, or partially obscured environments. At the same time, it can record the same leftover food from multiple perspectives through multiple fixed or adjustable shooting angles, thereby improving the diversity and information content of image acquisition.

[0018] In one approach, the preprocessing further includes multi-scale feature enhancement of the image and adaptive adjustment of preprocessing parameters according to changes in ambient lighting to further enhance food edges and texture information, and retain the distinguishing features between the remaining food and the background to the maximum extent, thereby providing a more stable and high-quality image input for subsequent segmentation and recognition steps.

[0019] In one approach, a boundary-aware factor is introduced for superpixel segmentation, ensuring that the superpixel units closely follow the actual outline of the food. By fusing multi-scale superpixel segmentation results, the diversity of remaining food size, shape, and fragments can be effectively addressed. Furthermore, by using shape prior constraints based on graph cuts, the stacked and mixed regions are segmented in a refined manner, improving segmentation accuracy and robustness.

[0020] In one approach, a sparse representation fusion network is used to fuse texture, color, shape, and multispectral features. After feature dimension compression and redundancy removal, the feature is further input into a transfer learning capsule network to achieve hierarchical relationship modeling for multiple food categories and high-precision category recognition in mixed and overlapping cases, effectively improving the automatic classification of food residues.

[0021] In one approach, the two-dimensional area of ​​the segmented food region is calculated by combining pixel count with image resolution, and the corresponding three-dimensional point cloud is reconstructed based on the depth information obtained from stereo vision. The region is mapped to three-dimensional space. By randomly sampling points within the bounding box, the number of points actually falling in the food region is counted to estimate the volume. In the volume estimation process, depth distribution histograms and local fitting corrections are used to address occlusion and stacking issues, thereby achieving dynamic and accurate statistics of the remaining food volume and weight.

[0022] In one scheme, the Monte Carlo sampling method for inferring volume includes: inferring the area, volume, or weight of the remaining food by pixel count and image depth information, and reconstructing the three-dimensional contour using depth data obtained from stereo vision;

[0023] The volume is inferred by combining Monte Carlo sampling, and the volume estimation results are optimized by stacking and occlusion correction factors;

[0024] By combining known density models of various foods, the weight of leftover food can be obtained, enabling comprehensive quantitative statistics on food waste.

[0025] Furthermore, an image processing-based food waste detection system, the system being applicable to the method described above, includes:

[0026] A high-resolution, multispectral imaging device is installed in a food waste bin or dish recycling area to automatically acquire images of leftover food at multiple times, angles, and scenes through timed shooting and sensor triggering mechanisms. The imaging device has automatic exposure, HDR, and fill light functions, and can add timestamps and spatial metadata information to the acquired images.

[0027] The image preprocessing module is configured to apply a fusion of the improved Retinex algorithm and a spatial adaptive denoising method to the acquired image of the remaining food, performing illumination correction, noise suppression, and edge texture enhancement, and outputting a preprocessed image with stable quality.

[0028] The image segmentation module is configured to use multi-scale boundary-aware superpixel segmentation and fully variational regularized graph cut algorithm on the preprocessed image to achieve fine segmentation of the remaining food region and background, and to optimize the segmentation of food regions in the case of stacking or occlusion by combining shape prior.

[0029] The food recognition module is configured to automatically identify multiple categories of food regions based on a sparse representation fusion network that integrates texture, color, shape, and multispectral features, and in conjunction with a transfer learning capsule network.

[0030] The statistical analysis module is configured to infer the area, volume, or weight of the remaining food based on the segmentation and recognition results of the food region, through pixel count and image depth information, to reconstruct the three-dimensional contour using depth data obtained from stereo vision, and then infer the volume by combining Monte Carlo sampling. The volume estimation results are optimized by using stacking and occlusion correction factors, and finally the weight of the remaining food is obtained by combining the known density models of various foods, so as to achieve a comprehensive quantitative statistics of food waste.

[0031] It also includes a data management and interface module, configured to uniformly associate and manage timestamps and spatial metadata of all collected and analyzed data in the detection process, and can interface with food waste management systems or nutrition assessment systems to provide data support for subsequent management, recycling, assessment and decision-making.

[0032] Beneficial effects of this invention:

[0033] This invention overcomes the problems of inaccurate segmentation, low recognition accuracy, and large volume estimation errors in existing technologies under complex lighting, mixed stacking, and occlusion conditions by introducing high-resolution multispectral imaging, depth vision technology, and a multimodal fusion recognition network. By employing an improved area-volume inference algorithm combined with a density model, it achieves high-precision segmentation, classification, and weight estimation of remaining food areas, improving the automation and accuracy of identification and measurement. This method reduces manual intervention, improves statistical efficiency, and provides strong data support for applications such as food waste reduction, resource utilization, intelligent management, and nutritional assessment, demonstrating promising prospects for widespread adoption and practical application. Attached Figure Description

[0034] Figure 1 This is a flowchart of the method of the present invention;

[0035] Figure 2 This is a system block diagram of the present invention.

[0036] In the diagram, 1-High-resolution, multispectral imaging equipment, 2-Image preprocessing module, 3-Image segmentation module, 4-Food recognition module, 5-Statistical analysis module, and 6-Data management and interface module. Detailed Implementation

[0037] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Typical embodiments of the invention are shown in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.

[0038] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. To facilitate understanding, the invention will now be described more fully with reference to the accompanying drawings. Typical embodiments of the invention are shown in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to make the disclosure of the invention more thorough and complete.

[0039] The relevant technologies involved in this patent belong to the following categories:

[0040]

[0041] like Figure 1 As shown, the implementation steps of the image processing-based food waste detection method are as follows:

[0042] Step 1: Acquire images of leftover food. Using high-resolution, multispectral imaging equipment, images of waste in food waste bins or dish recycling areas are acquired at different times and angles, ensuring coverage of various lighting and occlusion scenarios.

[0043] In the first step of implementing food waste detection—collecting images of leftover food—high-resolution, multispectral imaging equipment is deployed at key locations in food waste bins or dish recycling areas. This imaging equipment must not only be capable of visible light imaging but also equipped with near-infrared or other wavelength sensing modules to capture details of the different textures and physical properties of the food. Regarding image acquisition strategy, the system develops a timed shooting schedule, combined with sensor triggering mechanisms (such as automatic shooting when infrared movement is detected) to ensure coverage of all times when waste is disposed of.

[0044] To enhance image diversity and anti-interference capabilities, the equipment employs multiple shooting angles, such as through a robotic arm or multi-camera collaboration, allowing the same leftover food to be recorded from multiple perspectives, including front, side, and oblique views. Furthermore, addressing practical challenges such as fluctuating ambient light and obstruction from debris in the recycling area, the imaging equipment integrates automatic exposure, HDR (High Dynamic Range), and extreme low-light compensation functions, ensuring clear and information-rich images of leftover food even under conditions of strong backlight, low light, or partial obstruction. For ease of subsequent processing, each image is timestamped and imbued with spatial metadata, enabling precise data matching and traceability. These measures allow for the efficient and comprehensive acquisition of diverse and highly usable images of leftover food, laying a solid data foundation for subsequent analysis.

[0045] Step 2: Preprocess the acquired images. The improved Retinex algorithm and spatial adaptive denoising method are combined to perform illumination correction and noise suppression on the images, automatically enhancing food texture and edge features.

[0046] When preprocessing the acquired images of remaining food, an improved Retinex algorithm is first used to correct for uneven lighting and complex environments. Specifically, the system is based on a multi-scale Retinex model (MSR), treating the input image I(x,y) as the product of the true reflectance component R(x,y) and the illumination component L(x,y), i.e., I(x,y) = R(x,y)·L(x,y). To enhance the effect, spatially adaptive weights are introduced. By extending the traditional MSR, the Retinex operation is combined with a weighting factor guided by local variance, and normalization is performed using the following formula:

[0047]

[0048] Among them, F n (x,y) is a Gaussian filter kernel with scale n, w n (x, y) represents the weights adaptively adjusted based on local structural changes, effectively suppressing artifacts caused by abrupt changes in illumination. For noise that may arise during the acquisition process, a spatially adaptive nonlocal means (NLM) denoising method is employed, dynamically adjusting the denoising intensity by analyzing the similarity between pixel blocks. Its mathematical expression is:

[0049]

[0050] Among them, P x,y and P i,jLet h be the pixel block features centered at (x,y) and (i,j) respectively, and C(x,y) be the normalization factor to ensure that details are preserved in texture-rich areas and noise is effectively suppressed in smooth areas. Finally, combined with gradient-based edge enhancement operations, the Laplacian operator is introduced to further highlight the food texture structure and boundaries.

[0051]

[0052] Where λ is the adjustment coefficient. After the above fusion processing, the resulting image not only significantly improves the problem of uneven illumination, but also has low noise and high texture details, providing a more robust and high-quality input for subsequent food region segmentation and category recognition.

[0053] Step 3: Use a variant of the multi-scale superpixel segmentation algorithm (an improved version of the SLIC algorithm with boundary awareness) to over-segment the image region, generating continuous and structurally consistent food fragments; and use an improved graph cut algorithm to accurately separate the background from the food, improving the ability to distinguish mixed and overlapping remaining food.

[0054] After image preprocessing, the key step in oversegmenting the image region is to employ a variant of the multi-scale superpixel segmentation algorithm. This process is based on the SLIC (Simple Linear Iterative Clustering) superpixel algorithm, and improved by introducing a boundary-aware mechanism to address the complex boundaries of remaining food. Specifically, the input image is first mapped to the CIE-Lab color space, and the distance between pixels and cluster centers is calculated using the following metrics:

[0055]

[0056] in, It is the Euclidean distance in the Lab color space. Here, m is the distance in coordinate space, S is the compactness factor, and m is the desired superpixel size. To improve the accuracy of edge segmentation, the algorithm introduces a weight adjustment mechanism based on image gradients into the clustering distance. If a pixel is located in a boundary region, the distance metric is multiplied by a higher weight, pushing the superpixel division to strictly fit the true contour of the food fragment. Multi-scale implementation achieves a balance between large particles and small fragments by gradually decreasing the S value and fusing segmentation results from different scales, ensuring that the generated superpixels are both delicate and continuous.

[0057] After obtaining the structured superpixel fragments, to further separate the background and food regions while addressing the mutual occlusion and blending of remaining food items, an improved graph cut algorithm based on Total Variation regularization is employed. The entire image is defined as an undirected graph G = (V, E), where V is the set of superpixel nodes and E is the edge between adjacent superpixels. Each edge is assigned a weight.

[0058]

[0059] Where f i and f j The color and texture features of adjacent superpixels, Let λ be the boundary gradient and λ be the boundary sensitivity factor. The maximum flow-minimum cut algorithm is used to find the minimum cost splitting path, minimizing the energy function.

[0060]

[0061] Where L i For node labels (food / background), D i (·) represents the discrimination cost of the topic model. To address the mixing and overlap issues, the algorithm introduces prior shape constraints based on the fusion of global spectral information, assigning priority segmentation weights to typical food shapes, and using multi-resolution adaptive optimization iterations to further improve the separation effect of heterogeneous fragments. The final segmentation results can not only accurately distinguish the remaining food from the background, but also meticulously analyze different types of food fragments mixed in the same area, laying a solid foundation for subsequent identification and statistics.

[0062] Step 4: Employ a feature fusion network based on sparse representation to integrate texture, color, shape, and spectral features for food type identification. Utilize transfer learning to model the hierarchical relationship of food structures using a Capsule Network structure, effectively distinguishing between multiple categories of leftover food such as staple foods, vegetables, and meats.

[0063] After achieving high-quality food region segmentation, the next step is to perform fine-grained category identification on the segmented food fragments. First, for each superpixel fragment, multi-dimensional features are extracted, including fine-grained texture features described by Local Binary Pattern (LBP) and Gray-Level Co-occurrence Matrix (GLCM), color features composed of CIE-Lab and multispectral channel statistics, shape features such as Hu invariant moments and contour Fourier descriptors, and spectral features modeled from the multispectral response. To fully integrate these high-dimensional, heterogeneous feature information, a Sparse Representation Feature Fusion Network (SRFFN) is employed. Specifically, it is assumed that each feature mode corresponds to a data vector x.(k) Construct a joint feature representation X = [x (1) ,x (2) ,…,x (K) The sparse representation layer achieves correlation filtering and information redundancy removal between feature dimensions by solving the following optimization problem:

[0064]

[0065] Where D is the overcomplete dictionary of feature atoms, α is the sparse coefficient, and λ is the regularization parameter. The final fused feature z strengthens the significantly discriminative patterns by weighted combination of the sparse coefficient and the basic feature vector, making the downstream classification process more robust.

[0066] In terms of network architecture, the SRFFN fusion layer is followed by a Capsule Network structure based on the transfer learning paradigm to fully explore the hierarchical relationships of food fragments. The Capsule Network models the spatial information and compositional relationships of food instances in vector form, with its basic unit, the "capsule," outputting an activity vector v. j , representing the probability of existence and attributes of category j. Through a dynamic routing algorithm, capsule networks can achieve hierarchical modeling from fragments to the overall category. The dynamic routing mechanism predicts the vector from each lower-level capsule i to the higher-level capsule j. Calculate using the following formula:

[0067]

[0068] Where u i For input capsule output, W ij This is a learnable weight matrix. The final output is the activity vector v. j Based on the weighted sum of all inputs and after normalization:

[0069]

[0070] Where c ij For routing coefficients. Through end-to-end training, the network automatically learns how to hierarchically organize and categorize food fragments into food categories such as staple foods, vegetables, and meats. It can also effectively model the mixed and overlapping structures of food, demonstrating superior recognition capabilities in complex scenarios that traditional convolutional neural networks struggle to distinguish. Finally, through the linkage of fused features and capsule networks, the system can accurately output the category label of each segmented food fragment, achieving efficient and intelligent recognition of multiple categories of leftover food.

[0071] Step 5: Statistically analyze the types and quantities of remaining food. Based on the segmented food regions, apply an improved area-volume inference algorithm (combining Monte Carlo sampling and stereo vision depth estimation) to calculate the volume or weight of each type of food residue, and dynamically correct errors caused by occlusion and stacking.

[0072] After identifying and classifying the food areas, the next step is to accurately count the quantity and volume (or weight) of the remaining food in each category. First, based on the aforementioned precise region segmentation, the two-dimensional projected area of ​​the remaining food in each category is extracted; this area A... i The depth is obtained by combining pixel counting with image resolution conversion. To overcome the potential errors in estimating volume solely based on area, this scheme integrates depth information acquired from stereo vision, specifically using binocular or multi-view cameras to obtain the depth d(x,y) of each pixel through disparity matching. Each identified food region is mapped to a 3D point cloud space to obtain its spatial contour S. i The preliminary volume estimation uses the Monte Carlo sampling method: within the 3D bounding box of each food region, N points are randomly generated, and it is checked whether each point falls within the point cloud envelope formed by the food region. The number of hit points n is counted, and the volume is estimated accordingly.

[0073]

[0074] Where V box Let represent the bounding box volume. To further correct for biases introduced by stacking and occlusion between different remaining food items, a combination of depth distribution histogram and local fitting of occlusion regions is used. This is achieved by interpolating and filling depth abrupt changes and hole regions using the mean depth of adjacent pixels, while also introducing a statistical correction factor γ. i Dynamically adjust the estimated volume:

[0075] V i′ =V i ·γ i

[0076] Where γ i The weight is obtained through adaptive optimization based on typical stacking characteristics of each category, detected occlusion ratios, and historical sample statistical models. If further weight conversion is needed, it is combined with known food densities ρ for each category. i Obtain the residual weight

[0077] M i =V i′ ·ρ i

[0078] Throughout the process, the algorithm continuously optimizes parameters based on actual distribution and dynamic sampling errors, ensuring that the statistical results have high accuracy and robustness in the context of multi-category, multi-form, and complexly stacked food waste, achieving comprehensive quantification of the types and quantities of leftover food, and providing a reliable data foundation for subsequent resource recycling and nutritional assessment.

[0079] The food waste detection method based on image processing proposed in this invention demonstrates many significant technical achievements, specifically in the following aspects:

[0080] First, by employing multispectral imaging and multi-angle data acquisition methods, the problems of blurred images and loss of detail in complex lighting and occlusion scenarios, as encountered with traditional single-camera methods, are effectively overcome. Images acquired using this invention exhibit richer food textures and edge details, effectively improving the quality of the foundational data for subsequent recognition and segmentation.

[0081] Secondly, this invention comprehensively applies an improved Retinex algorithm and spatial adaptive denoising technology in the image preprocessing stage, significantly enhancing the image's illumination uniformity and signal-to-noise ratio. Experimental results show that this method can significantly improve the discriminability of food regions and reduce the negative impact of noise on segmentation accuracy under low illumination and high dynamic environments.

[0082] In the image segmentation and recognition stage, a multi-scale superpixel fusion and boundary-aware segmentation algorithm is employed to effectively solve the challenges of segmenting stacked, overlapping, and fragmented food residues. Compared with existing single-scale segmentation or traditional thresholding methods, the method of this invention can significantly improve the segmentation accuracy of food regions, especially in complex backgrounds and multi-category mixed scenes, where segmentation edges are smoother and the error rate is significantly reduced.

[0083] Furthermore, in estimating the volume and weight of food residue, this invention introduces for the first time a joint estimation method combining stereo vision depth reconstruction and Monte Carlo sampling, along with an adaptive correction model for occlusion areas, effectively overcoming the problem of large errors in estimating volume from the area of ​​a single image. Experiments show that, using this method, the relative error in estimating the volume and weight of remaining food can be reduced to less than 10%, far superior to the accuracy of traditional empirical methods and single-view methods.

[0084] By combining the above technical measures, this invention not only significantly improves the automation level and accuracy of food waste detection, but also provides a reliable data foundation for subsequent resource recovery and nutritional assessment of kitchen waste. Compared with existing technologies, this solution demonstrates stronger adaptability and robustness to complex scenarios involving food waste, possessing significant practical application value and innovative advantages, and contributing to the promotion of intelligent management and sustainable development in the catering industry.

[0085] like Figure 2 As shown, the image processing-based food waste detection system includes:

[0086] A high-resolution, multispectral imaging device 1 is installed in a food waste bin or dish recycling area to automatically acquire images of leftover food at multiple time periods, angles, and scenes through timed shooting and sensor triggering mechanisms. The imaging device has automatic exposure, HDR, and fill light functions, and can add timestamps and spatial metadata information to the acquired images.

[0087] Image preprocessing module 2 is configured to perform illumination correction, noise suppression, and edge texture enhancement on the acquired image of remaining food by fusing the improved Retinex algorithm with a spatial adaptive denoising method, and output a preprocessed image with stable quality.

[0088] Image segmentation module 3 is configured to use multi-scale boundary-aware superpixel segmentation and fully variational regularized graph cut algorithm on the preprocessed image to achieve fine segmentation of the remaining food region and background, and to optimize the segmentation of food regions in the case of stacking or occlusion by combining shape prior.

[0089] The food recognition module 4 is configured to use a sparse representation fusion network that integrates texture, color, shape and multispectral features, and is combined with a transfer learning capsule network to automatically identify multiple categories of food regions obtained from segmentation.

[0090] The statistical analysis module 5 is configured to infer the area, volume or weight of the remaining food based on the segmentation and recognition results of the food region, through pixel count and image depth information, to reconstruct the three-dimensional contour using the depth data obtained by stereo vision, and then to infer the volume by combining Monte Carlo sampling. The volume estimation results are optimized by using stacking and occlusion correction factors, and finally the weight of the remaining food is obtained by combining the known density models of various foods, so as to achieve a comprehensive quantitative statistics of food waste.

[0091] Data Management and Interface Module 6 is configured to centrally manage image data, processing results, and statistical information collected during the detection process, and to uniformly associate timestamps with spatial metadata; it also provides external interfaces to enable data interaction and integration with food waste management platforms or nutrition assessment systems.

[0092] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0093] It should be understood that the above detailed description of the technical solutions of the present invention with reference to preferred embodiments is illustrative and not restrictive. Those skilled in the art can modify the technical solutions described in the embodiments or make equivalent substitutions for some of the technical features based on reading this specification; however, these modifications or substitutions do not cause the essence of the corresponding technical solutions to depart from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A food waste detection method based on image processing, characterized in that: The method includes: Images of leftover food are collected using high-resolution, multispectral imaging equipment placed in food waste bins or dish recycling areas. Through timed shooting and sensor triggering mechanisms, images of leftover food are automatically collected in multiple time periods, from multiple angles, and in multiple scenarios. The collected images contain timestamps and spatial metadata information. The collected images of remaining food were preprocessed by using a fusion of the improved Retinex algorithm and a spatial adaptive denoising method to correct the illumination and suppress noise in the images. The preprocessed image is segmented into regions. Multi-scale boundary-aware superpixel segmentation and fully variational regularized graph cut algorithm are used to achieve fine segmentation of the remaining food region and background. For food stacking or occlusion, shape prior is used to further optimize the segmentation results. Food type identification is performed on the segmented food regions. A sparse representation fusion network based on the fusion of texture, color, shape and multispectral features is used, and a transfer learning capsule network is combined to perform multi-category intelligent identification of the remaining food. Based on the identified food regions, the quantity, volume, or weight of the remaining food in each category is statistically analyzed. Using the area information of the segmented regions and the depth information obtained from binocular or multi-view stereo vision, the food regions are mapped to three-dimensional space. The volume is inferred using the Monte Carlo sampling method, and the volume estimation results are optimized by combining stacking and occlusion correction factors.

2. The food waste detection method based on image processing according to claim 1, characterized in that: The high-resolution, multispectral imaging device is equipped with automatic exposure, HDR, and extreme low-light illumination supplementation functions. It can automatically adjust shooting parameters under different lighting conditions to adapt to the image acquisition needs of strong backlight, low light, or partially obscured environments. At the same time, it can acquire multi-angle images through multiple fixed or adjustable shooting angles to enhance the information coverage.

3. The food waste detection method based on image processing according to claim 1, characterized in that: The preprocessing also includes multi-scale feature enhancement of the image and adaptive adjustment of preprocessing parameters according to changes in ambient lighting to further enhance food edge and texture information and retain the distinguishing features between the remaining food and the background to the maximum extent, thereby providing a more stable and high-quality image input for subsequent segmentation and recognition steps.

4. The food waste detection method based on image processing according to claim 1, characterized in that: By introducing a boundary-aware factor for superpixel segmentation, the superpixel units are made to closely follow the actual outline of the food. By fusing multi-scale superpixel segmentation results, the diversity characteristics of the remaining food can be adapted; Furthermore, by incorporating shape prior constraints based on graph cut, the stacked mixed region is finely segmented.

5. The food waste detection method based on image processing according to claim 1, characterized in that: A sparse representation fusion network is used to fuse texture, color, shape and multispectral features. After feature dimension compression and redundancy removal, the feature is further input into a transfer learning capsule network to achieve hierarchical relationship modeling of multiple food categories and category recognition in mixed and overlapping cases.

6. The food waste detection method based on image processing according to claim 1, characterized in that: The two-dimensional area of ​​the segmented food region is calculated by combining pixel count with image resolution, and the three-dimensional point cloud is reconstructed by combining depth information obtained from stereo vision. Monte Carlo sampling is used to estimate the volume of the remaining food, and the weight is calculated based on the food density model.

7. The method according to claim 6, characterized in that: In the Monte Carlo sampling volume estimation process, depth distribution histograms and local fitting methods are used to correct errors caused by occlusion or stacking, and the volume estimation accuracy is optimized by random sampling within the bounding box and hit statistics of the point cloud space.

8. A food waste detection system based on image processing, characterized in that, The system includes: High-resolution multispectral imaging equipment is used to acquire images of leftover food from multiple time periods, angles, and scenes through timed shooting and sensor triggering mechanisms. The image preprocessing module is used to perform illumination correction, noise suppression, and edge enhancement on the acquired images; The image segmentation module is used to extract the remaining food region based on multi-scale boundary-aware superpixel segmentation and graph cut algorithms; The food recognition module is used to automatically identify multiple categories of food regions obtained from segmentation. The statistical analysis module is used to estimate the volume of remaining food by combining image depth information and to convert the weight based on the food density model. The data management and interface module is used to uniformly manage the test data and interface with the food waste management system or nutrition assessment system.

Citation Information

Patent Citations

  • Food waste detection method, device and system

    CN118521950A

Cited By

  • Method for monitoring residual feed in cattle and sheep feed trough based on vision

    CN121747037A

  • Machine vision-based food waste analysis method and system

    CN122365415A