Community organic garbage volume and fresh weight estimation method
By combining Mask R-CNN and binocular vision depth sensing technology, the problem of accuracy in estimating the volume and fresh weight of organic waste in the community was solved, and efficient organic waste identification and data estimation were achieved.
Patent Information
- Application Number
- CN202410565397.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-09
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies struggle to accurately identify and segment the volume and fresh weight of organic waste from community households, impacting the efficiency of its harmless and resource-based treatment.
A Mask R-CNN deep neural network was used to train an instance segmentation model for community organic waste images. Combined with binocular vision depth sensing technology, a waste volume prediction algorithm was constructed by combining the instance segmentation model with the binocular vision volume algorithm. The volume was then estimated using a univariate linear regression model of waste fresh weight.
The average relative error in estimating the volume and fresh weight of organic waste in the community was achieved between 5.98% and 18.71%, with most errors around 10%, thus improving the accuracy of organic waste identification and estimation.
Smart Images

Figure CN120931706A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method for estimating the volume and fresh weight of organic waste in a community. Background Technology
[0002] In recent years, the amount of household waste generated has increased rapidly, highlighting the growing importance of waste sorting and its harmless and resource-based treatment. Thanks to the rapid development and widespread application of electronic science and technology and deep learning, deep learning network models can now be used to classify and segment community waste images, achieving intelligent identification, classification, and automatic sorting of waste. As a crucial component of community household waste, organic waste, such as vegetable and fruit scraps and leftover food, is difficult to accurately identify and segment, and methods for estimating its volume and fresh weight are rarely studied. Estimating the volume and fresh weight of community organic waste is a crucial step in achieving its harmless and resource-based treatment, providing technical support for the automated and intelligent processing of waste. Summary of the Invention
[0003] In view of this, the present invention proposes a method for estimating the volume and fresh weight of community organic waste. By using deep learning and binocular vision depth sensing technology, organic waste can be identified and its volume and fresh weight estimated.
[0004] On the one hand, this invention proposes a method for estimating the volume and fresh weight of community organic waste, including the following steps:
[0005] S1: Use Mask R-CNN deep neural network to train community organic waste images to obtain a community organic waste instance segmentation model, which is used for target recognition and target instance segmentation;
[0006] S2: Use a binocular camera to acquire depth images of the community organic waste before and after processing;
[0007] S3: Use an instance segmentation model to segment the depth image of community organic waste and obtain the segmented waste depth data, which comes from a stereo camera;
[0008] S4: Combining instance segmentation model and binocular vision volume algorithm to construct waste volume prediction algorithm, which can obtain waste volume estimate;
[0009] S5: By combining the univariate linear regression model of waste fresh weight, a model of the relationship between waste volume and fresh weight can be obtained, and the fresh weight of waste can be estimated.
[0010] This invention combines instance segmentation model and binocular depth sensing technology to estimate the volume and fresh weight of community organic waste. The average relative error of the estimation results is between 5.98% and 18.71%, with most of them around 10%. It can effectively identify community organic waste and estimate its volume and fresh weight data. Attached Figure Description
[0011] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:
[0012] Figure 1 This is a flowchart for estimating the volumetric fresh weight of community organic waste according to the present invention;
[0013] Figure 2 Line graphs showing the model metrics for different learning rates in this invention;
[0014] Figure 3 Line graphs showing training metrics for different batch sizes in this invention;
[0015] Figure 4 This is a diagram showing the segmentation results of a community organic waste example according to the present invention;
[0016] Figure 5 This is a diagram of the volume calculation model based on depth data of the present invention;
[0017] Figure 6 This is a graph showing the relationship between the calculated values and actual measured values of the projected pixel area algorithm of this invention.
[0018] Figure 7 This is a relative deviation diagram for identifying the fresh weight of several types of community organic waste according to the present invention. Detailed Implementation
[0019] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0020] like Figure 1 As shown, this invention proposes a method for estimating the volume and fresh weight of organic waste in a community, comprising the following steps:
[0021] S1: Train community organic waste images to obtain a community organic waste instance segmentation model based on Mask R-CNN;
[0022] The computing device used was a microcomputer equipped with an Intel Core i9 10920x CPU, 64GB of RAM, an NVIDIA TITAN RTX graphics card, and Keras as the deep learning framework. Keras is an open-source artificial neural network library written in Python.
[0023] Images of community organic waste can be sourced from official sources, or taken with cameras, mobile phones, or downloaded from the internet. The performance evaluation in this invention used 136 images taken with a mobile phone in multiple residential communities, and 264 images downloaded from the internet. These images include eight common types of community organic waste (peanut shells, cabbage leaves, lettuce leaves, orange peels, banana peels, carrot peels, potato peels, and leftover food). All images were cropped to squares and their resolution was uniformly adjusted to 1024*1024, resulting in a dataset of 400 images representing common community organic waste.
[0024] The collected images were labeled using annotation tools, and image augmentation methods were also used to enrich the image sample data. Following the common ratio for small datasets, the training, validation, and test sets were divided into an 8:1:1 ratio. The obtained dataset was fed into a Mask R-CNN network model for training, thus obtaining a community organic waste instance segmentation model. The learning rate and batch size were selected as the hyperparameters for network training in experiments. Univariate experiments were used to find the optimal learning rate and batch size parameters. Other important hyperparameters, such as the network structure (ResNet-101), momentum (0.9), learning rate decay (0.0001), batch size (2), stride settings in FPN (5 options: [4, 8, 16, 32, 64]), anchor box side length settings (5 options: [32, 64, 128, 256, 512]), and aspect ratio settings for generated anchor boxes (3 options: [0.5, 1, 2]). The number of anchor boxes used for training in each image was 256. The weights of the loss functions for each category were all set to 1. The number of epochs was 60, and the number of training steps per epoch was set to 500.
[0025] After determining the optimal learning rate, adjust it to find a suitable batch size. First, set the learning rate to the range lr = [0.0005, 0.001, 0.002, 0.003, 0.004]. The training results are as follows... Figure 2 As shown.
[0026] Depend on Figure 2It can be seen that the larger the learning rate, the steeper the rise of the various evaluation metrics, and the more intense the oscillations in the values. To evaluate the final performance metrics of the model, the mean values of various metrics for the last 5 epochs of the 60-epoch model training are shown in Table 1.
[0027] Table 1. Model performance metrics at different learning rates
[0028]
[0029] In Table 1, AP (Average Precision) is a commonly used evaluation metric for object detection and recognition models. The AP values corresponding to threshold values of 0.5 and 0.75 are denoted as AP50 and AP75, respectively. As shown in Table 1, the learning rate is l... r When the learning rate is 0.001, the AP, AP50, and AP75 metrics are the highest among all models, therefore the optimal learning rate value is determined. r =0.001 is a better learning rate choice.
[0030] Learning rate l r The batch size is set to 0.001, with a range of [1, 2, 4]. The number of training steps per epoch is adjusted to 1000, 500, and 250 respectively based on the batch size of 1, 2, and 4 to ensure that the total amount of data used for training in each epoch remains constant. Other hyperparameters remain unchanged. Experimental metrics are as follows: Figure 3 As shown in Table 2, the average values of the metrics for the last 5 epochs are as follows.
[0031] Table 2 Performance Indicators for Different Batch Sizes
[0032]
[0033] The results show that when the batch size is 2, all indicators reach their peak values, making it the optimal choice.
[0034] In this embodiment, the learning rate of the Mask R-CNN instance segmentation model is set to 0.001 and the batch size is set to 2. The instance segmentation test results are as follows: Figure 4 As shown, the model achieves correct identification and segmentation of most targets.
[0035] S2: Use a binocular camera to acquire depth images of the community organic waste before and after the organic waste needs to be processed.
[0036] In this embodiment, a ZED binocular camera is used to acquire images of the community organic waste to be processed. The image output resolution is set to 1242*2208, the exposure parameter is set to 32 (6.34 milliseconds), and the calibration is completed using the camera's calibration software ZEDCalibration.
[0037] S3: Use an instance segmentation model to segment the depth image of community organic waste and obtain the segmented waste depth data. The waste depth data comes from a stereo camera; the waste depth data is the distance from a point on the object to the stereo camera.
[0038] S4: By combining the instance segmentation model with the binocular vision volume algorithm, a waste volume prediction algorithm can be constructed to obtain the waste volume estimation result;
[0039] like Figure 5 As shown, the garbage volume prediction algorithm adopts the principle of volume algorithm based on depth data. The distance from a point on an object to the camera is also called depth. If the depth of the object corresponding to each pixel in the image is calculated, a depth map can be obtained. The object is regarded as a combination of many tiny cylinders, the number of which is the number of pixels h×w. The base area A of each small cylinder is calculated. u,v With height H u,v The total volume of the measured object can be calculated. The formula for calculating volume V is as follows (1).
[0040]
[0041] In formula (1), A u,v H represents the base area of each small cylinder that makes up the object. u,v This indicates the height value of each column.
[0042] Height H u,v The calculation method is as follows: Depth data Z' is obtained before and after the organic waste is placed in the garbage community. u,v Z u,v The height from the waste placement platform in front of the community's organic waste to the binocular camera is the depth data Z'. u,v The height of the waste from the binocular camera after the organic waste in the garbage community is the depth data Z. u,v Each frame of depth data has a shape of h×w×3, meaning that each pixel stores the x, y, and z coordinates of the object. Taking the z-coordinate component yields a depth map Z' with a shape of h×w×1. u,v and Z u,v Using depth data Z' u,v The reference height G of the waste placement platform was calculated using the average method. u,v Height H u,v The calculation formula (2) is as follows:
[0043]
[0044] In formula (2), u = 0, 1, 2, ..., h-1; v = 0, 1, 2, ..., w-1; σ min and σ maxThese are the set minimum and maximum depth thresholds. Pixels whose depth data exceeds these thresholds are treated as outliers. G u,v This is the reference used in the formula to calculate the height of the object. Two methods are used to calculate this reference, one of which involves using the depth data Z' in front of where the object is placed. u,v The average value is used in the calculation, and the calculation method is as shown in formula (3).
[0045]
[0046] Another calculation method is to first reorganize the depth data, combining its horizontal and vertical pixel coordinates u and v and the depth value z of each pixel into a coordinate point (u,v,z), and then perform least-squares fitting on the data composed of all coordinate points to obtain the general equation of the plane.
[0047] Au + BV + Cz + D = 0 (4)
[0048] In formula (4), A, B, C, and D are the coefficients of the plane equation, which can be obtained through the fitting results. The depth value of each pixel on the fitted plane is further calculated by changing the form of the plane equation as shown in formula (5).
[0049]
[0050] Calculate pixel area A u,v The method is as follows: based on the projection relationship of binocular vision data, such as... Figure 5 As shown, the calculated pixel length and width differ for pixels with different depth values. Therefore, the base area A of the small cylinder needs to be adjusted according to the depth value of each pixel. u,v The pixel located at point (u,v) corresponds to the height m of the small square on the actual object. u,v and width n u,v The projection relationship between depth and projection is expressed as follows:
[0051]
[0052]
[0053] fov in equations (6) and (7) v and fov H These are the vertical and horizontal field of view of the binocular camera, respectively.
[0054] Therefore, the base area A of the small cylinder corresponding to the pixel at point (u,v) can be obtained. u,v As in formula (8):
[0055] A u,v =m u,v n u,v (8)
[0056] To accommodate the instance segmentation model used, depth data was truncated to a central 1024×1024 pixel region of interest for the above calculations.
[0057] Combining instance segmentation with binocular vision volumetric algorithms yields a volume prediction algorithm Vi for various types of waste, as shown in formula (9):
[0058]
[0059] In equation (9), u = 0, 1, 2, ..., h-1; v = 0, 1, 2, ..., w-1; i = 1, 2, ..., k, where k is the number of categories, and M... i It is the mask of the i-th class of samples generated by the Mask R-CNN model. It is a matrix of shape h×w in which each element takes the value [0,1].
[0060] In this example, an evaluation experiment was conducted on the volume prediction algorithm described above. The camera was mounted with its optical axis perpendicular to the plane where the garbage was placed. Depth data of the plane was acquired at different distances, and the estimated pixel area at the center of the image was calculated using formula (8). This was repeated three times, and the average value was taken. A total of seven imaging distances were selected for data acquisition, ranging from 0.4m to 0.7m with a step size of approximately 5cm. At the same time, the actual length and width of the 1024×1024 pixel area on the imaging platform were manually measured with a tape measure, and the average pixel area at that imaging distance was calculated as the measured value. The results are as follows: Figure 6 As shown, there is a high correlation between the pixel area measurement and prediction results, with R² = 0.999 and RMSE = 0.0046. The average relative error is 2.43%, and the maximum error is 3.12%. This indicates that the projection pixel area estimation method proposed in this invention can calculate the pixel area at different heights with relatively high accuracy.
[0061] S5: Combining the univariate linear regression model of waste fresh weight can yield the estimated results of waste fresh weight.
[0062] In this embodiment, in order to estimate the fresh weight of each type of waste, it is necessary to establish a relationship model between the detection volume and the actual fresh weight of each type of waste, as shown in formula (10).
[0063] m i =f i (V i (10)
[0064] In equation (10), f i With m i These represent the volume-fresh weight model and the fresh weight estimation results for the i-th type of waste, respectively.
[0065] A univariate linear regression model is performed on the calculated sample data volume x and the corresponding actual measured fresh weight y. The model expression is as follows, where p1 and p2 are the coefficients of the univariate linear regression model, as shown in formula (11).
[0066] y = p1x + p2 (11)
[0067] In this embodiment, the coefficient of determination (R²) and root mean square error (RMSE) are used to evaluate the regression performance of the model. The coefficient of determination measures the extent to which the variation of the dependent variable can be explained by the independent variables, and can be used to judge the explanatory power of the fitted regression model. The coefficient of determination is calculated as shown in formula (12), where y i This is the actual value. Indicates the predicted value. This represents the average number of observations, with a total sample size of n.
[0068]
[0069] The root mean square error (RMSE) is a commonly used indicator to measure the difference between the predicted values of a model and the actual data. In the fitting results, the lower the value of this indicator, the lower the deviation between the fitted model and the actual data, that is, the better the fitting effect. The calculation method of the root mean square error is as shown in formula (13).
[0070]
[0071] To evaluate the performance of the method for estimating the fresh weight of municipal solid waste, relative error (RE) and relative average error (RAE) were used for assessment. For any sample, the actual value y and its predicted value were compared. The relative error is calculated as shown in formula (14).
[0072]
[0073] The relative average error is calculated using formula (15).
[0074]
[0075] In equations (13), (14), and (15), n is the sample size. y represents the predicted value of a certain sample. i This is the actual value.
[0076] The performance of the proposed method for estimating the fresh weight of community organic waste was tested. Five samples of each type of waste were weighed, resulting in a total of 40 samples. Images and depth data of each sample were acquired under three different placement postures, resulting in a total of 120 sets of data. Each set of data was then tested. Figure 7The figure shows the box plots of relative errors for each category. The freshness weight estimation error mainly appears as negative values, indicating that the estimated value tends to be smaller than the true value. This phenomenon is primarily caused by the slightly lower resolution (28×28) of the predicted mask obtained by the established Mask R-CNN instance segmentation model. This results in the instance segmentation result being smaller relative to the garbage contour, leading to an overall bias in the prediction. This phenomenon is visible... Figure 7 Furthermore, some outliers in the results for carrot peels and banana peels were due to missed detections by the model.
[0077] Table 3 shows the relative average error of the fresh weight estimation results for each category of waste.
[0078]
[0079] A comprehensive evaluation was conducted by combining the relative average error and standard deviation, as shown in Table 3. Peanut shells and leftover food showed relatively better results, with relative average errors less than 10% and small standard deviations. This is because the structural characteristics of these two types of waste do not easily produce stacking gaps that affect the estimation results. The categories with slightly worse prediction results were cabbage leaves, potato peels, and lettuce leaves, with relative average errors less than 12% and small standard deviations. The main error originated from the gaps created by the stacking of waste. Orange peels, carrot peels, and banana peels showed relatively larger errors, with relative average errors ranging from 15% to 19% and large standard deviations. These errors were caused by a combination of stacking gaps and occasional missed detections.
[0080] The above performance evaluation results show that the average relative error of the method for estimating the volume and fresh weight of community organic waste proposed in this invention is as low as 5.98% and as high as 18.71%, with most of them being less than 11%, which proves the feasibility of the proposed method.
[0081] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for estimating the volume and fresh weight of organic waste in a community, characterized in that, Includes the following steps: S1: Use Mask R-CNN deep neural network to train community organic waste images to obtain a community organic waste instance segmentation model, which is used for target recognition and target instance segmentation; S2: Use a binocular camera to acquire depth images of the community organic waste before and after processing; S3: Use an instance segmentation model to segment the depth image of community organic waste and obtain the segmented waste depth data, which comes from a stereo camera; S4: Combining instance segmentation model and binocular vision volume algorithm to construct waste volume prediction algorithm, which can obtain waste volume estimate; S5: By combining the univariate linear regression model of waste fresh weight, a model of the relationship between waste volume and fresh weight can be obtained, and the fresh weight of waste can be estimated.
2. The method for estimating the volume and fresh weight of community organic waste according to claim 1, characterized in that, In step S1, building the instance segmentation model specifically includes the following steps: Collect images of common solid organic waste in communities, label them using annotation tools, and enrich the image sample data using image augmentation methods; Following the common ratio for small datasets, the dataset was divided into training, validation, and test sets in an 8:1:1 ratio. The resulting dataset was then used to train the Mask R-CNN deep neural network model, thereby obtaining an instance segmentation model for community organic waste.
3. The method for estimating the volume and fresh weight of community organic waste according to claim 1, characterized in that, In step S4, the binocular vision volumetric algorithm involves calculating the depth of the object corresponding to each pixel in the depth image to obtain a depth map. The distance from a point on the measured trash to the camera is also called depth. The measured trash is considered as a combination of many tiny cylinders, the number of which is equal to the number of pixels h×w in the measured trash image after instance segmentation. The base area A of each small cylinder corresponding to each pixel is calculated. u,v and height H u,v The total volume V of the measured waste can then be calculated. The formula for calculating the total volume V is as follows: In the above formula, u = 0, 1, 2, ..., h-1; v = 0, 1, 2, ..., w-1; Height H u,v The calculation method is as follows: Depth data Z' is obtained before and after the organic waste is placed in the garbage community. u,v Z u,v The height from the waste placement platform in front of the community's organic waste to the binocular camera is the depth data Z'. u,v The height of the waste from the binocular camera after the organic waste in the garbage community is the depth data Z. u,v ; Using depth data Z' u,v The reference height G of the waste placement platform was calculated using the average method. u,v Height H u,v The calculation formula is as follows: In the above formula, u = 0, 1, 2, ..., h-1; v = 0, 1, 2, ..., w-1, σ min and σ max These are the minimum and maximum depth thresholds, respectively. Pixels whose depth data exceeds these thresholds are treated as outliers. The base area A of the small cylinder u,v The calculation method is as follows: Based on the projection relationship of binocular vision data, the calculated pixel length and width differ for pixels with different depth values. Therefore, the base area A corresponding to the small cylinder needs to be adjusted according to the depth value of each pixel. u,v The pixel located at point (u,v) corresponds to the length m of the small square. u,v and width n u,v The projection relationship between depth and projection is expressed as follows: fov in the above formula v and fov H These are the vertical and horizontal field of view of the binocular camera, respectively. Therefore, the base area A of the small cylinder corresponding to the pixel at point (u,v) can be obtained. u,v for: A u,v =m u,v n u,v 。 4. The method for estimating the volume and fresh weight of community organic waste according to claim 1, characterized in that, In step S5, the garbage volume f i (V i ) and fresh weight m i The relational model is as follows: m i =f i (V i ) A univariate linear regression model is performed on the calculated sample data volume x and the corresponding actual measured fresh weight y. The model expression is as follows: y = p1x + p2 Where p1 and p2 are the coefficients of the univariate linear regression model. The coefficient of determination (R²) and root mean square error (RMSE) were used to evaluate the regression performance of the univariate linear regression model. The R² is used to measure the extent to which the variance of the dependent variable can be explained by the independent variables, and can be used to assess the explanatory power of the quasi-univariate linear regression model. The formula for calculating the coefficient of determination is as follows: The root mean square error (RMSE) is a commonly used metric to measure the difference between the predicted values of a model and the actual data. In fitting results, a lower RMSE value indicates a lower degree of deviation between the fitted model and the actual data, meaning a better fit. The formula for calculating the RMSE is as follows: To evaluate the performance of the method for estimating the fresh weight of municipal solid waste, relative error (RE) and relative average error (RAE) were used for assessment. For any sample, the actual value y and the predicted value... The formula for calculating relative error is as follows: The formula for calculating the relative average error (RAE) is as follows. The above y i This is the actual value. Indicates the predicted value. This represents the average number of observations, with a total sample size of n.