Intelligent classification method and system for kitchen waste based on AI image recognition
Patent Information
- Application Number
- CN202512004044.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2045-12-29
AI Technical Summary
然而,这些解决方案存在一个显著的技术问题,即对复杂环境下的垃圾图像识别准确率较低
[0006] In summary, the technical solution of this application, by acquiring multi-view visible light image sequences of kitchen waste to be classified and simultaneously obtaining ambient light intensity values and container background identification information, can comprehensively and accurately acquire image features and environmental information of waste samples from different perspectives. Based on this information, an initial feature set is generated, followed by feature fusion processing and removal of low-confidence feature dimensions to construct a multi-source heterogeneous feature vector, making the features richer and more accurate, and better reflecting the essential characteristics of waste. The multi-source heterogeneous feature vector is input into a pre-trained image recognition model for category matching calculation to obtain the initial confidence score of each candidate waste category, improving the accuracy of category matching. The judgment threshold for the current classification task is dynamically calculated based on the confidence score distribution of the most recent N historical classification tasks, and the final classification result is determined by combining the initial confidence score. This approach can adapt to classification tasks under different environments, further improving classification accuracy. Finally, the final classification result is mapped into kitchen waste sorting control instructions and sent to the execution mechanism to complete the physical sorting operation, realizing intelligent classification of kitchen waste and improving classification efficiency and accuracy.
Smart Images

Figure CN122066997B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of environmental protection and information technology, and in particular to a method and system for intelligent sorting of kitchen waste based on AI image recognition. Background Technology
[0002] In today's society, the effective sorting and treatment of kitchen waste is an important issue in the fields of urban environmental sanitation and resource recycling. With the growth of the urban population and the improvement of living standards, the amount of kitchen waste generated is increasing day by day. Traditional kitchen waste treatment methods mainly rely on manual sorting, which is not only inefficient but also labor-intensive, making it difficult to meet the needs of large-scale waste treatment.
[0003] To address the challenges of manual sorting, several machine vision-based waste sorting technologies have emerged. These technologies capture images of waste using cameras and then analyze them using image processing algorithms to classify the waste. For example, some systems employ simple color and shape feature extraction methods for preliminary waste image classification. However, these solutions suffer from a significant technical problem: low accuracy in recognizing waste images in complex environments. Due to the diversity of kitchen waste, variations in ambient lighting, and interference from the background of waste disposal containers, existing machine vision technologies struggle to accurately extract effective features from waste, resulting in unreliable classification results. Summary of the Invention
[0004] The main purpose of this application is to provide a method and system for intelligent classification of kitchen waste based on AI image recognition, which can improve the accuracy and efficiency of kitchen waste classification in complex environments.
[0005] To achieve the above objectives, embodiments of the present invention provide an intelligent sorting method for kitchen waste based on AI image recognition, the method comprising the following steps: Collect a multi-view visible light image sequence of kitchen waste to be sorted, and simultaneously acquire the corresponding ambient light intensity value and container background identification information. The multi-view visible light image sequence includes continuous image frames from three fixed angles: front view, side view, and top view. Based on the multi-view visible light image sequence, the original pixel distribution features are extracted, and an initial feature set is generated by combining the ambient light intensity value and the container background identification information. The initial feature set includes color channel mean vector, texture gradient direction histogram and background interference factor weight. The initial feature set is subjected to feature fusion processing, and low-confidence feature dimensions are removed according to the training sample screening strategy to construct a multi-source heterogeneous feature vector, wherein the multi-source heterogeneous feature vector includes HSV color space component statistics, local structure similarity index and container background correlation coefficient. The multi-source heterogeneous feature vectors are input into a pre-trained image recognition model for category matching calculation to obtain an initial confidence score corresponding to each candidate waste category, wherein each candidate waste category includes food scraps, fruit peels and vegetable leaves, expired food and other biodegradable organic matter. Based on the confidence score distribution of the most recent N historical classification tasks, the judgment threshold for the current classification task is dynamically calculated, and the final classification result is determined based on the judgment threshold and the initial confidence score of each candidate waste category, where N is a positive integer; The final classification result is mapped into a kitchen waste sorting control instruction and sent to the execution mechanism to complete the physical sorting operation.
[0006] In summary, the technical solution of this application, by acquiring multi-view visible light image sequences of kitchen waste to be classified and simultaneously obtaining ambient light intensity values and container background identification information, can comprehensively and accurately acquire image features and environmental information of waste samples from different perspectives. Based on this information, an initial feature set is generated, followed by feature fusion processing and removal of low-confidence feature dimensions to construct a multi-source heterogeneous feature vector, making the features richer and more accurate, and better reflecting the essential characteristics of waste. The multi-source heterogeneous feature vector is input into a pre-trained image recognition model for category matching calculation to obtain the initial confidence score of each candidate waste category, improving the accuracy of category matching. The judgment threshold for the current classification task is dynamically calculated based on the confidence score distribution of the most recent N historical classification tasks, and the final classification result is determined by combining the initial confidence score. This approach can adapt to classification tasks under different environments, further improving classification accuracy. Finally, the final classification result is mapped into kitchen waste sorting control instructions and sent to the execution mechanism to complete the physical sorting operation, realizing intelligent classification of kitchen waste and improving classification efficiency and accuracy. Attached Figure Description
[0007] Figure 1 This is a schematic diagram of a scenario for the intelligent sorting method for kitchen waste based on AI image recognition in the embodiments of this application; Figure 2 A flowchart of an AI image recognition-based intelligent sorting method for kitchen waste is provided for embodiments of this application; Figure 3 This is a schematic flowchart of the feature processing provided in the embodiments of this application; Figure 4 This is a schematic diagram illustrating the process of generating a weighted interference graph provided in an embodiment of this application. Figure 5 This is a schematic diagram illustrating the process of forming multi-source heterogeneous feature vectors provided in an embodiment of this application. Figure 6 A schematic diagram illustrating the process of generating the initial confidence score provided in the embodiments of this application; Figure 7 This is a schematic diagram of the final waste sorting and processing flow provided in the embodiments of this application; Figure 8 A schematic diagram of the structure of the AI image recognition-based intelligent sorting system for kitchen waste provided in this application embodiment; Figure 9 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0008] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0009] This application provides a method and system for intelligent sorting of kitchen waste based on AI image recognition, which will be described in detail below.
[0010] In this embodiment, AI-based image recognition-based intelligent sorting of kitchen waste is a comprehensive method for kitchen waste treatment that utilizes artificial intelligence technology. It involves a comprehensive consideration of multiple factors, aiming to intelligently determine the category of kitchen waste based on its image features and environmental information.
[0011] As shown in Figure 1, a scenario for intelligent sorting of kitchen waste based on AI image recognition is provided. In the kitchen waste treatment center scenario, it includes an image acquisition device, an ambient light sensor, a recognition device, an actuator, and a control center. The image acquisition device, ambient light sensor, recognition device, and actuator are connected to the control center via wired or wireless networks.
[0012] Taking a food waste treatment center in a large city as an example, this center receives a large amount of food waste from various areas every day. This food waste is placed in different waste disposal containers, each with different materials and backgrounds.
[0013] The image acquisition device is installed at a specific location in the waste conveying channel, and it has the function of photographing food waste samples at preset angles. For example, during the waste conveying process, when a food waste sample passes through the imaging area of the image acquisition device, the device will sequentially photograph it from three fixed angles: front view, side view, and top view, generating a sequence of image frames arranged in chronological order. These image frame sequences contain appearance information of food waste from different perspectives, and can more comprehensively reflect the characteristics of the waste.
[0014] An ambient light sensor is installed at a suitable location within the waste treatment center to monitor ambient light intensity in real time. The ambient light intensity at the waste treatment center varies depending on the time of day and weather conditions, which can affect the quality of image acquisition. The ambient light sensor reads the analog signal of the ambient light intensity, converts it into a digital light intensity value, and then synchronizes this value with the timestamp of the image frame sequence. This allows for subsequent image correction based on the light intensity value during processing, reducing the impact of lighting changes on image feature extraction.
[0015] The identification device processes and analyzes the acquired images. Specifically, it first identifies the graphic encoding information on the surface of the waste disposal container, parses out the unique container background identification information, and establishes an index association between this information and the corresponding image frame sequence. For example, containers of different materials and colors may interfere with the waste features in the image to varying degrees; this interference can be assessed and addressed using the container background identification information. Simultaneously, the identification device verifies the image frame sequence. If a blurred or occluded frame is detected, the image acquisition operation for the corresponding angle is retried to ensure that at least one clear image frame is retained from each viewpoint. Finally, the identification device integrates the clear image frames from all viewpoints, the synchronized illumination intensity values, and the container background identification information to form a structured data packet, which is then sent to the control center.
[0016] The control center receives structured data packets from the identification device and processes them further. It extracts raw pixel distribution features from multi-view visible light image sequences and combines them with ambient light intensity values and container background identification information to generate an initial feature set. This initial feature set is then fused, and low-confidence feature dimensions are removed according to a training sample selection strategy to construct a multi-source heterogeneous feature vector. This multi-source heterogeneous feature vector is input into a pre-trained image recognition model for category matching calculation, obtaining an initial confidence score corresponding to each candidate waste category. Based on the confidence score distribution of the most recent N historical classification tasks, the decision threshold for the current classification task is dynamically calculated, and the final classification result is determined based on the decision threshold and the initial confidence scores of each candidate waste category. Finally, the control center maps the final classification result into a food waste sorting control instruction and sends it to the execution mechanism.
[0017] The executing mechanism performs physical sorting of food waste according to received control commands. For example, multiple sorting ports are set up on the waste conveyor channel, each corresponding to a specific waste category. The executing mechanism uses robotic arms or other sorting equipment to accurately sort different categories of food waste into the corresponding sorting ports, achieving the classification and processing of food waste. Furthermore, throughout the entire process, the system records relevant information for each sorting task in real time, including sorting results and confidence scores, for subsequent statistical analysis and model optimization.
[0018] refer to Figure 2 , Figure 2 This is a flowchart illustrating an AI-based image recognition-based intelligent sorting method for kitchen waste provided in this application embodiment. The execution entity of this method can be a computer device (such as a control center), such as a server, etc. The AI-based image recognition-based intelligent sorting method for kitchen waste provided in this application embodiment specifically includes:
[0019] S10: Collect a multi-view visible light image sequence of kitchen waste to be sorted, and simultaneously acquire the corresponding ambient light intensity value and container background identification information. The multi-view visible light image sequence includes continuous image frames from three fixed angles: front view, side view, and top view.
[0020] In this embodiment, the multi-view visible light image sequence refers to a sequence of visible light images of kitchen waste to be classified, acquired from multiple different angles. These images provide information about the appearance of kitchen waste from different perspectives, helping to gain a more comprehensive understanding of the waste's characteristics. For example, a front view image can show the frontal features of the waste, a side view image can provide information about the side profile of the waste, and a top view image can present the top features of the waste. Ambient light intensity refers to the strength of light in the environment, which affects the brightness and color of the image. Container background identification information refers to the relevant identification of the waste disposal container, such as the material and color of the container. This information can be used to assess the degree of interference of the container background on the waste features in the image.
[0021] In one embodiment, step S10 can be implemented as follows: F1: Trigger the image acquisition device to sequentially capture kitchen waste samples at preset angle positions, generate a sequence of image frames arranged in chronological order, label the image frame sequence with corresponding viewpoint tags and store them in the cache area.
[0022] In this embodiment, the image acquisition device is a device used to capture images of kitchen waste samples, such as a camera. The preset angle position is a pre-set shooting angle, typically including three fixed angles: front view, side view, and top view. The image frame sequence is a series of image frames arranged in chronological order. Viewpoint labels are labels used to identify the shooting perspective of each image frame, such as front view, side view, and top view. The buffer area is an area used to temporarily store the image frame sequence.
[0023] For example, in a food waste treatment center, when a food waste sample is transported to a designated location, the system triggers cameras installed at preset angles to take pictures sequentially. Assuming the first camera is positioned at a direct line of sight, it will capture a series of image frames from this angle, arranged chronologically to form an image frame sequence. The system then labels each image frame with a "direct line of sight" label and stores it in a buffer.
[0024] F2: Read the analog signal output by the ambient light sensor, convert it into a digital light intensity value, and synchronize the light intensity value with the timestamp of the image frame sequence.
[0025] In this embodiment, the ambient light sensor is a device used to detect ambient light intensity, and its output is an analog signal. The digital light intensity value is a numerical value that can be processed by a computer after converting the analog signal. The timestamp is information used to identify the time when an image frame was captured.
[0026] For example, an ambient light sensor is installed at a location in a waste treatment center. It detects ambient light intensity in real time and outputs an analog signal. The system periodically reads this analog signal and converts it into digital light intensity values using an analog-to-digital converter. Assuming that when capturing a sequence of image frames, the start and end times of the sequence are recorded as timestamps, the corresponding light intensity values within that time period are then synchronized and bound to the image frame sequence.
[0027] F3: Identify the graphic encoding information on the surface of the waste disposal container, parse out the unique container background identification information, and establish an index association between the container background identification information and the corresponding image frame sequence.
[0028] In this embodiment, the graphic encoding information refers to the graphic encoding used to identify container information, such as QR codes or barcodes, set on the surface of the waste disposal container. The container background identification information is the unique information parsed from the graphic encoding information that identifies the container background, such as the container's material and color. Index association refers to establishing a relationship between the container background identification information and the corresponding image frame sequence, so that the image frame sequence can be processed based on the container background identification information in subsequent processing.
[0029] For example, a QR code is affixed to the surface of the waste disposal container, containing information such as the container's material and color. When a sample of food waste is placed in this container and passes through the image acquisition area, the system uses image recognition technology to identify the QR code and extract the container's background identification information. Then, it establishes an index association between this container background identification information and the sequence of image frames of food waste captured in the container.
[0030] F4: Check if there are blurry or occluded frames in the image frame sequence. If so, re-trigger the image acquisition operation at the corresponding angle to ensure that at least one clear image is retained for each viewpoint.
[0031] In this embodiment, a blurred or occluded frame refers to an image frame that is either blurry or partially occluded, which can affect subsequent feature extraction and classification results. Re-triggering the image acquisition operation at the corresponding angle means that when a blurred or occluded image frame is detected at a certain viewpoint, the image acquisition device at that viewpoint is triggered again to capture the image.
[0032] For example, when capturing a sequence of images from a frontal viewpoint, if one frame is found to be blurry due to lighting conditions or object obstruction, the system will detect this and re-trigger the frontal viewpoint image acquisition device to capture the image until a clear frontal viewpoint image is obtained.
[0033] In one embodiment, an image sharpness evaluation algorithm can be used to evaluate each frame in the image frame sequence. If the evaluation result shows that the sharpness of a certain frame is lower than a preset threshold, the frame is considered a blurry frame. For occluded frames, image segmentation and object detection algorithms can be used to detect whether there are abnormal occluded areas in the image. When a blurry or occluded frame is detected, the system sends a re-trigger signal to the image acquisition device at the corresponding angle to re-capture the image.
[0034] F5: Integrates clear image frames from all perspectives, synchronized illumination intensity values, and container background identification information to form a structured data packet.
[0035] In this embodiment of the application, a structured data packet refers to a data set that organizes clear image frames from all perspectives, synchronized illumination intensity values, and container background identification information according to a certain structure, so as to facilitate subsequent transmission and processing.
[0036] For example, clear image frames from three perspectives—front, side, and top—are stored in different fields of the data packet, and synchronized illumination intensity values and container background identification information are also stored in the corresponding fields. This creates a structured data packet containing all the necessary information.
[0037] S20: Extract the original pixel distribution features based on the multi-view visible light image sequence, and generate an initial feature set by combining the ambient light intensity value and the container background identification information, wherein the initial feature set includes the color channel mean vector, texture gradient direction histogram and background interference factor weight.
[0038] In this embodiment, the original pixel distribution feature refers to the characteristics reflected by the distribution of pixels in the image, which can reflect information such as color and texture of the image. The color channel mean vector is a vector composed of the pixel mean values of each color channel in the image, which can describe the overall color distribution of the image. The texture gradient direction histogram refers to the statistical information of the gradient direction distribution of texture in the image, which can reflect the texture features of the image. The background interference factor weight is a quantified value of the degree of background interference determined based on the container background identification information, used to evaluate and process background interference in subsequent processing.
[0039] In one embodiment, for calculating the color channel mean vector, the pixel mean and variance of the RGB three channels can be calculated for each image frame from each viewpoint to generate the color channel mean vector, which is then normalized and used as the basic color feature. For constructing the texture gradient direction histogram, gradient operators can be used to perform edge detection on the image frames, and the distribution frequency of gradient magnitudes in each direction can be statistically analyzed to obtain the texture gradient direction histogram. For determining the background interference factor weights, a pre-stored background interference level table can be queried based on the container background identification information to obtain the corresponding background interference factor weights. By merging these features, an initial feature set containing color, texture, and background interference information can be generated, providing richer feature information for subsequent feature fusion and classification.
[0040] S30: Perform feature fusion processing on the initial feature set and remove low-confidence feature dimensions according to the training sample screening strategy to construct a multi-source heterogeneous feature vector, wherein the multi-source heterogeneous feature vector includes HSV color space component statistics, local structure similarity index and container background correlation coefficient.
[0041] In this embodiment, feature fusion processing refers to integrating multiple features in the initial feature set to generate more representative and discriminative features. The training sample selection strategy refers to determining which feature dimensions contribute less to the classification result based on the features and classification results of the training samples, and then removing them. Multi-source heterogeneous feature vectors are vectors composed of features from different sources and types, which can more comprehensively describe the features of the object to be classified. HSV color space component statistics refer to the results obtained by statistically analyzing color components in the HSV color space, which can more intuitively reflect the hue, saturation, and brightness information of colors. Local structural similarity index refers to a quantitative indicator of the degree of structural similarity between adjacent viewpoint image frames, which can reflect the local structural features of the image. Container background correlation coefficient refers to a quantitative value of the semantic consistency between the container background and the current image content, which can be used to evaluate the impact of the container background on image features.
[0042] In one embodiment, for feature fusion processing, color features can be converted to the HSV color space, and statistical moments of hue, saturation, and brightness can be calculated to generate HSV color space component statistics. The structural similarity index between adjacent viewpoint image frames is calculated, and the minimum value is taken as the local structural similarity index. Based on the semantic consistency between the container background identification information and the current image content, the container background correlation coefficient is calculated. These features are included in a feature pool, and then each feature dimension in the feature pool is traversed, comparing its discrimination score to the corresponding category in the training set. If the discrimination score is lower than a preset baseline, it is marked as a low-confidence feature dimension. All marked low-confidence feature dimensions are removed from the feature pool, and the remaining features constitute a multi-source heterogeneous feature vector. In this way, a more representative and discriminative multi-source heterogeneous feature vector can be constructed, improving classification accuracy.
[0043] S40: Input the multi-source heterogeneous feature vector into the pre-trained image recognition model to perform category matching calculation and obtain the initial confidence score corresponding to each candidate waste category, wherein each candidate waste category includes food scraps, fruit peels and vegetable leaves, expired food and other biodegradable organic matter.
[0044] In this embodiment, the pre-trained image recognition model refers to a model trained on a large amount of image data, which can classify and recognize the features of the input image. Category matching calculation refers to matching the input multi-source heterogeneous feature vector with the features of each predefined candidate waste category in the model, calculating the matching degree of each candidate waste category, i.e., the initial confidence score. The initial confidence score refers to the probability that each candidate waste category will be matched before final classification.
[0045] In one embodiment, the multi-source heterogeneous feature vectors are divided into three subsets: color sub-vectors, structure sub-vectors, and background correlation sub-vectors. These three subsets are then input into the corresponding parallel convolutional branch networks of a pre-trained image recognition model for feature response calculation. Cross-branch feature weighting and fusion are performed at the output of the parallel convolutional branch networks to generate a fused feature representation. Based on the fused feature representation, an initial confidence score for each candidate waste category is calculated through a fully connected classification layer. The initial confidence scores are normalized so that the sum of the initial confidence scores for all candidate waste categories equals 1. The normalized initial confidence scores are then used as the output of the image recognition model, thus obtaining the initial confidence score corresponding to each candidate waste category. In this way, a pre-trained image recognition model can be used to accurately classify and recognize multi-source heterogeneous feature vectors.
[0046] S50: Based on the confidence score distribution of the most recent N historical classification tasks, dynamically calculate the judgment threshold for the current classification task, and determine the final classification result based on the judgment threshold and the initial confidence score of each candidate waste category, where N is a positive integer.
[0047] In this embodiment, the confidence score distribution of the most recent N historical classification tasks refers to the distribution of the initial confidence scores of each candidate waste category in the most recent N classification tasks. The decision threshold is a critical value used to determine whether a candidate waste category is the final classification result. The final classification result refers to the actual category of the kitchen waste to be classified, determined based on the decision threshold and the initial confidence scores of each candidate waste category.
[0048] In one embodiment, the highest confidence score output from each of the most recent N historical classification tasks is obtained, forming a confidence score sequence. A sliding window statistical analysis is performed on the confidence score sequence to calculate the mean and standard deviation of the confidence score within the current window. Based on the mean and standard deviation of the confidence score and the container background identification information, the decision threshold for the current classification task is calculated. The initial confidence scores of each candidate waste category are traversed, and candidate categories with initial confidence scores greater than or equal to the decision threshold are selected. If a unique candidate category meets the selection criteria, that candidate category is determined as the final classification result; if multiple candidate categories meet the selection criteria, the candidate category with the highest initial confidence score is selected as the final classification result. In this way, the decision threshold can be dynamically adjusted according to the historical classification task situation, improving the accuracy and reliability of the classification results.
[0049] S60: Map the final classification result into a kitchen waste sorting control instruction and send it to the execution mechanism to complete the physical sorting operation.
[0050] In this embodiment, the final classification result mapping refers to converting the determined final classification result into control instructions that the executing mechanism can understand and execute. The food waste sorting control instructions are instructions used to control the executing mechanism to physically sort food waste. The executing mechanism refers to the equipment responsible for the actual sorting operation of food waste, such as robotic arms, conveyor belts, etc.
[0051] In one embodiment, the final classification result is converted into a corresponding food waste sorting control instruction according to a pre-set mapping rule. For example, if the final classification result is food scraps, a corresponding control instruction is generated, instructing the actuator to sort the food waste to the corresponding sorting port for food scraps. The generated control instruction is sent to the actuator via a communication network. Upon receiving the instruction, the actuator performs physical sorting of the food waste according to the instruction's requirements. In this way, automatic classification and sorting of food waste can be achieved, improving the efficiency and accuracy of waste disposal.
[0052] In one embodiment, reference Figure 3 Step S20 may include steps S21-S25, which will be described in detail below: S21: Calculate the pixel mean and variance of the RGB three channels for each image frame from each viewpoint, generate a color channel mean vector, and normalize the color channel mean vector as the basic color feature.
[0053] In this embodiment, the RGB three channels refer to the three color channels of red (R), green (G), and blue (B) in an image. The pixel mean is the average value of all pixels in each color channel, and the variance is the degree of dispersion of pixel values relative to the mean. The color channel mean vector is a vector composed of the pixel means of the three color channels. Normalization refers to adjusting the value range of the color channel mean vector to a specific interval, typically [0, 1]. The basic color feature refers to the normalized color channel mean vector, which can be used to describe the overall color distribution of the image.
[0054] For example, for an image frame viewed from the front, calculate the pixel mean and variance of its R, G, and B channels respectively. Assuming the pixel mean of the R channel is 120, the pixel mean of the G channel is 100, and the pixel mean of the B channel is 80, the generated color channel mean vector is [120, 100, 80]. Then, normalize this vector, adjusting its value range to [0, 1], to obtain the normalized basic color features.
[0055] In one embodiment, functions from an image processing library can be used to calculate the pixel mean and variance of the RGB three channels. For example, using Python's OpenCV library, the pixel mean can be obtained by iterating through each pixel of an image frame, calculating the sum of the pixel values for the R, G, and B channels, and then dividing by the total number of pixels. The variance can be calculated by summing the squares of the differences between each pixel value and the mean, and then dividing by the total number of pixels. Normalization can be achieved by dividing each element of the color channel mean vector by its maximum value.
[0056] S22: Use gradient operators to perform edge detection on image frames, count the distribution frequency of gradient magnitudes in each direction, and construct a texture gradient direction histogram.
[0057] In this embodiment, the gradient operator is used to calculate the rate of change of pixel grayscale values in an image; common examples include the Sobel operator and the Prewitt operator. Edge detection refers to detecting regions in an image where pixel grayscale values change significantly, i.e., the edges of the image, using gradient operators. Gradient magnitude refers to the magnitude of the gradient vector, which represents the degree of change in pixel grayscale values. The texture gradient direction histogram is a histogram obtained by statistically analyzing the frequency distribution of gradient magnitudes in various directions in an image; it can reflect the texture features of the image.
[0058] For example, for an image frame viewed from the side, the Sobel operator is used for edge detection. The Sobel operator calculates the gradient values of each pixel in the image in the horizontal and vertical directions, and then calculates the gradient magnitude and gradient direction using formulas. The frequency distribution of gradient magnitudes in different gradient directions is then statistically analyzed. For example, the gradient direction can be divided into 0°-180° intervals, with each interval representing 10°, and the number of gradient magnitudes within each interval is counted to construct a texture gradient direction histogram.
[0059] In one embodiment, gradient operators and texture gradient orientation histograms can be constructed using functions from image processing libraries. For example, using Python's OpenCV library, the Sobel function can be called to calculate the gradient values of the image, and then a custom statistical function can be used to count the frequency distribution of gradient magnitudes in each direction to construct a texture gradient orientation histogram.
[0060] S23: Query the pre-stored background interference level table according to the container background identification information, obtain the corresponding background interference factor weight, and multiply the background interference factor weight with the image frame region mask to generate a weighted interference map.
[0061] In this embodiment, the background interference level table is a pre-stored table that records the background interference levels corresponding to different container background identification information. The background interference factor weight is a numerical value determined based on the background interference level, used to represent the degree of interference of the container background on the garbage features in the image. The image frame region mask is a binary image used to identify the main garbage region in the image, where the pixel value of the main garbage region is 1, and the pixel value of the background region is 0. The weighted interference map is the image obtained by multiplying the background interference factor weight by the image frame region mask, and it is used to evaluate and process background interference in subsequent processing.
[0062] For example, by querying a pre-stored background interference level table based on the container background identification information, it is found that the interference level corresponding to the container background is medium, and the corresponding background interference factor weight is 0.5. Foreground segmentation is performed on the image frame to generate an image frame region mask. This mask is then multiplied by the background interference factor weight of 0.5 to obtain a weighted interference map.
[0063] In one embodiment, a database management system can be used to store and query the background interference level table. After obtaining the container background identification information, the corresponding background interference factor weights are retrieved using a database query. For generating the image frame region mask, image segmentation algorithms, such as threshold-based or deep learning-based segmentation algorithms, can be used. The background interference factor weights are multiplied element-wise by the image frame region mask to obtain a weighted interference map.
[0064] In one embodiment, reference Figure 4 Step S23 may include steps S231-S235, which will be described in detail below: S231: Parse the material type code and color code in the container background identification information, match the predefined background interference level table entries, and read the corresponding interference level value.
[0065] In this embodiment, the material type code refers to the code used to identify the material type of the container, such as "01" for plastic and "02" for metal. The color code refers to the code used to identify the color of the container, such as "FF0000" for red and "0000FF" for blue. The predefined background interference level table entries are each item in a pre-set table, which contains the interference level values corresponding to different combinations of material type codes and color codes.
[0066] For example, if the container background identification information contains the material type code "01" and the color code "FF0000", it indicates that the container is a red plastic container. Query the predefined background interference level table, find the entry with the material type code "01" and the color code "FF0000", and read the corresponding interference level value, let's say it's 3.
[0067] In one embodiment, a string parsing algorithm can be used to parse the container background identification information to extract the material type code and color code. Then, by traversing a predefined background interference level table, a matching entry is found, and the corresponding interference level value is read.
[0068] S232: Convert the interference level value into background interference factor weights through a non-linear mapping function, ensuring that the weight values are within a preset range (0 to 1).
[0069] In this embodiment, the nonlinear mapping function is used to convert interference level values into background interference factor weights. It can nonlinearly adjust the weights according to different interference levels. The preset interval (0 to 1) refers to the range of values for the background interference factor weights, ensuring that the weight values are within this range for convenient subsequent calculations and processing.
[0070] For example, assuming the interference level value is 3, it can be converted into a background interference factor weight using a sigmoid nonlinear mapping function. The sigmoid function can map the input interference level value to the interval [0, 1], and assume the converted background interference factor weight is 0.6.
[0071] In one embodiment, a nonlinear mapping function can be implemented using Python's mathematical library. For example, an sigmoid function can be defined, taking the interference level value as input, and calculating the background interference factor weights. During the calculation, it is necessary to ensure that the weight values are within a preset interval (0 to 1); if they exceed the interval, truncation is performed.
[0072] In one embodiment, step S232 can be implemented as follows: A1: Set the upper and lower limits of the input range for the interference level value, and truncate values that exceed the range.
[0073] In this embodiment, the upper and lower limits of the input range are preset ranges of interference level values, for example, the lower limit is 1 and the upper limit is 5. Truncation processing refers to adjusting the interference level value to the upper or lower limit value when it exceeds the input range.
[0074] For example, if the input interference level value is 6, which exceeds the preset upper limit of 5, it will be truncated to 5. If the input interference level value is 0, which is lower than the preset lower limit of 1, it will be truncated to 1.
[0075] A2: The truncated interference level values are normalized and mapped using an S-shaped function to generate preliminary weight values.
[0076] In this embodiment, the sigmoid function is a commonly used nonlinear function with a shape resembling the letter "S," which can map input values to the interval [0, 1]. Normalization mapping refers to converting the truncated interference level values into preliminary weight values within the interval [0, 1] using the sigmoid function.
[0077] For example, if the truncated interference level is 4, it can be used as the input to an S-shaped function to calculate a preliminary weight value, which is assumed to be 0.8.
[0078] A3: Apply an upward slope constraint to the initial weight values to prevent drastic fluctuations in weights due to minor changes in level.
[0079] In this embodiment of the application, the rising slope constraint refers to limiting the rate of change of the initial weight value to prevent the weight value from fluctuating drastically when the interference level value changes slightly.
[0080] For example, when the interference level changes from 4 to 4.1, the initial weight value might change from 0.8 to 0.9. Without applying an upward slope constraint, the change in weight value could be too drastic. By applying an upward slope constraint, the range of weight value variation can be limited. Assuming the upward slope is limited to within 0.1, the change in weight value will not exceed 0.1, i.e., from 0.8 to 0.81.
[0081] A4: Verify whether the initial weight value is within the closed interval of 0 to 1. If not, force it to be limited to the boundary value.
[0082] In this embodiment, the closed interval from 0 to 1 refers to the legal range of the initial weight value, i.e., [0, 1]. Forcing the limit to the boundary value means that when the initial weight value exceeds the [0, 1] interval, it is adjusted to 0 or 1.
[0083] For example, if the initial weight value is 1.2, which is outside the range [0, 1], it will be forced to be 1. If the initial weight value is -0.1, which is below 0, it will be forced to be 0.
[0084] A5: Output the weights of background interference factors that conform to the interval constraints.
[0085] In this embodiment of the application, the background interference factor weight that meets the interval constraint refers to the background interference factor weight that is in the interval [0, 1] after the above processing.
[0086] For example, after truncation, normalization mapping, rising slope constraint, and boundary value limitation, the background interference factor weight is 0.7. This weight conforms to the interval constraint and can be used for subsequent calculations and processing.
[0087] S233: Perform foreground segmentation on the image frame to generate a binary region mask, preserving the pixel markers of the main area of kitchen waste.
[0088] In this embodiment, foreground segmentation refers to the operation of separating the foreground (kitchen waste) from the background in an image. Binarization region masking refers to converting the foreground-segmented image into a binary image, where the pixel value of the foreground region is 1 and the pixel value of the background region is 0.
[0089] For example, for an image frame viewed from above, a deep learning-based foreground segmentation algorithm is used to segment the image, resulting in a foreground (the main body of food waste) and a background. Then, the pixel values of the foreground region are set to 1, and the pixel values of the background region are set to 0, generating a binarized region mask.
[0090] In one embodiment, a deep learning model, such as the U-Net model, can be used to segment the foreground of an image frame. The image is input into a trained U-Net model, which outputs a segmentation result, which is then converted into a binary image to obtain a binary region mask.
[0091] S234: Multiply the background interference factor weights element by element with the region mask to obtain the attenuation coefficient matrix that only applies to the background region.
[0092] In this embodiment, the attenuation coefficient matrix is a matrix obtained by multiplying the background interference factor weights and the region mask element by element, which is used to attenuate the pixel values of the background region.
[0093] For example, suppose the background interference factor weight is 0.5, and the region mask is a binary image where the pixel value of the foreground region is 1 and the pixel value of the background region is 0. Multiplying the background interference factor weight element-wise with the region mask, the resulting attenuation coefficient matrix has a pixel value of 0.5 × 1 = 0.5 for the foreground region and a pixel value of 0.5 × 0 = 0 for the background region.
[0094] S235: The attenuation coefficient matrix is superimposed on the pixel values of the original image frame to generate a weighted interference map.
[0095] In this embodiment, the overlay operation refers to adding the pixel values of the attenuation coefficient matrix to the pixel values of the original image frame. The weighted interference map is the image obtained after the overlay operation, which takes into account the influence of background interference factors.
[0096] For example, the pixel values of the attenuation coefficient matrix are added element-by-element to the pixel values of the original image frame to obtain the weighted interference map. Assuming a pixel in the original image frame has a pixel value of 100 and the corresponding pixel in the attenuation coefficient matrix has a pixel value of 20, then the pixel value of that pixel in the weighted interference map is 100 + 20 = 120.
[0097] S24: Map ambient light intensity values to light compensation coefficients, perform linear correction on the color channel mean vector, and generate light-corrected color features.
[0098] In this embodiment, the illumination compensation coefficient is a coefficient determined based on the ambient light intensity value, used to correct the color channel mean vector for illumination. Linear correction refers to adjusting the color channel mean vector through linear operations to compensate for the impact of changes in light intensity on color features. The color features after illumination correction are the color channel mean vectors after illumination compensation processing, which can more accurately reflect the true color information of the image.
[0099] For example, when the ambient light intensity is high, the image may appear brighter, and the values in the color channel mean vector will also be correspondingly larger. By mapping this light intensity value to a light compensation coefficient less than 1, such as 0.8, and then multiplying each element of the color channel mean vector by this coefficient, a linear correction of the color channel mean vector can be achieved. Assuming the color channel mean vector is [120, 100, 80], after correction it becomes [120×0.8, 100×0.8, 80×0.8] = [96, 80, 64], thus obtaining the color features after light correction.
[0100] In one embodiment, a mapping table between light intensity values and light compensation coefficients can be established, allowing the corresponding compensation coefficient to be looked up based on the real-time acquired ambient light intensity value. Alternatively, a linear function can be used for mapping; for example, a base light intensity value and a corresponding standard compensation coefficient can be set, and the actual compensation coefficient can be calculated based on the ratio between the current light intensity value and the base value. The obtained light compensation coefficient is then multiplied element-wise with the color channel mean vector to perform a linear correction of the color channel mean vector, generating the light-corrected color features. This reduces the impact of light variations on color feature extraction and improves the accuracy of subsequent classification.
[0101] S25: Combine the illumination-corrected color features, texture gradient direction histogram, and weighted interference map to form the initial feature set.
[0102] In this embodiment, the merging operation combines three different types of features—the illumination-corrected color features, the texture gradient direction histogram, and the weighted interference map—to form a feature set containing multiple types of information. This initial feature set forms the basis for subsequent feature fusion and classification processing, integrating information from color, texture, and background interference.
[0103] For example, the illumination-corrected color features can be represented as vectors, the texture gradient direction histogram as an array, and the weighted interference map as an image matrix. Feature merging can be achieved by arranging these features in a specific order or storing them in a data structure. For instance, the color feature vectors can be placed first, followed by the texture gradient direction histogram array, and finally the weighted interference map matrix stored in a specific format to form the initial feature set.
[0104] In one embodiment, data structures such as dictionaries or lists in Python can be used to store and merge these features. The illumination-corrected color features, texture gradient direction histograms, and weighted interference maps can be used as different keys in the dictionary, or they can be added sequentially to a list. This allows for convenient access and use of these features in subsequent processing, providing rich information for building a more effective classification model. By merging these features, the characteristics of the kitchen waste to be classified can be more comprehensively described, improving the accuracy and reliability of the classification.
[0105] In one embodiment, reference Figure 5 Step S30 may include steps S31-S35, which will be described in detail below: S31: Convert the color features to the HSV color space and calculate the statistical moments of hue, saturation, and brightness to generate HSV color space component statistics, and include the HSV color space component statistics into the feature pool.
[0106] In this embodiment, the HSV color space is a color representation method that better conforms to human visual perception, consisting of three components: hue (H), saturation (S), and lightness (V). Statistical moments are numerical values describing the distribution characteristics of data; common examples include mean and variance. HSV color space component statistics refer to the statistical moments calculated from the hue, saturation, and lightness components in the HSV color space, which can more comprehensively reflect the characteristics of color. A feature pool is a collection used to store various features; incorporating HSV color space component statistics into the feature pool can provide more feature information for subsequent feature fusion and classification.
[0107] For example, for color features after illumination correction, they are converted from the RGB color space to the HSV color space. Assuming the conversion yields a set of HSV values, the mean and variance of the hue component, the mean and skewness of the saturation component, and the mean and kurtosis of the lightness component are calculated. These statistical values constitute the HSV color space component statistics. These statistics are then added to the feature pool.
[0108] In one embodiment, color features can be converted from the RGB color space to the HSV color space using functions from an image processing library. For example, in Python's OpenCV library, the corresponding functions can be called to perform the conversion. Then, the required statistical moments are calculated for each of the converted HSV components. Statistical functions from the NumPy library can be used to calculate statistics such as mean, variance, skewness, and kurtosis. The calculated HSV color space component statistics are stored in an array or list and added to the feature pool. In this way, richer information can be extracted from the color features, improving classification accuracy.
[0109] S32: Calculate the structural similarity index between adjacent viewpoint image frames, take the minimum value as the local structural similarity index, and add the local structural similarity index to the feature pool.
[0110] In this embodiment, adjacent viewpoint image frames refer to image frames taken from different but adjacent viewpoints, such as frontal and side viewpoints. The structural similarity index is an indicator used to measure the degree of structural similarity between two images, taking into account factors such as brightness, contrast, and structure. The local structural similarity index is the minimum value selected from the structural similarity indices of adjacent viewpoint image frames, reflecting the differences in local structure between the images. Adding the local structural similarity index to the feature pool can provide information about the local structure of the image for subsequent feature fusion and classification.
[0111] For example, frontal and side view image frames are selected to form the first pair of viewpoints, and the structural similarity index between them is calculated, assumed to be 0.8; side view and top view image frames are selected to form the second pair of viewpoints, and the calculated structural similarity index is 0.7; frontal and top view image frames are selected to form the third pair of viewpoints, and the calculated structural similarity index is 0.6. The magnitudes of these three similarity indices are compared, and the minimum value of 0.6 is selected as the local structural similarity index, which is then written into the feature pool in floating-point format.
[0112] In one embodiment, step S32 can be implemented as follows: B1: Select frontal and side view image frames to form the first pair of viewpoint combinations, calculate the structural similarity index between the two, and record it as the first similarity value.
[0113] In this embodiment, the frontal and side view image frames are image frames taken from different perspectives, providing information about the appearance of kitchen waste from the front and side, respectively. The structural similarity index is a numerical value that measures the degree of structural similarity between the two images. The first similarity value refers to the calculated result of the structural similarity index between the frontal and side view image frames.
[0114] For example, in a food waste disposal scenario, an image acquisition device captures image frames of the same food waste sample from both frontal and side views. The Structural Similarity Index (SSIM) algorithm is used to calculate the similarity between these two images. Assuming the calculated SSIM is 0.85, this value is recorded as the first similarity value.
[0115] B2: Select side view and top view image frames to form a second pair of viewpoint combinations, calculate the structural similarity index between the two, and record it as the second similarity value.
[0116] In this embodiment, the side view and top view image frames respectively show the side and top information of the kitchen waste. The second similarity value is the calculated result of the structural similarity index between the side view and top view image frames.
[0117] For example, an image acquisition device captures image frames of a food waste sample from side and top views. The SSIM algorithm is used to calculate the structural similarity index between the two images, assuming a value of 0.78, which is recorded as the second similarity value.
[0118] B3: Select frontal and top-view image frames to form a third pair of viewpoints, calculate the structural similarity index between the two, and record it as the third similarity value.
[0119] In this embodiment, the frontal and top-view image frames provide feature information about the front and top of the kitchen waste. The third similarity value is the calculated result of the structural similarity index between the frontal and top-view image frames.
[0120] For example, the SSIM algorithm is used to calculate the structural similarity index of food waste images taken from frontal and top-down angles. Assuming the obtained structural similarity index is 0.82, this value is recorded as the third similarity value.
[0121] B4: Compare the first similarity value, the second similarity value, and the third similarity value, and select the minimum value as the local structural similarity index.
[0122] In this embodiment, by comparing the magnitudes of these three similarity values, the case with the lowest structural similarity of the image under different viewpoint combinations can be identified. The local structural similarity index is the minimum value selected from these three similarity values, which can more sensitively reflect the differences in the local structure of the image.
[0123] For example, the first similarity value is 0.85, the second similarity value is 0.78, and the third similarity value is 0.82. Comparing these three values, the second similarity value of 0.78 is the smallest, and it is used as the local structural similarity index.
[0124] B5: Write the local structural similarity index into the feature pool in floating-point format.
[0125] In this embodiment, the floating-point format is a digital format used to represent real numbers, which can accurately represent the value of the local structural similarity index. The feature pool is a collection that stores various features; writing the local structural similarity index into the feature pool facilitates subsequent feature fusion and classification processing.
[0126] For example, a local structural similarity metric of 0.78 can be stored in the feature pool as a floating-point number. The feature pool can be stored using data structures such as lists or dictionaries in Python, adding the local structural similarity metric as an element to the list or storing it as a key-value pair in the dictionary.
[0127] S33: Based on the semantic consistency between the container background identifier information and the current image content, calculate the container background correlation coefficient and write the container background correlation coefficient into the feature pool.
[0128] In this embodiment, semantic consistency refers to the degree of agreement between the meaning represented by the container background identifier information and the meaning expressed by the current image content. The container background correlation coefficient is a value calculated based on the semantic consistency degree, which is used to measure the degree of association between the container background and the image content. Writing the container background correlation coefficient into the feature pool can provide information about the association between the container background and the image content for classification.
[0129] For example, the container background information indicates that the container is used to store fruit, while the current image content shows fruit peels and vegetable leaves. The semantic consistency between them can be calculated using semantic analysis algorithms. Assuming the consistency is 0.6, it can be converted into a container-background correlation coefficient according to certain calculation rules, such as 0.6, and this correlation coefficient can be written into the feature pool.
[0130] In one embodiment, natural language processing and image recognition techniques can be used to calculate semantic consistency. Text analysis is performed on the container background identifier information to extract key semantic information, while the image content is simultaneously identified and classified to obtain its semantic information. Then, the semantic consistency degree is obtained by calculating the similarity between these two semantic information sets. According to a preset conversion rule, the semantic consistency degree is converted into a container background correlation coefficient and stored in a feature pool. By considering the container background correlation coefficient, the interference of the container background on the classification results can be reduced, improving the accuracy of classification.
[0131] S34: Traverse each feature dimension in the feature pool and compare its discrimination score with the corresponding category in the training samples. If the discrimination score is lower than the preset baseline, it is marked as a low-confidence feature dimension.
[0132] In this embodiment, each feature dimension in the feature pool represents different aspects of information, such as color features and texture features. The discrimination score is a quantitative indicator of a feature's ability to distinguish between different categories in training samples. The preset baseline is a pre-set threshold used to determine whether the discrimination of a feature is sufficient. Low-confidence feature dimensions refer to feature dimensions with discrimination scores below the preset baseline; these features contribute little to classification.
[0133] For example, the feature pool includes feature dimensions such as color features, texture features, and container-background correlation coefficients. For each feature dimension, its discrimination score against different categories (such as food scraps, fruit peels, and vegetable leaves) is calculated in the training samples. Assuming the discrimination score for color features is 0.8, the discrimination score for texture features is 0.6, and the discrimination score for container-background correlation coefficients is 0.3, with a preset baseline of 0.5, then the discrimination score for the container-background correlation coefficient feature dimension is lower than the preset baseline, and it is marked as a low-confidence feature dimension.
[0134] In one embodiment, feature selection methods from machine learning algorithms can be used to calculate the discriminative score for each feature dimension. For example, methods such as analysis of variance (ANOVA) can be used to analyze each feature dimension in the training samples and calculate its ability to distinguish between different categories. The calculated discriminative score is compared with a preset baseline; if it is lower than the baseline, it is marked as a low-confidence feature dimension. By marking low-confidence feature dimensions, these features that do not contribute much to classification can be removed in subsequent processing, reducing computational load and improving classification efficiency and accuracy.
[0135] S35: Remove all labeled low-confidence feature dimensions from the feature pool and construct a multi-source heterogeneous feature vector from the remaining features.
[0136] In this embodiment, the removal operation refers to deleting feature dimensions marked as low confidence from the feature pool. A multi-source heterogeneous feature vector is a vector composed of the remaining features, containing different types of features from different sources (such as color, texture, container background, etc.), which can more comprehensively describe the characteristics of the kitchen waste to be classified.
[0137] For example, the feature pool originally contained color features, texture features, container-background correlation coefficients, and other feature dimensions. After the labeling in step S34, it was found that the container-background correlation coefficient was a low-confidence feature dimension. It was removed from the feature pool, and the remaining color features, texture features, and other retained features constituted a multi-source heterogeneous feature vector.
[0138] In one embodiment, reference Figure 6 Step S40 may include steps S41-S46, which will be described in detail below: S41: Divide the multi-source heterogeneous feature vector into three subsets: color sub-vector, structure sub-vector, and background correlation sub-vector.
[0139] In this embodiment, the multi-source heterogeneous feature vector contains feature information from different aspects, such as color, structure, and background association. The color sub-vector is a subset of color-related features in the multi-source heterogeneous feature vector, which can reflect the color characteristics of food waste. The structure sub-vector is a subset containing features related to the structure of food waste, such as texture and shape. The background association sub-vector is a subset of features related to the container background, which can reflect the influence of the container background on food waste classification.
[0140] For example, multi-source heterogeneous feature vectors include color features (such as HSV color space component statistics), texture features (such as texture gradient direction histograms), and container-background correlation coefficients. The color features are extracted to form a color sub-vector, the texture features are used to form a structure sub-vector, and the container-background correlation coefficients are used to form a background correlation sub-vector.
[0141] S42: Input the color sub-vector, structure sub-vector, and background association sub-vector into the corresponding parallel convolutional branch network in the pre-trained image recognition model to calculate the feature response.
[0142] In this embodiment, the pre-trained image recognition model is a model trained on a large amount of image data. It contains multiple parallel convolutional branch networks, each of which can process different types of features. Feature response calculation refers to inputting the input feature vector into the convolutional branch network, where the network performs operations such as convolution and pooling on the features to obtain the feature response result.
[0143] For example, the color sub-vector can be input into a parallel convolutional branch network in a pre-trained image recognition model that specifically processes color features. This branch network performs convolution operations on the color sub-vector to extract deep-level feature information of the color features and calculates the feature response result. Similarly, the structure sub-vector and background correlation sub-vector can be input into their respective parallel convolutional branch networks for feature response calculation.
[0144] In one embodiment, a pre-trained image recognition model is loaded using a deep learning framework such as TensorFlow or PyTorch. Color sub-vectors, structure sub-vectors, and background correlation sub-vectors are input into the input layers of the corresponding parallel convolutional branch networks in the model. The model automatically performs forward propagation computation to obtain the feature response results of each branch network. By processing different types of features in parallel, the model's computational power can be fully utilized, improving the efficiency and accuracy of feature extraction.
[0145] S43: Perform cross-branch feature weighting and fusion at the output of the parallel convolutional branch network to generate a fused feature representation.
[0146] In this embodiment, cross-branch feature weighted fusion refers to fusing different types of feature response results output by parallel convolutional branch networks and assigning different weights to each feature based on its importance. The fused feature representation is the fused feature vector, which integrates information from different types of features and can more comprehensively describe the characteristics of the kitchen waste to be classified.
[0147] For example, the parallel convolutional branch network outputs feature responses for color, structure, and background correlation sub-vectors, respectively. Different weights are assigned to the feature responses of each branch; for instance, the weight of the color feature response is 0.4, the weight of the structure feature response is 0.4, and the weight of the background correlation feature response is 0.2. These feature responses are then weighted and summed to obtain the fused feature representation.
[0148] In one embodiment, cross-branch feature weighted fusion can be implemented using fully connected layers or a custom fusion function. The output of the parallel convolutional branch network is used as the input to the fully connected layer, which then performs a weighted sum of the inputs according to pre-trained weights to obtain the fused feature representation. Alternatively, a custom fusion function can be defined, with weights manually set based on the importance of different features, to perform weighted fusion of the feature response results. Cross-branch feature weighted fusion can fully utilize information from different types of features, improving classification accuracy.
[0149] S44: Based on the fused feature representation, calculate the initial confidence score of each candidate waste category through a fully connected classification layer.
[0150] In this embodiment, the fully connected classification layer is a layer in the pre-trained image recognition model. It is used to map the fused feature representations onto each candidate waste category and calculate the initial confidence score for each category. The initial confidence score refers to the probability that each candidate waste category will be matched before the final classification.
[0151] For example, the fused feature representation is a vector that is input into a fully connected classification layer. The fully connected classification layer contains multiple neurons, each corresponding to a candidate waste category. Through linear transformations and activation functions in the fully connected layer, the fused feature representation is converted into a score for each candidate waste category. Assuming candidate waste categories include food scraps, fruit and vegetable peels, expired food, and other biodegradable organic matter, the fully connected layer calculates a score for each category and normalizes these scores to obtain an initial confidence score.
[0152] In one embodiment, this process is implemented using a fully connected layer function from a deep learning framework. The fused feature representation is input into the fully connected layer, and the output dimension of the fully connected layer is set to the number of candidate garbage categories. The fully connected layer automatically performs linear transformations and activation function calculations to obtain a score for each candidate garbage category. These scores are then normalized using a softmax function, converting them into initial confidence scores.
[0153] S45: Normalize the initial confidence scores so that the sum of the initial confidence scores of all candidate garbage categories equals 1.
[0154] In this embodiment, the normalization process adjusts the initial confidence scores to ensure that the sum of the initial confidence scores for all candidate waste categories is 1. This converts the initial confidence scores into a probability distribution, facilitating subsequent classification decisions.
[0155] For example, the initial confidence scores for each candidate waste category calculated by the fully connected classification layer are 0.2, 0.3, 0.4, and 0.1, respectively. These scores are normalized using the softmax function. The softmax function exponentially calculates each score and then divides these exponent values by the sum of all exponent values to obtain the normalized initial confidence scores. Assume the normalized scores are 0.15, 0.25, 0.5, and 0.1, and their sum is 1.
[0156] In one embodiment, the softmax function in Python is used for normalization. The initial confidence score is used as input to the softmax function, which automatically calculates the normalized score. Normalization converts the initial confidence score into probability values, more intuitively representing the likelihood of each candidate waste category being selected, thus improving the accuracy of classification decisions.
[0157] S46: Use the normalized initial confidence score as the output of the image recognition model to obtain the initial confidence score corresponding to each candidate waste category.
[0158] In this embodiment, the output of the image recognition model refers to the final result obtained after a series of processing steps. Here, the normalized initial confidence score is used as the output. By outputting the normalized initial confidence score, the probability of each candidate waste category being matched can be clearly determined.
[0159] For example, after normalization, the initial confidence scores for each candidate waste category are 0.15, 0.25, 0.5, and 0.1, respectively. These scores are used as the output of the image recognition model. In subsequent processing, classification decisions can be made based on these scores to determine the final classification result.
[0160] In one embodiment, within the deep learning framework, the normalized initial confidence score is used as the output of the model's final layer. Thus, when multi-source heterogeneous feature vectors are input into the model, the model automatically outputs the initial confidence score corresponding to each candidate waste category. By outputting these scores, subsequent classification processing can be conveniently performed, improving the accuracy and efficiency of classification.
[0161] The pre-trained image recognition model in this application mainly comprises multiple parallel convolutional branch networks and fully connected classification layers. The parallel convolutional branch networks are used to process different types of features, including color sub-vectors, structure sub-vectors, and background correlation sub-vectors. Each parallel convolutional branch network consists of multiple convolutional layers, pooling layers, and activation functions.
[0162] Convolutional layers extract local features from the input features by sliding convolution kernels across the input features to generate feature maps. Pooling layers downsample the feature maps, reducing their size while preserving important feature information. Activation functions introduce non-linearity to enhance the model's expressive power; commonly used activation functions include ReLU.
[0163] The fully connected classification layer, located at the end of the model, maps the fused feature representations to each candidate garbage category and calculates the initial confidence score for each category. Each neuron in the fully connected layer is connected to all neurons in the previous layer, transforming the input features into a probability distribution through linear transformations and activation functions (such as the softmax function).
[0164] The training steps for the pre-trained image recognition model are as follows: First, data preparation is performed by collecting a large amount of image data of kitchen waste covering different categories (such as food scraps, fruit peels and vegetable leaves, expired food, and other biodegradable organic matter). Preprocessing operations are then performed on this image data, including image cropping, scaling, and normalization, to ensure the consistency and standardization of the image data. Simultaneously, each image is labeled with its corresponding waste category, forming a training dataset. Next, feature extraction is performed. Following the method described in the claims, feature extraction is performed on the images in the training dataset to generate multi-source heterogeneous feature vectors, which are then divided into three subsets: color sub-vectors, structure sub-vectors, and background correlation sub-vectors. Finally, model initialization is performed by randomly initializing the weights W and bias parameters b of the parallel convolutional branch network and the fully connected classification layer.
[0165] During training, feature vectors from the training dataset are input into the pre-trained image recognition model for forward propagation. Parallel convolutional branch networks calculate feature responses for the color, structure, and background correlation sub-vectors respectively, and perform cross-branch feature weighting and fusion at the output to generate a fused feature representation. The fully connected classification layer calculates the initial confidence score for each candidate waste category based on the fused feature representation. The loss function L between the initial confidence score ypred output by the model and the true label ytrue is calculated; the cross-entropy loss function is commonly used, and its formula is... , where n is the number of candidate garbage categories.
[0166] The gradient of the loss function with respect to the model parameters is calculated using the backpropagation algorithm. The weights and bias parameters of the model are updated based on the gradient. Taking the stochastic gradient descent (SGD) optimization algorithm as an example, the update formula is as follows: , where α is the learning rate. Repeat the above steps until the model's loss function converges or the preset number of training epochs is reached.
[0167] In one embodiment, reference Figure 7 Step S50 may include steps S51-S55, which will be described in detail below: S51: Obtain the highest confidence score of each task output in the most recent N historical classification tasks, and form a confidence score sequence.
[0168] In this embodiment, the most recent N historical classification tasks refer to the N classification tasks preceding the current classification task. The highest confidence score refers to the maximum initial confidence score among all candidate waste categories in each classification task. The confidence score sequence is a sequence composed of the highest confidence scores from the most recent N historical classification tasks.
[0169] For example, suppose N = 5, and the highest confidence scores of the last 5 historical classification tasks are 0.8, 0.7, 0.9, 0.75, and 0.85, respectively. Arrange these scores in order to form the confidence score sequence [0.8, 0.7, 0.9, 0.75, 0.85].
[0170] In one embodiment, a database can be used to store the results of each classification task, including the initial confidence scores for each candidate waste category. When performing the current classification task, the results of the N most recent historical classification tasks are retrieved from the database, the highest confidence score for each task is extracted, and these scores are stored in a list to form a confidence score sequence. By obtaining the confidence score sequence, the confidence distribution of historical classification tasks can be understood, providing data support for dynamically calculating the decision threshold.
[0171] S52: Perform sliding window statistical analysis on the confidence score sequence to calculate the mean and standard deviation of the confidence score within the current window.
[0172] In this embodiment, sliding window statistical analysis is a method for analyzing time series data. It involves sliding a fixed-size window across the data series to perform statistical analysis on the data within the window. The confidence mean refers to the average confidence score within the current window, reflecting the overall level of the confidence score within the window. The standard deviation is an indicator of data dispersion, representing the degree of dispersion of the current window's confidence score relative to the mean.
[0173] For example, given a confidence score sequence of [0.8, 0.7, 0.9, 0.75, 0.85], and assuming a sliding window size of 3, when the window slides to the first three scores [0.8, 0.7, 0.9], the mean of these three scores is calculated as (0.8 + 0.7 + 0.9) / 3 = 0.8, and the standard deviation is calculated using the corresponding formula.
[0174] S53: Calculate the judgment threshold for the current classification task based on the confidence mean and standard deviation and the container background identification information.
[0175] In this embodiment, the decision threshold is a critical value used to determine whether a candidate waste category is the final classification result. The mean and standard deviation of the confidence score reflect the confidence distribution of historical classification tasks, and the container background identification information represents the container background characteristics of the current classification task. By comprehensively considering these factors, a decision threshold more suitable for the current classification task can be calculated.
[0176] For example, with a confidence level of 0.8 and a standard deviation of 0.05, the container background identification information indicates a moderate level of interference in the container background. This can be determined using a preset calculation formula, such as threshold = confidence level / standard deviation. The standard deviation (k) is determined based on the degree of background interference in the container; assuming moderate interference, k = 1. The resulting judgment threshold is... .
[0177] In one embodiment, a linear or non-linear calculation formula can be used to calculate the decision threshold. Based on extensive experimental data, the relationship between the mean confidence score, standard deviation, container background identification information, and the decision threshold is determined. The current mean confidence score, standard deviation, and container background identification information are then substituted into the formula to calculate the decision threshold for the current classification task. In this way, the decision threshold can be dynamically adjusted based on historical data and the current container background conditions, improving classification accuracy.
[0178] S54: Iterate through the initial confidence scores of each candidate waste category and filter out candidate categories whose initial confidence scores are greater than or equal to the determination threshold.
[0179] In this embodiment, the traversal operation refers to sequentially accessing the initial confidence score of each candidate waste category. The filtering operation involves filtering candidate categories based on a decision threshold to find candidate categories whose initial confidence score is greater than or equal to the decision threshold.
[0180] For example, the initial confidence scores for each candidate waste category are 0.6, 0.7, 0.8, and 0.9, respectively, with a decision threshold of 0.75. Each score is checked sequentially, and since 0.8 and 0.9 are greater than or equal to 0.75, the corresponding candidate categories are selected.
[0181] In one embodiment, a loop statement in Python can be used to iterate through the initial confidence scores of each candidate waste category. A conditional statement is used to determine whether each score is greater than or equal to a decision threshold; if the condition is met, the corresponding candidate category is added to a list. This filtering operation narrows down the range of candidate categories, providing more accurate information for determining the final classification result.
[0182] S55: If there is a unique candidate category that meets the screening criteria, then that candidate category is determined as the final classification result; if there are multiple candidate categories that meet the screening criteria, then the candidate category with the highest initial confidence score is selected as the final classification result.
[0183] In this embodiment, the final classification result refers to the category of kitchen waste determined after screening and judgment. When only one candidate category meets the screening criteria, it means that the category best meets the classification requirements, and it is determined as the final classification result. When multiple candidate categories meet the screening criteria, selecting the candidate category with the highest initial confidence score can improve the accuracy of classification.
[0184] For example, if only one candidate category meets the criteria after screening, with an initial confidence score of 0.8, this candidate category is selected as the final classification result. If two candidate categories meet the criteria after screening, with initial confidence scores of 0.8 and 0.9 respectively, the candidate category with the initial confidence score of 0.9 is selected as the final classification result.
[0185] Accordingly, to better implement the above methods, this application also provides an intelligent kitchen waste sorting system based on AI image recognition. For example... Figure 8 As shown, the AI image recognition-based intelligent kitchen waste sorting system 80 includes:
[0186] The acquisition module 801 is used to acquire a multi-view visible light image sequence of kitchen waste to be classified, and simultaneously acquire the corresponding ambient light intensity value and container background identification information. The multi-view visible light image sequence includes continuous image frames from three fixed angles: front view, side view, and top view. The feature extraction module 802 is used to extract the original pixel distribution features based on the multi-view visible light image sequence, and generate an initial feature set by combining the ambient light intensity value and the container background identification information, wherein the initial feature set includes color channel mean vector, texture gradient direction histogram and background interference factor weight; The feature fusion module 803 is used to perform feature fusion processing on the initial feature set and remove low-confidence feature dimensions according to the training sample screening strategy to construct a multi-source heterogeneous feature vector, wherein the multi-source heterogeneous feature vector includes HSV color space component statistics, local structure similarity index and container background correlation coefficient. The confidence processing module 804 is used to input the multi-source heterogeneous feature vector into a pre-trained image recognition model to perform category matching calculation and obtain an initial confidence score corresponding to each candidate waste category, wherein each candidate waste category includes food scraps, fruit peels and vegetable leaves, expired food and other biodegradable organic matter. The classification determination module 805 is used to dynamically calculate the judgment threshold of the current classification task based on the confidence score distribution of the most recent N historical classification tasks, and determine the final classification result based on the judgment threshold and the initial confidence score of each candidate waste category, where N is a positive integer; The instruction generation module 806 is used to map the final classification result into kitchen waste sorting control instructions and send them to the execution mechanism to complete the physical sorting operation.
[0187] The implementation details of each module are provided in the preceding method embodiments and will not be repeated here. The technical effects achieved by each module and device are described in the foregoing method embodiments.
[0188] like Figure 9 As shown, this application embodiment also provides a computer device 90, which includes a processor 901 and a memory 902, wherein the memory 902 stores a computer program, and when the computer program is executed by the processor 901, the processor 901 performs the steps of any of the methods described above.
[0189] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application are still within the scope of this application.
Claims
1. A method for intelligent sorting of kitchen waste based on AI image recognition, characterized in that, include: Collect a multi-view visible light image sequence of kitchen waste to be sorted, and simultaneously acquire the corresponding ambient light intensity value and container background identification information. The multi-view visible light image sequence includes continuous image frames from three fixed angles: front view, side view, and top view. Based on the multi-view visible light image sequence, the original pixel distribution features are extracted, and an initial feature set is generated by combining the ambient light intensity value and the container background identification information. The initial feature set includes color channel mean vector, texture gradient direction histogram and background interference factor weight. The initial feature set is subjected to feature fusion processing, and low-confidence feature dimensions are removed according to the training sample screening strategy to construct a multi-source heterogeneous feature vector, wherein the multi-source heterogeneous feature vector includes HSV color space component statistics, local structure similarity index and container background correlation coefficient. The multi-source heterogeneous feature vectors are input into a pre-trained image recognition model for category matching calculation to obtain an initial confidence score corresponding to each candidate waste category, wherein each candidate waste category includes food scraps, fruit peels and vegetable leaves, expired food and other biodegradable organic matter. Based on the confidence score distribution of the most recent N historical classification tasks, the judgment threshold for the current classification task is dynamically calculated, and the final classification result is determined based on the judgment threshold and the initial confidence score of each candidate waste category, where N is a positive integer; The final classification result is mapped into a kitchen waste sorting control instruction and sent to the execution mechanism to complete the physical sorting operation.
2. The method according to claim 1, characterized in that, The process of collecting multi-view visible light image sequences of kitchen waste to be classified, and simultaneously acquiring corresponding ambient light intensity values and container background identification information, includes: The image acquisition device is triggered to take pictures of kitchen waste samples sequentially at preset angle positions, generating a sequence of image frames arranged in chronological order. The image frame sequence is labeled with corresponding viewpoint tags and stored in the cache area. The analog signal output by the ambient light sensor is read, converted into a digital light intensity value, and then the light intensity value is synchronously bound to the timestamp of the image frame sequence. Identify the graphic encoding information on the surface of the waste disposal container, parse out the unique container background identification information, and establish an index association between the container background identification information and the corresponding image frame sequence; Check if there are blurry or occluded frames in the image frame sequence. If so, re-trigger the image acquisition operation at the corresponding angle to ensure that at least one clear image is retained for each viewpoint. The system integrates clear image frames from all perspectives, synchronized illumination intensity values, and container background identification information to form a structured data packet.
3. The method according to claim 1 or 2, characterized in that, The step of extracting original pixel distribution features based on the multi-view visible light image sequence and generating an initial feature set by combining the ambient light intensity value and container background identification information includes: For each image frame from each viewpoint, the pixel mean and variance of the RGB three channels are calculated to generate a color channel mean vector. The color channel mean vector is then normalized and used as the basic color feature. Gradient operators are used to perform edge detection on image frames, and the distribution frequency of gradient magnitudes in each direction is statistically analyzed to construct a texture gradient direction histogram. The background interference level table is queried according to the container background identification information to obtain the corresponding background interference factor weight. The background interference factor weight is multiplied with the image frame region mask to generate a weighted interference map. The ambient light intensity value is mapped to the light compensation coefficient, and the color channel mean vector is linearly corrected to generate the light-corrected color features. The initial feature set is formed by merging the illumination-corrected color features, texture gradient direction histogram, and weighted interference map.
4. The method according to claim 3, characterized in that, The step of performing feature fusion processing on the initial feature set and removing low-confidence feature dimensions according to the training sample selection strategy to construct a multi-source heterogeneous feature vector includes: The color features are converted to the HSV color space and the statistical moments of hue, saturation and lightness are calculated to generate HSV color space component statistics. The HSV color space component statistics are then included in the feature pool. Calculate the structural similarity index between adjacent viewpoint image frames, take the minimum value as the local structural similarity index, and add the local structural similarity index to the feature pool. Based on the semantic consistency between the container background identifier information and the current image content, the container background correlation coefficient is calculated, and the container background correlation coefficient is written into the feature pool. Iterate through each feature dimension in the feature pool and compare its discrimination score with the corresponding category in the training set. If the discrimination score is lower than the preset baseline, it is marked as a low-confidence feature dimension. Remove all labeled low-confidence feature dimensions from the feature pool, and use the remaining features to form a multi-source heterogeneous feature vector.
5. The method according to claim 4, characterized in that, The multi-source heterogeneous feature vectors are input into a pre-trained image recognition model for category matching calculation to obtain initial confidence scores corresponding to each candidate waste category, including: The multi-source heterogeneous feature vectors are divided into three subsets: color sub-vectors, structure sub-vectors, and background correlation sub-vectors. The color sub-vector, structure sub-vector, and background association sub-vector are respectively input into the corresponding parallel convolutional branch network in the pre-trained image recognition model to calculate the feature response. Cross-branch feature weighting and fusion are performed at the output of the parallel convolutional branch network to generate a fused feature representation. Based on the fused feature representation, the initial confidence score of each candidate waste category is calculated through a fully connected classification layer; The initial confidence scores are normalized so that the sum of the initial confidence scores of all candidate garbage categories equals 1. The normalized initial confidence score is used as the output of the image recognition model to obtain the initial confidence score corresponding to each candidate waste category.
6. The method according to claim 5, characterized in that, Based on the confidence score distribution of the most recent N historical classification tasks, the decision threshold for the current classification task is dynamically calculated, and the final classification result is determined based on the decision threshold and the initial confidence score of each candidate waste category, including: Obtain the highest confidence score output for each of the most recent N historical classification tasks, and construct a confidence score sequence; Perform sliding window statistical analysis on the confidence score sequence to calculate the mean and standard deviation of the confidence score within the current window; The decision threshold for the current classification task is calculated based on the mean and standard deviation of the confidence level and the container background identification information. Iterate through the initial confidence scores of each candidate waste category and filter out candidate categories whose initial confidence scores are greater than or equal to the judgment threshold; If there is a unique candidate category that meets the screening criteria, then that candidate category is determined as the final classification result; if there are multiple candidate categories that meet the screening criteria, then the candidate category with the highest initial confidence score is selected as the final classification result.
7. The method according to claim 6, characterized in that, The step of querying a pre-stored background interference level table based on the container background identification information to obtain the corresponding background interference factor weights, and multiplying the background interference factor weights by the image frame region mask to generate a weighted interference map includes: Parse the material type code and color code in the container background identification information, match them with the predefined background interference level table entries, and read the corresponding interference level value; The interference level value is converted into a background interference factor weight through a non-linear mapping function to ensure that the weight value is within a preset range. Perform foreground segmentation on the image frame to generate a binarized region mask, while preserving the pixel markers of the main area of kitchen waste; Multiply the background interference factor weights element by element with the region mask to obtain the attenuation coefficient matrix that only applies to the background region; The attenuation coefficient matrix is superimposed on the pixel values of the original image frame to generate a weighted interference map.
8. The method according to claim 6, characterized in that, The calculation of the structural similarity index between adjacent viewpoint image frames, taking the minimum value as the local structural similarity index, and adding the local structural similarity index to the feature pool includes: Select frontal and side view image frames to form the first pair of viewpoint combinations, calculate the structural similarity index between the two, and record it as the first similarity value; Select side view and top view image frames to form a second pair of viewpoint combinations, calculate the structural similarity index between the two, and record it as the second similarity value; Select frontal and top-view image frames to form a third pair of viewpoints, calculate the structural similarity index between the two, and record it as the third similarity value; Compare the first, second, and third similarity values, and select the minimum value as the local structural similarity index; The local structural similarity index is written into the feature pool in floating-point format.
9. The method according to claim 6, characterized in that, The interference level values are converted into background interference factor weights using a non-linear mapping function, ensuring that the weight values are within a preset range, including: Set the upper and lower limits of the input range for the interference level value, and truncate values that exceed the range; An S-shaped function is used to normalize and map the truncated interference level values to generate preliminary weight values. Apply an upward slope constraint to the initial weight values to prevent drastic fluctuations in weights due to minor changes in level. Verify whether the initial weight value is within the closed interval of 0 to 1; if not, force it to be limited to the boundary value. Output the background interference factor weights that conform to the interval constraints.
10. A smart kitchen waste sorting system based on AI image recognition, characterized in that, The system includes: The acquisition module is used to acquire a multi-view visible light image sequence of kitchen waste to be classified, and simultaneously acquire the corresponding ambient light intensity value and container background identification information. The multi-view visible light image sequence includes continuous image frames from three fixed angles: front view, side view, and top view. The feature extraction module is used to extract the original pixel distribution features based on the multi-view visible light image sequence, and combine the ambient light intensity value and container background identification information to generate an initial feature set, wherein the initial feature set includes color channel mean vector, texture gradient direction histogram and background interference factor weight; The feature fusion module is used to perform feature fusion processing on the initial feature set and remove low-confidence feature dimensions according to the training sample screening strategy to construct a multi-source heterogeneous feature vector, wherein the multi-source heterogeneous feature vector includes HSV color space component statistics, local structure similarity index and container background correlation coefficient. The confidence processing module is used to input the multi-source heterogeneous feature vector into a pre-trained image recognition model to perform category matching calculation and obtain an initial confidence score corresponding to each candidate waste category, wherein each candidate waste category includes food scraps, fruit peels and vegetable leaves, expired food and other biodegradable organic matter. The classification determination module is used to dynamically calculate the judgment threshold of the current classification task based on the confidence score distribution of the most recent N historical classification tasks, and determine the final classification result based on the judgment threshold and the initial confidence score of each candidate waste category, where N is a positive integer; The instruction generation module is used to map the final classification result into kitchen waste sorting control instructions and send them to the execution mechanism to complete the physical sorting operation.
Citation Information
Patent Citations
Garbage classification and recognition system for smart city
CN120259783A
Garbage classification multistage verification method based on dynamic confidence threshold and storage medium
CN120823429A