Multi-model garbage identification classification method and system based on deep learning
Patent Information
- Application Number
- CN202611198667.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-08
- Publication Date
- 2026-09-22
AI Technical Summary
一方面,纯视觉模型在面对垃圾堆叠、物品变形、遮挡及透明材质等复杂场景时,识别准确率显著下降;另一方面,单一视觉模态仅能获取垃圾的表观信息,无法感知垃圾的重量、密度、材质成分等物理属性,导致对材质相近或外观相似但类别不同的垃圾极易产生误判
[0022]The aforementioned deep learning-based multi-model waste identification and classification method and system constructs a multi-source heterogeneous data input foundation by acquiring visual image data, weight sensor data, and infrared material reflectance spectrum data of the waste to be identified, overcoming the deficiency of insufficient information from a single visual modality. Visual image data is input into a convolutional network to extract global contour features of the waste and generate a visual feature map, capturing the overall shape and appearance structure of the waste. The visual feature map is segmented to obtain multiple local image regions, which are then input into an attention network to extract local features and mapped back to the visual feature map to generate enhanced visual feature vectors, improving the model's adaptability to complex visual scenes such as stacking, deformation, and occlusion. Physical property features of the waste are extracted based on weight sensor data and infrared material reflectance spectrum data to obtain physical property feature vectors, quantifying the dynamic weight characteristics and material spectral properties of the waste. Multi-modal feature modeling based on the enhanced visual feature vectors and physical property feature vectors yields a virtual waste feature model, deeply integrating visual appearance information with physical property information. The virtual waste feature model is input into a classification network to generate waste identification and classification results, enabling accurate classification by comprehensively integrating visual and physical multi-dimensional information.
Smart Images

Figure CN122796629A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent waste classification and identification, and in particular relates to a multi-model waste identification and classification method and system based on deep learning. Background Technology
[0002] With the development of artificial intelligence technology, deep learning-based image recognition technology has gradually been introduced into the field of waste sorting. This technology automatically extracts visual features such as shape, texture, and color of waste through convolutional neural networks, enabling automatic identification and classification of waste types, which has significantly improved efficiency compared to traditional manual sorting methods.
[0003] In traditional technologies, waste sorting mainly relies on manual sorting or automatic identification methods based on a single vision sensor. Manual sorting is not only inefficient and costly, but also poses risks to the working environment of workers; while identification methods based on a single vision sensor typically only utilize RGB image data and perform feature extraction and classification through deep learning models.
[0004] The aforementioned traditional methods or existing approaches have significant technical shortcomings. On the one hand, pure visual models experience a significant drop in recognition accuracy when faced with complex scenarios such as piled-up waste, deformed objects, occlusion, and transparent materials. On the other hand, a single visual modality can only acquire superficial information about the waste and cannot perceive its physical properties such as weight, density, and material composition, leading to a high likelihood of misclassification for waste with similar materials or appearances but different categories. Furthermore, existing fusion methods often employ simple post-fusion strategies, lacking in-depth collaborative modeling of visual features and physical property features, making it difficult to fully leverage the complementary advantages of multi-source heterogeneous data. Summary of the Invention
[0005] Therefore, it is necessary to provide a deep learning-based multi-model garbage identification and classification method and system that can solve the above problems.
[0006] Firstly, this application provides a deep learning-based multi-model garbage identification and classification method, including:
[0007] Acquire visual image data, weight sensor data, and infrared material reflectance spectrum data of the waste to be identified;
[0008] Visual image data is input into a convolutional network to extract global contour features of garbage and generate a visual feature map;
[0009] The visual feature map is segmented to obtain multiple local image regions. Each local image region is input into the attention network to extract local features, and the local features are mapped onto the visual feature map to generate an enhanced visual feature vector.
[0010] Based on weight sensor data and infrared material reflectance spectrum data, physical property features of waste are extracted to obtain physical property feature vectors;
[0011] Based on enhanced visual feature vectors and physical attribute feature vectors, multimodal feature modeling is performed to obtain a virtual waste feature model;
[0012] The virtual waste feature model is input into the classification network to generate waste identification and classification results.
[0013] Secondly, this application also provides a deep learning-based multi-model waste identification and classification system, including:
[0014] The waste data acquisition module is used to acquire visual image data, weight sensor data, and infrared material reflectance spectrum data of the waste to be identified;
[0015] The visual feature extraction module is used to input visual image data into the convolutional network, extract global contour features of garbage, and generate visual feature maps;
[0016] The local feature enhancement module is used to segment the visual feature map to obtain multiple local image regions. Each local image region is input into the attention network to extract local features and map the local features onto the visual feature map to generate an enhanced visual feature vector.
[0017] The physical property extraction module is used to extract physical property features of waste based on weight sensor data and infrared material reflectance spectrum data, and obtain physical property feature vectors.
[0018] The multimodal modeling module is used to perform multimodal feature modeling based on enhanced visual feature vectors and physical attribute feature vectors to obtain a virtual waste feature model;
[0019] The classification result output module is used to input the virtual waste feature model into the classification network to generate waste identification and classification results.
[0020] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described deep learning-based multi-model garbage identification and classification method.
[0021] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described deep learning-based multi-model garbage identification and classification method.
[0022] The aforementioned deep learning-based multi-model waste identification and classification method and system constructs a multi-source heterogeneous data input foundation by acquiring visual image data, weight sensor data, and infrared material reflectance spectrum data of the waste to be identified, overcoming the deficiency of insufficient information from a single visual modality. Visual image data is input into a convolutional network to extract global contour features of the waste and generate a visual feature map, capturing the overall shape and appearance structure of the waste. The visual feature map is segmented to obtain multiple local image regions, which are then input into an attention network to extract local features and mapped back to the visual feature map to generate enhanced visual feature vectors, improving the model's adaptability to complex visual scenes such as stacking, deformation, and occlusion. Physical property features of the waste are extracted based on weight sensor data and infrared material reflectance spectrum data to obtain physical property feature vectors, quantifying the dynamic weight characteristics and material spectral properties of the waste. Multi-modal feature modeling based on the enhanced visual feature vectors and physical property feature vectors yields a virtual waste feature model, deeply integrating visual appearance information with physical property information. The virtual waste feature model is input into a classification network to generate waste identification and classification results, enabling accurate classification by comprehensively integrating visual and physical multi-dimensional information. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart of the deep learning-based multi-model garbage identification and classification method of the present invention;
[0025] Figure 2 This diagram illustrates the steps involved in constructing a virtual waste feature model in an implementation of the deep learning-based multi-model waste identification and classification method of the present invention.
[0026] Figure 3 This is a structural diagram of the deep learning-based multi-model waste identification and classification system of the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0028] In one embodiment, such as Figure 1As shown, a deep learning-based multi-model waste identification and classification method is provided. This embodiment illustrates the application of this method to a sorting terminal. It is understood that this method can also be applied to a server, or to an architecture including a sorting terminal and a server, and implemented through the interaction between the terminal and the server. The terminal hardware of this application may include: an RGB visual acquisition camera, an infrared spectral sensor, a pressure and weight sensor, a sorting terminal and / or a server, and a sorting execution mechanism. Application scenarios can include: smart waste disposal cabinets in residential communities, automated sorting lines in factories, and centralized sanitation sorting stations. In the above scenarios, when there is a need for accurate classification due to issues such as waste stacking and obstruction, similar appearance but different materials, and difficulty in identifying transparent / deformed waste, the RGB visual acquisition camera, infrared spectral sensor, and pressure and weight sensor simultaneously acquire multi-source data and transmit it to the sorting terminal in real time. The sorting terminal generates enhanced visual feature vectors and physical attribute feature vectors, completes multi-modal virtual waste feature modeling and dual-branch classification network verification, obtains the classification results, and controls the sorting execution mechanism to complete waste diversion.
[0029] In this embodiment, the method includes the following steps:
[0030] S01, acquire visual image data, weight sensor data and infrared material reflectance spectrum data of the waste to be identified.
[0031] Optionally, visual image data refers to two-dimensional color image data acquired by an RGB visual acquisition camera when photographing the waste to be identified, including apparent visual information such as the shape, texture, and color of the waste. Weight sensor data refers to mass signal data of the waste to be identified acquired by a pressure weight sensor, reflecting the dynamic weight changes of the waste. Infrared material reflectance spectral data refers to the reflectance spectral data of the waste to be identified in the infrared band acquired by an infrared spectral sensor, such as a near-infrared spectrometer or a hyperspectral imager. Different materials have characteristic reflectance spectral fingerprints for infrared light, and infrared material reflectance spectral data can be used to identify the material composition of the waste.
[0032] Optionally, when the waste to be identified enters the identification area, the sorting terminal can trigger the RGB visual acquisition camera, pressure and weight sensor, and infrared spectral sensor to collect data. The RGB visual acquisition camera captures visual image data of the waste; the pressure and weight sensor collects the weight signal of the waste in real time and sorts it according to the data acquisition timestamp to generate a weight time series; the infrared spectral sensor collects the reflectance spectrum data of the waste in the infrared band. The visual image data, weight sensor data, and infrared material reflectance spectrum data serve as the input basis for subsequent multimodal feature extraction and fusion, overcoming the deficiency of insufficient information from a single visual modality.
[0033] S02, input the visual image data into the convolutional network, extract the global contour features of the garbage, and generate a visual feature map.
[0034] Optionally, a convolutional network is a feedforward neural network structure containing convolutional layers, pooling layers, and nonlinear activation layers, used to automatically extract spatially hierarchical features from local to global in the input data. In this application, it specifically refers to a dilated convolutional network pre-trained with junk samples. A visual feature map refers to a two-dimensional feature matrix generated by the intermediate or output layers of a convolutional network. Each spatial location corresponds to a local receptive field in the input visual image, encoding the distribution information of visual response intensity such as edges, textures, and shapes within that local region.
[0035] Optionally, the sorting terminal can use visual image data as input, feeding it into a pre-trained convolutional network for forward propagation computation. This convolutional network performs sliding window convolution operations on the image using multiple convolutional kernels of different scales, extracting local structural responses such as edges, corners, and textures of the waste. By progressively expanding the receptive field of view through stacked convolutional and pooling layers, the convolutional network can integrate local responses to form a global feature description of the overall boundary orientation, shape contour, and spatial layout of the waste. It outputs a two-dimensional feature tensor as a visual feature map, which retains the spatial coordinate correspondence of the original image, providing a structured visual representation basis for subsequent local region segmentation and multimodal fusion.
[0036] S03, segment the visual feature map to obtain multiple local image regions, input each local image region into the attention network, extract local features, and map the local features onto the visual feature map to generate an enhanced visual feature vector.
[0037] Optionally, an attention network is a neural network structure based on a self-attention mechanism, used to learn the long-range dependencies between any two locations in the input features and adaptively assign attention weights to highlight important task-related information. In this embodiment, it specifically refers to a multi-head self-attention network capable of feature aggregation of local image regions. A local image region refers to a set of sub-feature blocks with defined sizes and boundaries, divided according to spatial location from the visual feature map. An enhanced visual feature vector refers to a one-dimensional feature representation obtained by integrating the enhanced features of each local region according to the original spatial arrangement order.
[0038] Optionally, the sorting terminal can segment the visual feature map, dividing it into multiple non-overlapping or partially overlapping local image regions using methods such as sliding windows or grid partitioning. Each region corresponds to a local field of view in the original image. The sorting terminal can input each local image region into an attention network, which calculates the correlation weight between any two locations within the region, and aggregates the features at each location based on the weight to enhance discriminative local texture, edge, or structural responses, generating local enhanced features for each local image region.
[0039] Optionally, the sorting terminal can backfill the corresponding positions of the local enhanced features into the visual feature map by interpolation or direct mapping based on the original spatial coordinates of each local image region in the visual feature map, forming a region mapping feature map. The sorting terminal can then stitch all the region mapping feature maps together in spatial order to obtain an enhanced visual feature vector. This retains global spatial structure information while incorporating local detail discrimination information enhanced by the attention mechanism, improving its adaptability to complex visual scenes such as garbage stacking, deformation, and occlusion.
[0040] S04. Based on weight sensor data and infrared material reflectance spectrum data, the physical property features of the waste are extracted to obtain a physical property feature vector.
[0041] Optionally, physical property features refer to a set of quantitative parameters extracted from weight sensor data and infrared material reflectance spectrum data that can characterize the essential physical properties of waste. A physical property feature vector is a one-dimensional numerical array formed by arranging physical property features in a preset order.
[0042] In practice, the sorting terminal can perform time-series analysis on weight sensor data, sorting it by acquisition timestamp to obtain a weight time-series sequence and calculating the weight change rate. It identifies abrupt change points from the weight change rate sequence to determine the start and end times of the waste disposal process, and statistically obtains dynamic weight characteristic parameters within this time interval. Using the start and end times as boundaries, the sorting terminal can extract effective spectral segments from the infrared material reflectance spectrum data for the corresponding time interval, integrate them to obtain the total reflectance energy of each channel, and arrange them into a spectral energy sequence. Then, it statistically analyzes the spectral energy sequence to obtain the material spectral response characteristic parameters. The sorting terminal can concatenate the dynamic weight characteristic parameters with the material spectral response characteristic parameters to generate a physical property feature vector, enabling the system to quantify the dynamic weight characteristics and spectral properties of waste, thus overcoming the limitation of purely visual models in perceiving physical properties.
[0043] S05. Based on the enhanced visual feature vector and physical attribute feature vector, multimodal feature modeling is performed to obtain the virtual waste feature model.
[0044] Optionally, multimodal feature modeling refers to the process of establishing a unified multidimensional feature representation space by fusing or jointly embedding heterogeneous feature vectors from different perceptual modalities, such as visual and weight modalities. A virtual waste feature model refers to a structured multidimensional feature representation obtained after multimodal feature modeling, capable of comprehensively representing the visual appearance and physical attributes of the waste to be identified.
[0045] In implementation, the sorting terminal can align and fuse enhanced visual feature vectors and physical attribute feature vectors at the feature level. This involves spatially mapping the physical attribute feature vectors based on the spatial structure information encoded by the enhanced visual feature vectors, establishing a spatial correspondence between visual and physical modal features. The sorting terminal can use feature cascading or gating fusion mechanisms to interactively fuse visual features with corresponding physical features according to spatial location or semantic dimensions, generating a multi-dimensional joint feature representation that includes both visual morphological and physical attribute descriptions. The sorting terminal can then perform feature transformation or structural recombination using this multi-dimensional joint feature representation to form a virtual waste feature model. This virtual waste feature model retains the visual appearance structure of the waste while embedding dynamic weight characteristics and material spectral response characteristics, enabling subsequent sorting decisions to utilize both visual and physical information.
[0046] S06. Input the virtual waste feature model into the classification network to generate waste identification and classification results.
[0047] Optionally, a classification network is a supervised learning model that receives feature representations and outputs category predictions. It typically contains at least one fully connected layer and an output layer. The output layer uses a Softmax activation function to map the category scores output by the fully connected layer to a probability distribution between 0 and 1. The waste identification and classification result refers to the predicted label and its corresponding confidence score output by the classification network, indicating the category to which the waste to be identified belongs.
[0048] In practice, the sorting terminal can use a virtual waste feature model as input to the classification network. This virtual waste feature model integrates visual appearance information and physical attribute information to provide a multi-dimensional joint feature representation for classification decisions.
[0049] Optionally, during the forward propagation of the classification network, the virtual waste feature model undergoes linear transformation and nonlinear activation processing in a fully connected layer, progressively compressing the high-dimensional features to the dimension of the number of categories, obtaining initial scores for each preset waste category. The initial scores are then normalized into category probability vectors using the Softmax function of the output layer. Each element in the probability vector represents the probability confidence that the waste to be identified belongs to the corresponding category. The sorting terminal can determine the waste identification and classification result based on the category index corresponding to the highest probability value, and simultaneously output the category label and the corresponding confidence score. The waste identification and classification results can be used to guide the generation of subsequent sorting execution control instructions.
[0050] The aforementioned deep learning-based multi-model waste identification and classification method overcomes the deficiency of insufficient information from a single visual modality by acquiring visual image data, weight sensor data, and infrared material reflectance spectrum data of the waste to be identified. It inputs the visual image data into a convolutional network to extract global contour features of the waste and generate a visual feature map, capturing the overall shape and appearance structure of the waste. Segmenting the visual feature map to obtain multiple local image regions and inputting them into an attention network to extract local features, then mapping them back to the visual feature map to generate enhanced visual feature vectors, improves the model's adaptability to complex visual scenes such as stacking, deformation, and occlusion. It also utilizes weight sensor data and infrared material reflectance spectrum data. The physical property features of waste are extracted from the data to obtain physical property feature vectors, which quantify the dynamic characteristics of waste weight and the spectral characteristics of materials, thus compensating for the deficiency of pure visual models in being unable to perceive physical properties. Based on the enhanced visual feature vectors and physical property feature vectors, multimodal feature modeling is carried out to obtain a virtual waste feature model, which deeply integrates visual appearance information and physical property information, giving full play to the complementary advantages of multi-source heterogeneous data. The virtual waste feature model is input into the classification network to generate waste identification and classification results, which can accurately classify waste by comprehensively integrating visual and physical multidimensional information, avoiding misjudgment of waste with similar materials or appearances, and improving the accuracy and robustness of waste identification and classification in complex scenarios.
[0051] In one embodiment, the convolutional network includes a dilated convolutional layer, a gradient entropy modulation layer, a contour continuation layer of the same material, and a background suppression aggregation layer;
[0052] Visual image data is input into a convolutional network to extract global contour features of garbage and generate visual feature maps, including:
[0053] S11, input the visual image data into the dilated convolutional layer, and convolve the visual image data through multiple convolutional kernels pre-configured based on the optical properties of garbage material to obtain the convolutional feature map;
[0054] S12, based on the convolutional feature map, generates a material-related edge response map by calculating pixel gradients in multiple directions;
[0055] S13, input the material-associated edge response map into the gradient entropy modulation layer, extract the gradient direction of each pixel in the material-associated edge response map, generate the gradient direction matrix, and extract the response intensity of each pixel in the material-associated edge response map to generate the response intensity matrix.
[0056] S14. Based on the gradient direction matrix, the distribution entropy within multiple preset neighborhood windows is statistically analyzed to obtain the direction entropy matrix; according to the preset mapping relationship of the optical properties of the waste material, the direction entropy matrix is mapped to the modulation weight of each neighborhood window; according to the modulation weight, the material-related edge response map is weighted and modulated to obtain the modulation edge map.
[0057] S15, based on the response intensity matrix, threshold segmentation is performed on the modulation edge map, and pixels with response intensity greater than the preset response intensity threshold are marked as high response pixels. Based on the high response pixels, an edge attention region mask is generated.
[0058] S16, perform color space transformation on the convolutional feature map to generate a chroma channel feature map; input the edge attention region mask and the chroma channel feature map into the contour continuation layer of the same material, perform endpoint detection on the modulated edge map, and identify the set of breakpoint pixels; for each breakpoint pixel in the set of breakpoint pixels, take multiple neighborhood blocks of different sizes with the breakpoint pixel as the center, calculate the mean of multiple different color channels of the pixels in the neighborhood block in the color space, and obtain the breakpoint chroma feature vector;
[0059] S17. Based on the chromaticity feature vector of each breakpoint pixel, calculate the Euclidean distance and chromaticity distance between any two breakpoint pixels. When the Euclidean distance between two breakpoint pixels is less than a preset distance threshold, the chromaticity distance is less than a preset chromaticity distance threshold, and both breakpoint pixels are located within the edge attention region mask or the overlap rate between the connecting line and the edge attention region mask is greater than a preset overlap rate threshold, they are determined to be a connectable breakpoint pair. For each connectable breakpoint pair, interpolation filling is performed along the connecting line of the breakpoint pixels to generate a continuous contour feature map.
[0060] S18: Input the continuous contour feature map into the background suppression aggregation layer. Through connected component analysis, obtain the connected component set. Calculate the area, shape compactness, and edge intensity of each connected component in the connected component set, and remove connected components whose area value is less than a preset area threshold, shape compactness value is less than a preset shape compactness threshold, and average edge intensity value is less than a preset edge intensity threshold. Use the remaining connected components as the foreground connected component set. Perform feature aggregation on the foreground connected component set to generate a visual feature map.
[0061] For example, a convolutional network may specifically include dilated convolutional layers, gradient entropy modulation layers, contour continuation layers of the same material, and background suppression aggregation layers. Among them, dilated convolutional layers refer to convolutional layer structures that use convolutional kernels with a dilation rate greater than 1 for feature extraction, which expand the receptive field to capture global contextual information of garbage materials without increasing the number of parameters;
[0062] The gradient entropy modulation layer refers to a network layer that calculates modulation weights based on the information entropy of the gradient direction distribution in the local neighborhood and uses these modulation weights to weight and enhance the edge response map. It is used to distinguish between material edges and non-material edges by taking advantage of the gradient direction uncertainty caused by the difference in optical reflection characteristics of the surface of garbage material.
[0063] The same material contour continuation layer refers to a network layer that pairs and interpolates the detected edge breakpoints according to the color consistency criterion in the color space. It is used to restore the contour break caused by uneven lighting or weak texture areas and maintain the contour integrity of the same waste material area.
[0064] Background suppression aggregation layer refers to a network layer that uses connected component analysis and combines geometric and photometric features such as area, shape compactness, and edge strength to filter out non-garbage background areas and aggregate features of the retained foreground areas.
[0065] Distribution entropy refers to the information entropy value calculated based on the statistical probability of the direction interval to which each pixel gradient direction belongs within a preset neighborhood window. It is used to quantify the degree of disorder in the gradient direction distribution within the window. Modulation weight can be obtained based on the direction entropy through a preset negative correlation mapping relationship. It is used to attenuate or maintain the response intensity of the corresponding pixel within the window in the material-related edge response map. The breakpoint chromaticity feature vector is a multi-dimensional vector formed by concatenating the mean values of multiple color channels in the color space in multiple neighborhood blocks of different sizes with the breakpoint pixel as the center. It is used to characterize the color distribution characteristics around the breakpoint. Connectable breakpoint pairs refer to a pair of breakpoints that satisfy the spatial proximity, chromaticity similarity, and spatial position constraints relative to the edge attention region mask.
[0066] Optionally, in S11, the sorting terminal can input visual image data into a dilated convolutional layer. This dilated convolutional layer contains multiple convolutional kernels pre-configured based on the optical characteristics of the waste material, such as its reflectivity curve and absorption peak position in the visible light band. Each convolutional kernel has a different dilation rate and orientation selectivity. Using these convolutional kernels, a sliding window convolution operation is performed on the spatial dimension of the visual image data. The sum of the pointwise product of the convolutional kernel weight and the corresponding image pixel value at each spatial location is calculated to generate a convolutional feature map containing multi-channel deep semantic information.
[0067] Optionally, the dilated convolutional layer uses the ResNet-50 skeleton and consists of 5 stages. Stages 3 to 5 use dilated convolutions with dilation rates of 2, 4, and 2, respectively, to expand the receptive field. After pre-training on ImageNet, it is fine-tuned on garbage image data.
[0068] Optionally, in S12, the sorting terminal can calculate the gray-level change rate (i.e., pixel gradient) of each pixel in the convolutional feature map in the horizontal, vertical, and diagonal directions using the Sobel operator or the Sher operator, and then weight and combine the gradient magnitudes in each direction to generate a material-related edge response map that can highlight the edge response intensity of the waste material.
[0069] Optionally, in S13, the sorting terminal can input the material-associated edge response map into the gradient entropy modulation layer, extract the gradient direction angle of each pixel in the material-associated edge response map, quantize each direction angle into multiple preset direction intervals to construct a gradient direction matrix, and extract the gradient magnitude corresponding to each pixel as the response intensity, and arrange them according to the pixel position to construct a response intensity matrix.
[0070] Optionally, in S14, the sorting terminal can define multiple preset neighborhood windows of different scales for each pixel position based on the gradient direction matrix, count the frequency of pixels whose gradient direction falls into each preset direction interval within each neighborhood window and convert it into probability, calculate the information entropy value of each window to obtain the direction entropy matrix, and based on the optical characteristics of the waste material, i.e., the high direction entropy value of the specular material and the low direction entropy value of the diffuse material, pre-determine the mapping relationship, map each entropy value in the direction entropy matrix to the modulation weight corresponding to the neighborhood window, multiply the modulation weight point by point with the response intensity of each pixel in the corresponding window in the material-related edge response map to complete the weighted modulation and output the modulation edge map.
[0071] Optionally, the directional entropy can be mapped to modulation weights using a 2-layer fully connected network (64-dimensional hidden layers). This mapping network is trained jointly with the entire model, and in addition to the main classification loss, an additional gradient consistency loss (the gradient magnitude (MSE) between the predicted edge map and the ground truth edge map) is added to improve edge localization accuracy.
[0072] Optionally, for each preset neighborhood window The size is The gradient direction angle of all pixels within the range is quantized into B uniform intervals, where B = 8 or 16, and the pixel frequency in each interval is counted. And calculate the probability. Then the directional entropy of the window Calculated according to Shannon's entropy formula:
[0073] ;
[0074] in To prevent small constants from being undefined in the logarithm. All windows... Arranging them according to spatial location yields the direction entropy matrix. .
[0075] Based on the optical properties of waste materials, such as specular reflection producing high randomness in gradient direction and diffuse reflection producing low randomness, a predefined monotonically decreasing mapping function is used to convert the direction entropy value into modulation weights. This embodiment employs a piecewise linear mapping:
[0076] ;
[0077] in , The preset threshold is determined through statistical calibration using material samples. The modulation weights for each window are then... Assigning weights to all pixels within the window yields the full-image modulation weight matrix. Then associate the material with the edge response map. and Pixel-by-pixel multiplication yields the modulation edge map. , This represents the Hadamard product.
[0078] Optionally, in S15, the sorting terminal can perform global threshold segmentation on the modulation edge map based on the response intensity matrix, marking pixels with response intensity values greater than a preset response intensity threshold as high-response pixels and the remaining pixels as low-response pixels. A binary edge attention region mask is generated based on the spatial distribution of the high-response pixels, where the high-response pixel positions are set to 1 and the low-response pixel positions are set to 0. Optionally, the formula for generating the binary mask is: ,in Using an Otsu adaptive threshold can further improve robustness.
[0079] Optionally, in S16, the sorting terminal can perform a conversion from RGB color space to CIELAB or HSV color space on the convolutional feature map, extract the chroma components to generate a chroma channel feature map, input the edge attention region mask and the chroma channel feature map to the contour continuation layer of the same material, perform endpoint detection on the modulation edge map, detect the start and end pixel positions of the connected path in the modulation edge map, identify all breakpoint pixels and classify them into a breakpoint pixel set, for each breakpoint pixel in the breakpoint pixel set, take a square neighborhood block with multiple odd-numbered pixel sizes as the center of the breakpoint pixel, calculate the pixel mean of each color channel in the chroma channel feature map in each neighborhood block, and concatenate all neighborhood blocks and the mean of all color channels in a predetermined order to form the breakpoint chroma feature vector of the breakpoint pixel.
[0080] Optionally, in S17, the sorting terminal can calculate the Euclidean distance between any two breakpoint pixels and the chromatic distance between the feature vectors based on the breakpoint chromatic feature vectors of each breakpoint pixel. When the Euclidean distance between two breakpoint pixels is less than a preset distance threshold, the chromatic distance is less than a preset chromatic distance threshold, and both breakpoint pixels are located within the area where the edge attention region mask value is 1, or the proportion of pixels covered by the line connecting the two breakpoint pixels that belong to the edge attention region mask value is 1 (i.e., the overlap rate) is greater than a preset overlap rate threshold, the sorting terminal can determine the two breakpoint pixels as a connectable breakpoint pair. For each connectable breakpoint pair, the sorting terminal can perform linear interpolation filling at the integer pixel position of the line connecting the two breakpoint pixels based on the response intensity of the adjacent existing pixels, connecting the broken contours into a complete contour, and generating a continuous contour feature map.
[0081] Optionally, in S18, the sorting terminal can input the continuous contour feature map into the background suppression aggregation layer, perform connected component analysis on the continuous contour feature map, mark all four-connected or eight-connected pixel connected regions and classify them into a connected component set, calculate the total number of pixels contained in each connected component in the connected component set as the area, calculate the ratio of the square of the perimeter of the bounding rectangle of the connected component to the area as the shape compactness, and calculate the average response intensity of the pixels on the boundary of the connected component in the modulation edge map as the edge intensity. Connected components with an area value less than a preset area threshold, a shape compactness value less than a preset shape compactness threshold, and an average edge intensity less than a preset edge intensity threshold are identified as background clutter regions and removed from the connected component set. The connected components that are not removed are retained as the foreground connected component set, generating a visual feature map for subsequent classification or localization tasks.
[0082] In one embodiment, the visual feature map is segmented to obtain multiple local image regions. Each local image region is input into an attention network to extract local features, and the local features are mapped onto the visual feature map to generate an enhanced visual feature vector, including:
[0083] S21, by using a multi-scale sliding window, the visual feature map is traversed according to a preset step size, the variance of pixel response intensity within the coverage area of each sliding window is calculated, and the coverage area of the sliding window with a pixel response intensity variance value greater than the variance threshold is selected to obtain multiple local image regions.
[0084] S22, each local image region is input into the attention network. Through the self-attention mechanism of the attention network, each local image region is linearly mapped to generate the query matrix, key matrix and value matrix of each local image region. Based on the query matrix and key matrix, the correlation weight between any two pixel positions in each local image region is calculated. Based on the correlation weight, the value matrix of each local image region is weighted and summed to generate the local aggregated feature vector of each local image region.
[0085] S23, based on the position coordinates of each local image region in the visual feature map, the local aggregated feature vector of each local image region is mapped back to the original image region corresponding to the position coordinates by bilinear interpolation, to obtain the region mapping feature map.
[0086] S24. For the region mapping feature maps corresponding to each position coordinate, concatenate them according to the arrangement order of the position coordinates in the visual feature map to generate an enhanced visual feature vector.
[0087] Specifically, the pixel response intensity variance refers to the degree of dispersion of the response intensity values of all pixels on the visual feature map within the sliding window coverage area relative to the mean of that area, used to characterize the intensity of fluctuations in texture or edge response within that area; the query matrix, key matrix, and value matrix are three feature matrices obtained by the attention network projecting the local image region through three sets of independent linear transformation weights, where the query matrix represents the feature description to be matched at each pixel position, the key matrix represents the feature description that can be matched at each pixel position, and the value matrix stores the original feature information to be aggregated at each pixel position; the relevance weight refers to the weight coefficients obtained by performing a dot product operation between the query matrix and the key matrix and then scaling and normalizing with Softmax, used to measure the feature similarity or dependency strength between any two pixel positions within the local image region; the local aggregated feature vector is the feature vector obtained by weighted summation of the value matrices at each position according to the relevance weight, used to characterize the compact feature description of the local image region after being enhanced by the self-attention mechanism; bilinear interpolation is a mapping method that uses the known feature values of the four nearest neighbor positions around the target position to perform linear interpolation in two directions to calculate the feature value of the target position.
[0088] Optionally, in S21, the sorting terminal can traverse and sample the spatial plane of the visual feature map using sliding windows of various sizes with a preset pixel step size. Within each sliding window coverage area, the sorting terminal can calculate the mean response intensity of all pixels in that area, calculate the sum of squares of the differences between the response intensity of each pixel and the mean, and then divide by the total number of pixels in the area to obtain the variance value of the pixel response intensity of that area. The sorting terminal can compare the calculated variance value with a preset variance threshold, retaining only the sliding window coverage areas with variance values greater than the variance threshold as local image regions, and discarding flat, textureless background areas, thus concentrating the computational resources of the attention network on local regions with strong discriminative power.
[0089] Optionally, in S22, the sorting terminal can input each local image region into the attention network separately. Through three independent linear mapping matrices, the feature matrices of the local image regions are linearly transformed to generate corresponding query matrices, key matrices, and value matrices. The sorting terminal can perform matrix multiplication on the transpose of the query matrix and the key matrix to obtain a similarity score matrix between any two pixel positions within the local image region. Each element in the similarity score matrix is divided by a preset scaling factor, and then Softmax normalization is performed along the row containing each query position to obtain a relevance weight matrix. The sorting terminal can perform matrix multiplication on the relevance weight matrix and the value matrix, and then perform a weighted summation of the features at each position in the value matrix to obtain the local aggregated feature vector for that local image region. This process, through a self-attention mechanism, establishes a long-range dependency between any two pixels within the region, enabling the local aggregated feature vector to integrate information from the entire region.
[0090] Optionally, the attention network employs an 8-head self-attention architecture, stacking 6 Transformer encoder layers. Cross-entropy loss is used during training, with the optimization objective being to ensure that the aggregated features of each local region correctly reflect its semantic category.
[0091] Optionally, for each local image region's feature tensor It can be achieved through three sets of linear projection matrices (d is the projection dimension) Calculate the query matrix respectively Key matrix Value matrix and reshape them into ,in This represents the total number of pixels within the region. Calculate the positions of any two pixels. Correlation weights between for:
[0092] ;in To query the i-th row of the matrix, Let j be the j-th row of the key matrix. Based on all... Constructing the correlation weight matrix Weighted aggregation of the value matrix yields locally aggregated eigenvectors, which are then flattened as follows: : ;Will According to the original space dimensions The features are rearranged to form an aggregated feature map of the local image region.
[0093] Optionally, in S23, the sorting terminal can determine the spatial location corresponding to each local aggregated feature vector based on the original position coordinates of each local image region in the visual feature map. For each local aggregated feature vector, the sorting terminal uses a bilinear interpolation method to interpolate and expand the feature sample value in the spatial dimension, taking it as the discrete feature sample value of the center position of that local image region. It then calculates the feature value of each pixel position after interpolation, obtaining a region mapping feature map with the same size as the local image region. This process continuously maps the discrete local aggregated feature vectors back to the original image region, maintaining the correspondence between features and spatial locations.
[0094] Optionally, in S24, the sorting terminal can concatenate all region mapping feature maps end-to-end along the feature channel dimension according to the original spatial arrangement order of each local image region in the visual feature map, forming a one-dimensional enhanced visual feature vector. This enhanced visual feature vector retains the global spatial structure information of the visual feature map, while incorporating local detail discrimination information enhanced by the attention mechanism.
[0095] In one embodiment, based on weight sensor data and infrared material reflectance spectrum data, physical property features of the waste are extracted to obtain a physical property feature vector, including:
[0096] S31, sort the weight sensor data according to the data acquisition timestamp to obtain the weight time series sequence, and calculate the weight change rate sequence based on the weight time series sequence;
[0097] S32, identify the set of abrupt change points in the weight change rate sequence that exceed the preset change rate threshold, and based on the set of abrupt change points, distinguish the start time and end time of waste disposal; using the start time and end time as boundaries, extract the spectral data of the corresponding time interval in the infrared material reflectance spectrum data to obtain the effective spectral segment;
[0098] S33, Integrate the effective spectral segment according to the preset wavelength channel to obtain the total reflection energy of each wavelength channel, and arrange the total reflection energy in wavelength order to generate a spectral energy sequence;
[0099] S34: For all rate of change values in the weight change rate sequence within the interval from the start time to the end time, calculate the maximum value, mean, standard deviation, and cumulative change to obtain the weight statistical feature vector; For the spectral energy sequence, calculate the peak position, half-width at half-maximum, and energy ratio of the peak value to the mean value of the entire band for each wavelength channel to obtain the spectral statistical feature vector.
[0100] S35, splices the weight statistical feature vector and the spectral statistical feature vector to generate the physical property feature vector.
[0101] For example, the weight change rate sequence refers to the sequence of weight change rates between adjacent time points obtained by performing a difference operation on the weight time series, used to reflect the transient fluctuation characteristics of weight during waste disposal. The set of abrupt change points refers to the set of discrete time points in the weight change rate sequence where the change rate value exceeds a preset change rate threshold, used to identify the start and end boundaries of waste disposal behavior.
[0102] An effective spectral segment refers to a subset of spectral data extracted from the complete infrared material reflectance spectral data that corresponds to the actual time interval of the waste disposal process. It is used to eliminate invalid spectral interference outside the disposal period. A spectral energy sequence is a sequence formed by arranging the total reflectance energy of each channel after integrating the effective spectral segment according to a preset wavelength channel, in wavelength order. It is used to characterize the reflectance energy distribution characteristics of waste materials in the infrared band.
[0103] The weight statistical feature vector is a multi-dimensional feature vector formed by statistically analyzing the weight change rate sequence from the start to the end of waste disposal, including its maximum value, mean, standard deviation, and cumulative change. It is used to quantify the dynamic weight characteristics of waste. The spectral statistical feature vector is a multi-dimensional feature vector formed by calculating the peak position, full width at half maximum (FWHM), and the energy ratio of the peak value to the mean across the entire spectral band from the spectral energy sequence. It is used to quantify the spectral response characteristics of waste materials. The physical property feature vector is a one-dimensional numerical array obtained by concatenating the weight statistical feature vector and the spectral statistical feature vector. It is used to comprehensively characterize the physical properties of waste.
[0104] Optionally, in S31, the sorting terminal can sort the weight sensor data collected by the pressure weight sensor in ascending order according to the data collection timestamps corresponding to each data point, forming a weight time series sequence. Based on the weight time series sequence, the sorting terminal can perform a difference operation on the weight values at adjacent time points, that is, calculate the difference between the weight value at the later time point and the weight value at the previous time point, divide it by the corresponding time interval, obtain the weight change rate at each time point, and arrange the weight change rates at each time point in chronological order to generate a weight change rate sequence.
[0105] Optionally, in S32, the sorting terminal can compare each weight change rate value in the weight change rate sequence with a preset change rate threshold, and identify all time points with change rate values greater than the preset change rate threshold as mutation points and classify them into a mutation point set. The sorting terminal can perform cluster analysis on the time points in the mutation point set, dividing the temporally continuous mutation points into the same mutation cluster, taking the timestamp corresponding to the earliest mutation point in each mutation cluster as the start time of waste disposal, and taking the timestamp corresponding to the latest mutation point in each mutation cluster as the end time of waste disposal. The sorting terminal can use the start time and end time as time boundaries to extract spectral data located between the start time and the end time from the complete time series of infrared material reflectance spectral data, and take the extracted data as a valid spectral segment to eliminate environmental background spectral interference during non-waste disposal periods.
[0106] Optionally, in S33, the sorting terminal can perform integration calculations on the effective spectral segment according to multiple preset wavelength channels, and sum up all the reflection spectral intensity values of each wavelength channel within the effective spectral segment time length to obtain the total reflection energy of that wavelength channel. The sorting terminal can arrange the total reflection energy of all wavelength channels in ascending order of wavelength to form a spectral energy sequence.
[0107] Optionally, in S34, the sorting terminal can extract all rate of change values within the interval from the start time to the end time from the weight change rate sequence. For all rate of change values within the interval, it calculates their arithmetic maximum, arithmetic mean, sample standard deviation, and the cumulative change from the start time to the end time (i.e., the sum of the products of each rate of change value and the time interval). These four statistics are then arranged in a preset order to obtain a weight statistical feature vector. The sorting terminal can also detect the wavelength position corresponding to the wavelength channel with the highest energy value in the spectral energy sequence as the peak position. It can detect the two wavelength positions corresponding to the points where the energy values on both sides of the peak position drop to half the peak value and calculate the difference between them as the half-peak width. It can calculate the ratio of the channel energy value at the peak position to the arithmetic mean of the energy values of all channels in the spectral energy sequence as the energy ratio. These three statistics are then arranged in a preset order to obtain a spectral statistical feature vector.
[0108] Optionally, the weight change rate sequence can be set at the initial time. Until the end time The sampling points inside are Its arithmetic maximum value mean Sample standard deviation Cumulative change , The sampling interval is [specified]. The weight statistical feature vector is obtained by arranging them in order.
[0109] For spectral energy sequences (L is the number of wavelength channels), maximum detection value and its corresponding channel index The peak position is ( (Channel spacing). Half-peak width To find the first time the energy is lower on both sides of the peak. Channel index Calculate the wavelength difference Full-band mean Energy ratio .Will The spectral statistical feature vectors are obtained by arranging them in order.
[0110] Optionally, in S35, the sorting terminal can concatenate the weight statistical feature vector and the spectral statistical feature vector end to end in the feature dimension to form a physical attribute feature vector. This physical attribute feature vector contains both the weight dynamic characteristics of the waste and the material spectral response characteristics, providing a quantitative input of the physical dimension for subsequent multimodal feature modeling.
[0111] like Figure 2 As shown, in one embodiment, multimodal feature modeling is performed based on enhanced visual feature vectors and physical attribute feature vectors to obtain a virtual waste feature model, including:
[0112] S41, based on enhanced visual feature vectors, identifies several individual waste areas through instance segmentation, and generates virtual waste volume data based on the volume estimation of each individual waste area;
[0113] S42, based on all waste individual regions and the pre-acquired waste individual material prior data, performs blind source separation on the spectral energy sequence in the physical property feature vector to obtain multiple individual spectral sequences, and maps each individual spectral sequence to the corresponding waste individual region to generate waste material distribution data;
[0114] S43, Based on the waste material distribution data and the preset waste material density benchmark, set the material density coefficient;
[0115] S44: Based on the virtual waste volume data of all waste individual areas, calculate the volume ratio, and according to the volume ratio and material density coefficient, map the weight statistical feature vector in the physical property feature vector to each waste individual area to generate waste density distribution data.
[0116] S45 integrates the virtual waste volume data, waste material distribution data, and waste density distribution data of all individual waste areas to obtain a virtual waste feature model.
[0117] Specifically, a virtual waste feature model refers to a structured, multi-dimensional feature representation obtained through multimodal feature modeling, capable of comprehensively characterizing the visual appearance and physical properties of the waste to be identified. Instance segmentation is a technique that classifies each pixel in a visual feature map into different object instances, used to distinguish regions in an image belonging to different waste individuals. Virtual waste volume data refers to the estimated volume of a single waste unit in three-dimensional space based on information such as the two-dimensional projected area and visual feature intensity of the waste unit region. Prior material data for waste units refers to pre-acquired benchmark data on the infrared reflectance spectral characteristics of various waste materials. Blind source separation is a signal processing technique that separates the independent spectral sequences corresponding to each waste unit from a mixed spectral energy sequence when individual spectral information of each unit is lacking.
[0118] Waste material distribution data refers to the spatial distribution information describing the material type of each waste unit after mapping the spectral sequence of each waste unit to the corresponding waste unit region. Waste density distribution data refers to the spatial distribution data describing the density information of each waste unit after allocating the weight statistical feature vector to each waste unit region according to the volume proportion and material density coefficient.
[0119] Optionally, in S41, the sorting terminal can classify pixels in the visual feature map pixel by pixel based on the enhanced visual feature vector using an instance segmentation algorithm, clustering pixels belonging to the same waste individual into the same waste unit region, thus identifying several waste unit regions. For each waste unit region, the sorting terminal can calculate the estimated volume of the waste unit region using a volume estimation formula based on parameters such as the projected physical area of the region in the visual feature map, the arithmetic mean of the pixel values of all pixels in the region on the visual feature map, the standard deviation of the pixel values in the visual feature map of the region, the global weighted average of all waste unit regions, and the maximum value of the weight change rate sequence, thereby generating virtual waste volume data.
[0120] Optionally, instance segmentation is based on Mask R-CNN, with ResNet-101 and FPN as the backbone. The total loss function includes the binary cross-entropy and Smooth L1 regression loss of RPN, the classification cross-entropy and regression loss of the detection branch, and the pixel-by-pixel binary cross-entropy loss of the mask branch.
[0121] Optionally, in S42, the sorting terminal can apply a blind source separation algorithm to the spectral energy sequence in the physical property feature vector based on all waste unit areas and the pre-acquired waste unit material prior data. Under the condition of knowing the mixed spectrum and the prior spectra of various materials, the mixed spectral energy sequence is decomposed into multiple independent individual spectral sequences through methods such as independent component analysis or non-negative matrix factorization. Each individual spectral sequence is then mapped to the corresponding waste unit area according to its spatial location, generating waste material distribution data describing the material type of each waste unit.
[0122] Optionally, nonnegative matrix factorization (NMF) is used for spectral unmixing. The optimization objective is to minimize the reconstruction error and add row sparsity regularization of the abundance matrix. The multiplication update rule is used to iteratively converge and separate the spectral vectors of each monomer.
[0123] Optionally, let the mixed spectral energy sequence be... (L is the number of wavelength channels), given that the number of waste monomer regions is K, the goal is to separate the spectral sequences of each monomer. Based on the assumptions of the linear mixture model: ,in The mixing coefficient, To reduce noise. Utilizing the pre-acquired prior spectral matrix of the waste monomer material. (M represents known material types) is used as a dictionary, and non-negative sparse encoding is employed for solution:
[0124] ;in The coefficient matrix, The separated monomer spectral matrices are obtained. The spectral sequences of each monomer are obtained by iteratively solving the problem using the Alternating Directional Multiplier Method (ADMM). (Right now (the kth column). Assigned to the corresponding k-th garbage unit region, forming a material type probability distribution vector (which will...) (Matching with the standard spectra of each material using least squares) generates waste material distribution data.
[0125] Optionally, in S43, the sorting terminal can find the corresponding standard density value from the preset waste material density benchmark based on the material type of each waste unit area indicated by the waste material distribution data, and set a material density coefficient for each waste unit area.
[0126] Optionally, in S44, the sorting terminal can calculate the proportion of the volume of each waste unit area to the total volume of all waste unit areas based on the virtual waste volume data of all waste unit areas, obtain the volume proportion, and according to the volume proportion and material density coefficient, the total weight represented by the weight statistical feature vector in the physical attribute feature vector is weighted and allocated according to the volume proportion and material density coefficient of each area, and mapped to each waste unit area to generate waste density distribution data describing the density information of each waste unit.
[0127] Optionally, suppose there are k individual waste zones with virtual volumes of respectively. Volume ratio The material density coefficient of each individual region is... The total weight can be obtained by looking up the table in step S43. Take the cumulative change in the weight statistical eigenvector. (Alternatively, it could be the final value minus the initial value of the weight time series). The weight of each monomer is allocated by volume and density weighting:
[0128] The density estimate of the i-th monomer is then... All individual units The data is combined and mapped back to image coordinates according to spatial location to obtain garbage density distribution data, and each pixel location is assigned a density value to its respective unit.
[0129] Optionally, in S45, the sorting terminal can fuse the virtual waste volume data, waste material distribution data, and waste density distribution data of each waste unit area at the feature level to form a structured feature representation containing multi-dimensional information on the volume, material, and density of each waste unit area, which is then output as a virtual waste feature model.
[0130] In one embodiment, the volume estimation formula is:
[0131]
[0132] in, Let i be the estimated volume of the i-th waste unit region. Let be the projected physical area of the i-th waste unit region in the visual feature map. The preset standard height scale coefficient, Let be the arithmetic mean of the pixel values of all pixels within the i-th garbage cell region on the visual feature map. For all individual waste areas The global weighted average, Let be the standard deviation of the pixel values of the visual feature map within the i-th garbage unit region. This represents the maximum value of the weight change rate sequence. This is a reference value for the rate of weight change. , and α represents the preset weighting coefficients, where α is used to adjust the nonlinear contribution of average visual intensity to volume. Enhancement effects used to control texture undulations The correction force used to adjust the weight peak.
[0133] For example, the volume estimation formula is used to map two-dimensional visual features to three-dimensional space, thereby obtaining virtual volume data for each waste unit region, providing a quantitative basis for subsequent density distribution calculation and multimodal feature fusion. Estimated volume of waste unit region Its projected physical area Standard height scale factor The product of these factors determines the reference volume of the region in a perfectly flat state; it can be determined by the average visual intensity ratio. Nonlinear correction is performed using the α power, where Let be the mean response of the pixel in the i-th region on the visual feature map. The ratio is a global weighted average of the response mean of all individual regions. This ratio reflects the visual saliency of the region relative to the whole. The response intensity on the visual feature map is related to the light reflection energy of the waste surface. The reflection energy is positively correlated with the object thickness or bulk density. Therefore, introducing this ratio can enable visual intensity-driven adaptive adjustment of the baseline volume.
[0134] At the same time, the formula introduces a texture undulation enhancement term. ,in The standard deviation of pixel values within a region characterizes the intensity of texture and edge undulations. The rougher the surface texture or the more irregular the stacking pattern, the more discrete the spatial distribution of visual response intensity, and the larger its equivalent volume. Therefore, the ratio of the standard deviation to the mean, combined with the weighting coefficient β, achieves positive volume compensation. The formula introduces a weight peak correction term. ,in This represents the maximum value of the weight change rate sequence. As a preset reference value, the peak value of the weight change rate during waste disposal reflects the impact kinetic energy and compaction degree of the waste when it falls onto the sensor. A higher peak value indicates a denser waste accumulation or a larger volume. By correcting the peak value in the volume estimation, the estimation results can be made to better reflect the actual physical state. The weighting coefficients α, β, and γ are all pre-calibrated based on experimental data and are used to adjust the contribution of visual intensity, texture undulation, and weight peak value to the volume estimation, respectively, so that the estimation results have good generalization and accuracy under different types of waste scenarios.
[0135] In one embodiment, the classification network includes a main classification branch and a physical verification branch;
[0136] The virtual waste feature model is input into the classification network to generate waste identification and classification results, including:
[0137] S51, the virtual waste feature model is input into the main classification branch, and the initial category probability vector is generated through mapping by the fully connected layer;
[0138] S52, input the waste material distribution data and waste density distribution data in the virtual waste feature model into the physical verification branch, perform cosine similarity matching with the pre-configured material-density correlation benchmark matrix, and generate class-by-class confidence correction coefficients;
[0139] S53. Based on the class-by-class confidence correction coefficient, the initial category probability vector is corrected class by class to obtain the waste category probability vector; based on the waste category probability vector and the virtual waste feature model, the waste identification and classification results are generated.
[0140] Specifically, the main classification branch refers to the forward propagation path in the classification network used to perform semantic category mapping on the virtual waste feature model to generate initial category probability predictions. The physical verification branch refers to the parallel processing path in the classification network used to correct the confidence of the prediction results of the main classification branch based on the material-density physical correlation constraints. The initial category probability vector refers to the probability distribution vector corresponding to each preset waste category obtained by the main classification branch through a linear transformation of the virtual waste feature model via a fully connected layer and then normalized by Softmax. The material-density correlation benchmark matrix refers to a reference matrix pre-constructed based on the physical correspondence between the material type and density value of various standard waste samples, where each row corresponds to a waste category, and each column corresponds to the standard correlation quantization value of the material spectral features and density features under that category.
[0141] The class-by-class confidence correction coefficient refers to the confidence adjustment factor calculated for each waste category after matching the waste material distribution data and waste density distribution data in the virtual waste feature model with the material-density correlation benchmark matrix using cosine similarity. This factor is used to differentiate the probability values corresponding to each category in the initial category probability vector. The waste category probability vector is the final probability distribution vector that incorporates physical attribute verification information, obtained after class-by-class correction of the initial category probability vector based on the class-by-class confidence correction coefficient.
[0142] Optionally, the main classification branch is a 3-layer fully connected network with 1024 and 512 hidden layers, outputting initial probabilities. The physical verification branch combines material and density features and calculates cosine similarity with the baseline matrix, generating class-by-class correction coefficients via temperature softmax. The total loss consists of cross-entropy loss and physical consistency loss (negative logarithmic correction coefficient).
[0143] Optionally, in S51, the sorting terminal can use the virtual waste feature model as input features and pass it to the main classification branch. The main classification branch contains at least one fully connected layer. The sorting terminal can use the fully connected layer to map the high-dimensional feature representation contained in the virtual waste feature model to a feature space with a dimension equal to the preset total number of waste categories through a linear transformation matrix, thereby obtaining the initial score for each category. Then, the initial score is converted into a probability distribution form in the interval of 0 to 1 with the sum of each component being 1 through the Softmax activation function, thereby obtaining the initial category probability vector. This initial category probability vector reflects the preliminary classification confidence distribution under the pure visual-physical fusion features.
[0144] Optionally, in S52, the sorting terminal can input waste material distribution data and waste density distribution data from the virtual waste feature model to the physical verification branch. The physical verification branch forms a feature pair by combining the material type indicated by the waste material distribution data and the density value indicated by the waste density distribution data. It then performs cosine similarity calculation with the standard material-density correlation quantization value corresponding to each category in the pre-configured material-density correlation benchmark matrix to obtain a similarity score between the virtual waste feature model and the standard correlation patterns of each category. The sorting terminal can generate a class-by-class confidence correction coefficient based on this similarity score. A higher similarity score indicates that the material-density combination of the waste better conforms to the physical attribute rules of that category, and the corresponding correction coefficient is larger; conversely, a lower similarity score results in a smaller correction coefficient, thus achieving confidence adjustment based on physical consistency.
[0145] Optionally, the pre-configured material-density correlation reference matrix is set as follows: Where C represents the total number of waste categories, and each row stores the standard material feature vector for that category, such as the material spectral principal component score and standard density value. Waste material distribution data extracted from the virtual waste feature model (which can be fused into a material feature vector) (and waste density distribution data (which can be fused into a density feature scalar)) Combined into physical feature vectors Let the unified dimension be D. For each category c, calculate the cosine similarity:
[0146] ;in Let be the vector corresponding to the c-th row of the baseline matrix. The class weights are obtained by Softmax normalization of the similarity scores. , Temperature coefficient (can be taken as...) Class-by-class confidence correction coefficient Take directly as Or further integrate with the initial confidence level, such as In this embodiment, let This is used for subsequent steps to weight the initial class probability vector class by class.
[0147] Optionally, in S53, the sorting terminal can perform weighted correction on the probability values corresponding to each category in the initial category probability vector according to the category-by-category confidence correction coefficient. For example, the initial probability value is multiplied by the corresponding confidence correction coefficient and then re-normalized to obtain the waste category probability vector. The sorting terminal can determine the identification and classification result of the waste to be identified based on the category index corresponding to the maximum probability value in the waste category probability vector, and output the waste identification and classification result and its corresponding confidence score by combining the spatial structure information in the virtual waste feature model. This embodiment, through a dual-branch collaborative classification architecture of the main classification branch and the physical verification branch, can further introduce material-density physical correlation constraints on the basis of visual-physical fusion features to perform physical rationality verification of the classification results, avoiding misjudgment of waste categories with similar materials or similar appearances but different physical properties by the pure data-driven model, thereby improving the accuracy and physical interpretability of the classification results.
[0148] The aforementioned deep learning-based multi-model waste identification and classification method constructs a multi-source heterogeneous data input foundation by acquiring visual image data, weight sensor data, and infrared material reflectance spectrum data of the waste to be identified, overcoming the deficiency of insufficient information from a single visual modality. The visual image data is input into a convolutional network, where the global contour features of the waste are extracted and a visual feature map is generated through the collaborative processing of dilated convolutional layers, gradient entropy modulation layers, same-material contour continuation layers, and background suppression aggregation layers. This captures the overall shape and appearance structure of the waste and effectively suppresses background interference, reconnects broken contours, and improves contour integrity in complex scenes. The visual feature map is segmented to obtain multiple local image regions, which are then input into an attention network to extract local features and mapped back to the visual feature map to generate enhanced visual feature vectors. This strengthens discriminative local textures and edge responses, improving the model's adaptability to complex visual scenes such as stacking, deformation, and occlusion.
[0149] Physical property feature vectors are extracted from waste based on weight sensor data and infrared material reflectance spectral data. This quantifies the dynamic weight characteristics and material spectral response characteristics of waste, overcoming the limitation of pure visual models in perceiving physical properties. A virtual waste feature model is obtained by multimodal feature modeling based on enhanced visual feature vectors and physical property feature vectors. This model deeply integrates visual appearance information and physical property information at the feature level, leveraging the complementary advantages of multi-source heterogeneous data. The virtual waste feature model is then input into a classification network containing a main classification branch and a physical verification branch to generate waste identification and classification results. This enables accurate classification by integrating visual and physical multidimensional information, and avoids misjudgment of waste with similar materials or appearances through physical consistency verification, improving the accuracy and robustness of waste identification and classification in complex scenarios.
[0150] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0151] Based on the same inventive concept, this application also provides a deep learning-based multi-model waste identification and classification system for implementing the aforementioned deep learning-based multi-model waste identification and classification method. The solution provided by this system is similar to the implementation scheme described in the above method. Therefore, the specific limitations of one or more deep learning-based multi-model waste identification and classification system embodiments provided below can be found in the limitations of the deep learning-based multi-model waste identification and classification method described above, and will not be repeated here.
[0152] In one exemplary embodiment, such as Figure 3 As shown, a deep learning-based multi-model waste identification and classification system is provided, including:
[0153] The waste data acquisition module 101 can be used to acquire visual image data, weight sensor data, and infrared material reflectance spectrum data of the waste to be identified;
[0154] The visual feature extraction module 102 can be used to input visual image data into a convolutional network, extract global contour features of garbage, and generate a visual feature map.
[0155] The local feature enhancement module 103 can be used to segment the visual feature map to obtain multiple local image regions, input each local image region into the attention network, extract local features, and map the local features onto the visual feature map to generate an enhanced visual feature vector.
[0156] The physical property extraction module 104 can be used to extract physical property features of waste based on weight sensor data and infrared material reflectance spectrum data to obtain physical property feature vectors.
[0157] The multimodal modeling module 105 can be used to perform multimodal feature modeling based on enhanced visual feature vectors and physical attribute feature vectors to obtain a virtual garbage feature model;
[0158] The classification result output module 106 can be used to input the virtual waste feature model into the classification network to generate waste identification and classification results.
[0159] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the deep learning-based multi-model garbage identification and classification method as described above.
[0160] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0161] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0162] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.
Claims
1. A multi-model garbage identification and classification method based on deep learning, characterized in that, The method includes: Acquire visual image data, weight sensor data, and infrared material reflectance spectrum data of the waste to be identified; The visual image data is input into a convolutional network to extract global contour features of the garbage and generate a visual feature map. The visual feature map is segmented to obtain multiple local image regions. Each local image region is input into an attention network to extract local features. The local features are then mapped onto the visual feature map to generate an enhanced visual feature vector. Based on the weight sensor data and the infrared material reflectance spectrum data, the physical property features of the waste are extracted to obtain a physical property feature vector; Based on the enhanced visual feature vector and the physical attribute feature vector, multimodal feature modeling is performed to obtain a virtual waste feature model; The virtual waste feature model is input into the classification network to generate waste identification and classification results.
2. The method according to claim 1, characterized in that, The convolutional network includes a dilated convolutional layer, a gradient entropy modulation layer, a contour continuation layer of the same material, and a background suppression aggregation layer; The step of inputting the visual image data into a convolutional network to extract global contour features of garbage and generate a visual feature map includes: The visual image data is input into the dilated convolutional layer, and the visual image data is convolved by multiple convolutional kernels pre-configured based on the optical properties of garbage material to obtain a convolutional feature map. Based on the convolutional feature map, a material-related edge response map is generated by calculating pixel gradients in multiple directions; The material-associated edge response map is input into the gradient entropy modulation layer, the gradient direction of each pixel in the material-associated edge response map is extracted to generate a gradient direction matrix, and the response intensity of each pixel in the material-associated edge response map is extracted to generate a response intensity matrix. Based on the gradient direction matrix, the distribution entropy within multiple preset neighborhood windows is statistically analyzed to obtain the direction entropy matrix; according to the preset mapping relationship of the optical properties of the waste material, the direction entropy matrix is mapped to the modulation weight of each neighborhood window; according to the modulation weight, the material-related edge response map is weighted and modulated to obtain the modulation edge map. Based on the response intensity matrix, the modulation edge map is segmented by thresholding, and pixels with response intensity greater than a preset response intensity threshold are marked as high response pixels. Based on the high response pixels, an edge attention region mask is generated. The convolutional feature map is transformed into a color space to generate a chroma channel feature map; the edge attention region mask and the chroma channel feature map are input into the same material contour continuation layer, and endpoint detection is performed on the modulation edge map to identify the set of breakpoint pixels; for each breakpoint pixel in the set of breakpoint pixels, multiple neighborhood blocks of different sizes are taken with the breakpoint pixel as the center, and the mean values of the pixels in the neighborhood blocks in multiple different color channels in the color space are calculated to obtain the breakpoint chroma feature vector; Based on the breakpoint chromaticity feature vector of each breakpoint pixel, calculate the Euclidean distance and chromaticity distance between any two breakpoint pixels. When the Euclidean distance between the two breakpoint pixels is less than a preset distance threshold, the chromaticity distance is less than a preset chromaticity distance threshold, and both breakpoint pixels are located within the edge attention region mask or the overlap rate between the connecting line and the edge attention region mask is greater than a preset overlap rate threshold, they are determined to be a connectable breakpoint pair. For each connectable breakpoint pair, interpolation filling is performed along the connecting line of the breakpoint pixels to generate a continuous contour feature map. The continuous contour feature map is input into the background suppression aggregation layer. A connected component set is obtained through connected component analysis. The area, shape compactness, and edge intensity of each connected component in the connected component set are calculated. Connected components whose area value is less than a preset area threshold, whose shape compactness value is less than a preset shape compactness threshold, and whose average edge intensity value is less than a preset edge intensity threshold are removed. The remaining connected components are used as the foreground connected component set. Feature aggregation is performed on the foreground connected component set to generate the visual feature map.
3. The method according to claim 2, characterized in that, The process involves segmenting the visual feature map to obtain multiple local image regions, inputting each local image region into an attention network to extract local features, and mapping these local features onto the visual feature map to generate an enhanced visual feature vector. This includes: By using a multi-scale sliding window, the visual feature map is traversed according to a preset step size. The variance of pixel response intensity within the coverage area of each sliding window is calculated. The sliding window coverage area with the variance of pixel response intensity greater than the variance threshold is selected to obtain the multiple local image regions. Each of the local image regions is input into the attention network. Through the self-attention mechanism of the attention network, each of the local image regions is linearly mapped to generate a query matrix, a key matrix, and a value matrix for each local image region. Based on the query matrix and the key matrix, the correlation weight between any two pixel positions within each local image region is calculated. Based on the correlation weight, the value matrix of each local image region is weighted and summed to generate a local aggregated feature vector for each local image region. Based on the position coordinates of each local image region in the visual feature map, the local aggregated feature vector of each local image region is mapped back to the original image region corresponding to the position coordinates using bilinear interpolation to obtain the region mapping feature map. The region mapping feature maps corresponding to each of the said position coordinates are concatenated according to the arrangement order of the position coordinates in the visual feature maps to generate the enhanced visual feature vector.
4. The method according to claim 1, characterized in that, The step of extracting physical property features of waste based on the weight sensor data and the infrared material reflectance spectrum data to obtain a physical property feature vector includes: The weight sensor data is sorted according to the data acquisition timestamp to obtain a weight time series sequence, and a weight change rate sequence is calculated based on the weight time series sequence. Identify the set of abrupt change points in the weight change rate sequence that exceed a preset change rate threshold. Based on the set of abrupt change points, distinguish the start time and end time of waste disposal. Using the start time and end time as boundaries, extract the spectral data of the corresponding time interval from the infrared material reflectance spectral data to obtain effective spectral segments. The effective spectral segment is integrated according to a preset wavelength channel to obtain the total reflected energy of each wavelength channel, and the total reflected energy is arranged in wavelength order to generate a spectral energy sequence. For all rate of change values in the weight change rate sequence within the interval from the start time to the end time, calculate the maximum value, mean, standard deviation, and cumulative change to obtain a weight statistical feature vector; for the spectral energy sequence, calculate the peak position, half-width at half-maximum, and energy ratio of the peak value to the mean value of the entire wavelength channel to obtain a spectral statistical feature vector. The physical property feature vector is generated by concatenating the weight statistical feature vector and the spectral statistical feature vector.
5. The method according to claim 4, characterized in that, The process of performing multimodal feature modeling based on the enhanced visual feature vector and the physical attribute feature vector to obtain a virtual waste feature model includes: Based on the enhanced visual feature vector, several individual waste areas are identified through instance segmentation, and virtual waste volume data is generated based on the volume estimation of each individual waste area. Based on all the waste individual regions and the pre-acquired waste individual material prior data, blind source separation is performed on the spectral energy sequence in the physical property feature vector to obtain multiple individual spectral sequences, and each individual spectral sequence is mapped to the corresponding waste individual region to generate waste material distribution data; Based on the waste material distribution data and the preset waste material density benchmark, a material density coefficient is set; Based on the virtual waste volume data of all the waste individual areas, the volume ratio is calculated, and according to the volume ratio and the material density coefficient, the weight statistical feature vector in the physical property feature vector is mapped to each of the waste individual areas to generate waste density distribution data; The virtual waste feature model is obtained by integrating the virtual waste volume data, waste material distribution data, and waste density distribution data of all the aforementioned waste individual regions.
6. The method according to claim 5, characterized in that, The formula for calculating the volume estimation is: in, Let i be the estimated volume of the i-th waste unit region. Let be the projected physical area of the i-th waste unit region in the visual feature map. The preset standard height scale coefficient, Let $\frac{i}{i}$ be the arithmetic mean of the pixel values of all pixels within the $i$-th garbage unit region on the visual feature map. For all individual waste areas The global weighted average, Let be the standard deviation of the pixel values of the visual feature map within the i-th garbage unit region. The maximum value of the weight change rate sequence. This is a reference value for the rate of weight change. , and α represents the preset weighting coefficients, where α is used to adjust the nonlinear contribution of average visual intensity to volume. Enhancement effects used to control texture undulations The correction force used to adjust the weight peak.
7. The method according to claim 6, characterized in that, The classification network includes a main classification branch and a physical verification branch; The step of inputting the virtual waste feature model into the classification network to generate waste identification and classification results includes: The virtual waste feature model is input into the main classification branch, and an initial category probability vector is generated through mapping via a fully connected layer. The waste material distribution data and waste density distribution data in the virtual waste feature model are input into the physical verification branch, and cosine similarity matching is performed with the pre-configured material-density correlation benchmark matrix to generate class-by-class confidence correction coefficients. Based on the class-by-class confidence correction coefficient, the initial category probability vector is corrected class by class to obtain the waste category probability vector; based on the waste category probability vector and the virtual waste feature model, the waste identification and classification result is generated.
8. A multi-model waste identification and classification system based on deep learning, characterized in that, The system includes: The waste data acquisition module is used to acquire visual image data, weight sensor data, and infrared material reflectance spectrum data of the waste to be identified; The visual feature extraction module is used to input the visual image data into a convolutional network, extract global contour features of garbage, and generate a visual feature map; The local feature enhancement module is used to segment the visual feature map to obtain multiple local image regions, input each local image region into the attention network, extract local features, and map the local features onto the visual feature map to generate an enhanced visual feature vector. The physical property extraction module is used to extract physical property features of waste based on the weight sensor data and the infrared material reflectance spectrum data, and obtain a physical property feature vector. The multimodal modeling module is used to perform multimodal feature modeling based on the enhanced visual feature vector and the physical attribute feature vector to obtain a virtual garbage feature model; The classification result output module is used to input the virtual waste feature model into the classification network to generate waste identification and classification results.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.