A method, apparatus, medium, and electronic device for classifying a biomass feedstock

Through multimodal sensors and fine segmentation technology, combined with the pollutant area ratio and category sensitivity parameters, the classification model feature information is adjusted to solve the accuracy and adaptability problems of the existing waste classification system and achieve more accurate waste classification.

CN120388246BActive Publication Date: 2025-10-10HANGZHOU TIANYAN ZHILIAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510889093.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-10
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

The existing waste classification system relies on manual or traditional image recognition technology, which has the disadvantages of low classification accuracy and large errors. It cannot adapt to the influence of factors such as surface contamination and morphological changes of waste, resulting in inaccurate classification.

Method used

A multimodal sensor is used to obtain a multimodal image set of recycled materials. A coarse regional sub-image set is obtained through pollution monitoring and positioning. Fine segmentation processing is performed to obtain a global pollution boundary mask. Combined with the pollutant area ratio and predefined category sensitivity parameters, the feature information of the classification model is adjusted to improve classification accuracy.

Benefits of technology

The accuracy and efficiency of waste classification are improved, human intervention is reduced, the sensitivity and robustness of the model to pollutants are enhanced, and the reliability of classification decisions under complex pollution conditions is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388246B_ABST
    Figure CN120388246B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method, device, medium and electronic equipment for classifying recycled material. A specific embodiment of the method comprises: acquiring an image set in multiple modalities for the recycled material by using a multi-modal sensor; performing pollution monitoring and positioning on the images of a first target modality in the image set to obtain a sub-image set of a rough pollution area; performing fine segmentation processing on the sub-image set of the rough pollution area to obtain a global pollution boundary mask; determining a pollution area proportion according to the global pollution boundary mask; inputting multi-modal fusion feature information corresponding to the image set into a pre-trained classification model to generate predicted feature information; adjusting the predicted feature information according to the pollution area proportion and a category sensitivity parameter set to obtain adjusted feature information; and inputting the adjusted feature information into a classification layer to obtain an impurity classification result. The embodiment can improve classification accuracy, reduce manual intervention, and improve waste recycling efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of computer technology, and more particularly to a classification method, device, medium, and electronic device for recycled materials. Background Art

[0002] In the process of waste sorting and recycling, accurate classification is the key to improving recycling efficiency, reducing pollution and increasing resource reuse rate.

[0003] However, existing waste sorting systems often rely on manual labor or traditional image recognition technology, which often suffers from the following technical issues: low classification accuracy, large errors, and an inability to adapt to factors such as surface contamination and morphological changes in waste. Therefore, developing an intelligent sorting method that can adapt to environmental changes such as waste morphology and surface contamination is a key technology to improve waste treatment efficiency. Summary of the Invention

[0004] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] Some embodiments of the present disclosure provide a method, device, medium, and electronic device for classifying recycled materials to solve one or more of the technical problems mentioned in the above background technology section.

[0006] In a first aspect, some embodiments of the present disclosure provide a classification method for recycled materials, including: using a multimodal sensor to obtain a multimodal image set for recycled materials; performing pollution monitoring and positioning on the image of the first target modality in the above image set to obtain a sub-image set of coarse polluted areas; performing fine segmentation processing on the above sub-image set of coarse polluted areas to obtain a global pollution boundary mask after fine segmentation; determining the pollutant area ratio based on the above global pollution boundary mask; inputting the multimodal fusion feature information corresponding to the above image set into a pre-trained classification model to generate predicted feature information, wherein the above predicted feature information is a vector consisting of impurity categories and impurity ratios, wherein the above predicted feature information is the output result of the previous layer of the classification layer in the above classification model; adaptively adjusting the impurity ratio corresponding to each impurity category in the above predicted feature information based on the above pollutant area ratio and a predefined set of category sensitivity parameters for each impurity category to obtain adjusted feature information; inputting the above adjusted feature information into the above classification layer to obtain an impurity classification result.

[0007] In a second aspect, some embodiments of the present disclosure provide a classification device for recycled materials, comprising: an acquisition unit configured to acquire a multimodal image set for recycled materials using a multimodal sensor; a coarse segmentation unit configured to perform pollution monitoring and positioning on the image of the first target modality in the above image set to obtain a sub-image set of coarse polluted areas; a fine segmentation unit configured to perform fine segmentation processing on the sub-image set of coarse polluted areas to obtain a global pollution boundary mask after fine segmentation; a determination unit configured to determine the area proportion of pollutants according to the above global pollution boundary mask; a generation unit configured to generate the multimodal image set corresponding to the above image set. The modal fusion feature information is input into a pre-trained classification model to generate predicted feature information, wherein the predicted feature information is a vector consisting of impurity categories and impurity proportions, wherein the predicted feature information is the output result of the previous layer of the classification layer in the classification model; the adjustment unit is configured to adaptively adjust the impurity proportion corresponding to each impurity category in the predicted feature information according to the above pollutant area ratio and a predefined category sensitivity parameter set for each impurity category to obtain adjusted feature information; the result output unit is configured to input the adjusted feature information into the above classification layer to obtain an impurity classification result.

[0008] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.

[0009] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner in the first aspect is implemented.

[0010] The aforementioned embodiments of the present disclosure have the following beneficial effects: The recycled material classification methods of some embodiments of the present disclosure can improve classification accuracy, reduce manual intervention, and enhance waste recycling efficiency. Specifically, inaccurate recycled material classification is caused by the fact that existing waste classification systems often rely on manual or traditional image recognition techniques, which suffer from low classification accuracy, large errors, and an inability to adapt to factors such as surface contamination and morphological changes in the waste. Based on this, the recycled material classification methods of some embodiments of the present disclosure first utilize a multimodal sensor to acquire a multimodal image set of the recycled material. By acquiring this multimodal image set, the different spectral characteristics are fused to enhance the ability to characterize waste surface contamination through complementary information. Then, contamination monitoring and location are performed on the images of the first target modality in the image set, generating a sub-image set of coarse contamination areas. Processing the first target modality image separately allows for rapid localization of contaminated areas, reducing the computational overhead of multimodal data fusion and improving processing efficiency. By analyzing the targeted modality characteristics, a preliminary contamination range is screened, reducing the complexity of subsequent refined processing. Next, the coarsely contaminated sub-image set is refined by segmentation to generate a global contamination boundary mask. This refined segmentation eliminates misjudgments from coarse detection, accurately locates contamination boundaries, and, by integrating global image information, avoids local omissions and ensures a complete outline of the contaminated area. Secondly, the contamination area percentage is determined based on the global contamination boundary mask. This provides a quantitative basis for dynamic adjustment of the classification model, accurately reflecting the influence of contaminants on the classification results and enhancing the model's sensitivity to contamination levels. Furthermore, the multimodal fusion feature information corresponding to the image set is input into a pre-trained classification model to generate predicted feature information. This predicted feature information is a vector consisting of impurity categories and impurity proportions, and is the output of the previous layer in the classification model. The multimodal fusion feature information is input into the classification model to enhance the characterization of recycled material materials and the thermal radiation and fluorescence properties of contaminants. Using the intermediate feature output of the previous layer in the classification layer preserves spatial location information for contaminated area location while avoiding global computational redundancy in the fully connected layer. This significantly improves model inference efficiency and provides high-resolution feature support for subsequent dynamic confidence adjustment. Secondly, based on the above-mentioned pollutant area ratio and the predefined category sensitivity parameter set for each impurity category, the impurity ratio corresponding to each impurity category in the above-mentioned predicted feature information is adaptively adjusted to obtain the adjusted feature information. This step quantifies the impact weight of pollutants on different categories and dynamically corrects the model output deviation to make the classification results more in line with the actual pollution scenario. Combined with the predefined sensitivity parameters, the global interference of a single pollution in the classification decision is avoided, which significantly improves the robustness of the method and the reliability of classification decisions under complex pollution conditions. Finally, the above-mentioned adjusted feature information is input into the above-mentioned classification layer to obtain the impurity classification result.Inputting the adjusted feature information into the classification layer can make the impurity classification results more accurate and more in line with the actual pollution scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0012] Figure 1 is a flow chart of some embodiments of a method for sorting recycled materials according to the present disclosure;

[0013] Figure 2 1 is a schematic structural diagram of some embodiments of a classification device for recycled materials according to the present disclosure;

[0014] Figure 3 It is a structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0016] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.

[0017] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0018] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0019] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0020] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0021] refer to Figure 1 , shows a process 100 of some embodiments of the method for classifying recycled materials according to the present disclosure. The method for classifying recycled materials includes the following steps:

[0022] Step 101: Using a multimodal sensor, a multimodal image set of recycled materials is acquired.

[0023] In some embodiments, the entity that executes the aforementioned classification of recycled materials can be hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is software, it can be installed in the hardware devices listed above. It can be implemented as multiple software programs or software modules to provide distributed services, or as a single software program or software module. This is not specifically limited here.

[0024] In other embodiments, the execution entity may utilize a multimodal sensor to acquire a multimodal image set of recycled materials. The multimodal sensor may include an RGB camera, an infrared camera, and a UV camera. The recycled materials may be recycled waste resources that can be reprocessed and reused, such as scrap metal, plastic, and paper. The multimodal image set is a collection of image data collected simultaneously by multiple sensors, such as an RGB camera, an infrared camera, and a UV camera.

[0025] Step 102 : performing pollution monitoring and positioning on the image of the first target modality in the image set to obtain a sub-image set of rough pollution areas.

[0026] In some embodiments, the execution entity may perform contamination monitoring and location on the image of the first target modality in the image set to obtain a set of sub-images of a rough contaminated area. The image of the first target modality may be an RGB image captured by an RGB camera. The sub-images of the rough contaminated area may be image segments obtained through preliminary screening to identify suspected contaminated areas. Contamination monitoring and location may be a process of rapidly identifying contaminants (e.g., oil stains, rust) on the surface of waste materials and locating their rough areas using image processing techniques (e.g., HSV color space conversion, adaptive threshold segmentation, and morphological optimization).

[0027] In some optional implementations of some embodiments, the execution entity may perform pollution monitoring and positioning on the image of the first target modality in the image set to obtain a rough pollution area sub-image set, which may include the following steps:

[0028] The first step is to perform color space conversion on the image in the first target modality to obtain a converted image. The image in the first target modality can be an RGB image, and the converted image can be an HSV image. In practice, an open source library (for example, OpenCV's cv2.cvtColor(src, cv2.COLOR_BGR2HSV)) can be used to map the RGB values ​​of each pixel in the RGB image to the HSV space to obtain an HSV image.

[0029] The second step is to perform adaptive threshold segmentation on the brightness channel of the converted image to obtain a binary mask image. In practice, first, a sliding window (e.g., 5×5 pixels) can be used to traverse the V channel image. Then, the local mean within each window is determined ( ) and standard deviation ( ). According to the formula Dynamically generate a threshold. Next, compare each pixel value with the threshold. If the pixel value is greater than the threshold, it is considered an uncontaminated area and can be set to 0. If the pixel value is less than or equal to the threshold, it is considered a contaminated area and can be set to 1. Finally, a binary mask image is obtained.

[0030] The third step is to denoise the binary mask image using the target denoising algorithm to obtain a denoised binary mask image. In practice, the binary mask image is first eroded and then dilated using a structuring element (e.g., a 3×3 rectangular kernel) to remove isolated noise points and glitches, resulting in the first denoised result. Then, the dilation and erosion operations are performed on this first denoised result to fill small holes and connect adjacent regions to avoid fragmentation of contaminated areas. Finally, the denoised binary mask image is obtained.

[0031] The fourth step is to use the target algorithm to mark the connected regions of the denoised binary mask image to obtain a connected region image set. In practice, the 8-neighborhood or 4-neighborhood rule can be used to traverse each pixel of the binary image to determine whether it is connected to the surrounding pixels. Each newly discovered connected region is assigned a unique label, and pixels in the same region are given the same label, thereby obtaining a connected region image set. Each connected region image represents an independent contaminated region, and the final output is a set of multiple sub-images, each of which retains only a complete contaminated region and its boundary information. The target algorithm can be the cv2.connectedComponents() function in OpenCV.

[0032] In the fifth step, the image in the first target modality is cropped using the bounding box coordinates corresponding to each connected region image in the connected region image set to obtain a set of subimages of the contaminated coarse region. First, for each subimage in the connected region image set, its minimum enclosing rectangular bounding box is determined. Then, the bounding box coordinates are mapped back to the original RGB image (the first target modality) to ensure that the cropped range accurately corresponds to the contaminated region. Finally, the original RGB image is cropped according to the bounding box coordinates. For example, OpenCV's cv2.rectangle() method can be used to crop the original RGB image.

[0033] Step 103 : performing fine segmentation processing on the sub-image set of the rough polluted area to obtain a finely segmented global polluted boundary mask.

[0034] In some embodiments, the execution entity may perform fine segmentation on the contaminated coarse region sub-image set to obtain a finely segmented global contaminated boundary mask, wherein the global contaminated boundary mask may be a binary image generated by performing pixel-level fine segmentation on the contaminated coarse region sub-image set.

[0035] In some optional implementations of some embodiments, the execution entity may perform fine segmentation processing on the rough contaminated area sub-image set to obtain a finely segmented global contaminated boundary mask, which may include the following steps:

[0036] In the first step, each contaminated coarse region sub-image in the contaminated coarse region sub-image set is preprocessed to obtain a preprocessed sub-image set. In practice, each contaminated coarse region sub-image in the contaminated coarse region sub-image set can be normalized, for example, by scaling each contaminated sub-image to a fixed size (e.g., 256×256 pixels).

[0037] In the second step, the preprocessed sub-images are fed into a pre-trained segmentation model to generate a segmentation probability map. Each pixel in this segmentation probability map represents the probability that the corresponding image region contains pollutants. This pre-trained segmentation model can be a U-Net or FCN model, consisting of an encoder (to extract multi-scale features) and a decoder (to restore spatial resolution), integrating contextual information via skip connections. The encoder uses ResNet-50 to extract features, while the decoder generates the segmentation probability map through upsampling and convolution.

[0038] The third step is to convert the segmentation probability map into a global binary mask image using a target conversion algorithm. The target conversion algorithm can be the Otsu adaptive threshold segmentation method, or it can be a method that converts the contamination probability (0-1) of each pixel in the probability map into a binary value (0 or 1) by setting a threshold or algorithm.

[0039] The fourth step is to optimize the global binary mask image to obtain an optimized global binary mask image. The specific implementation method is described in the generation of the denoised binary mask image and will not be repeated here.

[0040] In the fifth step, the optimized global binary mask image is subjected to boundary optimization and spatial alignment to obtain a global contamination boundary mask.

[0041] As an example, the optimized global binary mask image is first dilated and eroded to extract the edge contours of the contaminated area, resulting in a global binary image with continuous edge contours. The Zhang-Suen thinning algorithm is then used to remove non-skeleton pixels on the boundary of the global binary image with continuous edge contours, retaining a single-pixel-wide continuous skeleton to obtain a refined single-pixel boundary. A median filter is then applied along the boundary to smooth the jagged edges and preserve geometric continuity, resulting in a global binary mask image with optimized boundaries. Finally, geometric distortion is corrected through feature matching and perspective transformation, and bilinear interpolation resampling is combined to fully align the optimized global binary mask image with the original image spatial position, resulting in a global contamination boundary mask.

[0042] Step 104: Determine the pollutant area ratio based on the global pollution boundary mask.

[0043] In some embodiments, the execution entity may determine the polluted area ratio based on a global pollution boundary mask. The polluted area ratio may be calculated by calculating the ratio of polluted pixels in the global pollution boundary mask (binary image) to the total image area (e.g., oil pollution accounts for 28% of the total image area) to quantify the degree of pollution.

[0044] In some optional implementations of some embodiments, the execution entity may determine the pollutant area ratio based on the global pollution boundary mask, which may include the following steps:

[0045] The first step is to determine the total number of pixels with fixed mask values ​​in the global pollution boundary mask. In practice, we can traverse the binary array corresponding to the global pollution boundary mask image and count the number of polluted pixels (pixel value is 1).

[0046] The second step is to determine the polluted area percentage based on the total number of pixels in the fixed mask value and the total number of pixels in the image under the first target modality. In practice, the total number of pixels in the image under the first target modality can be determined by the image's dimensions (height × width). Then, the ratio of the number of polluted pixels to the total number of pixels in the image under the first target modality is determined. Finally, this ratio is used to determine the polluted area percentage.

[0047] Step 105: Input the multimodal fusion feature information corresponding to the image set into a pre-trained classification model to generate prediction feature information.

[0048] In some embodiments, the execution entity may input the multimodal fusion feature information corresponding to the image set into a pre-trained classification model to generate predicted feature information, wherein the predicted feature information is a vector consisting of impurity categories and impurity ratios, and wherein the predicted feature information is the output result of the previous layer of the classification layer in the classification model. Using the intermediate feature output of the previous layer of the classification layer not only preserves the spatial location information for locating the contaminated area, but also avoids the global computational redundancy of the fully connected layer. The impurity category may refer to the type of pollutants or impurities (such as oil, rust, plastic, etc.) that may be present in the recycled material as predicted by the classification model, usually expressed in the form of discrete labels. The impurity ratio may refer to the classification model's quantitative estimate of the coverage area or volume ratio of each type of impurity in the recycled material.

[0049] As an example, the pre-trained classification model may be a pre-trained convolutional neural network (CNN), such as ResNet-50 and VGGNet.

[0050] Step 106 : Adaptively adjust the impurity ratio corresponding to each impurity category in the predicted feature information according to the pollutant area ratio and a predefined category sensitivity parameter set for each impurity category to obtain adjusted feature information.

[0051] In some embodiments, the execution entity may adaptively adjust the impurity ratio corresponding to each impurity category in the predicted feature information according to the pollutant area ratio and a predefined set of category sensitivity parameters for each impurity category to obtain adjusted feature information.

[0052] In some optional implementations of some embodiments, the execution entity may adaptively adjust the impurity ratio corresponding to each impurity category in the predicted feature information based on the pollutant area ratio and a predefined category sensitivity parameter set for each impurity category to obtain the adjusted feature information, which may include the following steps:

[0053] The first step is to determine the pollutant impact factor based on the area ratio of the above pollutants.

[0054] In practice, the pollutant impact factor can be obtained by the following formula: .

[0055] in, The pollutant impact factor. is the percentage of pollutant area.

[0056] The second step is to determine the standard deviation of the vector corresponding to the above-mentioned prediction feature information. The vector corresponding to the above-mentioned prediction feature information can be , where N is the number of categories. First, according to Determine the mean of the vector corresponding to the predicted feature information , then, determine the standard deviation based on the mean and logits vector.

[0057] In practice, the standard deviation of the vector corresponding to the above predicted feature information can be obtained by the following formula: .

[0058] in, is the standard deviation of the vector corresponding to the predicted feature information, Represents the confidence score of the i-th impurity category in the corresponding vector of the predicted feature information. For each , Represents the mean of the vector corresponding to the predicted feature information.

[0059] The third step is to use the above-mentioned category sensitivity parameter set to determine the sensitivity parameters corresponding to the above-mentioned prediction feature information, and obtain a sensitivity parameter vector. Among them, the above-mentioned category sensitivity parameter set can be based on the existing classification labels, and the basic sensitivity parameters of each recycled material to the impurity category are pre-set. For example, "clean high-grade aluminum sheet" is inversely proportional to the pollutants, that is, the more pollutants, the lower the probability that it may be "clean high-grade aluminum sheet"; "oil-containing aluminum casting" is directly proportional to the pollutants, that is, the more pollutants, the greater the probability that it may be "oil-containing aluminum casting". For categories that are inversely proportional to pollutants, set the basic sensitivity parameter to -1, for categories that are directly proportional to pollutants, set the basic sensitivity parameter to 1, and for categories that are not related to pollutants, set the basic sensitivity parameter to 0. In practice, the parameter α can be used i Represents the basic sensitivity parameters of the i-th category. The sensitivity parameter vector can be [α1, α2, α3]. For example, the sensitivity parameter vector can be [1, -1, 0], corresponding to the sensitivity parameters of oiled aluminum castings, clean high-grade aluminum sheets, and rusted aluminum sheets, respectively.

[0060] The fourth step is to use the above-mentioned pollution impact factor, the sensitivity parameter corresponding to each predicted feature information in the above-mentioned sensitivity parameter vector and the above-mentioned standard deviation to adjust the impurity ratio corresponding to the corresponding impurity category to obtain the adjusted impurity ratio corresponding to each impurity category and obtain the adjusted feature information.

[0061] In practice, the adjusted feature information can be obtained by the following formula: .

[0062] in, Represents the confidence score of the i-th impurity category in the corresponding vector of the predicted feature information before adjustment, represents the basic sensitivity parameter of the i-th impurity category, σ is the standard deviation of the corresponding vector of predicted feature information, and F is the impact factor of the pollutant. is the confidence score of the i-th impurity category in the vector corresponding to the adjusted predicted feature information.

[0063] When using the above solution to adaptively adjust the predicted characteristic information corresponding to the recycled materials, if the recycled materials are high-value samples, the following technical problem often arises: "How to ensure the accuracy of the adjusted characteristic information corresponding to the recycled materials?" The factors that lead to this technical problem are often as follows: Using a linear model to adjust the predicted characteristic information corresponding to the recycled materials lacks an understanding of the correlation between characteristic information. Therefore, the following solution can be decided:

[0064] In the first step, a class sensitivity matrix is ​​constructed based on a predefined class sensitivity parameter set for each impurity class. The class sensitivity matrix can be a parameter matrix that quantifies the strength of the association between the impurity class and the pollutant, which can be .

[0065] As an example, the vector corresponding to the sensitivity parameter set may be [1, -1, 0], corresponding to: oiled aluminum casting ( ): The more oil stains there are, the more likely it is an oily aluminum casting.

[0066] Clean aluminum ( ): The more oil stains there are, the less likely it is to be clean aluminum.

[0067] Rusted aluminum plate ( ): Oil pollution has no effect on it.

[0068] The constructed category sensitivity matrix can be: .

[0069] The second step is to construct a pollution association graph using the aforementioned pollutant area percentage, the aforementioned category sensitivity matrix, and the eigenvectors corresponding to the aforementioned predicted characteristic information. The pollution association graph is a graph-structured data structure consisting of nodes (impurity categories) and edges (associations between impurity categories). In practice, edge weights can be determined by combining the sensitivity matrix, the pollutant area percentage, and the eigenvectors corresponding to the predicted characteristic information.

[0070] In practice, the edge weight can be determined using the following formula: .

[0071] in, is the area ratio of pollutants, Represents the first prediction scores, sensitivity matrix [ ] indicates the The quantitative parameter of the influence direction (positive / inverse / irrelevant) of each category (such as "clean aluminum plate" and "oiled casting") on the pollutant. ] indicates the connection between “ The weight value of the edge between the "category node" and the "pollutant node".

[0072] As an example, the node feature can be a feature vector corresponding to the predicted feature information, for example, it can be [0.1, 0.8, 0.3], and the edge weight can be [-0.02, 0.16, 0], where the negative sign indicates inverse proportion, and the larger the value, the stronger the inverse proportion, and the positive sign indicates direct proportion, and the larger the value, the stronger the direct proportion.

[0073] In the third step, the contamination association graph is input into a pre-trained graph embedding feature generation model to generate a graph embedding feature vector. The pre-trained graph embedding feature generation model can be a GNN (Graph Neural Network), which can convert discrete impurity categories into continuous low-dimensional vector representations. The graph embedding feature vector can be a vector that captures the global association information of the graph. In practice, the graph embedding feature vector can be a vector of shape (1, D) (D is the embedding dimension, for example, 8). For example, the graph embedding feature vector can be [0.2, -0.1, 0.3, 0.0, 0.1, 0.2, -0.1, 0.0].

[0074] The fourth step is to perform a linear transformation on the graph embedding feature vector to obtain a linearly transformed feature vector. This linear transformation can be used to map the graph embedding feature vector to a task-relevant feature space (such as the classification confidence adjustment space). In practice, this can be done through a fully connected layer (linear layer) for dimensionality mapping. The linearly transformed feature vector can have a shape of (1, M) (where M is the target dimension; for example, M = 3 corresponds to three types of impurities). In practice, the linearly transformed feature vector can be [0.5, -0.2, 0.3].

[0075] Step 5: Apply the target activation function to the linearly transformed feature vector to filter features and obtain an activated feature vector. The activation function can be ReLU activation. Using a nonlinear activation function can filter important features and suppress noise or irrelevant features. For example, if the linearly transformed feature vector is [0.5, -0.2, 0.3], the ReLU activation function can be used to preserve non-negative features, resulting in an activated feature vector of [0.5, 0.0, 0.3].

[0076] Step 6: Normalize the activated feature vector to obtain the weight adjustment vector. In practice, the Softmax function can be used to normalize the activated features into a probability distribution. For example, the activated feature vector [0.5, 0.0, 0.3] can be normalized to obtain the weighted feature vector [0.625, 0, 0.375].

[0077] In the seventh step, the weight adjustment vector is used to adaptively adjust the impurity ratio corresponding to each impurity category in the predicted feature information to obtain adjusted feature information, wherein the adjusted feature information includes the impurity ratio corresponding to each impurity category.

[0078] As an example, the vector corresponding to the feature information before adjustment is [0.1, 0.8, 0.3], the weight feature vector is [0.625, 0, 0.375], and the vector corresponding to the feature information after adjustment is [0.0625, 0, 0.1125].

[0079] The above optional steps and their related content, as an inventive feature of the embodiments of this disclosure, address the aforementioned technical problem of "how to ensure the accuracy of the adjusted characteristic information corresponding to recycled materials." This technical problem is often caused by the use of linear models to adjust the predicted characteristic information corresponding to recycled materials, without understanding the correlations between these characteristics. However, this invention significantly improves the accuracy of the adjusted characteristic information of recycled materials by constructing a contamination correlation map.

[0080] Step 107: input the adjusted feature information into the above classification layer to obtain the impurity classification result.

[0081] In some embodiments, the execution entity may input the adjusted feature information into the classification layer to obtain an impurity classification result.

[0082] In some optional implementations of some embodiments, the execution entity may input the adjusted feature information into the classification layer to obtain an impurity classification result, which may include the following steps:

[0083] In the first step, the impurity ratio corresponding to each impurity category in the adjusted feature information is converted into a probability value using the objective function corresponding to the classification layer, thereby obtaining an impurity classification probability vector. The objective function corresponding to the classification layer can be a Softmax function, which can convert the corresponding impurity ratio into a probability value. The impurity classification probability vector can represent the confidence level of each impurity category. For example, the impurity classification probability vector can be [0.6, 0.1, 0.3] (the sum of probabilities is 1), where the probability of impurity category 1 is 0.6, the probability of impurity category 2 is 0.1, and the probability of impurity category 3 is 0.3.

[0084] The second step is to determine the maximum impurity classification probability in the impurity classification probability vector. For example, in the impurity classification probability vector [0.6, 0.1, 0.3], the maximum impurity classification probability is 0.6.

[0085] The third step is to map the maximum impurity classification probability to the corresponding impurity category label to obtain the impurity category with the highest probability. The impurity category label can be a classification label that identifies the impurity category in the recycled material. For example, if the impurity category label list is ["Oil-Resistant Aluminum Casting," "Clean High-Grade Aluminum Sheet," "Rusted Aluminum Sheet"], then the impurity category with the highest probability is "Oil-Resistant Aluminum Casting."

[0086] In the fourth step, the impurity category with the highest probability is determined as the impurity classification result. For example, the impurity category classification result is "oil-containing aluminum casting".

[0087] In some optional implementations of some embodiments, the execution entity may, after inputting the adjusted feature information into the classification layer and obtaining the impurity classification result, include the following steps:

[0088] The first step is to convert the global contamination boundary mask into a color mask of the same size as the image of the first target modality. In practice, this can be done by iterating over the mask pixel values ​​and mapping each pixel to a predefined color, for example, marking pixel values ​​of 0 with green and pixel values ​​of 1 with red.

[0089] The second step is to superimpose the color mask with the image of the first target modality using a target image blending technique to produce a superimposed image. In practice, this can be achieved using the image blending formula (e.g., result = original image matrix × (1-β) + mask × β), where β is the transparency (typically 0.3-0.7).

[0090] The third step is to generate a classification probability visualization chart based on the above impurity classification probability vector. In practice, the above visualization chart can be used to draw a bar chart or pie chart of the impurity classification probability vector using Matplotlib or Seaborn.

[0091] The fourth step is to generate a pollution indicator visualization component based on the pollutant area percentage and the impurity classification results. The pollution indicator visualization component can be a text box or a graphical label. In practice, OpenCV can be used to draw a text box on the image.

[0092] Step 5: Integrate the superimposed image, the classification probability visualization chart, and the pollution indicator visualization component to obtain a complete classification visualization report. In practice, you can use PIL (Python Imaging Library) or OpenCV to stitch images and components.

[0093] The sixth step is to generate recycling recommendations based on the aforementioned classification visualization report using a pre-created database. This pre-created database can be a structured data set that stores pollutant characteristics and corresponding treatment rules in recycling scenarios. In practice, each piece of structured data may include: pollutant type (e.g., oil, rust, mixed pollution), pollution level indicator (e.g., pollutant area percentage, pollution intensity level), material type (e.g., aluminum, steel, plastic), and treatment recommendations (e.g., cleaning method, repair process, recycling process).

[0094] Traditional methods for training classification models often rely on single-modality images (for example, using only RGB images). This often leads to the following technical issue: single-modality images cannot fully characterize impurity characteristics. This is often caused by the following: RGB images are susceptible to interference from lighting and reflections, infrared images only reflect thermal radiation characteristics, and ultraviolet images only capture the fluorescence characteristics of materials. Therefore, the following approach can be used to train the classification model:

[0095] The first step is to obtain a multimodal image sample set and corresponding label set of recycled materials, wherein the multimodal image sample set includes: a first target modality sample set, a second target modality sample set, and a third target modality sample set. The multimodal image sample set may refer to a collection of multiple types of images acquired by sensors of different modalities for the same batch of recycled materials. The corresponding label set may include: recycled material type, magazine type, and coordinate information of the contaminated area. In practice, the first target modality sample set may be an RGB image set, the second target modality sample set may be an infrared image set, and the third target modality sample set may be an ultraviolet image set.

[0096] The second step is to convert the first target modality sample set in the multimodal image sample set to obtain a converted sample set. The first target modality sample set may be an RGB image sample set, and the converted images may be an HSV image sample set. In practice, the RGB values ​​corresponding to each pixel of each RGB image sample can be mapped to the HSV space using an open source library (e.g., OpenCV's cv2.cvtColor(src, cv2.COLOR_BGR2HSV)) or a custom algorithm to obtain HSV image samples and a resulting HSV image sample set.

[0097] The third step is to traverse the brightness channel of each sample in the transformed sample set according to the target size window to obtain the target size window set. In practice, the target size window can be a 5×5 window.

[0098] The fourth step is to determine the target adaptive threshold based on the pixel mean and standard deviation of the corresponding pixel value set in each window in the above window set. The specific steps are similar to the steps for generating the adaptive threshold when generating the binary mask image above, and will not be repeated here.

[0099] The fifth step is to perform adaptive threshold segmentation on the brightness channel of the converted image to obtain a binary mask image. In practice, first, a sliding window (e.g., 5×5 pixels) can be used to traverse the V channel image. Then, the local mean within each window is determined ( ) and standard deviation ( ). According to the formula Dynamically generate a threshold. Next, compare each pixel value with the threshold. If the pixel value is greater than the threshold, it is considered an uncontaminated area and can be set to 0. If the pixel value is less than or equal to the threshold, it is considered a contaminated area and can be set to 1. Finally, a binary mask image is obtained.

[0100] In step 6, in response to a pixel value in the pixel value set being greater than the target adaptive threshold, the pixel value is replaced with the corresponding pixel mean in the pixel value set to generate a denoised sample, thereby obtaining a denoised sample set. In response to a pixel value in the pixel value set being less than the target adaptive threshold, the corresponding pixel value in the pixel value set is retained.

[0101] In the seventh step, the second target modality sample set, the third target modality sample set, and the denoised image set are fused to obtain a fused image set, wherein the fused image set is an image set after multimodal information fusion.

[0102] In the eighth step, the fused image set is input into a deep feature extraction model for deep feature extraction, generating a deep feature information set. The deep feature extraction model can be a pretrained convolutional neural network (CNN), such as ResNet-50, which extracts global pooling layer features from the image. In practice, each fused image in the fused image set is input into a ResNet-50. After convolution and pooling, a 2048-dimensional feature vector is generated. This is then reduced to 512 dimensions after global average pooling. The feature vector corresponding to the deep feature information in the deep feature information set can be a 512-dimensional feature vector.

[0103] In step nine, based on the aforementioned deep feature information set and the corresponding label set, the pre-built initial classification model is trained using a cross-validation method using a target loss function and a target optimizer to obtain a classification model. The target loss function can be a cross-entropy loss function. The target optimizer can be an Adam optimizer. The cross-validation method can be a five-fold cross-validation method. For example, the 1,000 fused images can be divided equally into five groups, each containing 200 images. Four of these groups (800 fused images) are used to train the classification model, and one group (200 fused images) is used to validate the classification model.

[0104] In step 10, in response to the impurity classification results corresponding to the classification model being manually labeled as incorrect a cumulative target number of times, the classification model is fine-tuned using the image set corresponding to the target number of incorrect results. The fine-tuned model is designated as the classification model and a backup of the pre-fine-tuned classification model is made. In practice, the cumulative target number of times can be 50. When fine-tuning the classification model, the first few convolutional layers can be frozen (to retain common features) and only the last few fully connected layers can be trained (to adapt to the image set corresponding to the incorrectly classified results).

[0105] In step 11, in response to reaching a preset time interval, the classification model is retrained based on the image set and the multimodal image sample set corresponding to the impurity classification result set corresponding to the classification model, and the retrained model is determined as the classification model. In practice, the preset time interval can be three months.

[0106] When training a classification model using the above steps, the following technical issues often arise: difficulty aligning multimodal images and low fusion quality. These issues are often caused by differences in sensor viewing angles, shooting angles, and radiation characteristics between different modal images (e.g., RGB, infrared, and UV). This leads to geometric misalignment (e.g., the position of the same impurity in the RGB and infrared images) and inconsistent radiation. Direct fusion can lead to feature confusion (e.g., impurity regions overlapping and misaligned in the fused image). Therefore, the following solutions can be used:

[0107] Optionally, the fusing of the second target modality sample set, the third target modality sample set, and the denoised image set to obtain a fused image set may include the following steps:

[0108] In a first step, a target feature extraction algorithm is used to extract feature points from the second target modality sample set and the third target modality sample set, to obtain a second feature point set and a second descriptor set corresponding to the second target modality sample set, and a third feature point set and a third descriptor set corresponding to the third target modality sample set. The target feature extraction algorithm can be a SIFT algorithm. The second descriptor set can be a set of descriptors corresponding to the feature points in the second target modality sample set. The third descriptor set can be a set of descriptors corresponding to the feature points in the third target modality sample set. The descriptor can be a "feature fingerprint" of an image, used for subsequent feature matching and multi-modal fusion. The descriptor can be a quantized representation of a local feature in an image, capturing the unique properties (e.g., texture, shape, color) of the region through a set of numerical values (usually a high-dimensional vector). For example, the descriptor corresponding to the vector can be [0.9, 0.1,..., 0.7].

[0109] In a second step, a target feature point matching method is used to match the denoised image set and the second descriptor set, to obtain a second matching pair set, wherein the matching pairs in the second matching pair set include the feature point coordinates and descriptors corresponding to the second matching pair set. The target feature point matching method can be an ORB lightweight feature point matching method. For example, the point (x1, y1) in the denoised image is matched with the point (x2, y2) in the infrared image.

[0110] In a third step, a target feature point matching method is used to match the denoised image set and the third descriptor set, to obtain a third matching pair set, wherein the matching pairs in the third matching pair set include the feature point coordinates and descriptors corresponding to the third matching pair set. The target feature point matching method can be an ORB lightweight feature point matching method. For example, the point (x1, y1) in the denoised image is matched with the point (x3, y3) in the infrared image.

[0111] In a fourth step, a target screening algorithm is used to screen the second matching pair set and the third matching pair set, to obtain a screened second matching pair set and a screened third matching pair set. The target screening algorithm can be a RANSAC algorithm, which estimates a homography matrix by using the RANSAC algorithm, and removes the mismatched point pairs that do not meet the matrix transformation. For example, in the matching pairs of the denoised image and the infrared image, most of the point pairs meet the homographic transformation (i.e., geometric consistency), and a small number of point pairs do not meet the homographic transformation and are removed.

[0112] In the fifth step, a perspective transformation is performed on the filtered second matching pair set, the filtered third matching pair set, and the denoised image set to obtain an aligned second target modality sample set and an aligned third target modality sample set. In practice, a perspective transformation matrix (e.g., a homography matrix H) can be estimated based on the matching pairs, and the second target modality sample images and the third target modality sample images can be transformed into the coordinate system corresponding to the first target modality images using H.

[0113] In the sixth step, radiometric calibration is performed on the aligned second and third target modality sample sets, resulting in calibrated second and third target modality sample sets. In practice, the pixel values ​​of the second and third target modality images can be mapped to the radiometric range of the first target modality image through the radiometric transfer function (RTF) or histogram matching. For example, the pixel value range of an infrared image (0, 255) (low to high temperature) can be calibrated to the pixel value range of an RGB image (0, 255) (dark to bright), making the two comparable at the same scale.

[0114] In the seventh step, the calibrated second target modality sample set, the calibrated third target modality sample set, and the denoised image set are fused to produce a fused image set. In practice, a weight can be assigned to each modality (e.g., 0.4 for RGB images, 0.3 for infrared images, and 0.3 for UV images). The grayscale images of each target modality are then treated as separate channels to generate a multi-channel fused image.

[0115] The above optional steps and their related contents serve as an inventive point of an embodiment of the present disclosure, which solves the above technical problems of "single-modal images cannot fully characterize the characteristics of impurities" and "multi-modal image alignment is difficult and the fusion quality is low". The factors that lead to the above technical problems are often as follows: RGB images are easily affected by light and reflection interference, infrared images can only reflect thermal radiation characteristics, and ultraviolet images can only capture the fluorescence characteristics of materials. Due to differences in sensor viewing angles, shooting angles, and radiation characteristics, images of different modalities have geometric dislocations and radiation inconsistencies, and direct fusion will lead to feature confusion. The present invention uses multi-modal images to train the classification model, and accurately aligns the multi-modal images during training, which can significantly improve the accuracy and robustness of the classification model.

[0116] The aforementioned embodiments of the present disclosure have the following beneficial effects: The recycled material classification methods of some embodiments of the present disclosure can improve classification accuracy, reduce manual intervention, and enhance waste recycling efficiency. Specifically, inaccurate recycled material classification is caused by the fact that existing waste classification systems often rely on manual or traditional image recognition techniques, which suffer from low classification accuracy, large errors, and an inability to adapt to factors such as surface contamination and morphological changes in the waste. Based on this, the recycled material classification methods of some embodiments of the present disclosure first utilize a multimodal sensor to acquire a multimodal image set of the recycled material. This multimodal image set allows for the fusion of different spectral characteristics, enhancing the ability to characterize waste surface contamination through complementary information. Then, contamination monitoring and location are performed on the images of the first target modality in the image set, generating a sub-image set of roughly contaminated areas. Processing the first modality image separately allows for rapid localization of contaminated areas, reducing the computational overhead of multimodal data fusion and improving processing efficiency. Targeted analysis of the target modality features allows for a preliminary screening of the contamination range, reducing the complexity of subsequent refined processing. Next, the coarsely segmented sub-image set of contaminated areas is refined to produce a global contamination boundary mask. Refined segmentation eliminates misjudgments from coarse detection and accurately locates contaminant boundaries. It also combines global image information to avoid missing local areas and ensure a complete outline of the contaminated area. Secondly, the contaminant area percentage is determined based on the global contamination boundary mask. This provides a quantitative basis for dynamic adjustment of the classification model, accurately reflecting the influence of contaminants on the classification results and enhancing the model's sensitivity to contamination levels. Furthermore, the multimodal fusion feature information corresponding to the image set is input into a pre-trained classification model to generate predicted feature information. This predicted feature information is a vector consisting of impurity categories and impurity proportions, and is the output of the previous layer in the classification model. The multimodal fusion feature information is input into the classification model to enhance the characterization of waste material, contaminant thermal radiation, and fluorescence properties. Using the intermediate feature output of the previous layer of the classification layer preserves spatial location information for contaminated area location while avoiding global computational redundancy in the fully connected layer. This significantly improves model inference efficiency and provides high-resolution feature support for subsequent dynamic confidence adjustment. Secondly, based on the above-mentioned pollutant area ratio and the predefined category sensitivity parameter set for each impurity category, the impurity ratio corresponding to each impurity category in the above-mentioned predicted feature information is adaptively adjusted to obtain the adjusted feature information. This step quantifies the impact weight of pollutants on different categories and dynamically corrects the model output deviation to make the classification results more in line with the actual pollution scenario. Combined with the predefined sensitivity parameters, the global interference of a single pollution in the classification decision is avoided, which significantly improves the robustness of the method and the reliability of classification decisions under complex pollution conditions. Finally, the above-mentioned adjusted feature information is input into the above-mentioned classification layer to obtain the impurity classification result.The adjusted feature information is input into the classification layer, the weight of the pollution adjustment feature information can be accurately fused, and the classification result is more suitable for the actual pollution scene.

[0117] Further reference Figure 2 , as an implementation of the method shown in the above figures, the disclosure provides some embodiments of a classification device for recycled material, which device embodiments correspond to those method embodiments shown in Figure 1 , the classification device for recycled material can be applied in various electronic devices.

[0118] As shown in Figure 2 , a classification device for recycled material 200 includes an acquisition unit 201, a coarse segmentation unit 202, a fine segmentation unit 203, a determination unit 204, a generation unit 205, an adjustment unit 206, and a result output unit 207. The acquisition unit 201 is configured to acquire an image set in multiple modalities for recycled material by using a multi-modal sensor. The coarse segmentation unit 202 is configured to perform pollution monitoring and positioning on images of a first target modality in the image set to obtain a pollution rough region sub-image set. The fine segmentation unit 203 is configured to perform fine segmentation processing on the pollution rough region sub-image set to obtain a global pollution boundary mask after fine segmentation. The determination unit 204 is configured to determine a pollutant area ratio according to the global pollution boundary mask. The generation unit 205 is configured to input multi-modal fusion feature information corresponding to the image set into a pre-trained classification model to generate predicted feature information, wherein the predicted feature information is a vector composed of impurity categories and impurity proportions, and the predicted feature information is an output result of a previous layer of a classification layer in the classification model. The adjustment unit 206 is configured to adaptively adjust the impurity proportion corresponding to each impurity category in the predicted feature information according to the pollutant area ratio and a pre-defined category sensitivity parameter set for each impurity category to obtain adjusted feature information. The result output unit 207 is configured to input the adjusted feature information into the classification layer to obtain an impurity classification result.

[0119] It can be understood that the units described in the classification device for recycled material 200 correspond to each step in the method described with reference to Figure 1 . Therefore, the operations, features and beneficial effects described above for the method also apply to the classification device for recycled material 200 and the units contained therein, which will not be described here.

[0120] Reference is made below to Figure 3 , which shows a structural schematic diagram of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the disclosure. Figure 3The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0121] like Figure 3 As shown, electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 302 or programs loaded from a storage device 308 into a random access memory (RAM) 303. RAM 303 also stores various programs and data required for the operation of electronic device 300. Processing device 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to bus 304.

[0122] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as needed.

[0123] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.

[0124] It should be noted that in some embodiments of the present disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. Furthermore, in some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wire, optical cable, RF (radio frequency), or any suitable combination thereof.

[0125] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0126] The computer-readable medium may be included in the electronic device, or may exist independently and not incorporated into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: acquire a multimodal image set of recycled materials using a multimodal sensor; perform pollution monitoring and location on the first target modality image in the image set to obtain a rough contaminated area sub-image set; perform fine segmentation processing on the rough contaminated area sub-image set to obtain a finely segmented global contamination boundary mask; determine the pollutant area ratio based on the global contamination boundary mask; input the multimodal fusion feature information corresponding to the image set into a pre-trained classification model to generate predicted feature information, wherein the predicted feature information is a vector consisting of impurity categories and impurity ratios, and wherein the predicted feature information is the output result of the previous layer of the classification layer in the classification model; adaptively adjust the impurity ratio corresponding to each impurity category in the predicted feature information based on the pollutant area ratio and a predefined set of class sensitivity parameters for each impurity category to obtain adjusted feature information; and input the adjusted feature information into the classification layer to obtain an impurity classification result.

[0127] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0128] The computer program product of the first aspect can include a computer readable storage medium. The computer readable storage medium can include instructions. The instructions can include one or both of: instructions for causing a computer to implement a method as described above; and instructions for causing a computer to operate based on a system as described above. The computer readable storage medium can include one or more of: a magnetic disk; a magnetic tape; a magneto-optical disk; a semiconductor memory (e.g., a RAM, ROM, Flash memory, etc.); and an optical disk.

[0129] The units described in some embodiments of the present disclosure can be implemented by means of software, or by means of hardware. The described units can also be implemented in a processor, for example, a processor can be described as including an acquisition unit, a coarse segmentation unit, a fine segmentation unit, a determination unit, a generation unit, an adjustment unit, and a result output unit. In some cases, the names of these units do not constitute a limitation on the units themselves, for example, the acquisition unit can also be described as a unit that acquires an image set in multiple modalities for the biological material using a multi-modal sensor.

[0130] The functions described above in the detailed description can be performed by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include: Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0131] The above description is merely illustrative of the application of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the application involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above inventive concept. For example, the above features can be replaced with technical features disclosed in the embodiments of the present disclosure (but not limited to) having similar functions to form technical solutions.

Claims

1. A classification method for recycled materials, comprising: Using a multimodal sensor to obtain a multimodal image set of recycled materials; Performing pollution monitoring and positioning on the image of the first target modality in the image set to obtain a sub-image set of rough pollution areas; Performing fine segmentation processing on the sub-image set of the rough polluted area to obtain a finely segmented global pollution boundary mask; Determining the pollutant area ratio according to the global pollution boundary mask; Inputting the multimodal fusion feature information corresponding to the image set into a pre-trained classification model to generate predicted feature information, wherein the predicted feature information is a vector consisting of impurity categories and impurity ratios, and wherein the predicted feature information is an output result of a previous layer of the classification layer in the classification model; Adaptively adjusting the impurity ratio corresponding to each impurity category in the predicted feature information according to the pollutant area ratio and a predefined category sensitivity parameter set for each impurity category to obtain adjusted feature information; The adjusted feature information is input into the classification layer to obtain an impurity classification result.

2. The method according to claim 1, wherein The performing pollution monitoring and positioning on the image of the first target modality in the image set to obtain a pollution rough area sub-image set includes: Performing color space conversion on the image in the first target modality to obtain a converted image; Performing adaptive threshold segmentation on the brightness channel of the converted image to obtain a binary mask image; Performing denoising on the binary mask image using a target denoising algorithm to obtain a denoised binary mask image; Using a target algorithm to mark the connected regions of the denoised binary mask image, to obtain a connected region image set; The image in the first target modality is cropped using the bounding box coordinates corresponding to each connected region image in the connected region image set to obtain a contaminated coarse region sub-image set.

3. The method according to claim 1, wherein The fine segmentation processing is performed on the contaminated coarse region sub-image set to obtain a finely segmented global contaminated boundary mask, including: Preprocessing each contaminated coarse region sub-image in the contaminated coarse region sub-image set to obtain a preprocessed sub-image set; Inputting the preprocessed sub-image set into a pre-trained segmentation model to obtain a segmentation probability map, wherein each pixel value in the segmentation probability map represents the probability that the content of the corresponding image area is a pollutant content; Converting the segmentation probability map into a global binary mask image using a target conversion algorithm; Optimizing the global binary mask image to obtain an optimized global binary mask image; Boundary optimization and spatial alignment are performed on the optimized global binary mask image to obtain a global contamination boundary mask.

4. The method according to claim 1, wherein Determining the pollutant area ratio according to the global pollution boundary mask includes: Determining the total number of pixels with fixed mask values ​​in the global contamination boundary mask; The pollutant area ratio is determined based on the total number of pixels of the fixed mask value and the total number of pixels corresponding to the image under the first target modality.

5. The method according to claim 1, wherein Adaptively adjusting the impurity ratio corresponding to each impurity category in the predicted feature information according to the pollutant area ratio and a predefined category sensitivity parameter set for each impurity category to obtain adjusted feature information includes: Determine the pollution impact factor based on the area ratio of the pollutants; Determining a standard deviation of a vector corresponding to the predicted feature information; Determining the sensitivity parameter corresponding to the predicted feature information using the category sensitivity parameter set to obtain a sensitivity parameter vector; The pollution impact factor, the sensitivity parameter corresponding to each predicted feature information in the sensitivity parameter vector, and the standard deviation are used to adjust the impurity ratio corresponding to the corresponding impurity category to obtain the adjusted impurity ratio corresponding to each impurity category and obtain the adjusted feature information.

6. The method according to claim 1, wherein Inputting the adjusted feature information into the classification layer to obtain an impurity classification result includes: The impurity ratio corresponding to each impurity category in the adjusted feature information is converted into a probability value using the objective function corresponding to the classification layer to obtain an impurity classification probability vector; determining a maximum impurity classification probability in the impurity classification probability vector; Mapping the maximum impurity classification probability to a corresponding impurity category label to obtain a maximum probability impurity category; The impurity category with the highest probability is determined as the impurity classification result.

7. The method according to claim 6, wherein: After inputting the adjusted feature information into the classification layer to obtain an impurity classification result, the method further includes: Converting the global contamination boundary mask to a color mask of the same size as the image of the first target modality; superimposing the color mask and the image of the first target modality by a target image blending technique to obtain a superimposed image; generating a classification probability visualization chart according to the impurity classification probability vector; Generate a pollution index visualization component according to the pollutant area ratio and the impurity classification result; Integrating the superimposed image, the classification probability visualization chart, and the pollution index visualization component to obtain a complete classification visualization report; A recycling recommendation is generated based on the classification visualization report using a pre-created database.

8. A device for sorting recycled materials, comprising: An acquisition unit is configured to acquire a multimodal image set of the recycled material using a multimodal sensor; a coarse segmentation unit configured to perform pollution monitoring and positioning on the image of the first target modality in the image set to obtain a sub-image set of coarse pollution areas; a fine segmentation unit configured to perform fine segmentation processing on the contaminated coarse region sub-image set to obtain a finely segmented global contaminated boundary mask; A determination unit, which determines a pollutant area ratio according to the global pollution boundary mask; a generating unit configured to input the multimodal fusion feature information corresponding to the image set into a pre-trained classification model to generate predicted feature information, wherein the predicted feature information is a vector consisting of impurity categories and impurity ratios, and wherein the predicted feature information is an output result of a previous layer of the classification layer in the classification model; an adjusting unit configured to adaptively adjust the impurity ratio corresponding to each impurity category in the predicted feature information according to the pollutant area ratio and a predefined class sensitivity parameter set for each impurity category, to obtain adjusted feature information; The result output unit is configured to input the adjusted feature information into the classification layer to obtain an impurity classification result.

9. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A computer-readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Waste paper category automatic identification method fusing classification detection segmentation

    CN112836704A

  • Image segmentation method and device, electronic equipment and storage medium

    CN117746035A