Ai large model-based mixed plastic sorting robot complete equipment

By combining a multi-label dataset based on a large AI model and a multimodal vector retrieval cascade algorithm with a full visible light solution, the problems of high hardware cost and insufficient recognition accuracy in the sorting of mixed plastics are solved, achieving high-precision and low-cost identification and sorting of plastic bottle materials.

CN121374913BActive Publication Date: 2026-05-12ZHEJIANG LIANYUN ZHIHUI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG LIANYUN ZHIHUI TECH CO LTD
Filing Date
2025-12-15
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing mixed plastic sorting solutions rely on hyperspectral or near-infrared sensors, which have high hardware costs, large system size, insufficient accuracy in identifying contaminated, faded, or overlapping bottles, and lack systematic implementation in industrial sorting scenarios.

Method used

The system employs a multi-label dataset construction and multi-label prompting technology based on a large AI model, combined with a multimodal vector retrieval cascade algorithm. It reduces hardware costs through a full visible light solution, utilizes industrial cameras and edge computing boxes for material classification, and uses an actuator for sorting.

Benefits of technology

It reduces system costs by 35-50%, improves recognition accuracy, and its accuracy in determining the material of contaminated or faded bottles is comparable to that of hyperspectral solutions. It is scalable, easy to maintain, and adaptable to new materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121374913B_ABST
    Figure CN121374913B_ABST
Patent Text Reader

Abstract

The application provides a mixed plastic sorting robot complete equipment based on an AI large model, which comprises a conveying belt assembly for bearing and conveying mixed plastic bottles, a light box assembly and an industrial camera are arranged above the conveying belt assembly, an executing mechanism is arranged in a sorting area of the conveying belt assembly, the industrial camera and the executing mechanism are electrically connected to an AI edge computing box, the AI edge computing box runs a target detection large model algorithm and a multi-modal vector retrieval cascade algorithm, images acquired by the industrial camera are subjected to material classification, and the executing mechanism sorts plastic bottles according to the classification results, and the application adopts technologies such as multi-label data set construction, a large model architecture and multi-label prompting, so that high-precision identification and sorting are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mixed plastic sorting technology, and in particular to a complete set of equipment for mixed plastic sorting robots based on AI large model. Background Technology

[0002] Referring to patent publication number CN118097248A, this method involves acquiring multispectral images of plastic bottles to be recycled, including multi-band infrared and multi-band visible light images. Based on the multispectral images, target detection is performed using the near-infrared band to obtain a multi-band sub-image of the plastic bottle target. The visible light band within the multi-band sub-image is selected to deconstruct the plastic bottle target and determine the presence of cap, body, and label features. If label features are present, feature matching is used to identify the brand of the plastic bottle to be recycled. If no label features are present, end-member pixels are extracted based on the body features and the multi-band sub-image to draw the target. Existing mixed plastic sorting solutions mainly rely on hyperspectral or near-infrared (NIR) sensors combined with traditional single-label detection models: high hardware costs, large system size; and insufficient accuracy in identifying contaminated, faded, deformed, or overlapping bottles. However, with the development of large-scale visual language pre-training (VLP) technology, visible light solutions based on large object detection models are expected to achieve or even surpass the recognition performance of hyperspectral solutions while reducing hardware costs. However, there is currently a lack of systematic implementation for industrial sorting scenarios. Summary of the Invention

[0003] This invention solves the problem of insufficient accuracy in sorting and identifying miscellaneous plastics. It proposes a complete set of equipment for sorting miscellaneous plastics based on an AI large model. It adopts technologies such as multi-label dataset construction, large model architecture and multi-label prompts to achieve high-precision identification and sorting.

[0004] To achieve the above objectives, the following technical solution is proposed:

[0005] A complete set of equipment for sorting miscellaneous plastics based on AI large model includes a conveyor belt assembly that carries and transports miscellaneous plastic bottles. A light box assembly and an industrial camera are provided above the conveyor belt assembly. An actuator is provided in the sorting area of ​​the conveyor belt assembly. The industrial camera and the actuator are electrically connected to an AI edge computing box. The AI ​​edge computing box runs a target detection large model algorithm and a multimodal vector retrieval cascade algorithm to classify the materials of the images acquired by the industrial camera. The actuator sorts the plastic bottles according to the classification results.

[0006] The edge box of this invention is embedded in the box body, reducing wiring complexity. This invention uses full visible light, eliminating the need for hyperspectral / near-infrared modules, reducing system costs by 35-50%. This invention has high recognition accuracy: the accuracy of determining the material of contaminated or faded bottles is comparable to that of hyperspectral solutions. This invention has a certain degree of scalability: at the algorithm level, only incremental fine-tuning is needed to adapt to new materials (such as PP, PS). At the hardware level, the modular design makes it easy to maintain.

[0007] Preferably, the light box assembly is equipped with matrix LEDs and a diffuser plate to improve illuminance uniformity, and the double side plates of the light box assembly are made of an alumina substrate with heat dissipation holes.

[0008] The matrix LED of this invention is used to illuminate a conveyor belt. The light box assembly has a double-sided diffuser plate inside, ensuring illuminance uniformity ≥ 0.9. The two side plates of the light box assembly are made of 1 mm alumina substrate, with heat dissipation holes having an opening of 25%.

[0009] Preferably, the conveyor belt assembly includes a PVC antistatic belt with flexible protrusions and a servo drive motor that provides power for the transport of the PVC antistatic belt. The surface of the PVC antistatic belt is provided with flexible protrusions to reduce the rolling of the bottle. The flexible protrusions on the conveyor belt surface have a diameter of 5 mm and a height of 4 mm to prevent the bottle from rolling out of position.

[0010] Preferably, the encoder of the servo drive motor is triggered synchronously with the shutter of the industrial camera. The encoder is a precision pulse encoder, and the industrial camera is equipped with a fixed-focus lens. The synchronous triggering with the conveyor belt encoder reduces motion blur.

[0011] Preferably, the actuator employs a pneumatic valve array with a response time in the millisecond range.

[0012] The sorting control process of this invention is as follows: the bottle enters the detection area of ​​the lightbox assembly; the industrial camera simultaneously captures images of the bottle, the AI ​​edge computing box performs inference on the images, and the target detection large model algorithm and the multimodal vector retrieval cascade algorithm output the material, color, shape, position information, and delay stamp of the bottle; the scheduling execution mechanism triggers the air valve to sort according to the time window Δt, and after sorting is completed, the bottle falls into the corresponding hopper through the air valve execution mechanism.

[0013] Preferably, the AI ​​edge computing box has a built-in GPU, which stores a multimodal vector knowledge base for bottle material recognition. The AI ​​edge computing box is equipped with a large-scale object detection algorithm and a cascaded multimodal vector retrieval algorithm. The large-scale object detection algorithm outputs candidate bounding boxes and coarse material categories for each bottle in an image labeled with multidimensional attributes. The multidimensional attributes are at least 7 categories. The cascaded multimodal vector retrieval algorithm uses the multimodal vector knowledge base to perform multimodal vector retrieval on the image within the candidate bounding boxes to obtain the accurate material category.

[0014] A method for sorting miscellaneous plastics based on an AI large-scale model, employing the aforementioned complete set of equipment for sorting miscellaneous plastics based on an AI large-scale model, includes the following steps:

[0015] S1, an industrial camera captures images containing plastic bottles;

[0016] S2, perform multi-dimensional attribute annotation for each bottle;

[0017] S3, using the large-scale object detection algorithm, outputs candidate bounding boxes for bottles and coarse material categories;

[0018] S4, the multimodal vector retrieval cascade algorithm uses a multimodal vector knowledge base to perform multimodal vector retrieval on the image within the candidate box to obtain the accurate material category;

[0019] S5 combines the results of S3 and S4 to output the final material determination and classification result.

[0020] Preferably, the target detection large model algorithm dynamically injects material information, color information, and shape information based on the text prompt vector to achieve multi-label parallel inference.

[0021] Preferably, the multimodal vector retrieval cascade algorithm uses a cosine similarity threshold method to reorder the confidence of precise material categories. This invention uses a cosine similarity threshold and selects the Top-K = 5 results for confidence reordering.

[0022] Preferably, step S2 specifically includes the following steps: using annotation software to simultaneously annotate each bottle instance in the image with attributes of 7 dimensions or more.

[0023] This invention simultaneously labels each bottle instance in an image with ≥ 7 dimensions of attributes: material / color / shape (can shape, flat shape, etc.) / size level / gloss / bottle mouth shape / transparency; and this invention uses hierarchical JSON tags and on-chip cache index to support online incremental learning.

[0024] The beneficial effects of this invention are:

[0025] 1. Full visible light: No need for hyperspectral / near-infrared modules, reducing system costs by 35-50%;

[0026] 2. High precision: The accuracy of determining the material of contaminated or faded bottles is comparable to that of hyperspectral methods;

[0027] 3. Scalable: At the algorithm level, only incremental fine-tuning is needed to adapt to new materials (such as PP and PS); at the hardware level, the modular design makes it easy to maintain. Attached Figure Description

[0028] Figure 1 This is a diagram illustrating the device configuration of the present invention. Detailed Implementation Example

[0029] This embodiment proposes a complete set of equipment for sorting mixed plastics based on an AI large model, referencing... Figure 1 The system includes a conveyor belt assembly that carries and transports miscellaneous plastic bottles. Above the conveyor belt assembly are a light box assembly and an industrial camera. The sorting area of ​​the conveyor belt assembly is equipped with an actuator. The industrial camera and the actuator are electrically connected to an AI edge computing box. The AI ​​edge computing box runs a large-scale object detection algorithm and a multimodal vector retrieval cascade algorithm to classify the materials of the images acquired by the industrial camera. The actuator sorts the plastic bottles according to the classification results.

[0030] The light box assembly is equipped with matrix LEDs and a diffuser plate to improve illuminance uniformity. The double side plates of the light box assembly are made of alumina substrate with heat dissipation holes.

[0031] The matrix LED of this invention is used to illuminate a conveyor belt. The light box assembly has a double-sided diffuser plate inside, ensuring illuminance uniformity ≥ 0.9. The two side plates of the light box assembly are made of 1 mm alumina substrate, with heat dissipation holes having an opening of 25%.

[0032] The conveyor belt assembly includes a PVC antistatic belt with flexible bumps and a servo drive motor that provides power for the transport of the PVC antistatic belt. The surface of the PVC antistatic belt is provided with flexible bumps to reduce the rolling of the bottle. The flexible bumps on the conveyor belt surface have a diameter of 5 mm and a height of 4 mm to prevent the bottle from rolling and shifting.

[0033] The encoder of the servo drive motor is triggered synchronously with the shutter of the industrial camera. The encoder is a precision pulse encoder. The industrial camera is equipped with a fixed-focus lens and is triggered synchronously with the conveyor belt encoder to reduce motion blur.

[0034] The actuator employs a pneumatic valve array with a response time in the millisecond range.

[0035] The sorting control process of this invention is as follows: the bottle enters the detection area of ​​the lightbox assembly; the industrial camera simultaneously captures images of the bottle, the AI ​​edge computing box performs inference on the images, and the target detection large model algorithm and the multimodal vector retrieval cascade algorithm output the material, color, shape, position information, and delay stamp of the bottle; the scheduling execution mechanism triggers the air valve to sort according to the time window Δt, and after sorting is completed, the bottle falls into the corresponding hopper through the air valve execution mechanism.

[0036] The AI ​​edge computing box has a built-in GPU, which stores a multimodal vector knowledge base for bottle material recognition. The AI ​​edge computing box is equipped with a large-scale object detection algorithm and a cascaded multimodal vector retrieval algorithm. The large-scale object detection algorithm outputs candidate bounding boxes and coarse material categories for each bottle in an image labeled with multidimensional attributes. The multidimensional attributes are at least 7 categories. The cascaded multimodal vector retrieval algorithm uses the multimodal vector knowledge base to perform multimodal vector retrieval on the image within the candidate bounding box to obtain the accurate material category.

[0037] This embodiment also proposes a method for sorting mixed plastics based on an AI large model, using the aforementioned complete set of equipment for sorting mixed plastics based on an AI large model, including the following steps:

[0038] S1, an industrial camera captures images containing plastic bottles;

[0039] S2, perform multi-dimensional attribute annotation on each bottle; S2 specifically includes the following steps: use annotation software to simultaneously annotate the bottle instances in each image with more than 7 dimensions of attributes.

[0040] This invention simultaneously labels each bottle instance in an image with ≥ 7 dimensions of attributes: material / color / shape (can shape, flat shape, etc.) / size level / gloss / bottle mouth shape / transparency; and this invention uses hierarchical JSON tags and on-chip cache index to support online incremental learning.

[0041] S3, using a large-scale object detection algorithm to output candidate bounding boxes for bottles and coarse material categories; the large-scale object detection algorithm dynamically injects material information, color information and shape information based on text prompt vectors to achieve multi-label parallel reasoning.

[0042] S4, the multimodal vector retrieval cascade algorithm utilizes a multimodal vector knowledge base to perform multimodal vector retrieval on the image within the candidate box to obtain the accurate material category; the multimodal vector retrieval cascade algorithm uses a cosine similarity threshold method to reorder the confidence of the accurate material categories. This invention uses a cosine similarity threshold and selects Top-K = 5 results for confidence reordering.

[0043] S5 combines the results of S3 and S4 to output the final material determination and classification result.

[0044] The working principle of the large-scale object detection algorithm is as follows:

[0045] Multi-label dataset construction: Each bottle instance in the image is simultaneously labeled with ≥ 7 dimensions of attributes: material / color / shape (can-shaped, flat, etc.) / size level / gloss / bottle mouth shape / transparency; hierarchical JSON tags and on-chip cached indexes are used to support online incremental learning.

[0046] The model architecture of the large-scale object detection algorithm is as follows:

[0047] Detection backbone: Based on the Vision-Language Transformer pre-trained with 3 B+ image-text pairs, a model with fewer parameters is distilled through LoRA, which has the ability to detect open categories.

[0048] Multi-label hints: Dynamically inject hint vectors such as "material, color, shape, surface features, etc." during the decoding stage to achieve multi-task reasoning.

[0049] Vector retrieval cascade:

[0050] Coarse segmentation stage: The large model outputs candidate boxes and coarse material probabilities.

[0051] In the precision segmentation stage: the features within the candidate bounding box are extracted into 512-dim semantic vectors using CLIP-ViT, and the Top-K bottle templates are retrieved from the FaissGPU index.

[0052] Fusion strategy: confidence reordering + multi-label consistency loss + weighted average.

[0053] The process includes: Confidence Reordering: By reordering the predicted values ​​for each label, and combining consistency constraints and prior knowledge, the accuracy of label predictions is improved. Multi-Label Consistency Loss: During training, the consistency loss is optimized to ensure logical consistency between labels and avoid prediction conflicts between different labels. Weighted Averaging: By weighting the prediction results of each label, the confidence and importance of each label are comprehensively considered to obtain the fused output of the multi-label task.

[0054] The purpose of confidence reordering is to improve the accuracy and reliability of inference results by re-ranking the prediction results for each label to ensure that the most likely labels are selected first. Specific steps: For each input image, the model outputs a probability distribution of multiple labels (such as material, color, shape, etc.). The prediction result for each label is a probability value obtained through the sigmoid activation function. Confidence reordering weights and ranks the initial confidence based on several factors, including: Original prediction confidence: The probability value Pk of each label (from the model output), representing the initial confidence of the label. Consistency constraints: The relationships between labels; for example, material (PET) is usually closely related to color (transparent), so consistency between labels needs to be considered during reordering. Prior knowledge: Based on the model's historical training or domain experience, prior weights can be set for certain labels. For example, certain materials (such as plastic) appear more frequently in specific production environments, and these labels can be given higher initial weights. Reordered output: Each label is reordered according to the weighted confidence to obtain the ranking of the most relevant labels. The final output labels will be selected according to this reordering order.

[0055] The goal of multi-label consistency loss is to ensure that predictions across different labels are logically consistent. This is crucial for handling multi-label tasks because some labels have significant dependencies on each other, such as material and color, shape and size.

[0056] Consistency loss can be designed as a contrastive loss constrained by labels, aiming to minimize the semantic consistency error among multiple labels in the same image. For example, material (PET) and color (transparency) are often associated, so the differences between related labels should be small when calculating consistency loss. Optimizing consistency loss helps the model consider the influence of other labels when learning each label, thereby enhancing the coordination between labels. For instance, the model should not only learn the prediction of material but also consider the prediction of color and shape simultaneously, ensuring the logical consistency of these labels. The weighted averaging strategy aims to fuse predictions from different tasks (such as multiple labels) in a weighted manner, making the final output more accurate and reliable. Its basic idea is to weight and fuse the predictions of multiple labels based on the importance or confidence of each label. For example, if the prediction of material is more important or accurate than that of color, the material label can be given a higher weight. After weighted averaging, the final output label will be fused based on the confidence and importance of each label, resulting in a final multi-label output prediction. This process can further improve the model's performance, especially in multi-label prediction tasks, preventing some labels from being ignored due to low confidence.

[0057] The implementation process of the large-scale object detection algorithm and the cascaded multimodal vector retrieval algorithm is as follows:

[0058] Data collection: Collect ≥10,000 visible light images at the production site;

[0059] Labeling toolchain: The optimized labeling software X-Anylabeling supports multi-label hierarchy management.

[0060] The optimized annotation software first uses GroundingDINO+SAM to implement object contour annotation in the code, generating annotation files. Then, it imports the images and annotation files into X-Anylabeling, and integrates its own algorithm model for AI-assisted annotation. Following standard naming rules, the name of each class is encoded into a hierarchical path, and then the script breaks it down into attribute fields: material / color / shape (jar-shaped, flat, etc.) / size level / gloss / bottle neck shape / transparency.

[0061] Model training: Based on Vision-Language Transformer for detection, CLIP-ViT for semantic encoder, and combined with contrastive learning;

[0062] After preparing the dataset, divide it into Train / Val / Test = 80% / 10% / 10%, and feed the data into the network model.

[0063] Cold start and freeze over 5-10 epochs, freezing the early layers of the VLP backbone and only unfreezing LoRA, the detection head, and the attribute head, thus stabilizing the initial alignment of the multi-label head and the detection head.

[0064] After 10 epochs, we conduct joint training, unfreeze the mid-to-high-level attention (LoRA is still in place), enable alignment loss, and sample positive / negative pairs from the same batch at each step, using bucket-based balanced sampling.

[0065] In the final 10 epochs, data augmentation and hyperparameter tuning are removed, allowing the model to revert to fine-tuning on "single images that closely approximate the true distribution" at the end of training. In this way, the model can gradually adapt to complex tasks at different stages, ultimately achieving high-precision recognition and classification, ensuring excellent performance in real-world applications.

[0066] Online inference: Deployed to edge boxes via TensorRT + INT8 quantization;

[0067] A multi-label inference algorithm deployed on edge boxes simultaneously handles multiple tasks such as material detection, color classification, and shape recognition. Taking a single image as input, it outputs multiple labels through a joint inference model. First, image preprocessing is performed at the front end. Then, inference is performed through a multi-branch network, with each branch responsible for one task. Finally, the results of multiple tasks are fused to obtain the final prediction. The specific process is as follows:

[0068] The first step, image preprocessing and input format: The raw images captured by the industrial camera are first received by the processing module in the edge box. The camera can trigger encoder signal synchronization to ensure that each image is synchronized with the conveyor belt position. The image resolution is typically 1280×960 or 1920×1080, depending on the equipment configuration. White balance correction is performed on the image to ensure color accuracy. The image is then scaled to the input size required by the model, and normalization is performed to convert pixel values ​​to the range [0, 1].

[0069] The second step is loading and executing the multi-label inference model: The model is converted onto an edge computing box, the converted model is loaded, and the preprocessed image data is input into the inference model, outputting a multi-label prediction result. For each label (such as material, color, shape, etc.), the model outputs a probability value representing the probability of that category. The output probability of each label is compared with a preset threshold; if it is greater than the threshold, the label is considered activated.

[0070] Cascaded search: Faiss GPU index, vector dimension 512, Top-K = 5.

[0071] The purpose of cascaded retrieval is to retrieve the most relevant labels from predictions across multiple categories and then fuse them to obtain more accurate inference results. Through cascaded retrieval, we can not only improve the model's retrieval accuracy but also filter out the optimal labels from a large number of candidate labels, reducing computational load and optimizing inference efficiency.

[0072] The model first performs an initial screening using coarse-grained category labels, then conducts a second round of retrieval using more refined features, ultimately obtaining labels with high confidence. In multi-label tasks, the output probabilities of each label (such as material, color, shape, etc.) are independent, but they have strong intrinsic correlations. Cascaded retrieval first filters out the relevance of candidate labels, then further optimizes the final prediction through refined retrieval. Although the output of each label is independent, in real-world tasks, there are often correlations between labels. Assuming we have obtained each output label that meets the threshold, for example, certain materials (such as PET) and colors (transparent) often appear together, and shapes (cylindrical) and transparency (high) may also be correlated. Based on our internally defined set of category output rules, there are mainly two rules: 'definite output' and 'probable output'. For example, if the output is PET, cylindrical, and transparent label, it means that we have determined that the bottle is made of three-color PET material; if the output is HDPE, cylindrical, and opaque label, it means that we have determined that the bottle is made of HDPE material; if the output is HDPE, cylindrical, and transparent label, it means that the bottle is suspected to be made of HDPE material. When we receive the suspected HDPE material bottle, we will combine it with other categories of labels such as color and bottle mouth shape to determine the category of the output.

[0073] Table 1 compares the performance of the target detection large model algorithm in this embodiment with that of the traditional detection model.

[0074] Table 1. Comparison of the performance of traditional detection models and large-scale target detection algorithms.

[0075] The edge box of this invention is embedded in the box body, reducing wiring complexity. This invention uses full visible light, eliminating the need for hyperspectral / near-infrared modules, reducing system costs by 35-50%. This invention has high recognition accuracy: the accuracy of determining the material of contaminated or faded bottles is comparable to that of hyperspectral solutions. This invention has a certain degree of scalability: at the algorithm level, only incremental fine-tuning is needed to adapt to new materials (such as PP, PS). At the hardware level, the modular design makes it easy to maintain.

Claims

1. A complete set of equipment for sorting miscellaneous plastics using an AI large-scale model, characterized in that, The device includes a conveyor belt assembly that carries and transports miscellaneous plastic bottles. A light box assembly and an industrial camera are mounted on top of the conveyor belt assembly. An actuator is provided in the sorting area of ​​the conveyor belt assembly. The industrial camera and the actuator are electrically connected to an AI edge computing box. The AI ​​edge computing box has a built-in GPU, and the GPU stores a multimodal vector knowledge base for bottle material recognition. The AI ​​edge computing box runs a target detection large model algorithm and a multimodal vector retrieval cascade algorithm to classify the materials of the images acquired by the industrial camera, and the actuator sorts the plastic bottles according to the classification results; The target detection large model algorithm outputs candidate bounding boxes and coarse material categories for bottles in each image labeled with multi-dimensional attributes. The multimodal vector retrieval cascade algorithm uses a multimodal vector knowledge base to perform multimodal vector retrieval on the image within the candidate box to obtain the accurate material category, and then uses the cosine similarity threshold method to rearrange the confidence of the accurate material category.

2. The complete set of equipment for sorting miscellaneous plastics based on an AI large model as described in claim 1, characterized in that, The light box assembly is equipped with matrix LEDs and a diffuser plate to improve illuminance uniformity. The double side plates of the light box assembly are made of alumina substrate with heat dissipation holes.

3. The complete set of equipment for sorting miscellaneous plastics based on an AI large model as described in claim 1, characterized in that, The conveyor belt assembly includes a PVC antistatic belt with flexible protrusions and a servo drive motor that provides power for the transport of the PVC antistatic belt.

4. The complete set of equipment for sorting miscellaneous plastics based on an AI large model as described in claim 3, characterized in that, The encoder of the servo drive motor is triggered synchronously with the shutter of the industrial camera.

5. The complete set of equipment for sorting miscellaneous plastics based on an AI large model as described in claim 1, characterized in that, The actuator employs a pneumatic valve array with a response time in the millisecond range.

6. A method for sorting miscellaneous plastics based on an AI large-scale model, employing the complete set of equipment for sorting miscellaneous plastics based on an AI large-scale model as described in claim 5, characterized in that, Includes the following steps: S1, an industrial camera captures images containing plastic bottles; S2, perform multi-dimensional attribute annotation for each bottle; S3, using the large-scale object detection algorithm, outputs candidate bounding boxes for bottles and coarse material categories; S4, the multimodal vector retrieval cascade algorithm uses a multimodal vector knowledge base to perform multimodal vector retrieval on the image within the candidate box to obtain the accurate material category; S5, combining the results of S3 and S4, outputs the final material determination and classification result; The target detection large model algorithm dynamically injects material, color, and shape information based on text prompt vectors; the multimodal vector retrieval cascade algorithm uses cosine similarity thresholding to reorder the confidence of precise material categories.

7. The method for sorting mixed plastics based on the AI ​​large model according to claim 6, characterized in that, S2 specifically includes the following steps: using annotation software to simultaneously annotate the bottle instances in each image with attributes of 7 dimensions or more.