Sorting robot complete equipment based on AI large model

Through a complete set of sorting robot equipment based on AI large models, accurate joint identification and dynamic grasping of object location, category, and material are achieved, solving the problems of low recognition accuracy and poor adaptability of traditional sorting equipment in complex material scenarios, and providing an efficient and reliable intelligent sorting solution.

CN120772149APending Publication Date: 2025-10-14ZHEJIANG LIANYUN ZHIHUI TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511096488.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Traditional sorting equipment has low recognition accuracy and poor adaptability when faced with complex, changeable, small-batch or multi-material mixed material scenarios, and cannot meet the needs of fine sorting. In addition, traditional visual systems lack the ability to distinguish highly similar heterogeneous objects.

Method used

A complete set of sorting robot equipment based on AI large models is used, combined with industrial camera arrays, central control systems and robotic arm execution modules. Multimodal large models are used to achieve accurate joint recognition of object location, category, and material, and drive the actuators for dynamic grasping, integrating physical signal auxiliary modules to improve recognition accuracy.

Benefits of technology

It achieves high-precision recognition of the location, category, and material of objects, improves the system's ability to identify unknown or new categories of objects, supports flexible sorting, adapts to high-speed sorting in complex environments, and reduces dependence on massive labeled data for specific objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120772149A_ABST
    Figure CN120772149A_ABST
Patent Text Reader

Abstract

The invention discloses sorting robot complete equipment based on an AI large model, the sorting robot complete equipment comprises a complete sorting assembly line, the assembly line is provided with an industrial camera array, a central control system, a mechanical arm execution module, a core calculation module based on the AI large model and the like, the industrial camera array collects video streams, and an encoder is in signal synchronization with an image collection module; the core calculation module converts video streams into information streams, the central control system makes decisions and plans grabbing tracks according to sorting category semantics and the information streams, a physical signal auxiliary evidence module and the like are further arranged, and all the modules work cooperatively. The technical effects that the positions, categories, materials and other information of the articles can be accurately recognized, efficient and accurate article sorting is achieved, the target category and the container distribution strategy can be adjusted in real time, and the motion parameters of the mechanical arm can be adjusted according to the material attributes of the articles are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of industrial automation sorting equipment, and particularly to a sorting robot complete equipment based on an AI large model. BACKGROUND

[0002] In the fields of industrial production, logistics sorting, and recycling processing, material sorting is an important and complex link. Traditional sorting equipment based on fixed rules (such as color, shape, and size) or single sensors (such as photoelectric and metal detection) has low recognition accuracy, poor adaptability, and cannot meet the needs of fine sorting when facing complex, variable, small batch, or multi-material mixed material scenarios.

[0003] With the development of computer vision and artificial intelligence technology, visual sorting systems based on traditional deep learning models (such as CNN) have been applied, but their training relies on massive specific labeled data, and the model has limited generalization ability. New classes and new materials need to be retrained, and the understanding and flexible configuration ability of semantic information (such as object name and material attribute) are insufficient.

[0004] In recent years, multi-modal large models (such as CLIP, GPT-4V, etc.) have strong cross-modal understanding (image + text) and few-shot / zero-shot learning ability, and can infer rich semantic content from limited information, providing a new technical path to solve the above problems.

[0005] Traditional CV-based material recognition methods have essential limitations, especially in distinguishing high-similarity heterogeneous objects, such as: Reflective confusion: chrome-plated plastic parts and stainless steel parts (reflectivity > 80%) Transparent confusion: acrylic and glass (transmittance > 90%) Coating confusion: PVC imitation leather and real leather (surface texture similarity > 85%) Such defects are due to the three constraints of CV algorithms: Feature expression limitation: relying on artificial design features (such as HOG, LBP), unable to model material microscopic physical properties Physical information missing: RGB images are difficult to capture refractive index, conductivity, and other physical properties Semantic understanding disconnection: "metal" is only considered as visual texture, lacking material science definition Experimental data evidence: In the FMD (Flickr Material Database) dataset test, the classification accuracy of traditional ResNet-50 on the above confusion pairs is only 62.3% ± 5.1% (n = 500 samples).

[0006] Due to the limitations of feature expression of the CV algorithm, missing of physical information and disconnection of semantic understanding, it is difficult to realize accurate joint identification of the position, category and material of the article.

[0007] How to seamlessly integrate such powerful perception capability into high-speed, real-time industrial sorting production lines to realize accurate joint identification of the position, category and material of the article, and drive the actuator to perform dynamic grabbing accordingly, forming a high-efficiency and reliable complete equipment, is still a technical problem to be solved. SUMMARY

[0008] The application solves the problems of limited model generalization ability and insufficient understanding and flexible configuration of semantic information, and proposes a sorting robot complete equipment based on AI large model, which seamlessly integrates the powerful perception capability of AI large model into high-speed, real-time industrial sorting production lines to realize accurate joint identification of the position, category and material of the article, and drives the actuator to perform dynamic grabbing accordingly, forming a high-efficiency and reliable complete equipment.

[0009] To achieve the above object, the following technical scheme is proposed: A sorting robot complete equipment based on AI large model, comprising a complete sorting flow line, wherein the complete sorting flow line is provided with an industrial camera array, a central control system and a mechanical arm execution module, the central control system is electrically connected with a core computing module based on AI large model, the industrial camera array transmits a video stream to the core computing module through an image acquisition module, the conveyor belt of the complete sorting flow line is provided with an encoder electrically connected with an information processing unit of the central control system, the image acquisition module is synchronized with the encoder signal, the core computing module converts the video stream into an information stream including the position, category and material of the article by using AI large model, the central control system is provided with a man-machine interaction unit receiving sorting category semantics, the central control system is provided with a grabbing decision unit screening a decision result and an article position according to the sorting category semantics and the information stream, and a grabbing trajectory planning module of the mechanical arm execution module according to the decision result and the article position.

[0010] The complete sorting pipeline of the application is the basic carrier of the whole sorting equipment, and the conveying belt is used to transport the articles to be sorted. The industrial camera array is an important component for obtaining image information of the articles, which includes industrial cameras located above and on the side of the conveying belt. The application realizes the synchronous and high-precision identification of the category of the articles on the conveying belt, including complex semantic categories such as "old clothes with zippers", "PET plastic bottles with printed characters on the surface", material such as plastic, metal, glass, paper, wood, and accurate three-dimensional position (X, Y, Z coordinates). The application allows users to flexibly check specific categories that need to be grabbed through a human-computer interaction interface, realizing a highly customizable dynamic sorting strategy. The application deeply integrates the powerful perception and understanding capabilities of AI large models with the real-time control of industrial robots, forming a high-integration, fast-response and flexible intelligent sorting solution. The application can reduce the dependence on massive labeled data of specific articles and improve the identification capability of the system for unknown or new category articles.

[0011] As a preferred, the core computing module is electrically connected with a physical signal auxiliary verification module, and the physical signal auxiliary verification module includes a millimeter wave radar, an infrared camera and a split-focus plane polarization camera.

[0012] As a preferred, the detection data of the physical signal auxiliary verification module is fused with the visible light image obtained by the industrial camera through the central control system to form a comprehensive determination result of the material property of the article.

[0013] As a preferred, the industrial camera array is provided with at least one industrial camera above and on the side of the conveying belt, and the industrial camera is provided with an industrial light source. The industrial camera located above the conveying belt is a 20 million pixel and above global shutter industrial camera with a resolution of ≥5120x3840, and a polarization lens and a coaxial light source are additionally provided. The industrial camera located on the side of the conveying belt is a 2 million pixel 3D structured light camera with a depth map resolution of 1280x720 and a frame rate of 30 fps, and the Z-axis measurement accuracy is ±0.5 mm. The industrial camera located above the conveying belt in the application can be a 20 million pixel and above global shutter industrial camera with high resolution, and a polarization lens and a coaxial light source are additionally provided, which can more clearly capture the top image of the article and obtain rich visual information. The global shutter industrial camera here can be replaced by other high-resolution and high-frame-rate industrial cameras as long as the image acquisition requirements are met. The industrial camera located on the side of the conveying belt can be a 2 million pixel 3D structured light camera, which can be replaced by other types of 3D cameras, such as time-of-flight (ToF) cameras, to realize accurate acquisition of three-dimensional information of the article. The industrial camera is provided with an industrial light source, such as a ring light source and a strip light source, for providing sufficient and uniform illumination to ensure the quality of image acquisition.

[0014] As preferred, the core computing module comprises a visual-semantic understanding engine and a target detection and segmentation unit, the visual-semantic understanding engine is constructed based on a pre-trained AI multi-modal large model, and the target detection and segmentation unit performs instance segmentation and target detection and multi-attribute recognition on an input image.

[0015] As preferred, the mechanical arm execution module adopts a high-speed and high-precision industrial robot arm including a SCARA or a six-axis articulated arm, which is installed on the side or above the conveying belt; and an end effector is connected to the tail end of the high-speed and high-precision industrial robot arm.

[0016] As preferred, the end effector comprises at least two kinds of end clamping devices, and the central control system automatically switches the clamping device type according to the material attribute output by the core computing module, including but not limited to a pneumatic suction cup, a flexible clamping jaw or a magnetic adsorption mechanism.

[0017] The end effector (clamp) of the present application can select a pneumatic suction cup, a two-finger / three-finger electric flexible clamping jaw, a magnetic clamp, a special-purpose clamping device, etc. according to the diversity of the materials to be sorted, such as size, weight, shape, material properties: fragile / hard / soft / slippery.

[0018] After receiving the planned trajectory and grasping instructions from the central control system, the mechanical arm accurately moves to the position and implements the grasping action to pick up the target object from the conveying belt.

[0019] As preferred, the grasping decision unit comprises a trajectory optimization subunit, which adjusts the motion parameters of the mechanical arm execution module according to the material attribute of the object output by the core computing module, including but not limited to a speed curve, an acceleration threshold or a compliance control parameter.

[0020] As preferred, the complete sorting pipeline is provided with a blanking frame management system and a communication network, the blanking frame management system comprises a plurality of blanking frames placed near the working area of the mechanical arm execution module, and the communication network is used for information transmission between modules.

[0021] As preferred, the central control system comprises a dynamic mapping module, the blanking frame management system and the classification label defined on the human-computer interaction unit establish an editable corresponding relationship, supporting real-time adjustment of the target category and the container allocation strategy by the user.

[0022] The present application has the following advantages: 1. The powerful cross-modal understanding ability of the multi-modal large model can simultaneously and accurately identify the category (rich semantics), material and accurate position of the object, far exceeding the traditional visual system. It has "zero sample / less sample" learning ability for unknown categories and new materials, greatly improving the system applicability and scalability.

[0023] 2、User can check any combination of target categories through a simple interface, realize the flexibility of sorting "what you see is what you grab", and meet the needs of multi-variety, small-batch customized production or complex sorting tasks.

[0024] 3, Combined with high-speed visual capture, efficient model reasoning (which can use model compression and hardware acceleration technology optimization) and accurate motion control, it can effectively improve the sorting speed and accuracy in complex environment. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 The system configuration diagram of the present application. DETAILED DESCRIPTION

[0026] Embodiment: The embodiment provides a sorting robot complete equipment based on AI large model, referring to Figure 1 , including a complete sorting pipeline, the complete sorting pipeline is provided with an industrial camera array, a central control system and a mechanical arm execution module, the central control system is electrically connected with a core computing module based on an AI large model, the industrial camera array transmits a video stream to the core computing module through an image acquisition module, an encoder of the central control system is electrically connected with an information processing unit of the central control system, the image acquisition module is synchronous with the encoder, the core computing module converts the video stream into an information stream including the position, category and material of the article by using an AI large model, the central control system is provided with a man-machine interaction unit for receiving a sorting category semantic, the central control system is provided with a grabbing decision unit for screening a decision result and an article position according to the sorting category semantic and the information stream, and a grabbing trajectory planning module for planning a grabbing trajectory of the mechanical arm execution module according to the decision result and the article position.

[0027] The conveyor belt is used for carrying the articles to be sorted and moving at a uniform or variable speed, and is provided with an encoder for acquiring position information such as displacement of the conveyor belt in real time, providing a reference for image triggering and article position tracking. The encoder adopts 17-bit absolute value type (131072 PPR), and sends a position signal through an EtherCAT interface. The industrial camera is powered by PoE+ and receives the encoder pulse, and the triggering delay is less than 100us. If there are multiple industrial cameras, the multiple industrial cameras realize time synchronization of us level through PTP protocol.

[0028] The core computing module includes a visual-semantic understanding engine and a target detection and segmentation unit, the visual-semantic understanding engine is constructed based on a pre-trained AI multi-modal large model, and the target detection and segmentation unit performs instance segmentation, target detection and multi-attribute recognition on the input image.

[0029] The hardware architecture of the AI multimodal large model adopts a multimodal large model server or an edge computing unit, namely, deploying high-performance computing hardware (GPU / TPU), and the software architecture of the AI multimodal large model includes a visual-semantic understanding engine and a target detection and segmentation unit. The AI multimodal large model includes an improved and optimized CLIP-ViT model, a VLMo model or a variant model of the two. The visual-semantic understanding engine receives a multi-angle image sequence, namely, a video stream, acquired by an image acquisition module. The target detection and segmentation unit: Instance segmentation and target detection: accurately positioning the bounding box (Bounding Box) or pixel-level mask (Mask) of each item in the image, and identifying the independent object. Combined with the position information of the conveyor belt, the real-time three-dimensional position (X, Y, Z) of the item in the conveyor belt coordinate system is calculated.

[0030] Multi-attribute recognition: simultaneously recognizing the class (Class) of each segmented and detected item: based on the powerful multimodal understanding ability of the large model, the specific semantic class of the item is recognized, which is not limited to a small number of predefined classes, and complex semantic descriptions can be understood; material (Material): based on the understanding of the image texture, reflection characteristics and the like of the large model, combined with possible context inference, the main material type of the item is recognized.

[0031] The present application has the ability of zero-shot or few-shot training, and the AI multimodal large model naturally has the ability to understand unseen classes. Users can significantly improve the recognition accuracy of new classes / special classes by setting natural language descriptions or providing a small number of example images on the console without the need for large-scale retraining of the model.

[0032] The central control system of the present application adopts a PLC or an industrial PC as the "brain" of the entire device, which includes: an information processing unit that receives the item position, class and material information stream from the AI large model; a human-machine interaction unit (HMI) that provides a graphical interface on which an operator can: real-time view the conveyor belt video stream and the item recognition results such as labeled boxes, class or material labels, flexibly check the target classes that need to be sorted, and the human-machine interaction unit can support multiple selection, conditional combination selection, such as "plastic bottle and green color" and "paper box with surface damage". A grasping decision unit that quickly filters and judges the items on the conveyor belt that reach the reachable area of the robot arm according to the currently checked target class list, and transmits the position coordinates, size, properties and other information of the items belonging to the target class to the motion planning unit. A motion planning unit that plans the optimal and collision-free grasping trajectory of the end effector (gripper) of the robot arm based on the grasping decision result and the position of the item, taking into account the material, size, center of gravity and other information of the item, such as planning a compliant grasping strategy for fragile items.

[0033] The application realizes the synchronous and high-precision identification of the category of the articles on the conveying belt, including complex semantic categories such as "old clothes with zippers", "PET plastic bottles with printed fonts on the surface", materials such as plastic, metal, glass, paper, wood and accurate three-dimensional positions (X, Y, Z coordinates). The application allows users to flexibly check specific categories that need to be grabbed through a human-computer interaction interface, realizing a highly customizable dynamic sorting strategy. The application deeply integrates the powerful perception and understanding capabilities of AI large models with the real-time control of industrial robots, forming a set of intelligent sorting solutions with high integration, fast response and strong flexibility. The application can reduce the dependence on massive labeled data of specific articles and improve the identification capability of the system for unknown or new category articles.

[0034] The core computing module is electrically connected with a physical signal auxiliary verification module, and the physical signal auxiliary verification module includes a millimeter wave radar, an infrared camera and a split-focus plane polarization camera.

[0035] The detection data of the physical signal auxiliary verification module is fused with the visible light image obtained by the industrial camera through the central control system, to form a comprehensive judgment result of the material property of the article.

[0036] The visual-semantic understanding engine realizes physical property fusion through a multi-modal semantic understanding engine, and the technical path is as follows: S1, cross-modal physical property modeling: constructing a material semantic space Vectorizing material science parameters: Among them: Indicates that the double-albedo respective function is learned through cross-modal alignment; f conductivity Indicates the electrical conductivity parameter (metal >10 6 S / m, plastic <10 -12 S / m); f thermal Indicates the conductivity parameter (glass ≈1.0W / mK, acrylic ≈0.2W / mK); v material is the vectorized representation of the material in the semantic space, which is a 768-dimensional mathematical vector, that is, Material semantic space Among them, through the vector, the physical properties of the material, such as light reflection characteristics, electrical conductivity and thermal conductivity, can be uniformly coded as a calculable feature in a high-dimensional semantic space, supporting subsequent cross-modal learning or physical simulation; Mathematical model representing the generation of material semantic vector Vmaterial; its essence is to fuse and project the three types of physical properties of materials through a multi-layer perception (MLP), and finally output a high-dimensional vector representation, whose complete logic chain is as follows: S2, visual-physical property joint reasoning, material determination function realizes semantic level decision: Wherein: P(y m |I) represents the conditional probability, which represents the probability of the occurrence of variable y m under the premise of input image I, corresponding to the output probability in the prediction task; y m represents the target variable, representing a certain label, such as material category; I represents the input variable, representing image data; Softmax(W m ·f v (I)) represents the activation function, which converts the output of W m ·f v (I) into probability values; W m represents the weight matrix, which is used for linear transformation, and is used in the Softmax function to map visual features to probability distribution, the dimension depends on the space of y m ; f v (I) represents the feature function, which extracts the visual features of image I, and outputs a feature vector representing the image content; λ represents the hyperparameter, a scalar value, used to weight the physical property alignment term, controlling the contribution of this part to the overall probability; cos(v material ,v text ) represents the calculation of the cosine similarity of two vectors, the value is between -1 and 1, representing their semantic similarity; v material represents the material semantic vector, which is a high-dimensional vector (e.g. 768 dimensions), representing the embedding of the physical properties of the material in the semantic space; v text represents the text semantic vector, which is a high-dimensional vector, representing the embedding of the text description in the semantic space; The probability P(y m |I) is composed of two parts: prediction based on visual features (Softmax part) and adjustment based on physical property alignment (cosine similarity part); the physical property alignment term adds a bias, making the probability tend to be consistent with the material described in the text; the overall goal is to fuse visual and physical information for more robust prediction.

[0037] S3, multi-spectral evidence fusion, to break through the visible light limit, increase the physical signal auxiliary module, reference table 1, table 1 physical signal auxiliary module judgment basis comparison table Fusion decision function: Among them: the weight β dynamic allocation: visible light (β1=0.4), microwave (β2=0.3), infrared (β3=0.2), polarization (β4=0.1); y final Indicates the final decision class, through argmax from the candidate class to select the weighted probability and the maximum class as the final result of the system prediction, such as material type: metal, plastic, etc.; Indicates the maximum value corresponding to the parameter, through the expression in the parentheses to traverse all candidate y m , select the y m that makes the value maximum, if y m There are 3 categories (metal, plastic, ceramic), then calculate the weighted probability sum of each class, take the class with the highest sum as y final ; P(y m |D k ) represents the conditional probability, that is, the probability of the class being y m Under the data source D k ; y m Indicates the candidate class (such as material type); D k Indicates the kth detection data (D1=visible light, D2=microwave, D3=infrared, D4=polarization); Each data source has an independent classification model, such as CNN for visible light image classification, microwave sensor model, etc., output its corresponding probability distribution.

[0038] The technical effect comparison of the embodiment and the prior art is shown in table 2: Table 2 Technical effect comparison table In addition to the industrial camera array and physical signal verification module, other sensors such as ultrasonic sensors and lidar can also be added. Ultrasonic sensors can detect the distance and shape of objects, while lidar can obtain high-precision three-dimensional point cloud data of objects. Fusion processing of these sensor data with data from the industrial camera and physical signal verification module enables more comprehensive and accurate information about an object's location, shape, material, and so on. For example, when identifying objects with uneven surfaces or complex shapes, data from ultrasonic sensors and lidar can provide more detailed information, helping the core computing module more accurately determine the object's category and location. Furthermore, the information processing unit of the central control system can utilize more advanced data fusion algorithms, such as Kalman filters and particle filters, to fuse and process data from multiple sensors, improving the reliability and stability of the information.

[0039] The industrial camera array is equipped with at least one industrial camera above and on the side of the conveyor belt. The industrial cameras are equipped with industrial light sources. The industrial camera located above the conveyor belt is a 20-megapixel or above global shutter industrial camera with a resolution ≥5120×3840, and is equipped with an additional polarization lens and coaxial light source. The industrial camera located on the side of the conveyor belt is a 2-megapixel 3D structured light camera with a depth map resolution of 1280×720, 30fps; the Z-axis measurement accuracy is ±0.5mm.

[0040] The present invention includes at least one set of high-resolution industrial cameras, including color or RGB-D cameras, which are installed at key positions above or on the side of the conveyor belt to form a visual sensing array. The present invention is equipped with high-frequency industrial light sources, such as ring lights, strip lights, backlights, etc., to provide stable and uniform lighting and reduce the impact of reflections and shadows on image quality. The image acquisition module is synchronized with the conveyor belt encoder signal to ensure accurate triggering of image capture when the item passes through the preset shooting position. Preferably, a 20-megapixel and above global shutter industrial camera, such as the Sony IMX535 sensor solution, with a resolution ≥ 5120 × 3840, is used to ensure the ability to capture tiny features. For transparent / reflective objects, such as glass bottles, a polarized lens and a coaxial light source are added. A 2-megapixel 3D structured light camera is deployed on the side of the conveyor belt, such as The RealSense DepthCamera D455 offers a depth map resolution of 1280×720, 30fps, and a Z-axis measurement accuracy of ±0.5mm. The industrial light source utilizes an 850nm infrared backlight (for contour segmentation) and a 6500K color temperature LED strip frontlight (for surface texture analysis). The illumination is adjustable from 1000-5000 lux and supports strobe triggering (μs-level response). The lens uses a low-distortion industrial fixed-focus lens (focal length 12mm, F-value 2.8, distortion rate <0.1%).

[0041] The mechanical arm execution module adopts a high-speed and high-precision industrial robot arm including a SCARA or a six-axis articulated arm, which is installed on the side or above the conveying belt; and an end effector is connected to the tail end of the high-speed and high-precision industrial robot arm.

[0042] The end effector includes at least two end clamping devices, and the central control system automatically switches the type of the clamping device according to the material property output by the core computing module, including but not limited to a pneumatic suction cup, a flexible clamping jaw or a magnetic adsorption mechanism.

[0043] The end effector (clamp) of the present application can select a pneumatic suction cup, a two-finger / three-finger electric flexible clamping jaw, a magnetic clamp or a special clamping device according to the diversity of the materials to be sorted, such as size, weight, shape and material properties: fragile / hard / soft / slippery.

[0044] After receiving the planned trajectory and grasping instruction from the central control system, the mechanical arm accurately moves to the position and implements the grasping action to pick up the target object from the conveying belt.

[0045] The grasping decision unit includes a trajectory optimization subunit, which adjusts the motion parameters of the mechanical arm execution module according to the material property of the object output by the core computing module, including but not limited to a speed curve, an acceleration threshold or a compliance control parameter.

[0046] The complete sorting pipeline is provided with a blanking frame management system and a communication network, the blanking frame management system includes a plurality of blanking frames placed near the working area of the mechanical arm execution module, and the communication network is used for information transmission between modules.

[0047] The central control system includes a dynamic mapping module, the blanking frame management system and the classification label defined on the man-machine interaction unit establish an editable corresponding relationship, and the user can adjust the target category and the container allocation strategy in real time.

[0048] The central control system of the present application automatically guides the mechanical arm execution module to place the object into the corresponding blanking frame according to the category information of the object when the object is grasped, and the blanking frame management system can dynamically update the mapping relationship between the blanking frame and the category. The communication network adopts a high-speed and low-delay industrial Ethernet, such as Profinet, EtherCAT, Ethernet / IP or a special field bus, to ensure real-time and reliable transmission of image data, position signals, identification results and control instructions between modules.

[0049] The specific working process of one kind of sorting robot complete equipment based on an AI large model in the embodiment is as follows: Object conveying and triggering: the object moves at a constant speed with the conveying belt. The encoder monitors the displacement of the conveying belt in real time.

[0050] Image acquisition: When the item passes through the preset recognition station, the encoder signal triggers the high-resolution industrial camera (including a top-view camera and at least one side-view camera) to take high-quality images synchronously. The flash light is synchronized to provide supplementary lighting.

[0051] Multi-modal recognition and position calculation: The captured image stream is transmitted to the multi-modal large model server in the core computing module through a high-speed network.

[0052] Real-time image processing by visual-semantic understanding engine: Perform instance segmentation to identify individual items, identify specific categories (identify new categories using large model generalization ability), and identify material attributes (such as glass, hard plastic, soft plastic, metal, paper, fabric).

[0053] Target detection and segmentation unit combined with shooting angle and conveyor belt displacement (precise item passing trigger point timing calculated by encoder), through coordinate system conversion algorithm (such as image coordinates -> conveyor belt coordinates -> robot world coordinates), real-time calculation of the precise position (X, Y, Z) of each item in three-dimensional space, as well as size and orientation.

[0054] Recognition result transmission: The large model server outputs structured recognition results (item ID, position XYZ coordinates, recognized category label and confidence, recognized material label and confidence) and sends them to the information processing unit of the central control system through a high-speed network.

[0055] Human-machine interaction and decision-making: Operators can monitor the operation in real time (video stream + superimposed recognition results) on the HMI interface. According to the needs, select the target categories currently needed to be grabbed (multiple selection is available) on the HMI check interface.

[0056] Grasp planning and control: The grasp decision unit of the central control system continuously receives the recognition result stream. It compares each item entering the reachable area of the robot according to the target category list checked by the user: if the category of the item belongs to any target category in the checked list, it is marked as a "to-be-grabbed item". For the "to-be-grabbed" item, the motion planning unit receives the position, size, category, and material information of the item: a. Task allocation: Plan a reasonable grasp task for each "to-be-grabbed" item (considering item movement prediction and robot action time).

[0057] b. Trajectory planning: According to the 3D coordinates of the item and the characteristics of the gripper, plan the optimal collision-free motion trajectory of the robot end effector. For fragile items (such as recognized material "glass"), automatically introduce soft control parameters (force / position hybrid control).

[0058] c. Clamp selection control: if multiple clamps are configured, the most suitable clamp is automatically selected (or the parameters of the current clamp are adjusted) according to the properties of the object (material, size, shape).

[0059] Performing grasping: the central control system generates precise grasping instructions (target pose, path points, speed / acceleration parameters, clamp control signals) and sends them to the robot arm controller. The robot arm drives the end effector to strictly follow the planned trajectory and action sequence to perform the grasping action and stably take the target object from the conveyor belt.

[0060] Placing and sorting: the robot arm places the object accurately into the corresponding drop box according to the target category information of the object (mapped from the HMI check list). The control system manages the state of the drop box (such as capacity warning).

[0061] Real-time update and cycle: the entire process is continuously and real-time performed for the continuous objects on the conveyor belt, forming an intelligent assembly line operation.

[0062] This embodiment gives a detailed explanation of the visual-semantic understanding engine core technology: S11, multi-modal data labeling and training strategy: Data construction method: Weakly supervised labeling framework: Use a pre-trained large model (CLIP-ViT) to automatically generate pseudo labels. Given an image-text pair dataset: Where: represents the complete image-text pair dataset, the core research object is the basic data carrier for multi-modal machine learning tasks, and is the input source for pre-trained models such as CLIP-ViT, used for cross-modal feature alignment; I i represents the i-th image sample in the dataset, the digital carrier of visual information (pixel matrix or feature tensor), in the CLIP framework, as the input of ViT (Vision Transformer); T i represents the text description paired with the image Ii, which is the semantic annotation or label (descriptive language) of Ii, in the CLIP framework, as the input of the text encoder (such as Transformer), which is a variable-length language sequence (needs to be unified as a token sequence), together with Ii to form a cross-modal training pair; N represents the total number of samples in the dataset, defines the cardinality parameter of the dataset size, determines the value range of the index i (i∈{1,2,...,N}); Generate initial labels through cross-modal alignment loss: wherein: is the cross-modal alignment loss, used to measure the matching degree between visual-textual features, the smaller the value, the better the alignment of the image-text pair; i and j are index variables; i is used to traverse the current batch of image-text pairs (positive sample pairs), and j is used to traverse all possible text samples (including negative samples); the summation of j in the denominator covers the entire training set of texts, which is used for contrastive learning; f v , f t is a pre-trained visual / textual encoder, f v (I i ) is a visual encoder that extracts an image I i into a vector; f t (T j ) is a text encoder that extracts a text T j into a vector. τ is a temperature coefficient used to scale the similarity distribution, which controls the sharpness of the probability, and the probability distribution is sharper as т→0, and its typical value range is: 0.01~0.5 (which needs to be adjusted by experiments); sim() represents a similarity calculation function, which is used to calculate the cosine similarity between two vectors, and is usually defined as u T v / (||u|||v||).

[0063] Extract the object bounding box (use pre-trained Mask R-CNN) and material label (filter terms such as "metal" and "plastic" through text keywords) for the high-score matching pair (I i , T i ); the calculation logic is as follows: numerator: calculate the similarity of the positive sample pair (I and T), and strengthen the matching relationship; denominator: calculate the similarity of the current image I and all texts T (including negative samples), and suppress false matches; overall goal: minimize which is equivalent to maximizing the probability of the positive sample pair (the numerator after Softmax).

[0064] S12, active learning optimization: for difficult samples (such as confusing materials "reflective plastic vs. metal"), define an uncertainty measure: wherein: U(I) represents the classification uncertainty of the input image I, and U(I) is defined as 1 minus the highest probability value predicted by the model among all classes; when the model's prediction probability for the image I is very high, i.e. close to 1, U(I) is close to 0, indicating low uncertainty (high classification confidence), and when the prediction probability is very low, U(I) is close to 1, indicating high uncertainty (classification may be wrong or ambiguous). P(c|I; θ) represents the probability (probability value) of the image belonging to class c given the input image I and model parameters θ, the maximum value of the probability of all classes c is extracted, which reflects the prediction confidence of the model for the most likely class; I represents the input image (input image), that is, the input data sample of the model, usually refers to the image sample (pixel matrix or feature tensor), and the image needs to be preprocessed into a format that can be processed by the model (such as 224×224 resolution); c represents a class label, which belongs to a class set C, and in the probability term P(c|I; θ), c is the target prediction variable, which represents the output class of the model, and the value range covers all possible classes (such as 10 classes, c traverses 1 to 10); θ represents the parameters of the model, such as the weights and biases obtained by training, and in the probability term P(c|I; θ), θ is used as a conditional variable, which means that the probability is calculated based on these parameters (such as neural network weights), and θ is located in the lower right corner of the formula, emphasizing that it is a training parameter inside the model (rather than an input or output), and the parameters are learned and optimized through training data (such as gradient descent method); Select samples with U(I)>0.3 to be manually labeled, and the labeling cost is reduced by 72%.

[0065] S13, model training method: S131, multi-task joint training: the model output head contains three branches: S132, class semantic branch: output object class probability S133, material attribute branch: output material probability S134, position regression branch: output bounding box offset δx, δy, δw, δh; S135, the total loss function is: Wherein: Ltotal represents the total loss function, which is the optimization target based on multi-task joint training, and the joint learning of the following capabilities of the model is coordinated through the weighted sum of the three sub-loss terms; Lcls represents the focal loss function of object class recognition, which is an improved cross-entropy loss, which solves the problem of class imbalance such as rare objects and improves the classification accuracy of object classes; Lmat represents the cross-entropy loss function of material attribute classification, which is a standard classification loss, which measures the difference between probability distributions and optimizes the classification performance of material attributes such as metal / plastic; Smooth L1 loss function representing target box position regression, regression loss robust to outliers, replacing L2 loss, accurately fitting the position and size of the target box; delta represents the boundary box position offset predicted by the model (such as center coordinate offset + width and height scaling); b gt represents the position parameters of the real box (Ground Truth); y c represents the real class label of the object (such as one-hot vector [0, 1, 0]); y m represents the real label of the material (such as the index value of "metal"); P class represents the object class probability distribution output by the model (such as [0.1, 0.8, 0.1]); P material represents the material class probability distribution output by the model (such as metal probability 0.9); Class recognition task weight λ1=0.8, material classification task weight λ2=0.5, position regression task weight λ3=1.2, and material classification adopts a difficult sample mining strategy, giving weights to easily confused material pairs (such as glass / crystal): where: W m represents the weight coefficient of difficult material samples, and in the material classification task, the loss weight of easily confused materials is dynamically adjusted, such as weighting glass / crystal and other easily confused materials, and the greater the weight value, the more the model focuses on the difficulty of distinguishing the material category; S14, real-time inference engine design: S141, time domain feature aggregation: for the conveyor belt continuous 3 frames of images {I t-2 , I t-1 , I t}, the features are fused by optical flow field weighting: where: F fused represents the aggregated feature vector after time domain fusion, which is generated by weighted fusion of visual features of consecutive three frames (t-2, t-1, t), and contains spatiotemporal consistency information of the object in the motion process, and its dimension is the same as that of single frame f v (I k ); f v (I k ) represents the deep visual features of the k-th frame image, which are extracted from the original image I kMid-level feature extraction, encoding semantic information of single frame (e.g. object shape, texture, material), invariant to occlusion and illumination changes; I k Ik represents the k-th frame of the original image in the input sequence, whose data attribute is a three-channel RGB tensor (e.g. 224x 224x3); Flow k→t Ik represents the k-th frame of the original image in the input sequence, whose data attribute is a three-channel RGB tensor (e.g. 224x 224x3); The weight ωk=exp(-η|t-k|), η=0.5 controls the decay strength, and the Warp operation is a feature alignment based on optical flow.

[0066] S142, position estimation algorithm: Combine 2D bounding box and depth map Dt to calculate 3D position in world coordinate system: Where: K is the camera intrinsic parameter; R, t is the hand-eye calibration matrix, d=Dt(u,v) is the depth value at pixel point (u,v), and the 2D pixel point (u,v) is converted to world coordinates [x_w,y_w,z_w] through the depth value d; S143, confidence calibration mechanism: multi-modal confidence fusion: final confidence S final Ss is the semantic confidence, Sg is the geometric consistency, and S is the final confidence. sem Ss is the semantic confidence, Sg is the geometric consistency, and S is the final confidence. geo Ss is the semantic confidence, Sg is the geometric consistency, and S is the final confidence. s final =α·s sem +(1-α)·s geo ; Where: α represents the weight coefficient, the value range is [0,1], used to adjust the relative importance of semantic confidence S sem and geometric consistency S geo in the final decision; s sem = P class × P material The geometric consistency score is: Where: p pred represents; P track represents; σ is set to 1 / 3 of the object size (e.g. if a bottle with a diameter of 50mm is detected, then σ≈16.7mm).

[0067] S15, industrial scene optimization technology: S151, hardware-aware compression for ViT: Structural pruning: remove the contribution measurement head with contribution less than 5% in the multi-head attention; Among them: Imp h The importance score of the attention head h is calculated by quantifying it through the Frobenius norm. The larger the value, the more important the information contribution of the attention head in the model. The smaller the value, the weaker the impact of the head on the result (it can be pruned and removed). Represents the i-th weight matrix in the h-th attention head, where the value range of i is: corresponding to the four types of weight matrices of Query, Key, Value, and Projection in multi-head attention, that is, i∈{Q,K,V,P}, Indicates the weight of the query vector (Query); Indicates the weight of the calculated key vector (Key); Represents the weight of the calculated value vector (Value); Represents the projection layer weight (output fusion); F represents the Frobenius norm, which is defined as follows: The calculation principle is as follows: sum the squares of all elements of the matrix W and then take the square root, which is equivalent to the L2 norm after expanding the matrix into a vector; Quantization-aware fine-tuning: DoReFa quantization is used for the fully connected layer: Where: W quant Represents the quantized weight matrix. Its physical meaning is: the discrete value matrix of the original weight W after 8-bit linear uniform quantization, whose value range is compressed from floating point numbers to integer domain (0 to 255) Its core functions are: reducing model storage and computing overhead (memory usage is reduced to 1 / 4 of the original); hardware-friendly (suitable for integer arithmetic unit acceleration); round() means: rounding function, its calculation logic is: round the value in the brackets to the nearest integer approximation; for example: round(127.3) = 127; round(127.7) = 128 Its technical necessity: forcing continuous floating-point values ​​to be discretized into integers (the core of the quantization operation); ensuring that weights can be encoded in binary with a fixed bit width (such as 8-bit); W represents the original weight matrix, which has the following characteristics: full-precision floating-point parameters (such as FP32) trained by a deep neural network (such as ViT); its value range: [min(W), max(W)] (usually not 0~1 uniformly distributed); S152, inference pipeline optimization: three-stage pipeline is realized by CUDA Stream, and the inference delay is reduced from 45ms to 22ms. The optimization effect is shown in Table 3: Table 3 Model industrial scene optimization effect table The embodiment also gives a feasible solution of model architecture selection and deployment optimization strategy of AI large model, as follows: the model architecture selection of AI large model is as follows: Visual encoder: ViT-Large (Vision Transformer) or ResNet-152 is used as the backbone network, and an input resolution of 448x448 is supported. ViT uses a multi-head self-attention mechanism to extract global features, and ResNet optimizes deep gradient transmission through a residual structure.

[0068] Cross-modal fusion: a fusion decoder based on Transformer (12 layers, hidden layer dimension 768) is used to perform cross-attention calculation on image features and text embedding (from a pre-trained text encoder such as BERT), realizing visual-semantic alignment. The output dimension of the fusion layer is aligned with the visual feature map space, supporting pixel-level attribute prediction.

[0069] The deployment optimization strategy of AI large model is as follows: Model distillation: the original multi-modal large model (teacher model) is compressed into a lightweight student model. Task-specific distillation is used: under the premise of preserving the generalization ability of the model, the logits distribution of the teacher model on key tasks (material classification, target detection) is migrated to the student model (MobileNetV3+ small Transformer) using the KL divergence loss function. After compression, the model parameter quantity is reduced to 1 / 8, and the inference delay is reduced by 60%.

[0070] Quantization acceleration: INT8 quantization scheme is used: post-training quantization (PTQ): dynamic range calibration is performed on weights and activation values, and a TensorRT optimization engine is used. Quantization-aware training (QAT): simulate quantization errors during fine-tuning to improve low-precision model accuracy. It has been verified that INT8 quantization improves the inference speed of ViT-Large by 3.2 times, and the accuracy loss is less than 0.5%.

[0071] Edge deployment: NVIDIA Jetson AGX Orin (64GB) embedded TensorRT engine, utilizing TensorCore to perform INT8 convolution and matrix operations.

[0072] Cloud collaboration: For few-shot new class recognition requests, call the cloud A100 cluster (80GB memory) to run the FP16 precision model through gRPC, with a response delay of <300ms.

[0073] Operator optimization: Rewrite the self-attention calculation kernel using the Cutlass library, combined with the FlashAttention-2 algorithm, to reduce the Transformer inference memory usage by 45%. Through OpenVINO TM Deploy ResNet branches to utilize CPU AVX-512 instruction sets for parallel processing of preprocessing operations.

[0074] The advantages of the present invention are as follows: Strong and flexible recognition ability: Utilizing the powerful cross-modal understanding ability of multi-modal large models, it can simultaneously and accurately identify the category (rich semantics), material, and precise location of an item, far surpassing traditional vision systems. It has "zero-shot / few-shot" learning ability for unknown categories and new materials, greatly improving system applicability and scalability.

[0075] Highly customizable sorting strategy: Users can select any combination of target categories through a simple interface, achieving flexible sorting of "what you see is what you get", meeting the needs of multi-variety, small-batch customized production or complex sorting tasks.

[0076] High degree of intelligence: From perception (vision + large model understanding) to decision-making (dynamic filtering) to execution (trajectory planning + grasping), forming a closed loop, with strong system autonomy.

[0077] High integration, easy deployment and maintenance: As a complete equipment design, software and hardware integration is optimized, reducing the complexity of on-site debugging. Model updates usually only require remote fine-tuning or providing new descriptions, reducing the technical threshold and cost of later maintenance.

[0078] Efficiency and accuracy improvement: Combined with high-speed visual capture, efficient model inference (optimized using model compression and hardware acceleration techniques), and precise motion control, it can effectively improve sorting speed and accuracy in complex environments.

[0079] Widely applicable: It can be widely applied to waste material recycling and sorting, fine parts quality inspection and packaging, agricultural product (fruit and vegetable) grading, e-commerce logistics package sorting, manufacturing product line sorting, and other fields.

Claims

1. A complete set of sorting robot equipment based on AI large model, characterized by: It comprises a complete set of sorting lines, which are provided with an industrial camera array, a central control system and a robotic arm execution module. The central control system is electrically connected to a core computing module based on an AI large model. The industrial camera array transmits the video stream to the core computing module through an image acquisition module. The conveyor belt of the complete set of sorting lines is provided with an encoder electrically connected to the information processing unit of the central control system. The image acquisition module is synchronized with the encoder signal. The core computing module uses the AI ​​large model to convert the video stream into an information stream including the location, category and material of the item. The central control system is provided with a human-computer interaction unit for receiving the sorting category semantics. The central control system is provided with a grasping decision unit for screening out decision results and item locations according to the sorting category semantics and the information flow, and plans the grasping trajectory of the robotic arm execution module according to the decision results and the item location.

2. The AI ​​large model-based sorting robot complete set of equipment according to claim 1 is characterized in that: The core computing module is electrically connected to a physical signal auxiliary verification module, and the physical signal auxiliary verification module includes a millimeter wave radar, an infrared camera and a focal plane polarization camera.

3. The complete set of sorting robot equipment based on AI large model according to claim 2 is characterized in that: The detection data of the physical signal auxiliary verification module is weightedly fused with the visible light image acquired by the industrial camera through the central control system to form a comprehensive judgment result of the material properties of the object.

4. The AI ​​large model-based sorting robot complete set of equipment according to claim 2 is characterized in that: The industrial camera array is equipped with at least one industrial camera above and on the side of the conveyor belt. The industrial cameras are equipped with industrial light sources. The industrial camera located above the conveyor belt is a 20-megapixel or above global shutter industrial camera with a resolution ≥5120×3840, and is equipped with an additional polarization lens and coaxial light source. The industrial camera located on the side of the conveyor belt is a 2-megapixel 3D structured light camera with a depth map resolution of 1280×720, 30fps; the Z-axis measurement accuracy is ±0.5mm.

5. The AI ​​large model-based sorting robot complete set of equipment according to claim 2 is characterized in that: The core computing module includes a visual-semantic understanding engine and a target detection and segmentation unit. The visual-semantic understanding engine is built based on a pre-trained AI multimodal large model. The target detection and segmentation unit performs instance segmentation, target detection, and multi-attribute recognition on the input image.

6. The AI ​​large model-based sorting robot complete set of equipment according to claim 2 is characterized in that: The robotic arm execution module adopts a high-speed and high-precision industrial robot arm including SCARA or a six-axis articulated arm, which is installed on the side or above the conveyor belt; the end of the high-speed and high-precision industrial robot arm is connected to an end effector.

7. The AI ​​large model-based sorting robot complete set of equipment according to claim 6 is characterized in that: The end effector includes at least two end clamping devices, and the central control system automatically switches the type of clamping device according to the material properties output by the core computing module, including but not limited to a pneumatic suction cup, a flexible clamp or a magnetic adsorption mechanism.

8. The AI ​​large model-based sorting robot complete set of equipment according to claim 6 is characterized in that: The grasping decision unit includes a trajectory optimization subunit, which adjusts the motion parameters of the robotic arm execution module according to the material properties of the object output by the core computing module, including but not limited to the speed curve, acceleration threshold or compliance control parameters.

9. A complete set of sorting robot equipment based on AI large model according to any one of claims 1 to 8, characterized in that: The complete set of sorting production lines is equipped with a blanking frame management system and a communication network. The blanking frame management system includes several blanking frames placed near the working area of ​​the robot arm execution module, and the communication network is used for information transmission between the modules.

10. The AI ​​large model-based sorting robot complete set of equipment according to claim 9 is characterized in that: The central control system includes a dynamic mapping module, and the blanking frame management system establishes an editable correspondence with the classification labels defined on the human-computer interaction unit, supporting users to adjust target categories and container allocation strategies in real time.

Citation Information

Cited By

  • Miscellaneous plastic sorting robot complete equipment based on Ai large model

    CN121374913A

  • Ai large model-based mixed plastic sorting robot complete equipment

    CN121374913B

  • Multi-category industrial detection sorting system and method based on AI large model and robot and storage medium

    CN121639621A