Multi-modal image mechanical arm sorting method
By using a multimodal image acquisition system and composite calibration technology, a spatiotemporally consistent multimodal image fusion dataset is constructed, which solves the problem of low efficiency in traditional sorting methods and enables efficient and accurate sorting of recyclable resources.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGXI SHENGHE RESOURCES RECYCLING TECH CO LTD
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional manual sorting methods are inefficient, costly, and environmentally unfriendly. In real industrial scenarios, automated sorting systems based on a single visual modality are difficult to accurately match the features of the same object in multimodal images due to physical parallax and time delay, which affects the accuracy of sorting recycled resources.
Visible light, infrared thermal imaging, and depth images are acquired simultaneously through a multimodal image acquisition system. Offline joint calibration is performed using a composite calibration board. Combined with a hardware synchronous trigger signal and timestamp alignment mechanism, a spatiotemporally consistent multimodal image fusion dataset is constructed. Motion compensation is then performed, and fusion features are extracted to identify the category and location of target objects.
It improves the accuracy and efficiency of recycling sorting. Through precise alignment and feature fusion of multimodal data, it ensures that the robotic arm can accurately identify and grasp target objects.
Smart Images

Figure CN122032891A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, specifically to a multimodal image robotic arm sorting method. Background Technology
[0002] With the rapid development of the global resource recycling industry, the demand for sorting and processing of recycled resources (such as scrap metal, plastics, glass, and paper) has increased dramatically. Traditional manual sorting methods are inefficient, costly, and operate in harsh environments. Automated sorting systems based on a single visual modality, in actual industrial scenarios, simply stitch together or fuse information from different modalities. Due to independent triggering of each sensor, physical parallax, and time delays, the same object features in different modal images are difficult to accurately correspond spatially and temporally, resulting in feature misalignment. Furthermore, this misalignment is further amplified when the target moves, affecting the accuracy of subsequent information fusion and recognition, and consequently impacting the accuracy of recycled resource sorting.
[0003] To address these issues, we propose a multimodal image-based robotic arm sorting method. Summary of the Invention
[0004] The purpose of this invention is to provide a multimodal image robotic arm sorting method to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a multimodal image-based robotic arm sorting method, the method comprising the following steps: Multimodal image data of the target area is acquired synchronously through a multimodal image acquisition system. The multimodal image data includes at least visible light images, infrared thermal imaging images, and depth images. The multimodal image data is registered to construct a multimodal image fusion dataset with spatiotemporal consistency; Based on the multimodal image fusion dataset, the fusion features of the target are extracted by the multimodal feature extraction module; the category of the target object is identified according to the fusion features and its position and pose in three-dimensional space are calculated; Based on the identification and positioning results, control commands for the robotic arm are generated to drive the robotic arm to perform sorting operations.
[0006] Preferably, the step of registering the multimodal image data to construct a multimodal image fusion dataset with spatiotemporal consistency includes: Based on the spatial transformation relationship between the various modal sensors obtained in advance through the composite calibration plate, the visible light image, infrared thermal imaging image and depth image are transformed to the same coordinate system to form a spatially aligned primary fusion image group; By using a hardware synchronization trigger signal and timestamp alignment mechanism, it is ensured that each modal image in the primary fused image group corresponds to the same acquisition time; Motion compensation is performed on the spatially aligned primary fused image group in response to the movement of the target object on the conveyor belt. The multimodal image fusion dataset is constructed by associating the spatially and temporally registered and dynamically compensated multimodal image data with the corresponding acquisition environment parameters and target location information.
[0007] Preferably, the step of transforming the visible light image, infrared thermal imaging image, and depth image to the same coordinate system based on the spatial transformation relationship between the various modal sensors obtained in advance through composite calibration plate, to form a spatially aligned primary fused image group includes: The multimodal image acquisition system is jointly calibrated offline based on a composite calibration plate, the surface of which is simultaneously provided with optical marks, temperature control marks and depth marks; Based on the offline joint calibration results, spatial transformation parameters from the infrared camera coordinate system and depth sensor coordinate system to the visible light camera coordinate system or unified world coordinate system are calculated. During the online sorting process, coordinate transformation and pixel-level resampling are performed on the real-time acquired modal images based on the spatial transformation parameters.
[0008] Preferably, the step of ensuring that each modal image in the primary fused image group corresponds to the same acquisition time through the hardware synchronization trigger signal and timestamp alignment mechanism includes: A unified hardware synchronization triggering device is configured for the visible light camera, infrared thermal imaging camera and depth sensor. The main controller sends a synchronization pulse signal to trigger all sensors to perform image exposure and acquisition at the same physical moment. Assign a unique timestamp to each frame of image data acquired; Based on the synchronization pulse signal and timestamp, frame matching is performed on image streams from different sensors, and image frames corresponding to the same trigger period are combined into the primary fused image group.
[0009] Preferably, the step of performing motion compensation on the spatially aligned primary fused image group in response to the motion of the target object on the conveyor belt includes: Acquire real-time speed and direction information of the conveyor belt carrying the target object; Based on the motion speed, direction information, and the inherent exposure delay of each sensor, the displacement deviation of the same target at the moment of adjacent modal image acquisition is calculated; Based on the displacement deviation, pixel translation or affine transformation is performed on the corresponding modal images in the primary fused image group to compensate for the registration error.
[0010] Preferably, the step of associating the spatially and temporally registered and dynamically compensated multimodal image data with corresponding acquisition environment parameters and target location information to construct the multimodal image fusion dataset includes: Assign a unique identifier to each set of registered multimodal image data, and record the acquisition time, ambient temperature, ambient light intensity, conveyor belt speed, and sensor operating parameters. The registered visible light image, infrared image, depth image, and corresponding environmental parameters and identification information are encapsulated into structured data units. Multiple structured data units are organized according to time series or task batches to form a multimodal image fusion dataset.
[0011] Preferably, the step of extracting the fusion features of the target using the multimodal feature extraction module based on the multimodal image fusion dataset includes: The spatially aligned subsets of visible light images, infrared images, and depth images from the multimodal image fusion dataset are input into the visible light feature extraction unit, the infrared feature extraction unit, and the depth feature extraction unit, respectively, to extract the primary feature representation of the target in each modality in parallel. Based on the spatiotemporal consistency information recorded in the multimodal image fusion dataset, a one-to-one correspondence between the visible light primary features, infrared primary features, and depth primary features is established at the pixel level or region level. The reliability of each modal primary feature is evaluated based on the current scenario conditions, and dynamic fusion weights are assigned to each modal primary feature based on the evaluation results. The weighted modal primary features are then spliced together or aggregated element by element to generate the target fusion feature.
[0012] Preferably, the step of identifying the category of the target object based on the fused features and calculating its position and orientation in three-dimensional space includes: The fused features are simultaneously input into a jointly learned target classification network and a 3D localization regression network; The target classification network outputs the material category, structure category, and physical attribute label of the target object based on the fusion features. The 3D localization regression network regresses the 3D bounding box of the target object in a predetermined coordinate system based on the fused features and the corresponding depth information. The 3D bounding box is defined by its center point coordinates, size and orientation angle. Based on the target object's category, physical attribute labels, and 3D bounding box, and according to a preset grasping strategy rule base, the optimal grasping point coordinates and the approach posture of the robotic arm's end effector are determined.
[0013] Compared with the prior art, the beneficial effects of the present invention are: A hardware-level synchronous triggering mechanism ensures that visible light, infrared, and depth sensors acquire data at the same physical moment. Offline joint calibration is performed using a composite calibration board to obtain accurate spatial transformation parameters. During online operation, pixel-level spatial alignment is achieved through coordinate transformation and resampling, and motion compensation is performed to dynamically eliminate registration errors caused by target motion, thus constructing a spatiotemporally consistent multimodal image fusion dataset. By setting up a triple mechanism of sensor synchronization, offline calibration, and dynamic motion compensation, the alignment of multimodal data in time and space is improved, thereby enhancing the accuracy of subsequent identification and positioning of regenerated resources and ultimately improving the accuracy of regenerated resource sorting. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example
[0017] Please see Figure 1 This invention provides a multimodal image robotic arm sorting method: a multimodal image robotic arm sorting method, comprising the following steps: S1: Simultaneously acquire multimodal image data of the target area through a multimodal image acquisition system. The multimodal image data includes at least visible light images, infrared thermal imaging images, and depth images. It should be noted that the multimodal image acquisition system includes: at least one visible light camera for acquiring color or grayscale images of the target area; at least one infrared thermal imaging camera for acquiring thermal radiation distribution images of the target object; at least one depth sensor for acquiring three-dimensional point cloud data or depth maps of the target area; a synchronous triggering device for ensuring that the multimodal image data are acquired synchronously in time; the optical axes of the visible light camera and the infrared thermal imaging camera are aligned with the detection center of the depth sensor by means of a beam splitter prism or close proximity mounting, so that the center of their fields of view is highly coincident with the main detection area in physical space to minimize parallax; multiple sensor nodes are arranged above and to the side of the conveyor belt or working area, each node containing a set of tightly coupled sensor units (visible light, infrared, depth) to achieve multi-angle coverage of the target object; Specifically, the main control unit sends a unified hardware synchronization trigger pulse signal to all sensors in the multimodal image acquisition system; after receiving the hardware synchronization trigger pulse signal, the visible light camera, infrared thermal imaging camera, and depth sensor start image exposure and data acquisition at the same physical moment; the visible light image frame, infrared thermal imaging image frame, and depth image frame acquired at the same physical moment are marked with the same timestamp information; the image data of each modality with the same timestamp, together with the sensor operating parameters at the time of acquisition, are packaged into a raw data packet with spatiotemporal correlation and output to the subsequent processing unit; S2: Register multimodal image data to construct a multimodal image fusion dataset with spatiotemporal consistency; The steps for registering multimodal image data and constructing a spatiotemporally consistent multimodal image fusion dataset include: based on the spatial transformation relationship between each modal sensor obtained through pre-calibration using a composite calibration plate, transforming visible light images, infrared thermal imaging images, and depth images to the same coordinate system to form a spatially aligned primary fusion image group; ensuring that each modal image in the primary fusion image group corresponds to the same acquisition time through a hardware synchronous trigger signal and timestamp alignment mechanism; performing motion compensation on the spatially aligned primary fusion image group for the motion of the target object on the conveyor belt to ensure the alignment accuracy of the multimodal images under motion conditions; and associating the multimodal image data after spatial and temporal registration and dynamic compensation with the corresponding acquisition environment parameters and target location information to construct a multimodal image fusion dataset. Based on the spatial transformation relationships between various modal sensors obtained through pre-calibration using a composite calibration plate, the steps of transforming visible light images, infrared thermal imaging images, and depth images to the same coordinate system to form a spatially aligned primary fused image group include: offline joint calibration of the multimodal image acquisition system based on the composite calibration plate, the surface of which is simultaneously provided with optical marks for visible light imaging, temperature control marks with specific thermal radiation characteristics, and depth marks with three-dimensional geometric structures; based on the offline joint calibration results, calculating the spatial transformation parameters from the infrared camera coordinate system and the depth sensor coordinate system to the visible light camera coordinate system or the unified world coordinate system; and during the online sorting process, performing coordinate transformation and pixel-level resampling on the real-time acquired modal images according to the spatial transformation parameters to achieve rigid spatial alignment. Specifically, upon starting or entering online sorting mode, pre-calibrated and stored spatial transformation parameters are loaded from memory. The validity of the spatial transformation parameters under the current hardware state is quickly verified by placing a calibration board or a verification object of known geometry in the field of view. For each frame of infrared and depth images acquired in real time, the theoretical projection coordinates of each pixel in the infrared image and each effective 3D measurement point in the depth image are calculated in the target coordinate system (visible light camera coordinate system or unified world coordinate system) based on the loaded spatial transformation parameters. A regular resampling grid is generated according to the imaging plane (or preset virtual imaging plane) parameters of the target coordinate system. A lookup table is established from the pixel position of the original infrared / depth image to the pixel position of the target image based on the mapping relationship between the theoretical projection coordinates and the resampling grid. The original infrared and depth images are resampled using the lookup table. For the infrared image, the temperature or grayscale value at the grid point of the target image is calculated. For the depth image, the original depth value or 3D coordinates are assigned to the grid point of the target image to generate a registered infrared image and a registered depth image spatially aligned with the visible light image. It should be noted that the spatial transformation parameters specifically include: a rigid transformation matrix from the infrared camera coordinate system to the target coordinate system, containing a 3x3 rotation matrix R_ir and a 3x1 translation vector t_ir; a rigid transformation matrix from the depth sensor coordinate system to the target coordinate system, containing a 3x3 rotation matrix R_depth and a 3x1 translation vector t_depth; and an intrinsic parameter matrix K_target of the virtual imaging plane in the target coordinate system, used to define the projection relationship from three-dimensional space to the two-dimensional image plane. The online coordinate transformation steps are as follows: For the infrared image: the (u_ir, v_ir) pixel in the infrared image is back-projected onto a normalized three-dimensional direction vector in the infrared camera coordinate system using the intrinsic parameter matrix K_ir of the infrared camera and the distortion coefficients; then, through the transformation matrix (R_ir, t_i... The r) is transformed to the target coordinate system, and finally projected onto the target image plane using the target coordinate system intrinsic parameter matrix K_target to obtain its theoretical projection coordinates (u_target, v_target); For the depth image: the (u_d, v_d) pixel in the depth image and its measured depth value z_d are reconstructed into the three-dimensional coordinates (X_d, Y_d, Z_d) of the point in the depth sensor coordinate system using the depth sensor intrinsic parameter matrix K_depth; then transformed to the target coordinate system using the transformation matrix (R_depth, t_depth) to obtain (X_t, Y_t, Z_t), and finally projected onto the target image plane using K_target to obtain its theoretical projection coordinates (u_target, v_target), and Z_t is used as the depth value of the point in the target coordinate system.
[0018] The steps to ensure that each modal image in the primary fused image group corresponds to the same acquisition time through the hardware synchronization trigger signal and timestamp alignment mechanism include: configuring a unified hardware synchronization trigger device for the visible light camera, infrared thermal imaging camera and depth sensor; the main controller sends a synchronization pulse signal to trigger all sensors to perform image exposure and acquisition at the same physical moment; assigning a unified timestamp to each frame of acquired image data; and performing frame matching on the image streams from different sensors based on the synchronization pulse signal and timestamp, combining image frames corresponding to the same trigger period into a primary fused image group. For the motion of a target object on a conveyor belt, the steps for motion compensation of a spatially aligned primary fused image group include: acquiring the real-time motion speed and direction information of the conveyor belt carrying the target object; calculating the displacement deviation of the same target at the moment of acquisition of adjacent modal images based on the motion speed, direction information, and the inherent exposure delay of each sensor; and performing pixel translation or affine transformation on the corresponding modal images in the primary fused image group based on the displacement deviation to compensate for the registration error. The steps for constructing a multimodal image fusion dataset by associating spatially and temporally registered and dynamically compensated multimodal image data with corresponding acquisition environment parameters and target location information include: assigning a unique identifier to each set of registered multimodal image data and recording the acquisition time, ambient temperature, ambient light intensity, conveyor belt speed, and sensor operating parameters; encapsulating the registered visible light image, infrared image, depth image, and corresponding environmental parameters and identifier information into structured data units; and organizing multiple structured data units according to time series or task batches to form a multimodal image fusion dataset. It also includes registration optimization steps for complex stacked and occluded scenes: segmenting the depth image to identify target regions at different height levels in the stacked scene; extracting the corresponding regions in the visible light and infrared images for target regions at different height levels; independently registering each layer region and calculating the relative spatial transformation between layers to correct multimodal feature misalignment caused by height differences and occlusion. Specifically, based on the depth image, the stacked target region is divided into N discrete height layers along the depth direction, where N≥2. Each layer contains a cluster of 3D point clouds with depth values within a specific range. For each height layer, based on its 3D point cloud projection range in the target coordinate system, corresponding 2D image regions are extracted from the visible light image, infrared image, and the pre-registered registered infrared and registered depth images, forming a set of local multimodal image patches associated with each layer. For each height layer, using its local depth patch in the registered depth image as a spatial reference, its corresponding local visible light image patch and local infrared image patch are independently finely registered to this reference. This fine registration includes... Calculate the local affine transformation parameters of the visible light / infrared image patch relative to the depth patch in the two-dimensional image plane at this level; based on the spatial positional relationship of each height level in the original depth image, and the local transformation parameters of each level obtained through the level-independent feature registration step, calculate the relative spatial transformation relationship between adjacent levels or any two levels; use the relative spatial transformation relationship between levels to perform position alignment correction on the multimodal features of the same object spanning different levels (such as the visible light edge and infrared hot spot of a metal can) caused by height difference and perspective difference; or perform unified spatial relationship correction on the feature positions of different objects on different levels, and finally generate a multimodal fusion feature map with higher spatial consistency in stacked scenes. It should be noted that the layer depth interval segmentation step adopts a multi-scale adaptive segmentation strategy: First, a rough histogram analysis is performed on the depth map of the entire target area to identify the main depth peaks and preliminarily determine the possible number of layers N and the approximate depth intervals; within each preliminarily determined depth interval, spatially continuous and depth-similar point clouds are clustered into the same physical object or part of the object's surface. Each height layer may contain one or more spatially separated but depth-similar object clusters. The method for obtaining local affine transformation parameters is as follows: In the local depth patch of each layer, multiple feature points with three-dimensional geometric features (such as corner points and edge intersections) are automatically selected; in the local visible light / infrared image patch of the corresponding layer, two-dimensional image feature points corresponding to the three-dimensional geometric feature points are found; based on at least three sets of corresponding two-dimensional-three-dimensional or two-dimensional-two-dimensional feature point pairs, the affine transformation matrix required to align the local visible light / infrared image patch of the layer to its local depth patch is directly calculated using a pre-built least squares solver. The specific steps for calculating the transformation relationship between layers are as follows: Define the coordinate system of the bottom reference layer (usually the layer with the largest depth value and closest to the conveyor belt) as the layer reference system; for other i-th layers, firstly, through the 3D point cloud registration between its local depth patch and the local depth patch of the bottom reference layer (such as using a hardware acceleration unit based on the ICP principle), calculate the rigid transformation T_i from the coordinate system of the i-th layer to the coordinate system of the bottom reference layer; combine the local affine transformation parameter A_i obtained by the i-th layer itself through the layer-independent feature registration step, and the local affine transformation parameter A_ref of the bottom reference layer, derive the composite transformation relationship used to correct feature misalignment, which maps the image feature position of the i-th layer to the image feature position of the bottom reference layer.
[0019] It also includes registration adaptability enhancement steps for multiple materials: In the calibration stage, a composite calibration sample set containing various typical recycled material materials such as metal, plastic, glass, and paper is used for calibration to obtain robust calibration parameters under the surface reflectivity of different materials; In the online stage, when a specific material (such as highly reflective metal or transparent glass) is detected in the current scene, it automatically switches to a subset of dedicated calibration parameters for that type of material for spatial registration; The temperature uniformity of the target material is detected by infrared image detection, and when infrared image feature distortion caused by the thermal characteristics of the material is detected, the registration confidence of that area is downweighted or re-registration is triggered; Specifically, before online sorting, the system needs to be jointly calibrated offline using a composite calibration board. This calibration board includes: an optical marking layer printed with a high-contrast checkerboard or dot array for visible light camera calibration; a temperature-controlled marking layer embedded with tiny resistive units that can be heated according to a preset pattern, forming a clear hotspot array in the infrared image; and a depth marking layer with precisely machined recessed or raised patterns (such as conical holes or columnar bodies) to provide rich three-dimensional features for the depth sensor. The calibration board is placed at different positions and angles on the conveyor belt, triggering the system to acquire multiple sets of synchronous multimodal images. By processing these images, the following parameters can be accurately calculated: the intrinsic parameters (intrinsic parameters, distortion) of each sensor; the rigid transformation matrix M_ir2vis from the infrared camera coordinate system to the visible light camera coordinate system; and the rigid transformation matrix M_depth2vis from the depth sensor coordinate system to the visible light camera coordinate system. These parameters are stored in memory.
[0020] During online sorting, for each set of synchronously acquired raw image data, the following process is executed to construct a fused dataset: First, pre-stored calibration parameters are loaded. Periodically (e.g., per shift) or when an anomaly is detected during self-inspection, a simplified verification object (e.g., a plastic block with feature blocks) can be placed in the field of view to quickly acquire a set of images and verify whether the registration error under the current parameters is within the threshold (e.g., pixel error < 2). Then, spatial registration is performed, specifically: For infrared images, based on the transformation matrix M_ir2vis and the camera model, the theoretical projection coordinates (u, v) of each pixel on the visible light camera imaging plane are calculated. For depth images, each pixel is first combined with its depth value and back-projected into three-dimensional space (depth sensor coordinate system), then transformed to a three-dimensional point in the visible light camera coordinate system using M_depth2vis, and finally projected onto the visible light imaging plane. Based on the imaging parameters of the visible light camera, a virtual, regular target image grid is generated. Based on the aforementioned projection relationships, a lookup table (LUT) is pre-calculated. This table records which position (potentially a subpixel) in the original infrared / depth image should be retrieved for each pixel on the target grid. Bilinear interpolation is used for the infrared image to ensure smooth temperature / grayscale transitions; the nearest neighbor method is used for the depth image to preserve the accuracy of the original measurements. The final output consists of a registered infrared image and a registered depth image aligned with the visible light image pixels. These three images together constitute a "primary fused image set." Due to the use of hardware synchronization triggering and unified timestamps, time alignment is completed during data acquisition, primarily ensuring frame index alignment of the data stream.
[0021] The specific details of dynamic motion compensation are as follows: Despite trigger synchronization, there is a minute, nanosecond-level delay difference between sensor exposure and readout. For a high-speed conveyor belt (e.g., speed V), this can cause a tiny displacement (Δd = V * Δt) of the same object in different modal images. The pulse signal from the conveyor belt encoder is read in real time to accurately obtain the instantaneous linear velocity V and direction. Through sensor calibration, the fixed exposure delay differences Δt_ir-vis and Δt_depth-vis relative to the trigger signal for each modality are known. The displacement compensation amounts are calculated as: Δd_ir = V * Δt_ir-vis, Δd_depth = V * Δt_depth-vis. For the registered infrared and registered depth images in the "primary fusion image group," corresponding reverse pixel shifts (-Δd_ir, -Δd_depth) are performed respectively, thereby eliminating the motion-induced misalignment at the image level.
[0022] The visible light image, infrared image, and depth image, after the above processing and alignment, are treated as a single data unit. A unique ID is generated for this unit, and the following metadata is associated and recorded: time information, environmental information (ambient light intensity and temperature at the moment of acquisition), motion information (instantaneous speed of the conveyor belt), the gain and exposure time of each sensor at that time, and the target location bounding box. The image data and metadata are encapsulated into a structured data unit. The continuously generated data units are appended to the data sequence in chronological order, thus forming a multimodal image fusion dataset with spatiotemporal consistency. By setting up a triple mechanism of hardware synchronization, offline calibration, and dynamic motion compensation, the alignment of multimodal data in time and space is improved, thereby enhancing the accuracy of subsequent identification and location of regenerated resources. S3: Based on a multimodal image fusion dataset, the fusion features of the target are extracted through a multimodal feature extraction module; the category of the target object is identified based on the fusion features, and its position and pose in three-dimensional space are calculated; Based on a multimodal image fusion dataset, the steps for extracting the fusion features of a target using a multimodal feature extraction module include: inputting spatially aligned subsets of visible light, infrared, and depth images from the multimodal image fusion dataset into the visible light feature extraction unit, infrared feature extraction unit, and depth feature extraction unit, respectively, and extracting primary feature representations of the target in each modality in parallel. Specifically, the visible light feature extraction unit is configured to extract the target's color histogram, edge gradient, texture statistics, and shape contour features; the infrared feature extraction unit is configured to extract the target's thermal radiation intensity distribution, thermal gradient direction, and temperature uniformity features relative to the background; and the depth feature extraction unit is configured to extract the target's 3D point cloud normal vector, curvature, height distribution, and volume features. Based on the spatiotemporal consistency information recorded in the multimodal image fusion dataset, a one-to-one correspondence between the visible light primary features, infrared primary features, and depth primary features is established at the pixel level or region level. The reliability of each modality's primary features is evaluated according to the current scene conditions, and dynamic fusion weights are assigned to each modality's primary features based on the evaluation results. The weighted primary features are then concatenated to generate the target's fusion features. The steps for evaluating the reliability of primary features of each modality based on the current scene conditions are as follows: For visible light reliability factor calculation: calculation is performed based on the average brightness, contrast, and presence of overexposed or underexposed areas in the current visible light image; For infrared reliability factor calculation: calculation is performed based on the average temperature difference between the target and the background, the infrared image signal-to-noise ratio, and thermal distribution stability; For depth reliability factor calculation: calculation is performed based on the density, integrity, and measurement noise level of the point cloud in the target area; Fusion weight allocation: the reliability factor of each modality is normalized and used as the weight of the primary feature of that modality during fusion. The fusion weights are adjusted accordingly: when the target material is identified as highly reflective metal, the fusion weight of visible light features is reduced, while the fusion weights of infrared and depth features are increased; when the target material is identified as transparent or translucent glass / plastic, the fusion weight of infrared features is significantly increased, taking advantage of its sensitivity to temperature distribution; when the target is identified as severely occluded or stacked, the fusion weight of depth features is significantly increased, utilizing its 3D structural information; for example, when a highly reflective area is detected, the weight of the visible light branch is reduced, while the weights of the infrared and depth branches are increased. The network ultimately outputs the category of each identified target (e.g., "PET bottle", "aluminum can", "transparent glass"), a 3D bounding box (center coordinates (x, y, z), dimensions (w, h, d), orientation angle), and a recommended grab point.
[0023] The steps for identifying the category of a target object and calculating its position and orientation in 3D space based on fused features include: simultaneously inputting the fused features into a jointly learned target classification network and a 3D localization and regression network; the target classification network outputting the material category, structure category, and physical attribute label of the target object based on the fused features; the 3D localization and regression network regressing the 3D bounding box of the target object in a predetermined coordinate system based on the fused features and corresponding depth information, the 3D bounding box being defined by its center point coordinates, size, and orientation angle; and combining the target object's category, physical attribute label, and 3D bounding box, determining the optimal grasping point coordinates and the approach posture of the robotic arm's end effector based on a preset grasping strategy rule base. The gripping point and attitude decision-making steps include: if the target object is classified as fragile, the area with the largest support surface and the most stable center of gravity is selected as the gripping point, and a vertical gripping attitude perpendicular to the support surface is adopted; if the target object is classified as magnetic metal, a flat surface area is preferentially selected as the gripping point, and an electromagnetic chuck is triggered as the end effector; if the 3D localization regression network output shows that the target object is in an unstable stacked state, a gripping point and extraction trajectory aimed at avoiding overall collapse are determined. Specifically, the input is an aligned visible light ROI image. After preprocessing, color features are calculated sequentially (a color histogram (16x16x16 bins) is calculated in HSV space), texture and edge features are calculated (texture energy is extracted using a multi-scale Gabor filter bank; Canny edge map is calculated and edge direction histogram is extracted), and shape features are calculated (7 Hu invariant moments are calculated on the binarized target contour, and polygon approximation is performed to obtain vertex features). These features are concatenated into a primary visible light feature vector F_vis. The input is an aligned infrared temperature distribution map (converted to grayscale or pseudo-color). Thermal distribution statistics (average temperature, temperature variance, highest / lowest temperature points in the target area), thermal gradient field (the temperature gradient of the image (∂T / ∂x, ∂T / ∂y) is calculated, and the histogram and principal direction of the gradient magnitude are statistically analyzed), and uniformity measure (the normalized standard deviation of temperature in the target area is calculated as a uniformity score). These features are concatenated into a primary infrared feature vector F_ir. The input is aligned target point cloud data. The computational geometric features (using PCA to calculate the principal components of the point cloud, obtaining normal vectors and curvature estimates; calculating the distribution of the point cloud along the three principal axes (height, width, depth)), product and density (calculating the volume of the point cloud bounding box; calculating the average number of points (density) per unit volume), and height map features (projecting the point cloud onto a 2D mesh to generate a height map, calculating statistics such as roughness of the height map) are concatenated into a primary depth feature vector F_depth. From the raw data, the average gradient magnitude of the visible light image ROI is calculated as a contrast measure C_vis; the proportion of overexposed pixels O_vis is calculated. The reliability factor R_vis = α1*C_vis / (1+α2*O_vis). The average temperature difference ΔT between the infrared image ROI and the background region and the signal-to-noise ratio SNR_ir are calculated. The reliability factor R_ir = β1*|ΔT|+β2*SNR_ir. The effective point ratio P_valid and the average fitting residual E_depth of the depth point cloud are calculated (through plane fitting). The reliability factor R_depth = γ1 * P_valid / (1 + γ2 * E_depth). (α, β, γ are empirical coefficients that can be obtained through offline optimization); the three reliability factors are normalized by Softmax to obtain the dynamic fusion weights: W_vis, W_ir, W_depth, and W_vis + W_ir + W_depth = 1.Each feature vector is weighted: F'_vis = W_vis * F_vis, F'_ir = W_ir * F_ir, F'_depth = W_depth * F_depth; The weighted feature vectors are then concatenated to form the final adaptive fusion feature vector: F_fusion = Concat(F'_vis, F'_ir, F'_depth); F_fusion is fed into a multi-task learning network, which includes a shared base layer (composed of 3 fully connected layers, responsible for learning high-level abstract features H_shared from F_fusion), a classification branch (attached after H_shared, outputting the material category probability P_mat, structure category probability P_str, and attribute label probability P_attr through two fully connected layers and Softmax), and a localization branch (also attached after H_shared, directly regressing the 7 bounding box parameters (x, y, z, l, w, h, θ) through a fully connected layer); A grasping strategy rule base is set and stored in "IF-THEN" format. For example: IF Material Category = "Glass" AND Attribute Tag contains "Fragile" THEN {End effector = "Adaptive Vacuum Suction Cup"; Grip Point = "Center of Top Flat Area"; Approach Attitude = "Vertical Down"; Adhesion Force = "Low (<50kPa)"}. IF Material Category = "Steel" AND Structure Category = "Can" THEN {End effector = "Electromagnet"; Grip Point = "Middle of Side"; Approach Attitude = "Perpendicular to Can Axis"}. IF Structure Category = "Membrane" AND Attribute Tag contains "Deformable" THEN {End effector = "Two-Finger Flexible Gripper"; Grip Point = "Above Estimated Center of Gravity"; Grip Mode = "Wrap-Around"; Gripping Force = "Medium"}. The decision engine takes (P_mat, P_str, P_attr) and (x, y, z, l, w, h, θ) output by the network as input, matches them with the rule base, and outputs specific grabbing instruction tuples (Grasp_Point, End_Effector_Pose, Tool_Selection, Force_Params). By dynamically evaluating the reliability of each modality and assigning fusion weights, it can adapt to scene changes and collaboratively optimize recognition and localization, thereby improving the accuracy and efficiency of sorting recyclable resources.
[0024] S4: Generate robotic arm control commands based on the identification and positioning results, and drive the robotic arm to perform sorting operations.
[0025] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0026] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multimodal image-based robotic arm sorting method, characterized in that, Includes the following steps: Multimodal image data of the target area is acquired synchronously through a multimodal image acquisition system. The multimodal image data includes at least visible light images, infrared thermal imaging images, and depth images. The multimodal image data is registered to construct a multimodal image fusion dataset with spatiotemporal consistency; Based on the multimodal image fusion dataset, the fusion features of the target are extracted by the multimodal feature extraction module; the category of the target object is identified according to the fusion features and its position and pose in three-dimensional space are calculated; Based on the identification and positioning results, control commands for the robotic arm are generated to drive the robotic arm to perform sorting operations.
2. The multimodal image robotic arm sorting method according to claim 1, characterized in that: The step of registering the multimodal image data to construct a multimodal image fusion dataset with spatiotemporal consistency includes: Based on the spatial transformation relationship between the various modal sensors obtained in advance through the composite calibration plate, the visible light image, infrared thermal imaging image and depth image are transformed to the same coordinate system to form a spatially aligned primary fusion image group; By using a hardware synchronization trigger signal and timestamp alignment mechanism, it is ensured that each modal image in the primary fused image group corresponds to the same acquisition time; Motion compensation is performed on the spatially aligned primary fused image group in response to the movement of the target object on the conveyor belt. The multimodal image fusion dataset is constructed by associating the spatially and temporally registered and dynamically compensated multimodal image data with the corresponding acquisition environment parameters and target location information.
3. The multimodal image robotic arm sorting method according to claim 2, characterized in that: The step of transforming the visible light image, infrared thermal imaging image, and depth image to the same coordinate system based on the spatial transformation relationship between the various modal sensors obtained through pre-calibration using a composite calibration plate, to form a spatially aligned primary fused image group includes: The multimodal image acquisition system is jointly calibrated offline based on a composite calibration plate, the surface of which is simultaneously provided with optical marks, temperature control marks and depth marks; Based on the offline joint calibration results, spatial transformation parameters from the infrared camera coordinate system and depth sensor coordinate system to the visible light camera coordinate system or unified world coordinate system are calculated. During the online sorting process, coordinate transformation and pixel-level resampling are performed on the real-time acquired modal images based on the spatial transformation parameters.
4. The multimodal image robotic arm sorting method according to claim 2, characterized in that: The step of ensuring that each modal image in the primary fused image group corresponds to the same acquisition time through the hardware synchronization trigger signal and timestamp alignment mechanism includes: A unified hardware synchronization triggering device is configured for the visible light camera, infrared thermal imaging camera and depth sensor. The main controller sends a synchronization pulse signal to trigger all sensors to perform image exposure and acquisition at the same physical moment. Assign a unique timestamp to each frame of image data acquired; Based on the synchronization pulse signal and timestamp, frame matching is performed on image streams from different sensors, and image frames corresponding to the same trigger period are combined into the primary fused image group.
5. A multimodal image robotic arm sorting method according to claim 2, characterized in that: The step of performing motion compensation on the spatially aligned primary fused image group in response to the motion of a target object on the conveyor belt includes: Acquire real-time speed and direction information of the conveyor belt carrying the target object; Based on the motion speed, direction information, and the inherent exposure delay of each sensor, the displacement deviation of the same target at the moment of adjacent modal image acquisition is calculated; Based on the displacement deviation, pixel translation or affine transformation is performed on the corresponding modal images in the primary fused image group to compensate for the registration error.
6. The multimodal image robotic arm sorting method according to claim 2, characterized in that: The step of associating the spatially and temporally registered and dynamically compensated multimodal image data with corresponding acquisition environment parameters and target location information to construct the multimodal image fusion dataset includes: Assign a unique identifier to each set of registered multimodal image data, and record the acquisition time, ambient temperature, ambient light intensity, conveyor belt speed, and sensor operating parameters. The registered visible light image, infrared image, depth image, and corresponding environmental parameters and identification information are encapsulated into structured data units. Multiple structured data units are organized according to time series or task batches to form a multimodal image fusion dataset.
7. The multimodal image robotic arm sorting method according to claim 1, characterized in that: The step of extracting the fusion features of the target using the multimodal feature extraction module based on the multimodal image fusion dataset includes: The spatially aligned subsets of visible light images, infrared images, and depth images from the multimodal image fusion dataset are input into the visible light feature extraction unit, the infrared feature extraction unit, and the depth feature extraction unit, respectively, to extract the primary feature representation of the target in each modality in parallel. Based on the spatiotemporal consistency information recorded in the multimodal image fusion dataset, a one-to-one correspondence between the visible light primary features, infrared primary features, and depth primary features is established at the pixel level or region level. The reliability of each modal primary feature is evaluated based on the current scenario conditions, and dynamic fusion weights are assigned to each modal primary feature based on the evaluation results. The weighted modal primary features are then spliced together or aggregated element by element to generate the target fusion feature.
8. The multimodal image robotic arm sorting method according to claim 1, characterized in that: The step of identifying the category of the target object based on the fused features and calculating its position and orientation in three-dimensional space includes: The fused features are simultaneously input into a jointly learned target classification network and a 3D localization regression network; The target classification network outputs the material category, structure category, and physical attribute label of the target object based on the fusion features. The 3D localization regression network regresses the 3D bounding box of the target object in a predetermined coordinate system based on the fused features and the corresponding depth information. The 3D bounding box is defined by its center point coordinates, size and orientation angle. Based on the target object's category, physical attribute labels, and 3D bounding box, and according to a preset grasping strategy rule base, the optimal grasping point coordinates and the approach posture of the robotic arm's end effector are determined.