A multi-sensor-based palletizing and dispensing control method
By constructing a three-dimensional semantic map through multi-sensor fusion technology, accurate perception and adaptive control of complex dynamic environments are achieved, solving the problems of high misjudgment rate and poor stability of existing palletizing systems in dynamic environments, and improving the safety and efficiency of palletizing operations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SENAD TECH CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-06-19
AI Technical Summary
Existing palletizing systems struggle to achieve multi-dimensional, high-precision perception in complex dynamic environments, resulting in high misjudgment rates, poor stacking stability, and an inability to effectively prevent collisions, leading to insufficient system reliability.
By employing multi-sensor fusion technology, multimodal perception data is simultaneously collected through infrared sensors, industrial cameras, and lidar to construct a three-dimensional semantic map, enabling risk assessment and adaptive control, and achieving accurate identification and dynamic adjustment of the stacking environment.
It improves the safety and efficiency of palletizing operations, enhances the system's adaptability in complex environments, and improves the accuracy of single operations and overall reliability through closed-loop control and model update optimization.
Smart Images

Figure CN121778431B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of industrial and logistics automation technology, and in particular to a palletizing and unloading control method based on multiple sensors. Background Technology
[0002] In modern logistics and manufacturing, automated palletizing is a key element in improving warehousing and production line efficiency. Currently, common palletizing and material placement control methods mostly rely on preset programs or single sensor guidance, such as using only infrared sensors for distance detection or relying solely on industrial cameras for visual positioning. These methods are applicable in static, orderly palletizing environments, but in actual production lines or warehouses, palletizing positions are often in a complex and dynamic state, presenting significant technical challenges.
[0003] Existing systems mostly use infrared sensors, which can only acquire distance information and cannot distinguish the nature of foreign objects. This leads to a high false alarm rate, frequently triggering unnecessary emergency shutdowns, or missing transparent or low-reflectivity foreign objects, creating potential collision hazards. Traditional methods cannot construct the three-dimensional geometric contour of the stack surface, thus failing to accurately determine the offset posture of the placed boxes, resulting in insufficient fine-tuning precision and poor stacking stability. Single sensors are susceptible to common interferences in industrial environments, such as dust, changes in lighting, and reflective surfaces, leading to decreased data quality, increased false alarm rates, and difficulty in guaranteeing system reliability.
[0004] Therefore, how to achieve multi-dimensional, high-precision, and robust perception of the palletizing environment, and on this basis, build a hierarchical risk assessment and adaptive control strategy, has become a key technical challenge to improve the intelligence level of the palletizing system and ensure operational safety and efficiency. Summary of the Invention
[0005] The purpose of this application is to provide a multi-sensor-based palletizing and unloading control method to solve the above-mentioned technical problems. This method aims to achieve accurate risk assessment and adaptive action planning of the working environment through multi-modal perception and intelligent decision-making, thereby ensuring the safety, stability and high efficiency of palletizing operations in complex dynamic scenarios.
[0006] In some embodiments of this application, a multi-sensor-based palletizing and unloading control method is provided, including:
[0007] The robotic arm is controlled to carry the box to a preset height above the target stack and hover, while simultaneously collecting multimodal sensing data of the target stack to generate a multimodal sensing dataset.
[0008] Spatiotemporal registration is performed on the multimodal sensing dataset to generate a three-dimensional semantic map containing stack position status information;
[0009] The three-dimensional semantic map is compared with a preset benchmark model, and a risk level is generated based on the comparison results.
[0010] Based on the risk level, control the robotic arm to perform the corresponding preset material feeding action;
[0011] After the robotic arm performs a preset action, it verifies the placement state of the box based on data collected by sensors, generates a verification result, and updates the benchmark model based on the verification result.
[0012] In some embodiments of this application, when controlling the robotic arm to move the box to a preset height above the target stack and hover, and simultaneously collecting multimodal sensing data of the target stack to generate a multimodal sensing dataset, the process includes:
[0013] Control the robotic arm to carry the box to be stacked to a preset hovering height above the target stacking position and hover there;
[0014] Simultaneously activate multiple sensor groups deployed on the robotic arm to collect data; the sensor groups include at least an infrared sensor, an industrial camera, and a lidar.
[0015] The infrared sensor collects spatial distance data of the target stack, the industrial camera collects two-dimensional visual image data of the target stack, and the lidar collects three-dimensional point cloud contour data of the target stack.
[0016] The spatial distance data, two-dimensional visual image data, and three-dimensional point cloud contour data constitute the multimodal perception dataset.
[0017] In some embodiments of this application, spatiotemporal registration is performed on the multimodal sensing dataset to generate a three-dimensional semantic map containing stack position status information, including:
[0018] The multimodal sensing data is time-synchronized and spatially aligned to generate standardized multimodal sensing data;
[0019] Feature extraction and fusion are performed on the standardized multimodal perception data to generate a three-dimensional semantic map;
[0020] The three-dimensional semantic map includes at least the surface three-dimensional contour of the stack, the three-dimensional position, size, and type of the foreign object, and the offset posture information of the placed box.
[0021] In some embodiments of this application, the three-dimensional semantic map is compared with a preset benchmark model, and a risk level is generated based on the comparison results, including:
[0022] The 3D position, size, type, and offset posture information of the foreign object contained in the 3D semantic map, as well as the offset posture information of the placed box, are compared with the preset benchmark model to generate comparison results;
[0023] Based on the comparison results, a comprehensive judgment is made based on whether there are foreign objects, whether the foreign objects are negligible, and whether the box offset exceeds the corresponding threshold, and a comprehensive judgment result is generated.
[0024] Based on the comprehensive judgment, the risk levels are divided into no risk, low risk, medium risk, and high risk.
[0025] In some embodiments of this application, based on the comparison results, a comprehensive judgment is generated according to the presence of foreign objects, whether the foreign objects are negligible, and whether the box offset exceeds a corresponding threshold, including:
[0026] The criteria for determining the risk-free level are as follows:
[0027] There are no foreign objects in the three-dimensional semantic map, and the offset of the placed box is less than or equal to the first preset threshold.
[0028] The criteria for determining a low-risk level include any one of the following:
[0029] The three-dimensional semantic map contains objects whose semantic type is negligible foreign objects, and whose size is less than or equal to the second preset threshold.
[0030] The offset of the placed box is greater than the first preset threshold and less than or equal to the third preset threshold;
[0031] The criteria for determining the medium-risk level include any one of the following:
[0032] The three-dimensional semantic map contains objects whose semantic type is "objects that need to be fine-tuned to avoid", and their size is greater than the second preset threshold and less than or equal to the fourth preset threshold.
[0033] The offset of the placed box is greater than the third preset threshold and less than or equal to the fifth preset threshold;
[0034] The criteria for determining a high-risk level include any one of the following:
[0035] The three-dimensional semantic map contains objects whose semantic type is "requiring emergency shutdown" or whose size is greater than the fourth preset threshold.
[0036] The offset of the placed box is greater than the fifth preset threshold.
[0037] In some embodiments of this application, the robotic arm is controlled to perform corresponding preset material feeding actions according to the risk level, including:
[0038] If the risk level is no risk, control the robotic arm to vertically lower the box at a preset first speed;
[0039] If the risk level is low, the robotic arm is controlled to make fine-tuning of its posture based on the three-dimensional semantic map, and then the box is lowered at a second speed lower than the first speed.
[0040] If the risk level is medium risk, then control the robotic arm to plan and execute the obstacle avoidance path based on the three-dimensional semantic map;
[0041] If the risk level is high, the robotic arm will be controlled to perform an emergency stop and retraction action, and an alarm signal will be triggered.
[0042] In some embodiments of this application, after the robotic arm performs a preset action, it verifies the placement state of the box based on sensor-collected data, generates a verification result, and updates the benchmark model based on the verification result, including:
[0043] After the robotic arm performs the lowering action and the box contacts the stack surface, the sensor collects real-time data of the target stack position again.
[0044] Based on the real-time data, verify whether the relative positional relationship between the placed box and the surrounding stacking environment meets the preset safety conditions;
[0045] If the preset safety conditions are met, the robotic arm is controlled to perform a release operation to complete the material feeding, and the multimodal perception data and three-dimensional semantic map corresponding to this operation are fused into the preset benchmark model.
[0046] In some embodiments of this application, based on real-time sensing data, the verification of whether the relative positional relationship between the placed container and the surrounding stacking environment meets preset safety conditions includes:
[0047] Reconstructing the three-dimensional contour of the stack after placement based on real-time sensing data;
[0048] Calculate the minimum distance between the placed box and the boundary of the adjacent box or stack;
[0049] Determine whether the minimum spacing is greater than or equal to a preset safe distance threshold.
[0050] In some embodiments of this application, if the preset safety conditions are met, the robotic arm is controlled to perform a gripper release operation to complete the material unloading, and the multimodal perception data and three-dimensional semantic map corresponding to this operation are fused into the preset baseline model, including:
[0051] The multimodal perception dataset corresponding to the successful completion of this task will be stored together with historical data in the model library;
[0052] Based on the data in the model library, the sensor data characteristics under different environments or enclosure types are dynamically analyzed.
[0053] Based on the analysis results, the offset threshold used for risk classification is updated.
[0054] Some embodiments of this application also include:
[0055] The data acquisition or fusion process also includes sensor anomaly detection and cross-validation steps:
[0056] When the data quality of any sensor is detected to be lower than a preset threshold, the cross-validation mechanism is activated.
[0057] Consistency checks are performed using data from at least two other sensors to eliminate outlier data points.
[0058] The subsequent fusion process continues based on the verified valid data.
[0059] Compared with the prior art, the palletizing and unloading control method based on multiple sensors proposed in this application has the following advantages:
[0060] By synchronously collecting multimodal sensing data and constructing a 3D semantic map containing details of the stacking positions, the system achieves accurate identification and quantitative assessment of risks such as foreign objects on the work surface and box misalignment. Based on comparison with a preset benchmark model, the system can adaptively drive the robotic arm to perform differentiated material feeding actions, from routine lowering to emergency retraction, according to the graded risk assessment results. This effectively prevents collisions and ensures stacking stability in complex dynamic environments. After each operation, the closed-loop control system achieves self-optimization through state verification and model updates. This not only improves the safety and accuracy of individual operations but also enhances the overall adaptability to diverse boxes and complex working conditions, ultimately achieving a comprehensive improvement in the efficiency and reliability of palletizing operations. Attached Figure Description
[0061] Figure 1 This is a flowchart illustrating a multi-sensor-based palletizing and unloading control method in a preferred embodiment of this application. Detailed Implementation
[0062] The specific embodiments of this application will be described in further detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate this application, but are not intended to limit the scope of this application.
[0063] In the description of this application, it should be understood that the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.
[0064] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0065] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0066] like Figure 1 As shown, a preferred embodiment of this application provides a multi-sensor-based palletizing and unloading control method, comprising:
[0067] The robotic arm is controlled to carry the box to a preset height above the target stack and hover, while simultaneously collecting multimodal sensing data of the target stack to generate a multimodal sensing dataset.
[0068] Spatiotemporal registration is performed on the multimodal sensing dataset to generate a three-dimensional semantic map containing stack position status information;
[0069] The three-dimensional semantic map is compared with a preset benchmark model, and a risk level is generated based on the comparison results.
[0070] Based on the risk level, control the robotic arm to perform the corresponding preset material feeding action;
[0071] After the robotic arm performs a preset action, it verifies the placement state of the box based on data collected by sensors, generates a verification result, and updates the benchmark model based on the verification result.
[0072] In some embodiments of this application, when controlling the robotic arm to move the box to a preset height above the target stack and hover, and simultaneously collecting multimodal sensing data of the target stack to generate a multimodal sensing dataset, the process includes:
[0073] Control the robotic arm to carry the box to be stacked to a preset hovering height above the target stacking position and hover there;
[0074] Simultaneously activate multiple sensor groups deployed on the robotic arm to collect data; the sensor groups include at least an infrared sensor, an industrial camera, and a lidar.
[0075] The infrared sensor collects spatial distance data of the target stack, the industrial camera collects two-dimensional visual image data of the target stack, and the lidar collects three-dimensional point cloud contour data of the target stack.
[0076] The spatial distance data, two-dimensional visual image data, and three-dimensional point cloud contour data constitute the multimodal perception dataset.
[0077] In this embodiment, the robotic arm first carries the box to be palletized to a position directly above the target pallet and hovers stably at a preset hovering height. After hovering, the system synchronously triggers a multi-sensor array deployed on the robotic arm to collect data: infrared sensors arranged in four directions on the gripper collect spatial distance data of the pallet in real time; an industrial camera mounted on the front end of the gripper captures a two-dimensional visual image of the pallet; and a lidar located on the wrist of the robotic arm scans to obtain the three-dimensional point cloud contour of the pallet. All sensors are synchronized in time to ensure data temporal consistency. After data collection, the system encapsulates the above three types of data into a unified multimodal sensing dataset for subsequent fusion processing. This dataset not only contains the original sensing information but can also include auxiliary data such as sensor status and timestamps to support subsequent spatiotemporal registration and fusion analysis.
[0078] In some embodiments of this application, spatiotemporal registration is performed on the multimodal sensing dataset to generate a three-dimensional semantic map containing stack position status information, including:
[0079] The multimodal sensing data is time-synchronized and spatially aligned to generate standardized multimodal sensing data;
[0080] Feature extraction and fusion are performed on the standardized multimodal perception data to generate a three-dimensional semantic map;
[0081] The three-dimensional semantic map includes at least the surface three-dimensional contour of the stack, the three-dimensional position, size, and type of the foreign object, and the offset posture information of the placed box.
[0082] In this embodiment, the generation of standardized multimodal sensing data includes: the system first performs spatiotemporal registration on the raw sensing data from infrared sensors, industrial cameras, and LiDAR: hardware trigger signals ensure synchronous acquisition by each sensor, with time errors controlled within 10ms; simultaneously, a precise timestamp is added to each frame of data for subsequent interpolation and alignment. Based on a pre-calibrated sensor extrinsic parameter matrix, the infrared distance data, image pixel coordinates, and LiDAR point cloud are uniformly transformed to the coordinate system of the robotic arm's end effector, achieving spatial coordinate consistency. After the above processing, a standardized multimodal sensing dataset is obtained, which already possesses the characteristics of time alignment and coordinate unification, and can be directly used for fusion calculations.
[0083] In this embodiment, the offset attitude information includes the offset angle and the offset distance.
[0084] In this embodiment, based on standardized multimodal perception data, the system sequentially performs feature extraction, multimodal fusion, and map construction to generate a 3D semantic map for subsequent judgment. First, visual features are extracted from industrial camera images to identify the type of foreign object. Simultaneously, geometric features are extracted from LiDAR point clouds to reconstruct the 3D contour of the stack and calculate the box offset. The feature positions are then verified using distance data from infrared sensors. Subsequently, visual semantic information and point cloud geometric information are fused through projection mapping and feature association to form a 3D object description with semantic labels. Based on this, the system reconstructs a 3D model of the stack based on the fusion results and adds semantic labels to each element in the model. Finally, it outputs a 3D semantic map containing the surface contour of the stack, the location / size / type of foreign objects, and the box offset posture, thereby providing accurate and structured environmental perception information for risk classification and action planning.
[0085] In some embodiments of this application, the three-dimensional semantic map is compared with a preset benchmark model, and a risk level is generated based on the comparison results, including:
[0086] The 3D position, size, type, and offset posture information of the foreign object contained in the 3D semantic map, as well as the offset posture information of the placed box, are compared with the preset benchmark model to generate comparison results;
[0087] Based on the comparison results, a comprehensive judgment is made based on whether there are foreign objects, whether the foreign objects are negligible, and whether the box offset exceeds the corresponding threshold, and a comprehensive judgment result is generated.
[0088] Based on the comprehensive judgment, the risk levels are divided into no risk, low risk, medium risk, and high risk.
[0089] In this embodiment, a preset benchmark model serves as a reference for risk assessment within the system. This model includes: an ideal three-dimensional profile of the stack surface, reflecting the geometric shape of the stack without any foreign objects or box offset; a standard box placement posture, representing the theoretical position and posture parameters of the box on the stack surface under perfect alignment; definitions of foreign object semantic categories and size thresholds, pre-classifying different foreign object categories and their corresponding size safety thresholds; and environmental adaptability parameters, which may include sensor data features under different lighting, dust, and other environmental conditions to enhance the model's robustness. This model is constructed through initial calibration, learning from historical successful box placement data, and manual setting, and it has the ability to dynamically update based on actual box placement results during operation.
[0090] In this embodiment, the system generates comparison results including: foreign object detection information: including the three-dimensional position, size, semantic category of the foreign object, and whether the foreign object belongs to the categories of negligible, requiring fine-tuning to avoid, or requiring emergency shutdown; box offset information: including the translational offset and attitude deflection angle of the currently placed box relative to the ideal position in the reference model; difference quantification data: specific numerical difference measures, such as foreign object height, box offset distance, tilt angle, etc.; spatial relationship judgment: such as whether the foreign object is located within the projection area of the box to be placed, whether the box interferes with the boundaries of adjacent boxes or stacking positions, etc.
[0091] In some embodiments of this application, based on the comparison results, a comprehensive judgment is generated according to the presence of foreign objects, whether the foreign objects are negligible, and whether the box offset exceeds a corresponding threshold, including:
[0092] The criteria for determining the risk-free level are as follows:
[0093] There are no foreign objects in the three-dimensional semantic map, and the offset of the placed box is less than or equal to the first preset threshold.
[0094] The criteria for determining a low-risk level include any one of the following:
[0095] The three-dimensional semantic map contains objects whose semantic type is negligible foreign objects, and whose size is less than or equal to the second preset threshold.
[0096] The offset of the placed box is greater than the first preset threshold and less than or equal to the third preset threshold;
[0097] The criteria for determining the medium-risk level include any one of the following:
[0098] The three-dimensional semantic map contains objects whose semantic type is "objects that need to be fine-tuned to avoid", and their size is greater than the second preset threshold and less than or equal to the fourth preset threshold.
[0099] The offset of the placed box is greater than the third preset threshold and less than or equal to the fifth preset threshold;
[0100] The criteria for determining a high-risk level include any one of the following:
[0101] The three-dimensional semantic map contains objects whose semantic type is "requiring emergency shutdown" or whose size is greater than the fourth preset threshold.
[0102] The offset of the placed box is greater than the fifth preset threshold.
[0103] In this embodiment, in the risk classification mechanism of the present invention, the system compares the type and size of foreign objects extracted from the three-dimensional semantic map with multiple preset thresholds to comprehensively determine the current operation risk level. Specifically: No risk requires that there are no foreign objects at the stacking position and that the offset of the placed box does not exceed the first preset threshold; in this case, the system will proceed with the box placement according to the normal process. Low risk: If a negligible foreign object with a size not exceeding the second preset threshold is detected, or the box offset is between the first and third preset thresholds, the system will continue operation after slight adjustments or deceleration. Medium risk: When a foreign object requiring fine-tuning is identified with a size between the second and fourth preset thresholds, or the box offset is between the third and fifth preset thresholds, the system will initiate obstacle avoidance planning or attitude compensation. High risk: Once a foreign object with a semantic type requiring emergency shutdown is detected, or the foreign object size exceeds the fourth preset threshold, or the box offset exceeds the fifth preset threshold, the system will immediately stop the operation, retract the robotic arm, and trigger an alarm. The above judgment conditions, through the triple constraint of semantic size offset, achieve a progressive risk response from low to high, balancing operational safety and efficiency.
[0104] In this embodiment, the preset thresholds used for risk level determination are systematic parameters comprehensively set based on the robotic arm's accuracy, the box size, safety regulations, and numerous scenario tests. To achieve system self-adaptation, all thresholds support dynamic adjustment: the system continuously learns from historical operation data through a benchmark model library and optimizes threshold recommendations using statistical methods; in environments with interference such as dust or strong light, the system can automatically fine-tune the thresholds and match corresponding speed reduction strategies based on sensor data quality; when changing the box type, the system can automatically update the entire set of threshold parameters proportionally based on the new box size, achieving rapid adaptation.
[0105] In this embodiment, the semantic classification of foreign objects is accurately determined based on the results of multimodal fusion. Negligible foreign objects are usually paper scraps or dust; foreign objects that need to be finely adjusted to avoid are mostly small tools; foreign objects that require emergency shutdown include human limbs or large obstacles.
[0106] In some embodiments of this application, the robotic arm is controlled to perform corresponding preset material feeding actions according to the risk level, including:
[0107] If the risk level is no risk, control the robotic arm to vertically lower the box at a preset first speed;
[0108] If the risk level is low, the robotic arm is controlled to make fine-tuning of its posture based on the three-dimensional semantic map, and then the box is lowered at a second speed lower than the first speed.
[0109] If the risk level is medium risk, then control the robotic arm to plan and execute the obstacle avoidance path based on the three-dimensional semantic map;
[0110] If the risk level is high, the robotic arm will be controlled to perform an emergency stop and retraction action, and an alarm signal will be triggered.
[0111] In this embodiment, the specific control logic and action implementation methods include: Action execution under no-risk level: When the system determines that the stacking environment is ideal, there are no foreign objects, and the offset of the placed box is within the allowable range. At this time, the system controls the robotic arm to smoothly lower the box to be stacked along the vertical direction at a first speed. This speed is set to the highest efficiency speed for normal operation to achieve fast and accurate placement. Action execution under low-risk level: When the risk level is low, it usually means that there are negligible small foreign objects, or the placed box has a slight offset. In this case, the system first calculates the required attitude fine-tuning amount based on the 3D semantic map and drives the end effector of the robotic arm to complete the attitude adjustment. The robotic arm lowers the box at a second speed lower than the first speed. The deceleration is intended to provide the system with more reaction time and further ensure the placement safety under slight abnormal conditions. Action execution under medium-risk level: When medium risk is triggered, it indicates that there are foreign objects on the stacking position that need to be actively avoided, or the offset of the box is relatively obvious. The system will use a path planning algorithm based on the 3D semantic map to calculate an obstacle avoidance lowering path in real time. This path planning ensures the robotic arm can avoid obstacles before final placement. Once planned, the robotic arm will perform the unloading action along the path at a low speed, completing the task while avoiding collisions. High-risk actions: A high-risk level indicates an immediate danger, such as the detection of a person's limb, a large foreign object, or a severe displacement of the container. In this case, the system immediately sends an emergency stop command to the robotic arm, interrupting the current unloading process. Subsequently, the robotic arm is controlled to retreat along a safe trajectory to a preset safe position. Simultaneously, the system triggers an audible and visual alarm, notifying on-site personnel to intervene and suspending subsequent palletizing operations until the risk is manually confirmed and eliminated. Through this graded control strategy strictly tied to risk levels, this invention achieves a progressive safety response from uninterrupted high-speed operation to fine-tuned deceleration, then to active obstacle avoidance, and finally to emergency shutdown, thereby dynamically balancing the efficiency, safety, and reliability of palletizing operations in complex and ever-changing industrial scenarios.
[0112] In some embodiments of this application, after the robotic arm performs a preset action, it verifies the placement state of the box based on sensor-collected data, generates a verification result, and updates the benchmark model based on the verification result, including:
[0113] After the robotic arm performs the lowering action and the box contacts the stack surface, the sensor collects real-time data of the target stack position again.
[0114] Based on the real-time data, verify whether the relative positional relationship between the placed box and the surrounding stacking environment meets the preset safety conditions;
[0115] If the preset safety conditions are met, the robotic arm is controlled to perform a release operation to complete the material feeding, and the multimodal perception data and three-dimensional semantic map corresponding to this operation are fused into the preset benchmark model.
[0116] In some embodiments of this application, based on real-time sensing data, the verification of whether the relative positional relationship between the placed container and the surrounding stacking environment meets preset safety conditions includes:
[0117] Reconstructing the three-dimensional contour of the stack after placement based on real-time sensing data;
[0118] Calculate the minimum distance between the placed box and the boundary of the adjacent box or stack;
[0119] Determine whether the minimum spacing is greater than or equal to a preset safe distance threshold.
[0120] In this embodiment, after the robotic arm performs the lowering action and the bottom surface of the box contacts the surface of the stack, the system synchronously triggers the deployed sensor group again to collect real-time multimodal perception data of the target stack. Based on this real-time data, the system accurately verifies the relative positional relationship between the placed box and the surrounding environment to determine whether it meets the preset safety conditions. Specifically, this includes: 3D contour reconstruction: The system uses real-time collected point cloud and visual data to reconstruct the instantaneous 3D contour of the stack after the placement operation, obtaining the actual posture of the box and the surrounding geometric environment after placement; Minimum distance calculation: Based on the reconstructed 3D contour, the system calculates the minimum spatial distance between each edge of the placed box and the nearest obstacle; Safety threshold judgment: The system compares the calculated minimum distance with the preset safety distance threshold. If the minimum distance is greater than or equal to the safety threshold, it is determined that the preset safety conditions are met, the box is successfully placed and there is no risk of collision; otherwise, it is marked as an abnormal placement.
[0121] In some embodiments of this application, if the preset safety conditions are met, the robotic arm is controlled to perform a gripper release operation to complete the material unloading, and the multimodal perception data and three-dimensional semantic map corresponding to this operation are fused into the preset baseline model, including:
[0122] The multimodal perception dataset corresponding to the successful completion of this task will be stored together with historical data in the model library;
[0123] Based on the data in the model library, the sensor data characteristics under different environments or enclosure types are dynamically analyzed.
[0124] Based on the analysis results, the offset threshold used for risk classification is updated.
[0125] In this embodiment, if the verification result meets the preset safety conditions, the system will send a release command to the robotic arm to release the box and complete the palletizing and unloading operation. The system will integrate the complete multimodal perception dataset collected in this operation cycle, the generated 3D semantic map, and the final verification result as a successful sample into a preset benchmark model library. This update process allows the benchmark model to continuously absorb actual operation data, optimize its internal threshold parameters and environmental feature expressions, thereby improving the system's adaptability and judgment accuracy to future operation scenarios and achieving progressive self-optimization. Through the above-described execution-verification-update closed-loop process, this invention not only ensures the safety of a single unloading operation, but also enables the system to continuously adapt to complex on-site environments and diverse operation objects through a data-driven learning mechanism, continuously improving the overall intelligence level and reliability of palletizing operations.
[0126] Some embodiments of this application also include:
[0127] The data acquisition or fusion process also includes sensor anomaly detection and cross-validation steps:
[0128] When the data quality of any sensor is detected to be lower than a preset threshold, the cross-validation mechanism is activated.
[0129] Consistency checks are performed using data from at least two other sensors to eliminate outlier data points.
[0130] The subsequent fusion process continues based on the verified valid data.
[0131] In this embodiment, when the data quality of any sensor is detected to be lower than a preset threshold, the cross-validation mechanism is initiated, including: Infrared sensor data quality assessment: whether the signal strength is higher than a preset threshold; whether the ranging value is within the effective range calibrated by the sensor; and whether the variance of fluctuations in multiple consecutive frames of data exceeds a stable threshold. Industrial camera image quality assessment: calculating image sharpness scores based on image gradient or sharpness algorithms to determine whether motion blur or defocus exists; analyzing image histograms to determine whether the brightness is within a suitable range; and determining whether the lens is damaged or partially obscured through edge detection or occlusion area analysis. LiDAR point cloud quality assessment: statistically analyzing the effective point cloud density to determine whether it is lower than the minimum density threshold required for the scene; detecting whether there are large areas of holes or non-physical noise points in the point cloud; and verifying whether the scanning frame rate is stable and whether there are any dropped frames. If the quality score of any sensor data is lower than its preset confidence threshold, the system determines that the data is abnormal and immediately triggers the multi-sensor cross-validation process.
[0132] In this embodiment, when a sensor is marked as abnormal, the system immediately locks the data stream. The system extracts data collected by the other two sensors at the current moment and compares the data within the same spatial area using coordinate system one and feature mapping. If the infrared sensor is abnormal, the system uses the average distance of the corresponding area point cloud extracted by the lidar and compares it with the distance value calculated by the industrial camera through stereo vision or geometric constraints. If the difference between the two is less than a preset tolerance, the data is considered consistent, and the weighted average of the two can be used as reliable distance information. If the industrial camera is abnormal, the system mainly relies on the contour information of the lidar point cloud and the multi-point ranging of the infrared sensor to deduce the visual semantic information of the key area through 3D reconstruction and projection. If the lidar is abnormal, the system uses the high-resolution image of the industrial camera to extract 2D features and combines them with the accurate distance data of the infrared sensor in multiple directions to estimate the approximate 3D structure of the scene through a multi-view geometric algorithm. Based on the consistency judgment result, the system removes abnormal data points that significantly deviate from the consensus. For critical but missing data, interpolation and repair algorithms based on historical data or physical models are used to supplement it. Finally, the system outputs a cleaned and validated high-quality multimodal perception dataset for subsequent spatiotemporal registration and 3D semantic map construction.
[0133] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and substitutions can be made without departing from the technical principles of this application, and these improvements and substitutions should also be considered within the scope of protection of this application.
Claims
1. A multi-sensor-based palletizing and unloading control method, characterized in that, include: The robotic arm is controlled to carry the box to a preset height above the target stack and hover, while simultaneously collecting multimodal sensing data of the target stack to generate a multimodal sensing dataset. Spatiotemporal registration is performed on the multimodal sensing dataset to generate a three-dimensional semantic map containing stack position status information; The three-dimensional semantic map is compared with a preset benchmark model, and a risk level is generated based on the comparison results. Based on the risk level, control the robotic arm to perform the corresponding preset material feeding action; After the robotic arm performs a preset action, it verifies the placement state of the box based on data collected by sensors, generates a verification result, and updates the benchmark model based on the verification result. Spatiotemporal registration is performed on the multimodal sensing dataset to generate a 3D semantic map containing stack position status information, including: The multimodal sensing data is time-synchronized and spatially aligned to generate standardized multimodal sensing data; Feature extraction and fusion are performed on the standardized multimodal perception data to generate a three-dimensional semantic map; The three-dimensional semantic map includes at least the surface three-dimensional contour of the stack, the three-dimensional position, size, and type of the foreign object, and the offset posture information of the placed box.
2. The multi-sensor-based palletizing and unloading control method as described in claim 1, characterized in that, When controlling the robotic arm to move the box to a preset height above the target stack and hover, and simultaneously collecting multimodal sensing data of the target stack to generate a multimodal sensing dataset, the process includes: Control the robotic arm to carry the box to be stacked to a preset hovering height above the target stacking position and hover there; Simultaneously activate multiple sensor groups deployed on the robotic arm to collect data; the sensor groups include at least an infrared sensor, an industrial camera, and a lidar. The infrared sensor collects spatial distance data of the target stack, the industrial camera collects two-dimensional visual image data of the target stack, and the lidar collects three-dimensional point cloud contour data of the target stack. The spatial distance data, two-dimensional visual image data, and three-dimensional point cloud contour data constitute the multimodal perception dataset.
3. The multi-sensor-based palletizing and unloading control method as described in claim 2, characterized in that, The three-dimensional semantic map is compared with a preset benchmark model, and a risk level is generated based on the comparison results, including: The 3D position, size, type, and offset posture information of the foreign object contained in the 3D semantic map, as well as the offset posture information of the placed box, are compared with the preset benchmark model to generate comparison results; Based on the comparison results, a comprehensive judgment is made based on whether there are foreign objects, whether the foreign objects are negligible, and whether the box offset exceeds the corresponding threshold, and a comprehensive judgment result is generated. Based on the comprehensive judgment, the risk levels are divided into no risk, low risk, medium risk, and high risk.
4. The multi-sensor-based palletizing and unloading control method as described in claim 3, characterized in that, Based on the comparison results, a comprehensive judgment is generated, taking into account the presence of foreign objects, whether the foreign objects are negligible, and whether the box offset exceeds the corresponding threshold. This comprehensive judgment result includes: The criteria for determining the risk-free level are as follows: There are no foreign objects in the three-dimensional semantic map, and the offset of the placed box is less than or equal to the first preset threshold. The criteria for determining a low-risk level include any one of the following: The three-dimensional semantic map contains objects whose semantic type is negligible foreign objects, and whose size is less than or equal to the second preset threshold. The offset of the placed box is greater than the first preset threshold and less than or equal to the third preset threshold; The criteria for determining the medium-risk level include any one of the following: The three-dimensional semantic map contains objects whose semantic type is "objects that need to be fine-tuned to avoid", and their size is greater than the second preset threshold and less than or equal to the fourth preset threshold. The offset of the placed box is greater than the third preset threshold and less than or equal to the fifth preset threshold; The criteria for determining a high-risk level include any one of the following: The three-dimensional semantic map contains objects whose semantic type is "requiring emergency shutdown" or whose size is greater than the fourth preset threshold. The offset of the placed box is greater than the fifth preset threshold.
5. The multi-sensor based palletizing put-away control method of claim 1, wherein, Based on the risk level, the robotic arm is controlled to perform corresponding preset material feeding actions, including: If the risk level is no risk, control the robotic arm to vertically lower the box at a preset first speed; If the risk level is low, the robotic arm is controlled to make fine-tuning of its posture based on the three-dimensional semantic map, and then the box is lowered at a second speed lower than the first speed. If the risk level is medium risk, then control the robotic arm to plan and execute the obstacle avoidance path based on the three-dimensional semantic map; If the risk level is high, the robotic arm will be controlled to perform an emergency stop and retraction action, and an alarm signal will be triggered.
6. The multi-sensor-based palletizing and unloading control method as described in claim 1, characterized in that, After the robotic arm performs a preset action, it verifies the placement state of the box based on sensor data, generates a verification result, and updates the baseline model based on the verification result, including: After the robotic arm performs the lowering action and the box contacts the stack surface, the sensor collects real-time data of the target stack position again. Based on the real-time data, verify whether the relative positional relationship between the placed box and the surrounding stacking environment meets the preset safety conditions; If the preset safety conditions are met, the robotic arm is controlled to perform a release operation to complete the material feeding, and the multimodal perception data and three-dimensional semantic map corresponding to this operation are fused into the preset benchmark model.
7. The multi-sensor-based palletizing and unloading control method as described in claim 6, characterized in that, Based on real-time sensing data, verify whether the relative positional relationship between the placed container and the surrounding stacking environment meets the preset safety conditions, including: Reconstructing the three-dimensional contour of the stack after placement based on real-time sensing data; Calculate the minimum distance between the placed box and the boundary of the adjacent box or stack; Determine whether the minimum spacing is greater than or equal to a preset safe distance threshold.
8. The multi-sensor based palletizing put-away control method of claim 6, wherein, If the preset safety conditions are met, the robotic arm is controlled to perform a release operation to complete the material unloading, and the multimodal perception data and three-dimensional semantic map corresponding to this operation are fused into the preset baseline model, including: The multimodal perception dataset corresponding to the successful completion of this task will be stored together with historical data in the model library; Based on the data in the model library, the sensor data characteristics under different environments or enclosure types are dynamically analyzed. Based on the analysis results, the offset threshold used for risk classification is updated.
9. The multi-sensor based palletizing put-away control method as claimed in claim 2, wherein, Also includes: The data acquisition or fusion process also includes sensor anomaly detection and cross-validation steps: When the data quality of any sensor is detected to be lower than a preset threshold, the cross-validation mechanism is activated. Consistency check is performed through data of at least two other sensors to exclude abnormal data points; Based on the valid data after the check, subsequent fusion processing is continued.
Citation Information
Patent Citations
Automatic stacking control system based on visual positioning
CN120246695A
Multi-camera vision stacking pose planning method and system for intelligent loading and unloading robot
CN120534645A