Goods handling method and system based on spatial relation reasoning
The method uses RGB-D cameras and advanced models to enhance the detection and sorting of irregularly stacked goods, addressing the challenge of obscured parts and unsafe grasping sequences, ensuring safe and efficient handling in complex environments.
Patent Information
- Application Number
- CN202510764315.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The prior art is difficult to effectively identify the shape and position relationship of the obstructed part in complex environments, resulting in unreasonable grabbing order and may lead to collapse or damage to the goods.
The cargo handling method based on spatial relationship reasoning is adopted. By obtaining color images and depth images, the spatial relationship information of effective goods is screened, the unconstrained cargo collection is identified, and the optimal handling sequence is generated based on depth information and height information is used to drive the mechanical equipment to perform the grab operation.
It realizes high-precision and stable cargo identification and capture, avoids the risk of cargo collapse, optimizes the unloading process, improves the safety and efficiency of the system, and reduces deployment costs.
Smart Images

Figure CN120318592A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and robot control, and particularly relates to a goods handling method and system based on spatial relationship reasoning, which are mainly applied to goods recognition and grasping in automated scenarios such as logistics sorting, warehousing management, and transportation scheduling. Background Art
[0002] With the rapid development of the e-commerce and logistics industries, the demand for automated loading, unloading, sorting, and warehousing management of goods is increasing day by day.
[0003] In practical applications, the main challenges faced by automatic unloading robots are as follows: on the one hand, goods are often stacked in an irregular manner, resulting in serious occlusion problems; on the other hand, the robot needs to determine a reasonable grasping order to avoid causing the collapse or damage of goods during the unloading process. For example, in the tobacco logistics scenario, when a batch of tobacco packaging boxes are randomly stacked in a container or a truck, traditional object detection algorithms can often only identify the packaging boxes visible on the surface, unable to infer the shape and positional relationship of the occluded parts, and it is even more difficult to determine a safe and feasible grasping order.
[0004] Existing methods lack effective modeling of the spatial dependence relationship between objects and cannot determine which boxes can be grasped first without causing the instability of other boxes. Summary of the Invention
[0005] The main purpose of the present invention is to solve the technical problems of insufficient object recognition accuracy and unreasonable grasping order planning in the prior art under complex environments. A goods handling method based on spatial relationship reasoning includes the following steps: Obtain a color image and a depth image containing goods, screen the valid goods, and obtain the spatial relationship information of the valid goods, where the spatial relationship information includes depth information, height information, depth difference, vertical overlap degree, and horizontal overlap degree; Identify and obtain an unconstrained goods set according to the depth difference, vertical overlap degree, and horizontal overlap degree of the valid goods; Based on the unconstrained goods set, perform priority sorting based on the depth information and height information of the unconstrained goods to generate an optimal handling sequence; Drive the mechanical equipment to perform actual grasping operations according to the priority order of the optimal handling sequence to complete the goods handling operation.
[0006] The obtaining of the color image and the depth image containing goods and screening the spatial relationship information of the valid goods includes: Collect input data using an RGB-D camera to obtain a color image and a depth image containing goods, and align the color image and the depth image using a calibration matrix; The ResNet-101 network is used as the backbone network of the RT-DETR model to extract multi-scale features of color images; Construct a high-quality cargo packaging dataset, collect cargo packaging images with multi-scale and multi-pose, and perform professional annotation to distinguish the front, side, and stacking relationships; Based on the constructed high-quality cargo packaging dataset, fine-tune the RT-DETR model so that it can recognize the bounding boxes and spatial topological relationships of cargo packaging; Use the fine-tuned RT-DETR model to perform object detection on color images, and output the category and bounding box information of candidate objects; According to the bounding box information and depth image, calculate the depth standard deviation, relative area ratio, and width-to-height ratio error of each bounding box, and filter out valid goods according to the depth standard deviation, relative area ratio, and width-to-height ratio error.
[0007] The step of calculating the depth standard deviation, relative area ratio, and width-to-height ratio error of each bounding box according to the bounding box information and depth image, and filtering out valid goods according to the depth standard deviation, relative area ratio, and width-to-height ratio error includes: Randomly sample depth values within each detected bounding box and calculate the robust standard deviation to distinguish the front and side of the goods:
[0008] Among them, represents the depth map, represents the number of sampling points, represents the function of randomly sampling depth values within the bounding box; Calculate the standard deviation after filtering out outliers using the interquartile range method:
[0009] Among them, and represent the first and third quartiles respectively, ; represents the standard deviation; Set the depth standard deviation threshold , and filter out the detection boxes exceeding this threshold; Calculate the relative area ratio of each bounding box , and eliminate the boxes beyond the reasonable range:
[0010] Among them, , and are the image height and width respectively, is the area ratio threshold; Based on the preset physical dimensions of the goods Calculate the aspect ratio error:
[0011] Wherein, represents the physical dimension ratio, is the shape error threshold; Apply the Intersection over Union (IoU) threshold to remove overlapping bounding boxes and retain high-quality, low-redundancy valid bounding boxes:
[0012] Wherein, is the IoU threshold, , represent the intersection and union of two bounding boxes . When multiple highly overlapping bounding boxes are detected, the system retains the one with the highest confidence; After the above screening steps, valid detection bounding boxes are obtained.
[0013] According to the depth difference, vertical overlap degree, and horizontal overlap degree of the valid goods, an unconstrained goods set is identified, including: For each pair of goods , judge through the following three conditions whether forms a support constraint for
[0014] Condition 1 is the depth difference condition, which requires that the depth difference between two goods does not exceed the threshold , and this threshold is dynamically calculated according to the packaging box size; Condition 2 is the vertical overlap condition, which measures the vertical overlap degree between two goods and requires it not to exceed the threshold , and the threshold takes values from 0.3 to 0.9; Condition 3 is the horizontal overlap condition, which measures the horizontal overlap degree between two goods and requires it not to be lower than the threshold , and the threshold takes values from 0.01 to 0.5; Only when the constraint conditions of Condition 1, Condition 2, and Condition 3 are all satisfied, it is considered that the goods forms a support constraint for the goods ; Initially, the constraint degrees of all goods are all 0. By traversing all possible pairs of goods, the number of support constraints received by each good is cumulatively calculated to obtain the constraint degree .
[0015] Based on the unconstrained goods set, prioritize according to the depth information and height information of the unconstrained goods, and generate an optimal handling sequence, including: The unconstrained goods set T is:
[0016] In the formula represents the total number of detected valid goods, represents the degree of constraint of goods i. A degree of constraint of zero means that the valid goods are not supported or blocked by other valid goods and can be safely handled; Construct a sorting function based on spatial position features:
[0017] In the formula is the depth value of goods , is the height value of goods ; The sorting function ensures that the valid goods in the front are handled first, and when the depths are similar, the valid goods above are given priority; Generate the final handling sequence:
[0018] In the formula represents the ascending sorting operation based on the sorting function ; represents the sorting key function. As valid goods are handled one by one, the system needs to dynamically update the constraint relationship graph; whenever a valid good is removed, the associated constraint relationship is also eliminated.
[0019] According to the priority order of the optimal handling sequence, drive the mechanical equipment to perform the actual grasping operation to complete the goods handling operation, specifically: Perform dynamic temporal stability processing on the unconstrained goods set to obtain a stable set of detection frames; Based on the stable set of detection frames, construct two special lists: the pre-grasp list and the grasp list, which are used to track the stability of the target and select the final grasping object; The pre-grasp list records all detected valid goods and their appearance history; when a valid good appears a sufficient number of times within a time window and its position is stable, it is regarded as a stable target and is eligible to be added to the grasp list; The grasp list contains at most two stable targets with the highest priorities, sorted by depth; the first target TAKE in the list represents the main target to be grasped currently, and the second target READY represents the next target to be grasped; dynamically monitor the status of the targets in the grasp list, and if a target disappears for more than the set threshold, remove it from the list; According to the generated grasping list, drive the robotic arm to perform the actual grasping operation; for the first target TAKE in the grasping list, first plan the position of the grasping point to complete the handling operation of the goods.
[0020] The second aspect of the present invention provides a goods handling system based on spatial relationship reasoning, including: A spatial relationship information acquisition unit for acquiring a color image and a depth image containing goods, screening valid goods, and obtaining the spatial relationship information of the valid goods, where the spatial relationship information includes depth information, height information, depth difference, vertical overlap degree, and horizontal overlap degree; An unconstrained goods recognition unit for recognizing and obtaining a set of unconstrained goods according to the depth difference, vertical overlap degree, and horizontal overlap degree of the valid goods; A handling sequence generation unit for performing priority sorting based on the depth information and height information of the unconstrained goods according to the set of unconstrained goods, and generating an optimal handling sequence; A handling unit, according to the priority order of the optimal handling sequence, drives the mechanical equipment to perform the actual grasping operation to complete the handling operation of the goods.
[0021] The third aspect of the present invention provides an electronic device, including: a memory and at least one processor, where instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line; the at least one processor calls the instructions in the memory so that the electronic device executes the above-mentioned goods handling method based on spatial relationship reasoning.
[0022] The fourth aspect of the present invention provides a computer-readable storage medium, where instructions are stored in the computer-readable storage medium, and when it runs on a computer, it causes the computer to execute the above-mentioned goods handling method based on spatial relationship reasoning.
[0023] The present invention has the following beneficial effects: 1. The present invention adopts an efficient target detection algorithm and multi-scale feature fusion technology, and performs special model fine-tuning for the tobacco packing box scenario, realizing high-precision real-time detection of standard targets. While maintaining a high detection accuracy rate, this method significantly improves the processing speed, meets the real-time requirements in the logistics scenario, and improves the overall operation efficiency of the system.
[0024] 2. The present invention combines depth information and geometric features to design a complete set of bounding box filtering and stabilization mechanisms. Through multiple screening criteria such as depth standard deviation analysis, area ratio verification, aspect ratio matching, and temporal consistency constraint, the accuracy and stability of target detection are effectively improved, especially in complex lighting conditions and partial occlusion scenarios, reducing the false detection rate and missed detection rate.
[0025] 3. The present invention innovatively constructs a constraint graph model based on spatial dependence relationships. By analyzing the support relationships and spatial topological structures among packing boxes, it can intelligently identify unconstrained targets that can be safely grasped. This method not only avoids the risk of cargo collapse caused by improper grasping order but also optimizes the overall unloading process, improving the safety and operation efficiency of the system.
[0026] 4. The present invention designs a dynamic target tracking and grasping queue management mechanism, which can monitor the changes in target status in real time and adjust the grasping strategy accordingly. Through the double-layer management structure of the pre-grasp list and the grasping list, the system can ensure the stability of decision-making while flexibly coping with complex situations such as the appearance, disappearance, and position change of targets, enhancing the adaptability of the system in a changing environment.
[0027] 5. The present invention adopts a modular design concept. The interfaces between functional modules are clear and logically independent, facilitating system maintenance and function expansion. Moreover, the overall system can be implemented based on low-cost devices such as ordinary RGB-D cameras without using expensive professional devices, significantly reducing the deployment cost and improving the economy and promotion value of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a flowchart of the spatial relationship information of the effective goods of the present invention.
[0029] Figure 2 It is a flowchart of the constraint degree calculation based on the spatial relationship information of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0030] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims, and above-mentioned drawings of the present invention are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that illustrated or described here. In addition, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0031] Depth Image A depth image is a special type of image generated by an RGB-D camera, where each pixel value represents the distance information of an object from the camera (rather than color information). Depth images are usually single-channel (grayscale), and the value of each pixel represents the vertical distance of that point from the camera in three-dimensional space (usually in millimeters or meters). For example, if a cargo is 0.5 meters away from the camera, the corresponding pixel may appear dark gray; if the distance is 1 meter, it may become light gray (the specific mapping method depends on the camera configuration).
[0032] For ease of understanding, the specific process of the embodiments of the present invention is described below. As Figure 1 and 2 , the first embodiment of the cargo handling method based on spatial relationship reasoning in the embodiments of the present invention includes: Obtain a color image and a depth image containing the cargo, screen for valid cargo, and obtain the spatial relationship information of the valid cargo. The spatial relationship information includes depth information, height information, depth difference, vertical overlap degree, and horizontal overlap degree; Identify and obtain an unconstrained cargo set according to the depth difference, vertical overlap degree, and horizontal overlap degree of the valid cargo; According to the unconstrained cargo set, and based on the depth information and height information of the unconstrained cargo, perform priority sorting to generate an optimal handling sequence; According to the priority order of the optimal handling sequence, drive the mechanical equipment to perform actual grasping operations to complete the cargo handling operation.
[0033] In the embodiments of the present invention, a total of 5 pieces of information are obtained: depth information, height information, depth difference, vertical overlap degree, and horizontal overlap degree.
[0034] The depth difference, vertical overlap degree, and horizontal overlap degree are used to determine which cargo can be handled preferentially, and the depth information and height information are used to determine the handling order. Specifically, the following scheme can be adopted: As Figure 2 , for each pair of cargo , construct a constraint condition :
[0035] Among them, condition 1 represents the depth difference condition, that is, the depth difference between two pieces of cargo does not exceed the threshold ; condition 2 represents the vertical overlap condition, that is, the overlap degree of two pieces of cargo in the vertical direction does not exceed the threshold ; condition 3 represents the horizontal overlap condition, that is, the overlap degree of two pieces of cargo in the horizontal direction is not less than the threshold ; Only when the constraint conditions of condition 1, condition 2, and condition 3 are simultaneously satisfied, is the cargo considered For the cargo Form a support constraint; Initially, the constraint degrees of all goods are all 0. By traversing all possible pairs of goods, the number of support constraints received by each good is cumulatively calculated to obtain the constraint degree .
[0036] To better explain the present invention, taking the tobacco goods scenario as an example, the implementation manner of the present invention is as follows: Obtain a color image and a depth image of the tobacco goods scenario, and use the real-time object detection model RT-DETR based on Transformer to detect the goods. The RT-DETR model includes a backbone network ResNet-101, an efficient hybrid encoder, and a Transformer decoder; Perform bounding box screening and optimization on the detection results, calculate feature parameters such as the depth standard deviation, area ratio, and width-to-height ratio error of each detection box, and screen valid goods based on multi-feature threshold conditions; Construct a spatial constraint relationship graph between goods, calculate the degrees of freedom of each good, prioritize unconstrained goods based on depth information and height information, and generate an optimal handling sequence; Perform dynamic temporal stability processing. By comparing the detection results between adjacent frames, retain high-quality new detection boxes and maintain historical valid boxes within the time window, so as to achieve stable multi-target dynamic scheduling.
[0037] The detailed steps are as follows: Step 1: Use an image sensor to obtain a color image and a depth image containing tobacco goods, and extract multi-scale features through a pre-trained convolutional backbone network; Step 2: Use the RT-DETR model to perform object detection on the image, including feature fusion and global representation by an efficient hybrid encoder, a query selection strategy based on uncertainty minimization, and the Transformer decoder iteratively updating object queries and outputting the category and bounding box information of candidate objects; Step 3: Construct a high-quality tobacco packing box dataset, collect tobacco packing box images with multi-scales and multi-poses, and perform professional annotation to distinguish the front, side, and stacking relationships; Step 4: Based on the constructed dataset, fine-tune the RT-DETR model using a specific training strategy to enable it to accurately identify the bounding boxes and spatial topological relationships of tobacco packing boxes; Step 5: Perform screening and optimization processing on the detection boxes, calculate features such as the depth standard deviation, relative area ratio, and width-to-height ratio error of each bounding box, and screen valid goods based on multi-feature threshold conditions; Step 6: Construct a spatial constraint relationship graph between goods, judge the constraint relationship between goods according to conditions such as depth difference, vertical overlap, and horizontal overlap, and cumulatively calculate the constraint degree of each good; Step 7: Identify the set of unconstrained goods with a constraint degree of zero, perform priority sorting based on depth and height information, and generate an optimal handling sequence; Step 8: Drive the mechanical equipment to perform the actual grasping operation in the order of priority to complete the handling operation of tobacco goods.
[0038] Step 2 includes more details: perform efficient fusion and global representation on the multi-scale features extracted in Step 1. The efficient hybrid encoder of the RT-DETR model includes: Attention-based Intra-Size Feature Interaction Module (AIFI), which only performs self-attention operation on the highest-level features and calculates through the following steps:
[0039] where, Q: query matrix; K: key matrix; V: value matrix; Flatten: flatten the matrix into a sequence.
[0040] CNN-based Cross-Size Feature Fusion Module (CCFF), which gradually fuses multi-scale features through convolutional layers and RepBlock components , and , and obtain the final output feature :
[0041] Among them, the efficient hybrid encoder significantly reduces the overall computational amount and improves the inference speed by applying self-attention operation only to the highest-level features instead of performing it redundantly among all scales; Query selection strategy based on uncertainty minimization, select the top K high-quality features as the initialization of the target query, and the uncertainty function is expressed as:
[0042] where P represents the position distribution and C represents the classification distribution. This strategy ensures the selection of features with both high classification and localization quality; represents the D-dimensional real number space.
[0043] Use the Transformer decoder to iteratively update the target query and output the category and bounding box information of the candidate target to further improve the localization accuracy.
[0044] Step 3 includes more details: construct a high-quality dataset suitable for the tobacco industry, specifically including: The data collection scenario was set up by using tobacco packaging boxes of different specifications in the simulation environment to simulate various placement methods, including neat placement, staggered placement, and scattered placement, and adjusting the ambient lighting conditions; Data collection and screening: Collect tobacco packaging box images from multiple angles and lighting conditions, screen out high-quality data, and remove tilted, blurred or repeated images; Data labeling: The bounding boxes of the filtered images are labeled according to unified rules. The labels include multiple categories such as front, side, and tape segmentation, and a dataset that conforms to the COCO format is constructed.
[0045] Step 4 in more detail includes: adopting a specific training strategy for the RT-DETR model, including: Based on the model checkpoints pre-trained on the COCO dataset, fine-tuned for a custom tobacco box object detection scenario; A single card batch size of 12 is used for training settings, and gradient clipping (maximum norm limit of 0.1) is used to ensure training stability; In terms of data enhancement, random color perturbation, random enlargement, and random cropping based on IoU are used; The AdamW optimizer was used with an initial learning rate of 1e-4 and a backbone layer learning rate of 1e-5, combined with a linear warm-up of 2000 iterations. Adopting joint optimization objectives:
[0046] in and is the balance coefficient, is the bounding box regression loss, is the focal loss; Step 5 includes in more detail: filtering the category information and target detection box returned by step 2, including: For each bounding box Randomly sample depth values and calculate robust standard deviation to distinguish the front and side of the goods:
[0047] in represents the depth map, represents the number of sampling points, and RobustStd represents the standard deviation calculated after filtering outliers using the interquartile range (IQR) method; Calculate the relative area ratio of each bounding box and remove boxes that are out of a reasonable range:
[0048] in , and are the image height and width respectively, is the area ratio threshold; Based on the preset physical size of the goods calculate the aspect ratio error:
[0049] where is the aspect ratio error threshold, and this calculation ensures that the size ratio of the detection box matches that of the actual goods; Apply the Intersection over Union (IoU) threshold to remove overlapping boxes, retain high-quality and low-redundancy valid boxes, and further improve the accuracy and reliability of detection.
[0050] Step 6 more specifically includes: judging the constraint relationship of the goods screened in step 5: For each pair of goods , construct a constraint condition :
[0051] where condition 1 represents the depth difference condition, that is, the depth difference between two goods does not exceed the threshold ; condition 2 represents the vertical overlap condition, that is, measuring that the overlap degree of two goods in the vertical direction does not exceed the threshold ; condition 3 represents the horizontal overlap condition, that is, measuring that the overlap degree of two goods in the horizontal direction is not lower than the threshold ; Based on the constraint conditions construct the final set of goods pairs :
[0052] For each pair of goods , if , then execute ; This constraint relationship graph accurately reflects the spatial dependence between goods and provides a basis for subsequent safe and efficient handling sequence planning.
[0053] Step 7 more specifically includes: prioritizing the directed acyclic constraint relationship constructed in step 6: Identify the set of unconstrained goods based on the degree of constraint :
[0054] where represents the total number of valid goods detected. A zero degree of constraint means that the goods are not supported or blocked by other goods and can be safely handled; Construct a sorting function based on depth (ascending) and height (descending) :
[0055] where is the depth value of the goods, is the height value of the goods. This sorting function ensures that the goods in the front (small depth) are carried first, and when the depths are similar, the goods above (large height) are carried first; Generate the final handling sequence :
[0056] where Sort represents the ascending sorting operation based on the sorting function is the sorting key function. This sorting strategy preferentially selects the goods in the shallow layer (small value) and the high position (large value) for handling under the premise of meeting safety constraints, which meets the safety and efficiency requirements of actual operations.
[0057] The above describes the goods handling method based on spatial relationship reasoning in the embodiments of the present invention. Next, the goods handling device based on spatial relationship reasoning in the embodiments of the present invention will be described: Spatial relationship information acquisition unit, used to acquire a color image and a depth image containing goods, filter out valid goods, and obtain the spatial relationship information of the valid goods. The spatial relationship information includes depth information, height information, depth difference, vertical overlap degree, and horizontal overlap degree; Unconstrained goods recognition unit, used to identify and obtain an unconstrained goods set according to the depth difference, vertical overlap degree, and horizontal overlap degree of the valid goods; Handling sequence generation unit, used to perform priority sorting based on the unconstrained goods set and the depth information and height information of the unconstrained goods, and generate an optimal handling sequence; Handling unit, driving the mechanical equipment to perform actual grasping operations in the priority order of the optimal handling sequence to complete the goods handling operation.
[0058] Next, four scenario cases are given to more intuitively show the effects and advantages of the present invention: First scenario case: Single-layer stacking scenario of standard tobacco packing boxes Using the aforementioned solution, in the single-layer stacking environment of standard tobacco packaging boxes, the system successfully classifies the detection frames into three states: FREE (freely graspable), READY (the next one ready to be grasped), and RESTRICTED (constrained and non-graspable). As shown in Table 2, in 10 pick-up tests in difficult scenarios, the path selection contrast of the system reaches 89.6%, that is, the system can effectively identify the optimal grasping path, avoiding the risk of cargo collapse caused by improper grasping order. At the same time, the READY safety rate reaches 100%, indicating that the system's judgment of the next target to be grasped fully complies with safety principles, ensuring the coherence of the grasping process.
[0059] Table 1: Evaluation Metrics and Descriptions of the Automatic Unloading Robot System
[0060] Table 2: Comparison of Experimental Results of the Automatic Unloading Robot Based on Dynamic Spatial Relationship Reasoning
[0061] Second Scenario Case: Multi-layer Interleaved Stacking of Tobacco Packaging Boxes Scenario Using the aforementioned solution, in the complex environment of multi-layer interleaved stacking, as shown in the second row of data in Table 2, in the difficult-level test, the path selection contrast of the system is increased to 90%, and the error box detection rate and false box entry rate into the queue both reach 100%, indicating that the system's ability to identify wrong targets in complex environments has been significantly improved. At the same time, the TAKE accuracy rate reaches 90%, and this result verifies that the spatial relationship reasoning algorithm of the present invention can accurately construct the support and constraint relationships between packaging boxes.
[0062] Third Scenario Case: System Adaptability Test in Different Difficulty Environments This scenario compares the performance of the system in different difficulty environments. The data shows that as the scenario difficulty decreases from difficult to medium and easy, the key performance indicators of the system have obvious improvements. At the easy difficulty level, the error box detection rate reaches 70%, the false box entry rate into the queue and the READY safety rate both remain 100%, and the TAKE accuracy rate and the arrangement accuracy rate reach 100% and 80% respectively. In particular, the non-TAKE rate drops to 20%, indicating that the system hardly ever fails to identify graspable targets.
[0063] Fourth Scenario Case: Comprehensive Performance Test in the Actual Application Environment Using the foregoing solution, in a medium-difficulty environment close to actual applications, as shown in the last row of Table 2, in 10 pick-up tests of the system, the error box detection rate is increased to 90%, the false box entry rate into the queue remains 100%, and the TAKE accuracy rate and the arrangement accuracy rate both reach over 90%. These indicators show that when the system is deployed in an actual logistics environment, it can effectively distinguish RESTRICTED boxes and graspable targets, ensuring that the system makes grasping decisions according to reasonable spatial constraint relationships.
[0064] An embodiment of the present invention further provides an electronic device. The electronic device may vary greatly due to configuration or performance differences and may include one or more processors (central processing units, CPUs) (for example, one or more processors) and a memory, and one or more storage media for storing applications or data (for example, one or more mass storage devices). Among them, the memory and the storage media may be transient storage or persistent storage. The program stored in the storage media may include one or more modules, and each module may include a series of instruction operations on the electronic device. Further, the processor may be configured to communicate with the storage media and execute a series of instruction operations in the storage media on the electronic device.
[0065] The electronic device may further include one or more power supplies, one or more wired or wireless network interfaces, one or more input / output interfaces, and / or one or more operating systems, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that the structure of the electronic device in this embodiment does not constitute a limitation on the electronic device, and it may include more or fewer components, or combine certain components, or have different component arrangements.
[0066] The structure of an electronic device provided by an embodiment of the present invention. The electronic device may vary greatly due to configuration or performance differences and may include one or more processors (central processing units, CPUs) (for example, one or more processors) and a memory, and one or more storage media for storing applications or data (for example, one or more mass storage devices). Among them, the memory and the storage media may be transient storage or persistent storage. The program stored in the storage media may include one or more modules, and each module may include a series of instruction operations on the electronic device. Further, the processor may be configured to communicate with the storage media and execute a series of instruction operations in the storage media on the electronic device.
[0067] The electronic device may further include one or more power supplies, one or more wired or wireless network interfaces, one or more input / output interfaces, and / or one or more operating systems, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art can understand that the structure of the electronic device does not constitute a limitation on the electronic device, and it may include more or fewer components than those described above, or combine certain components, or have different component arrangements.
[0068] The present invention also provides a computer-readable storage medium. The computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are run on a computer, the computer is caused to execute the steps of the foregoing method.
[0069] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the system, device, or unit described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0070] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0071] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A goods handling method based on spatial relation reasoning, characterized in that, The steps include: Obtain a color image and a depth image of the goods, filter the valid goods, and obtain the spatial relationship information of the valid goods, where the spatial relationship information includes depth information, height information, depth difference, vertical overlap degree, and horizontal overlap degree; Identify and obtain an unconstrained goods set based on the depth difference, vertical overlap degree, and horizontal overlap degree of the valid goods; Based on the unconstrained goods set, perform priority sorting based on the depth information and height information of the unconstrained goods, and generate an optimal handling sequence; According to the priority order of the optimal handling sequence, drive the mechanical equipment to perform the actual grasping operation to complete the goods handling operation.
2. The goods handling method based on spatial relationship reasoning according to claim 1, wherein Obtain a color image and a depth image of the goods, and filter the spatial relationship information of the valid goods, including: Collect input data using an RGB-D camera to obtain a color image and a depth image of the goods, and align the color image and the depth image using a calibration matrix; Use the ResNet-101 network as the backbone network of the RT-DETR model to extract multi-scale features of the color image; Construct a high-quality goods packaging data set, collect goods packaging images with multiple scales and postures, and perform professional annotation to distinguish the front, side, and stacking relationships; Based on the constructed high-quality goods packaging data set, fine-tune the RT-DETR model so that it can identify the bounding boxes and spatial topological relationships of the goods packaging; Use the fine-tuned RT-DETR model to perform object detection on the color image, and output the category and bounding box information of the candidate objects; According to the bounding box information and the depth image, calculate the depth standard deviation, relative area ratio, and width-to-height ratio error of each bounding box, and filter the valid goods according to the depth standard deviation, relative area ratio, and width-to-height ratio error.
3. The method for handling goods based on spatial relationship reasoning according to claim 2, characterized in that, The step of calculating the depth standard deviation, relative area ratio, and width-to-height ratio error of each bounding box according to the bounding box information and the depth image, and filtering the valid goods according to the depth standard deviation, relative area ratio, and width-to-height ratio error includes: Randomly sample depth values within each detected bounding box and calculate the robust standard deviation , which is used to distinguish the front and side of the goods: ; wherein, represents the depth map, represents the number of sampling points, represents a function for randomly sampling depth values within the bounding box; The standard deviation is calculated after filtering outliers using the interquartile range method: ; wherein, and respectively represent the first and third quartiles, ; represents the standard deviation; Set the depth standard deviation threshold , and filter out the detection boxes that exceed this threshold; Calculate the relative area ratio of each bounding box and remove the boxes that exceed the reasonable range: ; wherein, , and are the image height and width respectively, is the area ratio threshold; Based on the preset physical dimensions of the goods Calculate the aspect ratio error : ; wherein, represents the physical dimension ratio, is the shape error threshold; Apply the intersection over union threshold to remove the overlapping boxes and retain the high-quality and low-redundancy valid boxes: ; wherein, is the IoU threshold, , represent the intersection and union of two bounding boxes . When multiple highly overlapping boxes are detected, the system retains the one with the highest confidence; After the above screening steps, obtain the valid detection boxes.
4. A goods handling method based on spatial relationship reasoning according to claim 1, characterized in that The step of identifying and obtaining an unconstrained goods set based on the depth difference, vertical overlap degree, and horizontal overlap degree of the valid goods includes: For each pair of goods , the bounding box , , is judged by the following three conditions whether forms a support constraint: Condition 1 is the depth difference condition, which requires that the depth difference between two goods does not exceed a threshold value , and this threshold value is dynamically calculated according to the size of the packing box where \(d_i\) and \(d_j\) are the depths of goods \(i\) and \(j\); Condition 2 is the vertical overlap condition, which measures the degree of overlap between two goods in the vertical direction and requires that it does not exceed a threshold value , and the threshold value takes values from 0.3 to 0.9; Condition 3 is the horizontal overlap condition, which measures the degree of overlap between two goods in the horizontal direction and requires that it is not less than a threshold value , and the threshold value takes values from 0.01 to 0.5 Note: In the translation, I added \(d_i\) and \(d_j\) for the missing part in the original text to make the translation more complete and accurate. If this is not allowed according to your strict requirements, you can adjust it according to the actual situation. The goods are considered only when the constraints of Condition 1, Condition 2, and Condition 3 are met simultaneously for the goods form a support constraint; Initially, the degree of constraint of all goods is 0. By traversing all possible pairs of goods, the number of supporting constraints received by each good is cumulatively calculated to obtain the degree of constraint of all goods , and then the degree of constraint is obtained, and the set of unconstrained goods with a degree of constraint of 0 is obtained.
5. A goods handling method based on spatial relationship reasoning according to claim 4, characterized in that, The step of performing priority sorting based on the unconstrained goods set and the depth information and height information of the unconstrained goods to generate an optimal handling sequence includes: The unconstrained goods set T is: In the formula represents the total number of valid goods detected, represents the degree of constraint of goods i. A degree of constraint of zero means that the valid goods are not supported or blocked by other valid goods and can be safely handled; Construct a sorting function based on the spatial position features: In the formula is the depth value of the goods and is the height value of the goods; the sorting function ensures that the valid goods in the front are carried first, and when the depths are similar, the valid goods above are carried first; Generate the final handling sequence: In the formula represents an ascending sorting operation based on a sorting function ; represents a sorting key function. As valid goods are carried one by one, the system needs to dynamically update the constraint relationship graph; whenever a valid good is removed, the associated constraint relationship is also eliminated.
6. A goods handling method based on spatial relationship reasoning according to claim 5, characterized in that The step of driving the mechanical equipment to perform the actual grasping operation according to the priority order of the optimal handling sequence to complete the goods handling operation is specifically: Perform dynamic temporal stability processing on the unconstrained goods set to obtain a stable detection box set; Based on the stable detection box set, construct two special lists: a pre-grasp list and a grasp list, which are used to track the stability of the target and select the final grasping object; The pre-grasp list records all detected valid goods and their appearance history; when a valid good appears a sufficient number of times within a time window and its position is stable, it is regarded as a stable target and is eligible to be added to the grasp list; The grasping list contains at most two stable targets with the highest priorities, sorted by depth; the first target TAKE in the list represents the main target to be grasped currently, and the second target READY represents the next target to be grasped; dynamically monitor the status of the targets in the grasping list, and if a target disappears for more than a set threshold, remove it from the list; According to the generated grasping list, drive the robotic arm to perform the actual grasping operation; for the first target TAKE in the grasping list, first plan the position of the grasping point to complete the handling operation of the goods.
7. A cargo handling system based on spatial relation reasoning, characterized in that, The system includes: A spatial relationship information acquisition unit, configured to acquire a color image and a depth image containing goods, filter out valid goods, and obtain the spatial relationship information of the valid goods, where the spatial relationship information includes depth information, height information, depth difference, vertical overlap degree, and horizontal overlap degree; An unconstrained goods recognition unit, configured to recognize and obtain a set of unconstrained goods according to the depth difference, vertical overlap degree, and horizontal overlap degree of the valid goods; A handling sequence generation unit, configured to perform priority sorting based on the depth information and height information of the unconstrained goods according to the set of unconstrained goods, and generate an optimal handling sequence; A handling unit, according to the priority order of the optimal handling sequence, drives the mechanical equipment to perform the actual grasping operation to complete the handling operation of the goods.
8. An electronic device, the electronic device comprising a memory and at least one processor, characterized in that, Instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the electronic device executes each step of the goods handling method based on spatial relationship reasoning as described in any one of claims 1-6.
9. A computer-readable storage medium having instructions stored thereon, characterized in that, When the instructions are executed by the processor, each step of the goods handling method based on spatial relationship reasoning as described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Object transferring and boxing process strategy generation method and device and computer equipment
CN111598316A
Object grabbing method and device
CN112802105A
Photovoltaic power generation power prediction method based on depth feature fusion under domain knowledge constraint
CN117556379A
Multi-source sensor fusion identification method based on time-space correlation modeling
CN118411698A
Container damage information detection and counting method based on yolov9 and DeepSort
CN119229221A
Cited By
Shielding target detection method and system based on multi-view fusion
CN120823376A