A Three-dimensional Space Precise Positioning Method for Suspended Objects
Through the rotation target detection and depth completion algorithm, the problem of inaccurate three-dimensional positioning of the lifting object in the traditional method is solved, and the high-precision positioning of the lifting object in the three-dimensional space is realized, and the automation and intelligence of crane lifting operations are supported.
Patent Information
- Application Number
- CN202510502775.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Traditional object detection methods cannot effectively calculate the position and posture of the hanging object in three-dimensional space, resulting in incomplete positioning information dimensions. The existing depth camera-based methods cause incomplete or distorted depth maps due to hardware limitations and lighting changes, which affect the accuracy of positioning results.
The rotation object detection and depth completion algorithm are used to obtain the angle and position information of the hanging object through rotation object detection, and the depth completion algorithm is used to complete the missing or incomplete depth information to achieve three-dimensional precise positioning of the hanging object.
It realizes all-round and high-precision positioning of the lifting objects in three-dimensional space, improves the integrity and accuracy of positioning, provides key data support for crane lifting operations, and improves the safety and operation efficiency of industrial lifting operations.
Smart Images

Figure CN120014057B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and object depth completion, and particularly relates to a method for accurately positioning a suspended object in three-dimensional space. Background Art
[0002] In the industrial hoisting operation scenario, the accurate three-dimensional space positioning of the suspended object is the core technical requirement for ensuring operation safety and operation efficiency. However, traditional object detection methods usually focus on object recognition and positioning in a two-dimensional plane, ignoring the position and attitude information of the object in three-dimensional space. For the specific scenario of the suspended object, traditional object detection methods cannot effectively calculate the attitude of the suspended object, that is, the rotation angle in the two-dimensional plane and the three-dimensional space position, resulting in incomplete dimensional positioning information and being difficult to meet the automation operation requirements of the crane hoisting operation. In addition, existing three-dimensional positioning methods based on depth cameras can obtain the depth information of the scene, thereby obtaining the position of the suspended object in three-dimensional space. However, due to the limitations of camera hardware, changes in lighting conditions, and the influence of reflective materials on light, depth maps often have incomplete or distorted problems. For example, the depth information in some areas may be empty, resulting in inaccurate acquisition of three-dimensional positioning in these areas and thus affecting the accuracy of the overall positioning result.
[0003] Therefore, it is necessary to provide a method for accurately positioning a suspended object in three-dimensional space to break through the key technical barriers for the intelligent upgrade of industrial hoisting. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a method for accurately positioning a suspended object in three-dimensional space. This method uses rotational object detection to obtain information such as the angular position and category of the suspended object, and uses a depth completion algorithm to complete missing or incomplete depth information, thereby making the three-dimensional positioning of the suspended object more accurate.
[0005] The technical solution of the present invention is: a method for accurately positioning a suspended object in three-dimensional space, including the following steps:
[0006] Step 1: Use a depth camera to obtain the RGB image and depth map of the suspended object from a front top view, and perform registration, that is, ensure that each point in the depth map can be accurately mapped to the corresponding point in the RGB image;
[0007] Step 2: Input the RGB image of the suspended object into a trained rotational object detection model to detect the rotational bounding box of the suspended object. The geometric center point of the bounding box is (x, y), and its rotation angle is θ;
[0008] Step 3: Map the rotational bounding box of the suspended object obtained from the RGB image of the suspended object to the depth map, thereby determining the position of the rotational bounding box in the depth map;
[0009] Step 4: Check whether there are missing depth values within the rotated bounding box in the depth map, that is, check whether the depth values of n points near the geometric center point within the rotated bounding box of the depth map are zero; if there are points with a depth value of 0, proceed to Step 5 for depth completion, and if the depth values of all points are not 0, directly proceed to Step 6;
[0010] Step 5: Send the missing depth map into the trained depth completion model for depth completion to obtain a complete depth map;
[0011] Step 6: Randomly select n points near the geometric center point within the rotated bounding box of the suspended object in the depth map, and calculate the average value of the depth values of these points as the true depth value of the suspended object, that is, the distance z between the suspended object and the depth camera;
[0012] Step 7: Based on the position (x, y, z) and the corresponding rotation angle θ of the suspended object in the three-dimensional space obtained in Step 6, the precise positioning of the suspended object in the three-dimensional space is achieved.
[0013] Further, in the above Step 2, the training process in the rotation target detection model includes:
[0014] Obtain the RGB image of the suspended object from the front overhead view, and then annotate the RGB image of the suspended object as a training sample;
[0015] Perform data augmentation on the obtained training samples;
[0016] Construct a rotation target detection model;
[0017] Input the training samples after data augmentation into the constructed rotation target detection model, and train based on the loss of the rotation target detection model.
[0018] The format of the rotation target annotation box for the RGB image of the suspended object is [classid, x, y, longside, shortside, θ];
[0019] classid is the category of the suspended object;
[0020] x, y are the center point coordinates of the rectangle;
[0021] longside, shortside are the long side and the short side of the rectangle;
[0022] θ is the angle of the rectangle, and the θ angle is defined as the angle between the long side of the rectangle and the positive x-axis direction, negative counterclockwise and positive clockwise.
[0023] Further, the data augmentation includes selecting four different RGB images for Mosaic augmentation, and stitching them into a new image by means of scaling, cropping and random arrangement. Further, Mixup augmentation is used to perform linear interpolation on paired images and their labels during training to generate virtual training samples. Furthermore, Copypaste augmentation is used to randomly copy and paste the input images, and the richness of the training data is increased by pasting different objects of different sizes onto a new background image.
[0024] Further, the construction of the rotation object detection model includes:
[0025] The input end is used to receive the RGB image of the suspended object;
[0026] The backbone network consists of a Focus module, a CBH module, a CSP module, and an SPP module, and is used to extract features at different levels of the input image;
[0027] The neck network consists of an upsampling module, a CBH module, and a CSP module, and is used to fuse features at different levels;
[0028] And an output layer composed of convolutional layers, which is used to output the category, center point coordinates, width, height, confidence, and rotation angle of the predicted target.
[0029] Further, it also includes calculating the loss of the rotation object detection model, and the loss of the rotation object detection model includes classification loss , confidence loss , bounding box regression loss and rotation angle loss ;
[0030] Among them, the classification loss is:
[0031]
[0032] In the formula is the predicted class label, is the true class label, and N represents the total number of classes;
[0033] The confidence loss is:
[0034]
[0035] represents the predicted confidence of the prediction box, represents the true confidence of the prediction box. When there is a target object in the prediction box, then = 1. When there is no target object in the prediction box, then = 0, and N represents the total number of prediction boxes;
[0036] Bounding box regression loss is as follows:
[0037]
[0038] In the formula, is and 's intersection over union, is the predicted bounding box, is the ground truth bounding box;
[0039] is the Euclidean distance between the centers of the predicted bounding box and the ground truth bounding box;
[0040] is the diagonal length of the smallest bounding rectangle containing the predicted bounding box and the ground truth bounding box;
[0041] is the balance parameter, used to balance and the aspect ratio loss;
[0042] is a measure of the difference in aspect ratio between the predicted bounding box and the ground truth bounding box, is expressed as:
[0043]
[0044] where is the width of the ground truth bounding box, is the height of the ground truth bounding box, is the width of the predicted bounding box, is the height of the predicted bounding box;
[0045] Rotation angle loss is as follows:
[0046]
[0047] In the formula, is the ground truth angle, is the predicted angle, and N represents the total number of angle categories;
[0048] is defined as converting the ground truth angle in the sling label into a one-dimensional array T of length 180 using a Gaussian function; where the value of the array T at the angle is 1, and as the angle changes from to the left and right, the value of the array T continuously decreases until it becomes 0;
[0049] The conversion formula for converting the angle true value into a one-dimensional array of 180 degrees in length is:
[0050]
[0051] In the formula, is a window function, and the radius of the window function represents the true angle of the current diagonal frame. Here, a Gaussian function with a standard deviation of 2 is used as the window function; is the true angle value, is the angle range.
[0052] Furthermore, the training process of the depth completion network includes:
[0053] Obtain the RGB image and the true depth image of the suspended object taken from directly above at high altitude as training samples;
[0054] Construct a depth completion model, including a self-depth completion module and an RGB-guided completion module;
[0055] Input the training samples into the constructed self-depth completion module and train based on the loss ;
[0056] Fix the parameters of the trained self-depth completion module, and based on the loss , use the dynamic gradient adjustment strategy to train the RGB-guided completion module.
[0057] Furthermore, the self-depth completion module uses the original depth map to generate a preliminary depth completion map; and the self-depth completion module consists of an encoder, a decoder, and a cross-scale attention block;
[0058] The RGB-guided completion module further optimizes the preliminary depth completion map under the guidance of the RGB image to generate the final depth completion map; and the RGB-guided completion module consists of a CNN, a Transformer, and a cross-modal attention module.
[0059] Furthermore, The calculation formula of is:
[0060]
[0061] In the formula, is the true depth image of the suspended object, is the preliminary depth completion image output by the self-depth completion module, and are balance parameters;
[0062] The calculation formula of is:
[0063]
[0064] In the formula, is the final depth image output by the RGB-guided completion module.
[0065] Furthermore, the gradient dynamic adjustment algorithm is used to update the parameters of the RGB-guided completion module, specifically including:
[0066] Calculate the gradient of the parameters of the feature extraction network and the gradient of the parameters of the RGB feature extraction network : the gradient of the parameters :
[0067]
[0068]
[0069] The ratio of the L2 norm of the model parameter gradients is used to calculate the gradient difference ratio of the model parameters corresponding to each modality : :
[0070]
[0071]
[0072] where is the L2 norm;
[0073] Calculate the adjustment factor of the parameters of the feature extraction network and the adjustment factor of the parameters of the RGB feature extraction network :
[0074]
[0075]
[0076] where is the adjustment coefficient;
[0077] Use the gradient dynamic adjustment algorithm to update the parameters of the feature extraction network and the parameters of the RGB feature extraction network :
[0078]
[0079]
[0080] In the formula, is the model parameter of the feature extraction network at the th iteration, is the model parameter of the RGB feature extraction network at the th iteration, is the learning rate, and are the first exponential decay rate and the second exponential decay rate respectively, is the th power of the first exponential decay rate is the th power of the second exponential decay rate , , is the th iteration parameter adjustment factor.
[0081] The beneficial technical effects of the present invention are as follows:
[0082] 1. The method of the present invention can calculate the complete three-dimensional spatial position (x, y, z) and rotation angle θ of the suspended object, realizing all-round and high-precision positioning, and greatly improving the integrity of positioning.
[0083] 2. The present invention adopts a depth completion method, which can more efficiently utilize the complementary information of RGB images and depth images, making the measured depth value more accurate and significantly improving the depth completion effect.
[0084] 3. Obtaining the accurate position and attitude information of the suspended object in three-dimensional space completely provides key data support for the automation of crane lifting operations, making up for the defect that traditional target detection methods cannot meet the high-dimensional requirements of positioning information for automated operations, and can effectively improve the safety and operation efficiency of industrial lifting operation scenarios, and promote the development of lifting operations towards automation and intelligence.
[0085] The above description is only an overview of the technical solution of the present invention. In order to understand the technical means of the present invention more clearly and implement it according to the content of the specification, the following takes the preferred embodiments of the present invention and combines with the attached drawings to describe in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 is a schematic diagram of the rotation target detection method of the present invention;
[0087] Figure 2 is a schematic diagram of the camera position of the present invention;
[0088] Figure 3 Schematic diagram of the process of the present invention;
[0089] Figure 4 Schematic diagram of the rotation target detection model of the present invention;
[0090] Figure 5 Schematic diagram of the depth completion network model of the present invention;
[0091] Figure 6 Schematic diagram of the self-depth completion module of the present invention;
[0092] Figure 7 Schematic diagram of the RGB-guided completion module of the present invention;
[0093] Figure 8 Depth value near the geometric center point of the depth map rotation bounding box of the present invention. Detailed implementation manners
[0094] In order to be able to more clearly understand the technical means of the present invention and implement it in accordance with the content of the specification, the following combines the accompanying drawings and embodiments to further describe in detail the specific implementation manners of the present invention. The following embodiments are used to illustrate the present invention but are not used to limit the scope of the present invention.
[0095] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned accompanying drawings are used to distinguish similar objects and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of the present application described herein.
[0096] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship recorded in the embodiments and shown in the accompanying drawings, or the orientation or positional relationship in which the invention product is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention.
[0097] The present invention specifically relates to a method for accurately positioning a suspended object in a three-dimensional space, including the following steps:
[0098] As Figures 1 - 3 shown, step 1, use a depth camera to obtain the RGB image and depth map of the suspended object from a directly overhead perspective, and perform registration, that is, ensure that each point in the depth map can be accurately mapped to the corresponding point in the RGB image;
[0099] Step 2: Input the RGB image of the suspended object into the trained rotation object detection model to detect the rotation bounding box of the suspended object. The geometric center point of the bounding box is (x, y), and its rotation angle is θ;
[0100] Further, in the above Step 2, the training process in the rotation object detection model includes:
[0101] Obtain the RGB image of the suspended object from the positive top-down view, and then annotate the RGB image of the suspended object as a training sample;
[0102] Perform data augmentation on the obtained training samples;
[0103] Construct a rotation object detection model;
[0104] Input the training samples after data augmentation into the constructed rotation object detection model, and train based on the loss of the rotation object detection model.
[0105] The format of the rotation object annotation box for the RGB image of the suspended object is [classid, x, y, longside, shortside, θ];
[0106] classid is the category of the suspended object;
[0107] x, y are the center point coordinates of the rectangle;
[0108] longside, shortside are the long side and short side of the rectangle;
[0109] θ is the angle of the rectangle. The θ angle is defined as the angle between the long side of the rectangle and the positive x-axis direction, negative counterclockwise and positive clockwise.
[0110] Further, the data augmentation includes selecting four different RGB images for Mosaic augmentation, and splicing them into a new image by means of scaling, cropping and random arrangement. Further, Mixup augmentation is used to perform linear interpolation on paired images and their labels during training to generate virtual training samples. Furthermore, Copypaste augmentation is used to randomly copy and paste the input images, and the richness of the training data is increased by pasting different objects of different sizes onto a new background image.
[0111] Further, as Figure 4 shown, the construction of the rotation object detection model includes:
[0112] An input end, used to receive the RGB image of the suspended object;
[0113] A backbone network, composed of a Focus module, a CBH module, a CSP module, and an SPP module, used to extract features of different levels of the input image;
[0114] The neck network, consisting of an upsampling module, a CBH module, and a CSP module, is used to fuse features at different levels;
[0115] and an output layer composed of convolutional layers, which is used to output the category, center point coordinates, width, height, confidence, and rotation angle of the predicted target.
[0116] Furthermore, it also includes calculating the loss of the rotation object detection model, and the loss of the rotation object detection model includes classification loss , confidence loss , bounding box regression loss and rotation angle loss ;
[0117] Among them, the classification loss is:
[0118]
[0119] In the formula is the predicted class label, is the true class label, and N represents the total number of classes;
[0120] The confidence loss is:
[0121]
[0122] represents the predicted confidence of the prediction box, represents the true confidence of the prediction box. When there is an object in the prediction box, then = 1, and when there is no object in the prediction box, then = 0, and N represents the total number of prediction boxes;
[0123] The bounding box regression loss is:
[0124]
[0125] In the formula, is and 's intersection over union, is the prediction box, is the true box;
[0126] is the Euclidean distance between the center points of the prediction box and the true box;
[0127] is the diagonal length of the smallest bounding rectangle containing the prediction box and the true box;
[0128] is a balance parameter used to balance and the aspect ratio loss;
[0129] is a measure of the difference in aspect ratio between the predicted bounding box and the ground truth bounding box, expressed as:
[0130]
[0131] where is the width of the ground truth bounding box, is the height of the ground truth bounding box, is the width of the predicted bounding box, is the height of the predicted bounding box;
[0132] The rotation angle loss is:
[0133]
[0134] In the formula, is the ground truth angle, is the predicted angle, and N represents the total number of angle categories;
[0135] is defined as converting the ground truth angle in the lifting object label into a one-dimensional array T of length 180 using a Gaussian function; where the value of the array T at the angle is 1, and as the angle changes from to the left and right, the value of the array T continuously decreases until it becomes 0;
[0136] The conversion formula for converting the angle ground truth into a one-dimensional array of 180 degrees in length is:
[0137]
[0138] In the formula, is a window function, and the radius of the window function represents the ground truth angle of the current oblique bounding box. Here, a Gaussian function with a standard deviation of 2 is used as the window function; is the ground truth angle value, is the angle range.
[0139] Step 3, Map the lifting object rotation bounding box obtained from the lifting object RGB image to the depth map to determine the position of the rotation bounding box in the depth map;
[0140] As Figure 8As shown, in step 4, check whether there are missing depth values within the rotated bounding box in the depth map, that is, check whether the depth values of n points near the geometric center point within the rotated bounding box of the depth map are zero; if there are points with a depth value of 0, proceed to step 5 for depth completion, and if the depth values of all points are not 0, directly proceed to step 6;
[0141] As Figure 8 shown, randomly select the depth values of n points near the geometric center point of the rotated bounding box of the depth map and take the average to obtain the true depth value of the suspended object; in the figure, the red dot is the depth value of the geometric center point of the rotated bounding box of the depth map.
[0142] Step 5: Send the missing depth map into the trained depth completion model for depth completion to obtain a complete depth map;
[0143] When there are missing depth values within the rotated bounding box of the depth map, the registered depth map and RGB map are passed into the pre-trained depth completion network for depth value completion to obtain the completed depth map, and then the rotated bounding box of the suspended object is matched to the completed depth map. Randomly select the depth values of n points near the geometric center point of the rotated bounding box of the depth map and take the average to obtain the true depth value of the suspended object;
[0144] When there are no missing depth values within the rotated bounding box of the depth map, randomly select the depth values of n points near the geometric center point of the bounding box of the depth map and take the average to obtain the true depth value of the suspended object.
[0145] Furthermore, the training process of the depth completion network includes:
[0146] Obtain the RGB map and true depth map of the suspended object taken from directly above at high altitude as training samples;
[0147] As Figure 5 shown, construct a depth completion model, including a self-depth completion module and an RGB-guided completion module;
[0148] Input the training samples into the constructed self-depth completion module and train based on the loss for training;
[0149] Fix the parameters of the trained self-depth completion module, and based on the loss , use the dynamic gradient adjustment strategy to train the RGB-guided completion module.
[0150] Furthermore, as Figure 6 shown, the self-depth completion module uses the original depth map to generate a preliminary depth completion map; and the self-depth completion module consists of an encoder, a decoder, and a cross-scale attention block;
[0151] As Figure 7As shown, the RGB-guided completion module further optimizes the preliminary depth completion map under the guidance of the RGB map to generate the final depth completion map; and the RGB-guided completion module consists of a CNN, a Transformer, and a cross-modal attention module;
[0152] The cross-scale attention block solves the problem of completing large missing areas by guiding each other between features of different scales (such as high-resolution features guiding low-resolution feature upsampling, and low-resolution features completing missing areas in high-resolution regions), improving the depth map completion effect;
[0153] The cross-modal attention module is used to achieve feature fusion between the depth and RGB images of the suspended object, and fully utilize the information in the RGB image to complete the depth map.
[0154] Furthermore, The calculation formula of
[0155]
[0156] In the formula, is the true depth image of the suspended object, is the preliminary depth completion image output by the depth completion module, and are balance parameters;
[0157] The calculation formula of
[0158]
[0159] In the formula, is the final depth image output by the RGB-guided completion module.
[0160] Furthermore, the gradient dynamic adjustment algorithm is used to update the parameters of the RGB-guided completion module, specifically including:
[0161] Calculate The gradient of the parameters of the feature extraction network and the gradient of the parameters of the RGB feature extraction network : :
[0162]
[0163]
[0164] The gradient difference ratio of the model parameters corresponding to each modality is calculated using the ratio of the L2 norms of the model parameter gradients of :
[0165]
[0166]
[0167] Wherein, is the L2 norm;
[0168] Calculate the adjustment factor of the feature extraction network parameters and the adjustment factor of the RGB feature extraction network parameters : :
[0169]
[0170]
[0171] Wherein, is the adjustment coefficient;
[0172] Update the parameters of the feature extraction network and the parameters of the RGB feature extraction network using the gradient dynamic adjustment algorithm: :
[0173]
[0174]
[0175] Wherein, is the th iteration model parameters of the feature extraction network, is the th iteration model parameters of the RGB feature extraction network, is the learning rate, and are the first exponential decay rate and the second exponential decay rate respectively, is the th power of the first exponential decay rate, is the th power of the second exponential decay rate, is the th iteration parameter
[0176] Reducing the convergence speed difference between the depth map modality and the RGB color map modality using direct gradient dynamic modulation, effectively reducing the heterogeneity between different modalities, making full use of the complementarity of different modalities, and obtaining rich feature information.
[0177] Using a two-stage depth completion task, each stage having its specific goals and methods. Compared with other completion algorithms, through this phased approach, different types of features can be utilized more effectively: such as the self-information of the original depth image and the rich information of the RGB image, and the accuracy and robustness of completion can be improved. At the same time, using the direct gradient dynamic adjustment algorithm, the heterogeneity between different modalities is effectively reduced, the complementarity of different modalities is fully utilized, and rich feature information is obtained.
[0178] Step 6: In the depth map, randomly select n points near the geometric center point within the lifting object rotation bounding box, and calculate the average value of the depth values of these points as the true depth value of the lifting object, that is, the distance z between the lifting object and the depth camera;
[0179] Step 7: According to the position (x, y, z) and the corresponding rotation angle θ of the lifting object in the three-dimensional space obtained in Step 6, the precise positioning of the lifting object in the three-dimensional space is realized.
[0180] The above embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: Any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention.
Claims
1. A three-dimensional space precise positioning method for suspended objects, characterized in that, It includes the following steps: Step 1: Use a depth camera to obtain the RGB image and depth map of the suspended object from a directly overhead perspective, and perform registration, that is, ensure that each point in the depth map can be accurately mapped to the corresponding point in the RGB image; Step 2: Input the RGB image of the suspended object into the trained rotating object detection model to detect the rotating bounding box of the suspended object. The geometric center point of the rotating bounding box of the suspended object is (x, y), and its rotation angle is θ'; Step 3: Map the rotating bounding box of the suspended object obtained from the RGB image of the suspended object to the depth map, thereby determining the position of the rotating bounding box of the suspended object in the depth map; Step 4: Check whether there are missing depth values within the rotating bounding box of the suspended object in the depth map, that is, check whether the depth values of n points near the geometric center point within the rotating bounding box of the suspended object in the depth map are zero; if there are points with a depth value of 0, enter Step 5 for depth completion, and if the depth values of all points are not 0, directly enter Step 6; Step 5: Send the missing depth map into the trained depth completion model for depth completion to obtain a complete depth map; Step 6: Randomly select n points near the geometric center point within the rotating bounding box of the suspended object in the depth map, and calculate the average value of the depth values of these points as the true depth value of the suspended object, that is, the distance z between the suspended object and the depth camera; Step 7: According to the position (x, y, z) of the suspended object in the three-dimensional space and the corresponding rotation angle θ' obtained in Step 6, the precise positioning of the suspended object in the three-dimensional space is realized.
2. The three-dimensional space precise positioning method for a suspended object according to claim 1, characterized in that In Step 2, the training process in the rotating object detection model includes: Obtain the RGB image of the suspended object from a directly overhead perspective, and then annotate the RGB image of the suspended object as a training sample; Perform data augmentation on the obtained training samples; Construct a rotating object detection model; Input the training samples after data augmentation into the constructed rotating object detection model, and train based on the loss of the rotating object detection model.
3. The three-dimensional space precise positioning method of a suspended object according to claim 2, wherein The format of the rotating object annotation box of the RGB image of the suspended object is [classid, x, y, longside, shortside, θ]; classid is the category of the suspended object; x, y are the coordinates of the geometric center point of the rectangular box; longside, shortside are the width and height of the rectangular box; θ is the angle of the rectangular box. The θ angle is defined as the angle between the width of the rectangular box and the positive direction of the x-axis, negative counterclockwise and positive clockwise.
4. According to the method for precise three-dimensional positioning of a suspended object described in claim 3, characterized in that The data augmentation includes selecting four different RGB images for Mosaic augmentation, and splicing them into a new image by means of scaling, cropping and random arrangement. Further, Mixup augmentation is used to perform linear interpolation on paired images and their labels during training to generate virtual training samples. Furthermore, Copypaste augmentation is used to randomly copy and paste the input images, and increase the richness of the training data by pasting different objects of different sizes onto a new background image.
5. A three-dimensional space precise positioning method for a suspended object according to claim 3, characterized in that The construction of the rotating object detection model includes: The input end is used to receive the RGB image of the suspended object; The backbone network, which consists of a Focus module, a CBH module, a CSP module, and an SPP module, is used to extract features at different levels of the input image; The neck network, which consists of an upsampling module, a CBH module, and a CSP module, is used to fuse features at different levels; And an output layer composed of convolutional layers, which is used to output the category, center point coordinates, width, height, confidence, and rotation angle of the predicted target.
6. The three-dimensional space precise positioning method of a suspended object according to claim 2, characterized in that It also includes calculating the loss of the rotated object detection model, and the loss of the rotated object detection model includes classification loss , confidence loss , bounding box regression loss and rotation angle loss ; Among them, the classification loss is as follows: ; where is the predicted class label, is the true class label, N represents the total number of classes; Confidence loss is as follows: ; Indicates the predicted confidence of the prediction bounding box, Indicates the true confidence of the prediction bounding box. When there is a target object inside the prediction bounding box, = 1. When there is no target object inside the prediction bounding box, = 0, N Indicates the total number of prediction bounding boxes; Bounding box regression loss is as follows: ; In the formula, is and 's intersection over union, is the predicted bounding box, is the ground truth bounding box; is the Euclidean distance between the centers of the predicted bounding box and the ground truth bounding box; is the diagonal length of the minimum bounding rectangle that contains the predicted bounding box and the ground truth bounding box; is a balance parameter for balancing and the aspect ratio loss; It is a measure of the difference in aspect ratio between the predicted box and the ground truth box, expressed as: ; Among them is the width of the ground truth box, is the height of the ground truth box, is the width of the predicted box, is the height of the predicted box; Rotation angle loss is as follows: ; wherein, is the true angle, is the predicted angle, N represents the total number of angle categories; It is defined that the Gaussian function is used to convert the true angle in the lifting object label into a one-dimensional array T with a length of 180; where the value of the array T at the angle is 1, and as the angle changes from to the left and right, the value of the array T continuously decreases until it becomes 0; The conversion formula for converting the true angle value into a one-dimensional array with a length of 180 is: ; Wherein, is a window function, and the radius of the window function is , represents the true angle of the current rectangular frame, and here a Gaussian function with a standard deviation of 2 is used as the window function.
7. A three-dimensional space precise positioning method for a suspended object according to claim 1, characterized in that, The training process of the depth completion network includes: Obtaining the RGB image and the true depth image of the suspended object taken from directly above at high altitude as training samples; Constructing a depth completion model, including a self-depth completion module and an RGB-guided completion module; Input the training samples into the constructed self-depth completion module and train based on the self-depth completion loss for training; Fix the parameters of the trained self-depth completion module and train the RGB-guided completion module based on the RGB-guided completion loss , using the dynamic gradient adjustment strategy 8. The method for precise three-dimensional space positioning of a suspended object according to claim 7, characterized in that The self-depth completion module uses the original depth map to generate a preliminary depth completion image; and the self-depth completion module consists of an encoder, a decoder, and a cross-scale attention block; The RGB-guided completion module further optimizes the preliminary depth completion image under the guidance of the RGB image to generate the final depth completion image; and the RGB-guided completion module consists of a CNN, a Transformer, and a cross-modal attention module.
9. The method for precise three-dimensional space positioning of a suspended object according to claim 7, characterized in that The calculation formula is as follows: ; In the formula, is the true depth image of the suspended object, is the preliminary depth completion image output by the depth completion module, and are balance parameters; The calculation formula is as follows: ; In the formula, is the final depth completion image output by the RGB guidance completion module.
10. A method for accurately positioning a suspended object in three-dimensional space according to claim 9, characterized in that, Using the gradient dynamic adjustment algorithm to update the parameters of the RGB-guided completion module, specifically including: Calculation Feature extraction network Parameter Gradient And RGB feature extraction network Parameter Gradient : ; ; Calculate the model parameters corresponding to each modality using the ratio of the L2 norms of the model parameter gradients of the gradient difference ratio : ; ; wherein, is the L2 norm; Calculation Feature extraction network parameters Adjustment factor , RGB feature extraction network parameters Adjustment factor : ; ; In the formula, is the adjustment coefficient; Update using the gradient dynamic adjustment algorithm Feature extraction network parameters and the RGB feature extraction network parameters : ; ; Wherein, is the moment, the model parameters of the feature extraction network, is the moment, the model parameters of the RGB feature extraction network at is the learning rate, and are the first exponential decay rate and the second exponential decay rate respectively, is the power of the first exponential decay rate, is the power of the second exponential decay rate, , is the parameter of the th iteration adjustment factor.
Citation Information
Patent Citations
Single-depth camera depth map real-time enhancement method and device based on neural network
CN110211061A
Localization method and system based on deep learning
WO2020173036A1