Hanging object three-dimensional space accurate positioning method

Through the combination of the rotation object detection model and the depth completion algorithm, the problem that traditional methods are difficult to accurately locate hanging objects in three-dimensional space is solved, and the precise positioning of hanging objects in three-dimensional space is achieved, and automated operations are supported.

CN120014057AActive Publication Date: 2025-05-16SPECIAL EQUIP SAFETY SUPERVISION INSPECTION INST OF JIANGSU PROVINCE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510502775.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-16
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Traditional object detection methods are difficult to accurately locate objects in three-dimensional space, and the depth information obtained by the depth camera is often incomplete or distorted, affecting the accuracy of the positioning results.

Method used

The rotation target detection model is used to obtain the rotation bounding box information of the hanging object, and the depth information is completed in combination with the depth completion algorithm, so as to achieve the precise positioning of the hanging object in three-dimensional space.

Benefits of technology

It realizes all-round and high-precision positioning of the lifting objects in three-dimensional space, improves the integrity and accuracy of positioning, and supports the automation of crane lifting operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014057A_ABST
    Figure CN120014057A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional space accurate positioning method for a suspended object, which comprises the following steps of: 1, acquiring an RGB image and a depth map of the suspended object in a positive overlook view angle by using a depth camera, and performing registration; step 2, inputting the RGB image of the hanging object into a trained rotating target detection model to detect a rotating bounding box of the hanging object; step 3, mapping the rotation bounding box of the hanging object into the depth map, and determining the position of the rotation bounding box in the depth map; step 4, checking whether the depth value in the rotation bounding box in the depth map is missing or not, and if yes, performing depth completion; step 5, the missing depth map is sent to the trained depth completion model for depth completion, and a complete depth map is obtained; step 6, calculating the average value of the depth values of the points in the depth map as the real depth value z of the hanging object; and 7, according to the position and the corresponding rotation angle, obtained in the step 6, of the hoisted object in the three-dimensional space, precise positioning of the hoisted object in the three-dimensional space is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and object depth completion, and in particular to a method for accurately positioning a suspended object in three-dimensional space. Background Art

[0002] In industrial lifting operation scenarios, accurate three-dimensional spatial positioning of the hoisted objects is a core technical requirement to ensure operational safety and operational efficiency. However, traditional target detection methods usually focus on object recognition and positioning in a two-dimensional plane, ignoring the position and posture information of the object in three-dimensional space. For the specific scenario of hoisting objects, traditional target detection methods cannot effectively calculate the posture of the hoisted objects, that is, the rotation angle and three-dimensional spatial position in the two-dimensional plane, resulting in incomplete positioning information dimensions, which is difficult to meet the needs of automated crane lifting operations. In addition, the existing three-dimensional positioning method based on depth cameras can obtain the depth information of the scene, thereby obtaining the position of the hoisted objects in three-dimensional space. However, due to the limitations of camera hardware, changes in lighting conditions, and the influence of reflective materials on light, the depth map is often incomplete or distorted. For example, the depth information of some areas may be empty, resulting in the inability to accurately obtain the three-dimensional positioning of these areas, thereby affecting the accuracy of the overall positioning results.

[0003] Therefore, it is necessary to provide a method for accurate three-dimensional positioning of hoisted objects in order to break through the key technical barriers to the intelligent upgrading of industrial hoisting. Summary of the invention

[0004] In order to solve the above technical problems, the present invention provides a method for accurate three-dimensional positioning of suspended objects. The method uses rotating target detection to obtain information such as the angle, position and type of the suspended objects, and uses a depth completion algorithm to complete missing or incomplete depth information, thereby making the three-dimensional positioning of the suspended objects more accurate.

[0005] The technical solution of the present invention is: a method for accurately positioning a hanging object in three-dimensional space, comprising the following steps:

[0006] Step 1: Use a depth camera to obtain the RGB image and depth map of the hanging object from a top-down perspective, and perform registration, that is, to ensure that each point in the depth map can be accurately mapped to the corresponding point in the RGB image;

[0007] Step 2: Input the RGB image of the hanging object into the trained rotating object detection model to detect the rotating bounding box of the hanging object. The geometric center point of the bounding box is (x, y) and its rotation angle is θ.

[0008] Step 3: Map the object rotation bounding box obtained from the object RGB image to the depth map, thereby determining the position of the rotation bounding box in the depth map;

[0009] Step 4: Check whether there are any missing depth values ​​in the rotation bounding box in the depth map, that is, check whether the depth values ​​of n points near the geometric center point in the rotation bounding box of the depth map are zero; if there are points with a depth value of 0, go to step 5 for depth completion; if the depth values ​​of all points are not 0, go directly to step 6;

[0010] Step 5: Send the missing depth map to the trained depth completion model for depth completion to obtain a complete depth map;

[0011] Step 6: In the depth map, randomly select n points near the geometric center point in the rotation bounding box of the hanging object, and calculate the average depth value of these points as the real depth value of the hanging object, that is, the distance z between the hanging object and the depth camera;

[0012] Step 7: According to step 6, the position (x, y, z) of the hanging object in the three-dimensional space and the corresponding rotation angle θ are obtained, thereby realizing the precise positioning of the hanging object in the three-dimensional space.

[0013] Furthermore, in step 2, the training process in the rotating target detection model includes:

[0014] Obtain an RGB image of the suspended object from a top-down perspective, and then annotate the RGB image of the suspended object as a training sample;

[0015] Perform data enhancement on the acquired training samples;

[0016] Build a rotating object detection model;

[0017] The data-augmented training samples are input into the constructed rotation target detection model, and training is performed based on the loss of the rotation target detection model.

[0018] The format of the rotation target annotation box of the RGB image of the hanging object is [classid, x, y, longside, shortside, θ];

[0019] classid is the category of the hanging object;

[0020] x, y are the coordinates of the center point of the rectangular box;

[0021] longside, shortside are the long side and short side of the rectangular frame;

[0022] θ is the angle of the rectangular box. The θ angle is defined as the angle between the long side of the rectangular box and the positive direction of the x-axis. Counterclockwise is negative and clockwise is positive.

[0023] Furthermore, the data enhancement includes selecting four different RGB images using Mosaic enhancement, and splicing them into a new image by scaling, cropping and random arrangement, further using Mixup enhancement to perform linear interpolation on pairs of images and their labels during training to generate virtual training samples, and further using Copypaste enhancement to randomly copy and paste input images to increase the richness of training data by pasting different objects of different sizes onto a new background image.

[0024] Furthermore, the construction of the rotating target detection model includes:

[0025] An input end is used to receive an RGB image of a hanging object;

[0026] The backbone network consists of the Focus module, CBH module, CSP module, and SPP module, which is used to extract features at different levels of the input image;

[0027] The neck network consists of an upsampling module, a CBH module, and a CSP module, which is used to fuse features at different levels;

[0028] And the output layer composed of convolutional layers is used to output the category, center point coordinates, width, height, confidence and rotation angle of the predicted target.

[0029] Furthermore, it also includes calculating the loss of the rotation target detection model, wherein the loss of the rotation target detection model includes the classification loss. , confidence loss , bounding box regression loss and rotation angle loss ;

[0030] Among them, the classification loss for:

[0031]

[0032] In the formula To predict the class label, is the true category label, N represents the total number of categories;

[0033] Confidence loss for:

[0034]

[0035] Represents the prediction confidence of the prediction box, Indicates the true confidence of the prediction box. When there is a target object in the prediction box, =1, when there is no target object in the prediction box, =0, N represents the total number of prediction boxes;

[0036] Bounding Box Regression Loss for:

[0037]

[0038] In the formula, for and The intersection ratio of is the prediction box, is the real frame;

[0039] is the Euclidean distance between the center point of the predicted box and the true box;

[0040] is the diagonal length of the minimum enclosing rectangle containing the predicted box and the true box;

[0041] is a balance parameter used to balance and aspect ratio loss;

[0042] is a measure of the difference in aspect ratio between the predicted box and the true box, It is expressed as:

[0043]

[0044] in is the width of the real frame, is the height of the real frame, is the width of the prediction box, is the height of the prediction box;

[0045] Rotation angle loss for:

[0046]

[0047] In the formula, is the real angle, is the predicted angle, N represents the total number of angle categories;

[0048] It is defined as using a Gaussian function to convert the real angle in the hanging object label into a one-dimensional array T of length 180; The value of array T is 1, as the angle changes from Changing left and right, the value of array T keeps decreasing until it becomes 0;

[0049] The conversion formula for converting the true value of an angle into a one-dimensional array with a length of 180 degrees is:

[0050]

[0051] In the formula, is a window function, the radius of the window function Indicates the true angle of the current oblique frame. Here, a Gaussian function with a standard deviation of 2 is used as Window functions; is the true angle value, is the angle range.

[0052] Furthermore, the training process of the depth completion network includes:

[0053] Obtain the RGB image and true depth image of the suspended object taken from high altitude as training samples;

[0054] Build a depth completion model, including a self-depth completion module and an RGB-guided completion module;

[0055] The training samples are input into the constructed self-depth completion module, based on the loss Conduct training;

[0056] The trained self-depth completion module parameters are fixed, based on the loss , the RGB guided completion module is trained using a dynamic gradient adjustment strategy.

[0057] Furthermore, the self-depth completion module generates a preliminary depth completion map using the original depth map; and the self-depth completion module is composed of an encoder, a decoder, and a cross-scale attention block;

[0058] The RGB-guided completion module further optimizes the preliminary depth completion map under the guidance of the RGB map to generate a final depth completion map; and the RGB-guided completion module is composed of a CNN, a Transformer and a cross-modal attention module.

[0059] Furthermore, The calculation formula is:

[0060]

[0061] In the formula, is the real depth image of the hanging object, is the preliminary depth completion image output by the depth completion module, and Balance parameters;

[0062] The calculation formula is:

[0063]

[0064] In the formula, is the final depth image output by the RGB-guided completion module.

[0065] Furthermore, the parameters of the RGB guided completion module are updated using a gradient dynamic adjustment algorithm, including:

[0066] calculate Feature extraction network parameter Gradient and RGB feature extraction network Parameters Gradient :

[0067]

[0068]

[0069] The model parameters corresponding to each mode are calculated using the ratio of the L2 norm of the model parameter gradient The gradient difference ratio :

[0070]

[0071]

[0072] In the formula, is the L2 norm;

[0073] calculate Feature extraction network parameters The regulatory factor , RGB feature extraction network parameters The regulatory factor :

[0074]

[0075]

[0076] In the formula, is the adjustment coefficient;

[0077] Update using gradient dynamic adjustment algorithm Feature extraction network Parameters and RGB feature extraction network Parameters :

[0078]

[0079]

[0080] In the formula, For the At iteration Feature extraction network The model parameters, For the RGB feature extraction network at iteration The model parameters, is the learning rate, and are the first exponential decay rate and the second exponential decay rate, respectively. is the first exponential decay rate of Second power, is the second exponential decay rate of Second power, , For the Iteration parameters Adjustment factor.

[0081] The beneficial technical effects of the present invention are:

[0082] 1. The method of the present invention can calculate the complete three-dimensional spatial position (x, y, z) and rotation angle θ of the hanging object, realize all-round and high-precision positioning, and greatly improve the integrity of positioning.

[0083] 2. The present invention adopts a depth completion method, which can more efficiently utilize the complementary information of the two modalities of RGB image and depth image, so that the measured depth value is more accurate and the effect of depth completion is significantly improved.

[0084] 3. Completely obtain the precise position and posture information of the hoisted object in three-dimensional space, providing key data support for the automation of crane hoisting operations, making up for the defect that traditional target detection methods cannot meet the high-dimensional demand for positioning information in automated operations, and can effectively improve the safety and operational efficiency of industrial hoisting operation scenarios, and promote the development of hoisting operations towards automation and intelligence.

[0085] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] Figure 1 It is a schematic diagram of the rotating target detection method of the present invention;

[0087] Figure 2 A schematic diagram of the camera position of the present invention;

[0088] Figure 3 It is a schematic diagram of the process of the present invention;

[0089] Figure 4 It is a schematic diagram of a rotating target detection model of the present invention;

[0090] Figure 5 Schematic diagram of the deep completion network model of the present invention;

[0091] Figure 6 is a schematic diagram of a self-depth completion module of the present invention;

[0092] Figure 7 It is a schematic diagram of the RGB guide completion module of the present invention;

[0093] Figure 8 The depth map of the present invention is a depth value near the geometric center point of the rotation bounding box. DETAILED DESCRIPTION

[0094] In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the specific implementation methods of the present invention are further described in detail below in conjunction with the drawings and examples. The following examples are used to illustrate the present invention but are not used to limit the scope of the present invention.

[0095] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so as to describe the embodiments of the present application described herein.

[0096] In the description of the present invention, it should be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc. indicate directions or positional relationships based on the directions or positional relationships described in the embodiments and shown in the accompanying drawings, or are directions or positional relationships in which the inventive product is usually placed when used. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as a limitation on the present invention.

[0097] The present invention specifically relates to a method for accurately positioning a hanging object in three-dimensional space, comprising the following steps:

[0098] like Figure 1-Figure 3 As shown, step 1, use a depth camera to obtain the RGB image and depth map of the hanging object in a top-down perspective, and perform registration, that is, ensure that each point in the depth map can be accurately mapped to the corresponding point in the RGB image;

[0099] Step 2: Input the RGB image of the hanging object into the trained rotating object detection model to detect the rotating bounding box of the hanging object. The geometric center point of the bounding box is (x, y) and its rotation angle is θ.

[0100] Furthermore, in step 2, the training process in the rotating target detection model includes:

[0101] Obtain an RGB image of the suspended object from a top-down perspective, and then annotate the RGB image of the suspended object as a training sample;

[0102] Perform data enhancement on the acquired training samples;

[0103] Build a rotating object detection model;

[0104] The data-augmented training samples are input into the constructed rotation target detection model, and training is performed based on the loss of the rotation target detection model.

[0105] The format of the rotation target annotation box of the RGB image of the hanging object is [classid, x, y, longside, shortside, θ];

[0106] classid is the category of the hanging object;

[0107] x, y are the coordinates of the center point of the rectangular box;

[0108] longside, shortside are the long side and short side of the rectangular frame;

[0109] θ is the angle of the rectangular box. The θ angle is defined as the angle between the long side of the rectangular box and the positive direction of the x-axis. Counterclockwise is negative and clockwise is positive.

[0110] Furthermore, the data enhancement includes selecting four different RGB images using Mosaic enhancement, and splicing them into a new image by scaling, cropping and random arrangement, further using Mixup enhancement to perform linear interpolation on pairs of images and their labels during training to generate virtual training samples, and further using Copypaste enhancement to randomly copy and paste input images to increase the richness of training data by pasting different objects of different sizes onto a new background image.

[0111] Further, such as Figure 4 As shown, the construction of the rotating target detection model includes:

[0112] An input end is used to receive an RGB image of a hanging object;

[0113] The backbone network consists of the Focus module, CBH module, CSP module, and SPP module, which is used to extract features at different levels of the input image;

[0114] The neck network consists of an upsampling module, a CBH module, and a CSP module, which is used to fuse features at different levels;

[0115] And the output layer composed of convolutional layers is used to output the category, center point coordinates, width, height, confidence and rotation angle of the predicted target.

[0116] Furthermore, it also includes calculating the loss of the rotation target detection model, wherein the loss of the rotation target detection model includes the classification loss. , confidence loss , bounding box regression loss and rotation angle loss ;

[0117] Among them, the classification loss for:

[0118]

[0119] In the formula To predict the class label, is the true category label, N represents the total number of categories;

[0120] Confidence loss for:

[0121]

[0122] Represents the prediction confidence of the prediction box, Indicates the true confidence of the prediction box. When there is a target object in the prediction box, =1, when there is no target object in the prediction box, =0, N represents the total number of prediction boxes;

[0123] Bounding Box Regression Loss for:

[0124]

[0125] In the formula, for and The intersection ratio of is the prediction box, is the real frame;

[0126] is the Euclidean distance between the center point of the predicted box and the true box;

[0127] is the diagonal length of the minimum enclosing rectangle containing the predicted box and the true box;

[0128] is a balance parameter used to balance and aspect ratio loss;

[0129] is a measure of the difference in aspect ratio between the predicted box and the true box, It is expressed as:

[0130]

[0131] in is the width of the real frame, is the height of the real frame, is the width of the prediction box, is the height of the prediction box;

[0132] Rotation angle loss for:

[0133]

[0134] In the formula, is the real angle, is the predicted angle, N represents the total number of angle categories;

[0135] It is defined as using a Gaussian function to convert the real angle in the hanging object label into a one-dimensional array T of length 180; The value of array T is 1, as the angle changes from Changing left and right, the value of array T keeps decreasing until it becomes 0;

[0136] The conversion formula for converting the true value of an angle into a one-dimensional array with a length of 180 degrees is:

[0137]

[0138] In the formula, is a window function, the radius of the window function Indicates the true angle of the current oblique frame. Here, a Gaussian function with a standard deviation of 2 is used as Window functions; is the true angle value, is the angle range.

[0139] Step 3: Map the object rotation bounding box obtained from the object RGB image to the depth map, thereby determining the position of the rotation bounding box in the depth map;

[0140] like Figure 8As shown, step 4, check whether there is any missing depth value in the rotation bounding box in the depth map, that is, check whether the depth values ​​of n points near the geometric center point in the rotation bounding box of the depth map are zero; if there is a point with a depth value of 0, go to step 5 for depth completion, if the depth values ​​of all points are not 0, go directly to step 6;

[0141] like Figure 8 As shown, the depth values ​​of n points near the geometric center point of the rotation bounding box of the depth map are randomly selected and averaged to obtain the true depth value of the hanging object; in the figure, the red dot is the depth value of the geometric center point of the rotation bounding box of the depth map.

[0142] Step 5: Send the missing depth map to the trained depth completion model for depth completion to obtain a complete depth map;

[0143] When the depth value in the rotation bounding box of the depth map is missing, the registered depth map and RGB map are passed into the pre-trained depth completion network for depth value completion to obtain the completed depth map, and then the rotation bounding box of the hanging object is matched to the completed depth map, and the depth values ​​of n points near the geometric center point of the rotation bounding box of the depth map are randomly selected and averaged to obtain the true depth value of the hanging object;

[0144] When there is no missing depth value in the depth map rotation bounding box, the depth values ​​of n points near the geometric center point of the depth map bounding box are randomly selected and averaged to obtain the true depth value of the hanging object.

[0145] Furthermore, the training process of the depth completion network includes:

[0146] Obtain the RGB image and true depth image of the suspended object taken from high altitude as training samples;

[0147] like Figure 5 As shown, a depth completion model is constructed, including a self-depth completion module and an RGB-guided completion module;

[0148] The training samples are input into the constructed self-depth completion module, based on the loss Conduct training;

[0149] The trained self-depth completion module parameters are fixed, based on the loss , the RGB guided completion module is trained using a dynamic gradient adjustment strategy.

[0150] Further, such as Figure 6 As shown, the self-depth completion module generates a preliminary depth completion map using the original depth map; and the self-depth completion module is composed of an encoder, a decoder, and a cross-scale attention block;

[0151] like Figure 7As shown, the RGB-guided completion module further optimizes the preliminary depth completion map under the guidance of the RGB map to generate a final depth completion map; and the RGB-guided completion module is composed of a CNN, a Transformer, and a cross-modal attention module;

[0152] The cross-scale attention block solves the problem of completing large missing areas and improves the completion effect of the depth map through mutual guidance between features of different scales (such as high-resolution features guiding low-resolution features to upsample, and low-resolution features completing high-resolution missing areas);

[0153] The cross-modal attention module is used to achieve feature fusion between the depth of the hanging object and the RGB image, and fully utilize the information in the RGB image to complete the depth map.

[0154] Furthermore, The calculation formula is:

[0155]

[0156] In the formula, is the real depth image of the hanging object, is the preliminary depth completion image output by the depth completion module, and Balance parameters;

[0157] The calculation formula is:

[0158]

[0159] In the formula, is the final depth image output by the RGB-guided completion module.

[0160] Furthermore, the parameters of the RGB guided completion module are updated using a gradient dynamic adjustment algorithm, including:

[0161] calculate Feature extraction network parameter Gradient and RGB feature extraction network Parameters Gradient :

[0162]

[0163]

[0164] The model parameters corresponding to each mode are calculated using the ratio of the L2 norm of the model parameter gradient The gradient difference ratio :

[0165]

[0166]

[0167] In the formula, is the L2 norm;

[0168] calculate Feature extraction network parameters The regulatory factor , RGB feature extraction network parameters The regulatory factor :

[0169]

[0170]

[0171] In the formula, is the adjustment coefficient;

[0172] Update using gradient dynamic adjustment algorithm Feature extraction network Parameters and RGB feature extraction network Parameters :

[0173]

[0174]

[0175] In the formula, For the At iteration Feature extraction network The model parameters, For the RGB feature extraction network at iteration The model parameters, is the learning rate, and are the first exponential decay rate and the second exponential decay rate, respectively. is the first exponential decay rate of Second power, is the second exponential decay rate of Second power, , For the Iteration parameters Adjustment factor.

[0176] Use direct gradient dynamic modulation to reduce The difference in convergence speed between the depth image modality and the RGB color image modality effectively alleviates the heterogeneity between different modalities, fully utilizes the complementarity of different modalities, and obtains rich feature information.

[0177] The depth completion task is performed in two stages, each with its own specific goals and methods. Compared with other completion algorithms, this staged approach can more effectively utilize different types of features: such as the self-information of the original depth image and the rich information of the RGB image, and improve the accuracy and robustness of the completion. At the same time, the direct gradient dynamic adjustment algorithm is used to effectively reduce the heterogeneity between different modalities, make full use of the complementarity of different modalities, and obtain rich feature information.

[0178] Step 6: In the depth map, randomly select n points near the geometric center point in the rotation bounding box of the hanging object, and calculate the average depth value of these points as the real depth value of the hanging object, that is, the distance z between the hanging object and the depth camera;

[0179] Step 7: According to step 6, the position (x, y, z) of the hanging object in the three-dimensional space and the corresponding rotation angle θ are obtained, thereby realizing the precise positioning of the hanging object in the three-dimensional space.

[0180] The above embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the aforementioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the aforementioned embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention.

Claims

1. A method for accurately positioning a suspended object in three dimensions, characterized in that: The following steps are involved: Step 1: Use a depth camera to obtain the RGB image and depth map of the hanging object from a top-down perspective, and perform registration, that is, to ensure that each point in the depth map can be accurately mapped to the corresponding point in the RGB image; Step 2: Input the RGB image of the hanging object into the trained rotating object detection model to detect the rotating bounding box of the hanging object. The geometric center point of the bounding box is (x, y) and its rotation angle is θ. Step 3: Map the object rotation bounding box obtained from the object RGB image to the depth map, thereby determining the position of the rotation bounding box in the depth map; Step 4: Check whether there are any missing depth values ​​in the rotation bounding box in the depth map, that is, check whether the depth values ​​of n points near the geometric center point in the rotation bounding box of the depth map are zero; if there are points with a depth value of 0, go to step 5 for depth completion; if the depth values ​​of all points are not 0, go directly to step 6; Step 5: Send the missing depth map to the trained depth completion model for depth completion to obtain a complete depth map; Step 6: In the depth map, randomly select n points near the geometric center point in the rotation bounding box of the hanging object, and calculate the average depth value of these points as the real depth value of the hanging object, that is, the distance z between the hanging object and the depth camera; Step 7: According to step 6, the position (x, y, z) of the hanging object in the three-dimensional space and the corresponding rotation angle θ are obtained, thereby realizing the precise positioning of the hanging object in the three-dimensional space.

2. A method for accurate positioning of a suspended object in three dimensions according to claim 1, characterized in that: In step 2, the training process in the rotating target detection model includes: Obtain an RGB image of the suspended object from a top-down perspective, and then annotate the RGB image of the suspended object as a training sample; Perform data enhancement on the acquired training samples; Build a rotating object detection model; The data-augmented training samples are input into the constructed rotation target detection model, and training is performed based on the loss of the rotation target detection model.

3. A method for accurate positioning of hanging objects in three-dimensional space according to claim 2, characterized in that: The format of the rotation target annotation box of the RGB image of the hanging object is [classid, x, y, longside, shortside, θ]; classid is the category of the hanging object; x, y are the coordinates of the center point of the rectangular box; longside, shortside are the long side and short side of the rectangular frame; θ is the angle of the rectangular box. The θ angle is defined as the angle between the long side of the rectangular box and the positive direction of the x-axis. Counterclockwise is negative and clockwise is positive.

4. A method for accurate positioning of a suspended object in three dimensions according to claim 3, characterized in that: The data augmentation includes selecting four different RGB images using Mosaic augmentation, and splicing them into a new image by scaling, cropping and random arrangement, further using Mixup augmentation to perform linear interpolation on pairs of images and their labels during training to generate virtual training samples, and further using Copypaste augmentation to randomly copy and paste input images to increase the richness of training data by pasting different objects of different sizes onto a new background image.

5. The method for accurate positioning of a suspended object in three dimensions according to claim 3, characterized in that: The construction of the rotating target detection model comprises: An input end is used to receive an RGB image of a hanging object; The backbone network consists of the Focus module, CBH module, CSP module, and SPP module, which is used to extract features at different levels of the input image; The neck network consists of an upsampling module, a CBH module, and a CSP module, which is used to fuse features at different levels; And the output layer composed of convolutional layers is used to output the category, center point coordinates, width, height, confidence and rotation angle of the predicted target.

6. A method for accurate positioning of a suspended object in three dimensions according to claim 2, characterized in that: It also includes calculating the loss of the rotation target detection model, wherein the loss of the rotation target detection model includes the classification loss , confidence loss , bounding box regression loss and rotation angle loss ; Among them, the classification loss for: ; In the formula To predict the class label, is the true category label, N represents the total number of categories; Confidence loss for: ; Represents the prediction confidence of the prediction box, Indicates the true confidence of the prediction box. When there is a target object in the prediction box, =1, when there is no target object in the prediction box, =0, N represents the total number of prediction boxes; Bounding Box Regression Loss for: ; In the formula, for and The intersection ratio of is the prediction box, is the real frame; is the Euclidean distance between the center point of the predicted box and the true box; is the diagonal length of the minimum enclosing rectangle containing the predicted box and the true box; is a balance parameter used to balance and aspect ratio loss; is a measure of the difference in aspect ratio between the predicted box and the true box, It is expressed as: ; in is the width of the real frame, is the height of the real frame, is the width of the prediction box, is the height of the prediction box; Rotation angle loss for: ; In the formula, is the real angle, is the predicted angle, N represents the total number of angle categories; It is defined as using a Gaussian function to convert the real angle in the hanging object label into a one-dimensional array T of length 180; The value of array T is 1, as the angle changes from Changing left and right, the value of array T keeps decreasing until it becomes 0; The conversion formula for converting the true value of an angle into a one-dimensional array with a length of 180 degrees is: ; In the formula, is a window function, the radius of the window function Indicates the true angle of the current oblique frame. Here, a Gaussian function with a standard deviation of 2 is used as Window functions; is the true angle value, is the angle range.

7. The method for accurate positioning of a suspended object in three dimensions according to claim 1, characterized in that: The training process of the depth completion network includes: Obtain the RGB image and true depth image of the suspended object taken from high altitude as training samples; Build a depth completion model, including a self-depth completion module and an RGB-guided completion module; The training samples are input into the constructed self-depth completion module, based on the loss Conduct training; The trained self-depth completion module parameters are fixed, based on the loss , the RGB guided completion module is trained using a dynamic gradient adjustment strategy.

8. A method for accurate positioning of a suspended object in three dimensions according to claim 7, characterized in that: The self-depth completion module generates a preliminary depth completion map using the original depth map; and the self-depth completion module is composed of an encoder, a decoder, and a cross-scale attention block; The RGB-guided completion module further optimizes the preliminary depth completion map under the guidance of the RGB map to generate a final depth completion map; and the RGB-guided completion module is composed of a CNN, a Transformer and a cross-modal attention module.

9. A method for accurate positioning of a suspended object in three dimensions according to claim 7, characterized in that: The calculation formula is: ; In the formula, is the real depth image of the hanging object, is the preliminary depth completion image output by the depth completion module, and Balance parameters; The calculation formula is: ; In the formula, is the final depth image output by the RGB-guided completion module.

10. A method for accurate positioning of a suspended object in three dimensions according to claim 9, characterized in that: Use the gradient dynamic adjustment algorithm to update the parameters of the RGB guided completion module, including: calculate Feature extraction network parameter Gradient and RGB feature extraction network Parameters Gradient : ; ; The model parameters corresponding to each mode are calculated using the ratio of the L2 norm of the model parameter gradient The gradient difference ratio : ; ; In the formula, is the L2 norm; calculate Feature extraction network parameters The regulatory factor , RGB feature extraction network parameters The regulatory factor : ; ; In the formula, is the adjustment coefficient; Update using gradient dynamic adjustment algorithm Feature extraction network Parameters and RGB feature extraction network Parameters : ; ; In the formula, For the At iteration Feature extraction network The model parameters, For the RGB feature extraction network at iteration The model parameters, is the learning rate, and are the first exponential decay rate and the second exponential decay rate, respectively. is the first exponential decay rate of Second power, is the second exponential decay rate of Second power, , For the Iteration parameters Adjustment factor.

Citation Information

Patent Citations

  • Single-depth camera depth map real-time enhancement method and device based on neural network

    CN110211061A

  • Plane grabbing detection method based on computer vision and deep learning

    CN112906797A

  • Indoor dynamic environment map construction method and system based on image restoration and completion

    CN118822906A

  • Method, device and equipment for complementing holes in depth map of rotating target

    CN119444619A

  • Localization method and system based on deep learning

    WO2020173036A1