A target positioning method, device, robot, and program product
By acquiring the preset key points and orientation probability distribution values in the two-dimensional scene image and combining them with filtering, the problem of inaccurate localization of partially occluded targets by the robot was solved, achieving higher localization accuracy and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
- Filing Date
- 2025-03-26
- Publication Date
- 2026-07-24
Smart Images

Figure CN120307276B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of robotics technology, and in particular relates to a target localization method, device, robot, and program product. Background Technology
[0002] In the field of automated robotic operations, especially in complex environments such as airports and logistics centers, the detection and localization of targets is a key technology for robots to complete automated tasks. For example, in airports, robots can automatically collect scattered baggage carts and transport them to designated areas. This automated task not only reduces the need for manpower but also improves the efficiency of baggage cart collection.
[0003] However, existing target localization methods often fail to accurately identify and locate partially occluded targets, resulting in robots being unable to perform tasks effectively. Summary of the Invention
[0004] This application provides a target localization method, apparatus, robot, and program product, which can solve the problem of low localization accuracy when locating partially occluded targets using existing target localization methods.
[0005] In a first aspect, embodiments of this application provide a target localization method, the method comprising:
[0006] During the process of the robot retrieving the target device, a two-dimensional scene image containing the target device is acquired;
[0007] The target device in the two-dimensional scene image is identified to obtain the identification result related to the target device. The identification result includes the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions.
[0008] Based on the homogeneous two-dimensional coordinate values of the plurality of preset key points and the probability distribution values of the plurality of preset unit directions, the target attitude information of the target device is determined, wherein the target attitude information includes: the center point coordinate value and direction of the target device after filtering;
[0009] The robot sends the driving path information generated based on the target posture information to the robot. The driving path information includes control commands and a driving path. The robot is used to retrieve the target device along the driving path according to the control commands.
[0010] In one possible implementation of the first aspect, determining the target attitude information of the target device based on the homogeneous two-dimensional coordinate values of the plurality of preset key points and the probability distribution values of the plurality of preset unit directions includes:
[0011] The initial attitude information of the target device is obtained based on the homogeneous two-dimensional coordinate values of the plurality of preset key points and the probability distribution values of the plurality of preset unit directions; wherein, the initial attitude information includes: the initial center point coordinate values and the initial direction of the target device;
[0012] The initial center point coordinates and initial direction of the target device are filtered to obtain the target attitude information of the target device.
[0013] In one possible implementation of the first aspect, obtaining the initial attitude information of the target device based on the homogeneous two-dimensional coordinate values of the plurality of preset key points and the probability distribution values of the plurality of preset unit directions includes:
[0014] The three-dimensional coordinates of the multiple preset key points are determined based on the homogeneous two-dimensional coordinates of the multiple preset key points and the camera intrinsic parameter matrix.
[0015] The initial center point coordinates of the target device are determined based on the three-dimensional coordinates of the plurality of preset key points.
[0016] The preset unit direction with the highest probability distribution value among the plurality of preset unit directions is determined as the initial direction of the target device.
[0017] In one possible implementation of the first aspect, filtering the initial center point coordinates and initial direction of the target device to obtain the target attitude information of the target device includes:
[0018] Based on the initial center point coordinates and the initial direction in the initial attitude information, the updated moving average value is calculated using the improved moving average filter MMAF.
[0019] Based on the updated moving average value, outliers in the initial center point coordinates and initial direction are removed to obtain the target attitude information of the target device.
[0020] In one possible implementation of the first aspect, the step of identifying the target device in the two-dimensional scene image to obtain an identification result related to the target device includes:
[0021] The target device in the two-dimensional scene image is detected by the target detection module, and multiple target bounding boxes containing the target device are generated.
[0022] Based on multiple preset key points, the homogeneous two-dimensional coordinate values of the multiple preset key points of the target device are extracted from multiple target bounding boxes by the high-resolution network HRNet in the target detection module.
[0023] Based on multiple preset unit directions, the probability distribution values of the multiple preset unit directions of the target device are extracted from multiple target bounding boxes by the high-resolution network HRNet and the residual neural network ResNet in the target detection module.
[0024] In one possible implementation of the first aspect, the step of extracting the homogeneous two-dimensional coordinate values of the multiple preset key points of the target device from multiple target bounding boxes using a high-resolution network HRNet in the target detection module based on multiple preset key points includes:
[0025] Based on multiple preset key points, the heat map coordinate values of multiple preset key points are extracted from multiple target bounding boxes through the high-resolution network HRNet;
[0026] The heat map coordinate values of the multiple preset key points are transformed to obtain the homogeneous two-dimensional coordinate values of the multiple preset key points.
[0027] In one possible implementation of the first aspect, before detecting the target device in the two-dimensional scene image by the target detection module and generating multiple target bounding boxes containing the target device, the method further includes:
[0028] Obtain a training dataset, wherein the training dataset includes multiple sample images and target annotation results for the multiple sample images;
[0029] An initial target detection model is trained using the multiple sample images and the target annotation results of the multiple sample images to obtain the target detection model.
[0030] Secondly, embodiments of this application provide a target positioning device, the device comprising:
[0031] The acquisition module is used to acquire a two-dimensional scene image containing the target device during the process of the robot retrieving the target device;
[0032] The recognition module is used to recognize the target device in the two-dimensional scene image and obtain recognition results related to the target device. The recognition results include homogeneous two-dimensional coordinate values of multiple preset key points and probability distribution values of multiple preset unit directions.
[0033] The determination module is used to determine the target attitude information of the target device based on the homogeneous two-dimensional coordinate values of the plurality of preset key points and the probability distribution values of the plurality of preset unit directions, wherein the target attitude information includes: the center point coordinate values and directions of the target device after filtering;
[0034] The sending module is used to send the driving path information generated based on the target posture information to the robot, wherein the driving path information includes control commands and a driving path, and the robot is used to retrieve the target device along the driving path according to the control commands.
[0035] Thirdly, embodiments of this application provide a robot, the robot including a controller, the controller being used to execute the target localization method described in any of the preceding claims.
[0036] Fourthly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the target positioning method described in any of the above claims.
[0037] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the target localization method described in any of the preceding claims.
[0038] Sixthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the target positioning method described in any one of the first aspects.
[0039] The beneficial effects of the embodiments in this application compared with the prior art are:
[0040] This application provides a target localization method, comprising: acquiring a two-dimensional scene image containing the target device during the process of a robot retrieving a target device; identifying the target device in the two-dimensional scene image to obtain identification results related to the target device, wherein the identification results include homogeneous two-dimensional coordinate values of multiple preset key points and probability distribution values of multiple preset unit directions; determining the target posture information of the target device based on the homogeneous two-dimensional coordinate values of the multiple preset key points and the probability distribution values of the multiple preset unit directions, wherein the target posture information includes: the center point coordinate value and direction of the target device after filtering; and sending the driving path information generated based on the target posture information to the robot, wherein the driving path information includes control commands and a driving path, and the robot is used to retrieve the target device along the driving path according to the control commands. Using the above technical solution, by identifying the center point coordinate value of the target device based on the homogeneous two-dimensional coordinate values of multiple preset key points in the identification results, and identifying the direction of the target device based on the probability distribution values of multiple preset unit directions in the identification results, the position and direction of the target device are identified separately, so that even if the target device is partially occluded, the position and direction of the target device can still be located, improving the accuracy of the robot's target localization. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a flowchart illustrating a target localization method provided in an embodiment of this application;
[0043] Figure 2 This is a schematic diagram illustrating how to determine the position of a luggage cart based on preset key points, according to an embodiment of this application.
[0044] Figure 3 This is a schematic diagram illustrating the determination of the direction of a luggage cart according to an embodiment of this application;
[0045] Figure 4 This is a schematic diagram illustrating the process of a robot locating and retrieving a luggage cart, provided in one embodiment of this application.
[0046] Figure 5 This is a schematic diagram of the structure of a target positioning device provided in an embodiment of this application;
[0047] Figure 6This is a schematic diagram of the structure of a robot provided in one embodiment of this application;
[0048] Figure 7 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation
[0049] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0050] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0051] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0052] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0053] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0054] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0055] In the field of robotic automation, especially in complex environments such as airports and logistics centers, target detection and localization are key technologies for robots to complete automated tasks. For example, in airports, robots can automatically collect scattered baggage carts and transport them to designated areas. This automation not only reduces manpower requirements but also improves the efficiency of baggage cart collection. However, existing target localization methods often fail to accurately identify and locate partially occluded targets, resulting in robots being unable to perform tasks effectively.
[0056] Existing target localization methods include those based on visual models, such as Convolutional Neural Networks (CNNs) based on deep learning. These visual models perform well when the target is fully visible and unobstructed, but their performance degrades when the target is partially occluded. In addition, there are target localization methods that rely on multiple sensors (such as wireless signals, visual sensors, and LiDAR). While these methods can provide relatively accurate localization information under ideal conditions, in practical applications, especially when the target is partially occluded or the environment is complex and variable, they often fail to accurately determine the target's pose due to the inability to observe a sufficient number of key points, posing significant challenges to the accuracy and real-time performance of their localization.
[0057] To enhance the autonomous operation capabilities of robots in complex environments, especially when targets are partially occluded, a novel perception strategy is needed to accurately and efficiently locate targets using only limited keypoint information. This strategy must not only possess strong anti-occlusion capabilities but also be able to respond rapidly in real-time environments to support the robot's dynamic planning and precise operation. Based on this, embodiments of this application provide a target localization method applicable to robot target detection and localization in complex environments. By performing layered processing of the target device's position and orientation, the accuracy of target device localization is improved even when partially occluded. Furthermore, a progressive perception strategy enhances the robot's ability to continuously update the target device's posture in dynamic environments, improving the robustness of localization. Finally, by simplifying data requirements, the implementation cost of this method is reduced, improving its scalability and feasibility for practical application.
[0058] Please see Figure 1 , Figure 1 This is a schematic flowchart illustrating a target localization method according to an embodiment of this application. As an example and not a limitation, this method can be applied to terminal devices, such as robots. This method can be applied to the positioning and retrieval system of robots autonomously collecting baggage carts in airports, and also to the transportation system of robots autonomously transporting objects in logistics warehouses, hospitals, etc., where autonomous collection or transportation of objects is required. Figure 1 As shown, the method includes:
[0059] S11. During the process of the robot retrieving the target device, acquire a two-dimensional scene image containing the target device.
[0060] In this embodiment, the executing entity can be a terminal device such as a server, or a target positioning system in a robot. In some examples, the robot can be a terminal device used to collect or transport objects. The robot needs to identify and locate a target device, and after locating the target device, it travels to the target device to collect or transport it. The target device is the equipment or object that the robot needs to retrieve. For example, in an airport, the robot can be a device with a target positioning system capable of autonomously retrieving baggage carts, and the target device is the baggage cart that needs to be retrieved.
[0061] The two-dimensional scene image can be a two-dimensional image of the actual scene at the location of the target device. This image may include the target device, obstructions to the target device, and other objects. The two-dimensional scene image can be acquired by the robot through a camera or other image acquisition device; this embodiment does not limit this.
[0062] S12. Identify the target device in the two-dimensional scene image and obtain the identification result related to the target device.
[0063] The recognition results include homogeneous two-dimensional coordinates of multiple preset key points and probability distribution values of multiple preset unit directions.
[0064] After obtaining a two-dimensional scene image of the target device's location, it is necessary to identify the target device in the two-dimensional scene image. In this embodiment, a target detection model (such as the YOLOv5 detection model) can be used to identify the target device in the two-dimensional scene image, thereby obtaining identification results related to the target device. The identification results may include homogeneous two-dimensional coordinate values of multiple preset key points, probability distribution values of multiple preset unit directions, and other content related to the target device, such as the target device's usage status (occupied, idle) and usage environment (indoor, outdoor).
[0065] Preset key points are prominent location points pre-defined on the target device to identify its location. These key points allow for the determination of the target device's position. Multiple preset key points can be used, and their number depends on the specific application; this embodiment does not impose a specific limitation. For example, as shown... Figure 2 As shown, Figure 2 This application provides an embodiment of a schematic diagram illustrating the determination of a luggage cart position based on preset key points. During the robot's positioning and retrieval of the luggage cart, there can be six preset key points for identifying the luggage cart, specifically the positions where the four wheels of the luggage cart intersect with the cart body (e.g., ...). Figure 2 a, b, c, d) in the text, and the positions of both ends of the luggage cart handle (such as...). Figure 2 (e, f in the text)
[0066] In a two-dimensional plane, a pair of two-dimensional coordinates (x, y) is usually used to represent the exact position of a point in the two-dimensional plane, or a vector (x, y) is used to mark the exact position of a point in the two-dimensional plane. If an additional variable w is added, the two-dimensional coordinates (x, y) of the point in the homogeneous coordinate system become (x, y, w), and the coordinates (x, y, w) are the homogeneous two-dimensional coordinates corresponding to the point (x, y).
[0067] In this embodiment, the 360 degrees are divided into equal parts according to a preset division rule, and each unit obtained from this division is a preset unit direction. For example, dividing 360 degrees into 36 equal parts represents 36 preset unit directions, each with an angle of 10 degrees. That is, starting from 0 degrees, each direction is increased by 10 degrees until 360 degrees are reached. The probability distribution value of the preset unit direction can be the probability distribution value for a certain direction of the target device. This probability distribution value can be calculated based on features, colors, and other information in the two-dimensional scene image, and can reflect information such as the orientation or movement trend of the target device in the two-dimensional scene image. The initial direction of the target device can be preliminarily estimated through the probability distribution value of the preset unit direction.
[0068] In one possible implementation, before detecting and recognizing target devices in a 2D scene image using the target detection module, an initial target detection model needs to be built and trained to obtain a target detection module capable of recognizing target devices from 2D images. The training process specifically involves:
[0069] Obtain the training dataset, which includes multiple sample images and the target annotation results of the multiple sample images.
[0070] The initial object detection model is trained using multiple sample images and the object annotation results of multiple sample images to obtain the object detection model.
[0071] The initial object detection model is a newly constructed, untrained deep learning model. The training dataset is a subset of data used to train the initial object detection model. This dataset contains known input features and desired output labels, i.e., sample images and their object annotations. Sample images are collected 2D scene images containing target devices; the object annotations are the model's output recognition results. Continuing the example above, in the process of a robot locating and retrieving luggage carts, the object detection model can be trained first. The training dataset can be a dataset containing tens of thousands of 2D scene images. The sample images in this dataset can include different luggage cart states (occupied, idle), various occlusion levels (e.g., visibility of 40%, visibility of 80%, etc.), and different scene environments (indoor, outdoor). The object annotations in the sample images can include target bounding boxes, multiple preset key points, orientation (luggage cart's facing direction), usage status, etc. Through these object annotations, recognition results related to idle luggage carts can be obtained.
[0072] It should be understood that by training the initial target detection model with a large sample dataset, the target detection model can accurately identify the recognition results related to the target device from the new two-dimensional scene image, namely the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions, so as to provide accurate recognition results for subsequent processing of the position and orientation of the target device.
[0073] It should be understood that by annotating and training only on two-dimensional images, the complexity of data collection and annotation is reduced, and the scalability of the system and the feasibility of practical applications are improved.
[0074] In some embodiments, a target device in a two-dimensional scene image is identified to obtain an identification result related to the target device, including:
[0075] The target detection module detects target devices in a 2D scene image and generates multiple target bounding boxes containing the target devices.
[0076] Based on multiple preset key points, the homogeneous two-dimensional coordinate values of multiple preset key points of the target device are extracted from multiple target bounding boxes by the high-resolution network HRNet in the target detection module.
[0077] Based on multiple preset unit orientations, the probability distribution values of multiple preset unit orientations of the target device are extracted from multiple target bounding boxes through the high-resolution network HRNet and the residual neural network ResNet in the target detection module.
[0078] In this embodiment, the target bounding box can be understood as a rectangular box drawn around the target device in a two-dimensional scene image, representing the approximate location of the target device in the image. A High-Resolution Network (HRNet) is a deep learning network architecture used to extract high-resolution features from an image. In this embodiment, HRNet is used to extract the coordinates of preset key points from the target bounding box. A Residual Neural Network (ResNet) is also a deep learning network architecture that simplifies the training process of deep networks by introducing residual connections, thereby improving model performance. In this embodiment, ResNet is combined with HRNet to extract features of the target device in a preset unit direction within the target bounding box, thereby calculating the probability distribution value in the preset unit direction.
[0079] In a specific embodiment, firstly, the target detection module processes the two-dimensional scene image to identify multiple target bounding boxes containing the target device, thereby determining the approximate location of the target device in the two-dimensional scene image. After determining the approximate location of the target device in the two-dimensional scene image, the high-resolution network HRNet is used to extract features from the image regions within the target bounding boxes. Based on multiple preset key points, the positions of these preset key points are extracted in the two-dimensional scene image, obtaining the homogeneous two-dimensional coordinate values of these preset key points. These homogeneous two-dimensional coordinate values can be conveniently extended to three-dimensional space. Furthermore, the high-resolution network HRNet and the residual neural network ResNet are used to extract features and classify the image regions within the target bounding boxes, obtaining the probability distribution value of the target device in each preset unit direction. This probability distribution value represents the probability that the target device's direction is that preset unit direction.
[0080] In some embodiments, based on multiple preset key points, the homogeneous two-dimensional coordinate values of multiple preset key points of the target device are extracted from multiple target bounding boxes by a high-resolution network HRNet in the target detection module, including:
[0081] Based on multiple preset key points, the heat map coordinates of multiple preset key points are extracted from multiple target bounding boxes using the high-resolution network HRNet.
[0082] The heatmap coordinate values of multiple preset key points are transformed to obtain the homogeneous two-dimensional coordinate values of the multiple preset key points.
[0083] It should be noted that in the process of extracting homogeneous two-dimensional coordinate values of multiple preset key points of the target device from multiple target bounding boxes using the high-resolution network HRNet, the heatmap coordinate values of multiple preset key points are first extracted from the target bounding boxes based on these preset key points. Here, the heatmap coordinate values refer to the position coordinates of the preset key points on the heatmap. The heatmap can represent the position information of the preset key points. Then, the heatmap coordinate values are converted into homogeneous two-dimensional coordinate values of the multiple preset key points. In this embodiment, the conversion method for converting the heatmap coordinate values of the multiple preset key points into homogeneous two-dimensional coordinate values can be to map the heatmap coordinate values to the pixel coordinate system of the image and add homogeneous coordinates to obtain a third-dimensional vector; other conversion methods can also be used, and this embodiment does not specifically limit this method. It should be understood that converting the heatmap coordinate values into homogeneous two-dimensional coordinate values can conveniently extend the two-dimensional coordinate points into three-dimensional space, providing assistance for subsequent position processing of the target device.
[0084] Continuing the example above, during the robot's location and retrieval of the luggage cart, the robot uses a target detection model to perform real-time detection on the 2D scene image of the luggage cart, generating multiple target bounding boxes. Then, it extracts the homogeneous 2D coordinates of six preset key points of the luggage cart and the probability distribution values of the luggage cart in a preset unit direction from these target bounding boxes. Specifically: First, a high-resolution network HRNet is used to predict the heatmap coordinates p of the six preset key points of the luggage cart. i =[x i ,y i ] T Generate homogeneous two-dimensional coordinate values corresponding to preset key points. Subsequently, the high-resolution network HRNet and the residual neural network ResNet are combined to detect unit orientations, that is, the probability distribution values of n preset unit orientations are output through fully connected layers and softmax layers.
[0085] S13. Determine the target attitude information of the target device based on the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions.
[0086] The target attitude information includes: the center point coordinates and orientation of the target device after filtering.
[0087] Target pose information can refer to the center point coordinates and orientation (i.e., the target device's facing direction) of the target device after filtering, which is the final position and orientation of the target device. The center point coordinates of the target device can refer to the position of a fixed point on the target device within a 2D scene image. This point is usually determined based on the shape, size, or preset rules of the target device. For example, if the target device has a regular shape (such as a circle or square), the center point coordinates can be the coordinates of its geometric center; if the target device has an irregular shape or uneven mass distribution (such as a luggage cart), the center point coordinates can be its centroid; of course, the center point coordinates can also be a specific point preset according to the actual situation.
[0088] The center point coordinates and direction of the target device after filtering can be understood as the center point coordinates and direction obtained after filtering to eliminate outliers. Outliers can be points in the data that are significantly different from the majority of data points, and these outliers may be caused by detection errors, noise, or other reasons. Compared to the center point coordinates and direction without filtering, the center point coordinates and direction after filtering can more accurately describe the position and direction of the target device. Filtering is a method used to smooth data, reduce noise, and improve data accuracy. In this embodiment, the center point coordinates and direction are filtered to eliminate outliers caused by detection errors or data fluctuations.
[0089] S14. Send the driving path information generated based on the target posture information to the robot.
[0090] The travel path information includes control commands and the travel path, which the robot uses to retrieve the target equipment along the travel path according to the control commands.
[0091] Specifically, after obtaining the target device's posture information, a driving path information can be generated using a Multi-Risk-RRT algorithm. This driving path information can include control commands for the robot's movement and the path the robot will take to retrieve the target device. After generating the driving path information, the control commands are sent to the robot. The robot will then navigate along the generated path to the target device, avoiding pedestrians and obstacles, thereby accurately reaching the target device's location and completing the retrieval.
[0092] It should be noted that during the robot's movement towards the target device, the robot's target localization system continuously monitors the target device and updates its orientation information to ensure the robot can accurately reach and grasp it. Once the robot has grasped the target device, its localization efforts cease. This continuous updating of the target device's positioning information as the robot approaches gradually improves positioning accuracy, thereby reducing positioning errors caused by long-distance detection.
[0093] In this embodiment, the Multi-Risk-RRT algorithm is a path planning algorithm designed for multi-risk environments. It considers multiple potential risk factors and can plan safer and more reliable paths. The Multi-Risk-RRT algorithm can combine multi-directional search and heuristic sampling to efficiently integrate heuristic information from dynamic subtrees into the root tree, overcoming the limitations of Two-Point Boundary Value Problem (TBVP) solvers, thereby improving motion planning performance in both static and dynamic environments. In this embodiment, the specific process of generating driving path information using the Multi-Risk-RRT algorithm is not described in detail.
[0094] It is understood that the target localization method provided in this embodiment includes: acquiring a two-dimensional scene image containing the target device during the process of a robot retrieving a target device; identifying the target device in the two-dimensional scene image to obtain an identification result related to the target device, wherein the identification result includes homogeneous two-dimensional coordinate values of multiple preset key points and probability distribution values of multiple preset unit directions; determining the target posture information of the target device based on the homogeneous two-dimensional coordinate values of the multiple preset key points and the probability distribution values of the multiple preset unit directions, wherein the target posture information includes: the center point coordinate value and direction of the target device after filtering; and sending the driving path information generated based on the target posture information to the robot, wherein the driving path information includes control commands and a driving path, and the robot is used to retrieve the target device along the driving path according to the control commands. Using the above technical solution, the center point coordinates of the target device are obtained by identifying the homogeneous two-dimensional coordinates of multiple preset key points in the recognition results, and the orientation of the target device is obtained by identifying the probability distribution values of multiple preset unit directions in the recognition results. The position and orientation of the target device are identified and processed separately, so that even if the target device is partially occluded, the position and orientation of the target device can still be located, thus improving the accuracy of the robot's target positioning.
[0095] In some embodiments, the target attitude information of the target device is determined based on the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions, including:
[0096] The initial attitude information of the target device is obtained based on the homogeneous two-dimensional coordinates of multiple preset key points and the probability distribution values of multiple preset unit directions; wherein, the initial attitude information includes: the initial center point coordinates and initial direction of the target device.
[0097] The initial center point coordinates and initial direction of the target device are filtered to obtain the target attitude information of the target device.
[0098] Compared to target attitude information, initial attitude information can be understood as the position and orientation of the target device without filtering, i.e., the initial center point coordinates and initial orientation. In a specific embodiment, firstly, the initial center point coordinates of the target device are calculated based on the homogeneous two-dimensional coordinates of multiple preset key points. In this embodiment, the average or weighted average of the homogeneous two-dimensional coordinates of these preset key points can be used to calculate the initial centerline coordinates. Secondly, the initial orientation of the target device is calculated based on the probability distribution values of multiple preset unit directions. In this embodiment, the preset unit direction corresponding to the highest probability distribution value can be selected as the initial orientation of the target device, or the preset unit direction corresponding to the weighted average of the probability distribution values can be used as the initial orientation of the target device. Then, the initial center point coordinates and initial orientation are filtered using a modified moving average filter (MMAF) to eliminate outliers caused by detection errors or data fluctuations, thus obtaining the final position and orientation of the target device, i.e., the center point coordinates and orientation.
[0099] In some embodiments, the initial attitude information of the target device is obtained based on the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions, including:
[0100] The three-dimensional coordinates of multiple preset key points are determined based on the homogeneous two-dimensional coordinates of multiple preset key points and the camera intrinsic parameter matrix.
[0101] The initial center point coordinates of the target device are determined based on the three-dimensional coordinates of multiple preset key points.
[0102] The initial direction of the target device is determined by the preset unit direction with the highest probability distribution value among multiple preset unit directions.
[0103] The camera intrinsic matrix describes the internal characteristics of image acquisition devices such as cameras. It can include parameters such as focal length and coordinates. The camera intrinsic matrix enables the conversion of 2D coordinate values to camera coordinates, providing a foundation for obtaining subsequent 3D coordinate values. The 3D coordinates of multiple preset keypoints can be represented as their coordinates in the 3D camera coordinate system. These coordinates can be calculated using the camera intrinsic matrix and the 2D coordinates of the preset keypoints, reflecting their positional information in 3D space.
[0104] In a specific embodiment, after obtaining the recognition result by recognizing the two-dimensional scene image through the target detection model, the homogeneous two-dimensional coordinate values of multiple preset key points in the recognition result can first be input into the camera projection model in the robot. Combined with the projection transformation of the camera intrinsic parameter matrix, the homogeneous two-dimensional coordinate values of the multiple preset key points in the recognition result are converted into three-dimensional coordinate values of the multiple preset key points, thereby obtaining the position information of the multiple preset key points in three-dimensional space. Then, based on the position information of the multiple preset key points in three-dimensional space, the initial center point coordinate values of the target device are determined, i.e., the position of the center point of the target device is obtained. In this embodiment, this can be achieved through the prior model M = {Y}. i =(x i ,y i ,ζ i )∈R 3 |i=0,1,...,n} and ground plane constraints To calculate the initial centerline coordinates of the target device, where Y represents the initial center point coordinates of the target device. i This represents the coordinates of the preset key points, x. i y i These represent the horizontal position values of the preset key points relative to the initial center point of the target device, ζ and ζ, respectively. i The value represents the height of the preset key point relative to the initial center point of the target device, i represents multiple preset key points, λ represents the vertical distance from the camera to the ground in the camera projection model, λ is fixed, and N represents the actual height of the camera from the ground.
[0105] For example, such as Figure 2 As shown, Figure 2 This is a schematic diagram illustrating how the position of a luggage cart is determined based on preset key points, according to an embodiment of this application. Figure 2 The example illustrates the process of a robot locating and retrieving a luggage cart. It is assumed that the bottom of the luggage cart and the bottom of the robot are on the same horizontal plane, and the vertical height λ from the center point of the camera on the robot to each preset key point is known. First, the homogeneous two-dimensional coordinates of multiple preset key points (such as a, b, c, d, e, f) on the two-dimensional scene image plane of the luggage cart are obtained through an object detection model. Figure 2 In the diagram, (u, v) represents the normal vector, and N represents the actual height of the camera above the ground. Then, the inverse matrix K of the camera intrinsic parameter matrix K from the camera projection model can be used. -1 Calculate the ray direction ρ of these preset key points; then, combine the camera height value λ and the height values ζ of multiple preset key points in the camera projection model. i The actual X positions of multiple preset key points in three-dimensional space are calculated. iFinally, by calculating the actual positions of multiple preset key points in three-dimensional space and their corresponding positions Y in the aforementioned prior model M, the system... i The average deviation is used to determine the coordinates of the center point of the luggage cart, that is, the actual position (X) of each preset key point can be determined. i ,vis) and the position (Y) in the prior model M i Subtract vis from , and then take the average to obtain the position (x, y) of the center point of the luggage cart.
[0106] After determining the center point of the target device, the orientation of the target device is then determined. This is done by selecting the unit orientation with the highest probability distribution value from multiple preset unit orientations, and then setting this unit orientation as the initial orientation of the target device. Specifically, the unit orientation with the highest probability value can be selected from the multiple preset unit orientation probability distribution values provided by the target detection model. For example, a 360-degree discrete representation can be used for precise orientation estimation. A loss function is then defined. This is used to calculate the deviation between the predicted direction and the true direction, where... Let j represent the probability distribution values of multiple preset unit directions. This represents a circular Gaussian probability distribution. Finally, the initial direction θ of the target device can be determined by selecting the unit direction corresponding to the highest probability distribution value, i.e. like Figure 3 As shown, Figure 3 This is a schematic diagram illustrating the determination of the direction of a luggage cart according to an embodiment of this application. Figure 3 The symbol indicates that the initial orientation of the luggage cart is 35 degrees.
[0107] Thus, through this embodiment, the initial position (initial centerline coordinates) and initial orientation of the target device can be determined, thereby realizing the positioning of the initial attitude of the target device.
[0108] It should be understood that in this embodiment, by using independent modules to process the position and orientation of the target device separately, the position of the target device can be located and its orientation estimated even when some key points are detected (the target device is partially occluded), which improves the accuracy of the robot's target location and enhances the robot's operational capability and robustness in dynamic environments.
[0109] In some embodiments, the initial center point coordinates and initial direction of the target device are filtered to obtain the target attitude information of the target device, including:
[0110] Based on the initial center point coordinates and initial direction from the initial attitude information, the updated moving average value is calculated using the improved moving average filter MMAF.
[0111] Based on the updated moving average, outliers in the initial center point coordinates and initial direction are removed to obtain the target attitude information of the target device.
[0112] In a specific embodiment, after obtaining the initial target attitude information (initial center point coordinates and initial direction) of the target device, the initial target attitude information can be filtered by an improved moving average filter (MMAF) to eliminate outliers caused by detection errors or data fluctuations, thereby obtaining the target attitude information of the target device. Specifically, the initial attitude information (initial center point coordinates and initial direction) is input into the improved moving average filter (MMAF). The MMAF filter calculates a series of moving average values of data points based on its internal algorithm (such as adaptive window size, weight allocation, etc.). These moving average values can be used as the updated center point coordinates and direction. For example, the MMAF can be... The parameters of the AF filter are initialized as F = (Δ, Θ_z), where Δ represents the window size of the moving average and Θ_z represents the threshold for outlier detection using the standard score z-score. After obtaining these updated moving averages, they can be compared with the initial data points (initial center point coordinates and initial direction). If the difference between a data point and the moving average exceeds a preset threshold (e.g., the preset threshold is Θ_z, which can be set according to actual conditions), the data point can be removed as a discrete point to obtain the target attitude information of the target device. The above filtering process can reduce noise and errors in the recognition process and improve the accuracy and reliability of the target attitude information.
[0113] Following the example above, such as Figure 4 As shown, Figure 4 This is a schematic diagram illustrating the process of a robot locating and retrieving a luggage cart, provided in one embodiment of this application. Figure 4 In the process, the robot first acquires a 2D scene image containing luggage carts. This 2D scene image is then input into the luggage cart detection module of the target detection model. The luggage cart detection module identifies the luggage carts in the image and outputs the target bounding box. Subsequently, the key point and orientation detection module extracts the recognition results from the target bounding box, namely the homogeneous 2D coordinate values of multiple preset key points. Probability distribution values of multiple preset unit directions Then, the key point processing module processes the homogeneous two-dimensional coordinates of multiple preset key points to output the initial center point coordinates (x, y) of the luggage cart. The direction processing module processes the probability distribution values of multiple preset unit directions to output the initial direction (θ) of the luggage cart. Finally, a filter is applied to the initial center point coordinates (x, y) and the initial direction (θ) of the luggage cart to output the filtered target attitude information (x, y) of the luggage cart. f ,y f ,θ f Finally, the target attitude information (x) of the luggage cart is... f ,y f ,θ f The input is given to the trajectory planning module, which then uses the target attitude information (x) of the luggage cart to plan the trajectory. f ,y f ,θ f The robot outputs the driving path information (v,ω), and moves towards the luggage cart according to the driving path information (v,ω) to complete the retrieval of the luggage cart.
[0114] It is understood that the target localization method provided in this application embodiment can have the following beneficial effects:
[0115] (1) By processing the position and orientation of the target device separately, the accuracy of positioning is improved when the target device is partially occluded, and the problem of positioning failure under occlusion conditions in existing target positioning methods is solved.
[0116] (2) By identifying and updating the target device's posture information in real time during the robot's approach to the target device, the robot's ability to continuously update the target posture in a dynamic environment is enhanced, thereby improving the accuracy and robustness of target positioning.
[0117] (3) By detecting the two-dimensional image of the target device, the requirements for input data are simplified, the implementation cost of the target recognition method is reduced, the scalability and feasibility of the method are improved, and the method can be more widely applied to robot target localization tasks in various complex environments.
[0118] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0119] Corresponding to the target localization method in the above embodiments, Figure 5 A schematic diagram of a target positioning device according to an embodiment of this application is shown. For ease of explanation, only the parts related to the embodiment of this application are shown.
[0120] Reference Figure 5 The target positioning device 3 in this embodiment includes:
[0121] The acquisition module 31 is used to acquire a two-dimensional scene image containing the target device during the process of the robot retrieving the target device.
[0122] The recognition module 32 is used to recognize the target device in the two-dimensional scene image and obtain the recognition result related to the target device. The recognition result includes the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions.
[0123] The determination module 33 is used to determine the target attitude information of the target device based on the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions. The target attitude information includes the center point coordinate values and directions of the target device after filtering.
[0124] The sending module 34 is used to send the driving path information generated based on the target posture information to the robot. The driving path information includes control commands and driving path. The robot is used to retrieve the target device along the driving path according to the control commands.
[0125] It is understood that in this embodiment, the present application provides a target positioning device 3. During the robot's retrieval of a target device, the acquisition module 31 acquires a two-dimensional scene image containing the target device; the recognition module 32 identifies the target device in the two-dimensional scene image to obtain recognition results related to the target device, wherein the recognition results include homogeneous two-dimensional coordinate values of multiple preset key points and probability distribution values of multiple preset unit directions; the determination module 33 determines the target posture information of the target device based on the homogeneous two-dimensional coordinate values of the multiple preset key points and the probability distribution values of the multiple preset unit directions, wherein the target posture information includes: the center point coordinates and direction of the target device after filtering; and the sending module 34 sends the driving path information generated based on the target posture information to the robot, wherein the driving path information includes control commands and a driving path, and the robot uses this driving path to retrieve the target device according to the control commands. Using the target positioning device 3, the center point coordinates of the target device are obtained by identifying the homogeneous two-dimensional coordinates of multiple preset key points in the recognition results, and the orientation of the target device is obtained by identifying the probability distribution values of multiple preset unit directions in the recognition results. The position and orientation of the target device are identified and processed separately, so that even if the target device is partially occluded, the position and orientation of the target device can still be located, thus improving the accuracy of the robot's target positioning.
[0126] Optionally, the determining module 33 includes:
[0127] The determination submodule is used to obtain the initial attitude information of the target device based on the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions; wherein, the initial attitude information includes: the initial center point coordinate value and the initial direction of the target device.
[0128] The filtering submodule is used to filter the initial center point coordinates and initial direction of the target device to obtain the target attitude information of the target device.
[0129] Optionally, the submodules to be determined include:
[0130] The first determining unit is used to determine the three-dimensional coordinates of multiple preset key points based on the homogeneous two-dimensional coordinates of multiple preset key points and the camera intrinsic parameter matrix.
[0131] The second determining unit is used to determine the initial center point coordinates of the target device based on the three-dimensional coordinates of multiple preset key points.
[0132] The third determining unit is used to determine the preset unit direction with the highest probability distribution value among multiple preset unit directions as the initial direction of the target device.
[0133] Optionally, the filtering submodule includes:
[0134] The update unit is used to calculate the updated moving average value based on the initial center point coordinates and initial direction in the initial attitude information using the improved moving average filter MMAF.
[0135] The removal unit is used to remove outliers from the initial center point coordinates and initial direction based on the updated moving average, thereby obtaining the target attitude information of the target device.
[0136] Optionally, the recognition module 32 includes:
[0137] The detection submodule is used to detect target devices in a 2D scene image through the target detection module and generate multiple target bounding boxes containing the target devices.
[0138] The first extraction submodule is used to extract the homogeneous two-dimensional coordinate values of multiple preset key points of the target device from multiple target bounding boxes using the high-resolution network HRNet in the target detection module, based on multiple preset key points.
[0139] The second extraction submodule is used to extract the probability distribution values of multiple preset unit directions of the target device from multiple target bounding boxes using the high-resolution network HRNet and the residual neural network ResNet in the target detection module, based on multiple preset unit directions.
[0140] Optionally, the first extraction submodule is specifically used for:
[0141] Based on multiple preset key points, the heat map coordinates of multiple preset key points are extracted from multiple target bounding boxes using the high-resolution network HRNet.
[0142] The heatmap coordinate values of multiple preset key points are transformed to obtain the homogeneous two-dimensional coordinate values of the multiple preset key points.
[0143] Optionally, the target positioning device 3 further includes:
[0144] The training set acquisition module is used to acquire the training dataset, which includes multiple sample images and the target annotation results of the multiple sample images.
[0145] The model training module is used to train an initial object detection model using multiple sample images and the object annotation results of multiple sample images, thereby obtaining the object detection model.
[0146] It should be noted that the information interaction and execution process between the modules in the target positioning device 3 described above are based on the same concept as the method embodiment of this application. For details on their specific functions and the resulting technical effects, please refer to the method embodiment section. They will not be repeated here.
[0147] This application also provides a robot, such as... Figure 6 As shown, Figure 6 This is a schematic diagram of the structure of a robot provided in one embodiment of this application. (Refer to...) Figure 6 The robot 4 in this embodiment includes a controller 41, which is used to execute the steps in any of the above-described target localization method embodiments.
[0148] The robot in this embodiment can be an autonomous baggage cart recovery robot in airports. This robot can effectively detect and locate baggage carts using the target localization method, achieving good positioning even when the baggage carts are partially obscured, thus improving the accuracy of baggage cart positioning, reducing manpower requirements, and increasing the efficiency of baggage cart recovery. It can also be an autonomous parcel transport robot in logistics warehouses. This robot can effectively detect and locate parcels using the target localization method, thereby achieving parcel grabbing and transport, while reducing manpower requirements and improving parcel transport efficiency. It can also be a robot used in scenarios such as hospitals and shopping malls that require object localization, but this embodiment does not specifically limit it to this.
[0149] It should be noted that the information interaction and execution process between the modules in the above-mentioned robot 4 are based on the same concept as the method embodiment of this application. For details on their specific functions and technical effects, please refer to the method embodiment section, and they will not be repeated here.
[0150] This application also provides a terminal device, such as... Figure 7 As shown, Figure 7 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. (Refer to...) Figure 7 The terminal device 5 in this embodiment includes: a memory 51, a processor 52, and a computer program stored in the memory 51 and executable on the processor 52. When the processor 52 executes the computer program, it implements the steps in any of the above-described target positioning method embodiments.
[0151] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps described in the above-described method embodiments.
[0152] This application provides a computer program product that, when run on a mobile terminal, enables the mobile terminal to implement the steps described in the above-described method embodiments.
[0153] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographic device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0154] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0155] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0156] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0157] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0158] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A target localization method, characterized in that, include: During the process of the robot retrieving the target device, a two-dimensional scene image containing the target device is acquired; The target device in the two-dimensional scene image is identified to obtain the identification result related to the target device. The identification result includes the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions. The initial attitude information of the target device is obtained based on the homogeneous two-dimensional coordinate values of the plurality of preset key points and the probability distribution values of the plurality of preset unit directions; wherein, the initial attitude information includes: the initial center point coordinate values and the initial direction of the target device; The initial center point coordinates and initial direction of the target device are filtered to obtain the target attitude information of the target device, wherein the target attitude information includes: the center point coordinates and direction of the target device after filtering; The robot receives driving path information generated based on the target posture information. This driving path information includes control commands and a driving path. The robot then retrieves the target device along the driving path according to the control commands. The step of obtaining the initial attitude information of the target device based on the homogeneous two-dimensional coordinate values of the plurality of preset key points and the probability distribution values of the plurality of preset unit directions includes: The three-dimensional coordinates of the multiple preset key points are determined based on the homogeneous two-dimensional coordinates of the multiple preset key points and the camera intrinsic parameter matrix. The initial center point coordinates of the target device are determined based on the three-dimensional coordinates of the plurality of preset key points. The preset unit direction with the highest probability distribution value among the plurality of preset unit directions is determined as the initial direction of the target device.
2. The target localization method as described in claim 1, characterized in that, The step of filtering the initial center point coordinates and initial direction of the target device to obtain the target attitude information of the target device includes: Based on the initial center point coordinates and the initial direction in the initial attitude information, the updated moving average value is calculated using the improved moving average filter MMAF. Based on the updated moving average value, outliers in the initial center point coordinates and initial direction are removed to obtain the target attitude information of the target device.
3. The target localization method as described in any one of claims 1-2, characterized in that, The step of identifying the target device in the two-dimensional scene image to obtain an identification result related to the target device includes: The target device in the two-dimensional scene image is detected by the target detection module, and multiple target bounding boxes containing the target device are generated. Based on multiple preset key points, the homogeneous two-dimensional coordinate values of the multiple preset key points of the target device are extracted from multiple target bounding boxes by the high-resolution network HRNet in the target detection module. Based on multiple preset unit directions, the probability distribution values of the multiple preset unit directions of the target device are extracted from multiple target bounding boxes by the high-resolution network HRNet and the residual neural network ResNet in the target detection module.
4. The target localization method as described in claim 3, characterized in that, The step of extracting homogeneous two-dimensional coordinate values of the multiple preset key points of the target device from multiple target bounding boxes using the high-resolution network HRNet in the target detection module based on multiple preset key points includes: Based on multiple preset key points, the heat map coordinate values of multiple preset key points are extracted from multiple target bounding boxes through the high-resolution network HRNet; The heat map coordinate values of the multiple preset key points are transformed to obtain the homogeneous two-dimensional coordinate values of the multiple preset key points.
5. The target localization method as described in claim 3, characterized in that, Before detecting the target device in the two-dimensional scene image using the target detection module and generating multiple target bounding boxes containing the target device, the method further includes: Obtain a training dataset, wherein the training dataset includes multiple sample images and target annotation results for the multiple sample images; An initial target detection model is trained using the multiple sample images and the target annotation results of the multiple sample images to obtain the target detection model.
6. A target positioning device, characterized in that, include: The acquisition module is used to acquire a two-dimensional scene image containing the target device during the process of the robot retrieving the target device; The recognition module is used to recognize the target device in the two-dimensional scene image and obtain recognition results related to the target device. The recognition results include homogeneous two-dimensional coordinate values of multiple preset key points and probability distribution values of multiple preset unit directions. The determination module is used to determine the target attitude information of the target device based on the homogeneous two-dimensional coordinate values of the plurality of preset key points and the probability distribution values of the plurality of preset unit directions, wherein the target attitude information includes: the center point coordinate values and directions of the target device after filtering; A sending module is used to send travel path information generated based on the target posture information to the robot, wherein the travel path information includes control commands and a travel path, and the robot is used to retrieve the target device along the travel path according to the control commands; wherein... The module to be determined includes: The determination submodule is used to obtain the initial attitude information of the target device based on the homogeneous two-dimensional coordinate values of the plurality of preset key points and the probability distribution values of the plurality of preset unit directions; wherein, the initial attitude information includes: the initial center point coordinate value and the initial direction of the target device; The filtering submodule is used to filter the initial center point coordinates and initial direction of the target device to obtain the target attitude information of the target device. The determined submodules include: The first determining unit is used to determine the three-dimensional coordinate values of the plurality of preset key points based on the homogeneous two-dimensional coordinate values of the plurality of preset key points and the camera intrinsic parameter matrix; The second determining unit is used to determine the initial center point coordinates of the target device based on the three-dimensional coordinates of the plurality of preset key points; The third determining unit is used to determine the preset unit direction with the highest probability distribution value among the plurality of preset unit directions as the initial direction of the target device.
7. A robot, characterized in that, The robot includes a controller for performing the method as described in any one of claims 1 to 5.
8. A computer program product, characterized in that, Includes a computer program, which, when run, causes the method as described in any one of claims 1 to 5 to be performed.