Target positioning method and device, robot and program product

The method improves target positioning accuracy for robots by using high-resolution networks and filter processing to handle partially obscured targets, enhancing operational efficiency and robustness in complex environments.

CN120307276AActive Publication Date: 2025-07-15SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510364315.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-15
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

The existing target positioning method cannot accurately identify and locate when facing partially obscured targets, resulting in the robot being unable to perform tasks effectively.

Method used

By acquiring the two-dimensional scene image, identify the homogeneous two-dimensional coordinate values of multiple preset key points of the target device and the probability distribution value of the preset unit direction, combined with filtering processing, the target attitude information of the target device is determined, and the driving path information is generated to guide the robot to recover the target device.

Benefits of technology

It improves the accuracy of positioning of robots on partially obscured targets in complex environments, enhances the localization robustness in dynamic environments, reduces the cost of data requirements, and improves the scalability of the method and feasibility of practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120307276A_ABST
    Figure CN120307276A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of robots, and provides a target positioning method and device, a robot and a program product.The method comprises the steps that in the process that the robot recycles target equipment, a two-dimensional scene image containing the target equipment is obtained; identifying a target device in the two-dimensional scene image to obtain an identification result related to the target device; determining target attitude information of the target equipment according to the homogeneous two-dimensional coordinate values of the plurality of preset key points and the probability distribution values of the plurality of preset unit directions in the recognition result; and driving path information generated according to the target attitude information is sent to the robot, so that the robot recycles the target equipment along the driving path according to the control instruction. According to the method, the position and the direction of the target equipment are identified, so that even if the target equipment is partially shielded, the position and the direction of the target equipment can be positioned, and the target positioning accuracy of the robot is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of robotics, and particularly relates to a target positioning method, device, robot, and program product. Background Art

[0002] In the field of robotic automation operations, especially in complex environments such as airports and logistics centers, the detection and positioning of targets by robots are key technologies for them to complete automated tasks. For example, in an airport, a robot can automatically collect scattered luggage carts and transport them to a designated area. This automated task can not only reduce the manpower requirement but also improve the collection efficiency of luggage carts.

[0003] However, existing target positioning methods often cannot accurately identify and position targets when faced with partially occluded targets, resulting in the robot being unable to effectively execute tasks. Summary of the Invention

[0004] Embodiments of this application provide a target positioning method, device, robot, and program product, which can solve the problem of low positioning accuracy when using existing target positioning methods to position partially occluded targets.

[0005] In a first aspect, embodiments of this application provide a target positioning method, which includes:

[0006] During the process of a robot retrieving a target device, obtain a two-dimensional scene image containing the target device;

[0007] Identify the target device in the two-dimensional scene image to obtain an identification result related to the target device, where the identification result includes homogeneous two-dimensional coordinate values of multiple preset key points and probability distribution values of multiple preset unit directions;

[0008] According to the homogeneous two-dimensional coordinate values of the multiple preset key points and the probability distribution values of the multiple preset unit directions, determine the target pose information of the target device, where the target pose information includes: the center point coordinate value and direction of the target device after filtering processing;

[0009] Send the driving path information generated according to the target pose information to the robot, where the driving path information includes a control instruction and a driving path, and the robot is used to retrieve the target device along the driving path according to the control instruction.

[0010] In a possible implementation manner of the first aspect, the determining the target pose information of the target device according to the homogeneous two-dimensional coordinate values of the multiple preset key points and the probability distribution values of the multiple preset unit directions includes:

[0011] Obtain the initial pose information of the target device according to the homogeneous two-dimensional coordinate values of the multiple preset key points and the probability distribution values of the multiple preset unit directions; wherein, the initial pose information includes: the initial center point coordinate value and the initial direction of the target device.

[0012] Perform filtering processing on the initial center point coordinate value and the initial direction of the target device to obtain the target pose information of the target device.

[0013] In a possible implementation manner of the first aspect, the obtaining the initial pose information of the target device according to the homogeneous two-dimensional coordinate values of the multiple preset key points and the probability distribution values of the multiple preset unit directions includes:

[0014] Determine the three-dimensional coordinate values of the multiple preset key points according to the homogeneous two-dimensional coordinate values of the multiple preset key points and the camera internal parameter matrix.

[0015] Determine the initial center point coordinate value of the target device according to the three-dimensional coordinate values of the multiple preset key points.

[0016] Determine the preset unit direction with the highest value among the probability distribution values of the multiple preset unit directions as the initial direction of the target device.

[0017] In a possible implementation manner of the first aspect, the performing filtering processing on the initial center point coordinate value and the initial direction of the target device to obtain the target pose information of the target device includes:

[0018] Calculate the updated moving average value according to the initial center point coordinate value and the initial direction in the initial pose information through an improved moving average filter MMAF.

[0019] Remove the outliers in the initial center point coordinate value and the initial direction according to the updated moving average value to obtain the target pose information of the target device.

[0020] In a possible implementation manner of the first aspect, the identifying the target device in the two-dimensional scene image to obtain an identification result related to the target device includes:

[0021] Detect the target device in the two-dimensional scene image through a target detection module to generate multiple target bounding boxes containing the target device.

[0022] According to a plurality of preset key points, homogeneous two-dimensional coordinate values of the plurality of preset key points of the target device are extracted from the plurality of target bounding boxes through a High-Resolution Network (HRNet) in the target detection module;

[0023] According to a plurality of preset unit directions, probability distribution values of the plurality of preset unit directions of the target device are extracted from the plurality of target bounding boxes through the High-Resolution Network (HRNet) and Residual Neural Network (ResNet) in the target detection module.

[0024] In a possible implementation manner of the first aspect, the step of extracting, according to a plurality of preset key points, homogeneous two-dimensional coordinate values of the plurality of preset key points of the target device from the plurality of target bounding boxes through the High-Resolution Network (HRNet) in the target detection module includes:

[0025] According to a plurality of preset key points, heatmap coordinate values of the plurality of preset key points are extracted from the plurality of target bounding boxes through the High-Resolution Network (HRNet);

[0026] The heatmap coordinate values of the plurality of preset key points are converted to obtain the homogeneous two-dimensional coordinate values of the plurality of preset key points.

[0027] In a possible implementation manner of the first aspect, before the target device in the two-dimensional scene image is detected by the target detection module to generate a plurality of target bounding boxes including the target device, the method further includes:

[0028] Obtain a training data set, where the training data set includes a plurality of sample images and target annotation results of the plurality of sample images;

[0029] Use the plurality of sample images and the target annotation results of the plurality of sample images to train an initial target detection model to obtain the target detection model.

[0030] In a second aspect, an embodiment of the present application provides a target positioning device, where the device includes:

[0031] An acquisition module, configured to acquire a two-dimensional scene image including the target device during the process of the robot recovering the target device;

[0032] An identification module, configured to identify the target device in the two-dimensional scene image to obtain an identification result related to the target device, where the identification result includes homogeneous two-dimensional coordinate values of a plurality of preset key points and probability distribution values of a plurality of preset unit directions;

[0033] A determination module, configured to determine the target pose information of the target device according to the homogeneous two-dimensional coordinate values of the multiple preset key points and the probability distribution values of the multiple preset unit directions, where the target pose information includes: the center point coordinate value and direction of the target device after filtering processing;

[0034] A sending module, configured to send the driving path information generated according to the target pose information to the robot, where the driving path information includes a control instruction and a driving path, and the robot is configured to recycle the target device along the driving path according to the control instruction.

[0035] In a third aspect, an embodiment of the present application provides a robot, where the robot includes a controller, and the controller is configured to execute the target positioning method described in any one of the above.

[0036] In a fourth aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the target positioning method described in any one of the above is implemented.

[0037] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the target positioning method described in any one of the above is implemented.

[0038] In a sixth aspect, an embodiment of the present application provides a computer program product, where when the computer program product runs on a terminal device, the terminal device is caused to execute the target positioning method described in any one of the first aspects above.

[0039] The beneficial effects of the embodiments of the present application compared with the prior art are:

[0040] An embodiment of the present application provides a target positioning method, including: during the process of a robot retrieving a target device, acquiring a two-dimensional scene image containing the target device; identifying the target device in the two-dimensional scene image to obtain an identification result related to the target device, where the identification result includes homogeneous two-dimensional coordinate values of multiple preset key points and probability distribution values of multiple preset unit directions; determining target pose information of the target device according to the homogeneous two-dimensional coordinate values of the multiple preset key points and the probability distribution values of the multiple preset unit directions, where the target pose information includes: the center point coordinate value and direction of the target device after filtering processing; sending the driving path information generated according to the target pose information to the robot, where the driving path information includes a control instruction and a driving path, and the robot is used to retrieve the target device along the driving path according to the control instruction. By using the above technical solution, the center point coordinate value of the target device is identified according to the homogeneous two-dimensional coordinate values of the multiple preset key points in the identification result, and the direction of the target device is identified according to the probability distribution values of the multiple preset unit directions in the identification result. The position and direction of the target device are respectively identified and processed, so that even if the target device is partially blocked, the position and direction of the target device can be located, improving the accuracy of the robot in target positioning. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 is a schematic flowchart of a target positioning method provided by an embodiment of the present application;

[0043] Figure 2 is a schematic diagram of determining the position of a luggage cart according to preset key points provided by an embodiment of the present application;

[0044] Figure 3 is a schematic diagram of determining the direction of a luggage cart provided by an embodiment of the present application;

[0045] Figure 4 is a schematic diagram of the process of a robot positioning and retrieving a luggage cart provided by an embodiment of the present application;

[0046] Figure 5 is a schematic structural diagram of a target positioning device provided by an embodiment of the present application;

[0047] Figure 6It is a schematic structural diagram of a robot provided by an embodiment of the present application;

[0048] Figure 7 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners

[0049] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures, technologies, etc. are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0050] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0051] It should also be understood that the term "and / or" as used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0052] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined" or "in response to determining" or "once detecting [the described condition or event]" or "in response to detecting [the described condition or event]" according to the context.

[0053] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0054] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0055] In the field of robotic automated operations, especially in complex environments such as airports and logistics centers, the detection and localization of targets by robots are key technologies for them to complete automated tasks. For example, in an airport, a robot can automatically collect scattered luggage carts and transport them to a designated area. This automated task can not only reduce the need for manpower but also improve the collection efficiency of luggage carts. However, existing target localization methods often fail to accurately identify and locate targets when they are partially occluded, resulting in the robot being unable to effectively execute tasks.

[0056] Existing target localization methods include those based on visual models, such as convolutional neural networks (CNNs) based on deep learning. This visual model has good localization results when the target is fully visible without occlusion, but its localization performance will decline when the target is partially occluded. In addition, there are also some target localization methods that rely on multiple sensors (such as wireless signals, visual sensors, and lidar). These methods can provide relatively accurate localization information under ideal conditions, but in practical applications, especially when the target is partially occluded or the environment is complex and changeable, they often cannot observe a sufficient number of key points, resulting in the inability to accurately solve the pose of the target, making the accuracy and real-time performance of their localization face many challenges.

[0057] To improve the autonomous operation ability of robots in complex environments, especially when the target is partially occluded, a new perception strategy is needed that can accurately and efficiently locate the target while relying only on limited key point information. This strategy not only needs to have strong anti-occlusion ability, but also should be able to respond quickly in real-time environments to support the dynamic planning and precise operation of robots. Based on this, the embodiments of this application provide a target positioning method, which can be applied to the detection and positioning of targets by robots in complex environments. By hierarchically processing the position and orientation of the target device, the accuracy of target device positioning in the case of partial occlusion is improved; at the same time, through the progressive perception strategy, the ability of the robot to continuously update the pose of the target device in a dynamic environment is enhanced, and the robustness of positioning is improved; and, by simplifying the data requirements, the implementation cost of this method is reduced, and the scalability and practical application feasibility of this method are improved.

[0058] Please refer to Figure 1 , Figure 1 FIG. is a schematic flowchart of a target positioning method provided by an embodiment of this application. By way of example and not limitation, this method can be applied to terminal devices, such as robots. This method can be applied to the positioning and recovery system of a robot for autonomously retrieving luggage carts in an airport, and can also be applied to the transportation system of a robot for autonomously collecting or transporting objects in a logistics warehouse, hospital, etc. that need to autonomously collect or transport objects. As Figure 1 shown, this method includes:

[0059] S11. During the process of the robot retrieving the target device, obtain a two-dimensional scene image containing the target device.

[0060] In this embodiment, the execution subject can be a terminal device such as a server, such as the target positioning system in a robot. In some examples, the robot can be a terminal device used to collect or transport objects. The robot needs to identify and locate a certain target device, and after locating the target device, drive to the target device to collect or transport it. The target device is the device or object that the robot needs to retrieve. For example, in an airport, the robot can be a device with a target positioning system that can autonomously retrieve luggage carts, and the target device is the luggage cart that needs to be retrieved.

[0061] The two-dimensional scene image can be a two-dimensional image of the actual scene where the target device is located. The two-dimensional scene image can include the target device, the occluder of the target device, and other objects, etc. The two-dimensional scene image can be obtained by the robot through a camera or other image acquisition devices, and this is not limited in this embodiment.

[0062] S12. Identify the target device in the two-dimensional scene image to obtain an identification result related to the target device.

[0063] Among them, the recognition result includes the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions.

[0064] After obtaining the two-dimensional scene image of the location where the target device is located, it is necessary to recognize the target device in the two-dimensional scene image. In this embodiment, the target device in the two-dimensional scene image can be recognized by a target detection model (such as the YOLOV5 detection model), so as to obtain a recognition result related to the target device. The recognition result may include the homogeneous two-dimensional coordinate values of multiple preset key points, the probability distribution values of multiple preset unit directions, and may also include other content related to the target device, such as the usage status (occupied, idle) of the target device, the usage environment (indoor, outdoor), etc.

[0065] The preset key points are significant position points preset on the target device for identifying the target device. The position of the target device can be determined through these key points. Among them, there can be multiple preset key points, and the setting of the number is related to the actual application and is not specifically limited in this embodiment. Exemplarily, as Figure 2 shown Figure 2 is a schematic diagram for determining the position of a luggage cart according to preset key points provided by an embodiment of the present application. During the process of a robot positioning and retrieving a luggage cart, there can be six preset key points for recognizing the luggage cart, specifically the positions where the four wheels of the luggage cart intersect with the vehicle body (such as Figure 2 a, b, c, d in Figure 2 ), and the two ends of the luggage cart handle (such as

[0066] In a two-dimensional plane, a pair of two-dimensional coordinate values (x, y) is usually used to represent the exact position of a point in the two-dimensional plane, or a vector (x, y) is used to calibrate the exact position of a point in the two-dimensional plane. If an additional variable w is added on this basis, the two-dimensional coordinate values (x, y) of this point become (x, y, w) in the homogeneous coordinate system, and the coordinate values (x, y, w) are the homogeneous two-dimensional coordinate values corresponding to the point (x, y).

[0067] In this embodiment, 360 degrees is equally divided according to a preset division rule, and each unit obtained by the division is the preset unit direction. For example, if 360 degrees is evenly divided into 36 parts, it means there are 36 preset unit directions, and the angle of each preset unit direction is 10 degrees, that is, starting from 0 degrees, each time it increases by 10 degrees to represent a direction until it reaches 360 degrees. The probability distribution value of the preset unit direction can be the probability distribution value of a certain direction of the target device. This probability distribution value can be calculated based on information such as features and colors in the two-dimensional scene image, and can reflect information such as the orientation or movement trend of the target device in the two-dimensional scene image. The initial direction of the target device can be initially estimated through the probability distribution value of the preset unit direction.

[0068] In a possible implementation manner, before detecting and recognizing the target device in the two-dimensional scene image through the target detection module, it is necessary to construct an initial target detection model and perform model training on the constructed initial target detection model to obtain a target detection module capable of recognizing the target device from the two-dimensional image. The specific training process is as follows:

[0069] Obtain a training data set, where the training data set includes multiple sample images and the target annotation results of multiple sample images.

[0070] Use multiple sample images and the target annotation results of multiple sample images to train the initial target detection model to obtain the target detection model.

[0071] The initial target detection model is a newly constructed deep learning model that has not undergone model training. The training data set is a data subset used to train the initial target detection model, and these data contain known input features and expected output labels, that is, sample images and the target annotation results of sample images. The sample image is a two-dimensional scene image collected containing the target device; the target annotation result of the sample image is the recognition result output by the model. Continuing with the above example, during the process of the robot positioning and recovering the luggage cart, the target detection model can be trained first. The training data set constructed can be a data set containing tens of thousands of two-dimensional scene images. The sample images in this data set can contain different states of the luggage cart (occupied, idle), various occlusion levels (such as visibility of 40%, visibility of 80%, etc.) and different scene environments (indoor, outdoor). The target annotation results of the sample images can contain target bounding boxes, multiple preset key points, direction (the orientation of the luggage cart), usage status, etc. Through these target annotation results, the recognition results related to the idle luggage cart can be obtained.

[0072] It should be understood that the initial target detection model is trained with a large number of sample data sets, so that the target detection model can accurately identify the recognition results related to the target device from the new two-dimensional scene image, that is, the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions, providing accurate recognition results for subsequent processing of the position and direction of the target device.

[0073] It should be understood that by only annotating and training two-dimensional images, the complexity of data collection and annotation is reduced, and the scalability of the system and the feasibility of practical applications are improved.

[0074] In some embodiments, to identify the target device in the two-dimensional scene image and obtain the recognition results related to the target device, including:

[0075] The target detection module is used to detect the target device in the two-dimensional scene image and generate multiple target bounding boxes containing the target device.

[0076] According to multiple preset key points, the homogeneous two-dimensional coordinate values of multiple preset key points of the target device are extracted from multiple target bounding boxes through the High-Resolution Network (HRNet) in the target detection module.

[0077] According to multiple preset unit directions, the probability distribution values of multiple preset unit directions of the target device are extracted from multiple target bounding boxes through the High-Resolution Network (HRNet) and the Residual Neural Network (ResNet) in the target detection module.

[0078] In this embodiment, the target bounding box can be understood as a rectangular box drawn around the target device in the two-dimensional scene image, and this target bounding box can represent the approximate position of the target device in the image. The High-Resolution Network (HRNet) is a deep learning network architecture used to extract high-resolution features in images. In this embodiment, the HRNet is used to extract the coordinate values of preset key points from the target bounding box. The Residual Neural Network (ResNet) is also a deep learning network architecture. By introducing residual connections, it simplifies the training process of the deep network and can improve the performance of the model. In this embodiment, the Residual Neural Network ResNet is combined with the High-Resolution Network HRNet to extract the features of the target device in the preset unit direction from the target bounding box, so as to calculate the probability distribution value of the preset unit direction.

[0079] In a specific embodiment, first, the target detection module processes the two-dimensional scene image to identify multiple target bounding boxes containing the target device, so as to determine the approximate position of the target device in the two-dimensional scene image; after determining the approximate position of the target device in the two-dimensional scene image, the High-Resolution Network (HRNet) extracts features from the image region in the target bounding box, and extracts the positions of multiple preset key points in the two-dimensional scene image according to multiple preset key points, obtaining the homogeneous two-dimensional coordinate values of these preset key points. The homogeneous two-dimensional coordinate values can be used to conveniently extend the two-dimensional coordinate points into the three-dimensional space. In addition, the HRNet and the Residual Neural Network (ResNet) extract features and classify the image region in the target bounding box to obtain the probability distribution values of the target device in each preset unit direction. The probability distribution value can represent the probability that the direction of the target device is the preset unit direction.

[0080] In some embodiments, according to multiple preset key points, the HRNet in the target detection module extracts the homogeneous two-dimensional coordinate values of multiple preset key points of the target device from multiple target bounding boxes, including:

[0081] According to multiple preset key points, the HRNet extracts the heatmap coordinate values of multiple preset key points from multiple target bounding boxes.

[0082] Convert the heatmap coordinate values of multiple preset key points to obtain the homogeneous two-dimensional coordinate values of multiple preset key points.

[0083] It should be noted that in the process of extracting the homogeneous two-dimensional coordinate values of multiple preset key points of the target device from multiple target bounding boxes by the HRNet, first, according to multiple preset key points, the heatmap coordinate values of multiple preset key points are extracted from the target bounding box. Among them, the heatmap coordinate value refers to the position coordinate of the preset key point on the heatmap. The heatmap can be used to represent the position information of the preset key point. Then, the heatmap coordinate values are converted into the homogeneous two-dimensional coordinate values of multiple preset key points. In this embodiment, the conversion method of converting the heatmap coordinate values of multiple preset key points into homogeneous two-dimensional coordinate values can be to map the heatmap coordinate values into the pixel coordinate system of the image and add homogeneous coordinates to obtain the third-dimensional vector, or other conversion methods, which are not specifically limited in this embodiment. It should be understood that by converting the heatmap coordinate values into homogeneous two-dimensional coordinate values, the two-dimensional coordinate points can be conveniently extended into the three-dimensional space, which helps with the subsequent position processing of the target device.

[0084] Continuing with the above example, during the process of the robot locating and retrieving the luggage cart, the robot performs real-time detection on the two-dimensional scene image of the luggage cart through the target detection model, generating multiple target bounding boxes. Then, the homogeneous two-dimensional coordinate values of six preset key points of the luggage cart and the probability distribution values of the preset unit direction of the luggage cart are extracted from these target bounding boxes. Specifically: First, the high-resolution network HRNet is used to predict the heatmap coordinate values p i =[x i ,y i T of the six preset key points of the luggage cart, generating the homogeneous two-dimensional coordinate values corresponding to the preset key points After that, the high-resolution network HRNet and the residual neural network ResNet are combined to detect the unit direction, that is, the probability distribution values of n preset unit directions are output through the fully connected layer and the softmax layer

[0085] S13. According to the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions, determine the target pose information of the target device.

[0086] Among them, the target pose information includes: the center point coordinate value and the direction of the target device after filtering processing.

[0087] The target pose information can refer to the center point coordinate value and the direction (i.e., the orientation of the target device) of the target device after filtering processing, that is, the final position and direction of the target device obtained. Among them, the center point coordinate value of the target device can refer to the position of a fixed point of the target device in the two-dimensional scene image, and this point is usually determined based on the shape, size or preset rules of the target device. For example, if the shape of the target device is a regular shape (such as a circle, a square), the center point coordinate value of the target device can be the coordinate value of its geometric center; if the shape of the target device is an irregular shape or its mass distribution is uneven (such as the target device is a luggage cart), the center point coordinate value of the target device can be its centroid; of course, the center point coordinate value of the target device can also be a specific point preset according to the actual situation.

[0088] ​The center point coordinate value and direction of the target device after filtering processing can be understood as the center point coordinate value and direction obtained after excluding outliers through filtering processing. An outlier can refer to a point in the data that is significantly different from most data points, and this outlier may be caused by detection errors, noise, or other reasons. Compared with the center point coordinate value and direction without filtering processing, the center point coordinate value and direction after filtering processing can more accurately describe the position and direction of the target device. Filtering processing is a method for smoothing data, reducing noise, and improving data accuracy. In this embodiment, by performing filtering processing on the center point coordinate value and direction, outliers caused by detection errors or data fluctuations are eliminated.

[0089] S14. Send the driving path information generated according to the target pose information to the robot.

[0090] Among them, the driving path information includes control instructions and a driving path, and the robot is used to recycle the target device along the driving path according to the control instructions.

[0091] Specifically, after obtaining the target pose information of the target device, the driving path information can be generated by the Multi-Risk-RRT algorithm. Among them, the driving path information can include control instructions for controlling the robot to drive and the driving path for the robot to recycle the target device. After generating the driving path information, the control instructions are sent to the robot, and the robot will navigate to the target device along the generated driving path according to the control instructions, while avoiding pedestrians and obstacles, so as to accurately drive to the position of the target device to complete the grasping of the target device.

[0092] It should be noted that during the process of the robot driving towards the target device, the target positioning system of the robot will continuously detect the target device and continuously update the target pose information of the target device, so that the robot can accurately drive to the target device to complete the grasping of the target device. After the robot completes the grasping of the target device, the positioning of the target device by the robot will also stop. It should be understood that by continuously updating the positioning information of the target device when the robot approaches the target device, the positioning accuracy is gradually improved, thereby reducing the positioning error caused by long-distance detection.

[0093] In this embodiment, the Multi-Risk-RRT algorithm is a path planning algorithm designed for multi-risk environments. It considers various potential risk factors and can plan a safer and more reliable path. The Multi-Risk-RRT algorithm can combine multi-directional search and heuristic sampling to efficiently integrate the heuristic information of the dynamic sub-tree into the root tree, overcoming the limitations of the Two-Point Boundary Value Problem (TBVP) solver, thereby improving the motion planning performance in static and dynamic environments. In this embodiment, the specific process of generating the driving path information through the Multi-Risk-RRT algorithm will not be elaborated.

[0094] It can be understood that a target positioning method provided in this embodiment includes: during the process of the robot retrieving the target device, acquiring a two-dimensional scene image containing the target device; identifying the target device in the two-dimensional scene image to obtain an identification result related to the target device, where the identification result includes the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions; determining the target attitude information of the target device according to the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions, where the target attitude information includes: the center point coordinate value and direction of the target device after filtering; sending the driving path information generated according to the target attitude information to the robot, where the driving path information includes control instructions and a driving path, and the robot is used to retrieve the target device along the driving path according to the control instructions. By using the above technical solution, the center point coordinate value of the target device is identified according to the homogeneous two-dimensional coordinate values of multiple preset key points in the identification result, and the direction of the target device is identified according to the probability distribution values of multiple preset unit directions in the identification result. The position and direction of the target device are respectively identified and processed, so that even if the target device is partially blocked, the position and direction of the target device can be located, improving the accuracy of the robot's target positioning.

[0095] In some embodiments, determining the target attitude information of the target device according to the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions includes:

[0096] Obtaining the initial attitude information of the target device according to the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions; where the initial attitude information includes: the initial center point coordinate value and initial direction of the target device.

[0097] Performing filtering processing on the initial center point coordinate value and initial direction of the target device to obtain the target attitude information of the target device.

[0098] Relative to the target pose information, the initial pose information can be understood as the position and orientation of the target device that has not been filtered, that is, the initial center point coordinate value and the initial orientation. In a specific embodiment, first, the initial center point coordinate value of the target device is calculated according to the homogeneous two-dimensional coordinate values of multiple preset key points. In this embodiment, the average value or weighted average value of these homogeneous two-dimensional coordinate values of the preset key points can be calculated to obtain the initial center line point coordinate value; and the initial orientation of the target device is calculated according to the probability distribution values of multiple preset unit directions. In this embodiment, the preset unit direction corresponding to the highest probability distribution value can be selected as the initial orientation of the target device, or the preset unit direction corresponding to the weighted average value of the probability distribution values can be used as the initial orientation of the target device. Then, the initial center point coordinate value and the initial orientation are filtered by an improved moving average filter (MMAF) to exclude outliers caused by detection errors or data fluctuations, and the final position and orientation of the target device, that is, the center point coordinate value and the orientation, are obtained.

[0099] In some embodiments, according to the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions, the initial pose information of the target device is obtained, including:

[0100] According to the homogeneous two-dimensional coordinate values of multiple preset key points and the camera internal parameter matrix, the three-dimensional coordinate values of multiple preset key points are determined.

[0101] According to the three-dimensional coordinate values of multiple preset key points, the initial center point coordinate value of the target device is determined.

[0102] The preset unit direction with the highest value among the probability distribution values of multiple preset unit directions is determined as the initial orientation of the target device.

[0103] The camera internal parameter matrix describes the internal characteristics of image acquisition devices such as cameras. The camera internal parameter matrix may include parameters such as focal length and coordinates. Through the camera internal parameter matrix, the conversion from two-dimensional coordinate values to camera coordinates can be realized, providing a basis for obtaining three-dimensional coordinate values later. The three-dimensional coordinate values of multiple preset key points can be expressed as the coordinate values of the preset key points in the three-dimensional camera coordinate system, which can be calculated through the camera internal parameter matrix and the two-dimensional coordinate values of the preset key points, reflecting the position information of the preset key points in three-dimensional space.

[0104] In a specific embodiment, after obtaining the recognition result by recognizing the two-dimensional scene image through the target detection model, the homogeneous two-dimensional coordinate values of multiple preset key points in the recognition result can be first input into the camera projection model in the robot. Combining the projection transformation of the camera internal parameter matrix, the homogeneous two-dimensional coordinate values of multiple preset key points in the recognition result are converted into the three-dimensional coordinate values of multiple preset key points, so as to obtain the position information of multiple preset key points in the three-dimensional space; then, according to the position information of multiple preset key points in the three-dimensional space, the initial center point coordinate value of the target device is determined, that is, the position of the center point of the target device is obtained. In this embodiment, the initial center line point coordinate value of the target device can be calculated through the prior model M = {Y i =(x i ,y i ,ζ i )∈R 3 |i = 0,1,...,n} and the ground plane constraint , where Y represents the coordinate value of the initial center point of the target device, Y i represents the coordinate value of the preset key point, x i , y i respectively represent the horizontal position values of the preset key point relative to the initial center point of the target device, ζ i represents the height value of the preset key point relative to the initial center point of the target device, i represents multiple preset key points, λ represents the vertical distance from the camera to the ground in the camera projection model, λ is fixed, and N represents the actual height from the camera to the ground.

[0105] For example, as Figure 2 shown, Figure 2 is a schematic diagram of determining the position of a luggage cart according to preset key points provided in an embodiment of the present application. Figure 2 shows the process of the robot positioning and recovering the luggage cart. It is assumed that the bottom of the luggage cart is on the same horizontal plane as the bottom of the robot, and the vertical height λ from the camera center point on the robot to each preset key point is known. First, the homogeneous two-dimensional coordinate values of multiple preset key points (such as a, b, c, d, e, f) on the two-dimensional scene image plane of the luggage cart are obtained through the target detection model; Figure 2 shows that (u, v) represents the normal vector, and N represents the actual height from the camera to the ground; then, the ray direction ρ of these preset key points can be calculated by using the inverse matrix K -1 of the camera internal parameter matrix K of the camera projection model; then, combining the height value λ of the camera in the camera projection model and the height values ζ i of multiple preset key points, the actual positions X i of multiple preset key points in the three-dimensional space are calculated.; Finally, by calculating the average deviation between the actual positions of multiple preset key points in the three-dimensional space and the corresponding positions Y in the above-mentioned prior model M i the coordinate value of the center point of the luggage cart is determined. That is, the actual position (X i , vis) of each preset key point can be subtracted from the position (Y i , vis) in the prior model M, and then the average value is taken to obtain the position (x, y) of the center point of the luggage cart.

[0106] After determining the position of the center point of the target device, the direction of the target device is determined. That is, according to the probability distribution values of multiple preset unit directions, the unit direction with the highest probability distribution value is selected, and then this unit direction is determined as the initial direction of the target device. Specifically, the unit direction corresponding to the highest probability value can be selected from the probability distribution values of multiple preset unit directions provided by the target detection model. Exemplarily, 360-degree discrete representation can be used for accurate direction estimation. Define the loss function for calculating the deviation between the predicted direction and the true direction, where represents the probability distribution values of multiple preset unit directions, j represents multiple preset unit directions, represents the "circular" Gaussian probability distribution. Finally, the initial direction θ of the target device can be obtained by selecting the unit direction corresponding to the highest probability distribution value, that is as Figure 3 shown, Figure 3 FIG. is a schematic diagram of determining the direction of a luggage cart provided by an embodiment of the present application, Figure 3 in which, it is shown that the initial direction of the luggage cart is 35 degrees.

[0107] So far, through this embodiment, the initial position (initial center line point coordinate value) and the initial direction of the target device can be determined, realizing the positioning of the initial pose of the target device.

[0108] It should be understood that in this embodiment, by using independent modules to process the position and direction of the target device respectively, it is realized that even when some key points are detected (the target device is partially occluded), the position of the target device can be located and its direction can be estimated, improving the accuracy of the robot's target positioning. At the same time, the operation ability and robustness of the robot in a dynamic environment are improved.

[0109] In some embodiments, the initial center point coordinate value and the initial direction of the target device are filtered to obtain the target pose information of the target device, including:

[0110] According to the initial center point coordinate value and the initial direction in the initial pose information, the updated moving average value is calculated through the improved moving average filter MMAF.

[0111] Based on the updated moving average, outliers in the initial center point coordinate values and the initial direction are removed to obtain the target attitude information of the target device.

[0112] In a specific embodiment, after obtaining the initial target attitude information (initial center point coordinate values and initial direction) of the target device, the initial target attitude information can be filtered by an improved moving average filter MMAF to exclude outliers caused by detection errors or data fluctuations, so as to obtain the target attitude information of the target device. Specifically: the initial attitude information (initial center point coordinate values and initial direction) is input into the improved moving average filter MMAF, and the MMAF filter will calculate the moving average of a series of data points according to its internal algorithm (such as adaptive window size, weight assignment, etc.). These moving average values can be used as the updated center point coordinate values and direction. For example, the parameters of the MMAF filter can be initialized as F = (Δ, Θ_z), where Δ represents the window size of the moving average, and Θ_z represents the threshold for outlier detection of the standard score z-score; after obtaining these updated moving average values, they can be compared with the initial data points (initial center point coordinate values and initial direction). If the difference between a certain data point and the moving average value exceeds a preset threshold (such as the preset threshold is Θ_z, and this preset threshold can be set according to the actual situation), this data point can be removed as a discrete point, so as to obtain the target attitude information of the target device. Through the above filtering process, the noise and errors in the recognition process can be reduced, and the accuracy and reliability of the target attitude information can be improved.

[0113] Continuing with the above example, as Figure 4 shown, Figure 4 is a schematic diagram of a process for a robot to locate and recover a luggage cart provided in an embodiment of the present application. Figure 4 In it, the robot first obtains a two-dimensional scene image containing the luggage cart, inputs the two-dimensional scene image into the luggage cart detection module in the target detection model, and the luggage cart in the image is recognized by the luggage cart detection module to output a target bounding box. Then, the recognition result, that is, the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions Then, the homogeneous two-dimensional coordinate values of multiple preset key points are processed by the key point processing module to output the initial center point coordinate values (x, y) of the luggage cart. The probability distribution values of multiple preset unit directions are processed by the direction processing module to output the initial direction (θ) of the luggage cart. Then, the initial center point coordinate values (x, y) of the luggage cart and the initial direction (θ) of the luggage cart are filtered by a filter to output the target pose information (x f , y f , θ f ) of the filtered luggage cart. Finally, the target pose information (x f , y f , θ f ) of the luggage cart is input into the trajectory planning module. The trajectory planning module outputs the driving path information (v, ω) according to the target pose information (x f , y f , θ f ) of the luggage cart. The robot drives towards the luggage cart according to the driving path information (v, ω) to complete the recovery of the luggage cart.

[0114] It can be understood that a target positioning method provided by an embodiment of the present application may have the following beneficial effects:

[0115] (1) By processing the position and direction of the target device separately, the accuracy of positioning in the case where the target device is partially occluded is improved, and the problem of positioning failure of the existing target positioning method under occlusion conditions is solved.

[0116] (2) By continuously identifying and updating the pose information of the target device in real time during the process of the robot approaching the target device, the ability of the robot to continuously update the target pose in a dynamic environment is enhanced, and the accuracy and robustness of target positioning are improved.

[0117] (3) By detecting the two-dimensional image of the target device, the requirements for input data are simplified, the implementation cost of the target recognition method is reduced, the scalability of the method and the feasibility of practical application are improved, so that the method can be more widely applied to the robot target positioning tasks in various complex environments.

[0118] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0119] Corresponding to the target positioning method in the above embodiment, Figure 5 FIG. shows a schematic structural diagram of a target positioning device provided by an embodiment of the present application. For the sake of illustration, only the parts related to the embodiments of the present application are shown.

[0120] Refer to Figure 5 , the target positioning device 3 of this embodiment includes:

[0121] An acquisition module 31, configured to acquire a two-dimensional scene image including a target device during the process of the robot recycling the target device.

[0122] An identification module 32, configured to identify the target device in the two-dimensional scene image to obtain an identification result related to the target device, where the identification result includes homogeneous two-dimensional coordinate values of multiple preset key points and probability distribution values of multiple preset unit directions.

[0123] A determination module 33, configured to determine the target pose information of the target device according to the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions, where the target pose information includes: the center point coordinate value and direction of the target device after filtering processing.

[0124] A sending module 34, configured to send the driving path information generated according to the target pose information to the robot, where the driving path information includes control instructions and a driving path, and the robot is configured to recycle the target device along the driving path according to the control instructions.

[0125] It can be understood that in this embodiment, the present application provides a target positioning device 3. During the process of the robot recycling the target device, the acquisition module 31 acquires a two-dimensional scene image including the target device; the identification module 32 identifies the target device in the two-dimensional scene image to obtain an identification result related to the target device, where the identification result includes homogeneous two-dimensional coordinate values of multiple preset key points and probability distribution values of multiple preset unit directions; the determination module 33 determines the target pose information of the target device according to the homogeneous two-dimensional coordinate values of multiple preset key points and the probability distribution values of multiple preset unit directions, where the target pose information includes: the center point coordinate value and direction of the target device after filtering processing; the sending module 34 sends the driving path information generated according to the target pose information to the robot, where the driving path information includes control instructions and a driving path, and the robot is configured to recycle the target device along the driving path according to the control instructions. By using the target positioning device 3, the center point coordinate value of the target device is identified according to the homogeneous two-dimensional coordinate values of multiple preset key points in the identification result, and the direction of the target device is identified according to the probability distribution values of multiple preset unit directions in the identification result. The position and direction of the target device are respectively identified and processed, so that even if the target device is partially blocked, the position and direction of the target device can be located, improving the accuracy of the robot's target positioning.

[0126] Optionally, the determination module 33 includes:

[0127] A determination sub-module, configured to obtain initial pose information of a target device according to homogeneous two-dimensional coordinate values of a plurality of preset key points and probability distribution values of a plurality of preset unit directions; wherein, the initial pose information includes: an initial center point coordinate value and an initial direction of the target device.

[0128] A filtering sub-module, configured to perform filtering processing on the initial center point coordinate value and the initial direction of the target device to obtain target pose information of the target device.

[0129] Optionally, the determination sub-module includes:

[0130] A first determination unit, configured to determine three-dimensional coordinate values of a plurality of preset key points according to homogeneous two-dimensional coordinate values of the plurality of preset key points and an intrinsic camera matrix.

[0131] A second determination unit, configured to determine an initial center point coordinate value of the target device according to the three-dimensional coordinate values of the plurality of preset key points.

[0132] A third determination unit, configured to determine a preset unit direction with the highest value among the probability distribution values of the plurality of preset unit directions as the initial direction of the target device.

[0133] Optionally, the filtering sub-module includes:

[0134] An update unit, configured to calculate an updated moving average value through an improved moving average filter MMAF according to the initial center point coordinate value and the initial direction in the initial pose information.

[0135] An outlier removal unit, configured to remove outliers in the initial center point coordinate value and the initial direction according to the updated moving average value to obtain target pose information of the target device.

[0136] Optionally, the recognition module 32 includes:

[0137] A detection sub-module, configured to detect a target device in a two-dimensional scene image through a target detection module to generate a plurality of target bounding boxes including the target device.

[0138] A first extraction sub-module, configured to extract homogeneous two-dimensional coordinate values of a plurality of preset key points of the target device from the plurality of target bounding boxes through a high-resolution network HRNet in the target detection module according to the plurality of preset key points.

[0139] A second extraction sub-module, configured to extract probability distribution values of a plurality of preset unit directions of the target device from the plurality of target bounding boxes through a high-resolution network HRNet and a residual neural network ResNet in the target detection module according to the plurality of preset unit directions.

[0140] Optionally, the first extraction sub-module is specifically configured to:

[0141] According to multiple preset key points, obtain the heat map coordinate values of multiple preset key points from multiple target bounding boxes through the High-Resolution Network (HRNet).

[0142] Convert the heat map coordinate values of multiple preset key points to obtain the homogeneous two-dimensional coordinate values of multiple preset key points.

[0143] Optionally, the target positioning device 3 further includes:

[0144] A training set acquisition module, configured to acquire a training data set, where the training data set includes multiple sample images and target annotation results of multiple sample images.

[0145] A model training module, configured to train an initial target detection model using multiple sample images and target annotation results of multiple sample images to obtain a target detection model.

[0146] It should be noted that for the information interaction, execution process, etc. between the modules in the above target positioning device 3, since they are based on the same concept as the method embodiments of the present application, their specific functions and the technical effects brought about can be specifically referred to in the method embodiment part, and will not be elaborated here.

[0147] The embodiment of the present application further provides a robot, as Figure 6 shown, Figure 6 is a schematic structural diagram of a robot provided in an embodiment of the present application. Referring to Figure 6 , the robot 4 in this embodiment includes: a controller 41, and the controller 41 is configured to execute the steps in any one of the above-mentioned target positioning method embodiments.

[0148] The robot in the embodiment of the present application can be a robot for autonomous baggage cart recovery in an airport. This robot can effectively detect and locate the baggage cart through this target positioning method, and can also achieve good positioning even when the baggage cart is partially blocked, improving the accuracy of the robot's positioning of the baggage cart, thereby reducing the manpower requirement and improving the efficiency of baggage cart recovery; it can also be a robot for autonomous package transportation in a logistics warehouse. This robot can effectively detect and locate the package through this target positioning method, so as to achieve the grasping and transportation of the package, while reducing the manpower requirement and improving the efficiency of package transportation; it can also be a robot that needs to locate objects in scenarios such as hospitals and shopping malls. No specific limitation is made in this embodiment.

[0149] It should be noted that for the information interaction, execution process, etc. between the modules in the above robot 4, since they are based on the same concept as the method embodiments of the present application, their specific functions and the technical effects brought about can be specifically referred to in the method embodiment part, and will not be elaborated here.

[0150] The embodiments of the present application also provide a terminal device, as Figure 7 shown Figure 7 is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Referring to Figure 7 , the terminal device 5 of this embodiment includes: a memory 51, a processor 52, and a computer program stored in the memory 51 and executable on the processor 52. When the processor 52 executes the computer program, the steps in the embodiment of the above-mentioned target positioning method are implemented.

[0151] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0152] The embodiments of the present application provide a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal is enabled to implement the steps in the above-mentioned various method embodiments when executed.

[0153] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0154] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0155] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0156] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0157] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0158] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.

Claims

1. A target positioning method, characterized in that, Including: During the process of the robot recycling the target device, obtain a two-dimensional scene image including the target device; Identify the target device in the two-dimensional scene image to obtain an identification result related to the target device, where the identification result includes homogeneous two-dimensional coordinate values of multiple preset key points and probability distribution values of multiple preset unit directions; Determine the target pose information of the target device according to the homogeneous two-dimensional coordinate values of the multiple preset key points and the probability distribution values of the multiple preset unit directions, where the target pose information includes: the center point coordinate value and direction of the target device after filtering; Send the driving path information generated according to the target pose information to the robot, where the driving path information includes a control instruction and a driving path, and the robot is used to recycle the target device along the driving path according to the control instruction.

2. The target positioning method according to claim 1, wherein The determining the target pose information of the target device according to the homogeneous two-dimensional coordinate values of the multiple preset key points and the probability distribution values of the multiple preset unit directions includes: Obtain the initial pose information of the target device according to the homogeneous two-dimensional coordinate values of the multiple preset key points and the probability distribution values of the multiple preset unit directions; where the initial pose information includes: the initial center point coordinate value and initial direction of the target device; Perform filtering processing on the initial center point coordinate value and initial direction of the target device to obtain the target pose information of the target device.

3. The target positioning method according to claim 2, characterized in that The obtaining the initial pose information of the target device according to the homogeneous two-dimensional coordinate values of the multiple preset key points and the probability distribution values of the multiple preset unit directions includes: Determine the three-dimensional coordinate values of the multiple preset key points according to the homogeneous two-dimensional coordinate values of the multiple preset key points and the camera internal parameter matrix; Determine the initial center point coordinate value of the target device according to the three-dimensional coordinate values of the multiple preset key points; Determine the preset unit direction with the highest value among the probability distribution values of the multiple preset unit directions as the initial direction of the target device.

4. The target positioning method according to claim 2, characterized in that The performing filtering processing on the initial center point coordinate value and initial direction of the target device to obtain the target pose information of the target device includes: According to the initial center point coordinate value and the initial direction in the initial pose information, calculate the updated moving average value through the improved moving average filter MMAF; According to the updated moving average value, remove the outliers in the initial center point coordinate value and the initial direction to obtain the target pose information of the target device.

5. The target positioning method according to any one of claims 1-4, characterized in that, The identifying the target device in the two-dimensional scene image to obtain an identification result related to the target device includes: Detect the target device in the two-dimensional scene image through a target detection module to generate multiple target bounding boxes including the target device; According to multiple preset key points, homogeneous two-dimensional coordinate values of the multiple preset key points of the target device are extracted from the multiple target bounding boxes through the High-Resolution Network (HRNet) in the target detection module; According to multiple preset unit directions, probability distribution values of the multiple preset unit directions of the target device are extracted from the multiple target bounding boxes through the High-Resolution Network (HRNet) and the Residual Neural Network (ResNet) in the target detection module.

6. The target positioning method according to claim 5, characterized in that The step of extracting, according to multiple preset key points, homogeneous two-dimensional coordinate values of the multiple preset key points of the target device from the multiple target bounding boxes through the High-Resolution Network (HRNet) in the target detection module includes: According to multiple preset key points, heatmap coordinate values of the multiple preset key points are extracted from the multiple target bounding boxes through the High-Resolution Network (HRNet); The heatmap coordinate values of the multiple preset key points are converted to obtain the homogeneous two-dimensional coordinate values of the multiple preset key points.

7. The target positioning method according to claim 5, characterized in that Before detecting the target device in the two-dimensional scene image through the target detection module to generate multiple target bounding boxes containing the target device, the method further includes: Obtaining a training data set, where the training data set includes multiple sample images and target annotation results of the multiple sample images; Training an initial target detection model using the multiple sample images and the target annotation results of the multiple sample images to obtain the target detection model.

8. A target positioning device, characterized in that, It includes: An acquisition module, configured to acquire a two-dimensional scene image containing the target device during the process of the robot recycling the target device; An identification module, configured to identify the target device in the two-dimensional scene image to obtain an identification result related to the target device, where the identification result includes homogeneous two-dimensional coordinate values of multiple preset key points and probability distribution values of multiple preset unit directions; A determination module, configured to determine target pose information of the target device according to the homogeneous two-dimensional coordinate values of the multiple preset key points and the probability distribution values of the multiple preset unit directions, where the target pose information includes: the center point coordinate value and direction of the target device after filtering; A sending module, configured to send driving path information generated according to the target pose information to the robot, where the driving path information includes a control instruction and a driving path, and the robot is configured to recycle the target device along the driving path according to the control instruction.

9. A robot, characterized in that, The robot includes a controller, and the controller is configured to execute the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that, It includes a computer program, and when the computer program is run, the method according to any one of claims 1 to 7 is executed.

Citation Information

Patent Citations

  • Attitude estimation method and device, electronic equipment and storage medium

    CN113822102A

  • Machine room fusion positioning method, system and device and storage medium

    CN115060268A

  • Target positioning method and system and electronic equipment

    CN117095319A

  • Autonomous pickup and placement pose acquisition method for robot in disordered scene

    CN118081758A

  • Quadruped robot autonomous navigation method and system for special environment

    CN119469168A