Article positioning method and device for robot, electronic equipment and storage medium
By fusing wireless radio frequency signal strength, phase, and 3D point cloud information, and using a graph neural network for error compensation model, the problem of low robot positioning accuracy and poor reliability in complex environments was solved, achieving stable and high-precision object positioning.
Patent Information
- Application Number
- CN202511753872.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-01-27
AI Technical Summary
Existing technologies struggle to achieve centimeter-level precision in robot positioning of objects in complex environments, especially in non-line-of-sight scenarios where radio frequency identification and visual positioning methods suffer from large errors and poor reliability.
By acquiring the radio frequency signal strength, phase information, and 3D point cloud information of the target object, a graph neural network is used to perform a fusion perception error compensation model to predict the positioning information. Combined with the fixed conversion relationship between the radio frequency sensing device and the 3D vision sensor, deep fusion of multi-source information and accurate positioning are achieved.
Even when vision is obstructed or radio frequency signals are interfered with, stable and high-precision object positioning is achieved, improving the robot's positioning ability and success rate in complex environments.
Smart Images

Figure CN121403386A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device, and storage medium for robot positioning of objects. Background Technology
[0002] In fields such as intelligent robots, automated warehousing and logistics, accurate positioning of items is a key prerequisite for robots to complete tasks such as grasping and sorting.
[0003] Existing technical solutions mainly fall into two categories: one is positioning methods based on radio frequency identification (RFID) and UWB (Ultra Wide Band). While these methods can penetrate some obstructions, they suffer from multipath effects and large signal strength fluctuations in complex environments, resulting in positioning errors of several meters in non-line-of-sight scenarios, making it difficult to meet centimeter-level accuracy requirements. The other category is vision-based positioning methods. Although they offer high accuracy within the visible range, the vision system fails when the target is obstructed or in poor lighting conditions. Furthermore, traditional two-dimensional vision systems struggle to acquire accurate depth and pose information. Therefore, a solution to address these issues is urgently needed. Summary of the Invention
[0004] This application provides a method, apparatus, electronic device, and storage medium for robot positioning of objects, in order to address the deficiencies existing in the prior art.
[0005] This application provides a method for a robot to locate an object, including: Acquire the target object's radio frequency signal strength, phase information, and 3D point cloud information; The wireless radio frequency signal strength information, the phase information, and the 3D point cloud information are fused to generate fused perception information. Based on the fused sensing information, the location information of the target item is predicted by a fused sensing error compensation model; wherein, the fused sensing error compensation model is obtained by training a graph neural network; Based on the location information of the target item, the robot is controlled to move and navigate in order to locate the target item.
[0006] According to an embodiment of this application, a method for robot positioning of an object includes acquiring the target object's radio frequency signal strength information, phase information, and 3D point cloud information, comprising: The radio frequency signal strength and phase information of the target item are obtained through a wireless radio frequency sensing device; The robot acquires 3D point cloud information of the target object using its 3D vision sensor.
[0007] According to an embodiment of this application, a method for robot positioning of an object, before acquiring the radio frequency signal strength information, phase information, and 3D point cloud information of the target object, the method further includes: A fixed transformation relationship is determined between the wireless radio frequency sensing device and the 3D vision sensor; wherein the fixed transformation relationship includes a translation vector and a rotation matrix; Based on the fixed transformation relationship, the coordinates corresponding to the signal data collected by the wireless radio frequency sensing device are sequentially transformed to the 3D vision sensor coordinate system and the robot base coordinate system.
[0008] According to an embodiment of this application, a method for locating an object using a robot includes acquiring the radio frequency signal strength information and phase information of the target object through a radio frequency sensing device, comprising: The system receives a return signal from a tag and collects the signal strength indication and phase data of the return signal; wherein the tag is an RFID tag or a UWB tag attached to the target item; the tag is used to generate the return signal in response to the radio frequency signal emitted by the wireless radio frequency sensing device at a predetermined frequency; The acquired signal strength indication and phase data are filtered and denoised to obtain the wireless radio frequency signal strength information and the phase information.
[0009] A method for robot positioning of an object according to an embodiment of this application, wherein acquiring 3D point cloud information of the target object through the robot's 3D vision sensor includes: The target object is captured using a 3D camera in RGBD image format, and the RGBD image is converted into initial 3D point cloud data. The initial 3D point cloud data is preprocessed to obtain the 3D point cloud information of the target item.
[0010] According to an embodiment of this application, a method for robot positioning of an object includes fusing the radio frequency signal strength information, the phase information, and the 3D point cloud information to generate fused perception information, including: Based on the wireless radio frequency signal strength information, the phase information, the 3D point cloud information, and the position information when the robot collects information, a spatiotemporal graph structure is constructed; wherein, the spatiotemporal graph structure is used to: define the sampling points of the robot at different times as graph nodes, and define the spatial proximity and temporal continuity relationship between nodes as edges of the graph; The spatiotemporal graph structure is input into a pre-trained graph neural network model, which outputs the fused sensing information.
[0011] According to an embodiment of this application, a method for robot positioning of objects includes a training process for a fusion perception error compensation model, comprising: Acquire wireless radio frequency signal strength information, phase information, and 3D point cloud information in different scenarios, and align them with the actual position information of objects in different scenarios to construct a training dataset; The model parameters of the graph neural network are iteratively optimized based on the training dataset to minimize the error between the model's predicted position and the actual position of the item, thereby obtaining the fusion perception error compensation model.
[0012] A method for locating an object using a robot, according to an embodiment of this application, includes controlling the robot to move and navigate based on the object's location information to locate the object. If the target item is within the range of radio frequency sensing but not within the range of 3D vision, coarse positioning is performed based on the positioning information to obtain the candidate location range of the target item. Generate a spherical search region centered on the range of candidate locations; The robot is controlled to move and navigate to the spherical search area to locate the target item.
[0013] According to an embodiment of this application, a method for locating an object using a robot, wherein the robot is controlled to move and navigate to a spherical search area to locate the target object, includes: Control the robot's movement and navigation until the target object enters the 3D vision sensing range; The current point cloud data of the target object is acquired through the 3D vision sensor, and the coordinates of the target object in the camera coordinate system are calculated based on the current point cloud data using a 3D vision precision localization algorithm. The coordinates are transformed from the camera coordinate system to the robot base coordinate system to complete the positioning of the target object in the robot base coordinate system.
[0014] According to an embodiment of this application, a method for robot positioning of an object includes calculating the coordinates of the target object in the camera coordinate system based on the current point cloud data using a 3D visual precision positioning algorithm, comprising: The current point cloud data is segmented to obtain segmented point cloud data; From the segmented point cloud data, extract the independent point cloud clusters corresponding to the target item; Based on the independent point cloud cluster, the three-dimensional geometric center coordinates of the target object in the camera coordinate system are calculated.
[0015] According to an embodiment of this application, a method for locating an object using a robot, wherein the robot is controlled to move and navigate to a spherical search area to locate the target object, includes: The robot is controlled to move and navigate to an optimal distance range from the target object; wherein, the optimal distance range is a distance interval in which the positioning error after the fusion of radio frequency sensing and 3D vision is less than a preset threshold. Multiple sets of radio frequency positioning results are obtained at multiple robot locations within the preset distance range; Based on the multiple sets of radio frequency positioning results, the state estimation result is generated by estimating through the fusion sensing error compensation model. The state estimation results are corrected by iterative movement of the robot, and the location of the target item is finally determined.
[0016] According to an embodiment of this application, a method for locating an object using a robot, wherein controlling the robot to move and navigate based on the object's location information to locate the object, further includes: When the target item is within the range of radio frequency sensing and within the range of 3D vision, the precise positioning result based on 3D vision and the coarse positioning result based on radio frequency sensing are acquired simultaneously. The precise positioning result and the coarse positioning result are fused to generate fused coordinates in the robot's base coordinate system, thereby completing the positioning of the target object at the current moment.
[0017] This application also provides a device for robot positioning of objects, comprising: The acquisition module is used to acquire the wireless radio frequency signal strength information, phase information, and 3D point cloud information of the target item; The fusion module is used to fuse the wireless radio frequency signal strength information, the phase information, and the 3D point cloud information to generate fused perception information; The prediction module is used to predict the location information of the target item based on the fused sensing information and through a fused sensing error compensation model; wherein the fused sensing error compensation model is obtained based on graph neural network training; The positioning module is used to control the robot's movement and navigation to locate the target item based on the target item's positioning information.
[0018] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the robot positioning method as described above.
[0019] This application also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the robot positioning method as described above.
[0020] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the robot positioning method as described above.
[0021] This application provides a method, apparatus, electronic device, and storage medium for robot object localization. The method involves acquiring the radio frequency signal strength information, phase information, and 3D point cloud information of the target object; fusing the radio frequency signal strength information, phase information, and 3D point cloud information to generate fused perception information; predicting the target object's location information using a fused perception error compensation model based on the fused perception information; wherein the fused perception error compensation model is trained based on a graph neural network; and controlling the robot's movement and navigation to locate the target object based on its location information. Therefore, this application, by deeply fusing radio frequency signal strength, phase, and 3D point cloud information and using a graph neural network-based fused perception error compensation model for prediction, effectively overcomes the inherent limitations of low positioning accuracy and poor reliability of single sensor sources in non-line-of-sight or complex environments. Even when vision is obstructed or radio frequency signals are interfered with, stable and high-precision prediction of the target object's position can be achieved through model inference, thereby significantly improving the robot's object localization capability and success rate in complex real-world scenarios. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart illustrating the robot positioning method for objects provided in an embodiment of this application.
[0024] Figure 2 This is a system architecture diagram of a robot locating an object, provided in an embodiment of this application.
[0025] Figure 3 This is a flowchart illustrating the non-visual / visual object positioning method provided in the embodiments of this application.
[0026] Figure 4This is a schematic diagram of the structure of the robot positioning device provided in the embodiments of this application.
[0027] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the embodiments of this application.
[0029] Before describing the embodiments of this application, the terms and concepts involved in the embodiments of this application will be explained illustratively.
[0030] 3D vision, Three-dimensional Vision, 3D Vision.
[0031] Passive IoT sensing.
[0032] Radio Frequency Identification (RFID).
[0033] Ultra Wide Band (UWB)
[0034] Mobile robot.
[0035] Embodied Intelligence (EAI).
[0036] Signal strength, Received Signal Strength Indication, RSSI.
[0037] The following is combined Figures 1-5 This application describes a method, apparatus, electronic device, and storage medium for robot positioning of items, according to embodiments of the present application.
[0038] It should be noted that in today's increasingly automated industrial production, warehousing, and logistics sectors, accurate and efficient item positioning is crucial. Especially in complex working environments where non-line-of-sight scenarios frequently occur, related positioning technologies have revealed several shortcomings: Limitations of Traditional RFID or UWB Positioning Technologies: Traditional RFID / UWB tag-based positioning methods mostly rely on Signal Strength Indicator (RSSI). However, RSSI is highly susceptible to environmental interference. In non-line-of-sight environments, signal propagation is affected by obstacles such as metal and liquids, resulting in reflection, refraction, diffraction, and absorption, leading to significant signal strength fluctuations and making it difficult to guarantee positioning accuracy. Typically, RSSI-based positioning errors can reach several meters. In complex industrial plants or large warehouses, this level of accuracy is far from meeting practical needs. For example, in an industrial workshop filled with metal shelves and equipment, RFID signals are constantly reflected during propagation, causing a significant deviation between the RSSI-determined item location and the actual location. This can prevent robots from accurately grasping target items, reducing production efficiency.
[0039] The limitations of 2D vision positioning technology: Traditional 2D vision positioning technology requires that the vision sensor and the target object maintain a visual line of sight. Once the object is outside the line of sight, such as being occluded by other objects, the 2D vision system cannot acquire effective image information, thus failing to achieve positioning. Furthermore, 2D vision has an inherent deficiency in acquiring object depth information, making it difficult to accurately determine the object's position in three-dimensional space. Even under partial visibility conditions, when lighting conditions are poor, such as in a dimly lit corner of a warehouse, the accuracy and reliability of 2D vision positioning are significantly reduced, easily leading to misjudgments or missed detections.
[0040] The shortcomings of integrating relevant positioning technologies: Some techniques that attempt to combine RFID / UWB positioning with 2D visual positioning simply superimpose the results of the two methods or select one according to certain rules, without deeply exploring the potential spatial relationship between them. This results in the inability to fully leverage the complementary advantages of the two technologies in non-line-of-sight scenarios. For example, when 2D visual positioning fails due to occlusion, directly using the limited accuracy of RFID / UWB positioning results cannot achieve high-precision positioning of objects outside of line of sight. Furthermore, the fixed transformation relationships between the positioning devices on the robot are rarely considered, making it difficult to accurately correlate the positioning results with the robot's coordinate system, affecting the robot's actual handling of objects.
[0041] Lack of deep fusion between 3D vision and RFID: While 3D vision technology can acquire three-dimensional information of objects, its integration with RFID technology currently has shortcomings. Most methods do not fully utilize the 3D point cloud information acquired by 3D vision to achieve deep fusion with the wireless information strength and phase of RFID, failing to establish an effective fusion perception error compensation model and thus unable to achieve high-precision positioning prediction in non-line-of-sight scenarios. Based on this, embodiments of this application provide a method for robot object positioning to address at least one of the above problems.
[0042] Figure 1 This is a flowchart illustrating the method for robot positioning of items provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes the following: Step 100: Obtain the wireless radio frequency signal strength information, phase information, and 3D point cloud information of the target item.
[0043] Specifically, radio frequency signal strength information, such as RSSI and phase information, refers to physical layer parameters extracted by radio frequency sensing devices such as RFID or UWB mounted on the robot by receiving signals returned by electronic tags attached to the target item. Together, they provide a rough distance and azimuth clue of the target item relative to the robot, and their signals have a certain degree of penetration, so they can still be perceived when vision is blocked.
[0044] 3D point cloud information refers to a set of data consisting of a large number of three-dimensional coordinate points collected by 3D vision sensors (such as structured light cameras or lidar) on robots. It accurately describes the geometry and spatial distribution of the object's surface within the field of view, thus providing high-precision three-dimensional position and orientation of the target object in the camera coordinate system.
[0045] Step 200: The wireless radio frequency signal strength information, the phase information, and the 3D point cloud information are fused to generate fused perception information.
[0046] Specifically, the fusion of perceived information is not a simple splicing of the above information, but rather refers to a feature representation that can comprehensively reflect the spatial state of the target object after deep feature extraction and association of the aforementioned multi-source heterogeneous data through a specific fusion algorithm, such as a neural network. It simultaneously contains the anti-occlusion characteristics of radio frequency signals and the precise geometric characteristics of visual point clouds.
[0047] Step 300: Based on the fused sensing information, predict the location information of the target item using a fused sensing error compensation model; wherein the fused sensing error compensation model is obtained by training a graph neural network.
[0048] Step 400: Based on the location information of the target item, control the robot to move and navigate to complete the location of the target item.
[0049] Specifically, the fusion perception error compensation model is a graph neural network model trained with a large amount of data. Its function is to learn and model the systematic error relationship between coarse wireless radio frequency positioning and fine 3D vision positioning in complex real environments. In this way, it can predict and compensate for the positioning deviation generated under a single sensing mode or simple fusion method based on the real-time input fusion perception information, and finally output more accurate and reliable positioning information.
[0050] Based on this refined positioning information, the robot's main control system plans a path and drives the chassis and robotic arm to move and navigate, ultimately reaching the target location or completing operations such as grasping objects, thus achieving a complete closed loop from perception to execution.
[0051] It should be noted that the embodiments of this application propose a system for robot positioning of objects. Figure 2 This is a system architecture diagram of the robot positioning of items provided in the embodiments of this application, such as... Figure 2 As shown, this system is a mobile robot platform equipped with wireless radio frequency (RF) devices and a 3D stereo camera. The robot's main control module acts as the processing center. Its built-in RF sensing and 3D vision sensing units process the raw sensor data. The RF-vision fusion positioning unit executes the fusion steps of this embodiment, deeply fusing multi-source information to generate fused perception information, which in turn drives the fusion perception error compensation model to predict accurate positioning information. Subsequently, the mobile navigation module plans a path based on this positioning information, controlling the mobile chassis to move within the robot's activity area, ultimately approaching and locating target items such as item A and item C. When necessary, human-computer interaction and task execution feedback can be achieved through interactive operation modules (such as voice and audio interfaces), thus forming a complete closed loop from multimodal perception and intelligent fusion positioning to autonomous navigation execution.
[0052] The above describes the steps of the robot object positioning method provided in this application embodiment. As can be seen from the above description, the robot object positioning method provided in this application embodiment acquires the radio frequency signal strength information, phase information, and 3D point cloud information of the target object; fuses the radio frequency signal strength information, the phase information, and the 3D point cloud information to generate fused perception information; based on the fused perception information, predicts the positioning information of the target object using a fused perception error compensation model; wherein the fused perception error compensation model is obtained based on graph neural network training; and controls the robot to move and navigate to complete the positioning of the target object based on the positioning information of the target object. Therefore, this application embodiment, by deeply fusing radio frequency signal strength, phase, and 3D point cloud information and using a graph neural network-based fused perception error compensation model for prediction, effectively overcomes the inherent limitations of low positioning accuracy and poor reliability of a single sensor source in non-line-of-sight or complex environments. Even when vision is obstructed or radio frequency signals are interfered with, stable and high-precision prediction of the target object's position can be achieved through model inference, thereby significantly improving the robot's object positioning capability and success rate in complex real-world scenarios.
[0053] Based on the above embodiments, in this embodiment, step 100, acquiring the radio frequency signal strength information, phase information, and 3D point cloud information of the target item, includes: Step 110: Obtain the radio frequency signal strength and phase information of the target item through a wireless radio frequency sensing device.
[0054] Step 110 specifically includes: Step 111: Receive the return signal from the tag and collect the signal strength indication and phase data of the return signal; wherein, the tag is an RFID tag or UWB tag attached to the target item; the tag is used to generate the return signal in response to the radio frequency signal emitted by the wireless radio frequency sensing device at a predetermined frequency.
[0055] Step 112: Filter and denoise the collected signal strength indication and phase data to obtain the wireless radio frequency signal strength information and the phase information.
[0056] Step 120: Obtain the 3D point cloud information of the target object through the robot's 3D vision sensor.
[0057] Step 100 involves acquiring the 3D point cloud information of the target object using the robot's 3D vision sensor, including: Step 130: Acquire RGBD images of the target object using a 3D camera, and convert the RGBD images into initial 3D point cloud data.
[0058] Step 140: Preprocess the initial 3D point cloud data to obtain the 3D point cloud information of the target item.
[0059] Specifically, in the wireless radio frequency sensing link, the wireless radio frequency sensing device carried by the robot, namely the RFID / UWB reader, first periodically transmits radio frequency signals to the surrounding environment at a predetermined frequency (e.g., 100Hz) to activate RFID or UWB tags attached to target items. The returned signal reflected after the tag is activated is received by the device, which then collects the signal strength indication (RSSI) and phase data. To improve data reliability, the large amount of raw signal strength indication and phase data received is filtered and denoised, for example, using median filtering or Kalman filtering algorithms to suppress environmental noise interference, ultimately obtaining stable and reliable wireless radio frequency signal strength and phase information.
[0060] In one embodiment, the robot moves to at least two points in the scene, using RFID / UWB and triangulation to calculate the two-dimensional coordinates of the tags. =arg min ,in The distance from the label to the i-th robot location is calculated using RSSI: (A is the reference signal strength at 1m, and n is the path loss exponent). Alternatively, during the robot's linear motion, the relationship between the phase change rate and the robot's speed can be used, combined with inertial navigation-assisted positioning, to calculate the RFID item's location coordinates based on the collected RSSI and AOA angles of arrival.
[0061] Meanwhile, in the 3D vision perception chain, the robot uses its onboard 3D vision sensors, i.e., 3D cameras (specifically structured light 3D cameras or LiDAR), to acquire RGBD images (images containing color and depth information) of the target object, or directly scan to obtain point clouds. Subsequently, a specific algorithm converts the RGBD images into initial 3D point cloud data. This initial 3D point cloud data undergoes preprocessing, including outlier removal and filtering / smoothing, to improve data quality, eliminate noise points, and retain effective object surface geometry information, ultimately yielding high-quality 3D point cloud information suitable for subsequent fusion. These two streams of preprocessed multimodal perception data lay a solid data foundation for subsequent deep fusion and precise localization.
[0062] In one embodiment, unsupervised clustering is performed on the 3D point cloud after 3D visual preprocessing. RANSAC plane segmentation is used to separate objects from the background. Then, Euclidean clustering or region growing is used to extract the point cloud of individual objects, and the cluster center points of the target objects are extracted. Furthermore, methods such as ICP point cloud registration algorithm, using point cloud neural networks such as PointNet / PointNet++ / PointNeXt, or converting 3D depth maps into 2D images and reusing 2D detection networks (such as SSD-6D, Faster R-CNN) to detect targets, followed by depth calculation, or discretizing 3D space into voxel grids and using 3DCNN to extract features (VoxNet, 3D-RCNN), can directly obtain accurate 3D localization (object center coordinates x, y, z in the camera coordinate system), pose estimation (object rotation angle, such as rotation along the x / y / z axes), and size calculation (length, width, and height of the 3D bounding box) from 3D point clouds, predicting the 3D orientation of objects; and achieving accurate object category recognition based on RGB images.
[0063] The robot object positioning method provided in this application, through a standardized data acquisition and preprocessing process, constitutes a high-quality data source for fused perception, which significantly improves the accuracy, robustness and environmental adaptability of the entire positioning system from the source.
[0064] Based on the above embodiments, in this embodiment, before step 100 obtains the radio frequency signal strength information, phase information, and 3D point cloud information of the target item, the method further includes: Step S110: Determine the fixed transformation relationship between the wireless radio frequency sensing device and the 3D vision sensor; wherein the fixed transformation relationship includes a translation vector and a rotation matrix.
[0065] Step S120: Based on the fixed transformation relationship, the coordinates corresponding to the signal data collected by the wireless radio frequency sensing device are sequentially transformed to the 3D vision sensor coordinate system and the robot base coordinate system.
[0066] Specifically, determining the fixed transformation relationship refers to accurately measuring the relative position and orientation between the wireless radio frequency sensing device and the 3D vision sensor during the robot manufacturing or system integration phase, using high-precision measuring instruments or visual calibration methods. This relationship is fully described mathematically using translation vectors and rotation matrices, which together constitute a complete fixed transformation relationship (TF).
[0067] Based on this, the coordinates corresponding to the signal data collected by the wireless radio frequency sensing device in its own coordinate system are first transformed to the coordinate system of the 3D vision sensor using the above-mentioned fixed transformation relationship, so that the radio frequency positioning results can be associated and compared with the 3D point cloud data in the same visual space. Then, through the known fixed transformation relationship from the 3D vision sensor to the robot base coordinate system in the robot system, the coordinates are finally unified to the robot base coordinate system for the entire robot motion.
[0068] In one embodiment, the process of constructing a unified coordinate system for radio frequency and 3D vision (robot base coordinate system) is described.
[0069] The RFID / UWB antenna and 3D camera are fixed at a specific position on the robot body (such as the torso or robotic arm) to ensure that their relative positions remain unchanged. The translation vector and rotation matrix are obtained by calibrating and measuring the fixed positional relationship between the RFID / UWB antenna and the 3D camera.
[0070] 1) Measurement of the relationship between the position transformation of the wireless radio frequency RFID / UWB antenna and the 3D camera.
[0071] Method 1: Using high-precision measuring instruments, such as a laser tracker, accurately measure the relative position and orientation relationship between the RFID device and the 3D vision sensor, i.e., determine the fixed transformation relationship (TF). During the measurement process, high-precision target spheres or reflectors are installed on the RFID device and the 3D vision sensor, respectively. The three-dimensional coordinates of the target sphere or reflector are measured using a laser tracker, thereby calculating the translation vector t (Tx, Ty, Tz) and rotation matrix (R) between the two sensors, and determining the fixed transformation relationship (TF) between the RFID reader and the 3D vision sensor. To improve measurement accuracy, multiple measurements can be performed and the average value taken.
[0072] Method 2: Calibrate the intrinsic parameters of the 3D camera using Zhang's calibration method, and complete the extrinsic parameter calibration of the camera using AprilTag or checkerboard calibration board to obtain the rotation matrix R and translation vector t, forming the TF transformation matrix.
[0073] 2) Achieve robot base coordinate unification based on TF coordinate transformation algorithm.
[0074] Based on the fixed transformation relation (TF) obtained from measurements, a coordinate transformation algorithm is developed. This algorithm first converts the signal data collected by the RFID device into coordinates in the 3D vision sensor coordinate system using a fusion perception error compensation model. Then, it uses the fixed transformation relation (TF) to transform it into the robot's base coordinate system. Specifically, assuming the object position measured by the RFID device is Prfid in its own coordinate system, the position P3d in the 3D vision sensor coordinate system is obtained through the fusion perception error compensation model. Then, according to the fixed transformation relation (TF), P3d is converted into the position Probot in the robot's base coordinate system, i.e., Probot = TF × P3d. In the actual calculation process, the translation vector and rotation matrix are subjected to corresponding mathematical operations to achieve accurate coordinate transformation.
[0075] The robot object positioning method provided in this application establishes a fixed transformation relationship between the wireless radio frequency device and the 3D vision sensor, and based on this, achieves precise connection of coordinate systems. This fundamentally solves the problem of spatial heterogeneity of data from different sensor modalities, enabling the coarse positioning results of the wireless radio frequency device to be seamlessly associated and deeply fused with the precise point cloud of the 3D vision under a unified robot base coordinate system.
[0076] Based on the above embodiments, in this embodiment, step 200 fuses the wireless radio frequency signal strength information, the phase information, and the 3D point cloud information to generate fused perception information, including: Step 210: Based on the wireless radio frequency signal strength information, the phase information, the 3D point cloud information, and the position information when the robot collects information, construct a spatiotemporal graph structure; wherein, the spatiotemporal graph structure is used to: define the sampling points of the robot at different times as graph nodes, and define the spatial proximity and temporal continuity relationship between nodes as edges of the graph.
[0077] Step 220: Input the spatiotemporal graph structure into a pre-trained graph neural network model and output the fused perception information.
[0078] Specifically, the sampling points at different times during the robot's mobile navigation process—that is, the locations where data is collected when the robot stops or moves continuously—are abstracted as graph nodes in a graph structure. Each node encapsulates local features of the wireless radio frequency signal strength information (such as RSSI value), phase information, and 3D point cloud information collected synchronously at that point, as well as the robot's position information when collecting information. At the same time, the spatial proximity and temporal continuity relationship between nodes are defined. For example, based on the robot's movement trajectory, adjacent sampling points or sampling points within a certain spatiotemporal window are connected and defined as edges in the graph, thereby constructing a dynamic graph model that can reflect both the spatial distribution of data and its temporal evolution.
[0079] Subsequently, this spatiotemporal graph structure, rich in spatiotemporal correlation information, is input into a pre-trained graph neural network model. This model, trained on a large amount of data, has learned how to aggregate and transmit information from complex graph structures and uncover deep, non-linear complementary relationships between radio frequency signals and visual point clouds. Finally, the graph neural network processes and integrates the information from all nodes, outputting fused perception information that comprehensively and accurately describes the spatial state of the target object.
[0080] In one embodiment, when an object moves within the robot's visual field, the high-precision positioning of 3D vision is fused with the coarse positioning of passive sensing to output a baseline true value, providing a basis for subsequent calibration.
[0081] Fusion strategy: Employ weighted Kalman filter (EKF): assign high weights (e.g., 0.8) to 3D visual observations and low weights (e.g., 0.2) to passive sensing, outputting the coordinates of objects in the fused visible area; synchronously record the fusion error (e.g., the residual between visual and passive sensing) to train the passive sensing and visual fusion perception error compensation model (e.g., by training a neural network to fit the nonlinear relationship between RSSI and the real distance), and obtain the optimal distance range between the robot and the object (e.g., 2m-5m) when the error after passive sensing and visual fusion is within the expected threshold range (e.g., <5cm).
[0082] The robot object positioning method provided in this application realizes a leap from simple superposition to intelligent association of multi-source perception information by constructing a spatiotemporal graph structure and using graph neural networks for deep fusion.
[0083] Based on the above embodiments, in this embodiment, the training process of the fusion perception error compensation model includes: Acquire wireless radio frequency signal strength information, phase information, and 3D point cloud information in different scenarios, and align them with the actual position information of objects in different scenarios to construct a training dataset; The model parameters of the graph neural network are iteratively optimized based on the training dataset to minimize the error between the model's predicted position and the actual position of the item, thereby obtaining the fusion perception error compensation model.
[0084] Specifically, by systematically acquiring wireless radio frequency signal strength information, phase information, and 3D point cloud information in different scenarios, such as warehouses with varying layouts and industrial workshops with varying lighting conditions, and using high-precision measurement equipment (such as laser trackers) to simultaneously measure the actual location information of tagged items, the multi-source sensing data is then precisely aligned with the real location data to form a large-scale labeled dataset covering a variety of typical working conditions.
[0085] Subsequently, based on this dataset, iterative learning of the model is initiated: the model parameters of the graph neural network are iteratively optimized based on the training dataset. Specifically, in each iteration, a batch of multi-source sensing data is input into the graph neural network, which calculates the predicted location of the item; then, the error between the model's predicted location and the actual location of the item is calculated through a loss function; finally, the model parameters in the network are automatically adjusted using optimization techniques such as backpropagation algorithm, with the optimization objective being to minimize the aforementioned positioning error.
[0086] Through repeated forward prediction-error calculation-backward optimization cycles, the graph neural network gradually learns to infer the complex, nonlinear mapping relationship between the noisy and biased multi-source sensing data and the real location, ultimately obtaining a well-trained fusion sensing error compensation model with strong generalization capabilities.
[0087] The robot positioning method provided in this application constructs a fusion perception error compensation model through a systematic data-driven approach, providing core intelligent reasoning capabilities for high-precision positioning.
[0088] Based on the above embodiments, in this embodiment, step 400, according to the positioning information of the target item, controls the robot to move and navigate to complete the positioning of the target item, including: Step 410: If the target item is within the range of radio frequency sensing but not within the range of 3D vision, perform coarse positioning based on the positioning information to obtain the candidate location range of the target item.
[0089] Step 420: Generate a spherical search region centered on the candidate location range.
[0090] Step 430: Control the robot to move and navigate to the spherical search area to complete the location of the target item.
[0091] Step 430 specifically includes: Control the robot's movement and navigation until the target object enters the 3D vision sensing range; The current point cloud data of the target object is acquired through the 3D vision sensor, and the coordinates of the target object in the camera coordinate system are calculated based on the current point cloud data using a 3D vision precision localization algorithm. The coordinates are transformed from the camera coordinate system to the robot base coordinate system to complete the positioning of the target object in the robot base coordinate system.
[0092] Step 430 also includes: The robot is controlled to move and navigate to an optimal distance range from the target object; wherein, the optimal distance range is a distance interval in which the positioning error after the fusion of radio frequency sensing and 3D vision is less than a preset threshold. Multiple sets of radio frequency positioning results are obtained at multiple robot locations within the preset distance range; Based on the multiple sets of radio frequency positioning results, the state estimation result is generated by estimating through the fusion sensing error compensation model. The state estimation results are corrected by iterative movement of the robot, and the location of the target item is finally determined.
[0093] This embodiment specifically illustrates how, in non-line-of-sight scenarios, a robot utilizes positioning information and employs two intelligent strategies to ultimately achieve precise positioning of a target object.
[0094] When the system determines that the target object is within the range of radio frequency perception but not within the range of 3D vision (i.e. the object is completely occluded), it first performs coarse localization based on the localization information predicted by the fusion perception error compensation model to obtain a candidate location range with a certain degree of uncertainty. In order to guide the robot search, the system generates a spherical search area centered on the candidate location range, such as a spherical space with a radius of 1 meter. This area defines the robot's initial exploration range.
[0095] Subsequently, the robot executes a differentiated localization strategy based on environmental conditions: if the occlusion is visual (such as an angle problem), the robot is controlled to move and navigate until the target object enters the range of the 3D vision sensor. At this time, the current point cloud data is acquired through the 3D vision sensor, and the 3D vision fine localization algorithm is started. This algorithm first segments the point cloud to separate the object from the background, then extracts the independent point cloud clusters corresponding to the target object, and finally calculates the three-dimensional geometric center coordinates of the target object in the camera coordinate system based on the independent point cloud clusters, and then completes the localization in the robot's base coordinate system through coordinate system transformation.
[0096] If physical obstruction renders vision completely unusable (e.g., an item inside a cabinet), the robot is controlled to move and navigate to the optimal distance range from the target item. This optimal distance range, determined through numerous experiments, is a specific distance interval that minimizes the fusion positioning error, such as 2-5 meters. Multiple sets of radio frequency positioning results are acquired at various robot positions within this range. These results are then used as sequential observations and input into the fusion perception error compensation model for data fusion and state estimation, generating a more accurate state estimation result. This result is then iteratively corrected through robot movement, ultimately determining the precise location of the target item directly without visual intervention.
[0097] The robot object localization method provided in this application, by designing two adaptive localization strategies for non-line-of-sight scenarios, enables the robot system to effectively cope with various non-line-of-sight challenges, significantly improving the robustness, success rate and efficiency of finding and locating targets in real complex scenarios.
[0098] Based on the above embodiments, in this embodiment, step 400, which controls the robot to move and navigate according to the positioning information of the target item to complete the positioning of the target item, further includes: Step 440: When the target item is within the range of radio frequency sensing and within the range of 3D vision, simultaneously acquire the precise positioning result based on 3D vision and the coarse positioning result based on radio frequency sensing.
[0099] Step 450: The precise positioning result and the coarse positioning result are fused to generate fused coordinates in the robot base coordinate system to complete the positioning of the target object at the current moment.
[0100] This embodiment provides an efficient fusion positioning strategy executed by the system when the target object is simultaneously detectable by both radio frequency and 3D vision.
[0101] When the target object is within the range of radio frequency sensing and within the range of 3D vision, the system synchronously acquires the precise positioning result based on 3D vision in parallel. This means that the 6D pose of the object can reach the centimeter or even millimeter level with the precision obtained by the 3D vision fine positioning algorithm, and the coarse positioning result based on radio frequency sensing, which means that the approximate position of the object may have a few decimeter errors, mainly calculated based on signal parameters such as RSSI or phase.
[0102] Subsequently, the two types of positioning results are fused: specifically, a filtering algorithm, such as weighted Kalman filtering, is used to assign higher weights to the precise positioning results of vision, while the coarse positioning results of radio frequency are used as an auxiliary reference and supplement. An optimal fused coordinate is estimated through algorithm optimization.
[0103] The robot object positioning method provided in this application generates fused coordinates in the robot's base coordinate system. It not only inherits the high-precision characteristics of visual positioning, but also enhances stability and robustness under short-term visual jitter or slight interference due to the introduction of radio frequency information. This enables the system to output a more reliable and accurate final positioning at the current moment, providing a direct basis for the robot's immediate grasping or interactive operation, and achieving the optimization of positioning performance in visual scenarios.
[0104] Figure 3 This is a flowchart illustrating the non-visual / visual object positioning method provided in the embodiments of this application. The following is a summary of the process. Figure 3 The present application provides a detailed description of the non-visual / visual object positioning method provided in the embodiments.
[0105] In the non-visual area object prediction and positioning process, the system performs radio frequency (RFID / UWB) data acquisition and preprocessing, and then achieves initial coarse positioning by using RFID identification of the object and a mobile-based tag positioning algorithm; at the same time, in the visible area process, the system performs 3D visual data acquisition and preprocessing, and completes 3D visual object recognition and precise positioning.
[0106] All of this perception data is based on a spatial reference that unifies wireless radio frequency devices and 3D cameras into the robot's base coordinate system.
[0107] Within the visible area, the system generates high-precision positioning results through a wireless radio frequency (RF) + 3D vision fusion positioning module and simultaneously trains an error compensation model to achieve high-precision positioning of objects within the visible area. For non-visual areas, the system innovatively introduces a spatial relationship knowledge graph to represent prior environmental layouts. Through a multiplication correction module for spatial relationships in non-visual areas, the RF positioning results are fused with spatial knowledge for reasoning. Finally, the optimized predicted position is output through a non-visual area positioning prediction and calibration module.
[0108] The robot object positioning method provided in this application has a cross-regional collaborative correction mechanism formed by the error compensation model and spatial relationship knowledge, which achieves full scene coverage from the visible area to the non-visual area and from coarse positioning to accurate prediction.
[0109] The following describes the robot positioning device provided in the embodiments of this application. The robot positioning device described below and the robot positioning method described above can be referred to in correspondence.
[0110] Figure 4 This is a schematic diagram of the structure of the robot positioning device provided in the embodiments of this application, as shown below. Figure 4 As shown in the embodiment of this application, the robot positioning device for locating items includes: The acquisition module 401 is used to acquire the wireless radio frequency signal strength information, phase information, and 3D point cloud information of the target item; The fusion module 402 is used to fuse the wireless radio frequency signal strength information, the phase information and the 3D point cloud information to generate fused perception information; The prediction module 403 is used to predict the location information of the target item based on the fused sensing information and through a fused sensing error compensation model; wherein the fused sensing error compensation model is obtained based on graph neural network training; The positioning module 404 is used to control the robot to move and navigate to complete the positioning of the target item based on the positioning information of the target item.
[0111] The robot object positioning device provided in this application acquires the radio frequency signal strength information, phase information, and 3D point cloud information of the target object; fuses the radio frequency signal strength information, phase information, and 3D point cloud information to generate fused perception information; based on the fused perception information, it predicts the positioning information of the target object through a fused perception error compensation model; wherein the fused perception error compensation model is trained based on a graph neural network; and controls the robot to move and navigate according to the positioning information of the target object to complete the positioning of the target object. Therefore, this application embodiment, by deeply fusing radio frequency signal strength, phase, and 3D point cloud information and using a graph neural network-based fused perception error compensation model for prediction, effectively overcomes the inherent limitations of low positioning accuracy and poor reliability of a single sensor source in non-line-of-sight or complex environments. Even when vision is obstructed or radio frequency signals are interfered with, it can achieve stable and high-precision prediction of the target object's position through model inference, thereby significantly improving the robot's object positioning capability and success rate in complex real-world scenarios.
[0112] Based on the above embodiments, in this embodiment, the acquisition module 401 is specifically used for: The radio frequency signal strength and phase information of the target item are obtained through a wireless radio frequency sensing device; The robot acquires 3D point cloud information of the target object using its 3D vision sensor.
[0113] Based on the above embodiments, in this embodiment, the device further includes a conversion module, specifically used for: Before acquiring the wireless radio frequency signal strength information, phase information, and 3D point cloud information of the target item, A fixed transformation relationship is determined between the wireless radio frequency sensing device and the 3D vision sensor; wherein the fixed transformation relationship includes a translation vector and a rotation matrix; Based on the fixed transformation relationship, the coordinates corresponding to the signal data collected by the wireless radio frequency sensing device are sequentially transformed to the 3D vision sensor coordinate system and the robot base coordinate system.
[0114] Based on the above embodiments, in this embodiment, the acquisition module 401 is further configured to: The system receives a return signal from a tag and collects the signal strength indication and phase data of the return signal; wherein the tag is an RFID tag or a UWB tag attached to the target item; the tag is used to generate the return signal in response to the radio frequency signal emitted by the wireless radio frequency sensing device at a predetermined frequency; The acquired signal strength indication and phase data are filtered and denoised to obtain the wireless radio frequency signal strength information and the phase information.
[0115] Based on the above embodiments, in this embodiment, the acquisition module 401 is further configured to: The target object is captured using a 3D camera in RGBD image format, and the RGBD image is converted into initial 3D point cloud data. The initial 3D point cloud data is preprocessed to obtain the 3D point cloud information of the target item.
[0116] Based on the above embodiments, in this embodiment, the fusion module 402 is specifically used for: Based on the wireless radio frequency signal strength information, the phase information, the 3D point cloud information, and the position information when the robot collects information, a spatiotemporal graph structure is constructed; wherein, the spatiotemporal graph structure is used to: define the sampling points of the robot at different times as graph nodes, and define the spatial proximity and temporal continuity relationship between nodes as edges of the graph; The spatiotemporal graph structure is input into a pre-trained graph neural network model, which outputs the fused sensing information.
[0117] Based on the above embodiments, in this embodiment, the device further includes a training module, specifically used for: Acquire wireless radio frequency signal strength information, phase information, and 3D point cloud information in different scenarios, and align them with the actual position information of objects in different scenarios to construct a training dataset; The model parameters of the graph neural network are iteratively optimized based on the training dataset to minimize the error between the model's predicted position and the actual position of the item, thereby obtaining the fusion perception error compensation model.
[0118] Based on the above embodiments, in this embodiment, the positioning module 404 is specifically used for: If the target item is within the range of radio frequency sensing but not within the range of 3D vision, coarse positioning is performed based on the positioning information to obtain the candidate location range of the target item. Generate a spherical search region centered on the range of candidate locations; The robot is controlled to move and navigate to the spherical search area to locate the target item.
[0119] Based on the above embodiments, in this embodiment, the device further includes a control module, specifically used for: Control the robot's movement and navigation until the target object enters the 3D vision sensing range; The current point cloud data of the target object is acquired through the 3D vision sensor, and the coordinates of the target object in the camera coordinate system are calculated based on the current point cloud data using a 3D vision precision localization algorithm. The coordinates are transformed from the camera coordinate system to the robot base coordinate system to complete the positioning of the target object in the robot base coordinate system.
[0120] Based on the above embodiments, in this embodiment, the device further includes a computing module, specifically used for: The current point cloud data is segmented to obtain segmented point cloud data; From the segmented point cloud data, extract the independent point cloud clusters corresponding to the target item; Based on the independent point cloud cluster, the three-dimensional geometric center coordinates of the target object in the camera coordinate system are calculated.
[0121] Based on the above embodiments, in this embodiment, the control module is further specifically used for: The robot is controlled to move and navigate to an optimal distance range from the target object; wherein, the optimal distance range is a distance interval in which the positioning error after the fusion of radio frequency sensing and 3D vision is less than a preset threshold. Multiple sets of radio frequency positioning results are obtained at multiple robot locations within the preset distance range; Based on the multiple sets of radio frequency positioning results, the state estimation result is generated by estimating through the fusion sensing error compensation model. The state estimation results are corrected by iterative movement of the robot, and the location of the target item is finally determined.
[0122] Based on the above embodiments, in this embodiment, the positioning module 404 is further configured to: When the target item is within the range of radio frequency sensing and within the range of 3D vision, the precise positioning result based on 3D vision and the coarse positioning result based on radio frequency sensing are acquired simultaneously. The precise positioning result and the coarse positioning result are fused to generate fused coordinates in the robot's base coordinate system, thereby completing the positioning of the target object at the current moment.
[0123] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device can be a robot or other electronic device. This electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 550. The processor 510, communications interface 520, and memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions from the memory 530 to execute a method for the robot to locate objects, including: Acquire the target object's radio frequency signal strength, phase information, and 3D point cloud information; The wireless radio frequency signal strength information, the phase information, and the 3D point cloud information are fused to generate fused perception information. Based on the fused sensing information, the location information of the target item is predicted by a fused sensing error compensation model; wherein, the fused sensing error compensation model is obtained by training a graph neural network; Based on the location information of the target item, the robot is controlled to move and navigate in order to locate the target item.
[0124] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application embodiment, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in at least one embodiment of this application embodiment. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0125] On the other hand, embodiments of this application also provide a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the robot positioning method for objects provided by the above methods, including: Acquire the target object's radio frequency signal strength, phase information, and 3D point cloud information; The wireless radio frequency signal strength information, the phase information, and the 3D point cloud information are fused to generate fused perception information. Based on the fused sensing information, the location information of the target item is predicted by a fused sensing error compensation model; wherein, the fused sensing error compensation model is obtained by training a graph neural network; Based on the location information of the target item, the robot is controlled to move and navigate in order to locate the target item.
[0126] In another aspect, embodiments of this application also provide a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a method for robot positioning of an item provided by the methods described above, including: Acquire the target object's radio frequency signal strength, phase information, and 3D point cloud information; The wireless radio frequency signal strength information, the phase information, and the 3D point cloud information are fused to generate fused perception information. Based on the fused sensing information, the location information of the target item is predicted by a fused sensing error compensation model; wherein, the fused sensing error compensation model is obtained by training a graph neural network; Based on the location information of the target item, the robot is controlled to move and navigate in order to locate the target item.
[0127] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0128] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of this application, and are not intended to limit them; although the embodiments of this application have been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for locating objects using a robot, characterized in that, include: Acquire the target object's radio frequency signal strength, phase information, and 3D point cloud information; The wireless radio frequency signal strength information, the phase information, and the 3D point cloud information are fused to generate fused perception information. Based on the fused sensing information, the location information of the target item is predicted by a fused sensing error compensation model; wherein, the fused sensing error compensation model is obtained by training a graph neural network; Based on the location information of the target item, the robot is controlled to move and navigate in order to locate the target item.
2. The method for robot positioning of items according to claim 1, characterized in that, The acquisition of the target item's radio frequency signal strength information, phase information, and 3D point cloud information includes: The radio frequency signal strength and phase information of the target item are obtained through a wireless radio frequency sensing device; The robot acquires 3D point cloud information of the target object using its 3D vision sensor.
3. The method for robot positioning of items according to claim 2, characterized in that, Before acquiring the radio frequency signal strength information, phase information, and 3D point cloud information of the target item, the method further includes: A fixed transformation relationship is determined between the wireless radio frequency sensing device and the 3D vision sensor; wherein the fixed transformation relationship includes a translation vector and a rotation matrix; Based on the fixed transformation relationship, the coordinates corresponding to the signal data collected by the wireless radio frequency sensing device are sequentially transformed to the 3D vision sensor coordinate system and the robot base coordinate system.
4. The method for robot positioning of items according to claim 2, characterized in that, The step of acquiring the radio frequency signal strength and phase information of the target item through a radio frequency sensing device includes: The system receives a return signal from a tag and collects the signal strength indication and phase data of the return signal; wherein the tag is an RFID tag or a UWB tag attached to the target item; the tag is used to generate the return signal in response to the radio frequency signal emitted by the wireless radio frequency sensing device at a predetermined frequency; The acquired signal strength indication and phase data are filtered and denoised to obtain the wireless radio frequency signal strength information and the phase information.
5. The method for robot positioning of an item according to claim 2 or 3, characterized in that, The acquisition of 3D point cloud information of the target object through the robot's 3D vision sensor includes: The target object is captured using a 3D camera in RGBD image format, and the RGBD image is converted into initial 3D point cloud data. The initial 3D point cloud data is preprocessed to obtain the 3D point cloud information of the target item.
6. The method for robot positioning of items according to claim 1, characterized in that, The step of fusing the wireless radio frequency signal strength information, the phase information, and the 3D point cloud information to generate fused perception information includes: Based on the wireless radio frequency signal strength information, the phase information, the 3D point cloud information, and the position information when the robot collects information, a spatiotemporal graph structure is constructed; wherein, the spatiotemporal graph structure is used to: define the sampling points of the robot at different times as graph nodes, and define the spatial proximity and temporal continuity relationship between nodes as edges of the graph; The spatiotemporal graph structure is input into a pre-trained graph neural network model, which outputs the fused sensing information.
7. The method for robot positioning of items according to claim 1, characterized in that, The training process of the fusion-sensing error compensation model includes: Acquire wireless radio frequency signal strength information, phase information, and 3D point cloud information in different scenarios, and align them with the actual position information of objects in different scenarios to construct a training dataset; The model parameters of the graph neural network are iteratively optimized based on the training dataset to minimize the error between the model's predicted position and the actual position of the item, thereby obtaining the fusion perception error compensation model.
8. The method for robot positioning of items according to claim 1, characterized in that, The step of controlling the robot to move and navigate according to the location information of the target item to complete the location of the target item includes: If the target item is within the range of radio frequency sensing but not within the range of 3D vision, coarse positioning is performed based on the positioning information to obtain the candidate location range of the target item. Generate a spherical search region centered on the range of candidate locations; The robot is controlled to move and navigate to the spherical search area to locate the target item.
9. The method for robot positioning of items according to claim 8, characterized in that, The controlled robot moves and navigates to the spherical search area to locate the target item, including: Control the robot's movement and navigation until the target object enters the 3D vision sensing range; The current point cloud data of the target object is acquired through the 3D vision sensor, and the coordinates of the target object in the camera coordinate system are calculated based on the current point cloud data using a 3D vision precision localization algorithm. The coordinates are transformed from the camera coordinate system to the robot base coordinate system to complete the positioning of the target object in the robot base coordinate system.
10. The method for robot positioning of an item according to claim 9, characterized in that, The step of calculating the coordinates of the target object in the camera coordinate system based on the current point cloud data using a 3D visual precision localization algorithm includes: The current point cloud data is segmented to obtain segmented point cloud data; From the segmented point cloud data, extract the independent point cloud clusters corresponding to the target item; Based on the independent point cloud cluster, the three-dimensional geometric center coordinates of the target object in the camera coordinate system are calculated.
11. The method for robot positioning of an item according to claim 8, characterized in that, The controlled robot moves and navigates to the spherical search area to locate the target item, including: The robot is controlled to move and navigate to an optimal distance range from the target object; wherein, the optimal distance range is a distance interval in which the positioning error after the fusion of radio frequency sensing and 3D vision is less than a preset threshold. Multiple sets of radio frequency positioning results are obtained at multiple robot locations within the preset distance range; Based on the multiple sets of radio frequency positioning results, the state estimation result is generated by estimating through the fusion sensing error compensation model. The state estimation results are corrected by iterative movement of the robot, and the location of the target item is finally determined.
12. The method for robot positioning of an item according to claim 1, characterized in that, The step of controlling the robot to move and navigate based on the location information of the target item to complete the location of the target item also includes: When the target item is within the range of radio frequency sensing and within the range of 3D vision, the precise positioning result based on 3D vision and the coarse positioning result based on radio frequency sensing are acquired simultaneously. The precise positioning result and the coarse positioning result are fused to generate fused coordinates in the robot's base coordinate system, thereby completing the positioning of the target object at the current moment.
13. A device for locating objects by a robot, characterized in that, include: The acquisition module is used to acquire the wireless radio frequency signal strength information, phase information, and 3D point cloud information of the target item; The fusion module is used to fuse the wireless radio frequency signal strength information, the phase information, and the 3D point cloud information to generate fused perception information; The prediction module is used to predict the location information of the target item based on the fused sensing information and through a fused sensing error compensation model; wherein the fused sensing error compensation model is obtained based on graph neural network training; The positioning module is used to control the robot's movement and navigation to locate the target item based on the target item's positioning information.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for robot positioning of an article as described in any one of claims 1 to 12.
15. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for robot positioning of an article as described in any one of claims 1 to 12.
16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for locating an article by a robot according to any one of claims 1 to 12.