An intelligent carrying robot with pose recognition function for fabricated building
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 中一达建设集团有限公司
- Filing Date
- 2026-05-06
- Publication Date
- 2026-08-04
AI Technical Summary
传统遥控设备(如手持遥控器、操纵杆)仅能发送运动指令,操作员无法直观获知机器人的当前位姿、目标位姿以及构件与安装位置之间的实时偏差信息
Smart Images

Figure CN122500757A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot control technology, specifically to an intelligent handling robot with pose recognition function for prefabricated buildings. Background Technology
[0002] Prefabricated construction refers to a construction method in which prefabricated components are manufactured in a factory and then transported to the construction site for assembly. During construction, prefabricated components (such as wall panels, beams, columns, composite slabs, and stair sections) need to be handled and precisely positioned for installation on-site. Intelligent handling robots with posture recognition capabilities can automatically identify the position and posture of components, enabling autonomous handling of prefabricated components. They are key equipment for improving the efficiency and quality of prefabricated building construction.
[0003] Currently, handling robots used in prefabricated building construction sites mainly rely on technologies such as laser navigation, visual SLAM, and UWB positioning to achieve autonomous path planning and handling. However, in actual construction environments, there are complex situations such as numerous dynamic obstacles, limited space, and high precision requirements for component installation. A fully autonomous handling mode cannot cope with all unexpected scenarios. For example, when the handling robot encounters unforeseen obstacles, needs to precisely place components into narrow installation positions with pre-embedded sleeves, or needs to coordinate with on-site workers, manual intervention is often required for remote control or fine-tuning.
[0004] Existing methods of manual intervention mainly suffer from the following technical drawbacks: Traditional remote control devices (such as handheld remote controllers and joysticks) can only send motion commands. Operators cannot intuitively obtain information about the robot's current pose, target pose, and real-time deviation between the component and its installation position. Operators need to rely on visual observation and experience to operate the robot, resulting in low precision and long processing time for manual fine-tuning, which makes it difficult to meet the millimeter-level precision requirements for the placement of prefabricated components in prefabricated buildings.
[0005] In conclusion, there is an urgent need for an intelligent handling robot that can deeply integrate pose recognition with augmented reality interaction to improve the flexibility, accuracy, and human-machine collaboration efficiency of handling operations at prefabricated building construction sites. Summary of the Invention
[0006] The purpose of this invention is to provide an intelligent handling robot with pose recognition function for prefabricated buildings, so as to solve the problems raised in the prior art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: an intelligent handling robot with pose recognition function for prefabricated buildings, comprising a handling robot body, wherein the handling robot body is provided with a pose recognition module, a panoramic camera, a path planning module and a manual fine-tuning control module; The transport robot body is connected to the augmented reality interaction unit via a wireless communication link, and the augmented reality interaction unit includes AR glasses; The AR glasses use a combination of ultra-wideband positioning and visual markers to align the spatial pose of the transport robot and overlay an augmented reality interface in the operator's field of vision. The augmented reality interface includes: the planned path generated by the path planning module, the current pose and target pose of the transport robot body obtained by the pose recognition module, and the distance information between the prefabricated components and obstacles superimposed in the construction site image captured by the panoramic camera. The manual fine-tuning control module responds to the gesture commands recognized by the AR glasses or the operation of the handle paired with the AR glasses, generates a six-degree-of-freedom control signal, and drives the end effector of the handling robot body to perform fine-tuning of position and attitude. During manual fine-tuning, the augmented reality interface displays the deviation between the end effector and the target pose in real time and provides prompts through the augmented reality interface. The transport robot automatically performs coarse positioning and transport operations based on the target position and orientation specified by the operator in the augmented reality interface. After coarse positioning is completed, it switches to manual fine-tuning mode, and the operator completes fine alignment through the augmented reality interaction unit.
[0008] Furthermore, the spatial pose alignment includes: the AR glasses measuring distances with UWB anchor points arranged on the transport robot body through their built-in UWB tags, recognizing preset visual markers on the transport robot body using the downward-facing camera on the AR glasses, calculating the six-degree-of-freedom pose of the AR glasses relative to the transport robot body, and aligning the coordinate system of the augmented reality content with the coordinate system of the transport robot body.
[0009] Furthermore, the above pose alignment method has the following drawbacks: when the visual markers of multiple robots appear simultaneously within the field of view of the same AR glasses, or when UWB tags simultaneously measure distances with multiple anchor points, it is impossible to uniquely determine which robot the AR glasses are currently aligning with, resulting in virtual content being overlaid on the wrong robot; to solve this problem, a better implementation method is as follows: Step S1: Each UWB anchor point on the robot body broadcasts a ranging request message according to a predetermined time slot. The message contains the robot's unique identification ID. At the same time, on the same clock edge when the broadcast message is sent, the synchronous trigger controller drives the active dynamic visual marker array to flash once according to a preset encoding rule. The flashing pattern uniquely corresponds to the robot's ID.
[0010] Step S2: After the UWB tag on the AR glasses receives a ranging request message from a certain UWB anchor point: record the received signal strength and arrival time, and calculate the approximate distance between the tag and the corresponding robot; at the same time, the UWB tag sends a hardware synchronization pulse to the downward-facing camera of the AR glasses, forcing the downward-facing camera to start exposure within the microsecond-level time window at the moment of receiving the UWB message, so as to capture the flashing pattern of the LED array on the robot at this time.
[0011] Step S3: The image captured by the downward-facing camera is transmitted to the decoding processing unit; the decoding processing unit identifies the location and combination of LED bright spots in the image, and calculates the robot ID represented by the LED array according to the preset encoding rules.
[0012] Step S4: Compare the ID in the UWB message from step S1 with the ID obtained from decoding in step S3: If they match, it is confirmed that the current UWB ranging value and the current image frame belong to the same target robot; If there is an inconsistency or decoding failure, the frame data is discarded and the process waits for the next time slot.
[0013] Step S5: After identity verification, the decoding processing unit retrieves the robot's pre-stored 3D model, the spatial geometric layout of the LED array, and the robot's body coordinate system definition from the local database based on the robot ID. Then, using the image coordinates of at least three LED points in the downward-facing camera and the known world coordinates (in the robot's coordinate system), the six-degree-of-freedom pose of the AR glasses relative to the robot body is calculated using the PnP (Perspective-n-Point) algorithm.
[0014] Step S6: Based on the calculated six-degree-of-freedom pose, transform the augmented reality content to be displayed in the AR glasses from the world coordinate system or map coordinate system to the robot's coordinate system, and correctly overlay and display it in the AR glasses' field of view.
[0015] When multiple robots are present in the environment, Time Division Multiple Access (TDMA) or Frequency Hopping Service (FHSS) is used to coordinate UWB ranging requests: each robot's UWB anchor point sends ranging requests and LED flashing in its unique time slot, with guard intervals between time slots to avoid ranging request conflicts between different robots and overlap of LED optical signals. The UWB tag of the AR glasses receives messages from all time slots and selects the robot related to the current task for alignment based on the decoded ID.
[0016] Furthermore, the augmented reality interface is also configured to overlay and display the actual pose contour of the prefabricated component identified by the pose recognition module, and to mark the positions of preset lifting points or embedded sleeves on the prefabricated component with a highlight color, so as to assist the operator in judging the precise position of clamping or placing.
[0017] Furthermore, the manual fine-tuning control module includes a gesture recognition unit, which collects the operator's hand skeletal point movements through the depth camera on the AR glasses and maps predefined gestures to motion commands of the end effector. The predefined gestures include finger sliding corresponding to translational movement, wrist rotation corresponding to rotational movement, and fist clenching corresponding to clamping action.
[0018] The process of switching to manual fine-tuning mode is as follows: the handling robot body automatically plans a coarse positioning path and performs handling based on the target area selected by the operator in the augmented reality interface and the current pose of the prefabricated component output by the pose recognition module. When the handling robot body determines that the deviation between its own pose and the target pose enters the preset threshold range, it automatically sends a switching request to the AR glasses. The AR glasses emit vibration or voice prompts to notify the operator to take over manual fine-tuning.
[0019] Furthermore, the deviation includes translational deviation components (Δx, Δy, Δz) and rotational deviation components (Δroll, Δpitch, Δyaw) between the end effector's current pose and target pose. The augmented reality interface displays each deviation component in the form of digital labels next to the virtual model of the end effector and dynamically changes the background color of the digital labels according to the magnitude of the absolute value of the deviation.
[0020] Furthermore, it also includes a remote collaboration unit: the AR glasses transmit the augmented reality image in the operator's field of vision to a remote expert terminal in real time. The remote expert terminal receives the annotation information input by the expert and overlays the annotation information in the field of vision of the AR glasses. The annotation information includes at least one of virtual arrows, circled areas, and text descriptions.
[0021] Furthermore, the augmented reality interface is configured to display the historical travel path and future predicted path of the transport robot in the form of trajectory lines during the automatic transport process, and to distinguish between the planned path and the dynamically replanned path with different colors.
[0022] Furthermore, the manual fine-tuning control module includes a speed mapping curve that non-linearly maps the offset of the joystick to the motion speed of the end effector.
[0023] Furthermore, the AR glasses are equipped with an ambient light sensor, and the display brightness of the augmented reality interface is automatically adjusted according to the ambient light intensity at the construction site. The adjustment of the display brightness is linked to the exposure parameters of the panoramic camera, so that the superimposed virtual information is consistent with the brightness of the real scene.
[0024] Compared with the prior art, the beneficial effects of the present invention are: 1. By using a positioning method that combines UWB with visual markers, high-precision spatial pose alignment between AR glasses and handling robots is achieved, enabling operators to gain "see-through" perception capabilities through the AR interface.
[0025] 2. By combining automated coarse positioning with AR-based manual fine-tuning, the robot's efficiency advantage in large-scale path planning is leveraged, while the operator's precise judgment and operational flexibility in complex environments are utilized, thus meeting the high-precision installation requirements of prefabricated buildings.
[0026] 3. The AR interface intuitively displays key data such as planned path, positional deviation, and component information, and supports natural interaction with gestures / handheld devices, which greatly reduces the learning cost and cognitive burden of operators and improves the safety and efficiency of handling operations. Attached Figure Description
[0027] Figure 1 This is a schematic diagram of the system framework of an intelligent handling robot with pose recognition function for prefabricated buildings according to the present invention. Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] Example 1:
[0030] Please see Figure 1 This embodiment provides an intelligent handling robot with pose recognition function for prefabricated buildings. Its core components include: the handling robot body (e.g., an AGV chassis with autonomous navigation capabilities, on which a multi-degree-of-freedom robotic arm is installed as an end effector), a pose recognition module fixed to the body (which can adopt a SLAM system that integrates binocular vision and inertial measurement unit (IMU), a panoramic camera (multiple fisheye cameras stitched together to achieve 360° surround view), a path planning module (with built-in navigation algorithm), and a manual fine-tuning control module (for processing fine control commands).
[0031] The robot itself connects to the augmented reality interaction unit via a 5G or Wi-Fi 6 wireless communication link. At the heart of the augmented reality interaction unit is an AR glasses (such as a Microsoft HoloLens 2 or a similar device) worn by the operator.
[0032] The typical workflow of this embodiment is as follows: First, the operator wears AR glasses in a safe area of the construction site. Once activated, the AR glasses automatically establish a wireless connection with the transport robot and perform spatial pose alignment (see Example 2). After alignment, an augmented reality interface is overlaid in the operator's field of vision, integrating real construction site footage with virtual information. The operator uses gestures to select the target placement area (e.g., specifying the installation location of a precast floor slab) within the virtual interface. Upon receiving the target pose, the transport robot's path planning module automatically plans a coarse positioning path from its current position to the target location and automatically performs the transport. When the robot determines it has reached the vicinity of the target location (e.g., the Euclidean distance between the end effector and the target pose is less than 10cm, and the angular deviation is less than 5°), it automatically sends a switching request to the AR glasses. The AR glasses vibrate and provide a voice prompt, "Please perform manual fine-tuning." The operator then enters manual fine-tuning mode via gestures or a handle, finely adjusting the six-degree-of-freedom pose of the end effector until the deviation is zero, completing the final placement. Through this process, efficient human-machine collaboration is achieved through "automated coarse handling + manual fine alignment."
[0033] Example 2: To achieve precise alignment of the coordinate systems of the AR glasses and the transport robot, this embodiment employs a method of fusing UWB with visual markers.
[0034] Specifically, at least four UWB anchor points (with known precise coordinates in the body coordinate system) are arranged at the four top corners of the handling robot body and at the base of the robotic arm. A UWB tag is integrated into the AR glasses. The AR glasses obtain a rough position (accurate to approximately ±10cm) relative to the handling robot body by measuring the distance signals between themselves and each UWB anchor point using trilateration.
[0035] To further improve accuracy, a downward-looking camera (not shown) is also installed below the frame of the AR glasses. The surface of the transport robot (especially the top plane) is covered with visual markers such as QR codes or April Tags. The downward-looking camera captures images of these visual markers in real time, and the PnP (Perspective-n-Point) algorithm is used to calculate the precise six-DOF pose of the AR glasses relative to the markers (with an accuracy down to the millimeter level).
[0036] Using the coarse position provided by UWB as the predicted value of the Kalman filter and the precise pose calculated by vision as the observed value, the final stable and high-precision pose of the AR glasses relative to the transport robot is obtained by fusing the two. This achieves precise binding between the virtual world coordinate system of the AR glasses and the real world coordinate system of the robot, ensuring that all augmented reality information (such as path lines, deviation indicators, etc.) can be accurately superimposed on the correct position of the real object.
[0037] Example 3: Unlike Example 2, this example provides a solution for quickly identifying the single target robot when multiple transport robots appear within the field of view of the same AR glasses. The pose alignment method in this example includes the following steps: Step S1: Each UWB anchor point on the robot body broadcasts a ranging request message according to a predetermined time slot (e.g., using TDMA to avoid conflicts between multiple robots). The message contains the robot's unique identification ID. At the same time, on the same clock edge (or within a fixed time offset ΔT) when the broadcast message is sent, the synchronization trigger controller drives the active dynamic visual marker array to flash once according to a preset encoding rule. The flashing pattern uniquely corresponds to the robot's ID.
[0038] Step S2: After the UWB tag on the AR glasses receives a ranging request message from a certain UWB anchor point: record the received signal strength and arrival time, and calculate the approximate distance between the tag and the corresponding robot; at the same time, the UWB tag sends a hardware synchronization pulse to the downward-facing camera of the AR glasses, forcing the downward-facing camera to start exposure within the microsecond-level time window at the moment of receiving the UWB message, so as to capture the flashing pattern of the LED array on the robot at this time.
[0039] Step S3: The image captured by the downward-facing camera is transmitted to the decoding processing unit; the decoding processing unit identifies the position and combination of LED bright spots in the image, and calculates the robot ID represented by the LED array according to the preset encoding rules (such as Manchester encoding, time frame encoding or spatial position encoding).
[0040] Step S4: Compare the ID in the UWB message from step S1 with the ID obtained from decoding in step S3: If they match, it is confirmed that the current UWB ranging value and the current image frame belong to the same target robot; If there is an inconsistency or decoding failure, the frame data is discarded and the process waits for the next time slot.
[0041] Step S5: After identity verification, the decoding processing unit retrieves the robot's pre-stored 3D model, the spatial geometric layout of the LED array, and the robot's coordinate system definition from the local database (or cloud) based on the robot ID. Then, using the image coordinates of at least three LED points in the downward-facing camera and the known world coordinates (in the robot's coordinate system), the six-degree-of-freedom pose (position and orientation) of the AR glasses relative to the robot body is calculated using the PnP (Perspective-n-Point) algorithm.
[0042] Step S6: Based on the calculated six-degree-of-freedom pose, transform the augmented reality content to be displayed in the AR glasses (such as navigation arrows, safety warning boxes, and gripping point indicators) from the world coordinate system or map coordinate system to the robot's coordinate system, and correctly overlay and display it in the AR glasses' field of view.
[0043] When multiple robots are present in the environment, Time Division Multiple Access (TDMA) or Frequency Hopping Service (FHSS) is used to coordinate UWB ranging requests: each robot's UWB anchor point sends ranging requests and LED flashing in its unique time slot, with guard intervals between time slots to avoid ranging request conflicts between different robots and overlap of LED optical signals. The UWB tag of the AR glasses receives messages from all time slots and selects the robot related to the current task for alignment based on the decoded ID.
[0044] 1. Achieved unique identity binding: By verifying the synchronization consistency between the ID in the UWB message and the LED blinking code, the problem of AR glasses misidentifying other robots in multi-robot scenarios is fundamentally solved, ensuring that virtual content is only superimposed on the target robot.
[0045] 2. Enhanced anti-interference capability: The TDMA / FHSS mechanism avoids ranging conflicts between multiple UWB nodes, while the LED blinking uses short-time high-brightness encoding, which can still decode stably under ambient light interference.
[0046] 3. Precise time synchronization: The UWB ranging pulse directly triggers the camera exposure, eliminating the dynamic error caused by the asynchrony between image acquisition and ranging in traditional solutions, and improving the dynamic alignment accuracy.
[0047] Fast solution: Utilizing the pre-stored robot prior model and the known geometric layout of the LED array, PnP solution requires fewer points (minimum 3 points), has a small computational load, and can run in real time on the AR glasses embedded platform.
[0048] 4. Scalability: The spatial layout of the LED array can be flexibly designed (such as triangles, rectangles, and asymmetrical patterns). Different robots can use different layouts as auxiliary recognition features to further reduce the probability of misidentification.
[0049] Example 4: This embodiment describes in detail the display content of the augmented reality interface. When the operator looks at the transport robot and its surrounding environment, the following information will be overlaid in the AR glasses' field of view: Path planning: A green 3D trajectory line extends from the robot's current position to the target placement area, representing the coarse positioning path generated by the path planning module.
[0050] Pose information: A status panel floats above the robot's virtual model, containing the current pose (e.g., X: 1.234m, Y: 2.345m, Yaw: 12.3°) and the target pose (e.g., X: 1.500m, Y: 2.800m, Yaw: 0.0°).
[0051] Component Information: The actual outline of the prefabricated component to be transported, as identified by the pose recognition module, will be displayed as a semi-transparent, highlighted red outline. Simultaneously, the four preset lifting points on the component will be marked with flashing gold spheres, indicating to the operator that the grippers should align with these positions.
[0052] Distance information: On the real-time images captured by the panoramic camera, the distance between the robot body and surrounding obstacles (such as scaffolding, other components) is dynamically marked with colored lines (green represents safety, yellow represents approach, and red represents danger).
[0053] Historical and Predicted Trajectory: During automated handling, the robot's historical paths are displayed as light gray dashed lines, while the predicted path for the next second based on its current motion state is displayed as blue dotted lines. When encountering temporary obstacles and requiring replanning, the replanned path is highlighted in orange, contrasting with the original green path.
[0054] Operators can perform manual fine-tuning in two ways. The first is through gestures: the depth camera on the AR glasses captures the skeletal points of the hand; when the operator extends their index finger and slides it to the right in the air, the end effector translates to the right; when the wrist rotates clockwise, the end effector rotates clockwise around the vertical axis; when the five fingers are brought together and then clenched into a fist, the gripper performs a clamping action. The second method is through paired controllers. The joystick offset of the controllers is mapped using a speed-mapping curve (e.g., an exponential curve: small offsets in the low-speed range correspond to extremely low speeds for precise alignment; offsets in the medium-to-high-speed range are approximately linear with the speed for rapid adjustments), making the operation smoother.
[0055] Example 5: This embodiment describes in detail the specific process and judgment logic for switching from automatic coarse positioning to manual fine-tuning mode.
[0056] Step S1: After the operator selects the target area through the AR interface, the transport robot body starts the automatic mode.
[0057] Step S2: The path planning module plans the global path, and the robot begins to move. During this process, the pose recognition module calculates the deviation of the end effector relative to the final target pose in real time.
[0058] Step S3: The controller determines whether the deviation falls within a preset threshold range. In this embodiment, the threshold is a translational deviation of <5cm and a rotational deviation of <2°. If the deviation does not fall within the threshold range, the automatic handling process continues; if the deviation does fall within the threshold range, proceed to step S4.
[0059] Step S4: The robot controller sends a "fine-tuning request" signal to the AR glasses via a wireless link.
[0060] Step S5: After the AR glasses receive the signal, they trigger the vibration motor (double pulse vibration for 0.5 seconds) and a prompt box pops up in the center of the field of vision: "Enter manual fine-tuning". At the same time, a voice prompt "Please take over control" is played through the bone conduction speaker.
[0061] Step S6: The operator raises their hand or picks up the handle to make the first fine-tuning motion. After detecting the gesture or handle signal, the manual fine-tuning control module 5 automatically takes over control, and the robot switches to manual fine-tuning mode. At this time, the information on the augmented reality interface automatically switches to fine-tuning assistance mode (see Example 5).
[0062] If the operator does not take any action within 10 seconds, the robot will automatically pause and issue an alarm.
[0063] Example 6: In manual fine-tuning mode, to help operators complete alignment quickly and accurately, the augmented reality interface displays the real-time deviation in a very intuitive way.
[0064] The deviation is defined as the difference between the current pose of the end effector and the target pose: translational deviation components (Δx, Δy, Δz) and rotational deviation components (Δroll, Δpitch, Δyaw). In the AR view, a semi-transparent "deviation indicator panel" appears next to the virtual model of the end effector. The panel displays each component as a numerical label, such as: "Δx: +2.3 mm", "Δy: -0.8 mm", "Δz: +5.1 mm", "Δroll: +0.2°", etc.
[0065] Dynamic prompts: The background color of the numerical labels changes dynamically based on the absolute value of the deviation. For example, the background is red when a component deviation is >10mm, yellow when it's 5-10mm, and green when it's <5mm. Simultaneously, a semi-transparent target model is displayed at the target pose, clearly showing the overlap deviation between the current end effector model and this model. When all deviation components are less than 0.5mm and 0.05°, the interface displays a green "Alignment Complete" confirmation icon and indicates that the placement action can proceed.
[0066] Example 7: This embodiment adds a remote collaboration unit to the existing embodiment 1 to handle complex or abnormal working conditions.
[0067] When on-site operators encounter situations that make difficult decisions (such as the target placement location being partially obstructed by temporary materials, or the prefabricated component having an unusual shape), they can activate the "Remote Assistance" function through the menu on the AR glasses. At this time, the AR glasses will transmit the augmented reality image (including the real scene and all overlaid virtual information) in the operator's field of vision in real time as a low-latency video stream to a remote expert terminal (such as a PC with a touch screen) located in the office or in the cloud.
[0068] After viewing the video feed on their devices, experts can use a mouse or stylus to annotate the screen: for example, drawing a red arrow to indicate "Clear the clutter here first," or circling an area and adding the text "Try entering from the side." These annotations are encoded in real-time and transmitted back to the AR glasses, precisely overlaid on the corresponding 3D location in the operator's field of vision (the annotations remain anchored in the real scene even when the operator's head rotates). Operators can then perform actions based on the expert's remote guidance. All remote sessions can be recorded and saved for subsequent training or incident analysis.
[0069] Example 8: Considering the drastic changes in lighting at the construction site (from indoor shadows to outdoor bright light), this embodiment optimizes the display of the AR glasses. The AR glasses integrate an ambient light sensor. The sensor detects the ambient illuminance value (Lux) in real time. The AR glasses' operating system automatically adjusts the display brightness based on the Lux value: reducing brightness in dark areas to avoid glare, and increasing brightness in bright light to ensure the virtual content is visible.
[0070] Simultaneously, the AR glasses transmit this Lux value wirelessly to the panoramic camera on the transport robot, which then adjusts the camera's exposure parameters (exposure time, ISO, aperture, etc.). In this way, when the brightness of the real scene in the operator's field of vision changes, the brightness of the image transmitted back by the panoramic camera also adjusts adaptively. Ultimately, this ensures that the brightness of the virtual information superimposed on the AR glasses (such as distance tags and path lines) remains consistent with the brightness of the real scene in terms of human visual perception, preventing the virtual information from being overexposed or underexposed and thus unrecognizable.
[0071] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A smart carrying robot with pose recognition function for fabricated buildings, characterized in that: The system includes a transport robot body, which is equipped with a pose recognition module, a panoramic camera, a path planning module, and a manual fine-tuning control module. The transport robot body is connected to the augmented reality interaction unit via a wireless communication link, and the augmented reality interaction unit includes AR glasses; The AR glasses use a combination of ultra-wideband positioning and visual markers to align the spatial pose of the transport robot and overlay an augmented reality interface in the operator's field of vision. The augmented reality interface includes: the planned path generated by the path planning module, the current pose and target pose of the transport robot body obtained by the pose recognition module, and the distance information between the prefabricated components and obstacles superimposed in the construction site image captured by the panoramic camera. The manual fine-tuning control module responds to the gesture commands recognized by the AR glasses or the operation of the handle paired with the AR glasses, generates a six-degree-of-freedom control signal, and drives the end effector of the handling robot body to perform fine-tuning of position and attitude. During manual fine-tuning, the augmented reality interface displays the deviation between the end effector and the target pose in real time and provides prompts through the augmented reality interface. The transport robot automatically performs coarse positioning and transport operations based on the target position and orientation specified by the operator in the augmented reality interface. After coarse positioning is completed, it switches to manual fine-tuning mode, and the operator completes fine alignment through the augmented reality interaction unit.
2. The intelligent carrying robot with pose recognition function for fabricated building according to claim 1, characterized in that: The spatial pose alignment includes: the AR glasses measuring distances with UWB anchor points arranged on the transport robot body through their built-in UWB tags, recognizing preset visual markers on the transport robot body using the downward-facing camera on the AR glasses, calculating the six-degree-of-freedom pose of the AR glasses relative to the transport robot body, and aligning the coordinate system of the augmented reality content with the coordinate system of the transport robot body.
3. The smart carrying robot with pose recognition function for fabricated building according to claim 1, characterized in that: The augmented reality interface is also configured to overlay the actual pose contour of the prefabricated component identified by the pose recognition module, and mark the positions of the preset lifting points or embedded sleeves on the prefabricated component with a highlight color to assist the operator in judging the precise position of clamping or placing.
4. The smart carrying robot with pose recognition function for fabricated building according to claim 1, characterized in that: The manual fine-tuning control module includes a gesture recognition unit, which collects the operator's hand skeletal point movements through the depth camera on the AR glasses and maps predefined gestures to motion commands of the end effector. The predefined gestures include finger sliding corresponding to translational movement, wrist rotation corresponding to rotational movement, and fist clenching corresponding to clamping action.
5. The smart carrying robot with pose recognition function for fabricated building according to claim 1, characterized in that: The process of switching to manual fine-tuning mode is as follows: the handling robot body automatically plans a coarse positioning path and performs handling based on the target area selected by the operator in the augmented reality interface and the current pose of the prefabricated component output by the pose recognition module. When the handling robot body determines that the deviation between its own pose and the target pose enters the preset threshold range, it automatically sends a switching request to the AR glasses. The AR glasses emit vibration or voice prompts to notify the operator to take over manual fine-tuning.
6. The smart carrying robot with pose recognition function for fabricated building according to claim 1, characterized in that: The deviations include translational deviations (Δx, Δy, Δz) and rotational deviations (Δroll, Δpitch, Δyaw) between the end effector's current pose and the target pose. The augmented reality interface displays each deviation component as a digital label next to the virtual model of the end effector and dynamically changes the background color of the digital label according to the magnitude of the absolute value of the deviation.
7. The smart carrying robot with pose recognition function for fabricated building according to claim 1, characterized in that: It also includes a remote collaboration unit: the AR glasses transmit the augmented reality image in the operator's field of vision to a remote expert terminal in real time. The remote expert terminal receives the annotation information input by the expert and overlays the annotation information in the field of vision of the AR glasses. The annotation information includes at least one of virtual arrows, circled areas, and text descriptions. 8.The prefabricated building oriented intelligent carrying robot with pose recognition function according to claim 1, characterized in that: The augmented reality interface is also configured to display the historical travel path and future predicted path of the transport robot in the form of trajectory lines during the automatic transport process, and to distinguish between the planned path and the dynamically replanned path with different colors. 9.The prefabricated building oriented intelligent carrying robot with pose recognition function according to claim 1, characterized in that: The manual fine-tuning control module includes a speed mapping curve, which non-linearly maps the offset of the joystick to the movement speed of the end effector.
10. The smart carrying robot with pose recognition function for fabricated building according to claim 1, characterized in that: The AR glasses are equipped with an ambient light sensor, and the display brightness of the augmented reality interface is automatically adjusted according to the ambient light intensity at the construction site. The adjustment of the display brightness is linked to the exposure parameters of the panoramic camera, so that the superimposed virtual information is consistent with the brightness of the real scene.