Control method, apparatus, system, and storage medium

By overlaying contact points and orientation markers onto the end effector's display, the problem of lack of visual feedback in teleoperated grasping tasks is solved, thereby improving the grasping success rate and operational efficiency.

CN122172630APending Publication Date: 2026-06-09INDEPENDENT VARIABLE ROBOT TECHNOLOGY (SHENZHEN) CO LTD +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610646296.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

In teleoperated grasping tasks, users lack intuitive visual feedback when controlling the robot's end effector through control devices such as handles, resulting in a low success rate in grasping objects.

Method used

Overlaying contact point markers and/or orientation markers on the screen displaying the end effector allows the user to know the predicted contact point position and orientation with the object when the end effector performs a closing action, facilitating precise adjustment of the end effector's position and orientation.

Benefits of technology

It improves the success rate of user-controlled end effector grasping objects, reduces the number of times users need to adjust the position and orientation, and improves operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122172630A_ABST
    Figure CN122172630A_ABST
Patent Text Reader

Abstract

The method comprises: controlling a display device to display a picture, the picture displaying an end effector; determining a contact point identifier and / or an orientation identifier according to pose information of the end effector, wherein the contact point identifier is used to represent a position of a predicted contact point of the end effector with an object when the end effector performs a closing action, and the orientation identifier is used to represent an orientation of the end effector; and controlling the display device to superimpose and display the contact point identifier and / or the orientation identifier in the picture. The technical solution of the embodiment of the present disclosure improves the success rate of the user controlling the end effector to grasp the object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of robotics, and more particularly to a control method, apparatus, system, and storage medium. Background Technology

[0002] With the rapid development of robotics technology, teleoperated robots are widely used in industrial manufacturing, medical surgery, and hazardous environment operations. However, in teleoperated grasping tasks, the lack of intuitive visual feedback when users control the robot's robotic arm and end effector to perform grasping actions via control devices such as handles results in a low success rate for users controlling the end effector to grasp objects. Summary of the Invention

[0003] This disclosure provides a control method, apparatus, system, and storage medium designed to improve the success rate of user-controlled end effector grasping objects.

[0004] In a first aspect, embodiments of this disclosure provide a control method, including: The control display device displays a screen showing an end effector; Based on the pose information of the end effector, a contact point identifier and / or an orientation identifier are determined, wherein the contact point identifier indicates the position of the predicted contact point between the end effector and the object when the end effector performs a closing action, and the orientation identifier indicates the orientation of the end effector; and Control the display device to overlay the contact point identifier and / or the orientation identifier on the screen.

[0005] Secondly, embodiments of this disclosure also provide a control device, including: At least one processor; At least one memory containing computer program code; The at least one memory, the computer program code, and the at least one processor are configured together to implement the control method described in the first aspect.

[0006] Thirdly, embodiments of this disclosure also provide a system, including: The control device described in the second aspect; Input device, which is communicatively connected to the control device; and An end effector or a robot including an end effector, the end effector or the robot including an end effector being communicatively connected to the control device, the input device and the control device being used to control the end effector to perform grasping.

[0007] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the control method described in the first aspect.

[0008] This disclosure provides a control method, apparatus, system, and storage medium. In this embodiment, a contact point identifier and / or orientation identifier are overlaid on a screen displaying an end effector. This allows the user to know the predicted contact point position between the end effector and the object when the end effector performs a closing action via the contact point identifier, and / or to know the orientation of the end effector via the orientation identifier. This facilitates precise adjustment of the end effector's position and / or orientation by the user, improving the success rate of the user controlling the end effector to grasp objects. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. The accompanying drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a schematic diagram of an architecture for implementing the control method provided in the embodiments of this disclosure; Figure 2 This is another schematic diagram of the architecture for implementing the control method provided in the embodiments of this disclosure; Figure 3 This is a schematic block diagram of the structure of a control device provided in an embodiment of this disclosure; Figure 4 This is a flowchart illustrating a control method provided in an embodiment of this disclosure; Figure 5 This is a schematic block diagram of a system provided in an embodiment of the present disclosure; Figure 6 This is a schematic block diagram of another system provided in this embodiment. Detailed Implementation

[0011] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. The described embodiments are some, but not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0012] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0013] It should be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure. As used in this disclosure and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0014] With the rapid development of robotics technology, teleoperated robots are widely used in industrial manufacturing, medical surgery, and hazardous environment operations. However, in teleoperated grasping tasks, the lack of intuitive visual feedback when users control the robot's robotic arm and end effector to perform grasping actions via control devices such as handles results in a low success rate for users controlling the end effector to grasp objects.

[0015] To address the aforementioned issues, this disclosure provides a control method, apparatus, system, and storage medium. In this embodiment, a contact point identifier and / or orientation identifier are overlaid on a screen displaying an end effector. This allows the user to determine the predicted contact point between the end effector and the object during a closing action via the contact point identifier, and / or the orientation of the end effector via the orientation identifier. This facilitates precise adjustment of the end effector's position and / or orientation, preventing users from blindly grasping objects and improving the success rate of the user controlling the end effector to grasp objects.

[0016] The following detailed description of some embodiments of this disclosure is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0017] Please see Figure 1 , Figure 1 This is a schematic diagram of an architecture for implementing the control method provided in the embodiments of this disclosure.

[0018] like Figure 1 As shown, the architecture includes a control device 110, an input device 120, an end effector 130, and a display device 200. The control device 110 is communicatively connected to the input device 120, the end effector 130, and the display device 200. The control device 110 and the input device 120 are used to control the end effector 130 to perform grasping. For example, the control device 110 receives grasping control information from the input device 120 and sends the grasping control information to the end effector 130 to control the end effector 130 to perform grasping. Figure 1In the illustrated architecture, the end effector 130 is mounted on a robotic arm, which is mounted on a base that cannot move autonomously. It should be noted that the control device 110 and the input device 120 can be integrated or separate; this embodiment does not impose a specific limitation on either.

[0019] Please see Figure 2 , Figure 2 This is another schematic diagram of the architecture for implementing the control method provided in the embodiments of this disclosure.

[0020] like Figure 2 As shown, the architecture includes a control device 110, an input device 120, a robot 140 including an end effector, and a display device 200. The control device 110 is communicatively connected to the input device 120, the robot 140, and the display device 200. The control device 110 and the input device 120 are used to control the robot 140. For example, the control device 110 receives control information from the input device 120 and sends the control information to the robot to control the robot to perform the action corresponding to the control information. It should be noted that the control device 110 and the input device 120 can be integrated or separate; this embodiment does not specifically limit this.

[0021] In this embodiment of the disclosure, Figure 1 or Figure 2 The input device 120 shown includes an operating handle device and / or an exoskeleton device. Figure 1 The end effector 130 shown or Figure 2 The robot 140 shown includes an end effector comprising a gripper, which may be a two-finger gripper, a three-finger gripper, or a five-finger gripper, etc. Figure 1 or Figure 2 The display device 200 shown includes smartphones, tablets, laptops, desktop computers, personal digital assistants, and head-mounted display devices, including virtual reality (VR) head-mounted display devices, augmented reality (AR) head-mounted display devices, and mixed reality (MR) head-mounted display devices.

[0022] In this embodiment of the disclosure, Figure 2 The robot 140 shown can be of different types and can be used in industrial, commercial, or household fields to perform different operations. This disclosure does not specifically limit the type of robot. For example, the robot in this disclosure can be a embodied robot or a non-embodied robot. Depending on the mode of locomotion, the robot in this disclosure can be a wheeled robot, a tracked robot, or a legged robot.

[0023] In some embodiments, such as Figure 3 As shown, Figure 1 or Figure 2 The control device 110 shown includes at least one processor 111 and at least one memory 112 containing computer program code. The at least one processor 111 and at least one memory 112 are connected via a bus 113, such as an I2C (Inter-integrated Circuit) bus.

[0024] Specifically, the processor 111 can be a microcontroller unit (MCU), a central processing unit (CPU), or a digital signal processor (DSP), etc.

[0025] Specifically, the memory 112 can be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a portable hard drive, etc.

[0026] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to this disclosure and does not constitute a limitation on the control device to which the solution of this disclosure is applied. A specific control device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0027] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0028] In this embodiment, at least one of the memory 112 and computer program code, together with at least one of the processors 111, are configured to enable the control device 110 to implement any of the control methods provided in the embodiments of this disclosure.

[0029] In some embodiments, at least one of the memory 112 and computer program code, together with at least one of the processors 111, are configured to enable the control device 110 to perform: The control display device 200 displays a screen showing the end effector; Based on the pose information of the end effector, a contact point identifier and / or an orientation identifier are determined, wherein the contact point identifier indicates the position of the predicted contact point between the end effector and the object when the end effector performs a closing action, and the orientation identifier indicates the orientation of the end effector; and The control display device 200 overlays the contact point identifier and / or the orientation identifier on the screen.

[0030] In some embodiments, at least one of the memory 112 and the computer program code, together with at least one of the processor 111, are configured to cause the control device 110 to further implement: Based on the confidence level of the predicted contact point and the location of the predicted contact point, control the display device to display the contact point identifier; or The display device is controlled to display the confidence level of the predicted contact point.

[0031] In some embodiments, at least one of the memory 112 and the computer program code, together with at least one of the processor 111, are configured to cause the control device 110 to further implement: The confidence level is determined based on the deviation between the predicted contact point positions obtained from at least two different viewpoints.

[0032] In some embodiments, the end effector is mounted on a robot, a first camera is mounted on the robot's head, and a second camera is mounted on the robot's chest. The first camera and the second camera have different perspectives, and the at least two different perspectives include the perspectives of the first camera and the second camera; or The end effector is mounted on a robot, which is equipped with a first camera and a second camera with different viewing angles, wherein the at least two different viewing angles include the views of the first camera and the second camera; or The end effector is mounted on the robot, and a first camera is mounted on the end effector. A second camera is mounted on the robot. The first camera and the second camera have different perspectives, and the at least two different perspectives include the perspectives of the first camera and the second camera.

[0033] In some embodiments, at least one of the memory 112 and the computer program code, together with at least one of the processor 111, are configured to cause the control device 110 to further implement: Based on the perpendicularity of the end effector's orientation to the surface of the object and the orientation of the end effector, the display device is controlled to display the orientation indicator; or The display device is controlled to show the degree of perpendicularity between the orientation of the end effector and the surface of the object.

[0034] In some embodiments, at least one of the memory 112 and the computer program code, together with at least one of the processor 111, are configured to cause the control device 110 to further implement: Control the display device to display the orientation of the end effector towards a target point, the target point being located on the object; and / or The display device is controlled to display the distance between the end effector and the object.

[0035] In some embodiments, at least one of the memory 112 and the computer program code, together with at least one of the processor 111, are configured to cause the control device 110 to further implement: The display device is controlled to display a size matching identifier, which indicates the degree of matching between the opening amplitude of the end effector and the size of the object.

[0036] In some embodiments, at least one of the memory 112 and the computer program code, together with at least one of the processor 111, are configured to cause the control device 110 to further implement: The display device is controlled to display a grasping ready flag, which indicates that the end effector is in a grasping ready state.

[0037] In some embodiments, at least one of the memory 112 and the computer program code, together with at least one of the processor 111, are configured to cause the control device 110 to further implement: In response to the end effector meeting the preset grasping ready condition, it is determined that the end effector is in the grasping ready state; The end effector satisfies at least one of the following preset grasping ready conditions: the confidence level of the predicted contact point position is greater than or equal to a preset confidence level; the perpendicularity of the end effector's orientation to the surface of the object is greater than or equal to a preset perpendicularity level; and the degree of matching between the end effector's opening amplitude and the size of the object is greater than or equal to a preset matching degree.

[0038] In some embodiments, at least one of the memory 112 and the computer program code, together with at least one of the processor 111, are configured to cause the control device 110 to further implement: The display device is controlled to display grasping force prompt information, which is used to provide the user with a prompt of the force required for the end effector to grasp the object. The force is estimated based on the material of the object and the surface curvature of the object at the predicted contact point.

[0039] In some embodiments, at least one of the memory 112 and the computer program code, together with at least one of the processors 111, are configured to cause the control device 110 to perform the following when determining contact point identifiers and / or orientation identifiers based on the pose information of the end effector: Based on the opening / closing state information and the pose information of the end effector, determine the closing path of the end effector; determine the point where the closing path intersects with the object as the predicted contact point; determine the contact point identifier based on the position of the predicted contact point; and / or Based on the pose information of the end effector, the orientation of the end effector is determined; based on the orientation of the end effector, the orientation identifier is determined.

[0040] In some embodiments, at least one of the memory 112 and the computer program code, together with at least one of the processors 111, are configured to enable the control device 110 to control the display device to overlay the contact point identifier and / or the orientation identifier on the screen, for the purpose of: The display device is controlled to overlay the orientation indicator onto the screen; In response to the existence of the predicted contact point with the object when the end effector performs a closing action, the display device is controlled to overlay the contact point identifier on the screen.

[0041] It should be noted that those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the control device described in the embodiments of this disclosure can be referred to the corresponding process in the following control method embodiments, and will not be repeated here.

[0042] The following will combine Figure 1 or Figure 2 The architecture described in this disclosure provides a detailed overview of the control methods provided by the embodiments. It should be noted that... Figure 1 or Figure 2 The architecture described herein is only used to explain the control methods provided in the embodiments of this disclosure, but does not constitute a limitation on the application architecture of the control methods provided in the embodiments of this disclosure.

[0043] Please see Figure 4 , Figure 4 This is a flowchart illustrating a control method provided in an embodiment of this disclosure.

[0044] like Figure 4 As shown, the control method includes steps S101 to S103.

[0045] Step S101: Control the display device to display the screen, and the screen displays the end effector.

[0046] Step S102: Determine the contact point identifier and / or orientation identifier based on the pose information of the end effector, wherein the contact point identifier is used to indicate the position of the predicted contact point between the end effector and the object when the end effector performs a closing action, and the orientation identifier is used to indicate the orientation of the end effector.

[0047] Step S103: Control the display device to overlay the contact point mark and / or orientation mark on the screen.

[0048] This embodiment overlays a contact point marker and / or orientation marker onto the screen displaying the end effector. This allows the user to know the predicted contact point position between the end effector and the object when the end effector performs a closing action through the contact point marker, and to know the orientation of the end effector through the orientation marker. This facilitates the user to accurately adjust the position and orientation of the end effector in advance, improving the success rate of the user controlling the end effector to grasp the object, and also reducing the number of times the user needs to adjust the position and orientation of the end effector, thus improving operational efficiency.

[0049] In some embodiments, controlling the display device to display an image includes: controlling the display device to display a view captured in real time by a camera; when the end effector is not mounted on the robot, the end effector is mounted with the camera; when the end effector is mounted on the robot, the end effector is mounted with the camera or the camera is mounted on a location on the robot other than the end effector.

[0050] In some embodiments, the predicted contact point between the end effector and the object when the end effector performs a closing action includes the point where the end effector intersects with the object during the simulated closing action. In this case, the end effector does not actually perform a closing action.

[0051] In some embodiments, determining the contact point identifier and / or orientation identifier based on the pose information of the end effector includes: determining the closed path of the end effector based on the opening / closing state information and pose information of the end effector; determining the point where the closed path intersects with the object as the predicted contact point; and determining the contact point identifier based on the position of the predicted contact point. This embodiment simulates a scenario where the end effector grasps an object in its current pose by using collision detection between the closed path of the end effector and the object, thereby accurately predicting the contact point between the end effector and the object.

[0052] In some embodiments, the opening / closing state information of the end effector includes the opening amplitude of the end effector, and the pose information of the end effector includes the position and orientation of the end effector. Determining the closing path of the end effector based on the opening / closing state information and pose information includes: determining the path along the closing direction of the end effector, from the opening position of the end effector to the closing direction of the end effector, which is half of the opening amplitude of the end effector, as the closing path of the end effector.

[0053] In some embodiments, determining the point where the closed path intersects with the object as the predicted contact point includes: determining the point closest to the closed path in the point cloud data of the object within the environment where the end effector is located; and determining the point closest to the closed path in the point cloud data as the predicted contact point in response to the point cloud data being less than a preset distance threshold between the point closest to the closed path and the closed path. The preset distance threshold can be set based on actual conditions, and this embodiment does not specifically limit it.

[0054] In some embodiments, determining the point where the closed path intersects with the object as a predicted contact point includes: acquiring target point cloud data of the object within the environment where the end effector is located; determining points in the target point cloud data whose distance from the closed path is less than a preset distance threshold as collision points; determining the collision point as a predicted contact point in response to the number of collision points being 1; and determining the first determined collision point or any collision point as a predicted contact point in response to the number of collision points being greater than or equal to 2.

[0055] In some embodiments, acquiring target point cloud data of objects within the environment of the end effector includes: acquiring first point cloud data of the environment of the end effector; and acquiring second point cloud data of objects within the environment of the end effector from the first point cloud data. The first point cloud data is generated based on images captured by at least two cameras in the robot on which the end effector is installed.

[0056] In some embodiments, determining the contact point identifier and / or orientation identifier based on the pose information of the end effector includes: determining the orientation of the end effector based on the pose information of the end effector; and determining the orientation identifier based on the orientation of the end effector.

[0057] In some embodiments, a first camera is mounted on the robot's head and a second camera is mounted on the robot's chest. Acquiring first point cloud data of the environment where the end effector is located includes: controlling the first camera to acquire a first RGB-D image and simultaneously controlling the second camera to acquire a second RGB-D image; performing timestamp alignment processing on the first RGB-D image and the second RGB-D image; and generating first point cloud data of the environment where the end effector is located based on the timestamp-aligned first RGB-D image and the second RGB-D image.

[0058] In this specification, to clearly describe the robot's structure, the anthropomorphic terms "head" and "chest" are used to define the upper and middle components of the robot. For example, the robot includes a head, chest, and waist arranged vertically from top to bottom. The head is located at the top of the robot and its shape mimics a human head. The chest is located directly below the head and forms the upper half of the robot's torso. The two are connected by a cervical joint, either integrally or separately, allowing the head to perform pitch and / or horizontal rotation relative to the chest. The lower end of the head is fixedly connected to the upper end of the cervical joint, and the lower end of the cervical joint is fixed to the upper surface of the chest. It should be noted that the terms "head" and "chest" are used only to refer to different structural locations and component sets, and do not limit the robot to having a completely anthropomorphic appearance; any component located at the top and capable of movement or fixation relative to the lower torso can be considered the head, and any component located below the head and forming the upper part of the torso can be considered the chest.

[0059] In some embodiments, the display device displays a viewfinder image from a camera. Determining the contact point identifier based on the predicted contact point position includes: projecting the predicted contact point position onto the camera's image plane based on the camera's intrinsic and extrinsic parameter matrices to obtain the projected position of the predicted contact point in the camera's image plane; and generating a contact point identifier at the projected position. The style of the contact point identifier can be set based on actual conditions, and this disclosure does not impose specific limitations on it. For example, the contact point identifier may be a solid dot in a circular shape, and the radius of the solid dot may be positively correlated with the size of the fingertip of the end effector.

[0060] In some embodiments, controlling the display device to overlay contact point identifiers on a screen includes: controlling the display device to overlay at least two contact point identifiers on the same screen, wherein the at least two contact point identifiers are used to indicate the positions of the predicted contact points with the object in at least two different directions when the end effector performs a closing action. For example, the end effector is a two-finger gripper, which includes a first gripper finger and a second gripper finger, the first gripper finger and the second gripper finger being disposed opposite each other, and the at least two contact point identifiers include contact point identifier A and contact point identifier B, where contact point identifier A is used to indicate the position of the predicted contact point between the first gripper finger and the object when the gripper performs a closing action, and contact point identifier B is used to indicate the position of the predicted contact point between the second gripper finger and the object when the gripper performs a closing action.

[0061] In some embodiments, before the control display device overlays the orientation marker on the screen, the control method provided in this disclosure further includes: obtaining the position and orientation of the end effector in the robot coordinate system; generating a ray along the orientation direction of the end effector in the robot coordinate system, using the position of the end effector in the robot coordinate system as the target starting point; determining the nearest intersection point of the ray and the object in the environment where the end effector is located as the target ending point in response to the ray intersecting with the object in the environment where the end effector is located, and determining a point along the ray at a preset distance from the target starting point as the target ending point in response to the ray not intersecting with the object in the environment where the end effector is located; projecting the target starting point and the target ending point onto the image plane corresponding to the camera according to the intrinsic and extrinsic parameter matrices of the camera that captures the screen, obtaining the target projection starting point and the target projection ending point; generating a ray segment marker according to the target projection starting point and the target projection ending point, and using the ray segment marker as the orientation marker. The preset distance can be set based on actual conditions, and this disclosure does not specifically limit it. For example, the preset distance is 1 meter.

[0062] In some embodiments, controlling the display device to display images includes: controlling the display device to display at least two images, the at least two images displaying the end effector, and the at least two images corresponding to at least two cameras having different viewing angles; controlling the display device to overlay contact point identifiers and / or orientation identifiers on the images includes: controlling the display device to overlay contact point identifiers and / or orientation identifiers on at least two images. In this embodiment, by using contact point identifiers and / or orientation identifiers provided by at least two different viewing angles, the user can more accurately know the position of the predicted contact point between the end effector and the object when the end effector performs a closing action, as well as the orientation of the end effector. This makes it easier for the user to accurately adjust the position and orientation of the end effector, further improving the success rate of the user controlling the end effector to grasp the object.

[0063] For example, an end effector is mounted on a robot. A first camera is mounted on the robot's head, and a second camera is mounted on the robot's chest. The first viewpoint corresponding to the first camera and the second viewpoint corresponding to the second camera are different. The control display device displays at least two images, including: the control display device displays the first image captured in real time by the first camera and the second image captured in real time by the second camera; the control display device overlays contact point identifiers and / or orientation identifiers on the at least two images, including: the control display device overlays and displays the first contact point identifier and / or the first orientation identifier on the first image, where the first contact point identifier indicates the position of the predicted contact point between the end effector and the object when performing a closing action from the first viewpoint, and the first orientation identifier indicates the orientation of the end effector from the first viewpoint; and the control display device overlays and displays the second contact point identifier and / or the second orientation identifier on the second image, where the second contact point identifier indicates the position of the predicted contact point between the end effector and the object when performing a closing action from the second viewpoint, and the second orientation identifier indicates the orientation of the robot's end effector from the second viewpoint.

[0064] It should be noted that the specific generation method of the contact point identifier for each frame is the same; the difference lies in the transformation matrix used to generate the contact point identifier for each frame. For example, controlling the display device to overlay the first contact point identifier in the first frame includes: obtaining the position and orientation of the end effector in the robot coordinate system, and obtaining the first transformation matrix from the robot's base to the first camera; transforming the position and orientation of the end effector in the robot coordinate system to the coordinate system of the first camera according to the first transformation matrix; determining the path that closes half of the end effector's opening amplitude along the closing direction of the end effector as the first closing path of the end effector; determining the position of the point where the first closing path intersects with the object as the first position of the predicted contact point between the end effector and the object in the first viewpoint; and controlling the display device to overlay the first contact point identifier in the first frame according to the first position.

[0065] Similarly, controlling the display device to overlay the second contact point identifier on the second screen includes: obtaining the position and orientation of the end effector in the robot coordinate system, and obtaining the second transformation matrix from the robot's base to the second camera; transforming the position and orientation of the end effector in the robot coordinate system to the coordinate system of the second camera according to the second transformation matrix; determining the path that closes half the opening amplitude of the end effector along the closing direction of the end effector as the second closing path of the end effector; determining the position of the point where the second closing path intersects with the object as the second position of the predicted contact point between the end effector and the object in the second view; and controlling the display device to overlay the second contact point identifier on the second screen according to the second position.

[0066] It should be noted that the specific generation method for the orientation markers corresponding to each frame is the same; the difference lies in the intrinsic and extrinsic parameter matrices of the camera used to generate the orientation markers for each frame. For example, the generation method for the first orientation marker corresponding to the first frame includes: obtaining the position and orientation of the end effector in the robot coordinate system; using the position of the end effector in the robot coordinate system as the target starting point, generating a ray along the orientation direction of the end effector in the robot coordinate system; in response to the ray intersecting with an object in the environment where the end effector is located, determining the nearest intersection point of the ray and the object in the environment where the end effector is located as the target endpoint; in response to the ray not intersecting with an object in the environment where the end effector is located, determining the point along the ray at a preset distance from the target starting point as the target endpoint; transforming the target starting point and target endpoint to the coordinate system of the first camera to obtain the first starting point and the first endpoint; based on the intrinsic parameter matrix of the first camera that captures the first frame, projecting the first starting point and the first endpoint onto the image plane corresponding to the first camera to obtain the first projection starting point and the first projection endpoint; generating a first ray segment marker based on the first projection starting point and the first projection endpoint, and using the first ray segment marker as the first orientation marker.

[0067] Similarly, the generation method of the second orientation identifier corresponding to the second image includes: obtaining the position and orientation of the end effector in the robot coordinate system; taking the position of the end effector in the robot coordinate system as the target starting point, generating a ray along the orientation direction of the end effector in the robot coordinate system; in response to the ray intersecting with an object in the environment where the end effector is located, determining the nearest intersection point of the ray and the object in the environment where the end effector is located as the target endpoint; in response to the ray not intersecting with an object in the environment where the end effector is located, determining the point along the ray at a preset distance from the target starting point as the target endpoint; transforming the target starting point and the target endpoint to the coordinate system of the second camera to obtain the second starting point and the second endpoint; according to the intrinsic parameter matrix of the second camera that captures the second image, projecting the second starting point and the second endpoint onto the image plane corresponding to the second camera to obtain the second projection starting point and the second projection endpoint; generating a second ray segment identifier based on the second projection starting point and the second projection endpoint, and using the second ray segment identifier as the second orientation identifier.

[0068] In some embodiments, the position of the contact point identifier in the image is updated as the end effector and / or object moves; and, or, the orientation identifier includes an end effector orientation indicator segment, the position of the starting point of which in the image is updated as the end effector moves.

[0069] In some embodiments, controlling the display device to display an image includes: controlling the AR display device, VR display device, or MR display device to display an image; controlling the display device to overlay and display touch point identifiers and / or orientation identifiers on the image includes: controlling the AR display device, VR display device, or MR display device to overlay and display touch point identifiers and / or orientation identifiers on the image. This embodiment, by displaying an image and touch point identifiers and / or orientation identifiers on an AR display device, VR display device, or MR display device, provides a better sense of immersion and a better user experience.

[0070] In some embodiments, the control method provided in this disclosure further includes: while controlling the display device to overlay and display contact point identifiers and / or orientation identifiers on the screen, controlling the input device or display device to output contact point prompting sounds and / or orientation prompting sounds, wherein the contact point prompting sounds are used to prompt the user that the end effector will contact or not contact the object when performing a closing action, and the orientation prompting sounds are used to prompt the user that the end effector is aligned with or not aligned with the object; and / or, while controlling the display device to overlay and display contact point identifiers and / or orientation identifiers on the screen, controlling the input device to output contact point prompting vibration signals and / or orientation prompting vibration signals, wherein the contact point prompting vibration signals are used to prompt the user that the end effector will contact or not contact the object when performing a closing action, and the orientation prompting vibration signals are used to prompt the user that the end effector is aligned with or not aligned with the object.

[0071] In some embodiments, controlling the display device to overlay contact point markers and / or orientation markers on the screen includes: controlling the display device to overlay orientation markers on the screen; and controlling the display device to overlay contact point markers on the screen in response to a predicted contact point with the object when the end effector performs a closing action. This embodiment effectively reduces the user's cognitive and spatial calculation burden by displaying orientation markers first and then contact point markers in a phased, progressive visual guidance manner, realizing a control flow from coarse positioning to fine adjustment. Thus, when the end effector is far from the object, the orientation marker assists the user in controlling the end effector to move closer to the object; when the end effector is close to the object, the contact point marker assists the user in grasping the object, thereby improving the success rate and consistency of the grasping task.

[0072] In some embodiments, controlling the display device to overlay a contact point identifier on the screen includes: controlling the display device to display the contact point identifier based on the confidence level and the predicted position of the predicted contact point. In this embodiment, the contact point identifier is displayed based on the confidence level and the predicted position of the predicted contact point. This allows the user to not only know the predicted position of the contact point between the end effector and the object when performing a closing action, but also whether the predicted position is reliable. This avoids misleading the user with the displayed contact point identifier, making it easier for the user to accurately adjust the position and orientation of the end effector, resulting in a better user experience.

[0073] In some embodiments, the position of the contact point identifier in the screen indicates the predicted contact point between the end effector and the object when the end effector performs a closing action, and the visual representation of the contact point identifier in the screen indicates the confidence level of the predicted contact point. The visual representation of the contact point identifier differs depending on the confidence level of the predicted contact point. For example, if the confidence level of the predicted contact point is greater than or equal to a preset confidence threshold, the contact point identifier is displayed in green; if the confidence level is less than the preset confidence threshold, the contact point identifier is displayed in yellow; and if the predicted contact point is determined to be invalid, the contact point identifier is displayed in red.

[0074] In some embodiments, the control method provided in this disclosure further includes: controlling a display device to display the confidence level of the predicted contact point. The confidence level of the predicted contact point is displayed on the screen or in an area of ​​the graphical user interface other than the area displaying the screen. This embodiment not only displays the contact point identifier but also the confidence level of the predicted contact point, making it easier for users to accurately adjust the position and orientation of the end effector, resulting in a better user experience.

[0075] In some embodiments, the control method provided in this disclosure further includes: determining the confidence level of the predicted contact point based on the deviation between the predicted contact point positions obtained from at least two different viewpoints. The greater the deviation between the predicted contact point positions obtained from at least two different viewpoints, the lower the confidence level of the predicted contact point; conversely, the smaller the deviation between the predicted contact point positions obtained from at least two different viewpoints, the higher the confidence level of the predicted contact point. This embodiment, based on the deviation between the predicted contact point positions obtained from at least two different viewpoints, can accurately determine the confidence level of the predicted contact point.

[0076] It should be noted that the specific method for determining the predicted contact point between the end effector and the object is the same in each viewpoint; the difference lies in the transformation matrix used for each viewpoint. For example, if the end effector is mounted on a robot, with a first camera mounted on the robot's head and a second camera mounted on its chest, and the first viewpoint of the first camera and the second viewpoint of the second camera are different, determining the predicted contact point between the end effector and the object in the first viewpoint includes: obtaining the position and orientation of the end effector in the robot coordinate system, and obtaining the first transformation matrix from the robot's base to the first camera; transforming the position and orientation of the end effector in the robot coordinate system to the coordinate system of the first camera based on the first transformation matrix; determining the first closed path of the end effector as the path from the open position of the end effector to the closing direction, closing half of the opening amplitude of the end effector along the closing direction; and determining the position of the point where the first closed path intersects the object as the first position of the predicted contact point between the end effector and the object in the first viewpoint.

[0077] Similarly, determining the position of the predicted contact point between the end effector and the object from the second perspective includes: obtaining the position and orientation of the end effector in the robot coordinate system, and obtaining the second transformation matrix from the robot's base to the second camera; transforming the position and orientation of the end effector in the robot coordinate system to the coordinate system of the second camera according to the second transformation matrix; determining the second closing path of the end effector as the path from the opening position of the end effector to the closing direction of the end effector, which is half of the opening amplitude of the end effector; and determining the position of the point where the second closing path intersects with the object as the second position of the predicted contact point between the end effector and the object from the second perspective.

[0078] In some embodiments, an end effector is mounted on a robot, which is equipped with a first camera and a second camera with different viewing angles, wherein at least two different viewing angles include the perspectives of the first camera and the second camera. For example, the first camera is mounted on the robot's head, and the second camera is mounted on the robot's chest. Another example is that the first camera is mounted on the robot's head, and the second camera is mounted on the robot's waist. Yet another example is that the first camera is mounted on the robot's head, and the second camera is mounted on the robot's robotic arm. Yet another example is that the robot's head is equipped with both the first and second cameras.

[0079] In some embodiments, an end effector is mounted on a robot, a first camera is mounted on the end effector, and a second camera is mounted on the robot. The first and second cameras have different perspectives, and at least two different perspectives include the perspectives of the first and second cameras. For example, the first camera is mounted on the end effector, and the second camera is mounted on the robot's head. Another example is that the first camera is mounted on the end effector, and the second camera is mounted on the robot's waist.

[0080] In some embodiments, the control method provided in this disclosure further includes: controlling a display device to display an orientation indicator based on the perpendicularity of the end effector's orientation to the object's surface and the end effector's orientation itself. The perpendicularity of the end effector's orientation to the object's surface is determined by the angle between the vector corresponding to the end effector's orientation and the object's surface normal vector. A smaller angle between the vector corresponding to the end effector's orientation and the object's surface normal vector indicates a higher degree of perpendicularity; conversely, a larger angle indicates a lower degree of perpendicularity. In this embodiment, the orientation indicator is displayed based on the perpendicularity of the end effector's orientation to the object's surface and the end effector's orientation itself. This allows the user to not only know the end effector's orientation but also its perpendicularity to the object's surface, making it easier for the user to precisely adjust the end effector's orientation to ensure it is perpendicular to the object's surface, thereby improving the stability of object grasping.

[0081] In some embodiments, the control method provided in this disclosure further includes: controlling a display device to display the perpendicularity of the end effector's orientation to the object's surface. The perpendicularity of the end effector's orientation to the object's surface is displayed on the screen or in an area of ​​the graphical user interface other than the area displaying the screen. This embodiment not only displays an orientation indicator but also displays the perpendicularity of the end effector's orientation to the object's surface, making it easier for users to precisely adjust the end effector's orientation so that it is perpendicular to the object's surface, thereby improving the stability of object grasping.

[0082] In some embodiments, the control method provided in this disclosure further includes: controlling a display device to display a target point corresponding to the orientation of the end effector, wherein the target point is located on an object. The method for determining the target point includes: obtaining the position and orientation of the end effector in a robot coordinate system; generating a ray along the orientation direction of the end effector in the robot coordinate system, starting from the position of the end effector in the robot coordinate system; and determining the nearest intersection point between the ray and an object in the environment of the end effector as the target point in response to the ray intersecting with an object in the environment of the end effector. This embodiment, by displaying the target point corresponding to the orientation of the end effector, allows the user to know that the end effector's orientation is aligned with the object, resulting in a better user experience.

[0083] In some embodiments, controlling the display device to display the target point corresponding to the orientation of the end effector includes: controlling the display device to mark the target point in the image. Marking the target point in the image includes: projecting the position of the target point onto the image plane corresponding to the camera based on the intrinsic and extrinsic parameter matrices of the camera capturing the image, obtaining the projected position of the target point, and marking the projected position of the target point in the image to mark the obtained target point. It should be noted that the display device can also be controlled to mark objects in the image corresponding to the orientation of the end effector.

[0084] In some embodiments, the control method provided in this disclosure further includes: controlling a display device to display the distance between the end effector and the object. The distance between the end effector and the object is displayed on the screen or in an area of ​​the graphical user interface other than the area displaying the screen. This embodiment displays the distance between the end effector and the object, facilitating safe adjustment of the end effector's position relative to the object by the user, ensuring that the end effector does not collide with the object when it approaches, thus enhancing safety.

[0085] In some embodiments, the control method provided in this disclosure further includes: controlling a display device to display a target point corresponding to the orientation of the end effector, the target point being located on an object; and controlling the display device to display the distance between the end effector and the object. This embodiment not only displays the target point on the object corresponding to the orientation of the end effector, but also displays the distance between the end effector and the object. This allows the user to know not only that the end effector is aligned with the object, but also the distance between the end effector and the object, making it easier for the user to safely adjust the position of the end effector relative to the object, resulting in a better user experience and higher safety.

[0086] In some embodiments, the control method provided in this disclosure further includes: controlling a display device to display a size matching identifier, the size matching identifier indicating the degree of matching between the opening range of the end effector and the size of the object. The size matching identifier is displayed on the screen or in an area of ​​the graphical user interface other than the area displaying the screen. In this embodiment, the user can know the degree of matching between the opening range of the end effector and the size of the object through the size matching identifier. This allows the user to adjust the opening range of the end effector or reselect an object whose size matches the opening range of the end effector when there is a mismatch between the opening range of the end effector and the size of the object, thus improving the success rate of the end effector in grasping the object.

[0087] For example, if the opening amplitude of the end effector is greater than the width of the object, the degree of matching between the end effector's opening amplitude and the object's size is high. In this case, the opening amplitude of the end effector matches the object's size, and the condition for grasping the object is met. Conversely, if the opening amplitude of the end effector is less than or equal to the width of the object, the degree of matching between the end effector's opening amplitude and the object's size is low. In this case, the opening amplitude of the end effector does not match the object's size, and the condition for grasping the object is not met.

[0088] In some embodiments, the degree of matching between the opening range of the end effector and the size of the object varies, and the visual representation of the size matching indicator differs accordingly. For example, if the opening range of the end effector matches the size of the object, the size matching indicator is displayed in green; if the opening range of the end effector does not match the size of the object, the size matching indicator is displayed in red.

[0089] In some embodiments, the control method provided in this disclosure further includes: controlling a display device to display a grasping-ready indicator, the grasping-ready indicator indicating that the end effector is in a grasping-ready state. The grasping-ready indicator is displayed on the screen or in an area of ​​the graphical user interface other than the area displaying the screen. In this embodiment, the user can intuitively know that the end effector is in a grasping-ready state through the displayed grasping-ready indicator, avoiding blind grasping and improving the success rate of the user controlling the end effector to grasp objects.

[0090] In some embodiments, the visual presentation of the gripping ready indicator differs when the end effector is in a gripping ready state from its visual presentation when the end effector is not in a gripping ready state. For example, the gripping ready indicator includes a gripping ready indicator light, which is illuminated when the end effector is in a gripping ready state and off when the end effector is not in a gripping ready state.

[0091] In some embodiments, the control method provided in this disclosure further includes: determining that the end effector is in a grasping ready state in response to the end effector satisfying a preset grasping ready condition; wherein, the end effector satisfying the preset grasping ready condition includes at least one of the following: the confidence level of the predicted contact point position is greater than or equal to a preset confidence level, the perpendicularity of the end effector's orientation to the object's surface is greater than or equal to a preset perpendicularity level, and the matching degree between the end effector's opening amplitude and the object's size is greater than or equal to a preset matching degree. It should be noted that the preset confidence level, preset perpendicularity level, and preset matching degree can be set based on actual conditions, and this disclosure does not specifically limit them.

[0092] In some embodiments, the control method provided in this disclosure further includes: controlling a display device to display grasping force prompt information. The grasping force prompt information provides a user with a prompt regarding the force required for the end effector to grasp the object. The force is estimated based on the object's material and the surface curvature of the object at the predicted contact point. This embodiment provides the user with a prompt regarding the force required for the end effector to grasp the object, facilitating the user to adjust the force required for the end effector to grasp the object, and further improving the success rate of the user controlling the end effector to grasp the object.

[0093] In some embodiments, controlling the display device to display grasping force prompt information includes: acquiring the material of the object and the surface curvature of the object at the predicted contact point; determining the force required for the end effector to grasp the object based on the material of the object and the surface curvature of the object at the predicted contact point using a preset force prediction model; and controlling the display device to display grasping force prompt information according to the force required for the end effector to grasp the object. The force prediction model is obtained by pre-training a neural network model using multiple training samples, including sample material, sample surface curvature, and labeled force.

[0094] Please see Figure 5 , Figure 5 This is a schematic block diagram of the structure of a system provided in an embodiment of this disclosure.

[0095] like Figure 5 As shown, system 100 includes a control device 110, an input device 120, and an end effector 130. The control device 110 is communicatively connected to both the input device 120 and the end effector 130. The control device 110 and the input device 120 are used to control the end effector 130 to perform grasping. The control device 110 and the input device 120 can be integrated or separate; this embodiment does not specifically limit their configuration.

[0096] Please see Figure 6 , Figure 6 This is a schematic block diagram of another system structure provided in the embodiments of this disclosure.

[0097] like Figure 6 As shown, system 100 includes a control device 110, an input device 120, and a robot 140 including an end effector. The control device 110 is communicatively connected to both the input device 120 and the robot 140. The control device 110 and the input device 120 are used to control the end effector of the robot 140 to perform grasping. The control device 110 and the input device 120 can be integrated or separate; this embodiment does not specifically limit their configuration.

[0098] It should be noted that, as those skilled in the art will clearly understand, for the sake of convenience and brevity, Figure 5 or Figure 6 The specific working process of the system described can be referred to the corresponding process in the aforementioned control method embodiments, and will not be repeated here.

[0099] The following example illustrates the control method provided in this embodiment, using an end effector mounted on a robot, a first camera mounted on the robot's head, and a second camera mounted on the robot's chest, with the first viewpoint corresponding to the first camera and the second viewpoint corresponding to the second camera being different.

[0100] This disclosure provides a control method that uses dual cameras (head and chest cameras) to construct an environmental model, calculates the position and orientation of the end effector through forward kinematics, predicts and marks the predicted contact point (contact point identifier) ​​in the image, and displays the orientation identifier of the end effector to assist the user in accurately controlling the grasping action.

[0101] Specifically, the graphical user interface displayed by the display device includes the following functional areas: 1. The field of view of the first camera (first-person perspective, i.e., the main viewpoint): 1.1. Display images captured in real time by the first camera; 1.2. The first contact point marker (circular highlighted mark) is superimposed on the image. The first contact point marker is used to indicate the first position of the predicted contact point between the end effector and the object in the first view when the end effector performs the closing action. 1.3. A first orientation marker (with an arrowed line segment) is overlaid on the image. The first orientation marker is used to indicate the orientation of the end effector in the first viewpoint. 1.4. Display the distance between the end effector and the object it is facing from the first viewpoint.

[0102] 2. The field of view of the second camera (secondary view, i.e., auxiliary view): 2.1. Display images captured in real time by the second camera; 2.2. The second contact point identifier (circular highlighted mark) is superimposed on the image. The second contact point identifier is used to indicate the second position of the predicted contact point between the end effector and the object in the second view when the end effector performs the closing action. 2.3. A second orientation marker (with an arrowed line segment) is overlaid in the image. The second orientation marker is used to indicate the orientation of the end effector in the second viewpoint. 2.4. Through items 2.1-2.3 above, the image area of ​​the second camera provides supplementary spatial information from the top / bottom view perspective.

[0103] 3. Dual-view fusion prompt area: 3.1. When the end effector performs a closing action, if the deviation between the first position of the predicted contact point with the object in the first view and the second position of the predicted contact point with the object in the second view is less than or equal to a preset deviation threshold, it is determined that the positions of the predicted contact point are consistent in the first and second views, and the first contact point identifier and the second contact point identifier are displayed in green. 3.2. When the end effector performs a closing action, if the deviation between the first position of the predicted contact point with the object in the first view and the second position of the predicted contact point with the object in the second view is greater than a preset deviation threshold, it is determined that the positions of the predicted contact point in the first view and the second view are inconsistent. The first contact point marker and the second contact point marker are displayed in yellow, and a view prompt message is displayed to prompt the user to adjust the shooting angle of the first camera and / or the second camera. 3.3. Display the confidence score of the predicted contact point, which is determined based on the deviation between the first and second positions.

[0104] The above control method can be applied to a control device, which includes an environmental modeling and end effector position calculation module, a contact point prediction and marking module, an orientation ray prompting module, and a prompt display and status management module. Environment modeling and end effector position calculation module: Inputs: Image streams captured by the first camera mounted on the robot's head, image streams captured by the second camera mounted on the robot's chest, and the angles of the robot's joints.

[0105] Step 1: Dual-view image acquisition Synchronously control the first camera (first-view) and the second camera (second-view) to acquire RGB-D images; The two acquired RGB-D images are timestamped to ensure data synchronization between the two perspectives.

[0106] Step 2: 3D Environment Reconstruction Registration and fusion of dual-view depth images; A unified 3D environmental model (point cloud or voxel mesh) is generated by stitching together point clouds. Establish the coordinate system for the environmental model.

[0107] Step 3: Position calculation of the end effector (forward kinematics) Based on the robot's forward kinematics model, the transformation matrix from the robot's base to the left and right end effectors is calculated according to the current joint angles θ. Formula: T_ee = FK(θ), where FK is the positive kinematic function; Extract the position P_ee of the end effector and the orientation vector V_ee of the left and right end effectors.

[0108] Step 4: Coordinate System 1 The transformation matrix T_base_cam1 from the base to the first camera and the transformation matrix T_base_cam2 from the base to the second camera are obtained through hand-eye calibration.

[0109] Transform the position P_ee of the end effector to the coordinate system of the first camera: P_cam1 = T_base_cam1×P_ee, and transform the position P_ee of the end effector to the coordinate system of the second camera: P_cam2 = T_base_cam2×P_ee.

[0110] Output: A 3D model of the environment in a unified coordinate system, and the position and orientation of the end effector in the dual-camera coordinate system.

[0111] Contact point prediction and marking module: Input: 3D environmental model, position and orientation of the end effector in the dual-camera coordinate system, and opening range of the end effector.

[0112] Step 1: Closure simulation of the end effector Based on the opening amplitude of the end effector, the position and orientation of the end effector in the first camera coordinate system, the closing path of the end effector is simulated; Define the gripper closure path: along the closing direction of the end effector, move a distance d inward from the position of the end effector (d is half the opening amplitude of the end effector).

[0113] Step 2: Collision Detection Searching for the nearest point on the closed path of the end effector of an object in the point cloud data of the environment point cloud. Use KD-Tree to accelerate nearest neighbor search; Judgment condition: If there exists a point Q in the point cloud that satisfies distance(Q, closed path) < ε, then it is determined to be a potential contact point.

[0114] Step 3: Predicting and Determining Contact Points The first point that satisfies the collision condition along the closed path is the predicted contact point C_pred; The predicted contact point position C_head of the end effector in the first view and the predicted contact point position C_chest of the end effector in the second view are calculated using the above method. Verify the consistency between the two perspectives: If |C_head - C_chest| < δ, then determine that the confidence of the predicted contact point is high; otherwise, mark the confidence of the predicted contact point as low.

[0115] Step 4: Image tag generation The predicted contact point position is projected onto the image plane of the first camera to obtain the first projection position: c1 = K1 × [R|t]1 × C_head, where K1 is the intrinsic parameter matrix of the first camera and [R|t]1 is the extrinsic parameter matrix of the first camera. A circular contact point marker (first contact point marker) is generated at the first projection position, and the radius is positively correlated with the fingertip size of the end effector. The predicted contact point position is projected onto the image plane of the second camera to obtain the second projection position: c2 = K2 × [R|t]2 × C_chest, where K2 is the intrinsic parameter matrix of the second camera and [R|t]2 is the extrinsic parameter matrix of the second camera. A circular contact point marker (second contact point marker) is generated at the second projection position, and the radius is positively correlated with the fingertip size of the end effector. Set the color of the contact point marker based on the confidence level: High confidence (consistent perspective): Green Low confidence (large dual-view bias): Yellow Invalid prediction: Red Output: Contact point identifier of the predicted contact point in the dual-view image, confidence score of the predicted contact point, and 3D coordinates of the predicted contact point.

[0116] Orientation Ray Indication Module: Inputs: position and orientation of the end effector, parameters of the two cameras, and environmental depth information.

[0117] Step 1: Orientation Ray Generation Starting from the position P_center of the end effector, generate a ray along the direction of the end effector's orientation vector V; Ray parameter equation: R(t) = P_center + t × V, t ∈ [0, L_max], where L_max is the maximum display distance (e.g., 1 meter).

[0118] Step 2: Calculation of Ray-Environment Intersection Point Find the nearest point on a ray in the environmental point cloud; Calculate the distance between the intersection point of the ray and the object: D_intersect = min{t | R(t) points in the nearest point cloud}; If the ray does not intersect the object, then D_intersect = L_max.

[0119] Step 3: Dual-view ray projection Transform the ray start point P_center and end point P_end = P_center + D_intersect×V to the coordinate system of the first camera to obtain P_center_cam1 and P_end_cam1; using the intrinsic parameter matrix of the first camera, project P_center_cam1 and P_end_cam1 onto the image plane of the first camera: p_start 1 = K_head × P_center_cam1, p_end 1 = K_head × P_end_cam1, where K_head is the intrinsic parameter matrix of the first camera; Transform the ray start point P_center and end point P_end = P_center + D_intersect×V to the coordinate system of the second camera to obtain P_center_cam2 and P_end_cam2; using the intrinsic parameter matrix of the second camera, project P_center_cam2 and P_end_cam2 onto the image plane of the first camera: p_start 2 = K_chest × P_center_cam2, p_end 2 = K_chest × P_end_cam2, where K_chest is the intrinsic parameter matrix of the second camera.

[0120] Step 4: Generation and Display of Ray Segments Draw the first ray segment (first orientation identifier) ​​based on p_start 1 and p_end 1, and draw the second ray segment (second orientation identifier) ​​based on p_start 2 and p_end 2; The first ray segment (first orientation marker) is superimposed on the image captured in real time by the first camera, and the second ray segment (second orientation marker) is superimposed on the image captured in real time by the second camera.

[0121] Output: Dual-view image with ray cues, ray-object distance.

[0122] Prompt Display and Status Management Module: Functional Positioning: This module presents the predicted contact points and ray segments to the operator in a visual manner, and provides auxiliary status information. This module does not involve modification of handle control commands; visual cues are for informational reference only.

[0123] Processing flow: Step 1: Dual-view image overlay display Receive the first contact point identifier and the second viewpoint contact point identifier output by the contact point prediction module. Receive the first ray segment (first orientation marker) and the second ray segment (second orientation marker) output by the ray generation module; The first contact point marker and the first ray segment (first orientation marker) are superimposed on the image captured in real time by the first camera; The second contact point marker and the second ray segment (second orientation marker) are superimposed on the image captured in real time by the second camera; Output a dual-view image with prompts.

[0124] Step 2: Calculate the ready state for capture Inputs: Confidence of the predicted contact point, angle between the ray and the surface normal vector of the object, opening amplitude of the end effector, and size of the object.

[0125] Conditional judgment (AND logic): Condition A: High confidence in predicting the contact point (the predicted contact point location is consistent in both views); Condition B: The angle between the ray and the surface normal vector of the object is less than a threshold (e.g., 15°, indicating that the end effector is oriented perpendicular to the surface of the object). Condition C: The opening amplitude of the end effector matches the size of the object (within a reasonable range); Output: Ready state (a state is ready if at least one of conditions A, B, and C is met; otherwise, it is not ready).

[0126] Step 3: Warning and Prompt Management When the confidence level of the predicted contact point is low: a yellow warning icon is displayed with the message "Adjust viewing angle recommended"; When contact point prediction is invalid: a red warning is displayed, and automatic marking of the end effector is disabled; When ready to grab: the green ready indicator light will illuminate.

[0127] Step 4: Output Status Information Output to the display device: dual-view display image, ready status, and warning messages; Handle input is directly mapped to robot control commands; this module does not modify the control commands.

[0128] The control method provided in this disclosure has the following beneficial effects: 1. Improve the success rate of grasping: By predicting the contact point from a dual-view perspective, the contact position after the end effector closes can be displayed in advance, avoiding blind grasping and improving the success rate by more than 30%.

[0129] 2. Enhanced spatial perception: The head and chest dual-view perspective provides stereoscopic spatial information, reducing the depth illusion caused by a single viewpoint.

[0130] 3. Optimize gripping posture: The orientation indicator corresponding to the end effector helps users adjust the angle of the end effector so that the orientation of the end effector is perpendicular to the object surface, thereby improving gripping stability.

[0131] 4. Improved operational efficiency: Reduced the number of trial and error attempts, shortening the time for a single grabbing operation by more than 40%.

[0132] 5. Reduce cognitive load: Visual cues are more intuitive, significantly shortening training time for new users.

[0133] 6. Enhanced prediction reliability: The dual-view consistency verification mechanism ensures the credibility of the displayed contact point identification and avoids misleading users.

[0134] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement any of the control methods provided in the specification of this disclosure.

[0135] The storage medium can be volatile or non-volatile. It can be an internal storage unit of the head-mounted display device described in the foregoing embodiments, such as the hard drive or memory of the head-mounted display device. Alternatively, it can be an external storage device of the head-mounted display device, such as a plug-in hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., provided on the head-mounted display device.

[0136] Those skilled in the art will understand that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware embodiments, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0137] It should be understood that the term "and / or" as used in this disclosure and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. It should be noted that, herein, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0138] The sequence numbers of the embodiments disclosed above are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. The above descriptions are merely specific embodiments of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this disclosure, and these modifications or substitutions should all be covered within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A control method, characterized in that, include: The control display device displays a screen showing an end effector; Based on the pose information of the end effector, a contact point identifier and / or an orientation identifier are determined, wherein the contact point identifier is used to indicate the position of the predicted contact point between the end effector and the object when the end effector performs a closing action, and the orientation identifier is used to indicate the orientation of the end effector; as well as Control the display device to overlay the contact point identifier and / or the orientation identifier on the screen.

2. The control method according to claim 1, characterized in that, Also includes: Based on the confidence level of the predicted contact point and the location of the predicted contact point, the display device is controlled to display the contact point identifier; or The display device is controlled to display the confidence level of the predicted contact point.

3. The control method according to claim 2, characterized in that, Also includes: The confidence level is determined based on the deviation between the predicted contact point positions obtained from at least two different viewpoints.

4. The control method according to claim 3, characterized in that, The end effector is mounted on the robot, a first camera is mounted on the robot's head, and a second camera is mounted on the robot's chest. The first camera and the second camera have different perspectives, and the at least two different perspectives include the perspectives of the first camera and the second camera. or The end effector is mounted on the robot, and the robot is equipped with a first camera and a second camera with different viewing angles, wherein the at least two different viewing angles include the views of the first camera and the second camera; or The end effector is mounted on the robot, and a first camera is mounted on the end effector. A second camera is mounted on the robot. The first camera and the second camera have different perspectives, and the at least two different perspectives include the perspectives of the first camera and the second camera.

5. The control method according to claim 1, characterized in that, Also includes: Based on the perpendicularity of the end effector's orientation to the surface of the object and the orientation of the end effector, the display device is controlled to display the orientation indicator; or The display device is controlled to show the degree of perpendicularity between the orientation of the end effector and the surface of the object.

6. The control method according to claim 1, characterized in that, Also includes: The display device is controlled to show the orientation of the end effector to a target point, the target point being located on the object; and / or The display device is controlled to display the distance between the end effector and the object.

7. The control method according to claim 1, characterized in that, Also includes: The display device is controlled to display a size matching identifier, which indicates the degree of matching between the opening amplitude of the end effector and the size of the object.

8. The control method according to claim 1, characterized in that, Also includes: The display device is controlled to display a grasping ready flag, which indicates that the end effector is in a grasping ready state.

9. The control method according to claim 8, characterized in that, Also includes: In response to the end effector meeting the preset grasping ready condition, it is determined that the end effector is in the grasping ready state; The end effector satisfies at least one of the following preset grasping ready conditions: the confidence level of the predicted contact point position is greater than or equal to a preset confidence level; the perpendicularity of the end effector's orientation to the surface of the object is greater than or equal to a preset perpendicularity level; and the degree of matching between the end effector's opening amplitude and the size of the object is greater than or equal to a preset matching degree.

10. The control method according to claim 1, characterized in that, Also includes: The display device is controlled to display grasping force prompt information, which is used to provide the user with a prompt of the force required for the end effector to grasp the object. The force is estimated based on the material of the object and the surface curvature of the object at the predicted contact point.

11. The control method according to any one of claims 1-10, characterized in that, The step of determining the contact point identifier and / or orientation identifier based on the pose information of the end effector includes: Based on the opening / closing state information and the pose information of the end effector, determine the closing path of the end effector; determine the point where the closing path intersects with the object as the predicted contact point; determine the contact point identifier based on the position of the predicted contact point; and / or Based on the pose information of the end effector, the orientation of the end effector is determined; based on the orientation of the end effector, the orientation identifier is determined.

12. The control method according to any one of claims 1-10, characterized in that, The control of the display device to overlay the contact point identifier and / or the orientation identifier on the screen includes: The display device is controlled to overlay the orientation indicator onto the screen; In response to the existence of the predicted contact point with the object when the end effector performs a closing action, the display device is controlled to overlay the contact point identifier on the screen.

13. A control device, characterized in that, include: At least one processor; At least one memory containing computer program code; The at least one memory, the computer program code, and the at least one processor are configured together to implement the control method according to any one of claims 1-12.

14. A system, characterized in that, include: The control device as claimed in claim 13; An input device, which is communicatively connected to the control device; as well as An end effector or a robot including an end effector, the end effector or the robot including an end effector being communicatively connected to the control device, the input device and the control device being used to control the end effector to perform grasping.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to implement the control method according to any one of claims 1-12.