Representing surface profile of target object and controlling robotic arm
By generating spheres whose density corresponds to the geometric complexity of the area to represent the surface contour of the target object and optimizing the rendering pipeline, the problems of accuracy and computational cost in the existing technology are solved, and efficient surface contour representation and real-time visualization are achieved.
Patent Information
- Application Number
- CN202510845068.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies lack adaptability to the geometric complexity of surface contours in different regions when representing the surface contours of target objects, resulting in low accuracy or excessive computational cost. In addition, the existing GUI visualization pipeline is inefficient, affecting system performance.
By generating a preset number of spheres to represent the surface contours of different areas of the target object, the density of the spheres corresponds to the geometric complexity of the area, and optimizing the rendering pipeline for efficient surface representation and rendering, taking advantage of GPU memory sharing and parallel thread processing.
The accuracy of surface contour representation is improved, the computational cost is reduced, and the rendering efficiency of the GUI is optimized, thus realizing a real-time visualization and high frame rate visualization system.
Smart Images

Figure CN120765879A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computers, for example, to a method and system for representing a surface contour of a target object, a control method and system for a robotic arm, an electronic device, a computer readable storage medium, and a computer program product, etc. BACKGROUND
[0002] Intelligent robot manipulation is crucial for enabling robotic systems to accomplish a wider range of industrial and home tasks, not just the hard-coded robot trajectory programming commonly used in high-tech, high-precision (down to the micron level) industries such as semiconductor and automobile manufacturing. This can include tasks such as picking and placing, painting, polishing, and cutting different objects placed in random positions / postures.
[0003] One of the most impressive recent advances in this field is GraspNet-1Billion: A Large-Scale Benchmark for General Object Grasping by Shanghai Jiao Tong University presented at the 2020 IEEE Conference on Computer Vision and Pattern Recognition. It allows a robotic arm to view a tabletop scene as point cloud data (requiring RGB-D data from a depth camera) and automatically determine which point in 3D space is physically suitable for grasping after being trained using 1.1 billion grasp poses. Despite this impressive achievement, this approach is relatively slow. It requires processing the initial scene (10 seconds on a notebook GPU) and determining the final pose of the gripper after a period of processing (which is not real-time and is not suitable for handling dynamic scenes). A separate classical control algorithm is then used to reach the desired pose, but this algorithm cannot avoid collisions / dynamic changes in the scene in real-time, so it is not safe for future operation with humans. Because this approach does not actually understand the scene (e.g., it does not know what it is grasping, only that it looks physically like a good grasp according to a neural network), it cannot be used to perform delicate / dexterous manipulation of objects. This approach is good at grasping random objects in any way, but as a result, it can make undesirable behavior: for example, it can choose to grasp the blade of a knife instead of the handle to grasp the knife. In the following demonstration example, GraspNet does not even realize that the banana is a graspable object. It is likely that other objects on the table would need to be cleared away first before the banana becomes a valid grasp object candidate.
[0004] Another impressive recent development in this field is OpenVLA, a vision-language-action model developed by a multi-institutional collaboration led by Stanford University. This work leverages large language models (LLMs) that have gained popularity in recent years, such as ChatGPT. This particular variant uses the Llama27Billion LLM (a limited-featured, open-source LLM designed with a low parameter count of approximately 7 billion, allowing it to run on a desktop GPU). The model takes as input a 2D RGB camera image and language instructions. These inputs are converted into "input tokens" for the LLM model. The LLM's output tokens are then de-tokenized into the desired robot actions (how much rotation / translation / grasp the gripper should perform) for a specific, previously trained robot. This cycle is repeated multiple times for each updated image after the robot action. The advantage of this model is that, because it uses an LLM at its core, it can handle more general, human-like instructions. Because the entire algorithm is based on neural networks, it must be trained on scenarios likely to be encountered during robot operation, and the robot performs poorly on out-of-distribution images not previously seen in the training set. The model was trained using 1 million demonstration data points and took 14 days to train on 64 Nvidia A100 GPUs (training data is very expensive to acquire, making training costly). Because of the use of LLM, running this method in the cloud on a remote GPU server with a large memory footprint is more appropriate. To enable the method to run on a desktop GPU with two Nvidia RTX 4090 GPUs, the model's inference accuracy was reduced, and the model ultimately ran at an inference speed of 6 Hz (too slow for real-time applications). Furthermore, the model had never seen dynamic obstacles in the training data. Nevertheless, the model is impressive because it can take human spoken language, such as "Put the eggplant in the bowl," look at a 2D RGB image, and automatically move the gripper to the eggplant's location, grab it, and place it in the bowl. It may also be inclined to grab the teapot's handle, as clear human habits / preferences are recorded in the training data. To achieve this OpenVLA model training task, a large amount of end-to-end robot training data is required, and there is no way to ensure that the robot will not collide with objects along the way, because it is a black box neural network system and may not work when the test set distribution is significantly different from the training data environment.
[0005] The two workflows described above demonstrate ML-based approaches for controlling robotic manipulators to reach a target. GraspNet-1Billion addresses the problem of grasping pose generation but does not address how to control the robot's trajectory to avoid obstacles and reach the desired target pose. OpenVLA can reach the target point while potentially avoiding collisions, provided it is extensively trained using a diverse set of artificially generated trajectory training data. However, due to the iterative use of LLM to generate labeling (which is decomposed into robot action parameters at each time step), this approach is very expensive to infer. Other more classic approaches exist. For example, recent work from Samsung relies on generating control points (trajectory proposals) and corresponding cost functions for the proposed paths (based on the computationally expensive signed distance function (SDF). This approach, called the Reactive Action Motion Planner (RAMP), is a novel variant of the older method, the Model Predicted Path Integral (MPPI). This research is able to avoid both static and dynamic obstacles when fed a 3D point cloud of objects in the environment. However, it is very slow, with trajectory updates only being performed at a rate of 3Hz (control commands are executed at 20Hz). In other words, its dynamic decision-making capabilities are very slow (3Hz). Summary of the Invention
[0006] Some embodiments of the present application overcome the defects of lack of adaptability to the geometric complexity of the surface contours of different areas when representing the surface contour of a target object, resulting in low accuracy or excessive computational cost, and provide a method for representing the surface contour of a target object, a control method for a robotic arm, a system, an electronic device, a computer-readable storage medium and a computer program product.
[0007] Some embodiments provide a method for representing a surface contour of a target object, the method comprising:
[0008] Obtaining a three-dimensional model of the target object;
[0009] Based on the three-dimensional model, a preset number of balls are generated to represent the surface contours of different regions of the target object, wherein the density of the balls in the different regions corresponds to the geometric complexity of the surface contours of the different regions.
[0010] Preferably, the different regions include a first region and a second region, and the geometric complexity of the surface profile of the first region is greater than the geometric complexity of the surface profile of the second region;
[0011] The density of the balls in the different regions corresponds to the geometric complexity of the surface profiles of the different regions, including:
[0012] The density of balls in the first area is greater than the density of balls in the second area.
[0013] Preferably, the method comprises:
[0014] determining a value of a density weight of the candidate sampling point in the different region of the target object, wherein the value of the density weight represents information of geometric complexity of a surface profile of the different region.
[0015] Preferably, the step of determining the value of the density weight of the candidate sampling point in the different region of the target object comprises:
[0016] determining a neighborhood of the candidate sampling point based on the position of the candidate sampling point and a preset value of neighborhood size;
[0017] determining the value of the density weight of the candidate sampling point based on the neighborhood.
[0018] Preferably, the step of determining the value of the density weight of the candidate sampling point based on the neighborhood comprises:
[0019] determining the value of the density weight of the candidate sampling point according to an average distance between the candidate sampling point and other candidate sampling points in the neighborhood.
[0020] Preferably, the three-dimensional model is a polygon mesh model, wherein the polygon mesh model is composed of a set of vertices, edges and faces, and the candidate sampling point is a vertex of the polygon mesh model.
[0021] The step of determining the neighborhood of the candidate sampling point based on the position of the candidate sampling point and a preset value of neighborhood size comprises:
[0022] determining, as the neighborhood of the candidate sampling point, a range in which the number of edges of polygons separated from the candidate sampling point is not greater than n, wherein the preset value of neighborhood size is n, and n is a positive integer;
[0023] or,
[0024] The three-dimensional model is a voxel model, wherein the voxel model is composed of a set of voxels, and the candidate sampling point is a center of a voxel on an outermost layer of the voxel model.
[0025] The step of determining the neighborhood of the candidate sampling point based on the position of the candidate sampling point and a preset value of neighborhood size comprises:
[0026] determining, as the neighborhood of the candidate sampling point, a range in which the distance from the candidate sampling point is not greater than r, wherein the preset value of neighborhood size is r, and r is a positive number.
[0027] Preferably, the step of generating a preset number of balls based on the three-dimensional model to represent the surface contours of different areas of the target object includes:
[0028] Determining the preset number of sampling points from all candidate sampling points based on the value of the density weight;
[0029] The preset number of spheres are generated with the determined sampling points as sphere centers.
[0030] Preferably, the step of determining the preset number of sampling points from all candidate sampling points based on the value of the density weight includes:
[0031] Determine one of the candidate sampling points as the initial sampling point;
[0032] Determine a next sampling point from the remaining candidate sampling points based on a Euclidean distance between each remaining candidate sampling point and a last candidate sampling point determined as a sampling point, and a density weight value of each remaining candidate sampling point, until the predetermined number of sampling points is determined, wherein the remaining candidate sampling points are candidate sampling points that have not yet been determined as sampling points.
[0033] Preferably, the step of determining the next sampling point from the remaining candidate sampling points according to the Euclidean distance between each remaining candidate sampling point and the last candidate sampling point determined as the sampling point, and the density weight of each remaining candidate sampling point, comprises:
[0034] Determine, based on the Euclidean distance between each remaining candidate sampling point and the last candidate sampling point determined as the sampling point, and the density weight of each remaining candidate sampling point, a value of a comprehensive distance between each remaining candidate sampling point and the last candidate sampling point determined as the sampling point;
[0035] The remaining candidate sampling point with the largest corresponding integrated distance value is determined as the next sampling point.
[0036] Preferably, the method further comprises:
[0037] Based on the three-dimensional model, the sphere radii of the preset number of spheres are determined, wherein the sphere radii of the spheres in the different regions correspond to the geometric complexity of the surface profiles of the different regions.
[0038] Preferably, the spherical radius of the balls in the different regions corresponds to the geometric complexity of the surface profiles of the different regions, including:
[0039] The spherical radius of the ball in the first area is smaller than the spherical radius of the ball in the second area.
[0040] Preferably, the sampling points include a first sampling point and a second sampling point;
[0041] The step of determining the spherical radius of the preset number of balls based on the three-dimensional model includes:
[0042] The spherical radius of each sphere is determined according to a preset radius and a density weight of a sampling point corresponding to each sphere, wherein when the density weight of the first sampling point is greater than the density weight of the second sampling point, the spherical radius of the sphere corresponding to the first sampling point is smaller than the spherical radius of the sphere corresponding to the second sampling point.
[0043] Preferably, the method further comprises:
[0044] Normalizing the density weight values of the sampling points corresponding to each sphere to obtain the normalized density weight values of the sampling points corresponding to each sphere, wherein the normalized density weight values are positive numbers less than 1.
[0045] Preferably, the step of determining the value of the spherical radius of each sphere according to the value of the preset radius and the value of the density weight of the sampling point corresponding to each sphere includes:
[0046] The value of the spherical radius of each sphere is determined according to the value of the preset radius and the value of the normalized density weight of the sampling point corresponding to each sphere.
[0047] Preferably, the step of determining the value of the spherical radius of each sphere according to the value of the preset radius and the value of the normalized density weight of the sampling point corresponding to each sphere includes:
[0048] The spherical radius of each sphere is determined according to the value of the preset radius, the value of the normalized density weight of the sampling point corresponding to each sphere, and the value of the proportional coefficient, wherein the proportional coefficient is used to control the sensitivity of the spherical radius of each sphere to changes in the value of the normalized density weight of the sampling point corresponding to each sphere.
[0049] Preferably, the method further comprises:
[0050] Generate descriptor information of the preset number of spheres, wherein the descriptor information includes at least one of index information, position information, normal information and sphere radius, the index information is an index list, and the index list is an index list of candidate sampling points {1, 2, 3, ..., n IH} subset, n IH is the total number of balls representing the surface contour of the target object.
[0051] Preferably, the method further comprises:
[0052] Obtain the location information of all candidate sampling points in real time;
[0053] The position information of all sampling points is determined according to the position information of all candidate sampling points and the index information.
[0054] Preferably, the step of obtaining the three-dimensional model of the target object includes:
[0055] Collecting a two-dimensional image of the space where the target object is located in real time;
[0056] Inputting the two-dimensional image into an open vocabulary object detection model, identifying the target object in the two-dimensional image using a vocabulary corresponding to the target object, and determining the position and size of the target object in the two-dimensional image;
[0057] generating a bounding box in the two-dimensional image according to the position and size, so as to frame the target object in the bounding box;
[0058] A three-dimensional model of the target object is obtained according to the two-dimensional image with the bounding box.
[0059] Preferably, the target object is a hand,
[0060] The step of obtaining the three-dimensional model of the target object according to the two-dimensional image with the bounding box comprises:
[0061] Determining the 2D hand joints and 3D hand vertex positions of the target object respectively;
[0062] The three-dimensional model is obtained according to the 2D hand joints and the 3D hand vertex positions.
[0063] The present invention also provides a system for representing a surface contour of a target object, the system comprising:
[0064] One or more processors, wherein the one or more processors are configured to execute a computer program to implement the aforementioned method for representing the surface contour of a target object.
[0065] Preferably, the system further comprises:
[0066] An image acquisition module, which is used to acquire a two-dimensional image of the space where the target object is located in real time;
[0067] The one or more processors are further configured to receive the two-dimensional image sent by the image acquisition module and execute a computer program to implement the aforementioned method for representing the surface contour of the target object.
[0068] Some embodiments further provide a method for controlling a robotic arm, wherein the robotic arm includes a base, at least one connecting rod, and a joint connecting the at least one connecting rod and the base, the at least one connecting rod including an end effector located at an end of the robotic arm, the control method comprising:
[0069] Representing the surface contour of the target operation object according to the aforementioned method;
[0070] Selecting one of the balls representing the surface contour of the target operation object as a target ball;
[0071] Acquire a first coordinate, wherein the first coordinate is a coordinate representing the posture of the center of the target ball in a world coordinate system;
[0072] Acquire a second coordinate, wherein the second coordinate is a coordinate representing a real-time posture of the end effector in the world coordinate system;
[0073] The movement of the robotic arm is controlled according to the difference between the first coordinate and the second coordinate.
[0074] Preferably, the step of controlling the movement of the robotic arm according to the difference between the first coordinate and the second coordinate comprises:
[0075] Obtaining an initial value of the joint angle θ of the at least one link,
[0076] Constructing an objective function, wherein the objective function includes a first component, the first component represents a difference between the first coordinate and the second coordinate, and the objective function uses the joint angle θ as a variable;
[0077] The value of the joint angle θ is updated by a gradient descent algorithm.
[0078] Preferably, there is at least one obstacle in the space where the robotic arm is located, and the control method includes:
[0079] Representing the at least one obstacle and the robotic arm according to the aforementioned method, wherein the surface contour of the at least one obstacle is represented by M balls, and the surface contour of the robotic arm is represented by N balls;
[0080] Determine all ball pairs, wherein the ball pair consists of an i-th ball representing the surface profile of the at least one obstacle and a j-th ball representing the surface profile of the robotic arm, where i is a positive integer not greater than M and i is a positive integer not greater than N;
[0081] Determining a first distance value for each ball pair, wherein the first distance is a distance value between a center of a ball representing a surface profile of the at least one obstacle and a center of a ball representing a surface profile of the robotic arm in each ball pair;
[0082] Determining a value of a sum of sphere radii of each sphere pair, wherein the value of the sum of sphere radii is the sum of the values of the sphere radii of the sphere representing the surface profile of the at least one obstacle and the values of the sphere radii of the sphere representing the surface profile of the robotic arm in each sphere pair;
[0083] Determining the spherical distance of each ball pair according to the value of the first distance of each ball pair and the value of the sum of the sphere radius;
[0084] determining a first collision probability value of each ball pair according to a difference between the spherical surface distance and a preset threshold, wherein the first collision probability represents a probability of collision between the two balls in each ball pair;
[0085] Determine a value of a second probability that the robotic arm collides with the at least one obstacle, wherein the value of the second probability is the sum of the values of the first collision probabilities of all ball pairs
[0086] The movement of the robot arm is controlled according to the value of the second possibility.
[0087] Preferably, the objective function comprises the second component, wherein the second component represents the value of the second likelihood.
[0088] Preferably, the step of obtaining the first coordinate includes:
[0089] Acquire a third coordinate, wherein the third coordinate is a coordinate representing the posture of the center of the target ball in the target operation object coordinate system;
[0090] Obtaining the 6D pose of the target operation object;
[0091] Converting the third coordinate into the first coordinate according to the 6D pose of the target operation object;
[0092] The step of obtaining the second coordinate comprises:
[0093] Acquiring a fourth coordinate, wherein the fourth coordinate is a coordinate representing a posture of the end effector in an end effector coordinate system;
[0094] Acquiring forward kinematics information of the robotic arm, wherein the forward kinematics information includes joint translation and rotation angles;
[0095] constructing, according to the forward kinematics information, a corresponding forward kinematics function for each link, wherein the forward kinematics function takes joint angle θ as a variable for transforming coordinates in the link coordinate system to the world coordinate system;
[0096] obtaining a 6D pose of the base;
[0097] transforming the fourth coordinates into the second coordinates according to the forward kinematics function and the 6D pose of the base.
[0098] Preferably, the control method further comprises:
[0099] obtaining fifth coordinates, wherein the fifth coordinates are coordinates representing poses of the ball representing the surface profile of the at least one obstacle in each ball pair in an at least one obstacle coordinate system;
[0100] obtaining a 6D pose of the at least one obstacle;
[0101] transforming the fifth coordinates into sixth coordinates according to the 6D pose of the at least one obstacle, wherein the sixth coordinates are coordinates representing poses of the ball representing the surface profile of the at least one obstacle in each ball pair in a world coordinate system;
[0102] obtaining seventh coordinates, wherein the seventh coordinates are coordinates representing poses of the ball representing the surface profile of the robot arm in each ball pair in a corresponding link coordinate system;
[0103] obtaining forward kinematics information of the robot arm, wherein the forward kinematics information comprises joint translation and rotation angles;
[0104] constructing, according to the forward kinematics information, a corresponding forward kinematics function for each link, wherein the forward kinematics function takes joint angle θ as a variable for transforming coordinates in the link coordinate system to the world coordinate system;
[0105] obtaining a 6D pose of the base;
[0106] transforming the seventh coordinates into eighth coordinates according to the forward kinematics function and the 6D pose of the base, wherein the eighth coordinates are coordinates representing poses of the ball representing the surface profile of the robot arm in each ball pair in a world coordinate system.
[0107] Some embodiments also provide a control system of a robot arm, the control system comprising: one or more processors configured to execute a computer program to implement the control method of the robot arm as described above.
[0108] Some embodiments also provide an electronic device comprising a memory, a processor, and a computer program stored in the memory and for running on the processor, wherein when the processor executes the computer program, the aforementioned method for representing the surface contour of a target object or the aforementioned method for controlling a robotic arm is implemented.
[0109] Some embodiments further provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the aforementioned method for representing the surface contour of a target object or the aforementioned method for controlling a robotic arm.
[0110] Some embodiments further provide a computer program product, including a computer program, which, when executed by a processor, implements the aforementioned method for representing the surface contour of a target object or the aforementioned method for controlling a robotic arm.
[0111] Some embodiments further provide a computer program, which, when executed by one or more processors, can implement the aforementioned method for representing the surface contour of a target object or the aforementioned method for controlling a robotic arm.
[0112] The positive effect of some embodiments is that: a method for representing the surface contour of a target object is provided. This method represents the surface contour of the target object as a preset number of spheres based on the geometric complexity of the surface contours of different areas of the object, thereby improving the accuracy of representing the surface contour of the target object and reducing the computational cost of representing the surface contour of the target object. BRIEF DESCRIPTION OF THE DRAWINGS
[0113] Figure 1 A schematic flow chart of a method for representing the surface profile of a target object provided in embodiment 1 of the present disclosure.
[0114] Figure 2 A schematic diagram of a neighborhood provided for some embodiments of the present disclosure.
[0115] Figure 3 An overlapping image of a 3D model of a teapot and a ball representing the surface contour of the teapot in a static pose frame is provided for some embodiments of the present disclosure.
[0116] Figure 4 A schematic diagram of voxel-based sampling provided for some embodiments of the present disclosure.
[0117] Figure 5 A schematic diagram illustrating an example of the position of a sphere used to represent the surface contour of an object when superimposed on an actual underlying physical object in world coordinates, provided for some embodiments of the present disclosure.
[0118] Figure 6 A schematic diagram of a process for obtaining a three-dimensional model of a target object provided in some embodiments of the present disclosure.
[0119] Figure 7 A schematic diagram of a process for obtaining a real-time three-dimensional model of a hand provided in some embodiments of the present disclosure.
[0120] Figure 8 A schematic diagram of a mesh model of a hand provided for some embodiments of the present disclosure.
[0121] Figure 9 A schematic diagram of a system 800 for representing a surface profile of a target object provided in some embodiments of the present disclosure.
[0122] Figure 10 A flow chart of the control method of the robotic arm provided in Example 3 of the present disclosure.
[0123] Figure 11 A schematic diagram of an electronic device provided in Example 5 of the present disclosure. DETAILED DESCRIPTION
[0124] The present invention is further described below by way of examples, but the present invention is not limited to the scope of the examples.
[0125] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various places herein does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0126] It should be understood that the terms "device," "system," "unit," and / or "module" used herein are a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. However, if other terms can achieve the same purpose, the terms may be replaced by other expressions.
[0127] As used herein, unless the context clearly indicates otherwise, the terms "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include additional steps or elements.
[0128] The terms "have", "may have", "include", or "may include" as used herein indicate the presence of the corresponding function, operation, element, etc. described herein and do not limit the existence of other function, operation, element, etc. In addition, it should be understood that the terms "include" or "have" as used herein indicate the presence of the features, numbers, steps, operations, elements, components, or combinations thereof described in the specification, without excluding the presence or addition of one or more other features, numbers, steps, operations, elements, components, or combinations thereof.
[0129] Flowcharts herein are used to illustrate the operations performed by the systems according to the embodiments herein. It should be understood that the preceding or subsequent operations are not necessarily performed in sequence. Instead, the steps can be processed in reverse order or simultaneously. Meanwhile, other operations can be added to these processes, or one or more steps can be removed from these processes.
[0130] In some examples of the present disclosure, the method steps described below can be implemented by a program or a custom circuit, or a combination of a custom circuit and a program, for example, the method steps described below can be executed by one or more GPUs, one or more CPUs or any other technically feasible combination of one or more processors such as ASIC chips. In some examples of the present disclosure, the one or more processors such as ASIC chips can include a memory. Further, those skilled in the art can understand that the scope of protection of the present disclosure can include any system and device that can execute the method steps described below.
[0131] Embodiment 1
[0132] Figure 1 A method for representing a surface profile of a target object is shown in an example embodiment of the present disclosure, the method comprising the steps of:
[0133] S101, obtaining a three-dimensional model of the target object.
[0134] S102, generating a preset number of spheres based on the three-dimensional model to represent surface profiles of different regions of the target object.
[0135] The density of the spheres in the different regions corresponds to the geometric complexity of the surface profiles of the different regions.
[0136] The inventors have found that although the use of simplified geometric primitives (e.g. spheres) to represent three-dimensional surfaces has been applied in applications requiring collision detection, conventional techniques rely on uniform sampling or simple geometric approximations, which lack adaptability to local surface features such as curvature and vertex density that represent the geometric complexity of the surface profile. This can result in poor accuracy or excessive computational overhead when modeling complex geometries.
[0137] In some embodiments of the present disclosure, a preset number of balls are generated based on a three-dimensional model to represent the surface contours of different areas of the target object, and the density of the balls in different areas corresponds to the geometric complexity of the surface contours of different areas, thereby achieving accurate modeling of complex geometric bodies and balancing the computational overhead.
[0138] Graphical User Interface (GUI)
[0139] The graphical user interface (GUI) allows users to easily and intuitively control the robot. The GUI displays video from each camera and a rendering of the 3D scene. To control the robot, users simply select commands from the toolbar and click on the interactive video panel.
[0140] The inventors found that the initial implementation of the current GUI visualization pipeline was inefficient, severely impacting the frames per second (FPS) of the entire system, including tasks such as mask segmentation, pose estimation, and control processing. The main bottleneck was inefficient data transfer between the CPU and GPU, and frequent memory copying led to latency spikes and performance degradation. The inventors also found that the rendering pipeline ran in synchronous blocking mode, preventing other processing threads from executing in parallel and causing further delays. Using an unoptimized GUI framework exacerbated these problems because it failed to take advantage of modern GPU acceleration technologies (such as direct GPU memory sharing), resulting in excessive CPU utilization and slower rendering times for scene meshes and surface spheres.
[0141] Using the sphere generated in some embodiments of the present disclosure to represent the surface contours of the target object can optimize the rendering pipeline. For example, memory sharing can be achieved by utilizing the PyTorch-OpenGL integrated GPU. This method implements zero-copy data sharing between PyTorch tensors and OpenGL buffers, eliminating redundant memory transfers and enabling direct GPU rendering of visual output. The solutions provided in some embodiments of the present disclosure can implement custom OpenGL shaders, thereby enabling transformation and lighting effects to be processed directly on the GPU, reducing the computational burden on the CPU. The solutions provided in some embodiments of the present disclosure can utilize parallel threads to separate rendering tasks from data processing, ensuring that GUI updates no longer hinder critical system operations. In addition, the frame rate cap mechanism dynamically adjusts GUI updates based on system load to maintain FPS stability. Further optimizations include lazy loading and caching strategies for static meshes to reduce redundant calculations and memory usage. These enhancements together create a scalable, responsive, and real-time visualization system that can run seamlessly with other computationally intensive tasks.
[0142] In an optional embodiment, the different regions include a first region and a second region, and the geometric complexity of the surface profile of the first region is greater than the geometric complexity of the surface profile of the second region; the density of balls in different regions corresponds to the geometric complexity of the surface profile of different regions, including: the density of balls in the first region is greater than the density of balls in the second region.
[0143] In an optional embodiment, the method further includes:
[0144] S103: Determine density weight values of candidate sampling points in different areas of the target object.
[0145] The density weight value represents the geometric complexity of the surface contours of different regions.
[0146] In an optional implementation, step S103 includes:
[0147] S1031 : Determine a neighborhood of the candidate sampling point based on the position of the candidate sampling point and a preset neighborhood size value.
[0148] S1032: Determine the density weight value of the candidate sampling point based on the neighborhood.
[0149] In an optional implementation, step S1032 includes: determining a density weight value of the candidate sampling point according to an average distance between the candidate sampling point and other candidate sampling points in the neighborhood.
[0150] Specifically, the density weight of the candidate sampling point can be determined according to the following formula:
[0151]
[0152] Among them, w i is the density weight; is the average distance; ∈ is a constant.
[0153] Specifically, ∈ is a small constant that prevents the denominator of the above formula from being 0, resulting in w i The numerical value of is unstable. From the above formula, we can see that w i along with decreases and increases, because represents the average distance between the candidate sampling point and other candidate sampling points in the neighborhood, The distribution density of candidate sampling points in the neighborhood increases, so the above formula assigns a higher density weight to the area of candidate sampling points, ensuring that when determining sampling points from the candidate sampling points, priority is given to determining sampling points in areas with high geometric complexity of the surface contour.
[0154] The average distance can be determined by the following formula
[0155]
[0156] in, is the average distance, is the set of other vertices in the neighborhood, p i is the candidate sampling point as the neighborhood center; p j are other candidate sampling points in the neighborhood.
[0157] In an optional implementation, the three-dimensional model is a polygonal mesh model, wherein the polygonal mesh model is composed of a set of vertices, edges, and faces, and the candidate sampling points are vertices of the polygonal mesh model.
[0158] Step S1031 includes:
[0159] With the candidate sampling point as the center, a range of polygons separated from the candidate sampling point and having a number of edges no greater than n is determined as a neighborhood of the candidate sampling point, wherein the preset neighborhood size is n, and n is a positive integer.
[0160] Computer Graphics
[0161] Computer graphics is the field of computer-based image rendering techniques that generate photo-realistic two-dimensional images. In this field, three-dimensional objects are represented using a three-dimensional object model, which is typically composed of a point cloud triangulated mesh represented by the coordinates of vertices and triangular faces, in addition to a color texture map of the model. A virtual camera and lighting model are used to mathematically render the three-dimensional object into a corresponding 2D image. The rendering methods themselves range from computationally inexpensive methods (such as rasterization) to computationally expensive methods (such as ray tracing). In some embodiments of the present disclosure, rasterization-based image rendering techniques may be used.
[0162] Specifically, in computer graphics, polygonal mesh models define the shape and outline of each 3D object. Polygonal meshes are built from smaller interconnected planes (usually triangles or rectangles) that act like a 3D puzzle. Figure 1 Each vertex in a polygon mesh stores x, y, and z coordinate information. Each face of that polygon then contains surface information.
[0163] Neighborhood
[0164] A region with a preset neighborhood size of n is called an n-ring neighborhood, which defines the area around a vertex based on the connectivity of adjacent faces. Figure 2 As shown, vertex p iThe n-ring neighborhood of the vertex p i Starting from the grid edge, all vertices that can be reached within n traversal steps, for example, the 1-ring neighborhood consists of the vertices with the vertex p i The 2-ring neighborhood consists of directly connected vertices (sharing an edge), and the 2-ring neighborhood consists of the vertices with the vertex p i The number of sides between which the polygon is separated is not greater than 2.
[0165] The parameter n, which controls the size of the n-ring neighborhood, is configurable. Larger values of n expand the neighborhood, capturing a wider range of surface features, while smaller values of n reduce the neighborhood, focusing on local details.
[0166] The adjacency relationship of the n-ring neighborhood can be calculated using the sparse adjacency matrix A. The specific formula is as follows:
[0167]
[0168] The expansion of the neighborhood is achieved by iteratively multiplying the adjacency matrix. The specific formula is as follows:
[0169]
[0170] Among them, N i (k) is a vertex in the k-ring neighborhood, A k represents the kth power of matrix A, δ i is the vector representing the initial vertex.
[0171] By changing the parameter n that controls the size of the n-ring neighborhood, the scope of the analysis can be set to adapt to local or global geometry. Smaller ring neighborhoods limit the calculation to a closer neighborhood and retain finer details, while larger ring neighborhoods include a wider surface environment and can capture global structure at the expense of higher computational cost, thus balancing detail preservation and computational efficiency.
[0172] In an optional embodiment, the three-dimensional model is a voxel model.
[0173] The voxel model is composed of a set of voxels, and the candidate sampling point is the center of the outermost voxel of the voxel model.
[0174] Step S1031 includes:
[0175] With the candidate sampling point as the center, the range with a distance no greater than r from the candidate sampling point is determined as the neighborhood of the candidate sampling point.
[0176] The preset neighborhood size value is r, and r is a positive number.
[0177] Specifically, the voxel model simulates the surface geometry of the target object based on voxels, discretizing the 3D object into a structured grid of cubic units (voxels). The resolution of the voxel model is controlled by the voxel_resolution parameter, which determines the number of divisions along the longest axis of the bounding box. Higher resolution can provide finer sampling, but at the cost of increased computational complexity.
[0178] In some embodiments, a polygonal mesh model can be used to generate a voxel model. For example, the polygonal mesh model can be converted to a voxel model before sampling. Given that voxels in a voxel model are typically very dense, in some embodiments, only a portion of the surface voxels (e.g., the outermost voxels) can be used for further sampling. In some embodiments, when the target object whose surface contour is represented by a sphere is a flexible object (e.g., a human hand), the vertex indices of the polygonal mesh model can be used for sampling without converting the polygonal mesh model to a voxel model.
[0179] In an optional implementation, step S102 includes:
[0180] S1021: Determine a preset number of sampling points from all candidate sampling points based on the density weight value.
[0181] S1022: Generate a preset number of spheres with the determined sampling point as the sphere center.
[0182] In an optional implementation, step S1021 includes:
[0183] S10211. Determine one of the candidate sampling points as the initial sampling point.
[0184] S10212: Determine the next sampling point from the remaining candidate sampling points according to the Euclidean distance between each remaining candidate sampling point and the last candidate sampling point determined as the sampling point, and the density weight of each remaining candidate sampling point, until the predetermined number of sampling points is determined.
[0185] The remaining candidate sampling points are candidate sampling points that have not yet been determined as sampling points.
[0186] Density-Weighted Farthest Point Sampling
[0187] In an optional implementation, step S10212 includes: determining a comprehensive distance value between each remaining candidate sampling point and the candidate sampling point last determined as the sampling point based on the Euclidean distance value between each remaining candidate sampling point and the candidate sampling point last determined as the sampling point, and the density weight value of each remaining candidate sampling point; and determining the remaining candidate sampling point having the largest corresponding comprehensive distance value as the next sampling point.
[0188] For example, the weighted distance can be determined according to the following formula:
[0189]
[0190] Among them, d′(p i ,p j ) Alternative sampling point p i With the last determined sampling point p j The comprehensive distance between i ,p j ) is the alternative sampling point p i With the last determined sampling point p j The Euclidean distance between is the alternative sampling point p i The density weight of .
[0191] j next =max i (d′(p i ,p j ))
[0192] The next sampling point j determined or selected next The maximum comprehensive distance value d′(p i ,p j ) of the sampling point p i .
[0193] The above formula is the formula of the DWFPS (Density-Weighted Farthest Point Sampling) strategy. The DWFPS strategy iteratively selects the vertex farthest from the last determined center of the sphere, and combines the density weight to bias the process of selecting the next center of the sphere towards areas with higher vertex distribution density (that is, areas with more complex geometry). The DWFPS strategy ensures that the sampling frequency is higher in areas with high vertex distribution density, and the sampling frequency is lower in areas with low vertex distribution density. Therefore, this method realizes adaptive sampling, which can effectively capture fine details and large-scale structures without redundant calculations. The DWFPS strategy can ensure uniform spatial coverage while giving priority to areas with higher complexity.
[0194] Unlike the SDF (Signed Distance Function) method, which uses a lot of memory because it generates a surface distance field function that fits any voxel in the 3D space of interest, the sphere-based calculations in this disclosure are performed pairwise between only a small number of spheres represented as a 3D point cloud. Consequently, the disclosed method uses very little CPU / GPU memory, yet is sufficient to generate cost / reward functions that can significantly alter the robot's motion.
[0195] Adaptive sphere radius for area surface contour changes
[0196] Optionally, in some embodiments of the present disclosure, the size of the ball may be adjusted accordingly according to changes in the surface contour of the target object. Figure 3 This process is illustrated using, for example, an overlay image of a 3D model of a teapot and a ball representing the surface contour of the teapot in a still pose frame. Figure 3 (a) shows the use of an adaptive sphere radius that can be adjusted based on the geometric complexity of the local area. For comparison, Figure 3 (b) shows the representation of the fixed sphere radius.
[0197] In an optional embodiment, the method further includes:
[0198] S104: Determine the sphere radius of a preset number of balls based on the three-dimensional model.
[0199] The sphere radius of the sphere in different regions corresponds to the geometric complexity of the surface profile of the different regions.
[0200] In an optional embodiment, the spherical radius of the balls in different regions corresponds to the geometric complexity of the surface profiles of the different regions, including: the spherical radius of the balls in the first region is smaller than the spherical radius of the balls in the second region. Figure 3 As shown in (a), the radius of the sphere in the area 310 of the teapot is smaller than the radius of the sphere in the area 312, and the radius of the sphere in the area 314 of the teapot is smaller than the radius of the sphere in the area 310.
[0201] In an optional implementation, the sampling points include a first sampling point and a second sampling point.
[0202] Step S104 may include:
[0203] S1041 : Determine the radius of each sphere according to the preset radius and the density weight of the sampling points corresponding to each sphere.
[0204] When the density weight of the first sampling point is greater than the density weight of the second sampling point, the radius of the sphere corresponding to the first sampling point is smaller than the radius of the sphere corresponding to the second sampling point.
[0205] In an optional embodiment, the method further includes:
[0206] S105 : Normalize the density weight value of the sampling point corresponding to each sphere to obtain the normalized density weight value of the sampling point corresponding to each sphere.
[0207] The value of the normalized density weight is a positive number less than 1.
[0208] Specifically, to ensure that the value of the density weight remains bounded, the density weight can be normalized using the following formula:
[0209]
[0210] Among them, w i is the density weight, is the normalized density weight.
[0211] Through the above normalization process, the value of the density weight can be mapped to a range between 0 and 1, thereby preventing extreme changes in the value of the density weight and ensuring the stability of the value of the density weight.
[0212] In an optional implementation, step S1041 includes:
[0213] S10411. Determine the value of the spherical radius of each sphere according to the value of the preset radius and the value of the normalized density weight of the sampling point corresponding to each sphere.
[0214] In an optional implementation, S10411 includes:
[0215] The value of the spherical radius of each sphere is determined according to the value of the preset radius, the value of the normalized density weight of the sampling point corresponding to each sphere, and the value of the proportional coefficient.
[0216] The proportional coefficient is used to control the sensitivity of the value of the sphere radius of each sphere to the change of the value of the normalized density weight of the sampling point corresponding to each sphere.
[0217] Specifically, the value of the sphere radius can be determined according to the following formula:
[0218]
[0219] Among them, r0 is the preset radius, which can be defined by the user. is the normalized density weight, and s is the proportional coefficient.
[0220] The scaling factor s is used to control the sensitivity of the sphere radius to changes in the value of the density weight. x Ensure that the sphere radius r i The scale factor s can change smoothly and continuously without sudden jumps, allowing the sphere to seamlessly adapt to local geometric changes. In some embodiments, the scale factor s can be set to 10 and the result can be visually inspected by overlaying the generated sphere with the polygonal mesh model of the target object. If the sphere is too small, the scale factor s can be increased. If the sphere is too large, the scale factor s can be decreased. In some embodiments, the scale factor s can be 0.1, 1, 10, 100, 1000, or any value between 0.1 and 1000.
[0221] Through the above formula, the sphere used to represent the surface contour of different areas of the target object can shrink in areas where the candidate sampling points are densely distributed, and expand in areas where the candidate sampling points are sparsely distributed, thereby achieving adaptive coverage without introducing unnecessary computational costs.
[0222] Voxel-based sampling
[0223] In some embodiments, the 3D model may be, for example, a voxel model. In cases where the 3D model is a voxel model, the density weights may be normalized according to the above formula, and the sphere radius may be determined based on the normalized density weights and the above formula. Because the sphere radius scales with the local voxel density, it allows for finer detail in high-density areas and larger spheres in sparse areas. In cases where uniformity is prioritized over adaptability, a fixed sphere radius may be used instead of determining the sphere radius using the above method.
[0224] Using a voxel model as a 3D model has the following advantages over using a polygonal mesh model as a 3D model:
[0225] 1. Works independently of the mesh topology, allowing the ball to be placed outside the mesh to capture volume features and discontinuous areas.
[0226] 2. Provides uniform coverage, making it effective for sparse or low-resolution meshes.
[0227] Figure 4 A voxel-based sampling process is shown, where the object is subdivided into a voxel grid and the sphere is positioned at the center of the voxel using DWFPS.
[0228] In an optional embodiment, the method further includes:
[0229] S106: Generate descriptor information of a preset number of balls.
[0230] The descriptor information includes at least one of index information, position information, normal information and sphere radius. The index information is an index list, which is an index list of candidate sampling points {1, 2, 3, ..., n IH} subset, n IH is the total number of balls representing the surface contour of the target object.
[0231] Specifically, these descriptor information can be seamlessly integrated with downstream tasks such as collision detection systems.
[0232] In an optional embodiment, the method further includes:
[0233] S107. Obtain the location information of all candidate sampling points in real time.
[0234] S108: Determine the location information of all sampling points according to the location information and index information of all candidate sampling points.
[0235] Specifically, once a specific set of hand mesh vertex indices I is selected H (i.e. the complete hand mesh vertex index {1,2,3,…,n IH}Integer index i within the subset H List, n IH The total number of hand mesh vertices) is used as the mounting point of the hand surface sphere (this set of indexes only needs to be generated once and can be reused continuously), and the position of the hand surface sphere can be continuously calculated in real time. The complete hand mesh vertices generated by the simpleHand model in the world coordinate system are represented as V hand , the center position of the wrist surface sphere in the world coordinate system is expressed as S hand , then S can be easily computed at any time using array index selection operations hand =V hand [I H ].
[0236] Given the surface sphere representations generated for robotic components, rigid and flexible objects (such as human hands) in the world coordinate system, they can be directly used for different types of real-time robotic control operations. Figure 5 Shows an example of the position in world coordinates of a sphere representing the surface outline of an object when overlaid with the actual underlying physical object. Figure 5 The scene shows spheres representing the surface contours of rigid objects, robotic arms, and human hands, all generated in real time. Figure 5 (a) shows a three-dimensional scene including various rigid objects, a robotic arm 510 and a human hand model 520. Figure 5 (b) shows the scene superimposed with the ball 550, and the normal vector 560 of the ball.
[0237] Optionally, some embodiments provide a method S200 for obtaining a three-dimensional model of a target object. Figure 6 As shown, the method S200 includes:
[0238] S210: Obtain real-time video stream data. For example, a real-time video stream is obtained by capturing a camera. A two-dimensional image of the target object exists in the video stream, so the two-dimensional image of the target object can be captured in real time.
[0239] S220 : Generate a three-dimensional model of the target object according to the two-dimensional image of the target object.
[0240] Those skilled in the art will appreciate that there are different ways to generate a three-dimensional model of a target object based on a two-dimensional image of the target object. For example, in some embodiments, the following method may be used to generate a three-dimensional model.
[0241] S222. Determine a bounding box for the target object in the two-dimensional image. The bounding box for the target object in the two-dimensional image can be determined based on an open vocabulary object detection model and input information corresponding to the target object. For example, the two-dimensional image can be input into an open vocabulary object detection model (e.g., "hand"). The target object can be identified in the two-dimensional image using the vocabulary corresponding to the target object, and the position and size of the target object in the two-dimensional image can be determined. Based on the position and size, a bounding box is generated in the two-dimensional image to enclose the target object in the bounding box. The bounding box can indicate the position and size of the target object in the two-dimensional image.
[0242] S224 : Generate a three-dimensional model of the target object based on the two-dimensional image and the bounding box of the target object.
[0243] In some embodiments, the target object may be, for example, a hand. Hereinafter, an example will be given in which the target object is a hand.
[0244] In some embodiments, step S224 may include:
[0245] S226: Determine the 2D hand joints and 3D hand vertex positions of the target object.
[0246] S228. Obtain a three-dimensional model based on the 2D hand joints and the 3D hand vertex positions.
[0247] like Figure 7 An example of constructing a three-dimensional model of a hand from a camera image in real time is shown, wherein the three-dimensional model may be, for example, a polygonal mesh model.
[0248] Step 710: Input an image. For example, an RGB image captured by a camera can be passed to an open-vocabulary object detection model, such as YOLO-World or GroundingDINO. The open-vocabulary object detection model YOLO-World is described in detail in the document “Cheng, Tianheng, Lin Song, Yixiao Ge, Wenyu Liu, Xinggang Wang, and Ying Shan. "Yolo-world: Real-time open-vocabulary object detection." In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 16901-16911.2024,” the entire disclosure of which is incorporated into this disclosure. The document “Ren, Tianhe, Qing Jiang, Shilong Liu, Zhaoyang Zeng, Wenlong Liu, Han Gao, Hongjie Huang et al.”Grounding DINO 1.5: Advance the “Edge” of Open-Set Object Detection.” arXiv preprint arXiv:2405.10300(2024)” describes the open-set object detection model GroundingDINO in detail, and all the disclosed contents are incorporated into the present disclosure.
[0249] Step 720: Object detection. For example, an open vocabulary object detection model can be used to detect the possible location and size of a hand, i.e., a bounding box of a hand. Alternatively, the open vocabulary object detection model can use the keyword "hand" as input to locate and detect hands in the input image. These open vocabulary object detection models can detect a variety of object categories, including hands, without the need for retraining.
[0250] Step 730: Joint and vertex estimation. For example, the bounding box labeled image can be processed by the simpleHand model (consisting of a label generator and a mesh regressor) to estimate the 2D hand joints and 3D hand vertex positions, respectively. The simpleHand model is described in detail in the paper “Zhishan Zhou, Shihao Zhou, Zhi Lv, Minqiang Zou, Yao Tang, and Jiajun Liang. A simple baseline for efficient hand mesh reconstruction. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1367–1376, June 2024,” the entire disclosure of which is incorporated into this disclosure. Furthermore, the local camera coordinates of the hand vertices can be converted to 3D world coordinates together with the depth information.
[0251] Step 740: Mesh reconstruction. For example, a hand mesh can be created based on the hand vertices obtained in step 730 and a triangle mesh generated using a hand model (e.g., a MANO model). The MANO model is described in detail in the document "Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: modeling and capturing hands and bodies together. ACM Transactions on Graphics, pages 1–17, 2017," the entire disclosure of which is incorporated herein.
[0252] The human hand is a highly flexible and deformable structure with multiple joints and degrees of freedom. Unlike rigid objects, its geometry is constantly changing due to the bending, stretching, and rotational movements of the fingers and palm. Therefore, it is impractical to hard-code the position of the sphere used to represent the surface contour of the hand in a static frame, as it cannot accurately capture the deformation associated with the posture. To address this problem, the adaptive surface sphere representation mentioned above can be utilized to dynamically determine the vertex indices of the spheres used to represent the surface contour of the hand on the simpleHand mesh template. This approach ensures that the spheres used to represent the surface contour of the hand adapt to local geometric changes and deformations in real time. Figure 8 A mesh model 810 of a hand in a possible hand pose is shown superimposed with a sphere 820 representing the surface contours of the hand. Figure 8The example shown is a ball 820 and a mesh model 810 of a target object (such as a hand) superimposed together as an output display, but those skilled in the art will appreciate that in some embodiments, the ball 820 may be displayed alone as an output without displaying the mesh model 810 of the hand. This processing method may have many advantages. For example, after completing the modeling of a complex object using a ball, the amount of calculation may be reduced and efficiency may be improved when re-rendering the object.
[0253] Example 2
[0254] Some embodiments of the present disclosure provide a system 800 for representing the surface contour of a target object, such as Figure 9 The system 800 includes:
[0255] One or more processors 820. The one or more processors can, for example, execute a computer program to implement the method for representing the surface contour of a target object in the first embodiment.
[0256] In an optional embodiment, the system further includes:
[0257] An image acquisition module is used to acquire a two-dimensional image of the space where the target object is located in real time;
[0258] The one or more processors are further configured to receive the two-dimensional image sent by the image acquisition module and execute, for example, a computer program to implement the method for representing the surface contour of the target object in embodiment 1.
[0259] In some embodiments, the image acquisition module may include a camera.
[0260] In some embodiments, the one or more processors may also output the three-dimensional model of the target object (e.g., a hand) and the generated preset number of balls to a display, instructing the display to display the three-dimensional model of the target object and the preset number of balls. For example, the three-dimensional model of the target object and the preset number of balls may be displayed superimposed on the display, or the three-dimensional model of the target object or the preset number of balls may be displayed separately on the display.
[0261] Object pose estimation
[0262] Object pose estimation is the process of determining the position and orientation of an object in space (6D pose), typically from image data. This is a fundamental task in virtual and augmented reality (VR, AR) and robotics. Understanding the spatial arrangement of objects in an environment can help robots manipulate them.
[0263] Example 3
[0264] The embodiment provides a control method of a mechanical arm, wherein the mechanical arm comprises a base, at least one connecting rod and a joint connecting the at least one connecting rod and the base, the at least one connecting rod comprises an end effector at the end of the mechanical arm, and the control method comprises the following steps: Figure 10
[0265] S301, a surface profile of a target operation object is represented according to the method in the embodiment 1.
[0266] S302, one of the balls representing the surface profile of the target operation object is selected as a target ball.
[0267] S303, a first coordinate is obtained.
[0268] The first coordinate is a coordinate representing a pose of a ball center of the target ball in a world coordinate system.
[0269] S304, a second coordinate is obtained.
[0270] The second coordinate is a coordinate representing a real-time pose of the end effector in the world coordinate system.
[0271] S305, the mechanical arm is controlled to move according to a difference between the first coordinate and the second coordinate.
[0272] In an optional implementation, the step S305 comprises:
[0273] S3051, an initial value of a joint angle θ of the at least one connecting rod is obtained.
[0274] S3052, a target function is constructed.
[0275] The target function comprises a first component representing the difference between the first coordinate and the second coordinate, and the target function takes the joint angle θ as a variable.
[0276] S3053, the value of the joint angle θ is updated by a gradient descent algorithm.
[0277] Specifically, for a robot with a pre-generated hard-coded trajectory, the robot only needs to reach a specific absolute coordinate in the world coordinate system. However, for a smart robot pursuing autonomous action, it is more important that the robot arm / end effector can reach a specific position relative to the target object in its operation environment. For example, if the robot is to perform the action of grabbing a teapot randomly placed in the operation environment, it means that the robot should reach a specific 3D position relative to the teapot, rather than a specific absolute coordinate in the world coordinate system.
[0278] This can be achieved by simply specifying one of the surface sphere center positions we continuously compute in real-time in the world coordinate system as the real-time target for the robot to approach. This can be achieved in a number of ways, for example by trying to minimize the distance between the position of the end effector on the robot arm in the world coordinate R eef,WC
[0279] With the differentiable forward kinematics function, we can set the initial joint angles as learnable parameters and compute the corresponding position of the end effector, then we can use gradient descent to minimize the absolute norm to update the joint angles, thus generating the motion trajectory of the robot in joint space.
[0280] Forward kinematics
[0281] Forward kinematics is a series of physical transformations that describe the physical transformation available to each component in a robot. The physical pose of a robot component is determined not only by its directly connected joint angle or translational stage (relative pose with respect to the adjacent component), but also by the relative pose of that adjacent component with respect to its own adjacent component. Therefore, it is more direct to describe the physical transformation of a robot component as a series of sequential transformations. In computer graphics, such a composite transformation can be simply described as the multiplication of a component geometry point cloud coordinate grid and a 4x4 matrix. A unique 4x4 matrix is usually provided for each independent component of the robot to describe the complete forward kinematics.
[0282] In an alternative embodiment, there is at least one obstacle in the space where the robot arm is located, and the control method comprises:
[0283] S306, representing at least one obstacle and the robot arm according to the method of embodiment 1.
[0284] Wherein the surface profile of the at least one obstacle is represented by M spheres, and the surface profile of the robot arm is represented by N spheres.
[0285] S307, determining all sphere pairs.
[0286] Wherein the sphere pair is composed of the i-th sphere representing the surface profile of the at least one obstacle and the j-th sphere representing the surface profile of the robot arm, i is a positive integer not greater than M, and j is a positive integer not greater than N.
[0287] S308, determining the value of the first distance of each sphere pair.
[0288] The first distance is a value of a distance between a center of a sphere representing a surface profile of the at least one obstacle and a center of a sphere representing a surface profile of the robot arm in each sphere pair.
[0289] S309, determining a value of a sphere radius and of each sphere pair.
[0290] The value of the sphere radius and is a sum of a value of a sphere radius of the sphere representing the surface profile of the at least one obstacle and a value of a sphere radius of the sphere representing the surface profile of the robot arm in each sphere pair.
[0291] S310, determining a spherical distance of each sphere pair according to the value of the first distance and the value of the sphere radius and of each sphere pair.
[0292] S311, determining a value of a first collision possibility of each sphere pair according to a difference between the spherical distance and a preset threshold.
[0293] The first collision possibility represents a possibility of collision between the two spheres in each sphere pair.
[0294] S312, determining a value of a second possibility of collision between the robot arm and the at least one obstacle.
[0295] The value of the second possibility is a sum of the values of the first collision possibilities of all sphere pairs.
[0296] S313, controlling movement of the robot arm according to the value of the second possibility.
[0297] In an optional implementation, the objective function includes a second component, where the second component represents the value of the second possibility.
[0298] In an optional implementation, the step S303 includes:
[0299] S3031, obtaining a third coordinate.
[0300] The third coordinate is a coordinate representing a pose of a center of a target sphere in a target operating object coordinate system.
[0301] S3032, obtaining a 6D pose of the target operating object.
[0302] S3033, converting the third coordinate to the first coordinate according to the 6D pose of the target operating object.
[0303] The step S304 includes:
[0304] S3041, obtaining a fourth coordinate.
[0305] The fourth coordinate is a coordinate representing a pose of the end effector in an end effector coordinate system.
[0306] S3042. Obtain forward kinematics information of the robotic arm.
[0307] Among them, the forward kinematics information includes joint translation and rotation angles.
[0308] Specifically, a robot consists of a base and a robotic arm, which is composed of multiple links and joints connecting them. The joints can generate movement and rotation for the robotic arm, allowing the end of the robotic arm to reach a specified position.
[0309] In forward kinematics, each joint in a robotic arm has a coordinate system. The coordinate systems are set as follows: For each coordinate system, the z-axis is set to coincide with the joint axis, the x-axis points in the direction of the next joint, and is perpendicular to the z-axis of the coordinate system corresponding to the previous joint and the z-axis of the current coordinate system. The direction of the y-axis is determined using the right-hand rule. Rotation and translation occur between adjacent coordinate systems, and the coordinate changes between adjacent links can be described using the Denavit-Hartenberg (DH) parameter.
[0310] DH parameters consist of four basic elements: θ (Theta), d (Distance), a (Link Length), and α (Twist Angle); θ represents the rotation angle around the previous z-axis, d represents the translation distance along the previous z-axis, a represents the length of the common perpendicular from the previous z-axis to the current z-axis, and α represents the angle between the two z-axes. The DH parameter table uses a series of values for these four parameters to define the position and orientation of each joint of the robot, thereby describing the posture of the entire manipulator. For example, a planar manipulator with two joints might have a DH parameter table containing two rows, each containing four parameter values. These values can be used to determine the position and posture of each joint relative to the previous joint, ultimately determining the position and posture of the end effector. Specifically, the end effector can be a gripper.
[0311] The forward kinematics information can be obtained by loading a URDF (United Robotics Description Format) file.
[0312] Unified Robot Description Format (URDF)
[0313] URDF is an example of a data format used to describe the physical geometry of a robot and its components in the form of point cloud meshes, as well as the robot's degrees of freedom to implement its physical structure transformations, such as joint angle rotations and physical translations. Mesh components in a URDF file may include a visual mesh and a collision mesh; the visual mesh is a large number of 3D polygons used for visualization, while the collision mesh is a smaller number of 3D polygons used for efficient collision detection.
[0314] S3043、According to the forward kinematics information, a corresponding forward kinematics function is constructed for each link.
[0315] The pose of a robot component is determined not only by its neighboring component's θ or d, but also by the relative pose of the neighboring component relative to its own neighboring component. Therefore, it is more intuitive to describe the transformation of the robot pose as a composite transformation of a sequence of transformations. In computer graphics, such a composite transformation can be simply described as the multiplication of a component geometry point cloud coordinate grid and a 4x4 transformation matrix. A unique 4x4 transformation matrix is usually provided for each independent component of the robot to describe the complete forward kinematics.
[0316] For the transformation between coordinate systems, first consider the rotation, from the coordinate system of the i-th joint to the coordinate system of the i-1-th joint, the x-axis is rotated by α i-1 , and the z-axis is rotated by θ i ; then consider the translation, first along the z-axis by d i , and then along the x-axis by a i-1 . If the joints of the robot arm are revolute joints, then θ i is the variable, and the rest are constants, when the robot arm moves. If the joints of the robot arm are prismatic joints, then a i-1 is the variable, and the rest are constants, when the robot arm moves.
[0317] wherein the forward kinematics function takes the joint angle θ as a variable, and is used to convert coordinates in the coordinate system of a link to the world coordinate system.
[0318] After knowing the D-H parameters corresponding to each joint, the transformation matrix between the coordinate systems of adjacent joints can be obtained , thereby obtaining the position and attitude of a link relative to an adjacent link, the transformation matrix is determined by θ, d, a, and α as introduced above. By multiplying all the transformation matrices, the forward kinematics function FK can be obtained, thereby obtaining the position and attitude of the end effector of the robot arm in the robot root coordinate system. In the present disclosure, the joints of the robot arm are revolute joints, so each transformation matrix takes the parameter θ as a variable, and the forward kinematics function is FK(θ), and the specific formula is as follows:
[0319]
[0320] wherein, is the transformation matrix; P is the pose; i is the i-th movable link of the robot arm; when i is 0, it is the base of the robot; and FK(θ) is the forward kinematics function taking θ as a variable.
[0321] S3044. Obtain the 6D pose of the base.
[0322] S3045. Convert the fourth coordinate into the second coordinate according to the forward kinematics function and the 6D pose of the base.
[0323] In an optional embodiment, the control method further includes:
[0324] S314: Obtain the fifth coordinate.
[0325] The fifth coordinate is a coordinate representing the posture of the ball in each ball pair representing the surface contour of at least one obstacle in at least one obstacle coordinate system.
[0326] S315: Obtain a 6D pose of at least one obstacle.
[0327] S316 : Convert the fifth coordinate into a sixth coordinate according to the 6D pose of the at least one obstacle.
[0328] The sixth coordinate is a coordinate representing the posture of the ball in each ball pair representing the surface contour of at least one obstacle in the world coordinate system.
[0329] Specifically, the sphere used to represent the surface contour of at least one obstacle can be expressed as:
[0330]
[0331] in, is the vector representing the coordinates of the center of each ball in the obstacle coordinate system, and is the corresponding radius, and M is the number of spheres representing the surface contour of the object.
[0332] The fifth coordinate can be converted to the sixth coordinate by the following formula:
[0333]
[0334] in, is the vector representing the coordinates of the center of each ball in the world coordinate system, P is the 6D pose of the obstacle, and is a 4×4 transformation matrix.
[0335] Through the above formula, the ball can be transformed from the obstacle coordinate system to the world coordinate system, and the coordinates in the world coordinate system can be obtained
[0336] S317. Obtain the seventh coordinate.
[0337] The seventh coordinate is a coordinate representing the posture of the ball representing the surface profile of the robotic arm in each ball pair in the corresponding link coordinate system.
[0338] S318. Obtain forward kinematics information of the robotic arm.
[0339] Among them, the forward kinematics information includes joint translation and rotation angles.
[0340] S319. Construct a corresponding forward kinematics function for each connecting rod based on the forward kinematics information.
[0341] The forward kinematics function uses the joint angle θ as a variable and is used to transform the coordinates in the link coordinate system to the world coordinate system.
[0342] S320: Obtain the 6D pose of the base.
[0343] S321. Convert the seventh coordinate into the eighth coordinate according to the forward kinematics function and the 6D pose of the base.
[0344] The eighth coordinate is a coordinate representing the posture of the ball representing the surface profile of the robotic arm in each ball pair in the world coordinate system.
[0345] Specifically, the seventh coordinate can be converted to the eighth coordinate by the following formula:
[0346]
[0347] Since the robot's pose is changing dynamically, we must transform the center of each sphere from the local coordinate system to the world coordinate system to reflect the robot's current world coordinate configuration. This transformation is achieved using the robot's forward kinematics and the 6D pose of the robot base. Given a 6D pose P of the robot base base and joint angles θ, and the robot surface sphere model S with N links robot ={l n |n∈{1,2,…N}}, where M n To represent the number of balls on the surface profile of the link, the forward kinematics matrix of each link can be calculated based on the initial joint angle θ by solving the robot forward kinematics.
[0348] The spherical distance of each sphere pair, that is, the distance between the sphere of the sphere representing the surface profile of the manipulator and the sphere of the sphere representing the surface profile of the obstacle, can be calculated by the following formula:
[0349]
[0350] Among them, d ij is the spherical distance; is the eighth coordinate; is the sixth coordinate; the spherical radius of the sphere representing the surface profile of at least one obstacle in each sphere pair; is the radius of the sphere that represents the surface profile of the robot arm.
[0351] In some embodiments, the first collision probability value of each ball pair can be calculated using the following formula:
[0352]
[0353] Among them, Cost ij The value of the first collision probability is d ij is the spherical distance, d threshold is the preset threshold.
[0354] When the pairwise cost function Cost ij A real-time robot control trajectory algorithm that minimizes this collision cost, when summed over all pairs of balls, will start when the distance between pairs of balls on the surface reaches a preset threshold d. threshold Using differentiable forward kinematics, the collision cost function can be easily combined with other cost terms (such as the target distance cost) to achieve the goal of approaching the target while avoiding obstacles in the path.
[0355] Because the spheres representing the surface contours of the objects are pre-generated and can be reused, their position in the world coordinate system depends on the object pose estimation output. Accordingly, the sphere-related portion of the computation is very lightweight. Sphere placement in world coordinates involves a single pose matrix multiplication with the rigid object sphere, or a single index selection operation with the flexible object sphere. On the other hand, sphere collision calculations only involve pairwise sphere distance calculations between the robot sphere and the object spheres. This computational overhead grows linearly with the number of object spheres in the system, making it well-suited for real-time dynamic control systems. Unlike the SDF / MPPI approach used by Samsung, which runs at a 3Hz trajectory update rate, the sphere-based control loop can run at 40-70Hz.
[0356] Example 4
[0357] This embodiment provides a control system for a robotic arm, which includes: one or more processors, and the one or more processors are used to execute, for example, a computer program to implement the control method of the robotic arm as in Example 3.
[0358] Example 5
[0359] Figure 11This is a structural schematic diagram of an electronic device showing an example embodiment of the present disclosure, wherein the electronic device includes a memory, a processor, and a computer program stored in the memory and for running on the processor. When the processor executes the computer program, it implements the method of representing the surface contour of a target object or the method of controlling a robotic arm according to any of the above embodiments. Figure 11 The electronic device 50 shown is only an example. For example, those skilled in the art will appreciate that part of the processor, such as an ASIC processor, may include a memory. Figure 11 The electronic device 50 shown should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0360] like Figure 11 As shown, the electronic device 50 may be a general-purpose computing device, such as a server device. Components of the electronic device 50 may include, but are not limited to, the at least one processor 51, the at least one memory 52, and a bus 53 connecting different system components (including the memory 52 and the processor 51).
[0361] The bus 53 includes a data bus, an address bus, and a control bus.
[0362] The memory 52 may include a volatile memory, such as a random access memory (RAM) 521 and / or a cache memory 522 , and may further include a read-only memory (ROM) 523 .
[0363] The memory 52 may also include a program tool 525 (or utility) having a set (at least one) of program modules 524, such program modules 524 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0364] The processor 51 executes various functional applications and data processing by running the computer programs stored in the memory 52, such as the method for representing the surface contour of the target object or the method for controlling the robotic arm provided in any of the above embodiments.
[0365] The electronic device 50 can also communicate with one or more external devices 54 (e.g., a keyboard, pointing device, etc.). Such communication can be performed via an input / output (I / O) interface 55. Furthermore, the electronic device 50 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 56. As shown, the network adapter 56 communicates with other modules of the electronic device 50 via a bus 53. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 50, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (RAID) systems, tape drives, and data backup storage systems.
[0366] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.
[0367] Example 6
[0368] Some embodiments of the present disclosure also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for representing the surface contour of a target object or the method for controlling a robotic arm provided in any of the above embodiments.
[0369] Example 7
[0370] Some embodiments of the present disclosure also provide a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned methods for representing the surface contour of a target object or the method for controlling a robotic arm.
[0371] In some embodiments of the present disclosure, the readable storage medium and computer program product may more specifically include but are not limited to: a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0372] Some embodiments of the present disclosure also provide a computer program, which, when executed by one or more processors, can implement the above-mentioned method for representing the surface contour of a target object.
[0373] Some embodiments of the present disclosure also provide another computer program, which, when executed by one or more processors, can implement the above-mentioned control method of the robotic arm.
[0374] The program code for executing the computer program product of the present disclosure may be written in any combination of one or more programming languages, and the program code may be executed entirely on the user device, partially on the user device, as a standalone software package, partially on the user device and partially on a remote device, or entirely on the remote device.
[0375] Although specific embodiments of the present invention have been described above, those skilled in the art will appreciate that these are merely illustrative and that the scope of the present invention is defined by the appended claims. Those skilled in the art may make various changes or modifications to these embodiments without departing from the principles and essence of the present invention, and such changes and modifications are intended to fall within the scope of the present invention.
Claims
1. A method for representing the surface profile of a target object, characterized in that The method comprises: Obtaining a three-dimensional model of the target object; Based on the three-dimensional model, a preset number of balls are generated to represent the surface contours of different regions of the target object, wherein the density of the balls in the different regions corresponds to the geometric complexity of the surface contours of the different regions.
2. The method for representing the surface contour of a target object according to claim 1, It is characterized by: The different regions include a first region and a second region, and the geometric complexity of the surface profile of the first region is greater than the geometric complexity of the surface profile of the second region; The density of the balls in the different regions corresponds to the geometric complexity of the surface profiles of the different regions, including: The density of balls in the first area is greater than the density of balls in the second area.
3. The method for representing the surface contour of a target object according to claim 1 or 2, wherein: The method comprises: Determine density weight values of candidate sampling points in different regions of the target object, wherein the density weight values represent information about geometric complexity of surface contours of the different regions.
4. The method for representing the surface contour of a target object according to claim 3, wherein: The step of determining the density weight values of candidate sampling points in different areas of the target object includes: Determining a neighborhood of the candidate sampling point based on a position of the candidate sampling point and a preset neighborhood size value; Based on the neighborhood, a density weight value of the candidate sampling point is determined.
5. The method for representing the surface contour of a target object according to claim 4, wherein: The step of determining the density weight value of the candidate sampling point based on the neighborhood includes: The density weight value of the candidate sampling point is determined according to the average distance between the candidate sampling point and other candidate sampling points in the neighborhood.
6. The method for representing the surface contour of a target object according to claim 5, wherein: The three-dimensional model is a polygonal mesh model, wherein the polygonal mesh model is composed of a set of vertices, edges and faces, and the candidate sampling points are vertices of the polygonal mesh model; The step of determining the neighborhood of the candidate sampling point based on the position of the candidate sampling point and a preset neighborhood size value includes: Taking the candidate sampling point as the center, a range of polygons separated from the candidate sampling point and having a number of sides no greater than n is determined as a neighborhood of the candidate sampling point, wherein the preset neighborhood size is n, and n is a positive integer; or, The three-dimensional model is a voxel model, wherein the voxel model is composed of a set of voxels, and the candidate sampling point is the center of the outermost voxel of the voxel model; The step of determining the neighborhood of the candidate sampling point based on the position of the candidate sampling point and a preset neighborhood size value includes: Taking the candidate sampling point as the center, a range with a distance from the candidate sampling point not greater than r is determined as a neighborhood of the candidate sampling point, wherein the preset neighborhood size is r, and r is a positive number.
7. The method for representing the surface contour of a target object according to any one of claims 3 to 6, wherein: The step of generating a preset number of balls based on the three-dimensional model to represent the surface contours of different areas of the target object includes: Determining the preset number of sampling points from all candidate sampling points based on the value of the density weight; The predetermined number of spheres are generated with the determined sampling points as sphere centers.
8. The method for representing the surface contour of a target object according to claim 7, wherein: The step of determining the preset number of sampling points from all candidate sampling points based on the value of the density weight includes: Determine one of the candidate sampling points as the initial sampling point; The next sampling point is determined from the remaining candidate sampling points based on a Euclidean distance between each remaining candidate sampling point and a last candidate sampling point determined as a sampling point, and a density weight value of each remaining candidate sampling point, until the preset number of sampling points is determined, wherein the remaining candidate sampling points are candidate sampling points that have not yet been determined as sampling points.
9. The method for representing the surface contour of a target object according to claim 8, wherein: The step of determining the next sampling point from the remaining candidate sampling points according to the Euclidean distance between each remaining candidate sampling point and the last candidate sampling point determined as the sampling point, and the density weight of each remaining candidate sampling point, comprises: Determine, based on the Euclidean distance between each remaining candidate sampling point and the last candidate sampling point determined as the sampling point, and the density weight of each remaining candidate sampling point, a value of a comprehensive distance between each remaining candidate sampling point and the last candidate sampling point determined as the sampling point; The remaining candidate sampling point with the largest corresponding comprehensive distance value is determined as the next sampling point.
10. The method for representing the surface contour of a target object according to any one of claims 3 to 9, characterized in that: The method further comprises: Based on the three-dimensional model, the sphere radii of the preset number of spheres are determined, wherein the sphere radii of the spheres in the different regions correspond to the geometric complexity of the surface profiles of the different regions.
11. The method for representing the surface contour of a target object according to claim 10, wherein: The spherical radius of the balls in the different regions corresponds to the geometric complexity of the surface contours of the different regions, including: The spherical radius of the ball in the first area is smaller than the spherical radius of the ball in the second area.
12. The method for representing the surface contour of a target object according to claim 10 or 11, characterized in that: in, The sampling points include a first sampling point and a second sampling point; The step of determining the spherical radius of the preset number of balls based on the three-dimensional model includes: The spherical radius of each sphere is determined according to a preset radius and a density weight of a sampling point corresponding to each sphere, wherein when the density weight of the first sampling point is greater than the density weight of the second sampling point, the spherical radius of the sphere corresponding to the first sampling point is smaller than the spherical radius of the sphere corresponding to the second sampling point.
13. The method for representing the surface contour of a target object according to claim 12, wherein: The method further comprises: Normalizing the density weight values of the sampling points corresponding to each sphere to obtain the normalized density weight values of the sampling points corresponding to each sphere, wherein the normalized density weight values are positive numbers less than 1.
14. The method for representing the surface contour of a target object according to claim 13, wherein: The step of determining the value of the spherical radius of each sphere according to the value of the preset radius and the density weight of the sampling point corresponding to each sphere includes: The value of the spherical radius of each sphere is determined according to the value of the preset radius and the value of the normalized density weight of the sampling point corresponding to each sphere.
15. The method for representing the surface contour of a target object according to claim 14, wherein: The step of determining the value of the spherical radius of each sphere according to the value of the preset radius and the value of the normalized density weight of the sampling point corresponding to each sphere includes: The spherical radius of each sphere is determined according to the value of the preset radius, the value of the normalized density weight of the sampling point corresponding to each sphere, and the value of the proportional coefficient, wherein the proportional coefficient is used to control the sensitivity of the spherical radius of each sphere to changes in the value of the normalized density weight of the sampling point corresponding to each sphere.
16. The method for representing the surface contour of a target object according to any one of claims 12 to 15, characterized in that: The method further comprises: Descriptor information of the preset number of spheres is generated, wherein the descriptor information includes at least one of index information, position information, normal information, and sphere radius of the spheres on the three-dimensional model.
17. The method for representing the surface contour of a target object according to claim 16, wherein: The method further comprises: Obtain the location information of all candidate sampling points in real time; The position information of all sampling points is determined according to the position information of all candidate sampling points and the index information.
18. The method for representing the surface contour of a target object according to any one of claims 1 to 17, wherein: The obtaining of the three-dimensional model of the target object comprises: Get data of real-time video stream; A three-dimensional model of the target object is generated according to the two-dimensional image of the target object in the video stream.
19. The method for representing the surface contour of a target object according to claim 18, wherein: Generating a three-dimensional model of the target object according to the two-dimensional image of the target object in the video stream includes: determining a bounding box of the target object in the two-dimensional image based on an open vocabulary object detection model and input information corresponding to the target object, wherein the bounding box indicates a position and size of the target object in the two-dimensional image, optionally wherein the open vocabulary object detection model comprises Grounding DINO or YOLO-World; A three-dimensional model of the target object is generated according to the two-dimensional image and the bounding box.
20. The method for representing the surface contour of a target object according to claim 18 or 19, characterized in that: The target object is a hand, and the method includes: displaying a three-dimensional model of the hand and the preset number of balls.
21. A system for representing the surface contour of a target object, characterized in that: The system comprises: One or more processors, wherein the one or more processors are configured to implement the method for representing the surface contour of a target object according to any one of claims 1 to 17.
22. The system for representing the surface contour of a target object according to claim 21, wherein: The system further comprises: An image acquisition module, wherein the image acquisition module is used to obtain data of a real-time video stream; The one or more processors are further configured to process the two-dimensional image of the target object in the video stream obtained by the image acquisition module to implement the method according to claim 18 or 19.
23. The system for representing the surface contour of a target object according to claim 21 or 22, wherein: The target object is a hand, and the system further comprises: a display; The one or more processors are further configured to output the three-dimensional model of the hand and the preset number of balls to the display to implement the method of claim 20.
24. A method of controlling a robotic arm, wherein: The robotic arm includes a base, at least one connecting rod, and a joint connecting the at least one connecting rod and the base, wherein the at least one connecting rod includes an end effector located at an end of the robotic arm, and the method includes: According to the method of any one of claims 1 to 20, generating a first preset number of balls to represent the surface contour of the target operation object; Selecting one of the balls representing the surface contour of the target operation object as a target ball; The robotic arm is controlled according to the distance between the end effector and the center of the target ball.
25. The method for controlling a robotic arm according to claim 24, wherein: The controlling the robotic arm according to the distance between the end effector and the center of the target ball comprises: Acquire a first coordinate, wherein the first coordinate is a coordinate representing the center of the target ball in a world coordinate system; Acquire a second coordinate, wherein the second coordinate represents a coordinate of the end effector in the world coordinate system; Based on minimizing the distance between the first coordinate and the second coordinate, the robotic arm is controlled to approach the target operation object.
26. The method of claim 25, wherein: The controlling the robotic arm to approach the target operation object based on minimizing the distance between the first coordinate and the second coordinate includes: Based on an initial value of the joint angle θ of the end effector, a gradient descent method is used to iteratively update the value of the joint angle θ so as to minimize the distance difference between the first coordinate and the second coordinate.
27. The method according to claim 22 or 23, wherein: There is at least one obstacle in the space where the robotic arm is located, and the control method includes: The method according to any one of claims 1 to 20, generating spheres to represent the at least one obstacle and the robotic arm, wherein a surface contour of the at least one obstacle is represented by M spheres, and a surface contour of the robotic arm is represented by N spheres; Determine all ball pairs, wherein the ball pair consists of an i-th ball representing the surface profile of the at least one obstacle and a j-th ball representing the surface profile of the robotic arm, where i is a positive integer not greater than M and i is a positive integer not greater than N; Determining the probability of collision between the two balls based on comparing the spherical surface distance between the two balls in each ball pair with a preset threshold; The robotic arm is controlled according to the possibility of the collision.
28. The method of claim 27, wherein: The determining of the possibility of collision between the two balls according to comparing the spherical surface distance between the two balls in each ball pair with a preset threshold value includes: Determine a collision cosine cost function for each ball pair based on a comparison between a spherical surface distance between two balls in each ball pair and a preset threshold, wherein the collision cosine cost function represents a probability of collision between the two balls; The controlling the robotic arm according to the possibility of the collision includes: The robotic arm is controlled by minimizing the sum of the collision cosine cost functions for each team in all the ball pairs.
29. A system for controlling a robotic arm, characterized in that: The system includes: one or more processors, and the one or more processors are used to implement the control method of the robotic arm according to any one of claims 22 to 28.
30. An electronic device comprising one or more memories, one or more processors, and a computer program stored in the one or more memories and configured to run on the one or more processors, wherein: When the one or more processors execute the computer program, the method for representing the surface contour of a target object according to any one of claims 1 to 20 or the method for controlling a robotic arm according to any one of claims 24 to 28 is implemented.
31. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by one or more processors, the method for representing the surface contour of a target object according to any one of claims 1 to 20 or the method for controlling a robot arm according to any one of claims 24 to 27 is implemented.
32. A computer program product comprising a computer program, characterized in that When the computer program is executed by one or more processors, the method for representing the surface contour of a target object according to any one of claims 1 to 20 or the method for controlling a robot arm according to any one of claims 24 to 28 is implemented.