Visual positioning method, robot control method, related equipment and medium
By combining the acquisition of RGB images and point cloud data with the positioning control of the mechanical platform and robotic arm, the problems of high cost and complex point cloud processing of lidar were solved, and low-cost and high-precision visual positioning and operation control were achieved.
Patent Information
- Application Number
- CN202310130923.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-02
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2043-02-02
AI Technical Summary
The existing visual positioning systems of outdoor robots rely on lidar, which leads to high costs and computational loads, making it difficult to perform high-precision object recognition at close range, and the processing of point cloud data is complex, limiting their widespread application.
Visual sensors are used to collect RGB images, depth images and point cloud data. The center of mass and main direction of the target object are determined by combining images and point clouds. Combined with the positioning control of the mechanical platform and robotic arm, a visual positioning method that combines coarse positioning and fine positioning is realized.
It reduces the hardware requirements for visual positioning, improves positioning accuracy and operation precision, is suitable for outdoor operation robots, reduces overall costs and simplifies data processing.
Smart Images

Figure CN116237937B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of machine vision technology, and in particular relates to a visual positioning method, a robot control method, and related equipment and media. Background Art
[0002] With the development of computer and robotics technology, modern mobile robots are finding increasingly widespread application in industrial manufacturing, military, civilian applications, and scientific research. They can replace humans in performing tasks that are difficult or difficult to perform under harsh conditions. Mobile robot research lies at the intersection of multiple disciplines, providing a broad platform for the development of new theories and methods. Based on their operating environments, mobile robots can be broadly divided into two categories: outdoor robots and indoor robots.
[0003] The visual positioning system for outdoor robots consists of a camera system and a control system. The camera system includes a computer (with an image acquisition card) and an industrial camera or depth camera. It primarily collects visual images or 3D point clouds and uses machine vision technology to process these images or point clouds for guided positioning and pattern recognition. This allows for rapid acquisition of the center of mass and boundaries of objects, meeting the robot's self-positioning requirements and shortening the gap between its desired position and its end position. The control system, comprised of a control box and a computer, controls the specific position of the end-user. The camera captures the workspace, and the computer uses the image to identify tracking features, calculate and identify data, and use inverse kinematics to determine the error at each robot position. This then controls the high-precision end-effector module to precisely adjust the robot's position and posture.
[0004] Currently, the mainstream approach for designing vision systems for outdoor robots is to use lidar to scan the scene and then locate objects of interest based on the resulting point cloud, combined with feature extraction and pattern recognition. This is because lidar is insensitive to changes in outdoor lighting conditions and can maintain relatively consistent depth imaging performance under varying lighting conditions. However, lidar's advantage lies in long-range ranging. When it comes to object recognition within close ranges of only a few meters, lidar is almost inadequate. In other words, for close-range recognition and positioning, high-precision ranging radar must be used to scan and reconstruct the scene within a small area to obtain positioning information. However, the high cost of lidar itself has prevented the widespread adoption of this method for visual guidance of operational robots. Furthermore, the raw point cloud data acquired by lidar is complex and time-consuming to process, requiring higher computer performance and further increasing overall hardware costs. Summary of the Invention
[0005] In view of this, the embodiments of the present application provide a visual positioning method, a robot control method, related equipment and media, which are used to reduce the difficulty of visual positioning and improve the accuracy of visual positioning, thereby obtaining a universal visual positioning method, and the visual positioning method does not have high requirements on the equipment.
[0006] A first aspect of an embodiment of the present application provides a visual positioning method, the method comprising:
[0007] Collecting a data stream through a visual sensor camera, wherein the data stream includes an RGB image, a depth image, and a point cloud, and the data stream includes data of a target object to be located;
[0008] Determining a target point cloud corresponding to the target object from the point cloud according to the RGB image and the depth image;
[0009] Determining the center of mass and main direction of the target object according to the target point cloud;
[0010] The target object is positioned according to the center of mass and the main direction.
[0011] A second aspect of an embodiment of the present application provides a robot control method, which is applied to the robot, wherein the robot includes a mechanical platform and a mechanical arm, and the method includes:
[0012] Determining first positioning data, where the target object is a working object of the robot, and the first positioning data is used to represent a positional relationship of the mechanical platform relative to the target object;
[0013] Controlling the mechanical platform to move to a position corresponding to the first positioning data;
[0014] Determining second positioning data, where the second positioning data is used to represent a positional relationship of the robotic arm relative to the target object after the robotic platform reaches a predetermined position;
[0015] The positioning accuracy of the first positioning data is lower than the positioning accuracy of the second positioning data, and the robot determines the first positioning data and the second positioning data by the method described in the first aspect above.
[0016] A third aspect of the embodiments of the present application provides a visual positioning device, comprising:
[0017] A visual data acquisition module is used to acquire a data stream through a visual sensor camera, wherein the data stream includes an RGB image, a depth image, and a point cloud, and the data stream includes data of a target object to be located;
[0018] a target point cloud determination module, configured to determine a target point cloud corresponding to the target object from the point cloud according to the image, the RGB image and the depth image;
[0019] A center of mass and direction determination module, configured to determine the center of mass and main direction of the target object based on the target point cloud;
[0020] A positioning module is used to locate the target object according to the center of mass and the main direction.
[0021] A fourth aspect of the embodiments of the present application provides a robot control device, which is applied to the robot, wherein the robot includes a mechanical platform and a mechanical arm, and the device includes:
[0022] A first positioning module is configured to determine first positioning data, wherein the target object is a working object of the robot, and the first positioning data is used to represent a positional relationship of the mechanical platform relative to the target object;
[0023] a first control module, configured to control the mechanical platform to move to a position corresponding to the first positioning data;
[0024] a second positioning module, configured to determine second positioning data, wherein the second positioning data is used to represent a positional relationship of the robotic arm relative to the target object after the robotic platform reaches a predetermined position;
[0025] a second control module, configured to control the robotic arm to operate on the target object according to the second positioning data;
[0026] The positioning accuracy of the first positioning data is lower than the positioning accuracy of the second positioning data, and the robot determines the first positioning data and the second positioning data by the method described in the first aspect above.
[0027] The fifth aspect of the embodiments of the present application provides a robot, which includes a mechanical platform and a mechanical arm. The robot locates a target object using the method described in the first aspect above; the robot is controlled using the method described in the second aspect above to perform operations on the target object.
[0028] A sixth aspect of an embodiment of the present application provides a robotic arm equipped with a monocular depth camera. The robotic arm locates a target object using the method described in the first aspect above to perform operations on the target object.
[0029] A seventh aspect of an embodiment of the present application provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method described in the first or second aspect above is implemented.
[0030] An eighth aspect of the embodiments of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first or second aspect above is implemented.
[0031] A ninth aspect of the embodiments of the present application provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute the method described in the first or second aspect above.
[0032] Compared with the prior art, the embodiments of the present application have the following advantages:
[0033] In the embodiment of the present application, when performing visual positioning, the terminal device can collect data streams through visual sensors. The data streams may include RGB images, depth images, and point clouds. The data streams include data of the located target object. Based on the RGB images and depth images, the terminal device can determine the target point cloud corresponding to the target object from the point cloud image, thereby being able to determine the center of mass and main direction of the target object based on the target point cloud, and based on the center of mass and main direction of the target object, achieve positioning of the target object. In the embodiment of the present application, when performing visual positioning, RGB image data, depth image data, and point cloud data can be used simultaneously to determine the target object, and the complex point cloud contour extraction algorithm can be converted into image detection combined with basic point cloud data processing, which is convenient for transplantation and integration, and has a small number of parameters. Therefore, the visual positioning method in the embodiment of the present application is highly versatile and has low hardware requirements.
[0034] Based on the above-mentioned visual positioning method, an embodiment of the present application also provides a robot control method. The robot may include a mechanical platform and a mechanical arm. The robot may use the above-mentioned visual positioning method to coarsely locate a target object, obtain first positioning data, and then control the mechanical platform to move to within the operating radius of the mechanical arm based on the first positioning data. After the mechanical platform moves to the position corresponding to the first positioning data, the above-mentioned visual positioning method is used to achieve fine positioning of the target object, obtain second positioning data, and then control the mechanical arm to operate on the target object based on the first positioning data. The robot operates on the target object based on secondary positioning by combining coarse positioning and fine positioning, thereby improving the accuracy of the robot's positioning and correspondingly improving the accuracy of the robot's operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art.
[0036] Figure 1 This is a schematic diagram of the steps of a visual positioning method provided in an embodiment of the present application;
[0037] Figure 2 This is a schematic flow chart of the steps of a robot control method provided in an embodiment of the present application;
[0038] Figure 3 is a flow chart of another robot control method provided in an embodiment of the present application;
[0039] Figure 4 is a schematic diagram of a visual positioning device provided in an embodiment of the present application;
[0040] Figure 5 is a schematic diagram of a robot control device provided in an embodiment of the present application;
[0041] Figure 6 This is a schematic diagram of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0042] In the following description, specific details such as specific system structures and technologies are provided for the purpose of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obstructing the description of the present application with unnecessary details.
[0043] The technical solution of this application is described below through specific embodiments.
[0044] Reference Figure 1 , shows a schematic flow chart of the steps of a visual positioning method provided in an embodiment of the present application, which may specifically include the following steps:
[0045] S101, collecting a data stream through a visual sensor camera, wherein the data stream includes an RGB image, a depth image, and a point cloud, and the data stream includes data of a target object to be located.
[0046] The execution subject of the embodiment of the present application is a terminal device, which may include a visual sensor. Based on the data collected by the visual sensor, the terminal device can perform visual positioning. For example, the terminal device may be a robot, and the robot may include a mechanical platform and a mechanical arm. The mechanical platform and the mechanical arm of the robot may be respectively installed with visual sensors, and the mechanical platform and the mechanical arm may be positioned respectively based on the visual sensors. When the robot is far away from the target object, the robot may perform visual positioning, determine the positional relationship between the mechanical platform and the target object, and thereby control the mechanical platform to move to the vicinity of the target object based on the positioning result, and then determine the positional relationship between the mechanical arm and the target object, and thereby control the mechanical arm to operate the target object based on this positioning result. The visual sensor may be a camera, specifically, an active light binocular depth camera, a monocular depth camera, etc. There is no limitation on the type of visual sensor in this embodiment.
[0047] Visual sensors can capture visual image data. For example, they can capture RGB image data, depth images, and point clouds. RGB images are two-dimensional images captured by visual sensors, which can include images of the target object and, of course, images surrounding the target object. Depth images include the depth value of each pixel, which can represent the distance between the visual sensor and the target object. Point cloud images can be points in the image within the visual sensor's field of view, equivalent to a collection of multiple points on the surface of the target object. Of course, the point cloud also includes a point cloud collection of objects surrounding the target object.
[0048] In order to perform visual positioning, the target object needs to be determined from RGB image data, depth image and point cloud.
[0049] S102: Determine a target point cloud corresponding to the target object from the point cloud according to the image, the RGB image, and the depth image.
[0050] In this embodiment, the RGB image, depth image and point cloud can be aligned, and the depth image and point cloud can all be converted to the image coordinate system corresponding to the RGB image, so that they can correspond in the image coordinate system. The image coordinate system can be automatically calibrated and generated by the camera. The camera coordinate system is a three-dimensional rectangular coordinate system established with the camera's focal center as the origin and the optical axis as the Z axis; the origin of the camera coordinate system is projected into the image to obtain the origin of the image coordinate system, thereby establishing the image coordinate system. The conversion between the image coordinate system and the depth coordinate system can also be calibrated by the camera's own calibration system.
[0051] The terminal device can rearrange the point cloud based on the RGB image to obtain a rearranged point cloud, making the shape of the rearranged point cloud consistent with the shape of the RGB image, thereby achieving alignment between the RGB image and the point cloud. After alignment, it is equivalent to aligning each pixel in the RGB image with each pixel in the point cloud. In other words, the corresponding positions of the target object in the RGB image data and the point cloud are the same.
[0052] For example, a terminal device can determine a depth point cloud in a depth coordinate system. The depth coordinate system is a coordinate system established based on the depth image. Each pixel in the depth point cloud has a depth value, which is equivalent to aligning the point cloud to the depth image. The depth point cloud can then be converted to an image coordinate system to obtain a visual point cloud, which is equivalent to aligning the point cloud and the depth image with the RGB image. In other words, the visual point cloud is aligned with the RGB image, and the depth value of each point is also determined. After determining the visual point cloud, the visual point cloud can be preprocessed, namely, removing foreground and background pixels from the visual point cloud. Foreground pixels are pixels with depth values less than a first preset depth threshold, which means they are relatively close to the visual sensor. Background pixels are pixels with depth values greater than a second preset depth threshold, which means they are relatively far from the visual sensor. In other words, foreground and background pixels are pixels that are determined to be unlikely to be the target object. Removing these foreground and background pixels from the visual point cloud is equivalent to removing noise information, facilitating subsequent target recognition. The visual point cloud is rearranged to obtain a rearranged point cloud aligned with the RGB image.
[0053] In one possible implementation, the point cloud obtained in the depth coordinate system can be aligned to the RGB image coordinate system, and the depth coordinate system can be denoted as T depth , the corresponding point cloud in the depth coordinate system can be recorded as P depth , the RGB image coordinate system is T color The visual sensor can be automatically calibrated to obtain the conversion matrix from the depth coordinate system to the RGB image coordinate system: The expression of the aligned point cloud in the image coordinate system can be:
[0054]
[0055] For the aligned point cloud, the depth threshold can be set based on Remove the foreground and background. Specifically, the aligned point cloud can be depth-truncated according to the depth threshold, and the truncated point cloud is recorded as P c ' olor , which can be in the form of:
[0056]
[0057] Among them, z is the depth value of the pixel point, is the first depth threshold, is the second depth threshold. After removing the background and foreground, the point cloud can be rearranged so that the point cloud is consistent with the shape and size of the RGB image, thereby obtaining a rearranged point cloud corresponding to the RGB image. For example, if the resolution of the RGB image is W×H, W represents the length of the RGB image and H represents the height of the RGB image, the shape of the point cloud after rearrangement is consistent with the shape of the RGB image, that is:
[0058] P′ color· shape=Image color· shape = (H, W, 3)
[0059] Based on the RGB image and the rearranged point cloud, the target point cloud corresponding to the target object in the rearranged point cloud can be determined.
[0060] First, target detection may be performed on the RGB image, thereby determining at least one to-be-detected object frame from the RGB image; the to-be-detected object frame includes a detection frame of the target object.
[0061] Exemplarily, a prediction model can be obtained through deep learning, and the prediction model is used to detect a detection frame of an object of a corresponding category from an image. Specifically, a plurality of sample images can be collected within the operating range of a terminal device through a visual sensor; a preset model is trained based on the plurality of sample images to obtain a prediction model, and the prediction model is used to identify an object frame corresponding to an object of a target category from an image; then, the prediction model can be used to determine at least one object frame to be detected from an RGB image. Exemplarily, the category of the target object can be a cup, and the category of the cup is determined to be 0, and then input into the prediction model. Based on the prediction model, a detection frame corresponding to an object of category 0 can be identified from the image, that is, an object frame to be detected of an object of the target category corresponding to the target object, and the object frame to be detected has corresponding position information.
[0062] In one possible implementation, the visual sensor can collect 1,000 RGB images within the robot's operating range, and then use 800 of them as training data for training, and the remaining 200 as test data to verify the model. After 500 rounds of training, the model with the best performance on the test set is selected as the final prediction model. For real-time RGB images, after normalization, they are input into the trained model. The prediction results of the prediction model are in the form of: prediction category Prediction confidence p i , Bounding Box Represents the predicted center point, w i ,hi Indicates the width and height of the bounding box. For example, setting the confidence threshold to p thres =0.5, the category of the object to be detected is 0, then:
[0063]
[0064] Since the object to be detected frame corresponds to the target object's category, the object to be detected frame includes the target object's detection frame. After determining the object to be detected frame, the target object's detection frame can be determined based on the rearranged point cloud. Since the RGB image and the rearranged point cloud are already aligned, the corresponding object to be detected point cloud frame can be determined from the rearranged point cloud based on the corresponding position of the object to be detected frame.
[0065] Based on the depth image, the target point cloud frame corresponding to the target object can be determined from the point cloud frame to be detected. For example, the average depth value corresponding to the point cloud frame to be detected can be determined. The average depth value can be the average of the depth values of each pixel in the point cloud frame to be detected whose depth value is not zero; then the point cloud frame to be detected whose average depth value is within a preset range is used as the intermediate point cloud frame; if there is only one intermediate point cloud frame, the intermediate point cloud frame can be directly determined as the target point cloud frame; or, if there are multiple intermediate point cloud frames, the target point cloud frame can be determined from the multiple intermediate point cloud frames based on the imaging quality.
[0066] Image quality can be characterized by imaging ratio and imaging uniformity. The ratio of pixels with non-zero depth values in the middle point cloud frame is used as the imaging ratio. Specifically, the number of pixels with non-zero depth values in the middle point cloud frame can be determined. Based on the length and width of the middle point cloud frame, the total number of pixels corresponding to the middle point cloud frame can be calculated. The number of pixels with non-zero depth values divided by the total number of pixels is used as the imaging ratio.
[0067] Generally, the target object is usually a connected point cloud area. Therefore, the pixels in the middle point cloud frame can be clustered to obtain the number of categories corresponding to multiple pixels in the middle point cloud frame. The number of categories is used to represent the imaging uniformity. For example, if the number of categories is large, it can be determined that the object in the middle point cloud frame is fragmented and does not belong to the target object.
[0068] Based on this, the intermediate point cloud frame with an imaging ratio greater than a preset threshold and a number of categories less than a preset value can be determined as the target point cloud frame. Of course, if there are multiple intermediate point cloud frames with an imaging ratio greater than the preset threshold and a number of categories less than the preset value, the intermediate point cloud frame with the largest imaging ratio or the lowest number of categories can be selected as the target point cloud frame.
[0069] In a possible implementation, the object frame to be detected can be mapped to the rearranged point cloud, and the object frame to be detected can be mapped to the three-dimensional space to obtain the point cloud information corresponding to the object frame to be detected. The cropped point cloud frame to be detected is recorded as P″ color .
[0070] Calculate the average depth of pixels whose depth values are not 0 within each point cloud frame to be detected. The formula is as follows:
[0071]
[0072] Where n represents z i The sum of pixels that are ≠0, z i is the depth value of the pixel.
[0073] Candidate detection frames are further screened based on the average depth. Generally, they can be selected in combination with the motion range of the terminal device itself. The depth detection range suitable for the visual sensor is set to
[0074]
[0075] The depth imaging quality of the filtered point cloud frame to be detected is judged. The quality indicators for judging the depth imaging quality may include imaging ratio and imaging uniformity. The imaging ratio is the proportion of pixels with effective depth in the frame. The imaging ratio may have a threshold, for example, the threshold is set to 0.8. The imaging uniformity can be characterized by the clustering results of the point cloud in the point cloud frame to be detected. For example, the point cloud in the point cloud frame to be detected is clustered and segmented. If the clustering result is ≥2, there is a break in the point cloud imaging in the frame. Therefore, the point cloud frame to be detected with an imaging ratio greater than 0.8 and a clustering result of 1 can be used as the target point cloud frame B. * Then use the target point cloud box B * Perform positioning.
[0076] The point cloud in the target point cloud frame is the target point cloud corresponding to the target object. The target point cloud is equivalent to the collection of all points on the outer surface of the target object.
[0077] After determining the target point cloud, the target point cloud can be filtered to remove noise points in the target point cloud. Noise points can include outliers and cluttered points. Point cloud filtering can include point cloud sphere radius filtering and point cloud statistical filtering. Point cloud statistical filtering is used to remove obvious outliers in the target point cloud. Obvious outliers are sparse and have low information density, so they can be filtered. Point cloud statistical filtering can define a point cloud as invalid if it is less than a certain density. The average distance from each point in the target point cloud to its k nearest points is calculated. Based on a given mean and variance, points outside the given variance around the point can be removed, thereby removing obvious outliers. Point cloud sphere radius filtering can draw a sphere centered on a point and calculate the number of points falling within the sphere. If the number is greater than a given value, the point is retained; if the number is less than a given value, the point is removed. Point cloud sphere radius filtering is used to remove cluttered points in the target point cloud.
[0078] For example, for B * The 3D point cloud mapped in the frame is filtered to remove outliers and cluttered points, and finally a clean point cloud data of the object surface is obtained, which is recorded as P * .
[0079] S103: Determine the center of mass and main direction of the target object according to the target point cloud.
[0080] Determine the target pixel points with non-zero depth values in the target point cloud. The target pixel points are the locations on the surface of the target object. Based on the vector coordinates of each target pixel point, the center pixel point can be determined, and the center pixel point coordinates are used as the center of mass vector coordinates of the target object. The center of mass vector coordinates are used to represent the location of the center of mass. For example, the center of mass vector coordinates of the target object can be determined based on the vector coordinates of each target pixel point through the following formula:
[0081]
[0082] in, is the coordinate of the center of mass vector, p i is the vector coordinate of the i-th target pixel, and n is the number of target pixels.
[0083] Based on each target pixel point, the rotation matrix corresponding to the target object can be calculated. Specifically, the covariance matrix can be determined based on the vector coordinates of each target pixel point. The covariance matrix can be determined based on the vector coordinates of each target pixel point using the following formula:
[0084]
[0085] Where E is the covariance matrix, is the coordinate of the center of mass vector, p iis the vector coordinate of the i-th target pixel, and n is the number of target pixels. Determine the eigenvalues and eigenvectors of the covariance matrix; use the eigenvector corresponding to the maximum value among the eigenvalues as the target eigenvector; perform an orthogonal calculation on the target eigenvector to obtain an orthogonal basis for the target eigenvector in vector space; and calculate the rotation matrix between the image coordinate system and the coordinate axes corresponding to the orthogonal basis.
[0086] The rotation vector is obtained based on the rotation matrix transformation, and the rotation vector is used to represent the main direction of the target object.
[0087] S104: Position the target object according to the centroid and the main direction.
[0088] The center of mass and the principal directions are combined to determine the pose of the target object, enabling its localization. The pose can include both position and posture. The center of mass vector coordinates represent the position of the target object, while the principal directions represent its posture.
[0089] In this embodiment, when performing visual positioning, target detection can be performed based on RGB images, depth images and point clouds, thereby reducing the difficulty of target detection and improving the accuracy of target detection; the complex point cloud contour extraction algorithm is converted into image detection combined with point cloud for data processing, which facilitates data transplantation and integration, has a small number of parameters and strong versatility.
[0090] Reference Figure 2 , shows a schematic flow chart of the steps of another robot control method provided by an embodiment of the present application. The robot may include a mechanical platform and a mechanical arm, and may specifically include the following steps:
[0091] S201 : Determine first positioning data, where the target object is a working object of the robot, and the first positioning data is used to represent a positional relationship of the mechanical platform relative to the target object.
[0092] The mechanical platform can be equipped with an active optical binocular depth camera, and the end of the robotic arm can be equipped with a monocular depth camera.
[0093] Based on the active optical binocular depth camera, the visual positioning method in the previous embodiment can be used to determine the first positioning data. The first positioning data is the visual positioning of the mechanical platform relative to the target object, which is used to characterize the positional relationship of the mechanical platform relative to the target object and is coarse positioning data.
[0094] S202: Control the mechanical platform to move to a position corresponding to the first positioning data.
[0095] The robot controls the mechanical platform to move according to the first positioning data, so that the robot approaches the target object and moves into the operating range of the mechanical arm.
[0096] S203: Determine second positioning data, where the second positioning data is used to represent a positional relationship of the robotic arm relative to the target object after the robotic platform reaches a predetermined position.
[0097] The predetermined position may be a position corresponding to the first positioning data. After the mechanical platform reaches the predetermined position, the robot may determine the second positioning data of the mechanical arm relative to the target object.
[0098] Based on the monocular depth camera, the visual positioning method in the previous embodiment can be used to determine the second positioning data. The second positioning data is the visual positioning of the robotic arm relative to the target object, which is used to characterize the position relationship of the robotic arm relative to the target object after the mechanical platform reaches the position corresponding to the first positioning data, and is the precise positioning data.
[0099] S204: Control the robotic arm to operate on the target object according to the second positioning data.
[0100] The robot controls the movement of the robotic arm according to the second positioning data, so that the robotic arm operates on the target object.
[0101] The positioning accuracy of the above-mentioned first positioning data is lower than the positioning accuracy of the second positioning data. The first positioning data is coarse positioning data, which is used to move the robot close to the target object; the second positioning data is fine positioning data, which is used to control the robot arm to operate the target object.
[0102] In the embodiments of the present application, determining the first positioning data is equivalent to detecting a target at a long distance, while determining the second positioning data is equivalent to detecting a target at a close distance. In one possible implementation, target recognition can be performed based on a multi-scale recognition neural network. During the process of determining the positioning data, the scale of the neural network can be adjusted through parameters, thereby enabling the neural network to recognize targets in images at different distances.
[0103] In the embodiments of this application, a binocular depth camera combined with a monocular depth camera achieves precise three-dimensional positioning from far to near, improving positioning precision and operational accuracy. Furthermore, when performing visual positioning, the center of mass direction of the object under test can be quickly extracted by integrating image processing and point cloud processing. This approach has low hardware requirements and is versatile, thus facilitating the control of outdoor operating robots.
[0104] It should be noted that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0105] Reference Figure 3, shows a flow chart of another robot control method provided by an embodiment of the present application, in Figure 3 In the method shown, the robot can be equipped with an active optical binocular depth camera and a monocular depth camera. The active optical binocular depth camera can be installed on the mechanical platform itself, and the monocular depth camera can be installed at the end of the robotic arm. The detection range of the active optical binocular depth camera can be 0.4-4m, and the detection range of the monocular depth camera is 0.2-1.5m. It can be seen that the detection distance of the active optical depth camera is relatively long, and the detection range of the monocular depth camera is relatively small. In other words, the active optical depth camera is suitable for long-distance detection, and the monocular depth camera is suitable for close-range detection. When the robot is far away from the target object, the target object can be located based on the active optical depth camera of the mechanical platform, thereby obtaining coarse positioning data of the target object. Based on the coarse positioning data, the robot controls the mechanical platform to move to a position closer to the target object. When the robot is close to the target object, the monocular depth camera can be used for fine positioning, and the robotic arm can be controlled to operate the target object based on the fine positioning data.
[0106] When performing coarse positioning based on an active optical binocular depth camera, the active optical depth camera can be used to acquire data streams, which can include RGB images, depth images, and point clouds. The image coordinate system and the depth coordinate system are then aligned.
[0107] Align the point cloud obtained in the depth coordinate system to the RGB image coordinate system, and record the depth coordinate system as T depth , the corresponding point cloud in the depth coordinate system is P depth , the RGB image coordinate system is T color According to the binocular camera calibration, the conversion matrix from the depth coordinate system to the RGB image coordinate system can be obtained as follows: The expression of the aligned point cloud in the camera coordinate system is:
[0108]
[0109] When processing data, dual-thread processing can be enabled for RGB images and point cloud data, including deep learning image processing threads and point cloud data processing threads.
[0110] The deep learning image processing thread may include: using a depth camera to collect 1000 RGB images within the robot arm's operating range, 800 of which are used as training data for training, and the remaining 200 as test data to verify the model. After 500 rounds of training, the model with the best performance on the test set is selected as the final model. For real-time RGB images, they are normalized and then input into the trained model. The prediction results of the model can be in the form of: predicted category Prediction confidence pi , Bounding Box Represents the predicted center point, w i , h i Indicates the width and height of the bounding box.
[0111] The point cloud data processing thread may include: setting the initial depth threshold to A preliminary depth threshold is used to remove the foreground and background, and the aligned point cloud is depth-truncated according to the depth threshold. The truncated point cloud is recorded as P c ' olor , which has the form:
[0112]
[0113] The data obtained by the deep learning image processing thread and the point cloud data processing thread can be further filtered and processed. Specifically, the candidate boxes can be filtered from the prediction results obtained by the deep learning image processing thread, and the confidence threshold is set to p thres =0.5, the category of the object to be detected is 0, then:
[0114]
[0115] The resolution of the RGB image is W×H, where W represents the length of the image and H represents the height of the image. The point cloud obtained by the point cloud data processing thread is rearranged so that the shape of the rearranged point cloud is consistent with the shape of the RGB image:
[0116] P′ color· shape=Image color· shape = (H, W, 3)
[0117] Apply the candidate bounding box to the rearranged point cloud, map the bounding box to the three-dimensional space to obtain the corresponding point cloud information within the box, and the cropped point cloud is recorded as P″ color .
[0118] The average depth of each box in the truncated point cloud is calculated as follows:
[0119]
[0120] Where n represents z i The sum of pixels that are ≠0 (valid value).
[0121] The candidate detection frames are further screened based on the average depth. Generally, the selection is made based on the range of motion of the robot arm itself. The depth detection range of the camera suitable for the robot arm is set to
[0122]
[0123] The depth imaging quality of the retained bounding boxes is assessed, with quality metrics including imaging ratio and imaging uniformity. Imaging ratio refers to the number of pixels with valid depth within the candidate bounding box exceeding a certain threshold, set at 0.8. Imaging uniformity involves clustering and segmenting the point cloud within the candidate box. If the clustering result is ≥ 2, the point cloud within the box is considered to have discontinuities.
[0124] The bounding box that meets the above indicators is used as the final bounding box B corresponding to the object * Conduct positioning;
[0125] To B * The 3D point cloud mapped in the frame is filtered to remove outliers and cluttered points, including: point cloud sphere radius filtering; point cloud statistical filtering. Finally, the clean object surface point cloud data is obtained, denoted as P * .
[0126] For the acquired point cloud P * Extract the center of mass and direction. The specific steps include:
[0127] Centroid calculation formula: where p k is a three-dimensional vector represented by (x, y, z);
[0128] Main direction calculation process:
[0129] Calculate the covariance matrix, the calculation formula is: Where N represents P * the number of midpoints;
[0130] Calculate the eigenvalues and eigenvectors of the covariance matrix and solve the equation: where λ j represents the j-th eigenvalue of the covariance matrix, is the jth eigenvector;
[0131] The eigenvalue λ j Sort by, λ0≥λ1≥λ2;
[0132] Keep the eigenvector corresponding to the largest eigenvalue
[0133] right Perform Schmid orthogonalization calculation to obtain a set of orthogonal bases in the vector space, that is, a set of coordinate axes constructed along the main directions of the object;
[0134] Project the (x, y, z) coordinates of each point onto the above coordinate axes to obtain a compact bounding box for the entire point cloud;
[0135] Calculate the rotation matrix between the above coordinate axes and the camera's own coordinate system, convert it into a rotation vector, and use it as the main direction of the point cloud.
[0136] The obtained center of mass and main direction are combined together as the coarse positioning posture of the robot operation to guide the robot to move.
[0137] When the robot reaches the designated position, the active optical binocular depth camera is switched to a monocular depth camera to obtain the corresponding depth data stream. The above steps are repeated to finally obtain the object's precise 3D positioning posture. The robotic arm is then controlled based on this precise positioning data.
[0138] An embodiment of the present application further provides a robot comprising a mechanical platform and a mechanical arm, wherein the robot locates a target object using the method described in Example 1; the robot is controlled using the method described in Example 2 to perform operations on the target object. Both the mechanical platform and the mechanical arm may be equipped with visual sensors. For example, the mechanical platform may include an active optical binocular depth camera, and the end of the mechanical arm may be equipped with a monocular depth camera. Based on the visual sensor, data streams may be collected, and visual positioning may be performed based on the collected data streams.
[0139] The present application also provides a robotic arm equipped with a monocular depth camera. The robotic arm locates a target object using the method described in Example 1 to perform operations on the target object. The robotic arm can be equipped with a visual sensor, such as a camera. The present application does not limit the type of visual sensor.
[0140] Reference Figure 4 , shows a schematic diagram of a visual positioning device provided in an embodiment of the present application, which may specifically include a visual data acquisition module 41, a target point cloud determination module 42, a center of mass and direction determination module 43, and a positioning module 44, wherein:
[0141] A visual data acquisition module 41 is configured to acquire a data stream through a visual sensor camera, wherein the data stream includes an RGB image, a depth image, and a point cloud, and the data stream includes data of a target object to be located;
[0142] a target point cloud determining module 42, configured to determine a target point cloud corresponding to the target object from the point cloud according to the image, the RGB image, and the depth image;
[0143] A center of mass and direction determination module 43 is used to determine the center of mass and main direction of the target object based on the target point cloud;
[0144] The positioning module 44 is configured to locate the target object according to the centroid and the main direction.
[0145] In a possible implementation, the target point cloud determination module 42 includes:
[0146] a rearrangement submodule, configured to determine a rearranged point cloud based on the RGB image and the point cloud, wherein a shape of the rearranged point cloud is consistent with a shape of the RGB image;
[0147] A submodule for determining a frame of an object to be detected, configured to determine at least one frame of an object to be detected from the RGB image;
[0148] A submodule for determining a point cloud frame to be detected, configured to determine a point cloud frame to be detected corresponding to the object frame to be detected from the rearranged point cloud;
[0149] A target point cloud frame determination submodule is used to determine a target point cloud frame corresponding to a target object from the point cloud frame to be detected based on the depth image;
[0150] The target point cloud determination submodule is used to use the point cloud in the target point cloud frame as the target point cloud.
[0151] In a possible implementation, the rearrangement submodule includes:
[0152] a depth point cloud determining unit, configured to determine a depth point cloud in a depth coordinate system, wherein the depth coordinate system is established based on the depth image, and each pixel in the depth point cloud has a depth value;
[0153] a visual point cloud determining unit, configured to convert the depth point cloud into an image coordinate system to obtain a visual point cloud, wherein the image coordinate system is established based on the RGB image;
[0154] a foreground and background deletion unit, configured to delete foreground pixels and background pixels in the visual point cloud, wherein the foreground pixels are pixels having a depth value less than a first preset depth threshold, and the background pixels are pixels having a depth value greater than a second preset depth threshold;
[0155] A rearrangement unit is used to rearrange the visual point cloud to obtain the rearranged point cloud.
[0156] In a possible implementation, the above-mentioned submodule for determining the frame of the object to be detected includes:
[0157] A sample image acquisition unit, configured to acquire a plurality of sample images within the operating range of the terminal device through the visual sensor;
[0158] A prediction model training unit, configured to train a preset model based on a plurality of sample images to obtain a prediction model, wherein the prediction model is used to identify an object frame corresponding to an object of a target category from an image;
[0159] The to-be-detected object frame determining unit is configured to determine at least one to-be-detected object frame from the RGB image according to the prediction model.
[0160] In a possible implementation, the target point cloud frame determination submodule includes:
[0161] An average depth value determining unit, configured to determine an average depth value corresponding to the point cloud frame to be detected, wherein the average depth value is an average of the depth values of each pixel whose depth value is not zero in the point cloud frame to be detected;
[0162] an intermediate point cloud frame determining unit, configured to take the point cloud frame to be detected whose average depth value is within a preset range as an intermediate point cloud frame;
[0163] The target point cloud frame determination unit is used to determine the intermediate point cloud frame as the target point cloud frame if there is only one intermediate point cloud frame; or, if there are multiple intermediate point cloud frames, determine the target point cloud frame from the multiple intermediate point cloud frames according to imaging quality.
[0164] In a possible implementation, the imaging quality is characterized by imaging ratio and imaging uniformity, and the target point cloud frame determination unit includes:
[0165] an imaging ratio determination subunit, configured to take the ratio of pixels having non-zero depth values in the intermediate point cloud frame as the imaging ratio;
[0166] a category number determination subunit, configured to cluster the pixel points in the intermediate point cloud frame to obtain the category number corresponding to the plurality of pixel points in the intermediate point cloud frame, wherein the category number is used to characterize the imaging uniformity;
[0167] The target point cloud frame determination subunit is configured to determine the intermediate point cloud frame whose imaging ratio is greater than a preset threshold and whose category quantity is less than a preset value as the target point cloud frame.
[0168] In a possible implementation, the apparatus further includes:
[0169] The filtering module is used to perform filtering processing on the target point cloud to filter out noise points in the target point cloud.
[0170] In a possible implementation, the centroid and direction determination module 43 includes:
[0171] A target pixel point determination submodule is used to determine target pixel points with non-zero depth values in the target point cloud;
[0172] a centroid vector coordinate determination submodule, configured to determine the centroid vector coordinates of the target object based on the vector coordinates of each of the target pixel points, wherein the centroid vector coordinates are used to represent the position of the centroid;
[0173] A rotation matrix determination submodule, configured to calculate the rotation matrix corresponding to the target object based on each of the target pixel points;
[0174] The rotation vector determination submodule is configured to obtain a rotation vector based on the rotation matrix conversion, where the rotation vector is used to represent the main direction of the target object.
[0175] In one possible implementation, the centroid vector coordinate determination submodule includes:
[0176] The centroid vector coordinate determining unit is configured to determine the centroid vector coordinates of the target object based on the vector coordinates of each target pixel point using the following formula:
[0177]
[0178] in, is the coordinate of the centroid vector, p i is the vector coordinate of the i-th target pixel point, and n is the number of the target pixels.
[0179] In one possible implementation, the rotation matrix determination submodule includes:
[0180] a covariance matrix determining unit, configured to determine a covariance matrix based on the vector coordinates of each of the target pixel points;
[0181] a characteristic data determination unit, configured to determine the eigenvalues and eigenvectors of the covariance matrix;
[0182] a target eigenvector determining unit, configured to take the eigenvector corresponding to the maximum value among the eigenvalues as the target eigenvector;
[0183] an orthogonal basis determination unit, configured to perform an orthogonal calculation on the target eigenvector to obtain an orthogonal basis of the target eigenvector in a vector space;
[0184] The rotation matrix determination unit is used to calculate the rotation matrix between the camera coordinate system and the coordinate axes corresponding to the orthogonal basis.
[0185] In a possible implementation, the covariance matrix determining unit includes:
[0186] The covariance matrix determination subunit is configured to determine the covariance matrix based on the vector coordinates of each target pixel point using the following formula:
[0187]
[0188] Wherein, E is the covariance matrix, is the coordinate of the centroid vector, p i is the vector coordinate of the i-th target pixel point, and n is the number of the target pixels.
[0189] Reference Figure 5 , shows a schematic diagram of a visual positioning device provided in an embodiment of the present application. The device can be applied to a robot, wherein the robot includes a mechanical platform and a mechanical arm. The device can specifically include a first positioning module 51, a first control module 52, a second positioning module 53, and a second control module 54, wherein:
[0190] A first positioning module 51 is configured to determine first positioning data, wherein the target object is a working object of the robot, and the first positioning data is configured to represent a positional relationship between the mechanical platform and the target object;
[0191] A first control module 52 is configured to control the mechanical platform to move to a position corresponding to the first positioning data;
[0192] A second positioning module 53 is used to determine second positioning data, where the second positioning data is used to represent the position relationship of the robotic arm relative to the target object after the robotic platform reaches a predetermined position;
[0193] a second control module 54, configured to control the robotic arm to operate on the target object according to the second positioning data;
[0194] The positioning accuracy of the first positioning data is higher than the positioning accuracy of the second positioning data, and the robot determines the first positioning data and the second positioning data by the method described in the first aspect above.
[0195] In one possible implementation, an active optical binocular depth camera is installed on the mechanical platform of the robot, and a monocular depth camera is installed on the mechanical arm. The active optical binocular depth camera is used to collect data streams when determining the first positioning data, and the monocular depth camera is used to collect data streams when determining the second positioning data.
[0196] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment part.
[0197] Figure 6 This is a schematic diagram of the structure of a terminal device provided in an embodiment of the present application. Figure 6 As shown, the terminal device 6 of this embodiment includes: at least one processor 60 ( Figure 6Only one is shown in the figure) a processor, a memory 61, and a computer program 62 stored in the memory 61 and executable on the at least one processor 60, wherein the processor 60 implements the steps of any of the above-mentioned method embodiments when executing the computer program 62.
[0198] The terminal device may be a robot or other intelligent device. The terminal device may include, but is not limited to, a processor 60 and a memory 61. It will be understood by those skilled in the art that Figure 6 It is only an example of the terminal device 6 and does not constitute a limitation on the terminal device 6. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.
[0199] The processor 60 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.
[0200] In some embodiments, the memory 61 may be an internal storage unit of the terminal device 6, such as a hard disk or memory of the terminal device 6. In other embodiments, the memory 61 may also be an external storage device of the terminal device 6, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 6. Furthermore, the memory 61 may include both an internal storage unit of the terminal device 6 and an external storage device. The memory 61 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 61 may also be used to temporarily store data that has been output or is about to be output.
[0201] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.
[0202] An embodiment of the present application provides a computer program product. When the computer program product is run on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0203] The above embodiments are intended only to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they may still modify the technical solutions described in the above embodiments or replace some of the technical features therein with equivalents; and such modifications or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present application and should be included within the scope of protection of the present application.
Claims
1. A visual positioning method, characterized in that: The method comprises: Collecting a data stream through a visual sensor, wherein the data stream includes an RGB image, a depth image, and a point cloud, and the data stream includes data of a target object to be located; Determining a target point cloud corresponding to the target object from the point cloud according to the RGB image and the depth image; Determining the center of mass and main direction of the target object according to the target point cloud; Positioning the target object according to the centroid and the main direction; Determining a target point cloud corresponding to the target object from the point cloud according to the RGB image and the depth image includes: Determining a rearranged point cloud according to the RGB image and the point cloud, wherein a shape of the rearranged point cloud is consistent with a shape of the RGB image; Determine at least one object frame to be detected from the RGB image; Determine a point cloud frame to be detected corresponding to the object frame to be detected from the rearranged point cloud; Determine an average depth value corresponding to the point cloud frame to be detected, where the average depth value is the average of the depth values of each pixel whose depth value is not zero in the point cloud frame to be detected; The point cloud frame to be detected whose average depth value is within a preset range is used as an intermediate point cloud frame; If there is only one intermediate point cloud frame, the intermediate point cloud frame is determined as the target point cloud frame; or, if there are multiple intermediate point cloud frames, the target point cloud frame is determined from the multiple intermediate point cloud frames according to imaging quality; The point cloud in the target point cloud frame is used as the target point cloud; The imaging quality is characterized by imaging ratio and imaging uniformity, and determining the target point cloud frame from the plurality of intermediate point cloud frames according to the imaging quality includes: The proportion of pixels with non-zero depth values in the middle point cloud frame is used as the imaging ratio; Clustering the pixel points in the middle point cloud frame to obtain the number of categories corresponding to the plurality of pixel points in the middle point cloud frame, wherein the number of categories is used to characterize the imaging uniformity; The intermediate point cloud frame whose imaging ratio is greater than a preset threshold and whose category quantity is less than a preset value is determined as the target point cloud frame.
2. The method according to claim 1, wherein The determining, based on the RGB image and the point cloud, to rearrange the point cloud comprises: Determine a depth point cloud in a depth coordinate system, wherein the depth coordinate system is established based on the depth image, and each pixel in the depth point cloud has a depth value; Converting the depth point cloud into an image coordinate system to obtain a visual point cloud, wherein the image coordinate system is established based on the RGB image; Deleting foreground pixels and background pixels in the visual point cloud, wherein the foreground pixels are pixels having a depth value less than a first preset depth threshold, and the background pixels are pixels having a depth value greater than a second preset depth threshold; The visual point cloud is rearranged to obtain the rearranged point cloud.
3. The method according to claim 1, wherein Determining at least one to-be-detected object frame from the RGB image includes: Collecting a plurality of sample images by the visual sensor; Training a preset model based on the plurality of sample images to obtain a prediction model, wherein the prediction model is used to identify an object frame corresponding to an object of a target category from an image; At least one frame of the object to be detected is determined from the RGB image according to the prediction model.
4. The method according to any one of claims 1 to 3, wherein After the step of determining a target point cloud corresponding to the target object from the point cloud based on the image, the RGB image, and the depth image, and before the step of determining the center of mass and the main direction of the target object based on the target point cloud, the method further includes: The target point cloud is filtered to filter out noise points in the target point cloud.
5. The method according to claim 4, wherein Determining the center of mass and main direction of the target object according to the target point cloud includes: Determine a target pixel point having a non-zero depth value in the target point cloud; Based on the vector coordinates of each of the target pixel points, determining the center of mass vector coordinates of the target object, wherein the center of mass vector coordinates are used to represent the position of the center of mass; and calculating the rotation matrix corresponding to the target object according to each of the target pixel points; A rotation vector is obtained based on the rotation matrix conversion, and the rotation vector is used to represent the main direction of the target object.
6. The method according to claim 5, wherein The determining of the center of mass vector coordinates of the target object based on the vector coordinates of each of the target pixel points includes: The centroid vector coordinates of the target object are determined based on the vector coordinates of each target pixel point using the following formula: in, is the coordinate of the centroid vector, is the vector coordinate of the i-th target pixel point, and n is the number of the target pixels.
7. The method according to claim 5, wherein Calculating the rotation matrix corresponding to the target object according to each of the target pixels includes: Determining a covariance matrix based on the vector coordinates of each of the target pixel points; determining the eigenvalues and eigenvectors of the covariance matrix; The eigenvector corresponding to the maximum value among the eigenvalues is used as the target eigenvector; Performing an orthogonal calculation on the target eigenvector to obtain an orthogonal basis of the target eigenvector in a vector space; The rotation matrix between the image coordinate system and the coordinate axes corresponding to the orthogonal basis is calculated.
8. The method according to claim 7, wherein The determining of the covariance matrix based on the vector coordinates of each of the target pixel points includes: The covariance matrix is determined based on the vector coordinates of each target pixel point using the following formula: , Wherein, E is the covariance matrix, is the coordinate of the centroid vector, is the vector coordinate of the i-th target pixel point, and n is the number of the target pixels.
9. A robot control method, characterized in that: Applied to a robot, the robot comprising a mechanical platform and a mechanical arm, the method comprising: Determining first positioning data, where the target object is a working object of the robot, and the first positioning data is used to represent a positional relationship of the mechanical platform relative to the target object; Controlling the mechanical platform to move to a position corresponding to the first positioning data; Determining second positioning data, where the second positioning data is used to represent a positional relationship of the robotic arm relative to the target object after the robotic platform reaches a predetermined position; controlling the robotic arm to operate on the target object according to the second positioning data; The positioning accuracy of the first positioning data is lower than the positioning accuracy of the second positioning data, and the robot determines the first positioning data and the second positioning data by the method according to any one of claims 1 to 8.
10. The method according to claim 9, wherein An active optical binocular depth camera is installed on the mechanical platform of the robot, and a monocular depth camera is installed on the mechanical arm. The active optical binocular depth camera is used to collect data streams when determining the first positioning data, and the monocular depth camera is used to collect data streams when determining the second positioning data.
11. A robot, characterized in that: The robot includes a mechanical platform and a mechanical arm, and the robot locates the target object by the method described in any one of claims 1 to 8; the robot is controlled by the method described in any one of claims 9 to 10 to achieve operation on the target object.
12. A robotic arm, characterized in that: The robotic arm positions the target object by the method according to any one of claims 1 to 8, so as to perform operations on the target object.
13. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 8 or 9 to 10 is implemented.
14. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 or 9 to 10 is implemented.
Citation Information
Patent Citations
Visual recognition and positioning method for robot intelligent capture application
CN108171748A
Material sorting method and device based on three-dimensional vision
CN112509145A