Mechanical arm visual control system based on computing power chip

Image data is collected through multiple industrial cameras, combined with image preprocessing and improved YOLOv8 model and environmental modeling algorithm, the detection accuracy problem of the robotic arm visual positioning system in dynamic occlusion scenarios is solved, and the precise three-dimensional pose information acquisition and grabbing position of the target object is achieved.

CN120382500APending Publication Date: 2025-07-29CHANGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510807393.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The detection accuracy of the existing robotic arm visual positioning system has decreased in dynamic occlusion scenarios, making it difficult to accurately obtain the three-dimensional position information of the target object, resulting in poor accuracy of the grab position and angle, especially when facing objects with complex shapes and different placement postures, the grab failure rate is high.

Method used

Multi-channel industrial cameras are used to collect image data, and target data sets are constructed through image graying and denoising processing. Feature extraction is used using the improved YOLOv8 object detection model, and three-dimensional modeling is carried out in combination with environmental modeling algorithms to build motion paths, and precise motion planning is carried out in combination with kinematic models.

Benefits of technology

Improve the accuracy of the grab position and angle, and achieve accurate acquisition and stable grabbing effect of the target object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120382500A_ABST
    Figure CN120382500A_ABST
Patent Text Reader

Abstract

The invention relates to a mechanical arm visual control system based on a computing power chip, and the system comprises the steps: S1, collecting image data in an actual operation environment through a plurality of industrial cameras, and obtaining an initial data set; s2, preprocessing the training set by adopting an image graying and denoising mode to obtain a target data set; s3, using the YOLOv8 target detection model as a reference model to construct an improved YOLOv8 target detection model as a mechanical arm visual detection network model; s4, performing three-dimensional modeling on the training set by adopting an environment modeling algorithm, and constructing a motion path; s5, the test set is sent into the constructed system to complete grabbing of the mechanical arm; through a visual detection and three-dimensional modeling method, image recognition is carried out on a grabbed object, target three-dimensional pose information is accurately obtained, then accurate motion planning is carried out in combination with a kinematics model, a grabbing track is optimized, and the effect of effectively improving the accuracy of the grabbing position and angle is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent detection technology, and particularly to a robotic arm vision control system based on a computing power chip. Background Art

[0002] A robotic arm consists of an arm, a wrist, a gripper, a drive system, and a control system. The arm can be telescoped, bent, and rotated to achieve large-range movement; the wrist can move flexibly to ensure precise attitude adjustment of the gripper; the gripper has various forms to meet the requirements of grasping, handling, operating, etc. The drive system includes motors, cylinders, reducers, and transmission mechanisms to ensure precise and reliable movement of each joint. The control system, as the core, has sensors, controllers, software algorithms, etc., and is responsible for motion planning, trajectory control, force control, and vision servo.

[0003] Currently, robotic arm vision is a technology that obtains image data of the environment and target objects through vision sensors (such as industrial cameras, depth cameras, etc.) and extracts information using image processing and pattern recognition technologies. It provides key information such as the position, attitude, and shape of the target object for the robotic arm, enabling the robotic arm to perform precise grasping, handling, assembly, etc. operations. It also has the ability of real-time feedback and adaptive adjustment, which can help the robotic arm cope with complex and changing environments and task requirements.

[0004] However, the detection accuracy of existing robotic arm vision positioning systems drops significantly in dynamic occlusion scenarios, making it difficult to accurately obtain the three-dimensional pose information of the target object, resulting in poor accuracy of the grasping position and angle. Especially when facing objects with complex shapes and different placement postures, the grasping failure rate is relatively high.

[0005] Therefore, the present invention proposes a robotic arm vision control system based on a computing power chip. Summary of the Invention

[0006] Based on the above problems existing in the prior art, the purpose of the embodiments of the present invention is to provide a robotic arm vision control system based on a computing power chip.

[0007] To achieve the above purpose, the technical solution adopted by the present invention is: A robotic arm vision control system based on a computing power chip, comprising: S1, collecting image data in the actual operation environment through multiple industrial cameras to obtain an initial data set; S2, preprocessing the training set by image grayscale conversion and denoising to obtain a target data set; S3, constructing an improved YOLOv8 target detection model as the robotic arm vision detection network model with the YOLOv8 target detection model as the benchmark model; S4, adopting an environmental modeling algorithm to perform three-dimensional modeling according to the training set and construct a motion path; S5. Feed the test set into the constructed system to complete the grasping of the robotic arm.

[0008] The present invention is further configured such that in S2, the training set is preprocessed by means of image grayscale conversion and denoising to obtain a target data set, including: performing grayscale conversion on the collected image, converting the color image into a grayscale image, separating the image from the environment to extract target features, and then using a filter to reduce the amount of unwanted noise in the image. After the training set is preprocessed, a target data set is obtained.

[0009] The present invention is further configured such that in digital image processing, first, the images in the training set are converted into grayscale images, and then Gaussian filtering is used to perform weighted averaging on the entire image. By performing weighted averaging on its own and other pixel values in the neighborhood, a target data set is obtained.

[0010] The present invention is further configured such that an improved YOLOv8 target detection model is constructed based on the YOLOv8 target detection model as the robotic arm vision detection network model, including: Step S31. Perform lightweight design on the YOLOv8 model, and quantize the model parameters from the traditional 32-bit floating-point representation to 16-bit fixed-point representation; Step S32. By introducing an adaptive feature pyramid network (AFPN), fuse feature maps of different scales.

[0011] The present invention is further configured such that introducing an adaptive feature pyramid network (AFPN) to fuse feature maps of different scales includes: AFPN has a learnable weighting mechanism, and dynamically adjusts the contribution ratio of features of different scales through the learnable weight mechanism. Feature extraction is performed on the target data set through the adaptive feature pyramid network, and then the extracted feature maps are fused. The weights of different feature maps are dynamically allocated according to the scale and position of the target, and further convolution operations and activation function processing are performed on the fused feature maps to enhance the expression ability of the features. The adaptive feature pyramid network performs upsampling and downsampling operations on the feature maps, thereby obtaining feature maps of different resolutions, and further capturing targets of different sizes in the target data set.

[0012] The present invention is further configured such that in S4, an environmental modeling algorithm is used to perform three-dimensional modeling on the training set to construct a motion path, including: Step S41. Voxelize the depth image and lidar point cloud data according to the spatial position relationship of the pictures in the training set; Step S42. Perform three-dimensional modeling on the filtered point cloud data through a surface reconstruction module.

[0013] The present invention is further configured to voxelize the depth image and the lidar point cloud data according to the spatial position relationship of the pictures in the training set, including: mapping the pixel points in the depth image to the corresponding voxel units, fusing them with the lidar point cloud to obtain fused point cloud data containing color, depth and three-dimensional position information, voxelizing the fused point cloud data, dividing the three-dimensional space into small cubic voxels, calculating the average coordinate value of the points in each voxel, and using this value to represent all the points in the voxel, thereby reducing the density of the point cloud, removing some noise points, and completing the point cloud filtering processing of the training set.

[0014] The present invention is further configured to perform three-dimensional modeling on the filtered point cloud data through a surface reconstruction module, including: considering the point cloud data as the gradient field of a scalar function, estimating the gradient information of the scalar function in the neighborhood of the point cloud data, solving the scalar function in the entire definition domain by solving the Poisson equation, and then obtaining the reconstructed surface through isosurface extraction, and then processing the point cloud data through Poisson reconstruction to generate a closed three-dimensional model.

[0015] The embodiment of the present invention further provides a network-side server, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned computing chip-based robotic arm vision control system.

[0016] An embodiment of the present invention also provides a computer-readable storage medium storing a computer program, which implements the above-mentioned computing chip-based robotic arm vision control system when executed by a processor.

[0017] The beneficial effects of the present invention are: An embodiment of the present invention provides a robotic arm vision control system based on a computing power chip. The system collects image data in an actual operating environment through multiple industrial cameras to obtain an initial data set, and preprocesses the training set by image grayscale and denoising to obtain a target data set. An improved YOLOv8 target detection model is constructed using the YOLOv8 target detection model as a benchmark model as a robotic arm vision detection network model. An environmental modeling algorithm is used to perform three-dimensional modeling on the training set to construct a motion path. The test set is sent into the constructed system to complete the grasping of the robotic arm. Through visual detection and three-dimensional modeling methods, image recognition is performed on the grasped object to accurately obtain the three-dimensional position information of the target. Then, precise motion planning is performed in combination with the kinematic model to optimize the grasping trajectory, thereby effectively improving the accuracy of the grasping position and angle. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The present invention will be further described below with reference to the accompanying drawings and examples.

[0019] In the figure: Figure 1 This is a flow chart of a robotic arm vision control system based on a computing chip in the present invention.

[0020] Figure 2 It is a structural diagram of a network-side server provided according to a second embodiment of the present invention. DETAILED DESCRIPTION

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0022] The embodiments of the present invention provide a robotic arm vision control system based on a computing chip. The embodiment of the present invention collects image data in an actual operating environment through multiple industrial cameras to obtain an initial data set, and preprocesses the training set by image grayscale and denoising to obtain a target data set; an improved YOLOv8 target detection model is constructed with the YOLOv8 target detection model as a benchmark model as a robotic arm vision detection network model; an environmental modeling algorithm is used to perform three-dimensional modeling on the training set to construct a motion path; the test set is sent into the constructed system to complete the grasping of the robotic arm, and image recognition is performed on the grasped object through visual detection and three-dimensional modeling methods to accurately obtain the three-dimensional position information of the target, and then precise motion planning is performed in combination with the kinematic model to optimize the grasping trajectory, thereby effectively improving the accuracy of the grasping position and angle.

[0023] The following is a detailed description of the implementation details of a robotic arm vision control system based on a computing chip in this embodiment. The following content is only for the convenience of understanding the implementation details and is not necessary for the implementation of this solution. This invention utilizes the Rockchip RK3588 chip, integrating dual ISPs (supporting multiple sensor inputs), an NPU (supporting mixed-precision computation), and a multi-core CPU to enable parallel processing of image preprocessing, neural network inference, and motion control algorithms. The dual ISPs support simultaneous processing of data from multiple cameras, enabling panoramic stitching and high-dynamic-range image synthesis, adapting to complex lighting environments.

[0024] See Figure 1The first embodiment of the present invention provides a robotic arm vision control system based on a computing chip, comprising: S1, collects image data in the actual operating environment through multiple industrial cameras to obtain the initial data set.

[0025] Specifically, we use multiple industrial cameras deployed at the worksite to collect image data from the actual operating environment. These cameras, including RGB and depth cameras, capture data from different angles, lighting conditions, and background environments to ensure the diversity and representativeness of the dataset. We also accurately annotate the captured images with information such as the target object's category, bounding box location, and pose information within the image. The initial dataset is divided into training and test sets in an 8:2 ratio.

[0026] S2, preprocess the training set by image grayscale and denoising to obtain the target data set; The collected images are grayscaled, the color images are converted into grayscale images, the images are separated from the environment to extract the target features, and then filters are used to reduce the amount of unnecessary noise in the image. After the training set is preprocessed, the target data set is obtained.

[0027] Specifically, the color of each pixel in a color image is determined by its R, G, and B components, and each component ranges from 0 to 255, so a single pixel can have a range of 256*256*256 possible colors. A grayscale image is a special color image with identical R, G, and B components, where each pixel ranges from 0 to 255. By first converting images in the training set to grayscale in digital image processing, the computational effort for subsequent images is reduced. After grayscale conversion, just like color images, grayscale images can still reflect the distribution and characteristics of the overall and local chromaticity and brightness levels of the entire image.

[0028] The present invention uses Gaussian filtering to process the grayscale image. Gaussian filtering performs weighted averaging on the entire image. By taking weighted averaging of its own pixel values and other pixel values in the neighborhood, in the frequency domain, the filtering process is not interfered by high-frequency signals, thereby achieving a better effect of removing Gaussian noise.

[0029] The target data set is obtained by preprocessing the training set by grayscale processing and filtering.

[0030] S3, using the YOLOv8 target detection model as the baseline model to build an improved YOLOv8 target detection model as the robotic arm visual detection network model; Step S31, performing lightweight design on the YOLOv8 model, quantizing the model parameters from the traditional 32-bit floating point representation to 16-bit fixed point representation; Specifically, the model's network structure is optimized, reducing the number of layers and channels, lowering the model's computational complexity and making it more suitable for running on resource-limited embedded chips. A quantization strategy is also used to optimize the model's parameters. The model parameters are quantized from traditional 32-bit floating-point representation to 8-bit or 16-bit fixed-point representation, thereby reducing the model's storage space and computational complexity.

[0031] In step S32, feature maps of different scales are fused by introducing an adaptive feature pyramid network (AFPN).

[0032] Building on the original YOLOv8 PANet structure, the system introduces AFPN (Adaptive Feature Pyramid Network) to support the dynamic fusion of cross-scale feature information. AFPN features a learnable weighting mechanism that dynamically adjusts the contribution of features at different scales. The adaptive feature pyramid network extracts features from the target dataset and then fuses the extracted feature maps. During the fusion process, an adaptive weight adjustment mechanism is used to dynamically assign weights to different feature maps based on the target's scale and position. The fused feature maps are then subjected to further convolution operations and activation function processing to enhance the expressive power of the features. The adaptive feature pyramid network upsamples and downsamples the feature maps to obtain feature maps of varying resolutions, enabling it to capture objects of varying sizes within the target dataset.

[0033] S4, uses the environment modeling algorithm to perform three-dimensional modeling based on the training set and construct the motion path; The environment modeling algorithm consists of a point cloud filtering module and a surface reconstruction module. The training set is processed by the point cloud filtering module, and then a closed three-dimensional module is generated by the surface reconstruction module.

[0034] Step S41: voxelize the depth image and lidar point cloud data according to the spatial position relationship of the images in the training set.

[0035] The pixels in the depth image are mapped to the corresponding voxel units and fused with the lidar point cloud to obtain fused point cloud data containing color, depth and three-dimensional position information.

[0036] The fused point cloud data is voxelized, dividing the 3D space into small cubic voxels. The side length of the voxels is set. Statistical processing is performed on the point cloud within each voxel, calculating the average or median coordinate value of the points within each voxel. This value is used to represent all points within the voxel, thereby reducing the density of the point cloud and removing some noise points, thus completing the point cloud filtering of the training set.

[0037] Step S42: Perform three-dimensional modeling on the filtered point cloud data through a surface reconstruction module.

[0038] The filtered point cloud data is divided into multiple local areas. According to the density of the point cloud and the geometric characteristics of the target object, the area is divided into a fixed number of points (such as 100-500 points per area) or a fixed spatial range (such as a cube area with a side length of 0.1m-0.3m).

[0039] By calculating the coefficients of the polynomial surface, the fitting error between the surface and the point cloud data in the area is minimized, and the distance from each point to the fitting surface is calculated as the fitting error. If the error exceeds the set threshold, it is considered that the area is not suitable for fitting with the current order of polynomial, and the polynomial order is adjusted and fitted again.

[0040] The point cloud data is regarded as the gradient field of a scalar function. The gradient information of the scalar function is estimated in the neighborhood of the point cloud data. By solving the Poisson equation, the scalar function is solved in the entire definition domain. Then, the reconstructed surface is obtained by isosurface extraction. The point cloud data is then processed by Poisson reconstruction to generate a closed three-dimensional model.

[0041] S5, send the test set into the constructed system to complete the grasping of the robotic arm.

[0042] In the robotic arm's visual servo control system, the images in the test set are first analyzed using a YOLOv8-based object detection model. The target object's position in the image is located and bounding box information is output. The geometric features of the target object are extracted based on the coordinate parameters of the bounding box, and the approximate position of the target within the image plane is determined. The width and height of the bounding box are analyzed to calculate the aspect ratio of the target object, thereby estimating the shape and size of the target object. Simultaneously, the depth data obtained by the depth camera is combined to calculate the depth information of the target object, construct the target object's three-dimensional coordinates, and determine the target object's position in three-dimensional space. Furthermore, by analyzing the target object's contour and geometric features and combining the depth data to calculate the target object's posture information, the robotic arm provides accurate three-dimensional pose data for the grasping action.

[0043] During the grasping process, the robotic arm performs motion planning based on the calculated 3D position of the target, combined with its own kinematic model, to generate an appropriate grasping trajectory. As the robotic arm's end effector approaches the target object, it adjusts its posture based on the object's posture information to ensure accurate grasping. Furthermore, the system can adjust the grasping force in real time based on the target object's geometric features and position information. For smooth, easily slippery objects, the grasping force is appropriately increased. For irregularly shaped objects, the grasping angle and position are adjusted to improve grasping stability and success rate, ultimately enabling the target to be grasped.

[0044] An embodiment of the present invention provides a robotic arm vision control system based on a computing power chip. The system collects image data in an actual operating environment through multiple industrial cameras to obtain an initial data set, and preprocesses the training set by image grayscale and denoising to obtain a target data set. An improved YOLOv8 target detection model is constructed using the YOLOv8 target detection model as a benchmark model as a robotic arm vision detection network model. An environmental modeling algorithm is used to perform three-dimensional modeling on the training set to construct a motion path. The test set is sent into the constructed system to complete the grasping of the robotic arm. Through visual detection and three-dimensional modeling methods, image recognition is performed on the grasped object to accurately obtain the three-dimensional position information of the target. Then, precise motion planning is performed in combination with the kinematic model to optimize the grasping trajectory, thereby effectively improving the accuracy of the grasping position and angle.

[0045] The second embodiment of the present invention relates to a network side server, such as Figure 2 As shown, it includes at least one processor 302; and a memory 301 that is communicatively connected to the at least one processor 302; wherein the memory 301 stores instructions that can be executed by the at least one processor 302, and the instructions are executed by the at least one processor 302 to enable the at least one processor 302 to execute the above-mentioned data processing method. Memory 301 and processor 302 are connected using a bus. The bus can include any number of interconnected buses and bridges, connecting various circuits of one or more processors 302 and memory 301. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits. These are all well known in the art and are therefore not described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 302 is transmitted over a wireless medium via an antenna. Furthermore, the antenna receives data and transmits it to processor 302.

[0046] The processor 302 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory 301 can be used to store data used by the processor 302 when performing operations.

[0047] A third embodiment of the present invention relates to a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the robotic arm vision control system based on a computing chip in the first embodiment.

[0048] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of this patent. Making insignificant modifications to the algorithm or adding insignificant designs to the process, but without changing the core design of its algorithm and process, is within the protection scope of this patent.

[0049] The above are only embodiments of the present invention. Common knowledge such as the specific structures and characteristics known in the art are not described in detail here. Those of ordinary skill in the art know all the common general technical knowledge in the technical field to which the invention belongs before the filing date or the priority date, can know all the existing technologies in this field, and have the ability to apply the conventional experimental means before this date. Those of ordinary skill in the art can, under the inspiration given in this application, combine their own abilities to improve and implement this solution. Some typical well-known structures or well-known methods should not become obstacles for those of ordinary skill in the art to implement this application. It should be noted that for those skilled in the art, without departing from the structure of the present invention, several deformations and improvements can be made, which should also be regarded as within the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application should be based on the content of its claims, and the specific implementation manners described in the specification can be used to explain the content of the claims.

[0050] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A robotic arm vision control system based on a computing power chip, characterized in that, Including: S1. Collect image data in the actual operation environment through multiple industrial cameras to obtain an initial data set. S2. Preprocess the training set by means of image grayscale conversion and denoising to obtain a target data set. S3. Construct an improved YOLOv8 target detection model based on the YOLOv8 target detection model as the robotic arm vision detection network model. S4. Adopt an environment modeling algorithm to perform three-dimensional modeling based on the training set and construct a motion path. S5. Send the test set into the constructed system to complete the grasping of the robotic arm.

2. The robotic arm vision control system based on a computing power chip according to claim 1, wherein In S2, preprocessing the training set by means of image grayscale conversion and denoising to obtain a target data set includes: performing grayscale conversion on the collected images, converting color images into grayscale images, separating the images from the environment to extract target features, and then using a filter to reduce the amount of unwanted noise in the images. After preprocessing, the training set obtains a target data set.

3. The robotic arm vision control system based on a computing power chip according to claim 2, characterized in that, In digital image processing, first convert the images in the training set into grayscale images, and then use Gaussian filtering to perform weighted averaging on the entire image. By performing weighted averaging on its own and other pixel values in the neighborhood, a target data set is obtained.

4. A robotic arm vision control system based on a computing power chip according to claim 3, characterized in that, Constructing an improved YOLOv8 target detection model based on the YOLOv8 target detection model as the robotic arm vision detection network model includes: Step S31. Perform lightweight design on the YOLOv8 model, and quantize the model parameters from the traditional 32-bit floating-point representation to 16-bit fixed-point representation. Step S32. Introduce an adaptive feature pyramid network (AFPN) to fuse feature maps of different scales.

5. The robotic arm vision control system based on a computing power chip according to claim 4, characterized in that, Introducing an adaptive feature pyramid network (AFPN) to fuse feature maps of different scales includes: AFPN has a learnable weighting mechanism and dynamically adjusts the contribution ratio of feature maps of different scales through the learnable weight mechanism. Extract features from the target data set through the adaptive feature pyramid network, and then fuse the extracted feature maps. Dynamically allocate the weights of different feature maps according to the scale and position of the target, and perform further convolution operations and activation function processing on the fused feature maps to enhance the expression ability of the features. The adaptive feature pyramid network performs upsampling and downsampling operations on the feature maps, thereby obtaining feature maps of different resolutions, and then capturing targets of different sizes in the target data set.

6. The robotic arm vision control system based on a computing power chip according to claim 1, characterized in that, In S4, adopting an environment modeling algorithm to perform three-dimensional modeling based on the training set and construct a motion path includes: Step S41. Voxelize the depth image and lidar point cloud data according to the spatial position relationship of the pictures in the training set. Step S42. Perform three-dimensional modeling on the filtered point cloud data through a surface reconstruction module.

7. The robotic arm vision control system based on a computing power chip according to claim 6, characterized in that, Voxelization of depth images and lidar point cloud data for images in the training set according to their spatial position relationships includes: mapping pixel points in the depth image to corresponding voxel units, fusing them with the lidar point cloud to obtain fused point cloud data containing color, depth, and three-dimensional position information, voxelizing the fused point cloud data, dividing the three-dimensional space into small cubic voxels, calculating the average coordinate value of points within each voxel, and using this value to represent all points within the voxel, thereby reducing the density of the point cloud and removing some noise points, and further completing the point cloud filtering process for the training set.

8. A robotic arm vision control system based on a computing power chip according to claim 7, characterized in that, Three-dimensional modeling of the filtered point cloud data is performed through a surface reconstruction module, including: regarding the point cloud data as the gradient field of a certain scalar function, estimating the gradient information of the scalar function within the neighborhood of the point cloud data, solving the Poisson equation to solve the scalar function within the entire domain, then obtaining the reconstructed surface through isosurface extraction, and further processing the point cloud data through Poisson reconstruction to generate a closed three-dimensional model.

9. A network-side server, characterized in that, Including: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the robotic arm vision control system based on a computing power chip as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the robotic arm vision control system based on a computing power chip as described in any one of claims 1 to 8.