Method and device for controlling mechanical arm

By using the object detection model of the feature pyramid network and calibration of the hand-eye relationship in the visual robot arm, the problem of insufficient accuracy when dealing with small-sized targets is solved, and high-precision target recognition, positioning and grasping is achieved, which improves the automation level and reduces labor costs.

CN119952723APending Publication Date: 2025-05-09GUANGZHOU ON BRIGHT ELECTRONICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510353556.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Traditional visual robotic arm devices are difficult to effectively identify targets or accurately grasp due to insufficient accuracy when dealing with small-sized targets, resulting in inefficiency and requiring human intervention.

Method used

The target detection model of the structure-like pyramid network is adopted, and the target image information obtained by the image sensor is input to the model to obtain the target detection result. The target image coordinates are converted into the target three-dimensional coordinates in the robotic arm coordinate system by calibrating the hand-eye relationship, thereby controlling the robotic arm to perform accurate actions.

Benefits of technology

The identification and positioning accuracy of small-size targets is improved, the accurate positioning and grabbing of small-size targets is achieved, the level of production automation is improved, and labor costs are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119952723A_ABST
    Figure CN119952723A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for controlling a mechanical arm. The method comprises the following steps: inputting target image information about a target object acquired by an image sensor into a target detection model of which the structure is a feature pyramid network, so as to obtain a target detection result corresponding to the target image information; based on the target image coordinates and a calibrated hand-eye relation between a camera coordinate system and a mechanical arm coordinate system, target three-dimensional coordinates of the target object in the mechanical arm coordinate system are determined, and the mechanical arm coordinate system is a coordinate system describing the position relation between a mechanical arm and surrounding objects; and controlling the mechanical arm to act based on the target three-dimensional coordinate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot arm control, and in particular to a method and device for controlling a robot arm. Background Art

[0002] The visual robot arm device can grasp the target by visually identifying the target and driving the robot arm to move. Traditional visual robot arm devices perform better on larger targets. Once the target size is small, it is easy to fail to effectively identify the target or fail to accurately grasp the target due to its own insufficient precision, resulting in low work efficiency and the need for human intervention. Summary of the invention

[0003] According to an embodiment of the present invention, a method for controlling a robotic arm includes: inputting target image information about a target object acquired by an image sensor into a target detection model having a structure of a feature pyramid network to obtain a target detection result corresponding to the target image information; determining the target image coordinates of the target object in a camera coordinate system based on the target detection result, the camera coordinate system being a coordinate system used by the image sensor to describe the positional relationship between objects in the target image information; determining the target three-dimensional coordinates of the target object in the robotic arm coordinate system based on the target image coordinates and a calibrated hand-eye relationship between the camera coordinate system and the robotic arm coordinate system, the robotic arm coordinate system being a coordinate system that describes the positional relationship between the robotic arm and objects around it; and controlling the robotic arm to move based on the target three-dimensional coordinates.

[0004] According to an embodiment of the present invention, a device for controlling a robotic arm includes: a processor; and a memory on which computer executable instructions are stored, wherein the computer executable instructions, when executed by the processor, prompt the processor to execute the above-mentioned method for controlling the robotic arm.

[0005] According to the computer-readable storage medium of an embodiment of the present invention, computer-executable instructions are stored thereon, wherein these computer-executable instructions, when executed by a processor, prompt the processor to execute the above-mentioned method for controlling a robotic arm.

[0006] A computer program product according to an embodiment of the present invention comprises computer executable instructions, wherein when these computer executable instructions are executed by a processor, they prompt the processor to execute the above-mentioned method for controlling a robotic arm. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The present invention can be better understood from the following description of specific embodiments of the present invention in conjunction with the accompanying drawings, in which:

[0008] Figure 1 A schematic flow chart of a method for controlling a robot arm according to an embodiment of the present invention is shown.

[0009] Figure 2 A schematic structural diagram of a system for controlling a robot arm according to an embodiment of the present invention is shown.

[0010] Figure 3 A schematic structural diagram of a target detection model according to an embodiment of the present invention is shown.

[0011] Figure 4 A schematic flow chart of a method for determining a calibrated hand-eye relationship according to an embodiment of the present invention is shown.

[0012] Figure 5 A schematic diagram of a computer system that can implement the method and apparatus for controlling a robot arm according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0013] The features and exemplary embodiments of various aspects of the present invention will be described in detail below. In the detailed description below, many specific details are proposed to provide a comprehensive understanding of the present invention. However, it is obvious to those skilled in the art that the present invention can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present invention by illustrating examples of the present invention. The present invention is by no means limited to any specific configuration and algorithm proposed below, but covers any modification, replacement and improvement of elements, parts and algorithms without departing from the spirit of the present invention. In the accompanying drawings and the following description, known structures and technologies are not shown to avoid unnecessary ambiguity to the present invention.

[0014] Taking into account the problem of insufficient accuracy of traditional visual robotic arm devices, a method and device for controlling a robotic arm according to an embodiment of the present invention are proposed, wherein target image information is input into a network detection model to obtain a target detection result and determine the target image coordinates, and the target three-dimensional coordinates are determined based on the calibrated hand-eye relationship and the robotic arm movement is controlled. Since the network detection model with a feature pyramid network structure has high recognition accuracy, it can more accurately identify small-sized target objects, thereby achieving accurate positioning and grasping of small-sized target objects, improving the level of production automation, and reducing labor costs.

[0015] Figure 1 A schematic flow chart of a method for controlling a robot arm according to an embodiment of the present invention is shown. Figure 2 A schematic structural diagram of a system for controlling a robotic arm according to an embodiment of the present invention is shown. Figure 1 and Figure 2 , a method for controlling a robotic arm according to an embodiment of the present invention is described in detail.

[0016] like Figure 1 As shown, the method for controlling a robot arm according to an embodiment of the present invention includes the following steps S101-S104.

[0017] S101: inputting target image information about the target object acquired by the image sensor into a target detection model structured as a feature pyramid network to obtain a target detection result corresponding to the target image information;

[0018] S102: determining the target image coordinates of the target object in the camera coordinate system according to the target detection result, where the camera coordinate system is a coordinate system used by the image sensor to describe the positional relationship between various objects in the target image information;

[0019] S103: determining target three-dimensional coordinates of the target object in the robotic arm coordinate system based on the target image coordinates and the calibrated hand-eye relationship between the camera coordinate system and the robotic arm coordinate system, where the robotic arm coordinate system is a coordinate system that describes the positional relationship between the robotic arm and its surrounding objects;

[0020] S104: Control the robot arm to move based on the target three-dimensional coordinates.

[0021] like Figure 2 The system shown includes a perception layer, a decision-making and planning layer, a control layer, and an execution layer. The perception layer includes image sensors and other sensors. The image sensors are used to obtain image information of the surrounding environment. Other sensors include pressure sensors, tactile sensors, etc., which are mainly installed at the end of the robotic arm to perceive the positional relationship and state between the robotic arm and the object. The decision-making and planning layer includes an image data processor, which is used to receive target image information from the image sensor to generate target three-dimensional coordinates. The decision-making and planning layer also includes a feedback adjustment processor, which is used to make motion decisions for the robotic arm based on data feedback from other sensors. The control layer includes a controller, which is used to generate control signals for the robotic arm. The execution layer includes a joint drive motor and an actuator located at the end of the robotic arm, wherein the joint drive motor is located at each joint of the robotic arm, and acts based on the control signal sent by the controller to adjust the posture angle of the corresponding joint, so that the end of the robotic arm moves to the target three-dimensional coordinates; wherein the actuator is located at the end of the robotic arm and is used to perform a grasping action.

[0022] According to the method of the embodiment of the present invention, the target image information is input into the target detection model to obtain the target detection result. The target image coordinates of the target object can be determined based on all the target detection results of the target image information. Then, the target three-dimensional coordinates corresponding to the target image coordinates can be determined by calibrating the hand-eye relationship, and finally the robot arm is controlled to move based on the target three-dimensional coordinates. The target three-dimensional coordinates correspond to the position of the target object in the actual three-dimensional space. When the robot arm moves based on the target three-dimensional coordinates and its end moves to the target three-dimensional coordinates, the actuator at the end can directly grasp the target object. In order to ensure that the actuator at the end can accurately grasp the target object, the method of this embodiment improves the accuracy of identifying the target object in the target image information through the target detection model of the feature pyramid structure, and then determines the accurate target image coordinates, converts the target image coordinates into accurate target three-dimensional coordinates through reliable calibration of the hand-eye relationship, and controls the robot arm to accurately move based on the target three-dimensional coordinates. Therefore, the method of this embodiment can more accurately realize the positioning and grasping of small-sized target objects.

[0023] In some embodiments, the target detection model is configured to: generate multiple convolutional feature data of different sizes through multiple input convolutional layers that are connected layer by layer from bottom to top through target image information; perform an upsampling operation on the convolutional feature data of the highest layer to obtain the upsampled feature data of the highest layer; starting from the second highest layer that is one layer lower than the highest layer, iteratively perform the following operations downward until the fused feature data of a preset layer is obtained: perform a feature fusion operation on the convolutional feature data of the current layer and the upsampled feature data that is one layer higher than the current layer to obtain the fused feature data of the current layer; perform an upsampling operation on the fused feature data of the current layer to obtain the upsampled feature data of the current layer; and determine the target detection result based on the fused feature data of the preset layer.

[0024] Figure 3 A schematic structural diagram of a target detection model according to an embodiment of the present invention is shown, wherein a plurality of input convolutional layers are connected layer by layer from bottom to top, and each input convolutional layer performs a corresponding convolution operation on the input target image information or convolutional feature data to output convolutional feature data of different sizes; wherein a preset number of upsampling fusion layers, not less than two, are connected layer by layer from top to bottom, and each upsampling fusion layer is horizontally connected to the corresponding plurality of input convolutional layers, and each upsampling fusion layer is used to perform an upsampling operation on the convolutional feature data of a higher layer to obtain the upsampling feature data of a higher layer, and perform a feature fusion operation on the upsampling feature data of a higher layer and the convolutional feature data of a current layer to obtain the fusion feature data of the current layer. The lowest layer of the preset number of upsampling fusion layers (i.e., the upsampling fusion layer of the preset layer) is horizontally connected to an output layer, and the output layer determines the target detection result of the layer based on the feature fusion data of the preset layer, and is used to determine the target image coordinates of the target object.

[0025] like Figure 3 As shown, assuming that there are five input convolutional layers and three upsampling fusion layers, the order of the layers in the entire target detection model is based on the order of the input convolutional layers from bottom to top. Therefore, the five input convolutional layers are recorded from the first layer to the fifth layer from bottom to top, and the three upsampling fusion layers are recorded from the fourth layer to the second layer from top to bottom, which are horizontally connected to the three input convolutional layers from the fourth layer to the second layer, respectively. The bottom layer (preset layer) of the upsampling fusion layer is the second layer. The convolution feature data output according to the bottom-up path are P1-P5, and P5 is the convolution feature data of the highest layer (fifth layer). The upsampling operation can be performed on P5 to obtain the upsampling feature data C5 of the fifth layer. A feature fusion operation is performed on P4 and C5 to obtain the fused feature data C'4 of the fourth layer, and an upsampling operation is performed on C'4 to obtain the upsampled feature data C4 of the fourth layer; a feature fusion operation is performed on P3 and C4 to obtain the fused feature data C'3 of the third layer, and an upsampling operation is performed on the fused feature data C'3 to obtain the upsampled feature data C3 of the third layer; a feature fusion operation is performed on P2 and C3 to obtain the fused feature data C'2 of the second layer. Since the second layer is the preset layer, the iteration is stopped after the fused feature data C'2 of the second layer is obtained, and the target detection result is determined using the fused feature data C'2 of the second layer.

[0026] like Figure 3 As shown, multiple input convolutional layers connected layer by layer from bottom to top extract features from the target image information layer by layer to obtain multiple convolutional feature data of different sizes. The convolutional feature data corresponding to each layer contains feature information of different scales and different levels, indicating the abstract representation of the target image information at different scales. In the process of layer-by-layer extraction, the resolution of the convolutional feature data gradually decreases, while the semantic information gradually becomes richer. In some embodiments, the multiple input convolutional layers include at least one of a residual network, a visual geometry group (VGG) network, and a mobile network (Mobilenet).

[0027] The purpose of the upsampling operation is to restore the resolution so that the resolution of the convolutional feature data of the highest layer or the fused feature data of other layers is the same as the resolution of the convolutional feature data of the lower layer, for example Figure 3 The resolution of C5 is the same as that of P4, and the resolution of C4 is the same as that of P3, so as to perform subsequent feature fusion operations. In some embodiments, the upsampling operation includes a bilinear interpolation operation or a deconvolution operation.

[0028] The feature fusion operation can combine the high-resolution information of the convolution feature data with the rich semantic information of the upsampled feature data of a higher layer, which not only helps to transmit the detailed information of the convolution feature data, but also enhances the positioning capability of the upsampled feature data of a higher layer. In some embodiments, the feature fusion operation includes an element-by-element addition operation or a feature concatenation operation.

[0029] Furthermore, in order to ensure that the number of channels of the convolution feature data and the upsampled feature data for feature fusion operation is the same, a 1×1 convolution operation can also be performed on the convolution feature data of the lower layer, such as Figure 3 As shown in , P2-P5 are recorded as P'2-P'5 after the 1×1 convolution operation. Therefore, in some embodiments, the target detection model can also be configured to: before performing the feature fusion operation, perform a 1×1 convolution operation on the corresponding convolution feature data respectively, so that the number of channels of the convolution feature data of the current layer to perform the feature fusion operation is the same as the number of channels of the upsampled feature data of the layer higher than the current layer.

[0030] The target detection result can be determined based only on the fused feature data of the preset layer, because the preset layer is the bottom layer of the upsampling fusion layer. Compared with the fused feature data of the higher layer, the fused feature data of the bottom layer contains the richest feature information. Therefore, the fused feature data of the preset layer can be used to perform small target detection to determine the target detection result. In some embodiments, such as Figure 3 As shown in the figure, each layer can be set with a horizontally connected output layer to generate the target detection result of the layer based on the convolution feature data of the highest layer or the fusion feature data of other layers. The target detection result of each layer includes the bounding box and category label corresponding to the predicted anchor point corresponding to the layer. The output layer can be implemented using Fast RCNN or other networks related to target detection algorithms.

[0031] Furthermore, the feature convolution data can eliminate the aliasing effect in upsampling through a 3×3 convolution operation, so the target detection model is also configured to: perform a 3×3 convolution operation on the fused feature data of the preset layer before determining the target detection result.

[0032] To adapt to the real-time requirements of industrial scenarios, in some embodiments, the training of the target detection model can introduce transfer learning and data enhancement strategies to optimize the reasoning speed and detection accuracy, ensuring that the trained target detection model can achieve efficient small target detection with limited computing resources.

[0033] In the method of this embodiment, the hand-eye relationship calibration can be expressed as a transformation matrix, and the two-dimensional target image coordinates in the target image information can be converted into three-dimensional target three-dimensional coordinates in the working space of the robot arm using the transformation matrix. The hand-eye relationship calibration can be pre-calibrated using a calibration object (such as a calibration plate or a target object with known geometric features), that is, establishing a transformation relationship between the camera coordinate system and the robot arm coordinate system, specifically adjusting the parameter value of the transformation matrix of the initial hand-eye relationship according to feedback.

[0034] Figure 4 FIG. 2 is a schematic flow chart of a method for calibrating hand-eye relationship according to an embodiment of the present invention. Figure 4 As shown, in some embodiments, the calibration hand-eye relationship is determined by the following operations: inputting the calibration object image information about the calibration object acquired by the image sensor into the target detection model to obtain the calibration object detection result corresponding to the calibration object image information; determining the calibration object image coordinates of the calibration object in the camera coordinate system according to the calibration object detection result; determining the calibration object three-dimensional coordinates of the calibration object in the robotic arm coordinate system based on the calibration object image coordinates and the initial hand-eye relationship between the camera coordinate system and the robotic arm coordinate system; controlling the movement of the robotic arm to move the end of the robotic arm to the calibration object three-dimensional coordinates; judging whether the actual distance between the end and the calibration object is less than a threshold value based on the calibration object image information acquired again by the image sensor; adjusting the initial hand-eye relationship based on the judgment, and iteratively executing the operations of determining the calibration object three-dimensional coordinates in the robotic arm coordinate system and controlling the movement of the robotic arm until the actual distance is less than the threshold value; and determining the current initial hand-eye relationship as the calibration hand-eye relationship.

[0035] exist Figure 4 In the calibration object, the three-dimensional coordinates of the calibration object refer to the coordinates of the calibration object image converted according to the current initial hand-eye relationship. The more accurate the hand-eye relationship is, the closer the converted three-dimensional coordinates of the calibration object are to the actual coordinates of the calibration object, and the smaller the actual distance between the end of the robotic arm and the calibration object when the three-dimensional coordinates of the calibration object are used to control the robotic arm. The calibration object image information acquired again by the image sensor should include the image information of the calibration object and the end, so as to determine the actual distance between the end and the calibration object. When the actual distance between the end and the calibration object is less than the threshold, the actuator at the end can grasp the calibration object located at the three-dimensional coordinates of the calibration object; when the actual distance between the end and the calibration object is not less than the threshold, the actuator cannot grasp the calibration object, which means that the initial hand-eye relationship of the conversion coordinates is not accurate enough at this time, and the transformation matrix needs to be further corrected. The transformation matrix can be solved in combination with the least squares method.

[0036] In some embodiments, the type and quantity of the calibration objects can be selected and adjusted according to the requirements. In some embodiments, the initial hand-eye relationship can be preset or calculated based on the calibration object detection results.

[0037] In some embodiments, controlling the robot arm to move based on the target three-dimensional coordinates includes: controlling the joint drive motor of the robot arm based on kinematic calculation and feedback adjustment to move the end of the robot arm to the target three-dimensional coordinates; adjusting the posture of the robot arm through closed-loop control to enable the actuator at the end to perform a grasping action. Among them, the kinematic calculation specifically includes inverse kinematic calculation, and the inverse kinematic calculation is combined with motion planning to control the joint drive motor. The feedback adjustment specifically includes PID adjustment. The closed-loop control is performed based on various state signals of the robot arm (such as signals from other sensors), which further ensures the accuracy and stability of the grasping action.

[0038] Figure 5 A schematic diagram of a computer system that can implement the method and apparatus for controlling a robot arm according to an embodiment of the present invention is shown. Figure 5 The computer system 500 shown is only an example and should not bring any limitation to the functions and scope of use of the method and apparatus for controlling a robot arm according to the embodiments of the present invention.

[0039] like Figure 5 As shown, the computer system 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 to a random access memory (RAM) 503. Various programs and data required for the operation of the computer system 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0040] Typically, the following devices may be connected to the I / O interface 505: input devices 506 including, for example, a touch screen, a touch pad, a camera, an accelerometer, a gyroscope, a sensor, etc.; output devices 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, a motor, an electronic speed controller, etc.; storage devices 508 including, for example, a flash memory (FlashCard), etc.; and communication devices 509. The communication devices 509 may allow the computer system 500 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 5 The computer system 500 is shown with various devices, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have instead. Figure 5 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0041] In particular, according to some embodiments of the present invention, the process described above with reference to the flowchart may be implemented as a computer program. For example, a computer readable medium is provided on which a computer program is stored, the computer program comprising a method for executing Figure 1 The computer program is a program code of the method for controlling a robot arm shown in the figure. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functional units defined in the apparatus for controlling a robot arm according to the embodiment of the present invention are implemented.

[0042] It should be noted that the computer-readable medium according to the embodiment of the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium according to the embodiment of the present invention may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device, or device. In addition, the computer-readable signal medium according to the embodiment of the present invention may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0043] Computer program code for performing operations according to embodiments of the present invention may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet).

[0044] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0045] The present invention can be implemented in other specific forms without departing from its spirit and essential features. For example, the algorithms described in the specific embodiments can be modified, and the system architecture does not depart from the basic spirit of the present invention. Therefore, the current embodiments are regarded as exemplary and non-restrictive in all aspects, and the scope of the present invention is defined by the appended claims rather than the above description, and all changes falling within the meaning and equivalent scope of the claims are thus included in the scope of the present invention.

Claims

1. A method for controlling a robotic arm, comprising: Inputting target image information about the target object acquired by the image sensor into a target detection model having a feature pyramid network structure to obtain a target detection result corresponding to the target image information; Determine, according to the target detection result, the target image coordinates of the target object in a camera coordinate system, where the camera coordinate system is a coordinate system used by the image sensor to describe the positional relationship between various objects in the target image information; Determining target three-dimensional coordinates of the target object in the robotic arm coordinate system based on the target image coordinates and a calibrated hand-eye relationship between the camera coordinate system and the robotic arm coordinate system, wherein the robotic arm coordinate system is a coordinate system that describes a positional relationship between the robotic arm and objects around it; as well as The robot arm is controlled to move based on the target three-dimensional coordinates.

2. The method according to claim 1, wherein: The object detection model is configured as: The target image information is passed through a plurality of input convolutional layers connected layer by layer from bottom to top to generate a plurality of convolutional feature data of different sizes; Performing an upsampling operation on the convolution feature data of the highest layer to obtain upsampled feature data of the highest layer; Starting from the convolution feature data of the next highest layer one layer lower than the highest layer, the following operations are iteratively performed downward until the fusion feature data of the preset layer is obtained: Performing a feature fusion operation on the convolution feature data of the current layer and the up-sampled feature data of a layer higher than the current layer to obtain fused feature data of the current layer; as well as Performing an upsampling operation on the fused feature data of the current layer to obtain the upsampled feature data of the current layer; Based on the fused feature data of the preset layer, the target detection result is determined.

3. The method according to claim 2, wherein: The multiple input convolutional layers include at least one of a residual network, a visual geometry group (VGG) network, and a mobile network (Mobilenet).

4. The method according to claim 2, wherein: The up-sampling operation includes a bilinear interpolation operation or a deconvolution operation.

5. The method according to claim 2, wherein: The feature fusion operation includes an element-by-element addition operation or a feature concatenation operation.

6. The method according to claim 2, wherein: The target detection model is also configured as: Before performing the feature fusion operation, a 1×1 convolution operation is performed on the corresponding convolution feature data respectively so that the number of channels of the convolution feature data of the current layer on which the feature fusion operation is to be performed is the same as the number of channels of the upsampled feature data of a layer higher than the current layer.

7. The method according to claim 2, wherein: The target detection model is also configured as: Before determining the target detection result, a 3×3 convolution operation is performed on the fused feature data of the preset layer.

8. The method according to claim 1, wherein: The calibrated hand-eye relationship is determined by the following operations: Inputting the calibration object image information about the calibration object acquired by the image sensor into the target detection model to obtain the calibration object detection result corresponding to the calibration object image information; Determining the calibration object image coordinates of the calibration object in the camera coordinate system according to the calibration object detection result; Determine the calibration object three-dimensional coordinates of the calibration object in the robotic arm coordinate system based on the calibration object image coordinates and the initial hand-eye relationship between the camera coordinate system and the robotic arm coordinate system; Controlling the movement of the robotic arm so that the end of the robotic arm moves to the three-dimensional coordinates of the calibration object; Based on the image information of the calibration object acquired again by the image sensor, determining whether the actual distance between the terminal and the calibration object is less than a threshold; Adjusting the initial hand-eye relationship based on the judgment, and iteratively performing operations of determining the three-dimensional coordinates of the calibration object in the robotic arm coordinate system and controlling the movement of the robotic arm until the actual distance is less than the threshold; as well as The current initial hand-eye relationship is determined as the calibrated hand-eye relationship.

9. The method according to any one of claims 1 to 8, wherein: Controlling the robot arm to move based on the target three-dimensional coordinates includes: Based on kinematic calculation and feedback regulation, controlling the joint drive motor of the robotic arm so that the end of the robotic arm moves to the target three-dimensional coordinates; The posture of the robot arm is adjusted through closed-loop control so that the actuator at the end performs a grasping action.

10. A device for controlling a robotic arm, comprising: processor; as well as A memory having computer executable instructions stored thereon, wherein the computer executable instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 9.

11. A computer-readable storage medium having computer-executable instructions stored thereon, wherein: When the computer executable instructions are executed by a processor, the processor is caused to perform the method of any one of claims 1 to 9.

12. A computer program product comprising computer executable instructions, wherein: When the computer executable instructions are executed by a processor, the processor is caused to perform the method of any one of claims 1 to 9.