Control method and device of multi-mode intelligent manipulator
Through the control method of multimodal intelligent robot, binocular vision and deep learning technology, the robot's autonomous recognition and path optimization in complex environments is realized, which improves flexibility and intelligence, and is suitable for multiple application scenarios.
Patent Information
- Application Number
- CN202510847443.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-24
AI Technical Summary
Existing robots rely mostly on a single control mode, which is difficult to meet the dynamic needs in complex environments. The flexibility and intelligence level are insufficient, and the human-computer interaction capabilities need to be improved.
The control method of multimodal intelligent robot is adopted, combined with binocular vision technology, YOLO model, GPT-4o model and three-dimensional haptic sensor, to achieve the fusion of multiple modal information. Through deep learning and intelligent control algorithms, the robot can independently identify the target and optimize the execution path.
It realizes more efficient, more accurate and natural intelligent operation, and is suitable for industrial manufacturing, medical assistance, agricultural automation and home services and provides personalized operation methods.
Smart Images

Figure CN120347781A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of manipulators and artificial intelligence, and in particular, to a control method and device for a multi-modal intelligent manipulator. Background Art
[0002] Currently, in the fields of intelligent manufacturing, intelligent healthcare, service manipulators, agricultural automation, etc., the requirements for the intelligence and flexibility of manipulators are increasing day by day. Traditional manipulators mostly rely on single control modes, such as preset trajectory control or simple visual feedback control, and it is difficult to meet the dynamic requirements in complex environments. The multi-modal intelligent manipulator can more accurately understand the external environment, adapt to different tasks, and interact more naturally with humans by integrating multiple sensing methods.
[0003] However, most of the current manipulators on the market still mainly use single-modal control, and fail to fully utilize multi-modal fusion technology for information fusion, resulting in relatively large room for improvement in their flexibility, intelligence level, and human-computer interaction ability; the current research trend shows that manipulators integrating multiple sensing and control technologies have become an important development direction in the field of intelligent manipulators, but in the face of complex environments, their good adaptability and human-computer collaboration efficiency still need to be improved. Summary of the Invention
[0004] The purpose of the present invention is to provide a control method and device for a multi-modal intelligent manipulator. Through deep learning and intelligent control algorithms, the manipulator can autonomously identify targets, optimize the execution path, and adapt to environmental changes, can integrate multiple modal information, achieve more efficient, accurate, and natural intelligent operations, be applicable to multiple application scenarios such as industrial manufacturing, medical assistance, agricultural automation, and home services, and provide a more personalized and intelligent operation method according to user needs.
[0005] To achieve the above purpose, the present invention provides a control method for a multi-modal intelligent manipulator, including the following steps: S1. Capture the left and right images of a target object in the same scene through two cameras based on binocular vision technology; S2. Identify and locate the target object through the YOLO model, and output the category and bounding box coordinates of the target object; S3. Extract the center point coordinates of the target object in the left and right images based on S1 and S2; S4. Match the objects in the left and right images that have the same category as the target object and y similar horizontal axis coordinates; S5. Calculate the disparity between the left and right images and the depth information of the object to estimate the three-dimensional spatial coordinates of the target object; S6. Record the historical coordinates of S5, smooth the historical coordinates of each target object to calculate the actual three-dimensional space coordinates, combine them with the detection labels and bounding box coordinates, and output the object detection results of the YOLO model; S7. Input the object detection results of S6 into the task decision model to convert the task information into structured instructions; S8. The manipulator executes the grasping action according to the structured instructions of S7 and calculates and adjusts the grasping force.
[0006] Preferably, in S4, by comparing the categories and y-axis horizontal coordinates of the objects in the left and right images, it is determined whether they are the same object. If they are the same object, the difference in the y-axis horizontal coordinates of the objects in the left and right images is calculated, and the object with the smallest difference is selected for matching; if no matching object can be found, it is regarded as an unmatched object. At this time, through the two-dimensional coordinates of the object, the camera position needs to be adjusted and start from S1 again.
[0007] Preferably, the calculation process of the parallax and the depth information of the object between the left and right images in S5 is as follows: Parallax Calculation formula: ; (1) Where, and are the horizontal pixel coordinates of the target object in the left and right images respectively; Depth Z Calculation formula: ; (2) Where, B is the baseline of the camera, in millimeters, is the focal length of the camera, in millimeters; 100 is the unit conversion coefficient.
[0008] Preferably, the formula for calculating the actual three-dimensional space coordinates in S6 is as follows: ; (3) ; (4) Where, and are the center point coordinates detected of the target object in the image, is the pixel size, in millimeters / pixel, ( ) is the converted actual three-dimensional space coordinates.
[0009] Preferably, the specific content of the task decision model in S7 is as follows: S71. The GPT-4o model uses the Transformer structure to encode the input natural language task instructions, extract context semantic information, and combine pre-trained knowledge for task understanding. The VisionEncoder in the model extracts features from the input image, generates high-dimensional visual feature vectors, and annotates key targets, such as object detection and semantic segmentation. The text and visual features are aligned through the Shared Latent Space to ensure the correlation of information between the two modalities. S72. Based on the knowledge base pre-trained by the GPT-4o model, combined with Prompt-based Learning to parse the user's instructions, identify the core task objectives and constraints, and use dependency analysis and entity recognition NER methods to decompose the executable task information into structured instructions for the manipulator to execute.
[0010] Preferably, the specific process of the manipulator calculating and adjusting the grasping force in S8 is as follows: S81. Initially determine the opening degree of the gripper according to the geometric information of the target object obtained by the binocular vision technology in S1, and then dynamically estimate the mass of the target object in combination with the information feedback by the three-dimensional tactile sensor set on the manipulator. The specific process is as follows: In the static state, after the gripper lifts the target object, the three-dimensional tactile sensor can sense: ; (5) Among them, is the tangential force of the gripper, is the mass of the object, is the gravitational acceleration; Estimated mass is: ; (6) S82. The process and formula for calculating the grasping force are as follows: ; (7) Among them, is the normal force applied on each side of the gripper, is the friction coefficient between the gripper and the target object; S83. Determine whether the gripper slips off according to whether the tangential force exceeds the friction limit. The specific judgment formula is as follows: ; (8) Among them, is the current normal force, is the magnitude of the current tangential force, is the tangential force of the gripper in the axis direction, is the gripper in Tangential force in the axial direction; S84. If the determination result in S83 is that the hand slips off, adjust the grasping force according to the estimated mass in S71. The adjustment formula is as follows: ; (9) Wherein, is the safety factor, and the value range is .
[0011] The present invention provides a control device for a multi-modal intelligent manipulator, including a base. Above the base, a large arm fixing seat, a large arm, an intermediate arm, a small arm, and an end effector are sequentially connected. A fixture interface is provided on the end effector, and a power interface, a bus interface, a safety switch interface, and a digital output interface are respectively provided on the side of the base.
[0012] Preferably, the maximum horizontal rotation radius of the manipulator device is set to 661.5 mm, the vertical movement range is set to 1122 mm, and the interference radius of the base is set to 120 mm.
[0013] Therefore, the present invention adopts the above-mentioned control method and device for a multi-modal intelligent manipulator. Compared with the prior art, it has the following beneficial effects: 1. In the control method of the multi-modal intelligent manipulator in this application, multiple modal fusion technologies are used for information fusion, greatly improving its flexibility, intelligence level, and human-computer interaction ability. It can fuse multiple modal information to achieve more efficient, accurate, and natural intelligent operations, and is applicable to multiple application scenarios such as industrial manufacturing, medical assistance, agricultural automation, and home services; 2. The end effector of the manipulator device in this application can be installed with multiple end tools (such as grippers, suction cups, etc.) to achieve specific operation functions. The overall mechanical structure is designed compactly, the joints are flexible, suitable for desktop work scenarios or loading onto mobile platforms (such as AGVs or humanoid manipulators) for use. At the same time, it has high stability and easy operation. The modular design of this manipulator is convenient for maintenance and expansion, providing users with highly customizable application possibilities.
[0014] The following further describes the technical solutions of the present invention in detail through the drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a flowchart of a control method for a multi-modal intelligent manipulator of the present invention; Figure 2 is an overall structure diagram of a control device for a multi-modal intelligent manipulator of the present invention; Reference Numerals 1. Base; 2. Upper arm fixing seat; 3. Upper arm; 4. Middle arm; 5. Lower arm; 6. End effector; 7. Fixture interface; 8. Power interface; 9. Bus interface; 10. Safety switch interface; 11. Switch output interface. DETAILED DESCRIPTION
[0016] In the description of the present invention, it should be noted that the terms "upper", "lower", "inside", "outside", etc. indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings, or are directions or positional relationships in which the product of the invention is usually placed when in use. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as a limitation on the present invention.
[0017] Example
[0018] like Figure 1 As shown, a control method of a multi-modal intelligent manipulator of the present invention comprises the following steps: S1, based on binocular vision technology, two cameras are used to capture the left and right images of the target object in the same scene; S2, identify and locate the target object through the YOLO model, and output the category and bounding box coordinates of the target object; S3, extracting the center point coordinates of the target object in the left and right images based on S1 and S2; S4, matching objects in the left and right images that have the same category and similar y-axis horizontal coordinates as the target object through S2 and S3; by comparing the categories and y-axis horizontal coordinates of the objects in the left and right images, determine whether they are the same object. If they are the same object, calculate the difference in the y-axis horizontal coordinates of the objects in the left and right images, and select the object with the smallest difference for matching; if a matching object cannot be found, it is regarded as an unmatched object. At this time, the camera position needs to be adjusted and restarted from S1 through the two-dimensional coordinates of the object; S5. The three-dimensional space coordinates of the target object are estimated by using the disparity between the left and right images and the depth information of the object. The calculation process of the disparity between the left and right images and the depth information of the object is as follows: Parallax Calculation formula: ; (1) in, and These are the horizontal pixel coordinates of the target objects in the left and right images respectively; depth Z Calculation formula: ; (2) in,B is the baseline of the camera, in millimeters, is the focal length of the camera, in millimeters; 100 is the unit conversion coefficient; S6. Record the historical coordinates of S5, smooth the historical coordinates of each target object to calculate the actual three-dimensional space coordinates, and combine them with the detection label and bounding box coordinates to output the target detection results of the YOLO model; the formula for calculating the actual three-dimensional space coordinates is as follows: ; (3) ; (4) where and are the center point coordinates of the target object detected in the image, is the pixel size, in millimeters per pixel, ( ) is the converted actual three-dimensional space coordinate; S7. Input the target detection results of S6 into the task decision model to convert the task information into structured instructions; the specific content of the task decision model is as follows: S71. The GPT-4o model uses the Transformer structure to encode the input natural language task instructions, extract context semantic information, and combine pre-trained knowledge for task understanding. The vision encoder VisionEncoder in the model extracts features from the input image to generate high-dimensional visual feature vectors, and annotates key targets, such as object detection and semantic segmentation. Through the joint embedding space Shared Latent Space, the text and visual features are aligned to ensure the correlation of information between the two modalities; S72. Based on the knowledge base pre-trained by the GPT-4o model, combined with Prompt-based Learning to parse the user's instructions, identify the core task objectives and constraints, and use dependency analysis and entity recognition NER methods to decompose the executable task information into structured instructions for the manipulator to execute; S8. The manipulator executes the grasping action according to the structured instructions of S6 and calculates and adjusts the grasping force; The specific process of the manipulator calculating the grasping force is as follows: S81. Initially determine the jaw opening degree according to the geometric information of the target object obtained by the binocular vision technology in S1, and then dynamically estimate the mass of the target object in combination with the information feedback by the three-dimensional tactile sensor set on the manipulator. The specific process is as follows: In the static state, after the jaw lifts the target object, the three-dimensional tactile sensor can sense: ; (5) where is the tangential force of the jaw, is the mass of the object, is the gravitational acceleration; Estimated mass is: ; (6) S82. The process and formula for calculating the grasping force are as follows: ; (7) Among them, is the normal force applied on each side of the jaw, is the friction coefficient between the jaw and the target object; S83. Determine whether the jaw slips off by checking whether the tangential force exceeds the friction limit. The specific judgment formula is as follows: ; (8) Among them, is the current normal force, is the magnitude of the current tangential force, is the tangential force of the jaw in the axis direction, is the tangential force of the jaw in the axis direction; S84. If the determination result of S83 is that the jaw slips off, adjust the grasping force according to the estimated mass in S71. The adjustment formula is as follows: ; (9) Among them, is the safety factor, and its value range is .
[0019] Such as Figure 2As shown in the figure, the present invention also provides a control device for a multi-modal intelligent manipulator, which includes a base 1. Above the base 1, there are sequentially connected a large arm fixing seat 2, a large arm 3, an intermediate arm 4, a small arm 5 and an end effector 6. A fixture interface 7 is provided on the end effector 6. On the side of the base 1, there are respectively provided a power interface 8, a bus interface 9, a safety switch interface 10 and a digital output interface 11. The base 1 not only provides the support and stability functions of the manipulator, but also includes multiple interfaces, such as the power interface 8, the bus interface 9, the safety switch interface 10 and the digital output interface 11, so as to realize power supply, signal transmission and external device control. The power interface 8 is connected to a power adapter, and the safety switch interface 10 is connected to a short-circuit interface plug. Users can also lead out wires to externally connect a switch. When the safety switch interface 10 is short-circuited to GND, the manipulator can move normally. If the safety switch is open to GND, the manipulator immediately stops moving and waits to continue moving after the signal is short-circuited to the ground. The large arm 3 and the small arm 5 of the manipulator are connected by multiple joints, forming a flexible motion structure. Among them, the large arm 3 is embedded with an integrated controller, providing the manipulator with efficient motion control and signal processing capabilities. The small arm 5 is connected to the large arm 3 through the intermediate arm 4 and is connected to the fixture interface 7 of the end effector 6 at the end, and can install a variety of end tools (such as grippers, suction cups, etc.) to realize specific operation functions.
[0020] The maximum horizontal rotation radius of the manipulator device is set to 661.5 mm, the vertical movement range is set to 1122 mm, and the interference radius of the base 1 is set to 120 mm. The manipulator has excellent motion performance and high-precision control capabilities, as shown in Table 1 specifically.
[0021] Table 1 Details of the mechanical data of the manipulator device ;
[0022] During the specific implementation process, the operations for controlling the application of the manipulator through software are as follows; 1. Manipulator reset and calibration Open the manipulator control software, select the serial port number connected to the manipulator, and click to open the serial port; drag the "robot arm calibration" action in the action list to the selected action list; click the "Execute Action" button, and the manipulator will automatically start calibration. After starting the automatic calibration, each axis rotates in the predetermined direction and stops after reaching the corresponding positioning switch. After all axes reach the positioning points, the manipulator rotates to a fixed starting posture; After the manipulator reset is completed or the calibration is completed, it enters the waiting motion instruction state. If a motion instruction from the upper computer is received, the manipulator immediately calculates and interpolates from the current position to the target point. If there is no motion instruction, the manipulator remains in the current posture without moving. At the same time, the manipulator sends the status information including the positioning coordinates, axis angles, etc. to the upper computer in real time.
[0023] 2. Basic Process Motion Control
[0024] After opening the serial port using the software, the current 3D coordinates (X, Y, Z) of the robotic arm and the angles of each axis of the robotic arm (a1 - a6) will be displayed above the software interface. Users can control the robotic arm by changing the 3D coordinates or by changing the angle of each axis. The unit of the 3D coordinates is millimeters (mm), and the unit of the angle is degrees. Drag the "Change Coordinates" function from the action list into the selected action list. Enter the coordinates (X, Y, Z) that the user expects the robotic arm to reach in the input box of the function. The entered 3D coordinates need to be separated by English commas. Then click the "Execute Action" button below the software, and the robotic arm will automatically move to the corresponding angle.
[0025] 3. Angle Setting of the End Effector of the Robotic Arm
[0026] When using the robotic arm to grasp an object, sometimes it is necessary to grasp it from different angles. For example, when grasping a sphere such as a small ball, it can be directly grasped from above, while when grasping a cylinder such as a bottle, it needs to be grasped at an angle slightly tilted towards the horizontal direction. Move the "Change Gripper Terminal Angle" function to the front before using the "Change Coordinates" function; The angle in the "Change Gripper Terminal Angle" function represents the angle between the end effector of the robotic arm and the direction perpendicular to the ground. If it is set to 0 degrees, the robotic arm will grasp the object from directly above. If it is set to 90 degrees, the robotic arm will grasp the object from the horizontal direction.
[0027] 4. Single Axis Angle Setting of the Robotic Arm
[0028] The "Change Single Axis Angle" function can change the angle of a specified axis while keeping other angles unchanged. Drag the "Change Single Axis Angle" function from the action list into the selected action list, enter the axis to be transformed and the expected angle value to be transformed. Then click the "Execute Action" button below the software, and the robotic arm will transform the corresponding angle.
[0029] 5. Calling the Camera
[0030] The VB-1 vision platform configures a dedicated vision camera and algorithm for target detection and target 3D coordinate detection of the robotic arm. The steps to call the camera are as follows: (a) Drag the "Turn on the Camera for Target Detection" function from the action list into the selected action list; (b) Enter the name of the object to be detected in the action list. After clicking the execute action, wait for a period of time, and the detection result of the camera will pop up.
[0031] 6. Large Model + Vision Control of the Robotic Arm
[0032] The software supports using a binocular camera to provide visual information for the manipulator, and uses a large model to assist the user in controlling the manipulator. The hand follows the instruction to drag the "Visual Large Language Model Understanding Instruction Function Usage" function from the action list into the selected action list, and inputs "the manipulator follows the hand movement", and then executes the action; After waiting for a while, the images of the left and right cameras will pop up. In the left camera, the detected hand box, the detection confidence, and the three-dimensional coordinates of the hand will be displayed. At this time, press the "a" key on the keyboard, and the manipulator will start to follow the hand movement. Press the "q" key on the keyboard, and the manipulator will stop moving. The principle of this algorithm for detecting three-dimensional coordinates is to utilize the parallax between the two cameras. Therefore, it is necessary to ensure that the hand appears in both cameras simultaneously to obtain three-dimensional coordinates. Otherwise, only two-dimensional coordinates can be obtained.
[0033] 7. Semantic Understanding of Large Language Model
[0034] Drag the "Visual Large Language Model Understanding Instruction Function Usage" function from the action list into the selected action list, and input the user's command, such as "pour the water from the bottle into the cup", and then execute the action.
[0035] After waiting for a while, the images of the left and right cameras will pop up, and the objects expected by the user in the images, such as the bottle and the cup, will be detected. At this time, press the "a" key on the keyboard, and the manipulator will start to execute the user's instruction. Press the "q" key on the keyboard, and the manipulator will stop moving.
[0036] Therefore, the present invention adopts a control method and device for a multi-modal intelligent manipulator as described above. Through deep learning and intelligent control algorithms, the manipulator can autonomously identify targets, optimize the execution path, and adapt to environmental changes. It can integrate multiple modal information to achieve more efficient, accurate, and natural intelligent operations, and is applicable to multiple application scenarios such as industrial manufacturing, medical assistance, agricultural automation, and home services, and provides a more personalized and intelligent operation method according to user needs.
[0037] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A control method for a multi-modal intelligent manipulator, characterized in that: It includes the following steps: S1. Capture the left and right images of the target object in the same scene through two cameras based on binocular vision technology; S2. Identify and locate the target object through the YOLO model, and output the category and bounding box coordinates of the target object; S3. Extract the center point coordinates of the target object in the left and right images based on S1 and S2; S4. Match the objects in the left and right images that have the same category as the target object and y the closest horizontal axis coordinates; S5. Calculate the disparity between the left and right images and the depth information of the object to combine the three-dimensional spatial coordinates of the target object; S6. Record the historical coordinates of S5, smooth the historical coordinates of each target object to calculate the actual three-dimensional spatial coordinates, and combine them with the detection label and bounding box coordinates to output the target detection result of the YOLO model; S7. Input the target detection result of S6 into the task decision model to convert the task information into structured instructions; S8. The manipulator executes the grasping action according to the structured instructions of S7 and calculates and adjusts the grasping force.
2. The control method of a multi-modal intelligent manipulator according to claim 1, characterized in that: In S4, by comparing the categories of objects in the left and right images, it is determined whether they are the same object. If they are the same object, the y horizontal coordinate difference of the axes of the objects in the left and right images is calculated, and the object with the smallest difference is selected for matching; if no matching object can be found, it is regarded as an unmatched object. At this time, based on the two-dimensional coordinates of the object, the camera position needs to be adjusted and start over from S1.
3. The control method of a multi-modal intelligent manipulator according to claim 2, characterized in that: The calculation process of the disparity between the left and right images and the depth information of the object in S5 is as follows: Parallax Calculation formula: ;(1) Among them, and are the horizontal pixel coordinates of the target object in the left and right figures, respectively. Depth Z Calculation formula: ;(2) Among them, B is the baseline of the camera, in millimeters, is the focal length of the camera, in millimeters; 100 is the unit conversion factor.
4. The control method of a multi-modal intelligent manipulator according to claim 3, characterized in that: The formula for calculating the actual three-dimensional spatial coordinates in S6 is as follows: ;(3) ;(4) Among them, and are the center point coordinates detected for the target object in the image, is the pixel size, in millimeters per pixel, ( ) is the converted actual three-dimensional space coordinates.
5. The control method of a multi-modal intelligent manipulator according to claim 4, characterized in that: The specific content of the task decision model in S7 is as follows: S71. The GPT-4o model uses the Transformer structure to encode the input natural language task instructions, extract the context semantic information, and combine the pre-trained knowledge for task understanding. The Vision Encoder in the model extracts features from the input image to generate a high-dimensional visual feature vector, and annotates the key targets, such as object detection and semantic segmentation. The text and visual features are aligned through the Shared Latent Space to ensure the correlation of information between the two modalities; S72. Based on the knowledge base pre-trained by the GPT-4o model, combined with Prompt-based Learning to parse the user's instructions, identify the core task objectives and constraints, and use the dependency analysis and entity recognition NER methods to disassemble the executable task information into structured instructions for the manipulator to execute.
6. The control method of a multi-modal intelligent manipulator according to claim 5, characterized in that: The specific process of the manipulator calculating and adjusting the grasping force in S8 is as follows: S81. Initially determine the opening degree of the gripper according to the geometric information of the target object obtained by the binocular vision technology in S1, and then dynamically estimate the mass of the target object in combination with the information feedback by the three-dimensional tactile sensor set on the manipulator. The specific process is as follows: In the static state, after the gripper lifts the target object, the three-dimensional tactile sensor can sense: ; (5) Among them, is the tangential force of the jaw, is the mass of the object, is the gravitational acceleration; Estimated mass is as follows: ;(6) S82. The process and formula for calculating the grasping force are as follows: ;(7) Among them, is the normal force applied to both sides of the jaw, is the friction coefficient between the jaw and the target object; S83. Determine whether the gripper slips off according to whether the tangential force exceeds the friction limit. The specific determination formula is as follows: ;(8) Among them, is the current normal force, is the magnitude of the current tangential force, is the tangential force of the jaw in the axis direction, is the tangential force of the jaw in the axis direction; S84. If the determination result of S83 is that the gripper slips off, adjust the grasping force according to the estimated mass in S71. The adjustment formula is as follows: ;(9) Among them, is the safety factor, and its value range is .
7. A control device for a multi-modal intelligent manipulator, characterized in that: Apply a control method for a multi-modal intelligent manipulator as described in any one of claims 1-6. The device includes a base, above which a large arm fixing seat, a large arm, an intermediate arm, a small arm, and an end effector are sequentially connected. A fixture interface is provided on the end effector, and a power interface, a bus interface, a safety switch interface, and a digital output interface are respectively provided on the side of the base.
8. The control device of a multi-modal intelligent manipulator according to claim 7, characterized in that: The maximum horizontal rotation radius of the manipulator device is set to 661.5 mm, the vertical movement range is set to 1122 mm, and the interference radius of the base is set to 120 mm.
Citation Information
Patent Citations
Systems and methods for multilingual text generation
CN113228030A
Robot grabbing detection method based on multi-mode visual information fusion
CN115861999A
A method for intelligent grasping manipulator based on monocular depth estimation and multimodal positioning
CN119748449A
Robot control method and system based on large language model
CN119871428A
Cited By
Double-mechanical-arm collaborative goods taking method and system based on multi-source vision cross-view-angle fusion
CN121223808A