Perception, understanding and control method of robot
By dividing the embodied intelligent robot into two modules, "robot body" and "actuator unit", and optimizing computing resources and data transmission, the problem of low output frequency of multimodal large models is solved, and efficient perception, planning and control are achieved.
Patent Information
- Application Number
- CN202411685953.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-23
- Publication Date
- 2025-05-27
Smart Images

Figure CN120038739A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robots, and particularly to a method for perception, understanding, and control of a robot. Background Art
[0002] The perception, planning, decision-making, and control of embodied intelligence include an embodied brain system and an embodied cerebellum system. As Figure 2 shown, it demonstrates the working mode of the current system. Brain system: The robot command and dispatch center with a multimodal large model as the core, taking text and images as inputs and outputting voice interaction and task decision-making information. Cerebellum system: The specific task execution model, including a trajectory planning module and a motion controller.
[0003] The multimodal large model integrates the interaction capabilities output by the understanding, planning, and action models. Such a model is an end-to-end vision-language-action model that directly outputs the pose of the six degrees of freedom at the end of the robot that needs to be executed. For example, the RT-2 model of the Google DeepMind team can directly output actions from its model.
[0004] However, the output frequency of the multimodal large model is relatively low, while the requirements for the frequency of motion planning and motion control of embodied intelligence are much higher than the output frequency of the multimodal large model; and the computational resource cost and training data cost required to increase the output frequency are extremely expensive. Summary of the Invention
[0005] The present invention relates to a method and device for perception, understanding, and control of a robot. The device, as Figure 3 shown, includes a processor (510, 540), a memory (541), sensors (511, 542), an actuator (521), communication interfaces (512, 543), and a communication bus (531). For ease of understanding, Figure 3 and Figure 2 part of the schematic in
[0006] adopts a similar description, but these schematics do not limit the scope of application of the present invention. Figure 3 shown, components 540, 541, 542, 543 belong to the "robot body"; components 510, 511, 512, 521 belong to the "actuator unit". The "robot body" and the "actuator unit" are connected and communicate through the communication bus 531.
[0007] The types of "robot body" include, but are not limited to, single-arm robots, dual-arm robots, multi-arm robots, fixed-station robots, humanoid robots, wheeled mobile robots, multi-legged mobile robots, and other robot bodies that can be equipped with robotic arms. The "actuator unit" includes the end of the robotic arm actuator that works in coordination with the "robot body", the Gripper, and other end devices of the robot actuator. In the following text, the Gripper mentioned refers to the above-mentioned "actuator unit".
[0008] There is no need for a rigid connection between the "actuator unit" and the "robot body", that is, there is no need for pre-calibration, deployment, testing, and maintenance, so there is no need to meet the high-precision requirements required by common industrial robots and industrial robotic arms.
[0009] The embodied intelligent robot described in this method includes two main modules: the "robot body" and the "actuator unit". Figure 1 The simplest functional schematic diagram of this method is described, mainly including two main modules: the "robot body" (010) and the "actuator unit" (011). For the convenience of understanding, in Figure 1 the shown MVP schematic diagram, the "robot body" (010) is equipped with a two-dimensional sensor (012), while the end effector of the "actuator unit" (011) is equipped with a three-dimensional sensor (013).
[0010] The two-dimensional sensor (012) operates at a relatively low frequency and is used to understand the environment of the robot actuator's operating space. It is mentioned above that the brain system will perform environmental understanding, task understanding, and task decomposition. Since the brain system, that is, the multi-modal large model, needs a large amount of data training to obtain generality. Therefore, general sensors that are easier to obtain input data have the advantage of being easier to collect data. Therefore, the data of the two-dimensional sensor is easy to obtain and easy to process and train.
[0011] The three-dimensional sensor (013) operates at a relatively high frequency, can output its own position and attitude coordinates in the physical space, and can simultaneously sense the point cloud information coordinates of objects within a certain range.
[0012] Figure 1 In the method shown, the "robot body" can perceive the environment and calculate the pose of the six degrees of freedom of the robot end that needs to be executed; the "actuator unit" can quickly perceive the physical relationships in the space, return the real-time pose coordinates of the robot actuator and the changes in environmental obstacles during movement, and provide data input for motion planning (cerebellum system); thereby improving the accuracy of perception and understanding of the embodied intelligent robot, as well as high-speed planning, decision-making, and control.
[0013] In the design of the above method, different from general robot designs, different processors are allocated to the "robot body" and the "actuator unit". This effectively optimizes and allocates the computing resources, data transfer, and sensing devices between the two main modules of the "robot body" and the "actuator unit", improving the working efficiency and economy of the embodied intelligent robot.
[0014] The above-mentioned "general robot design" is, for example, the very typical Mobile ALOHA designed by Stanford University, which uses a single processor (A consumer-grade laptop with Nvidia 3070 Ti GPU and Intel i7-12800H).
[0015] The above description is for simplified expression and does not represent the limitations of the application of this method.
[0016] Combined Figure 3 , we will further describe the above method and device.
[0017] The sensor 511 includes multiple modules.
[0018] 1. High-speed spatial positioning module: This module can calculate the "position and pose" of the end effector in real time, that is, Tracking and Location, without relying on complex robot body designs and expensive high-performance motors to estimate values, thus solving the problem of cumulative errors in past estimated pose values. The sensors include but are not limited to one or more image sensors, acceleration sensors, gravity sensors, magnetic sensors, etc.
[0019] 2. High-speed spatial perception sensor array: The end of the gripper is equipped with a high-speed sensor array, and these sensors can output spatial features at high frequencies. Spatial features include but are not limited to point feature descriptions and coordinates, line feature descriptions and coordinates, point cloud information descriptions and coordinates, object recognition descriptions, etc.
[0020] 3. The above modules may reuse the same physical units considering economy.
[0021] The actuator 521 also has the ability of state perception. It has the state of grasping and the force feedback of grasping: One or more sensors are used to complete the detection and feedback work of the state of grasping and the force of grasping, including image sensors, force feedback sensors, analog force feedback sensors, six-axis force sensors, acceleration sensors, gravity sensors, magnetic sensors, etc.
[0022] The processor 510, which belongs to the edge computing type processor, provides limited computing power and its main function is to simplify signal transmission. It includes computing resources and algorithms, and can directly calculate the high-speed signals collected above into the target data required by the robot body.
[0023] The communication interface 512. Since the above processor reduces the need for high-speed signal transmission such as image information, it reduces the complex transmission cables. This means that existing buses such as RS485 and CAN can be reused for transmission, instead of using new high-speed signal transmission methods such as USB3.0, which greatly reduces the implementation cost.
[0024] The communication interface 543 is connected to the communication interface 512 through the communication bus 531. This bus connection method can achieve one master and multiple slaves. It enables the robot body to simultaneously and dynamically mount multiple execution units.
[0025] The sensor 542 is located on the robot body. Just like a person's eyes, it is generally located at the head position of the robot. As mentioned above, the generality required by the large model depends on a large amount of data training. Therefore, the sensor 542 provides the necessary data input for the large model. It is convenient for the model to identify the environment where the robot operates, the task state, target recognition, the position of the target, etc. The sensor 542 may be the most common two-dimensional sensor, 2D RGB Image Sensor; but depending on the model and the working environment, there may also be other sensors or sensor combinations, such as infrared sensors, structured light sensors, TOF sensor groups, etc.
[0026] The processor 540 and the memory 541 are the physical "brains" of the robot. The brain system and the cerebellum system are both deployed here. As mentioned in the above background, the multi-modal large model integrates the interaction capabilities of understanding, planning, and action model output. Such a model is an end-to-end vision-language-action model that directly outputs the pose of the six degrees of freedom of the robot end. At the same time, the cerebellum system calculates and outputs the motion trajectory, control parameters, etc. of the robot actuator according to the pose of the six degrees of freedom of the robot end mentioned above.
[0027] To further understand the working mode of the method of the present invention, two sub-modules of this method are further described: Figure 4 As shown, it is the "high-speed spatial positioning sub-module at the actuator end", whose function is to obtain the real-time position and pose of the Grippers in the main system of the robot body, and complete the coordinate binding and coordinate synchronization during this period.
[0028] Step S101: After the Gripper spatial positioning sensor module is powered on, it starts autonomous operation, calculates the "original pose coordinates", and outputs them externally at a frequency of 60HZ. The above spatial positioning module refers to the processor 510 and the sensor 511, and the calculated coordinates refer to the three-dimensional spatial coordinates and vectors that conform to the Cartesian coordinate system. The origin of the "original pose coordinates" is the initialization position of the spatial positioning module. At this time, the "original pose coordinates" already have a physical world reference, which is completed and takes effect after calibrating the camera and the module at the factory, that is, the pose coordinates and their continuous movements already have a physical reference and data accuracy, and this accuracy can reach the level of 0.1 mm.
[0029] Step S102: A rigid connection is adopted between the Gripper sensor and the execution end (fingertip) and is calibrated at the factory; the "original pose coordinates" are subjected to the first coordinate transformation according to the factory calibration to obtain the "execution end pose coordinates". A rigid connection means that two points A and B are directly connected by an undeformable physical method and are not affected by environmental changes such as time, temperature, and humidity.
[0030] The actuator 521 and the sensor 511 are directly processed in a rigid connection manner, and the relative position is calculated during the design. The accuracy of this relative position is guaranteed by the accuracy of the processing process. Also, because the designed distance between the actuator 521 and the sensor 511 is short, the actual error is very low and can be ignored. According to the above method, the pose coordinates of the actuator are calculated by conversion.
[0031] Step S103: The Gripper and the robot body are dynamically bound through machine vision algorithms. After the binding is synchronized, the relative coordinates are calculated; the "execution end pose coordinates" are subjected to the second coordinate transformation according to the above relative coordinates to obtain the "relative robot pose coordinates 1".
[0032] The above "dynamic binding of machine vision algorithms" can be implemented in the following multiple ways.
[0033] 1. Add a positioning identifier to the robot body. The sensor of the Gripper module is used to identify the special identifier on the robot body for positioning, such as a QR code, etc.; according to the relative relationship between the two, the coordinates of the robot body can be calculated, and the origin of the coordinates at this time is the origin of the coordinate system of the execution end pose coordinates.
[0034] 2. Add a positioning unit with the same technology as the above Gripper module to the robot body. The above positioning unit can calculate the coordinates of the robot body, and the origin of the coordinates at this time is the origin of the coordinate system initialized by the robot body. Since the two coordinate systems use positioning units with the same technology, the spatial correspondence relationship between the two origins can be calculated.
[0035] 3. Input the three-dimensional features of the Gripper with a specific shape described in this method into the large model. Through the input image information, the large model algorithm identifies the Gripper with the specific shape described in this method and its position in the robot space, and completes coordinate synchronization. Subsequently, the pose coordinates are continuously updated through the Gripper module.
[0036] The above synchronization process is calculated by processors 510 and 540 and synchronized through the communication bus 531. Therefore, processor 540 obtains the accurate pose information of actuator 521 in the robot operation space, "relative robot pose coordinate 1".
[0037] Step S104. Similarly, if there is another group of Grippers, "relative robot pose coordinate 2" is calculated in the same way.
[0038] This method adopts dynamic binding between the robot body and the actuator. Therefore, there is no need for prior fixation between them because the position can be adjusted at any time. The robot body and the actuator can be bound at any time according to the above method, even during the operation of the actuator. Therefore, this method supports the binding of one robot body to multiple actuators. A typical scenario is a dual-arm robot.
[0039] Therefore, this method also supports the actuator to switch to another robot body after binding to one robot body to complete the relay work of the actuator. A typical scenario is the pipeline operation between two fixed-station robots.
[0040] Step S105. Obtain and maintain the pose coordinates of multiple 60HZ Grippers in the robot system.
[0041] The above robot system refers to the processor 540 and the memory 541 of one robot body. This method is not limited to the mutual binding and exchange between multiple processors 540, memories 541 that make up the robot body, and multiple actuators 521.
[0042] The above is the workflow of the "high-speed spatial positioning sub-module at the actuator end".
[0043] Figure 6 As shown, it is the "actuator motion path planning sub-module", which describes how this method completes the motion path planning and execution of the actuator from point A to point B.
[0044] Step S301: The brain system outputs the latest six - degree - of - freedom position and pose B point to be executed. The above - mentioned brain system refers to a multi - modal large model including a processor 540 and a memory 541. It calculates the pose that the actuator 521 needs to move to, the B point. The B point is generally the target position of one of the subtasks after the task - completion action is disassembled.
[0045] Step S302: The Gripper module calculates the six - degree - of - freedom pose A point at the current position. For ease of description, the Gripper module mentioned below refers to the Gripper device including the sensor 511, the processor 510, and the actuator 521 of this method. For ease of description, the six - degree - of - freedom pose A point of the Gripper module refers to the relative robot pose coordinates obtained in the above calculation.
[0046] Step S303: The cerebellar system, the processor calculates the motion planning path from point A to point B.
[0047] The above - mentioned motion planning path has taken into account the known obstacles in the current operating space.
[0048] The above - mentioned motion planning path has taken into account the physically movable range of the robot actuator.
[0049] To improve the operating efficiency of the cerebellar system, common motion planning paths can be saved, so that there is no need to calculate again during execution, thus improving the efficiency.
[0050] Step S304: The cerebellar system performs motion control according to the above - mentioned motion planning and starts to execute. According to the motion trajectory planning in S303, the robot controls each executable unit to perform translation and rotation operations. There are many mature control methods for the motion control of this part of the robotic arm, which will not be elaborated here.
[0051] Step S305: Determine whether there is an abnormality in the verification between the real - time A1 point calculated by the Gripper module and the planned trajectory. During the execution of step S304, the Gripper module still works continuously and outputs the latest pose coordinate A1 point. The system checks the execution process of step S304; when the verification is incorrect, step S302 is executed again. When it is normal, step S304 continues to run. This method helps the cerebellar system to perform real - time inspection and feedback on the execution results during the execution of the actuator's motion, thus improving the accuracy of control.
[0052] Step S306: Determine whether the high-speed sensor senses environmental changes and obstacle changes. During the execution of Step S304, the Gripper module continues to work, outputting real-time environmental and obstacle data around the actuator, which is output to the processor 540 of the robot body. The processor 540 will analyze the data to determine whether the environment and obstacles have changed. The condition for whether there is a change refers to whether it affects the path planned for forward movement. If there is a change, the current action will be stopped, and Step S303 will be executed again; if there is no change, Step S304 will continue to run. This method helps the cerebellar system sense and analyze real-time environmental changes during the movement of the actuator, thus improving the real-time performance and accuracy of trajectory planning.
[0053] Step S307: Determine whether the brain system outputs a new pose at point B1 and replaces the original point B. During the execution of Step S304, the brain system may output new instructions and a new task pose at point B1. At this time, Step S304 will be terminated, and Step 301 will be executed again.
[0054] Step S308: Determine whether point B has been reached.
[0055] Step S309: Complete the execution of the movement from point A to point B.
[0056] The above is the working process of the "actuator movement path planning sub-module".
[0057] Additionally, in addition to the main method of directly outputting low-frequency actuator poses based on the multi-modal large model, this method can also have the following variants to further improve the economy of this method. The economy of the additional method is reflected in lower performance requirements for the processor 540 and memory 541 of the robot body.
[0058] The specific description of the additional method is as follows: As mentioned above, end-to-end multi-modal large models, such as OpenVLA_7B, RT-2_55B, etc., have much higher deployment costs than general LLM large language models. In the absence of the above end-to-end large models, that is, when the actuator pose coordinates for disassembling tasks cannot be directly obtained, this method can still work in the following way.
[0059] As Figure 5 shown, is the process of the above method working.
[0060] Step S201: After the Gripper sensor module is powered on, it starts autonomous operation, calculates the "sparse point cloud information" within the sensor range, and outputs it externally at a frequency of 60HZ. The above means that the sensor 511 sends the calculated data to the processor 510, and this process does not occupy the computing power of the processor 540, so it does not occupy the transmission bandwidth of the communication bus 531 either.
[0061] Step S202: The robot body sensor acquires a two-dimensional image; the processor performs target recognition based on the two-dimensional image to obtain the "target position matrix" of the target in the two-dimensional image, and outputs it at a frequency of 30 - 60HZ. The above means that the sensor 542 acquires the two-dimensional image and sends it to the processor 540 for operation. The target object can be recognized through traditional CV methods such as YOLO. Since it is the processing and recognition of a two-dimensional image, this process has relatively low requirements for the computing resources of the processor 540.
[0062] Step S203: Convert and align the above "sparse point cloud information" according to the obtained Gripper "relative robot pose coordinates". Since the information of the above "sparse point cloud information" is calculated and saved according to the original coordinates of the Gripper, their coordinate systems need to be converted into the same coordinate system as the robot body according to the method mentioned above, and thus the data alignment work is completed.
[0063] Step S204: Screen the above obtained "sparse point cloud information" according to the "target position matrix" to obtain the "effective target sparse point cloud information", with an output frequency of 60HZ. After completing Step S203, the data is already aligned. The processor 510 receives the "target position matrix" frame information from the processor 540. The processor 510 and the processor 540 perform timestamp synchronization. The processor 510 sorts the data of each frame of "sparse point cloud information" and each frame of "target position matrix" according to the timestamp arrangement. In the way of adjacent frames, the data of the "sparse point cloud information" is screened frame by frame according to the spatial range pointed to by the "target position matrix". The new "effective target sparse point cloud information" is obtained, and the output frequency is the same, which is 60HZ.
[0064] The above steps are mainly calculated in the processor 510, and this process does not occupy the computing power of the processor 540, so it does not occupy the transmission bandwidth of the communication bus 531 either.
[0065] Step S205: Perform fusion calculation on the obtained "effective target sparse point cloud information" to obtain "effective target point cloud information", with an output frequency of 20 - 40HZ. The above method effectively fuses multiple frames before and after the obtained high-frequency data, sacrificing the frame rate to improve the single-frame accuracy of each frame. At the same time, the data is also effectively screened and compressed.
[0066] Step S206: Calculate the estimated position and attitude of the target in three-dimensional space based on the "effective target point cloud information", with an output frequency of 20 - 40HZ. The "effective target point cloud information" is calculated in the processor 510. After effective compression, it is sent to the processor 540 for calculation through the communication bus 531. The processor 540 and the memory 541 realize the reconstruction of the three-dimensional space, calculate the point cloud contour of the target in the three-dimensional space, and simultaneously estimate the position and attitude.
[0067] After completing the above steps, the recognition, three-dimensional reconstruction, and pose coordinate estimation of the current object are realized.
[0068] Input the obtained target three-dimensional space information into the processor 540. According to the existing mature methods (grasping of objects with known contours), operations such as the grasping of the actuator can be completed, which will not be elaborated technically. Similarly, combined with the more easily obtained and deployed LLM large language model, the tasks of the robot actuator are completed. Description of the Drawings
[0069] To more clearly illustrate the present invention and the technical solutions in the embodiments, the drawings required for use in the description of the invention will be briefly introduced below. Obviously, the drawings in the following description are only some schematic diagrams and embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative labor.
[0070] Figure 1 Describes the simplest functional schematic diagram of the method; Figure 2 Describes the current embodied intelligence brain system and cerebellum system; Figure 3 Describes the device composition in the method and combines with Figure 2 part of the content for description; Figure 4 Describes the running steps of the "high-speed spatial positioning sub-module at the actuator end" of the method; Figure 5 Describes the running steps of the method for estimating the position and attitude of the target object in space; Figure 6 Describes the running steps and logic of the "actuator motion path planning sub-module" of the method; Figure 7 It describes an implementation case of the end - effector unit described in this method; Figure 8 It describes an implementation case of the embodied intelligent robot described in this method. Detailed implementation manners
[0071] Figure 7 As shown, it is an implementation case of the end - effector unit described in this method.
[0072] As Figure 7 As shown, the end - effector unit mainly consists of 4 objects in three parts: 103 Gripper, 101 and 102 Finger, and 104 Gripper Coupling. The 4 objects are fixed together by a mechanical structure.
[0073] 103 Gripper, which is the main part of the end - effector unit, includes various sensors such as 121, 122, 123, 124, 125, includes 131 servo motors, includes 141 computing chip SoC, and includes communication interfaces 151, 152.
[0074] 101 and 102 Finger are replaceable gripper fingertips that are compatible with 103. The fingertips are used in pairs and can replace shapes and materials according to different scenarios to achieve different application functions.
[0075] 104 Gripper Coupling is a replaceable flange that is compatible with 103 and is used to adapt and fix 103 to different robotic arms, making 103 compatible. 153 is the communication interface for connecting with the robotic arm.
[0076] As mentioned above, combined with Figure 3 For easy understanding by comparison, the above - mentioned 510 processor includes 141, the above - mentioned 511 sensors include 121, 122, 123, 124, 125, the above - mentioned 512 communication interfaces include 151, 152, 153, and the above - mentioned 521 actuators include 101, 102, and 131.
[0077] Among the sensors, 123 and 125 sensors are high - speed image sensors. They have a global shutter ability of 60 to 120 Hz and an image acquisition performance of more than 300,000 pixels, ensuring high frame rate and high clarity of image data.
[0078] The 124 sensor is a high - performance 6 - axis inertial measurement unit (IMU), and these sensors can provide accurate motion and orientation data.
[0079] 123, 124, and 125 together constitute the high-speed spatial computing module.
[0080] 123 and 125 together constitute the high-speed spatial perception sensor array.
[0081] The above-mentioned 123 and 125 sensors have a certain field of view angle and can sense the activity range of 101 and 102. Therefore, they can provide effective information output for the robot to grasp objects.
[0082] 121 and 122 are six-axis force sensors that provide key force feedback when the fingers 101 and 102 grasp an object. These fingers are controlled by a transmission device driven by a 131 servo motor, and the 131 servo motor itself is a high-performance servo motor with a force feedback detection function, ensuring the accuracy and controllability of the gripper's actions.
[0083] In terms of communication, the end effector is equipped with multiple communication interfaces 151, 152, and 153. Among them, interfaces 152 and 153 include interfaces such as RS485 bus and CAN bus for data communication with the robot body. In addition to supporting the above communication protocols, the 151 interface also provides a USB 2.0 / 3.0 debugging interface and an expansion interface for system debugging and function expansion.
[0084] Rigid connection, in the above method, a certain rigid connection design is required in the Gripper product to ensure normal operation; the rigid connection design includes the rigid connection between 123, 124, and 125, the rigid connection between 101, 102, and 103, and at the same time, the rigid connection between the 123, 124, 125 module and 103.
[0085] It should be added that the above-mentioned 103 refers to a part of the structure of the end effector, and the size of this structure only needs to support the above-mentioned rigid connection requirements, so the volume is small.
[0086] Figure 8 Shown is an implementation case of a dual-arm robot applying this method.
[0087] Among them, 100 and 300 are the end effectors shown above Figure 7 in the figure.
[0088] At the same time, 440 at the bottom is the robot fixed platform, and 410 is the operating object placement platform.
[0089] The two robotic arms of the dual-arm robot, 281 and 282, are fixed on the bottom platform 440.
[0090] At the same time, the robot body 200 is also fixed on the bottom platform 440.
[0091] The objects on the 440 platform are randomly fixed without fixed positional relationships, so there is no need for manual maintenance of fixed positions.
[0092] 100 and 300 are respectively connected to the 281 and 282 robotic arms through flanges.
[0093] In the 200 robot body, there are 241 and 242 processors and memories, 221 sensors, 252 communication interfaces, and 250 communication buses.
[0094] Meanwhile, referring to Figure 3 As shown, 540 and 541 correspond to 241 and 242, 543 corresponds to 252, 531 corresponds to 250, and 542 corresponds to 221.
[0095] In this example, both 152 and 352 are UART interfaces, 252 is a USB interface, and they are interconnected through the 250 bus (RS485).
[0096] The above is the static description part of this example; to better illustrate the implementation of this method, below we will start the above robot and execute an actual task to demonstrate how the robot works. The following is the operation process of this example.
[0097] The task executed by the robot is: grasping the 400 target object and moving it to the specified position.
[0098] Task input: The 241 and 242 robot body processors obtain the above task information.
[0099] Brain system.
[0100] Task understanding: The 241 and 242 robot body processors start to understand the process of the task. First, they understand that only one Gripper of the 100 actuator is required to execute the task this time. Then, they need to determine the current position of the 400 object, then confirm the basic shape of the 400 object, then confirm the position of the grasping point on the 400 object, then confirm the position where the 100 actuator should move to, then confirm whether the grasping is successful, and then move the 100 actuator to drive the 400 object to the specified position, and the task is completed.
[0101] Environmental understanding: The multi-modal large model deployed in the 241 and 242 robot body processors starts to calculate how the robot should complete the corresponding operations based on the continuous images input by the 221 sensors and the above understanding of the task.
[0102] Task Disassembly: After understanding the above tasks, disassemble the tasks in chronological order. First, the 100 end effector needs to be moved to position (B). The information of position (B) includes the three-dimensional coordinates in the robot operating space that the 103 needs to move to, the angle that the 103 needs to move to, and the opening and closing sizes of the 101 and 102.
[0103] The 100 end effector starts to activate.
[0104] First, the 100 end effector starts to initialize. The 123 and 125 image sensors and the 124 IMU sensor start to output the pose coordinates (a) for high-speed positioning, and use the initialized spatial position as the origin of the coordinate system.
[0105] Then, according to the known connection relationship between the module and the 103 Gripper structure, and the rigid connection relationship between the fingertips of the 101 and 102, the pose coordinates (a1) are calculated. The pose coordinates (a1) represent the pose coordinates of the end effector, that is, the representation of the grasping coordinates understood by the robot model.
[0106] The 252 identifier is fixed on the robot body, and its relative position relationship with the 221 sensor is known.
[0107] The 125 recognizes the 252 identifier. The 252 identifier is a positioning identifier in the coordinate system and is a special QR Code.
[0108] The above pose coordinates (a1) are converted to obtain the pose coordinates (A). The pose coordinates (A) are the same as the coordinate system understood by the robot body.
[0109] At this time, in the robot system, the coordinate system of the 400 object is synchronized with the coordinate system of the 100 end effector. At this time, the real-time pose coordinates of the 100 end effector are (A).
[0110] After the above synchronization is completed, it is no longer necessary to recognize the 252 identifier until it is necessary to synchronize again.
[0111] Cerebellum system.
[0112] Motion Planning: After knowing the pose coordinates (A) of the 100 end effector and the task coordinates (B), start planning the motion path of the 100 end effector from (A) to (B).
[0113] Through the information input by the 221 sensor, the environmental perception data calculated, and the large model can recognize the known obstacle objects.
[0114] The cerebellum system calculates the optimal motion path based on the above environmental perception data.
[0115] Motion control: According to the above motion path, the robot 241 processor starts to control different actuators (211, 212, 213, 214) to perform rotation tasks according to parameters.
[0116] Cause the 100 actuator to move from (A) to (B).
[0117] During this process, 123 and 125 sense that the point cloud information in the operation space has changed, and send the above information to the 241 processor through the 250 bus for analysis and calculation.
[0118] The cerebellar system in the 241 processor runs efficiently, analyzes that the above point cloud information will not interfere with the above motion path, and the operation control continues to execute.
[0119] The 100 actuator continuously outputs a new (A) and sends it to the 241 processor for verification. The result is consistent with the planned motion path, and the operation control continues to execute.
[0120] The 100 actuator continuously outputs a new (A) and sends it to the 241 processor for verification. When the value of (A) is the same as (B), it is determined that the current task is completed, and the 100 actuator reaches (B).
[0121] The brain system outputs a new position (B1) that the 100 actuator needs to move to.
[0122] The above task process runs repeatedly in the same way.
[0123] Then the robot executes the grasping subtask.
[0124] The large model outputs the suggested grasping point position (range) on the 400 object.
[0125] As mentioned above, 123 and 125 sense the point cloud information in the operation space and send the above information to the large model of the 241 processor through the 250 bus.
[0126] The 241 processor calibrates the grasping position according to the suggested grasping point position (range) and the collected real-time point cloud information, and finally outputs the grasping point position.
[0127] According to the finally output grasping point position, correct the current point (B2) of the 100 actuator in the above task.
[0128] The above task process runs repeatedly in the same way.
[0129] The 100 actuator reaches (B2), and at the same time the 131 servo motor also starts to work, moving the fingertips of the 101 and 102 grippers to the position required by (B2).
[0130] End the task (B2) until it is determined that the grasping is successful.
[0131] Determine that the grasping is successful through the 124 sensor, 121, and 122 sensors.
[0132] Drive the 400 object to the specified position with the 100 actuator.
[0133] The task is completed.
[0134] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A robot system, characterized in that: include: The robot body, which is equipped with two-dimensional sensors to understand the environment of the robot's actuator operation space; The actuator unit is equipped with a three-dimensional sensor for outputting the position and posture coordinates of the actuator in the physical space.
2. The robot system according to claim 1, characterized in that: The operating frequency of the three-dimensional sensor is higher than that of the two-dimensional sensor.
3. The robot system according to claim 1, characterized in that: The system allocates different processors to the "robot body" and the "actuator unit", wherein the "robot body" includes a higher-performance processor and memory, while the "actuator unit" includes an edge computing processor to reduce data transmission so that the communication bus does not need to be upgraded.
4. The robot system according to claim 1, characterized in that: There is no need for a rigid connection between the "actuator unit" and the "robot body".
5. The robot system according to claim 1, characterized in that: A dynamic binding method between a robot and an actuator is included, which allows the robot body to be bound and unbound with multiple actuator units, wherein the binding method is based on data exchange between a communication bus and a processor.
6. The robot system according to claim 1, characterized in that: The actuator unit sensor consists of several modules: A high-speed spatial positioning module that can calculate the "position and posture" of the end effector in real time, including one or more image sensors, acceleration sensors, gravity sensors, and magnetic sensors; A high-speed spatial perception sensor array that can output spatial features at high frequency, including point feature descriptions and coordinates, line feature descriptions and coordinates, point cloud information descriptions and coordinates, and object recognition descriptions; The modules may reuse the same physical units for economic reasons.
7. The robot system according to claim 1, characterized in that: The types of the robot body include but are not limited to single-arm robots, dual-arm robots, multi-arm robots, fixed-station robots, humanoid robots, wheeled mobile robots, multi-legged mobile robots and other robot bodies that can be equipped with robotic arms.
8. The robot system according to claim 1, characterized in that The two-dimensional sensor also includes an infrared sensor, a structured light sensor, a TOF sensor group, etc., depending on the model and the working environment.
9. A robot control method, characterized in that: The following steps are involved: Environmental understanding using 2D sensors; Use three-dimensional sensors to output position and posture coordinates; Use three-dimensional sensors to perceive point cloud information; Assign different processors to the "robot body" and "actuator unit" to reduce data transmission; Use the point cloud information of the end effector to assist the robot in grasping objects; The control method as described in claim 9 is characterized in that the method also includes the actuator unit adjusting the motion path according to the real-time perception data to adapt to environmental changes.
10. The control method according to claim 9, characterized in that: A method for dynamically binding a robot body and an actuator using a machine vision algorithm can be implemented in the following ways: Add positioning marks on the robot body, and use the actuator unit's sensor to identify special marks on the robot body for positioning, such as a QR code. A positioning unit with the same technology as the actuator unit is added to the robot body to calculate the coordinates of the robot body and complete the synchronization; The three-dimensional features of the Gripper of a specific shape are input into the large model, and the Gripper and its position in the robot space are identified through the large model algorithm, and coordinate synchronization is completed.
11. The robot system according to claim 1, characterized in that: The system is allowed to directly output motion coordinate manipulation without relying on an end-to-end large model, but instead only rely on a large language model and the robot system described in claim 1 to perform work in collaboration.