Robot and operation method based on multi-modal fusion in complex restricted environment
Through the multimodal fusion robot design, combined with depth cameras, tactile sensors and control algorithms, medium-sized operation problems in complex and constrained environments are solved, and high-precision and highly intelligent operation capabilities are achieved.
Patent Information
- Application Number
- CN202210737033.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-06-27
AI Technical Summary
Existing robots have difficulty completing medium-sized tasks such as handling and disassembly in complex and constrained environments, and are limited in movement and evacuation in environments where map modeling cannot be done in advance.
The robot design with multimodal fusion is adopted, including a depth camera system, multiple tactile sensors, wireless communication modules and controllers. The robot joint movement is coordinated through the motion control system algorithm, and combined with impedance control, deep reinforcement learning and multimodal information processing to realize the robot's work in complex environments.
It realizes high-precision operation ability in complex and constrained environments, can simulate human behavior and complete tasks, has the ability to work at the same time with high intelligence and dual robotic arms, and is suitable for environments where map modeling cannot be modeled in advance.
Smart Images

Figure CN115319764B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous operation of robots in complex restricted environments. Specifically, it relates to a robot and an operation method based on multi-modal fusion in complex restricted environments. Background Art
[0002] A complex restricted environment is an environment where map modeling cannot be carried out in advance, and entry and evacuation operations are restricted. The operation in restricted spaces involves a wide range of fields and industries, with a complex operation environment, many dangerous and harmful factors, and it is easy to occur production safety accidents, causing serious consequences; when operators are in distress, the rescue is difficult, and blind rescue or improper rescue methods are likely to cause an increase in casualties.
[0003] Most robots generally only apply to environments where map modeling can be carried out in advance and entry and evacuation operations are not restricted. For example, the patent document CN109202885A discloses a material handling mobile composite robot that can send materials into the specified material operation area with the least number of robots and the least number of processes. The material handling environment can be map modeled in advance, and entry and evacuation are not restricted. Therefore, the composite robot can autonomously perform global path planning according to the pre-determined tasks, and then complete the material handling and moving tasks.
[0004] Although there are already some robots for working in restricted spaces, the existing robots for working in restricted spaces can only collect relevant signals and cannot complete medium-sized tasks, such as handling and disassembly. Therefore, it is particularly important to study a robot that can complete medium-sized operation tasks in complex restricted environments. Summary of the Invention
[0005] Aiming at the defects in the prior art, the purpose of the present invention is to provide a robot and an operation method based on multi-modal fusion in complex restricted environments, so as to solve problems such as the harsh working environment in restricted spaces and the difficulty in carrying out work.
[0006] A robot based on multi-modal fusion in complex restricted environments according to the present invention includes a robot body, multiple tactile sensors, a wireless communication module, a human-machine interaction interface, and a controller;
[0007] A depth camera system is provided at the head of the robot body;
[0008] Multiple tactile sensors are provided at the manipulator of the robot body; the multiple sensors are used for the robot to adjust the operation force and operation position through data feedback when operating on an object;
[0009] The human-machine interaction interface is used to help the robot specify the target object and issue operation instructions to the robot;
[0010] The wireless communication module is used for the human - machine interaction interface to communicate with the robot;
[0011] The controller includes a motion control system algorithm;
[0012] Based on the environmental information obtained by the depth camera system, the communication information between the wireless communication module and the artificial interaction interface, and the data feedback of the multi - tactile sensor, the controller generates a motion equation of the robot through the motion control system algorithm, and sends signals to the servos and servo motors of each joint to control the coordinated movement of each joint of the robot.
[0013] Furthermore, the motion control system algorithm includes a dynamic whole - body mobile operation algorithm of the robot based on impedance control, a manipulator control algorithm based on deep reinforcement learning, a target recognition and positioning algorithm, and a multi - modal information processing and classification algorithm. In particular, it also includes a physical constraint condition algorithm for judging whether dangerous operations will occur when the robot is moving as a whole.
[0014] Furthermore, the robot includes a driving structure and an operating mechanism;
[0015] The driving structure realizes a mobile chassis for the robot to move; specifically, a wheeled mobile chassis such as a two - wheel drive structure, a four - wheel drive structure, or a crawler - type drive structure can be adopted;
[0016] The operating mechanism includes a manipulator configuration, and the manipulator configuration includes at least a shoulder joint, an elbow joint, a hand joint, a wrist joint, and a finger joint. Specifically, the operating mechanism can adopt configurations such as a single manipulator or a double manipulator;
[0017] When the driving structure drives the robot to move to a designated working location, the operating mechanism completes the coordinated movement of each joint of the robot under the control of the controller.
[0018] Furthermore, the multi - tactile sensor has at least four - fold tactile sensing of contact pressure, thermal conductivity, object temperature, and ambient temperature; specifically, the multi - tactile sensor can adopt an integrated sensor or can be composed of a combination of a pressure sensor, a heat conduction sensor, and a temperature sensor.
[0019] Furthermore, the robot installs an array of N multi - tactile sensors on each manipulator, where N is not less than 30, and at least 5 multi - tactile sensors are installed on each finger joint;
[0020] The output of the multi - tactile sensor array is an N×m matrix;
[0021] Wherein, N is the number of multi - tactile sensors in the array, and m is the number of types that the sensor can detect, and m is not less than 4.
[0022] Furthermore, the dynamic whole-body movement operation algorithm of the impedance control-based robot enables the robot to utilize its own structure and the environment while performing tasks, and autonomously handle physical constraints and collision avoidance problems;
[0023] The input is the coordinate position of the object and the coordinates of each current joint, as well as the steering angle and rotation angle of the driving wheels, and the output is the coordinates of the target shoulder joint, as well as the steering angle and rotation angle of the driving wheels; the n driving degrees of freedom are grouped according to subsystems and control interfaces; the dynamic equation is:
[0024]
[0025] The vector q ∈ R t represents the joint coordinates of the robot manipulator, w ∈ R s includes the steering angle and rotation angle of the wheels of the robot driving wheels; t represents the driving degrees of freedom of the upper part of the robot; s represents the driving degrees of freedom of the mobile chassis; where n = s + t. g q g(q) represents the gravitational torque that appears at the upper body joints, g b (w) is the reference value. Centripetal / Coriolis is represented by ; τ ext represents the external torque and force; the control inputs are τ w and τ b . M ww , M wq , M qq represent the inertial elements, and conflict avoidance related to preventing physical collisions is added to the main task command.
[0026] When the wheels are aligned with the instantaneous center of rotation (ICR), the consistent movement of the robot can be achieved, and the ICR is defined by the translational and rotational speeds of the mobile chassis: Ψ describes the relationship between the Cartesian velocity of the mobile chassis and the position of the ICR defined by the coordinates x ICR and y ICR . The wheel direction and speed v 1,w -v 4,w are the same as the ICR. By dividing a single low-priority block into multiple subtasks and projecting them into the space of the main task. By using the controller to separate the upper body from the mobile chassis.
[0027] Furthermore, the chassis path planning algorithm is to determine whether the distance between the robot and the target object is less than the threshold after converting the target object coordinates and the actual target coordinates of the robot, and according to the distance coordinates between the robot and the target object, through the kinematic solution equation, output the speed, acceleration, rotation angle, and movement time required for the robot's movement path;
[0028] The kinematic solution equation is as follows:
[0029]
[0030] where (x, y, z) represents the distance the robot needs to move in the world coordinate system, v1, v2, v3…v n represents the speeds required by n leg wheels, a1, a2, a3…a n represents the accelerations required by n leg wheels, θ1, θ2, θ3…θ n represents the desired rotation angles of n leg wheels, t1, t2, t3…t n represents the movement times of n leg wheels; K, P, T are transformation matrices; v 1-begin , v 2-begin , v 3-begin …v n-begin represents the initial speeds of n leg wheels, a 1-begin , a 2-begin , a 3-begin …a n-begin represents the initial accelerations of n leg wheels, θ 1-begin , θ 2-b e gin , θ 3-b e gin …θ n-begin represents the initial rotation angles of n leg wheels;
[0031] Transformation between the target object coordinates and the actual target coordinates of the robot: The target object coordinates are (x d , y d , z d ), where (x d , y d ) represents the horizontal plane coordinates of the target object in the world coordinate system, and z d represents the height of the target object relative to the center point of the robot's head; (x n , y n , z n ) represents the actual target of the robot, and (x n , y n ) represents the horizontal plane coordinates of the actual target of the robot in the world coordinate system, and z n represents the height of the actual target of the robot relative to the center point of the robot's head. The transformation relationship between the two is:
[0032] |(x d , y d , z d ) - (x n , y n , z n ) - (x, y, z)| ≤ (x max, y max , z max )
[0033] |(x d , y d , z d )-(x n , y n , z n )-(x, y, z)| ≥ (x min , y min , z min )
[0034] (x, y, z) represents the distance that the robot needs to move in the world coordinate system, and (x max , y max , z max ) represents the maximum threshold distance of the robot from the target in the world coordinate system; (x min , y min , z min ) represents the minimum threshold distance of the robot from the target in the world coordinate system.
[0035] Further, the target recognition and positioning algorithm is based on the YOLOV4 network. The input is the video stream image of the depth camera system, and the output is a tensor composed of the center coordinates, width and height values of the prediction box, the confidence of the prediction box, and eighty class scores; the obtained prediction box is used as the input to the HED edge detection network to output the coordinates of the object edge points in the image; the HED edge detection network outputs the output of the last convolutional layer of each layer of the five groups of convolutional feature extraction networks and combines them through transposed convolution; finally, the coordinates output by the HED edge detection network are combined with the depth map to output the three-dimensional coordinates of the edge points of the target object, and the pose of the object is obtained.
[0036] Further, the robotic arm control algorithm based on deep reinforcement learning: After completing the initialization of the algorithm weights and loading the algorithm training weights, the robotic arm makes decisions based on the joint coordinate positions returned by the sensors, the robotic arm joint angle states, and the target object coordinates, outputs the predicted control quantities of each joint angle of the robotic arm, and makes the next decision based on the environmental state after the robotic arm executes the motion; through continuous trial-and-error learning, the network parameters approach the direction that enables the robotic arm to learn to accurately approach the target faster until the robotic arm can accurately output a strategy to approach the target object according to the environmental state or the reward fluctuation tends to be stable, and then the training is terminated and the result is output.
[0037] Further, the multi-modal information processing and classification algorithm: After removing the information with large errors, dimension processing, and standardization of the output signals of the multiple tactile sensor arrays, an N×4 signal matrix detected by 4 types of multiple tactile sensors for the object is generated, where N represents the number of multiple tactile sensors in the array. The 4 types of multiple tactile sensor detection signals are: thermal conductivity, contact pressure, object temperature, and ambient temperature; and the time-series signals are imported into the LSTM neural network for training and testing to determine the type information of the grasped object.
[0038] The LSTM neural network includes 3 gates, namely, the input gate, the forget gate, the output gate, and the candidate memory cell and the memory cell component.
[0039] The present invention also provides an operation method for a wheeled humanoid robot based on multi-modal fusion. Using the robot based on multi-modal fusion in a complex and restricted environment described above, the method further includes the following steps:
[0040] Step 1: Robot startup: Click the startup button, and the intelligence of the robot system starts the device through the control button, so that all the electrical components of the intelligent control system on the entire device are powered on and started, and all the electrical components of the controller, the visual sensor of the depth camera system, the multiple tactile sensors, the drive wheels of the mobile chassis, the motor control drive device, and the remote communication module are in the working standby state;
[0041] Step 2: Robot initialization: All the electrical components on the entire device execute and complete the initialization according to the system set parameters, and each joint of the robot is reset to the initial working set state according to the system set;
[0042] Step 3: Manipulator reset: Determine whether each joint is reset. If not, the robot continues to execute the initialization command in Step 2. If it has been reset, then execute the next step;
[0043] Step 4: The robot starts to work: The robot arrives at the specified position, and through the video stream image and the human-machine interaction interface transmitted by the depth camera system, controls the robot to reach the specified working position, selects the target object and the relevant operation mode, and sends the position information and the operation signal of the selected object in the environmental image to the controller;
[0044] Step 5: Video frame capture: Capture the video stream frame to obtain the depth image and the environmental image;
[0045] Step 6: Load the YOLOV4 network: Load the environmental image into the YOLOV4 network for forward testing to obtain the information of the candidate boxes;
[0046] Step 7: Obtain the three-dimensional coordinates of the object: Determine the corresponding candidate box through the position information of the target in the environmental image, and obtain the coordinates of the object in the depth image;
[0047] Step 8: Obtain the pose of the target object: Perform edge detection and image segmentation on the image, and fuse it with the depth image information to obtain the pose of the target object and the three-dimensional coordinates of the edge points;
[0048] Step 9: Calculate obstacle information: Obtain the pose and center point coordinate information of the robot, the target object, and the surrounding obstacles through the depth image and the environmental image;
[0049] Step 10: Chassis movement: If the distance between the robot and the object is less than the threshold, jump to Step 11. Otherwise, combine the distance information between the robot and the object with the obstacle information. The controller generates the chassis movement trajectory of the robot through the chassis path planning algorithm, transmits a signal to the motor driver of the chassis, and performs feedback through the speed measurement module to guide the robot to reach the condition where the distance between the robot and the object is less than the threshold, and then return to Step 5;
[0050] Step 11: Operation preparation: According to the distance between the robot and the object, use the dynamic whole-body movement operation algorithm of the robot based on impedance control method to manipulate the robot's manipulator to approach the object's whole-body movement algorithm;
[0051] Step 12: Calculate physical constraint conditions: According to the physical constraint condition algorithm, judge whether there will be dangerous operations when the robot is moving the whole body. If so, return to the previous step; otherwise, jump to the next step;
[0052] Step 13: Information update request: Capture a video frame, obtain a video stream frame, and obtain a depth image and an environmental image;
[0053] Step 14: Load the YOLOV4 network: Load the environmental image into the YOLOV4 network for forward testing to obtain the information of the candidate boxes;
[0054] Step 15: Obtain the three-dimensional coordinates of the object: Determine the corresponding candidate box through the position information of the target in the environmental image, and obtain the coordinates of the object in the depth map;
[0055] Step 16: Obtain the pose of the object: Perform edge detection and image segmentation on the image, and combine it with the depth map information to obtain the pose of the object and the three-dimensional coordinates of the edge points;
[0056] Step 17: Start operation: According to the distance between the robot and the object and the set operation instructions, perform the specified operation on the object through the manipulator control algorithm based on deep reinforcement learning;
[0057] Step 18: Information feedback: Feed back the 4×30 information matrix measured by the multiple tactile sensors on the manipulator to the controller in real time. If it is determined that the operation is completed, jump to the next step; if modification is required, return to the previous step;
[0058] Step 19: Reset the robotic arm finger joints: The robotic arm control algorithm based on deep reinforcement learning enables the robotic arm finger joints to perform a reset operation and wait for the next operation.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] 1. The wheeled humanoid robot based on multi-modal fusion of the present invention adopts a humanoid design, has a high degree of intelligence, high precision, can work with two robotic arms simultaneously, and can simulate human behavior to complete operations on objects in complex and restricted environments.
[0061] 2. The present invention can be effectively applied to complex and restricted environments, especially environments where map modeling cannot be carried out in advance and the entry and evacuation operations are restricted.
[0062] 3. The robot of the present invention can complete operations on workpieces imitating humans through the robotic arm control algorithm based on deep reinforcement learning, and can send relevant control sequence signals through the controller to ensure that the robot completes the specified operations.
[0063] 4. The technical solution adopted by the present invention is to set the structure of the wheeled humanoid robot based on multi-modal fusion in the form of imitating the human body structure, and perform relevant operations on the control robot through the remote control device. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0065] Figure 1 Figure 1 It is a schematic front view of the overall structure of the present invention;
[0066] Figure 2 It is a schematic side view of the overall structure of the present invention;
[0067] Figure 3 It is a flow chart of the overall operation steps of the present invention;
[0068] Figure 4 It is the robotic arm control process of the present invention;
[0069] Figure 5 It is the system communication diagram of the present invention.
[0070] The figures show:
[0071] 1. Human-machine interface;
[0072] 2. Wireless communication module;
[0073] 3. Depth camera system;
[0074] 4. Power module;
[0075] 5. Robot body;
[0076] 6. Controller. Specific embodiments
[0077] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all fall within the protection scope of the present invention.
[0078] The present invention provides a robot based on multi-modal fusion in a complex restricted environment, including a robot body 5, multiple tactile sensors, a wireless communication module 2, a human-computer interaction interface 1, and a controller 6; a depth camera system 3 is provided at the head of the robot body 5, and multiple tactile sensors are provided at the manipulator of the robot body 5; the multiple sensors are used to adjust the operation force and operation position by data feedback when the robot operates on an object; the human-computer interaction interface 1 is used to help the robot specify the target object and issue operation instructions to the robot; the wireless communication module 2 is used for communication between the human-computer interaction interface 1 and the robot; the controller 6 is provided with a motion control system algorithm; the controller 6 generates a motion equation of the robot through the motion control system algorithm according to the environmental information obtained by the depth camera system 3, the communication information between the wireless communication module 2 and the artificial interaction interface, and the data feedback of the multiple tactile sensors, and sends signals to the servos and servo motors of each joint to control the coordinated movement of each joint of the robot. It also includes a power module 4 to provide power for the operation of the robot.
[0079] Specifically, the motion control system algorithm includes a dynamic whole-body mobile operation algorithm of the robot based on impedance control, a manipulator control algorithm based on deep reinforcement learning, a target recognition and positioning algorithm, and a multi-modal information processing and classification algorithm. In particular, it also includes a physical constraint condition algorithm for judging whether dangerous operations will occur when the robot moves as a whole.
[0080] The working principle of the present invention is as follows:
[0081] First of all, in a complex restricted environment, map modeling cannot be carried out in advance, and the environment where entry and evacuation operations are restricted. Further, a complex restricted environment can be understood as an unstructured environment that cannot be simulated in advance.
[0082] Such as Figure 1 and Figure 2As shown in the figure, the present invention provides a robot based on multi-modal fusion in a complex restricted environment, including a robot body 5, multiple tactile sensors, a wireless communication module 2, a human-machine interaction interface 1, and a motion control system algorithm. The motion control system algorithm includes a dynamic whole-body movement operation algorithm of the robot based on impedance control, a manipulator control algorithm based on deep reinforcement learning, a target recognition and positioning algorithm, and a multi-modal information processing and classification algorithm.
[0083] The wheeled robot can adopt various drive structures such as two-wheel and four-wheel, and can also adopt various structures such as a single manipulator or a double manipulator. The depth camera system 3 is integrated into the head of the robot body 5 to obtain environmental information for map acquisition. It has at least 5 joint structures including shoulder joint, elbow joint, hand joint, wrist joint, and finger joint.
[0084] An array of no less than 30 multiple tactile sensors is installed on each manipulator, with at least 5 multiple tactile sensors installed on each finger joint. The multiple tactile sensors have at least four functions: pressure, heat conduction, object temperature, and ambient temperature. They can be integrated sensors or composed of a combination of a pressure sensor, a heat conduction sensor, and a temperature sensor. The multiple sensors are used to adjust the applied force and working position during object operation through data feedback.
[0085] Generally speaking, the basic composition design of the integrated multiple tactile sensor includes a top thermal film, a porous material, and a cold film. The top thermal film of the multiple tactile is used to detect the thermal conductivity of the contacted object. When pressure is applied to the sensor, the material in the sensor undergoes elastic deformation, which can reduce the porosity of the porous material, thereby increasing its thermal conductivity. The cold film is located on the top and bottom layers and can detect the temperature of the contacted object and the ambient temperature. In addition, the sensor uses a constant temperature difference feedback circuit and the cold film, and the thermal film can achieve temperature compensation to ensure that the test process is not affected by changes in the object temperature and ambient temperature.
[0086] The output of the multiple tactile sensor array is an N×m matrix (N is the number of multiple tactile sensors in the array, and m is the number of types that the sensor can detect). For example, if the multiple tactile sensor can detect 4 types of detection, such as the thermal conductivity of the object, the applied pressure, and the temperature of the object and the ambient temperature, then the output of the multiple tactile sensor array is N×4 (N is the number of multiple tactile sensors in the array, N≥30).
[0087] The human-machine interaction interface 1 is mainly used to help the robot specify the target object and issue instructions such as grasping, disassembling, and returning to the robot. The wireless communication module 2 is used for communication between the human-machine interaction interface 1 and the wheeled humanoid robot based on multi-modal fusion.
[0088] The target recognition and localization algorithm is based on the YOLOV4 network. The YOLOV4 network mainly consists of four modules: the input end, the backbone network, the Neck network, and the Head output end. The input is the video stream image of the depth camera system 3, and the output is a tensor composed of the center coordinates, width, and height values of the prediction box, the confidence of the prediction box, and the scores of eighty categories. The obtained prediction box is used as the input to the HED edge detection network to output the coordinates of the object edge points in the image. The HED edge detection network outputs the output of the last convolutional layer of each layer of the five groups of convolutional feature extraction networks and combines them through transposed convolution. Finally, the coordinates output by the HED edge detection network are combined with the depth map to output the three-dimensional coordinates of the edge points of the target object, and the pose of the object is obtained.
[0089] The dynamic whole-body mobile operation algorithm of the robot based on the impedance control method enables the robot to autonomously handle physical limitations and collision avoidance while performing the main tasks. The input is the coordinate position of the object, the coordinates of each current joint, and the steering angle and rotation angle of the drive wheels. The output is the coordinates of the target shoulder joint and the steering angle and rotation angle of the drive wheels. The n driving degrees of freedom are grouped according to the subsystems and control interfaces. The dynamic equation is:
[0090]
[0091] The vector q ∈ R t represents the joint coordinates of the robot manipulator, and w ∈ R s includes the steering angle and rotation angle of the wheels of the robot drive wheels; t represents the driving degrees of freedom of the upper part of the robot; s represents the driving degrees of freedom of the mobile chassis; where n = s + t. g q (q) represents the gravitational torque that appears at the upper body joints, and gb(w) is the reference value. The centripetal / coriolis effect is represented by . τ ext represents the external torque and force. The control inputs are τ w and τ b . M ww , M wq , M qq represent the inertia elements, and the collision avoidance related to preventing physical collisions is added to the main task command.
[0092] When the wheels are aligned with the instantaneous center of rotation (ICR), the consistent motion of the robot can be achieved. The ICR is defined by the translational and rotational speeds of the mobile chassis: Ψ describes the relationship between the Cartesian velocity of the mobile chassis and the position of the ICR defined by the coordinates x ICR and y ICR . The wheel direction and speed v 1,w -v 4,wThe same as ICR. By dividing a single low-priority block into multiple subtasks and projecting them into the space of the main task. By using the controller 6 to separate the upper body from the mobile chassis.
[0093] The chassis path planning algorithm is as follows: The coordinates of the target object are (x d , y d , z d ), where (x d , y d , z d ) represent the three-dimensional coordinates of the target in the world coordinate system, and h d represents the height of the center point of the robot's head. (x n , y n , z n ) represents the actual target of the robot. (x n , y n ) represents the horizontal plane coordinates of the target in the world coordinate system, and z n represents the height of the center point of the robot's head. The conversion relationship between the two is:
[0094] |(x d , y d , z d ) - (x n , y n , z n ) - (x, y, z)| ≤ (x max , y max , z max )
[0095] |(x d , y d , z d ) - (x n , y n , z n ) - (x, y, z)| ≥ (x min , y min , z min )
[0096] (x, y, z) represents the distance that the robot needs to move in the world coordinate system, (x max , y max , z max ) represents the maximum threshold distance of the robot from the target in the world coordinate system. (x min , y min , z min ) represents the minimum threshold distance of the robot from the target in the world coordinate system. The kinematic solution equation is:
[0097]
[0098] Among them, v1, v2, v3, v4 represent the speeds required for the four leg wheels, a1, a2, a3, a4 represent the speeds required for the four leg wheels, θ1, θ2, θ3, θ4 represent the desired rotation angles of the four leg wheels, and t1, t2, t3, t4 represent the movement times of the four leg wheels. K, P, T are transformation matrices. v 1-begin , v 2-begin , v 3-begin , v 4-begin represents the initial speeds of the four leg wheels, a 1-begin , a 2-begin , a 3-begin , a 4-begin represents the initial speeds of the four leg wheels, θ 1-begin , θ 2-begin , θ 3-begin , θ 4-begin represents the initial rotation angles of the four leg wheels.
[0099] Manipulator control algorithm based on deep reinforcement learning: As Figure 5 shown, after completing the initialization of the algorithm weights and loading the algorithm training weights, the manipulator makes decisions based on the joint coordinate positions returned by the sensors, the manipulator joint angle states, and the target object coordinates, outputs the predicted control quantities of each joint angle of the manipulator, and makes the next decision based on the environmental states after the manipulator executes the action, such as the distance between the finger joint and the target object (step tracking), the manipulator joint angle states (angle tracking), the position of the target object, etc. Through continuous trial-and-error learning, the network parameters are approximated in the direction of enabling the manipulator to learn more quickly and accurately approach the target until the manipulator can accurately output a strategy to approach the target object according to the environmental state or the reward fluctuation tends to be stable, and then the training can be terminated in advance and the result can be output.
[0100] Multimodal information processing and classification algorithm: After removing the information with larger errors, dimension processing, and standardization of the output signals of the multiple tactile sensor arrays, a signal matrix of N×4 (N represents the number of multiple tactile sensors in the array, and 4 is the type that the sensor can detect) about the thermal conductivity, contact pressure, object temperature, and environmental temperature information of the object is generated, and it is imported into the LSTM neural network for training and testing according to the time-series signal to judge the type information of the grasped object. LSTM contains 3 gates, namely the input gate, forget gate, output gate, candidate memory cell, and memory cell component.
[0101] The wheeled humanoid robot based on multimodal fusion adopts a humanoid design, has high intelligence, has high precision, can work with both manipulators simultaneously, and can simulate human behavior to complete operations on objects in complex and restricted environments.
[0102] The overall operation steps of the present invention are asFigure 3 and Figure 5 as shown below:
[0103] Step 1: Robot startup: Click the startup button of the power module 4. The intelligence of the robot system starts the device through the control button, enabling the intelligent control system on the entire device to be powered on and started. All electrical components of the controller 6, vision sensor, multi-touch sensor, drive wheels of the mobile chassis, motor control drive device, and remote communication module are in a working standby state.
[0104] Step 2: Robot initialization: All electrical components on the entire device execute and complete initialization according to the system-set parameters, and each joint of the robot is reset to the initial working setting state according to the system setting.
[0105] Step 3: Manipulator reset: Determine whether each joint is reset. If not, the robot continues to execute the initialization command. If it has been reset, then perform the next operation.
[0106] Step 4: Robot starts working: The robot arrives at the specified position. Through the video stream image transmitted by the depth camera system 3 and the human-machine interface 1, the robot is manipulated to reach the specified working position, and the target object and related operation methods are selected. The position information and operation signal of the selected object in the environmental image are sent to the controller 6.
[0107] Step 5: Video frame capture: The obtained video stream is frame-captured to obtain a depth image and an environmental image.
[0108] Step 6: Load the YOLOV4 network: The environmental image is loaded into the YOLOV4 network for forward testing to obtain the information of the candidate boxes.
[0109] Step 7: Obtain the three-dimensional coordinates of the object: Determine the corresponding candidate boxes through the position information of the target in the environmental image, and obtain the coordinates of the object in the depth image.
[0110] Step 8: Obtain the pose of the target object: Perform edge detection and image segmentation on the image, and fuse it with the depth image information to obtain the pose of the target object and the three-dimensional coordinates of the edge points.
[0111] Step 9: Obstacle information calculation: Through the depth image and the environmental image, obtain the pose and the center point coordinate information of the obstacles between the robot and the target object and around them.
[0112] Step 10: Chassis movement: If the distance between the robot and the object is less than the threshold, jump to Step 1. Otherwise, combining the distance information between the robot and the object with the obstacle information, the controller 6 generates the chassis movement trajectory of the robot through the chassis path planning algorithm, transmits a signal to the motor driver of the chassis, and performs feedback through the speed measurement module to guide the robot to reach the condition where the distance between the robot and the object is less than the threshold, and then return to Step 5.
[0113] Step 1: Operation preparation: According to the distance between the robot and the object, the dynamic whole-body movement operation algorithm of the robot based on the impedance control method is used to manipulate the robotic arm of the robot close to the whole-body movement algorithm of the object.
[0114] Step 12: Physical constraint condition calculation: According to the physical constraint condition algorithm, it is judged whether dangerous operations will occur when the robot is moving the whole body. If so, return to the previous step; otherwise, jump to the next step.
[0115] Step 13: Information update request: Video frame capture, obtaining the video stream frame capture, and obtaining the depth image and the environmental image.
[0116] Step 14: Load the YOLOV4 network: Load the environmental image into the YOLOV4 network for forward testing to obtain the information of the candidate boxes.
[0117] Step 15: Obtain the three-dimensional coordinates of the object: Determine the corresponding candidate box through the position information of the target in the environmental image, and obtain the coordinates of the object in the depth map.
[0118] Step 16: Obtain the pose of the object: Perform edge detection and image segmentation on the image, and combine it with the depth map information to obtain the pose of the object and the three-dimensional coordinates of the edge points.
[0119] Step 17: Start operation: According to the distance between the robot and the object and the set operation instructions, perform the specified operation on the object through the robotic arm control algorithm based on deep reinforcement learning.
[0120] Step 18: Information feedback: Feed back the 30×4 information matrix measured in real time by the multiple tactile sensors on the robotic arm to the controller 6. If it is determined that the operation is completed, jump to the next step; if modification is required, return to the previous step.
[0121] Step 19: Reset the finger joints of the robotic arm: The robotic arm control algorithm based on deep reinforcement learning resets the finger joints of the robotic arm and waits for the next operation.
[0122] Those skilled in the art know that, in addition to implementing the systems, devices, and their respective modules provided by the present invention in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the systems, devices, and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers 6, and embedded microcontrollers 6, etc., to implement the same program. Therefore, the systems, devices, and their respective modules provided by the present invention can be regarded as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structure within the hardware component; the modules for implementing various functions can also be regarded as either software programs for implementing the method or the structure within the hardware component.
[0123] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific implementation manners, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A robot based on multimodal fusion in a complex restricted environment, characterized in that, It includes a robot body, multiple tactile sensors, a wireless communication module, a human-machine interaction interface, and a controller; A depth camera system is provided at the head of the robot body; Multiple tactile sensors are provided at the manipulator of the robot body; the multiple tactile sensors are used to adjust the operation force and operation position by data feedback when the robot operates on an object; The human-machine interaction interface is used to help the robot specify the target object and issue operation instructions to the robot; The wireless communication module is used for communication between the human-machine interaction interface and the robot; The controller includes a motion control system algorithm; Based on the environmental information obtained by the depth camera system, the communication information between the wireless communication module and the human-machine interaction interface, and the data feedback of the multiple tactile sensors, the controller generates a motion equation of the robot through the motion control system algorithm, and sends signals to the servos and servo motors of each joint to control the coordinated movement of each joint of the robot; The motion control system algorithm includes a dynamic whole-body mobile operation algorithm of the robot based on impedance control, a manipulator control algorithm based on deep reinforcement learning, a target recognition and positioning algorithm, and a multi-modal information processing and classification algorithm; The dynamic whole-body mobile operation algorithm of the robot based on impedance control enables the robot to utilize its own structure and environment while performing tasks, and autonomously handle physical limitations and collision avoidance problems; the input is the coordinate position of the object and the coordinates of each current joint, the steering angle and rotation angle of the driving wheels, and the output is the coordinates of the target shoulder joint, the steering angle and rotation angle of the driving wheels. The n driving degrees of freedom are grouped according to the subsystems and control interfaces, and the dynamic equation is: The vector q ∈ R t represents the joint coordinates of the robotic manipulator, and w ∈ R s includes the steering angle and the rotation angle of the wheels of the robot's drive wheels; t represents the drive degrees of freedom of the upper part of the robot; s represents the drive degrees of freedom of the mobile chassis; where n = s + t; g q (q) represents the gravitational torque that appears at the upper body joints, g b (w) is the reference value, and the centripetal effect is represented by τ ext represents the external torque and force, and the control input is τ w and τ q , M ww , M wq, M qq represents the inertial element, and add collision avoidance related to preventing physical collisions to the main task command; The target recognition and positioning algorithm is based on the YOLOV4 network. The input is the video stream image of the depth camera system, and the output is a tensor composed of the center coordinates, width and height values of the prediction box, the confidence of the prediction box, and eighty class scores; the obtained prediction box is used as the input to the HED edge detection network to output the coordinates of the object edge points in the image; the HED edge detection network outputs the output of the last convolutional layer of each layer of the five groups of convolutional feature extraction networks, and combines them through transposed convolution; finally, the coordinates output by the HED edge detection network are combined with the depth map to output the three-dimensional coordinates of the edge points of the target object, and the pose of the object is obtained; The multi-modal information processing and classification algorithm: After removing the information with large errors, dimension processing, and standardization of the output signals of the multiple tactile sensor array, an N×4 signal matrix detected by 4 types of multiple tactile sensors about the object is generated, where N represents the number of multiple tactile sensors in the array. The 4 types of multiple tactile sensor detection signals are: thermal conductivity, contact pressure, object temperature, and ambient temperature; and they are imported into the LSTM neural network according to the time series signal for training and testing to judge the type information of the grasped object.
2. The robot based on multi-modal fusion in a complex restricted environment according to claim 1, wherein The robot includes a driving structure and an operating mechanism; The driving structure includes a mobile chassis for realizing the position movement of the robot; The operating mechanism includes a manipulator configuration, and the manipulator configuration includes at least a shoulder joint, an elbow joint, a hand joint, a wrist joint, and a finger joint; When the driving structure drives the robot to move to a specified working location, the operating mechanism completes the coordinated movement of each joint of the robot under the controller.
3. The robot based on multi-modal fusion in a complex restricted environment according to claim 1, characterized in that The multiple tactile sensors have at least four types of tactile sensing, namely contact pressure, thermal conductivity, object temperature, and ambient temperature; The robot installs an array of N multiple tactile sensors on each manipulator, where N is not less than 30, and at least 5 multiple tactile sensors are installed on each finger joint; The output of the multiple tactile sensor array is an N×m matrix; Wherein, N is the number of multiple tactile sensors in the array, m is the number of types that the sensor can detect, and m is not less than 4.
4. The robot based on multi-modal fusion in a complex restricted environment according to claim 1, characterized in that, The robotic arm control algorithm based on deep reinforcement learning: After completing the initialization of the algorithm weights and loading the algorithm training weights, the robotic arm makes decisions based on the joint coordinate positions returned by the sensors, the robotic arm joint angle states, and the target object coordinates, outputs the predicted control quantities of each joint angle of the robotic arm, and makes the next decision based on the environmental state after the robotic arm executes the movement; Through continuous trial-and-error learning, the network parameters approach the direction that enables the robotic arm to learn more quickly and accurately approach the target until the robotic arm can accurately output a strategy to approach the target object according to the environmental state or the reward fluctuation tends to be stable, and then the training is terminated and the result is output.
5. The robot based on multi-modal fusion in a complex restricted environment according to claim 1, wherein The LSTM neural network includes an input gate, a forget gate, an output gate, a candidate memory cell, and a memory cell component.
6. A running method of a robot based on multi-modal fusion in a complex restricted environment, characterized in that, Using the robot based on multimodal fusion in a complex restricted environment according to any one of claims 1-5, further comprising: Step 1: Robot startup: Click the startup button, and the intelligence of the robot system starts the device through the control button, so that the intelligent control system on the entire device is powered on and started, and all electrical components of the controller, the visual sensor of the depth camera system, the multiple tactile sensors, the drive wheels of the mobile chassis, the motor control drive device, and the remote communication module are in a working standby state; Step 2: Robot initialization: All electrical components on the entire device execute and complete the initialization according to the system set parameters, and each joint of the robot is reset to the initial working set state according to the system set; Step 3: Robotic arm reset: Determine whether each joint is reset. If not, the robot continues to execute the initialization command in Step 2. If it has been reset, then execute the next step; Step 4: The robot starts to work: The robot arrives at the specified position, and through the video stream image transmitted by the depth camera system and the human-machine interaction interface, controls the robot to reach the specified working position, selects the target object and the relevant operation mode, and sends the position information and operation signal of the selected object in the environmental image to the controller; Step 5: Video frame capture: Capture the video stream to obtain a depth image and an environmental image; Step 6: Load the YOLOV4 network: Load the environmental image into the YOLOV4 network for forward testing to obtain the information of the candidate boxes; Step 7: Obtain the three-dimensional coordinates of the object: Determine the corresponding candidate box through the position information of the target in the environmental image, and obtain the coordinates of the object in the depth image; Step 8: Obtain the pose of the target object: Perform edge detection and image segmentation on the image, and fuse it with the depth image information to obtain the pose of the target object and the three-dimensional coordinates of the edge points; Step 9: Calculate obstacle information: Obtain the pose and center point coordinate information of the obstacles between the robot and the target object and around them through the depth image and the environmental image; Step 10: Chassis movement: If the distance between the robot and the object is less than the threshold, jump to Step 11. Otherwise, combine the distance information between the robot and the object with the obstacle information. The controller generates the chassis movement trajectory of the robot through the chassis path planning algorithm, transmits a signal to the motor driver of the chassis, and performs feedback through the speed measurement module to guide the robot to meet the condition that the distance between the robot and the object is less than the threshold, and then return to Step 5; Step 11: Operation preparation: According to the distance between the robot and the object, use the dynamic whole-body movement operation algorithm of the robot based on impedance control method to manipulate the robot's manipulator to approach the object's whole-body movement algorithm; Step 12: Calculate physical constraint conditions: According to the physical constraint condition algorithm, judge whether there will be dangerous operations when the robot is moving the whole body. If so, return to the previous step, otherwise jump to the next step; Step 13: Information update request: Video frame capture, capture the video stream frame to obtain the depth image and the environmental image; Step 14: Load the YOLOV4 network: Load the environmental image into the YOLOV4 network for forward testing to obtain the information of the candidate boxes; Step 15: Obtain the three-dimensional coordinates of the object: Determine the corresponding candidate box through the position information of the target in the environmental image, and obtain the coordinates of the object in the depth map; Step 16: Obtain the pose of the object: Perform edge detection and image segmentation on the image, and combine it with the depth map information to obtain the pose of the object and the three-dimensional coordinates of the edge points; Step 17: Start operation: According to the distance between the robot and the object and the set operation instructions, perform the specified operation on the object through the manipulator control algorithm based on deep reinforcement learning; Step 18: Information feedback: Feed back the 4×30 information matrix measured by the multiple tactile sensors on the manipulator to the controller in real time. If it is determined that the operation is completed, jump to the next step. If modification is required, return to the previous step; Step 19: Reset the finger joints of the manipulator: Use the manipulator control algorithm based on deep reinforcement learning to reset the finger joints of the manipulator and wait for the next operation.
Citation Information
Patent Citations
Material carrying and moving composite robot
CN109202885A
Grabbed object recognition method based on tactile vibration signal and visual image fusion
CN112388655A
Mechanical arm system oriented to 3C assembly scene and based on computer vision and machine learning
CN113119073A
Cited By
Variable height and foot wheel conversion humanoid robot mechanism and information processing control method
CN118907260B
Dual-arm humanoid robot based on physical artificial intelligence and method for controlling the same
KR103019927B1