Robot system

By introducing the transformation matrix of R-CNN and the visual coordinate system and the robot arm coordinate system into the robot system, the precise positioning and efficient operation of elevator buttons are achieved, and the accuracy and efficiency problems of traditional robot systems in the detection and operation of elevator buttons are solved, and the automation level and reliability are improved.

CN120503261AActive Publication Date: 2025-08-19GUANGZHOU ON BRIGHT ELECTRONICS
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510677420.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-08-19
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

Traditional robot systems have problems such as insufficient accuracy, inefficiency and excessive human intervention in elevator button detection and operation.

Method used

Region-based convolutional neural network (R-CNN) is used to combine vision sensors and robotic arms to achieve accurate positioning and state recognition of elevator buttons through the transformation matrix of visual coordinate system and robotic arm coordinate system, and drive the robotic arm to perform operations through the motor driver.

Benefits of technology

It improves the automation level and operation efficiency of the robot system, reduces manual intervention, and improves the reliability and accuracy of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120503261A_ABST
    Figure CN120503261A_ABST
Patent Text Reader

Abstract

There is provided a robotic system comprising a visual sensor, a master computer, a motor driver, and a robotic arm wherein: the visual sensor is configured to acquire a scene image of a current scene, the current scene comprising an elevator button; the main control computer is configured to acquire the button state of the elevator button and the space coordinate of the elevator button in a visual coordinate system based on the visual sensor by utilizing a region-based convolutional neural network based on a scene image of a current scene; based on the button state of a target elevator button on which the mechanical arm is to execute the operation in the elevator buttons and the space coordinate of the target elevator button in the visual coordinate system, a motion instruction used for controlling the mechanical arm to execute the operation on the target elevator button is generated; and the motor driver is configured to drive the mechanical arm to move based on the motion instruction, so that the mechanical arm executes operation on the target elevator button.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the field of automatic control, and more particularly to a robotic system. Background Art

[0002] Typically, a robot system includes a visual sensor, a robotic arm (including the robotic arm body and end effector), and a main control computer. Due to its operational flexibility, it has been widely used in industrial assembly, safety and explosion protection and other fields. Summary of the Invention

[0003] A robotic system according to an embodiment of the present invention includes a visual sensor, a main control computer, a motor driver, and a robotic arm, wherein: the visual sensor is configured to capture a scene image of a current scene, wherein the current scene includes an elevator button; the main control computer is configured to, based on the scene image of the current scene, use a region-based convolutional neural network to obtain the button state of the elevator button and the spatial coordinates of the elevator button in a visual coordinate system based on the visual sensor; and, based on the button state of a target elevator button among the elevator buttons on which the robotic arm is to perform an operation and the spatial coordinates of the target elevator button in the visual coordinate system, generate a motion instruction for controlling the robotic arm to perform an operation on the target elevator button; and the motor driver is configured to drive the robotic arm to move based on the motion instruction so that the robotic arm performs the operation on the target elevator button. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] The present invention can be better understood from the following description of specific embodiments of the present invention in conjunction with the accompanying drawings, in which:

[0005] Figure 1 is a block diagram illustrating functional modules of a robot system according to an embodiment of the present invention.

[0006] Figure 2 It shows Figure 1 The flowchart of the example control process of the robot arm by the host computer is shown.

[0007] Figure 3 It shows Figure 1 Flowchart showing an example recognition process of elevator buttons by a master computer.

[0008] Figure 4 is a flow chart illustrating an example calibration process for determining a transformation matrix between a vision coordinate system and a robotic arm coordinate system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0009] The features and exemplary embodiments of various aspects of the present invention will be described in detail below. In the detailed description below, many specific details are proposed in order to provide a comprehensive understanding of the present invention. However, it will be apparent to those skilled in the art that the present invention can be implemented without the need for some of these specific details. The following description of the embodiments is merely intended to provide a better understanding of the present invention by illustrating examples of the present invention. The present invention is by no means limited to any specific configuration and algorithm proposed below, but covers any modification, replacement and improvement of elements, components and algorithms without departing from the spirit of the present invention. In the accompanying drawings and the following description, well-known structures and technologies are not shown in order to avoid unnecessary ambiguity in the present invention. In addition, it should be noted that the term "A is connected to B" used herein can mean "A is directly connected to B" or "A is indirectly connected to B via one or more other elements."

[0010] Considering the problems of insufficient precision, low efficiency, and excessive human intervention in traditional robotic systems, a robotic system according to an embodiment of the present invention is proposed. By incorporating a region-based convolutional neural network (R-CNN) into robotic arm control, this system can achieve precise positioning, state recognition, and efficient pressing of elevator buttons, thereby improving the robotic system's automation level, operational efficiency, and system reliability.

[0011] Figure 1 FIG. 1 is a block diagram showing functional modules of a robot system according to an embodiment of the present invention. Figure 1 As shown, the robotic system 100 includes a vision sensor 102, a host computer 104, a motor driver 106, and a robotic arm 108, wherein: the vision sensor 102 is configured to capture a scene image of a current scene, wherein the current scene includes an elevator button; the host computer 104 is configured to use R-CNN to obtain, based on the scene image of the current scene, the button state (e.g., pressed state or unpressed state) of the elevator button and the spatial coordinates of the elevator button in a visual coordinate system based on the vision sensor 102; and based on the button state of a target elevator button among the elevator buttons on which the robotic arm 108 is to perform an operation and the spatial coordinates of the target elevator button in the visual coordinate system, generate a motion instruction for controlling the robotic arm 108 to perform an operation on the target elevator button; and the motor driver 106 is configured to drive the robotic arm to move based on the motion instruction so that the robotic arm performs the operation on the target elevator button.

[0012] In some embodiments, the master computer 104 is further configured to generate a motion decision regarding the movement of the robotic arm 108 to the target elevator button and the execution of an operation on the target elevator button, and a motion plan regarding the movement pattern of the robotic arm 108 to the target elevator button based on the button state of the target elevator button and the spatial coordinates of the target elevator button in the visual coordinate system, and generate a motion instruction based on at least one of the motion decision and the motion plan.

[0013] In some embodiments, the robotic arm 108 includes a robotic arm body and an end effector. The robotic arm body is configured to move the end effector to a target elevator button, and the end effector is configured to perform an operation on the target elevator button, such as a pressing operation. In this case, the main control computer 104 is further configured to generate spatial coordinates of the target elevator button in the robotic arm coordinate system based on a transformation matrix between the visual coordinate system and the robotic arm-based robotic arm coordinate system, as well as the spatial coordinates of the target elevator button in the visual coordinate system. Furthermore, based on the spatial coordinates of the target elevator button in the robotic arm coordinate system and the current position of the end effector in the robotic arm coordinate system, the main control computer 104 is configured to calculate joint angles of one or more joints of the robotic arm body using at least one of a forward kinematics algorithm and an inverse kinematics algorithm. The joint angles of the one or more joints of the robotic arm body are included in the motion command.

[0014] In some embodiments, the main control computer 104 is further configured to adjust the joint angles of one or more joints of the manipulator body of the manipulator 108 using a proportional, integral, differential (PID) control algorithm based on the distance between the end effector of the manipulator 108 and the target elevator button while the motor driver 106 drives the manipulator 108 to move.

[0015] Figure 2 It shows Figure 1 The flowchart of the example control process of the robot arm by the main control computer is shown in FIG. Figure 2As shown, in some embodiments, the control process of the main control computer 104 over the manipulator 108 includes: S202, generating the spatial coordinates of the target elevator button in the manipulator coordinate system based on the transformation matrix between the visual coordinate system and the manipulator coordinate system and the spatial coordinates of the target elevator button in the visual coordinate system; S204, calculating the joint angles of one or more joints of the manipulator body of the manipulator 108 by at least one of a forward kinematics algorithm and an inverse kinematics algorithm based on the spatial coordinates of the target elevator button in the manipulator coordinate system and the current position of the end effector of the manipulator 108 in the manipulator coordinate system; S206, controlling the movement of the manipulator body of the manipulator 108 based on the calculated joint angles; S208, adjusting the joint angles of one or more joints of the manipulator body of the manipulator 108 by a PID adjustment algorithm based on the distance between the end effector of the manipulator 108 and the target elevator button; and S210, controlling the end effector of the manipulator 108 to perform an operation on the target elevator button.

[0016] In some embodiments, the main control computer 104 is further configured to: utilize an R-CNN backbone network to extract multiple feature maps of different sizes from a scene image of the current scene; utilize an R-CNN region proposal network (RPN) to find a candidate button region on the smallest feature map among the multiple feature maps, and predict a bounding box and object score of the candidate button region, wherein the object score of each candidate button region is related to the probability that the candidate button region corresponds to a real elevator button; utilize an R-CNN context-aware feature enhancement (CAFE) module to extract fixed-size feature blocks from the multiple feature maps, and utilize a self-attention mechanism to perform feature fusion on the feature blocks, the candidate button regions, and the bounding boxes and object scores of the candidate button regions to generate a feature-enhanced feature block; utilize an R-CNN detection and state classification module to identify the button state of the elevator button and the spatial coordinates of the elevator button's bounding box in the visual coordinate system based on the feature-enhanced feature block.

[0017] In some embodiments, the backbone network of R-CNN can adopt EfficientNet-B0 to extract multiple feature maps {C1, C2, C3, C4} of different sizes from the scene image of the current scene. RPN can be used to find candidate button regions on feature map C4. For example, 200 candidate button regions can be found based on non-maximum suppression (NMS), and the bounding boxes and object scores of the candidate button regions can be predicted. The CAFE module can be used to extract fixed-size feature blocks from the feature maps {C1, C2, C3, C4} through the Region or Interest Align (ROIAlign) operation, and the self-attention mechanism is used to perform feature fusion on the feature blocks, candidate button regions, and the bounding boxes and object scores of the candidate button regions to enhance the robustness of elevator button positioning and state recognition.

[0018] In some embodiments, the detection and state classification module of R-CNN can adopt a lightweight model based on depthwise separable convolution. For example, the lightweight model based on depthwise separable convolution can include a 3x3 depthwise convolution layer, a 1x1 convolution layer based on a softmax activation function, and a 1x1 convolution layer based on a ReLU activation function, wherein the feature map generated by the 3x3 depthwise convolution layer processing the feature enhancement feature block is used for the following two processing branches: a button detection branch, wherein the 1x1 convolution layer based on a softmax activation function processes the feature map generated by the 3x3 depthwise convolution layer to recognize the elevator button, and the 1x1 convolution layer based on a ReLU activation function processes the feature map generated by the 1x1 convolution layer based on a softmax activation function to recognize the spatial coordinates of the elevator button's bounding box in the visual coordinate system; and a state classification branch, wherein the 1x1 convolution layer based on a softmax activation function processes the feature map generated by the 3x3 depthwise convolution layer to recognize the button state of the elevator button.

[0019] In some embodiments, a multi-task loss function (including RPN loss, button detection loss, and state classification loss) can be used to train the lightweight model based on depthwise separable convolution. The generalization ability of the lightweight model based on depthwise separable convolution can be improved through data augmentation (e.g., random flipping, scaling, color jittering) and transfer learning. The latency of each frame image can be ensured to be less than 60 milliseconds by pruning and quantizing the lightweight model based on depthwise separable convolution.

[0020] Figure 3 It shows Figure 1 The flowchart of the master computer's example recognition process of the elevator button is shown in FIG. Figure 3As shown, in some embodiments, the main control computer 104 recognizes the elevator button, including: S302, using EfficientNet-B0 to extract multiple feature maps {C1, C2, C3, C4} of different sizes from the scene image of the current scene; S304, using RPN to find candidate button regions on the feature map C4, and predict the bounding box and object score of the candidate button region; S306, using the CAFE module to extract fixed-size feature blocks from the multiple feature maps {C1, C2, C3, C4}, and using a self-attention mechanism to perform feature fusion on the feature blocks, candidate button regions, and the bounding box and object score of the candidate button region to generate a feature-enhanced feature block; S308, using the detection and state classification module to identify the button state of the elevator button and the spatial coordinates of the bounding box of the elevator button in the visual coordinate system based on the feature-enhanced feature block.

[0021] In some embodiments, the visual sensor 102 is further configured to capture an image of the calibration plate, and the main control computer 104 is further configured to obtain the spatial coordinates of multiple feature points of the calibration plate in the visual coordinate system based on the image of the calibration plate, and calculate the transformation matrix between the visual coordinate system and the robotic arm coordinate system based on the spatial coordinates of the multiple feature points in the visual coordinate system and the spatial coordinates of multiple predetermined positions of the end effector of the robotic arm 108 in the robotic arm coordinate system.

[0022] Figure 4 FIG. 1 is a flow chart illustrating an example calibration process for determining a transformation matrix between a visual coordinate system and a robotic arm coordinate system according to an embodiment of the present invention. Figure 4 As shown, the calibration process for determining the transformation matrix between the visual coordinate system and the manipulator coordinate system includes: S402, using the visual sensor 102 to capture an image of a calibration plate; S404, using the main control computer 104 to obtain the spatial coordinates of multiple feature points of the calibration plate in the visual coordinate system based on the image of the calibration plate; S406, combining the spatial coordinates of the multiple feature points of the calibration plate in the visual coordinate system with the spatial coordinates of multiple predetermined positions of the end effector of the manipulator 108 in the manipulator coordinate system, and calculating the transformation matrix between the visual coordinate system and the manipulator coordinate system using the least squares method; and S408, applying the calculated transformation matrix to the robot system 100 to test whether the manipulator 108 can be correctly moved to the corresponding position based on the spatial coordinates of the target elevator button detected by the visual sensor 102, and then adjusting the calibration parameters or adding sampling points based on the error. Repeating the above process can establish the transformation matrix between the visual coordinate system and the manipulator coordinate system.

[0023] In summary, the robotic system according to an embodiment of the present invention combines computer vision and robotic arm control technology to solve the problems of elevator button detection, state recognition, and operation. It is suitable for high-precision and high-efficiency automation scenarios, can reduce manual intervention and operating costs, and provide a practical solution for industrial automation.

[0024] The present invention may be implemented in other specific forms without departing from its spirit and essential characteristics. For example, the algorithms described in the specific embodiments may be modified without departing from the basic spirit of the present invention. Therefore, the present embodiments are to be considered in all respects as illustrative and not restrictive, and the scope of the present invention is defined by the appended claims rather than the foregoing description. All modifications that come within the meaning and scope of the claims and equivalents are intended to be included within the scope of the present invention.

Claims

1. A robotic system comprising a visual sensor, a main control computer, a motor driver, and a robotic arm, wherein: The visual sensor is configured to capture a scene image of a current scene, wherein the current scene includes an elevator button; The main control computer is configured to, based on the scene image of the current scene, use a region-based convolutional neural network to obtain the button state of the elevator button and the spatial coordinates of the elevator button in a visual coordinate system based on the visual sensor, and, based on the button state of a target elevator button among the elevator buttons on which the robotic arm is to perform an operation and the spatial coordinates of the target elevator button in the visual coordinate system, generate a motion instruction for controlling the robotic arm to perform an operation on the target elevator button; and The motor driver is configured to drive the robotic arm to move based on the motion instruction so that the robotic arm performs an operation on the target elevator button.

2. The robot system according to claim 1, wherein: The main control computer is further configured to: Extracting a plurality of feature maps of different sizes from a scene image of the current scene using the backbone network of the region-based convolutional neural network; Using the region proposal network of the region-based convolutional neural network, finding a candidate button region on a feature map with the smallest size among the multiple feature maps, and predicting a bounding box and an object score of the candidate button region, wherein the object score of each candidate button region is related to the probability that the candidate button region corresponds to a real elevator button; Utilizing the context-aware feature enhancement module of the region-based convolutional neural network, extracting fixed-size feature blocks from the multiple feature maps, and using a self-attention mechanism to perform feature fusion on the feature blocks, the candidate button regions, and the bounding boxes and object scores of the candidate button regions to generate feature-enhanced feature blocks; The detection and state classification module of the region-based convolutional neural network is utilized to identify the button state of the elevator button and the spatial coordinates of the bounding box of the elevator button in the visual coordinate system based on the feature enhancement feature block.

3. The robot system according to claim 2, wherein: The backbone network adopts EfficientNet-B0.

4. The robot system according to claim 2, wherein: The main control computer is further configured to utilize the context-aware feature enhancement module to extract the feature blocks through a region of interest alignment operation.

5. The robot system according to claim 2, wherein: The detection and state classification module adopts a lightweight model based on depthwise separable convolution.

6. The robot system according to claim 5, wherein: The lightweight model based on depthwise separable convolution includes a 3x3 depthwise convolution layer, a 1x1 convolution layer based on a softmax activation function, and a 1x1 convolution layer based on a ReLU activation function.

7. The robot system according to claim 1, wherein: The main control computer is further configured to generate, based on the button state of the target elevator button and the spatial coordinates of the target elevator button in the visual coordinate system, a motion decision for the robot arm to move to the target elevator button and perform an operation on the target elevator button, and a motion plan for a motion pattern for the robot arm to move to the target elevator button, and generate the motion instruction based on at least one of the motion decision and the motion plan.

8. The robot system according to claim 7, wherein: The robotic arm includes a robotic arm body and an end effector, and the main control computer is further configured to generate spatial coordinates of the target elevator button in the robotic arm coordinate system based on a transformation matrix between the visual coordinate system and a robotic arm coordinate system based on the robotic arm, and the spatial coordinates of the target elevator button in the visual coordinate system, and calculate joint angles of one or more joints of the robotic arm body using at least one of a forward kinematics algorithm and an inverse kinematics algorithm based on the spatial coordinates of the target elevator button in the robotic arm coordinate system and a current position of the end effector in the robotic arm coordinate system, wherein the joint angles of the one or more joints of the robotic arm body are included in the motion instruction.

9. The robot system according to claim 8, wherein: The main control computer is further configured to adjust the joint angles of one or more joints of the robotic arm body using a proportional, integral, and differential adjustment algorithm based on the distance between the end effector and the target elevator button during the process of the motor driver driving the robotic arm to move.

10. The robot system according to claim 8, wherein: The visual sensor is also configured to capture an image of a calibration plate, and the main control computer is further configured to obtain the spatial coordinates of multiple feature points of the calibration plate in the visual coordinate system based on the image of the calibration plate, and calculate the transformation matrix between the visual coordinate system and the robotic arm coordinate system based on the spatial coordinates of the multiple feature points in the visual coordinate system and the spatial coordinates of multiple predetermined positions of the end effector in the robotic arm coordinate system.

Citation Information

Patent Citations

  • Elevator door position and opening and closing state detection method and device, medium and terminal

    CN113435466A

  • Target detection method and device, equipment and medium

    CN115311542A

  • Robot control method and device and computer readable storage medium

    CN115625703A

  • Robot elevator taking method, taking device, robot and computer program product

    CN119217342A