Safety vision device and safety vision system

The safety vision system, which uses a 3D skeleton estimation model of the human body and robot, estimates the 3D joint points and joint axis angles of the operator and the robot, solving the problem of area sensor limitations and enabling the robot to decelerate or stop without the need for area sensors.

CN116745083BActive Publication Date: 2025-12-16FANUC LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202180084112.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-21
Filing Date
2021-12-14
Publication Date
2025-12-16
Estimated Expiration
2041-12-14

AI Technical Summary

Technical Problem

In existing technologies, area sensors need to be placed near the robot, which restricts the actions and movements of both the operator and the robot. It is impossible to slow down or stop the robot without area sensors when the operator intrudes into the robot's operating area.

Method used

Using a safety vision device, the system estimates the 3D joint data and joint axis angles of the operator and robot by combining the 3D skeleton estimation models of the human body and robot with the 2D images and distance slopes captured by an external camera, calculates their range and determines the overlap, and outputs deceleration or stop instructions.

Benefits of technology

This technology enables the robot to slow down or stop without the need for area sensors when an operator intrudes into the robot's operating area, thus avoiding restrictions on the actions of both the operator and the robot.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116745083B_ABST
    Figure CN116745083B_ABST
Patent Text Reader

Abstract

When an operator intrudes into a movement area of a robot, the robot can be slowed down or stopped without using an area sensor. A safety vision device has: a human 3D skeleton estimation model; a robot 3D skeleton estimation model; an input section that inputs a 2D image of an operator and a robot taken by an external camera, a distance between the camera and the robot, and a slope; an estimation section that inputs the input 2D image, the distance between the camera and the robot, and the slope to the human 3D skeleton estimation model and the robot 3D skeleton estimation model, estimates 3D joint data of 3D coordinate values representing joint positions of the operator, and angles of a plurality of joint axes included in the robot; and a proximity determination section that calculates an area of the operator and an area of the robot based on the 3D joint data and the angles of the plurality of joint axes, and outputs a slowdown or stop instruction for the robot based on an overlap of the area of the operator and the area of the robot.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a safety vision device and a safety vision system. BACKGROUND

[0002] There is known a technique in which, when a worker who is a safety monitoring object is likely to enter an action area of a robot, the action area of the robot is set around the worker, and safety action control, emergency stop control, and the like of the robot are performed when the robot intrudes into the action area. For example, refer to Patent Literature 1.

[0003] PRIOR ART DOCUMENTS

[0004] PATENT LITERATURE

[0005] Patent Literature 1: Japanese Patent Application Publication No. 2004-243427 SUMMARY

[0006] PROBLEMS TO BE SOLVED BY THE INVENTION

[0007] In the related art, a region sensor or the like is used in order to detect intrusion of a worker into an action area of a robot. However, the region sensor needs to be provided near the robot, and thus, there is a limitation on the action and movement of the worker and the robot.

[0008] Therefore, it is desired that, when a worker intrudes into an action area of a robot, the robot is decelerated or stopped without using a region sensor.

[0009] MEANS FOR SOLVING THE PROBLEMS

[0010] (1) One embodiment of the safety vision device of the present disclosure has: a human 3D skeleton estimation model that inputs a 2D image of a human and outputs 3D joint data of 3D coordinate values representing joint positions of the human; a robot 3D skeleton estimation model that inputs a 2D image of a robot, and a distance and a slope between a camera that has captured the 2D image of the robot and the robot, and outputs angles of a plurality of joint axes included in the robot; an input unit that inputs a 2D image of an operator and the robot captured by an external camera, and a distance and a slope between the external camera and the robot; an estimation unit that inputs the 2D image input by the input unit, and the distance and the slope between the external camera and the robot to the human 3D skeleton estimation model and the robot 3D skeleton estimation model, estimates 3D joint data of 3D coordinate values representing joint positions of the operator, and the angles of the plurality of joint axes included in the robot; and a proximity determination unit that calculates a region representing a range of the operator and a region representing a range of the robot based on the 3D joint data and the angles of the plurality of joint axes, and outputs a deceleration or a stop instruction for the robot based on an overlap of the calculated region of the operator and the region of the robot.

[0011] (2) One embodiment of the safety vision system of the present disclosure has: a robot, a camera, and the safety vision device of (1).

[0012] Effects of Invention

[0013] According to one embodiment, the robot can be caused to decelerate or stop without using a region sensor when the operator intrudes into a movement region of the robot. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 is a functional block diagram that represents a functional configuration example of a safety vision system of one embodiment.

[0015] Figure 2 is a diagram that represents one example of a relationship between a 2D skeleton estimation model and a joint angle estimation model as a robot 3D skeleton estimation model.

[0016] Figure 3 is a diagram that represents one example of joint positions of an operator estimated by an estimation unit.

[0017] Figure 4 is a diagram that represents one example of a region of a robot.

[0018] Figure 5 is a flowchart that explains a determination process of a safety vision device.

[0019] Figure 6is a diagram showing an example of a configuration of a safety vision system.

[0020] Figure 7 is a functional block diagram showing an example of a functional configuration of a machine learning device.

[0021] Figure 8A is a diagram showing an example of a frame image in which the angle of the joint axis J4 is 90 degrees.

[0022] Figure 8B is a diagram showing an example of a frame image in which the angle of the joint axis J4 is -90 degrees.

[0023] Figure 9 is a diagram showing an example for increasing the number of training data.

[0024] Figure 10 is a diagram showing an example of coordinate values of joint axes in standardized XY coordinates.

[0025] Figure 11 is a diagram showing an example of feature mapping of joint axes of a robot.

[0026] Figure 12 is a diagram showing an example of a comparison of a frame image and an output result of a 2-dimensional skeleton estimation model.

[0027] Figure 13 is a diagram showing an example of a joint angle estimation model. DETAILED DESCRIPTION

[0028] Hereinafter, an embodiment of the present disclosure will be described using the drawings.

[0029] <Embodiment>

[0030] Figure 1 is a functional block diagram showing an example of a functional configuration of a safety vision system of an embodiment. As shown in Figure 1 , the safety vision system 1 has a robot 10, a safety vision device 20, and a camera 40.

[0031] The robot 10, the safety vision device 20, and the camera 40, which is an external camera, can be connected to each other via a network not shown such as a wireless LAN (Local Area Network), Wi-Fi (registered trademark), and a mobile phone network complying with standards such as 4G, 5G, and the like. In this case, the robot 10, the safety vision device 20, and the camera 40 have a communication section not shown for communicating with each other by the connection. Further, the robot 10 and the safety vision device 20 can perform data transmission and reception via a communication section not shown, or can perform data transmission and reception via a robot control device (not shown) that controls the operation of the robot 10.

[0032] <Robot 10>

[0033] To those skilled in the art, robot 10 is, for example, a well-known industrial robot, which drives servo motors (not shown) on each of the plurality of joint axes (not shown) included in robot 10 according to drive instructions from a robot control device (not shown), thereby driving the movable parts (not shown) of robot 10.

[0034] Furthermore, robot 10 will be described below as a 6-axis vertical multi-joint robot with 6 joint axes J1 to J6, but it can also be a vertical multi-joint robot other than 6 axes, or a horizontal multi-joint robot, a parallel link robot, etc.

[0035] <Camera 40>

[0036] The camera 40, serving as an external camera, is, for example, a digital camera, and is fixedly mounted on a wall or pillar in a factory where the robot 10 is located, so as to capture images of the robot 10 and the operator U, who is a user carrying a safety vision device 20 described later. Alternatively, the camera 40 can also be a camera mounted on a smartphone, tablet, augmented reality (AR) glasses, mixed reality (MR) glasses, or the like.

[0037] Camera 40 captures images of robot 10 and operator U at a specified frame rate (e.g., 30 frames per second), generating a 2D image, or frame image, projected onto a plane perpendicular to the optical axis of camera 40. Camera 40 outputs the generated frame image to safety vision device 20. Furthermore, the frame image generated by camera 40 can be a visible light image such as an RGB color image or a grayscale image.

[0038] Furthermore, the robot coordinate system of robot 10 and the camera coordinate system of camera 40 are aligned in the world coordinate system through pre-calibration.

[0039] <Safety Vision Device 20>

[0040] Safety vision devices 20 include, for example, smartphones, tablets, augmented reality (AR) glasses, mixed reality (MR) glasses, etc.

[0041] like Figure 1 As shown, the safety vision device 20 includes a control unit 21, a communication unit 22, and a storage unit 23. Furthermore, the control unit 21 includes a 3D object recognition unit 211, a self-position estimation unit 212, an input unit 213, an estimation unit 214, a proximity determination unit 215, and a notification unit 216.

[0042] The communication unit 22 is a communication control device that transmits and receives data with networks such as wireless LAN (Local Area Network), Wi-Fi (registered trademark), and mobile phone networks conforming to standards such as 4G and 5G. For example, the communication unit 22 can communicate directly with the camera 40, or it can communicate with the robot 10 via a robot control device (not shown) that controls the movement of the robot 10.

[0043] The storage unit 23 may be, for example, ROM (Read Only Memory) or HDD (Hard Disk Drive), and stores system programs and safety vision applications executed by the control unit 21 (described later). Additionally, the storage unit 23 may also store the human 3D skeleton estimation model 231 (described later), the robot 3D skeleton estimation model 232 composed of the 2D skeleton estimation model 2321 (described later) and the joint angle estimation model 2322 (described later), and 3D recognition model data 233.

[0044] <Human 3D Skeletal Estimation Model 231>

[0045] The 3D human skeleton estimation model 231 is, for example, a learned model generated by performing supervised learning using training data through a machine learning device (not shown) based on a well-known method of 3D Pose Estimation (e.g., https: / / engineer.dena.com / posts / 2019.12 / cv-papers-19-3d-human-pose-estimation / ). The training data consists of input data of frame images of dynamic images of arbitrary human figures obtained from datasets such as Human3.6M (http: / / vision.imar.ro / human3.6m / description.php), and label data of 3D joint data representing the 3D coordinate values ​​of the joint positions of arbitrary human figures pre-annotated in the frame image.

[0046] <Robot 3D Skeletal Estimation Model 232>

[0047] like Figure 2 As shown, the robot's 3D skeleton estimation model 232 consists of a 2D skeleton estimation model 2321 and a joint angle estimation model 2322.

[0048] Figure 2 This is a diagram illustrating an example of the relationship between the 2D skeleton estimation model 2321 and the joint angle estimation model 2322, which serves as the 3D skeleton estimation model 232 of the robot.

[0049] like Figure 2As shown, the 2-dimensional skeleton estimation model 2321 is a convolutional neural network (CNN) that inputs frame images of the robot 10 captured by the camera 40 and outputs a 2-dimensional pose indicating pixel coordinates of the center positions of the joint axes J1 to J6 of the robot 10 in the frame images.

[0050] The 2-dimensional skeleton estimation model 2321 is generated by, for example, a machine learning device not shown, using training data, by performing supervised learning according to a deep learning model used by a publicly known markerless animal tracking tool (e.g., DeepLabCut) or the like, wherein the training data is input data of frame images of the robot 10 in various poses captured by the camera 40, and label data indicating values of 2-dimensional coordinates (pixel coordinates) of the center positions of the joint axes J1 to J6 in the frame images at the time of capturing each frame image.

[0051] On the other hand, the joint angle estimation model 2322 is a neural network or the like that inputs a distance and a slope between the camera 40 and the robot 10, and a 2-dimensional pose of pixel coordinates output from the 2-dimensional skeleton estimation model 2321, wherein the 2-dimensional pose of pixel coordinates indicates the center positions of the joint axes J1 to J6 of the robot 10 normalized with respect to the width and height of the frame image with the joint axis J1 as the origin, and outputs angles of the joint axes J1 to J6 of the robot 10, the joint axis J1 being a basic link of the robot 10.

[0052] The joint angle estimation model 2322 is generated by, for example, a machine learning device not shown, using training data, by performing supervised learning, wherein the training data is input data of a distance and a slope between the camera 40 and the robot 10, and a 2-dimensional pose indicating the center positions of the joint axes J1 to J6 normalized, and label data of the angles of the joint axes J1 to J6 of the robot 10 at the time of capturing the frame image.

[0053] Further, details of the machine learning device for generating the robot 3-dimensional skeleton estimation model 232 (the 2-dimensional skeleton estimation model 2321 and the joint angle estimation model 2322) will be described later.

[0054] <3-dimensional recognition model data 233>

[0055] The 3-dimensional recognition model data 233 stores, for example, in advance, the attitude and the direction change of the robot 10, and stores as a 3-dimensional recognition model the feature amount such as the edge amount extracted from each of the plurality of frame images of the robot 10 captured by the camera 40. In addition, the 3-dimensional recognition model data 233 can store, in association with each 3-dimensional recognition model, the 3-dimensional coordinate value of the origin of the robot coordinate system of the robot 10 in the world coordinate system at the time of capturing the frame image of each 3-dimensional recognition model, and the information indicating the direction of each of the X-axis, the Y-axis, and the Z-axis of the robot coordinate system in the world coordinate system.

[0056] <Control section 21>

[0057] The control section 21 has a CPU (Central Processing Unit), a ROM, a RAM, a CMOS (Complementary Metal-Oxide-Semiconductor) memory, and the like, which are configured to be able to communicate with each other via a bus, and are well known to those skilled in the art.

[0058] The CPU is a processor that controls the entire safety vision device 20. The CPU reads out the system program and the safety vision application program stored in the ROM via the bus, and controls the entire safety vision device 20 in accordance with the system program and the safety vision application program. Thus, as shown in FIG. 1, the control section 21 realizes the functions of the 3-dimensional object recognition section 211, the own position estimation section 212, the input section 213, the estimation section 214, the approach determination section 215, and the notification section 216. Various data such as temporary calculation data and display data are stored in the RAM. In addition, the CMOS memory is configured as a nonvolatile memory that is backed up by a battery not shown, and maintains the storage state even if the power of the safety vision device 20 is turned off. Figure 1

[0059] <3-dimensional object recognition section 211>

[0060] The 3-dimensional object recognition section 211 acquires, for example, via the communication section 22, the frame image of the robot 10 captured by the camera 40. The 3-dimensional object recognition section 211 extracts, for example, the feature amount such as the edge amount from the frame image of the robot 10 captured by the camera 40, using a well-known method of 3-dimensional coordinate recognition of a robot. In this regard, as a well-known method, for example, refer to "https: / / linx.jp / product / mvtec / halcon / feature / 3d_vision.html".

[0061] ​The 3-dimensional object recognition unit 211 matches the extracted feature amount with the feature amount of the 3-dimensional recognition model stored in the 3-dimensional recognition model data 233. The 3-dimensional object recognition unit 211 obtains, for example, the 3-dimensional coordinate value of the robot origin in the world coordinate system and the information indicating the directions of the X axis, the Y axis, and the Z axis of the robot coordinate system in the 3-dimensional recognition model having the highest agreement degree, based on the matching result.

[0062] <Own position estimation unit 212>

[0063] The own position estimation unit 212 obtains the 3-dimensional coordinate value of the origin of the camera coordinate system of the camera 40 in the world coordinate system (hereinafter, also referred to as "3-dimensional coordinate value of the camera 40") using, for example, a publicly known method of own position estimation. The information obtaining unit 301 calculates the distance and the slope between the camera 40 and the robot 10 based on the obtained 3-dimensional coordinate value of the camera 40 and the obtained 3-dimensional coordinate value of the robot origin.

[0064] Further, the robot 10 and the camera 40 are fixedly arranged in the factory, and thus the own position estimation unit 212 can calculate the 3-dimensional coordinate value of the robot origin in the world coordinate system and the 3-dimensional coordinate value of the camera 40, the distance and the slope between the camera 40 and the robot 10 only once at the start of the safety vision application, and store them in the storage unit 23.

[0065] <Input unit 213>

[0066] The input unit 213 inputs the frame image of the worker U and the robot 10 captured by the camera 40 and the distance and the slope between the camera 40 and the robot 10 calculated by the own position estimation unit 212.

[0067] <Estimation unit 214>

[0068] The estimation unit 214 inputs the frame image of the worker U and the robot 10 and the distance and the slope between the camera 40 and the robot 10 input by the input unit 213 to the human 3-dimensional skeleton estimation model 231 and the robot 3-dimensional skeleton estimation model 232.

[0069] Specifically, the estimation unit 214 estimates 3-dimensional joint data indicating the 3-dimensional coordinate value of the joint position of the worker U in the input frame image based on the output of the human 3-dimensional skeleton estimation model 231.

[0070] Figure 3 is a diagram indicating an example of the joint position of the worker U estimated by the estimation unit 214.

[0071] As shown in Figure 3 , the joint of the worker U estimated by the estimation unit 214 is indicated, for example, by a black dot.

[0072] Further, the estimation unit 214 estimates the angles of the joint axes J1 to J6 of the robot 10 in the input frame image based on the output of the 3-dimensional robot skeleton estimation model 232.

[0073] Further, as described above, the estimation unit 214 normalizes the pixel coordinates of the center positions of the joint axes J1 to J6 output from the 2-dimensional skeleton estimation model 2321, and inputs them to the joint angle estimation model 2322. Further, the estimation unit 214 can set the confidence c of the 2-dimensional pose output from the 2-dimensional skeleton estimation model 2321 to "1" when the confidence c is 0.5 or more, and set it to "0" when the confidence c is less than 0.5. i is 0.5 or more, and set it to "0" when the confidence c is less than 0.5.

[0074] Further, the estimation unit 214 can use the time series data of the 3-dimensional joint node data of the worker U estimated based on a plurality of frame images that are continuous in the time series before the frame image in which a part of the worker U is occluded by the robot 10 and that are continuous in the time series up to the time when the entire worker U and the entire robot 10 are imaged, to estimate the 3-dimensional joint node data of the worker U in the frame image in which a part of the worker U is occluded by the robot 10, when a part of the worker U is occluded by the robot 10 in the frame image input by the input unit 213.

[0075] Further, the estimation unit 214 can use the time series data of the angles of the joint axes J1 to J6 of the robot 10 estimated based on a plurality of frame images that are continuous in the time series before the frame image in which a part of the robot 10 is occluded by the worker U and that are continuous in the time series up to the time when the entire worker U and the entire robot 10 are imaged, to estimate the angles of the joint axes J1 to J6 of the robot 10 in the frame image in which a part of the robot 10 is occluded by the worker U, when a part of the robot 10 is occluded by the worker U in the frame image input by the input unit 213.

[0076] <Approach determination unit 215>

[0077] The approach determination unit 215 calculates a region representing the range of the worker U and a region representing the range of the robot 10 based on the 3-dimensional joint node data of the worker U and the angles of the joint axes J1 to J6 of the robot 10 estimated by the estimation unit 214, and outputs a deceleration or stop instruction for the robot 10 to a robot control device (not shown) based on the overlap of the calculated region of the worker U and the region of the robot 10.

[0078] Specifically, the approach determination unit 215, for example, arranges the joint nodes of the worker U in a 3-dimensional space of a world coordinate system based on the 3-dimensional joint node data of the worker U estimated by the estimation unit 214, and generates a skeleton of the worker U by connecting the joint nodes with straight lines. The approach determination unit 215 calculates a region representing the range of the worker U by modeling with a cuboid or the like having a length, a depth, and a height that are set in advance in accordance with the generated straight lines of the skeleton.

[0079] In addition, the approach determination section 215 uses a Denavit-Hartenberg (DH) parameter table defined in advance to solve forward kinematics based on the angles of the joint axes J1 to J6 estimated by the estimation section 214, and calculates 3-dimensional coordinate values of the center positions of the joint axes J1 to J6. Then, the approach determination section 215 configures the center positions of the joint axes J1 to J6 of the robot 10 calculated in the 3-dimensional space of the world coordinate system, and generates a skeleton of the robot 10. The approach determination section 215 calculates a region representing the range of the robot 10 by modeling with a cuboid or the like having lengths, depths, and heights set in advance in accordance with the generated skeleton of the robot 10.

[0080] Further, the DH parameter table is prepared in advance, for example, based on the specifications of the robot 10, and stored in the storage section 23.

[0081] Figure 4 Fig. 6 is a diagram showing an example of the region of the robot 10. Further, Figure 4 The region of the robot 10 shown in Fig. 6 is composed of two regions R1 and R2 having different sizes. That is, the lengths, depths, and heights in the region R1 are set to be larger than those in the region R2.

[0082] Then, the approach determination section 215 determines whether the region of the worker U in the world coordinate system calculated overlaps the region R1 or the region R2 of the robot 10. When the region of the worker U overlaps only the region R1 of the robot 10, the approach determination section 215 determines that there is time before the worker U collides with the robot 10, and outputs a deceleration instruction to the robot control device (not shown), thereby decelerating the motion of the robot 10.

[0083] On the other hand, when the region of the worker U also overlaps the region R2 of the robot 10, the approach determination section 215 immediately determines that there is a danger that the worker U collides with the robot 10, and outputs a stop instruction to the robot control device (not shown), thereby stopping the motion of the robot 10.

[0084] In this way, by providing two regions R1 and R2 having different sizes as the region of the robot 10, the approach determination section 215 can appropriately determine which of the deceleration or the stop instruction should be output.

[0085] <Notification Section 216>

[0086] When the approach determination section 215 outputs the deceleration or the stop instruction, the notification section 216 outputs a warning sound via a speaker (not shown) included in the safety vision device 20.

[0087] Further, the notification section 216 can display a message indicating a warning on a display device (not shown) such as an LCD (Liquid Crystal Display) included in the safety vision device 20.

[0088] <Judgment processing of safety vision device 20>

[0089] Next, the operation of the judgment processing of the safety vision device 20 according to the present embodiment will be described.

[0090] Figure 5 is a flowchart of the judgment processing of the safety vision device 20. The flow shown herein is repeatedly executed during the execution of the safety vision application program by the safety vision device 20.

[0091] In step S1, the 3-dimensional object recognition section 211 acquires a frame image of the operator U and the robot 10 taken by the camera 40 at a prescribed frame rate.

[0092] In step S2, the 3-dimensional object recognition section 211 acquires a 3-dimensional coordinate value of the robot origin in the world coordinate system and information indicating the directions of the X axis, the Y axis, and the Z axis of the robot coordinate system, based on the frame image acquired in step S1 and the 3-dimensional recognition model data 233.

[0093] In step S3, the self-position estimation section 212 acquires a 3-dimensional coordinate value of the camera 40 in the world coordinate system based on the frame image acquired in step S1.

[0094] In step S4, the self-position estimation section 212 calculates the distance and the slope between the camera 40 and the robot 10 based on the 3-dimensional coordinate value of the camera 40 acquired in step S3 and the 3-dimensional coordinate value of the robot origin of the robot 10 acquired in step S2.

[0095] In step S5, the input section 213 inputs the frame image acquired in step S1 and the distance and the slope between the camera 40 and the robot 10 calculated in step S4.

[0096] In step S6, the estimation section 214 inputs the frame image input in step S5 to the human 3-dimensional skeleton estimation model 231, and thereby estimates 3-dimensional joint data indicating 3-dimensional coordinate values of joint positions of the operator U in the input frame image. Further, the estimation section 214 inputs the frame image input in step S2 and the distance and the slope between the camera 40 and the robot 10 to the robot 3-dimensional skeleton estimation model 232, and thereby estimates the angles of the joint axes J1 to J6 of the robot 10 at the time of taking the input frame image.

[0097] In step S7, the approach determination section 215 calculates a region indicating the range of the worker U from the 3D joint data of the worker U estimated in step S6. In addition, the approach determination section 215 calculates the regions Rl, R2 of the robot 10 from the angles of the joint axes Jl to J6 of the robot 10 estimated in step S6.

[0098] In step S8, the approach determination section 215 determines whether the region of the worker U calculated in step S7 overlaps the region Rl of the robot 10 calculated in step S7. When the region of the worker U overlaps the region Rl of the robot 10, the process proceeds to step S9. On the other hand, when the region of the worker U does not overlap the region Rl of the robot 10, the safety vision device 20 ends the determination process.

[0099] In step S9, the approach determination section 215 determines whether the region of the worker U calculated in step S7 overlaps the region R2 of the robot 10 calculated in step S7. When the region of the worker U overlaps the region R2 of the robot 10, the process proceeds to step S10. On the other hand, when the region of the worker U does not overlap the region R2 of the robot 10, the process proceeds to step Sll.

[0100] In step S10, the approach determination section 215 outputs a stop instruction to the robot control device (not shown).

[0101] In step Sll, the approach determination section 215 outputs a deceleration instruction to the robot control device (not shown).

[0102] In step S12, the notification section 216 outputs a warning sound via the speaker (not shown) of the safety vision device 20.

[0103] According to the above, the safety vision device 20 of the embodiment inputs the frame images captured of the worker U and the robot 10, the distance and the slope between the camera 40 and the robot 10 to the human 3D skeleton estimation model 231 and the robot 3D skeleton estimation model 232 which are learning completed models, thereby estimating the 3D joint data indicating the 3D coordinate values of the joint positions of the worker U and the angles of the joint axes Jl to J6 of the robot 10. The safety vision device 20 calculates the region indicating the range of the worker U from the estimated 3D joint data, and the regions Rl, R2 indicating the range of the robot 10 from the angles of the joint axes Jl to J6. The safety vision device 20 determines whether the region of the worker U overlaps the region Rl or the region R2 of the robot 10, and outputs a deceleration or a stop instruction to the robot control device (not shown).

[0104] Thus, when the worker intrudes into the action region of the robot 10, the safety vision device 20 can decelerate or stop the robot without using the area sensor.

[0105] The safety vision device 20 is described above, and a machine learning device for generating a 3-dimensional skeleton estimation model 232 of the robot 3 is described.

[0106] Figure 7 is a functional block diagram showing a functional configuration example of the machine learning device 30.

[0107] As shown in Figure 7 , the machine learning device 30 has an information acquisition section 301, a 2-dimensional posture acquisition section 302, an input data acquisition section 303, a label acquisition section 304, a learning section 305, and a storage section 306.

[0108] In the following description, the machine learning device 30 acquires only data acquired at a time when all data can be synchronized as training data. For example, in a case where the camera 40 shoots frame images at 30 frames / second, the period of the angles of the plurality of joint axes included in the robot 10 can be 100 milliseconds, and the machine learning device 30 can acquire training data at a predetermined period of 100 milliseconds or the like at which other data can be synchronized.

[0109] <Storage section 306>

[0110] The storage section 306 is a RAM (Random Access Memory) or the like, and stores input data acquired by the input data acquisition section 303 described later, label data acquired by the label acquisition section 304 described later, and a learning completed model constructed by the learning section 305 described later. In addition, the storage section 306 can also store 3-dimensional recognition model data 3061.

[0111] Further, the 3-dimensional recognition model data 3061 is omitted from description, for example, like the 3-dimensional recognition model data 233 of the safety vision device 20.

[0112] <Information acquisition section 301>

[0113] The information acquisition section 301 acquires frame images of the robot 10 shot by the camera 40, for example, via a communication section not shown. The information acquisition section 301 acquires 3-dimensional coordinate values of the robot origin in the world coordinate system and information indicating the directions of the X-axis, the Y-axis, and the Z-axis of the robot coordinate system from the acquired frame images, for example, like the 3-dimensional object recognition section 211 of the safety vision device 20.

[0114] Further, the information acquisition unit 301 can acquire the 3-dimensional coordinate value of the camera 40 in the world coordinate system, calculate the distance and the slope between the camera 40 and the robot 10 from the acquired 3-dimensional coordinate value of the camera 40 and the acquired 3-dimensional coordinate value of the robot origin, as with the self-position estimation unit 212 of the safety vision device 20.

[0115] <2-dimensional posture acquisition unit 302>

[0116] The 2-dimensional posture acquisition unit 302 transmits a request to the robot 10 at a prescribed period synchronized by the above 100 milliseconds or the like, for example, via a communication unit not shown, acquires the angles of the joint axes J1 to J6 of the robot 10 at the time of capturing the frame image acquired by the information acquisition unit 301.

[0117] Then, the 2-dimensional posture acquisition unit 302 calculates the 3-dimensional coordinate value of the center positions of the joint axes J1 to J6, calculates the 3-dimensional posture of the robot 10 in the world coordinate system from the acquired angles of the joint axes J1 to J6, for example, using a predefined DH parameter table.

[0118] Further, the 2-dimensional posture acquisition unit 302 projects the center positions of the joint axes J1 to J6 of the robot 10 calculated by the forward kinematics in the 3-dimensional space of the world coordinate system from the viewpoint of the camera 40 determined from the distance and the slope between the camera 40 and the robot 10 calculated by the information acquisition unit 301 to the projection plane determined from the distance and the slope between the camera 40 and the robot 10, for example, using a publicly known method of projecting to a 2-dimensional plane, thereby generating 2-dimensional coordinates (pixel coordinates) (x i , y i ) of the center positions of the joint axes J1 to J6 as the 2-dimensional posture of the robot 10. Further, i is an integer of 1 to 6.

[0119] Further, as shown in Figs. 15 and 16, the joint axes can be hidden in the frame image due to the posture of the robot 10 and the direction of photographing. Figure 8A Figure 8B

[0120] Figure 8A Fig. 14 is a diagram showing an example of a frame image in which the angle of the joint axis J4 is 90 degrees. Figure 8B Fig. 15 is a diagram showing an example of a frame image in which the angle of the joint axis J4 is -90 degrees.

[0121] In the frame image of Fig. 14, the joint axis J6 is hidden and does not appear. On the other hand, in the frame image of Fig. 15, the joint axis J6 appears. Figure 8A Figure 8B

[0122] ​​​​Therefore, the 2D attitude acquisition unit 302 connects adjacent joint axes of the robot 10 to each other using line segments, and defines the thickness of each line segment using a pre-set link width of the robot 10. Based on the 3D attitude of the robot 10 calculated through forward kinematics and the optical axis direction of the camera 40 determined by the distance and slope between the camera 40 and the robot 10, the 2D attitude acquisition unit 302 determines whether other joint axes exist on the line segments. The 2D attitude acquisition unit 302 determines whether other joint axes are located in a depth direction opposite to the camera 40 side relative to the line segment. Figure 8A In that case, other joint axes Ji ( Figure 8A The confidence level of the joint axis J6) i Set to "0". On the other hand, the 2D pose acquisition unit 302 is located on the side of the camera 40 relative to the line segment of other joint axes Ji. Figure 8B In that case, other joint axes Ji ( Figure 8B The confidence level of the joint axis J6) i Set to "1".

[0123] That is, the 2D pose acquisition unit 302 can obtain the 2D coordinates (pixel coordinates) (x, y, z) of the center position of the projected joint axes J1 to J6. i y i The confidence level c represents whether each joint axis J1 to J6 is visible in the frame image. i The 2D pose of robot 10 is included.

[0124] In addition, it is expected that multiple training data for supervised learning will be prepared in the machine learning device 30.

[0125] Figure 9 This is a diagram illustrating an example of using increased training data.

[0126] like Figure 9 As shown, the 2D pose acquisition unit 302, for example, randomly assigns a distance and slope between the camera 40 and the robot 10 to increase training data, causing the 3D pose of the robot 10 calculated by forward kinematics to rotate. The 2D pose acquisition unit 302 can generate multiple 2D poses of the robot 10 by projecting the 3D pose of the rotated robot 10 onto a 2D plane determined by the randomly assigned distance and slope.

[0127] <Input Data Acquisition Unit 303>

[0128] The input data acquisition unit 303 acquires, for example via a communication unit not shown, frame images of an arbitrary person dynamic image as input data from a dataset such as Human3.6M in order to generate the human 3-dimensional skeleton estimation model 231 described above. In addition, the input data acquisition unit 303 acquires, as input data, a frame image acquired from the camera 40 and the distance and slope between the camera 40 and the robot 10 at the time of photographing the frame image acquired from the information acquisition unit 301 in order to generate the robot 3-dimensional skeleton estimation model 232. The input data acquisition unit 303 stores the acquired input data in the storage unit 306.

[0129] Further, the input data acquisition unit 303 can convert the 2-dimensional coordinates (pixel coordinates) (x i , y i ) of the center positions of the joint axes J1 to J6 included in the 2-dimensional pose generated by the 2-dimensional pose acquisition unit 302 into values of XY coordinates that are standardized to -1 < X < 1 with the joint axis J1 as the origin of the robot 10 and to -1 < Y < 1 with the width of the frame image and the height of the frame image, respectively, when generating the joint angle estimation model 2322 that constitutes the robot 3-dimensional skeleton estimation model 232, as shown in FIG. 6. Figure 10

[0130] <Label Acquisition Unit 304>

[0131] The label acquisition unit 304 acquires, for example via a communication unit not shown, 3-dimensional joint data that indicates 3-dimensional coordinate values of joint positions in a camera coordinate system that is pre-annotated in each of the frame images of the arbitrary person described above, as label data (correct answer data) from the dataset such as Human3.6M described above in order to generate the human 3-dimensional skeleton estimation model 231 described above. Further, the label acquisition unit 304 can also convert the 3-dimensional joint data from the camera coordinate system to the world coordinate system.

[0132] In addition, the label acquisition unit 304 acquires, as label data (correct answer data), the angles of the joint axes J1 to J6 of the robot 10 at the time of photographing the frame image at a predetermined period of time synchronized with the above-described 100 milliseconds or the like from the 2-dimensional pose acquisition unit 302 and the 2-dimensional pose that indicates the center positions of the joint axes J1 to J6 of the robot 10 in the frame image in order to generate the robot 3-dimensional skeleton estimation model 232 described above. The label acquisition unit 304 stores the acquired label data in the storage unit 306.

[0133] <Learning Unit 305>

[0134] ​The learning unit 305 receives a group of the input data and the label described above as training data. The learning unit 305 performs supervised learning using the received training data, thereby constructing the human 3D skeleton estimation model 231 and the robot 3D skeleton estimation model 232 composed of the 2D skeleton estimation model 2321 and the joint angle estimation model 2322.

[0135] Then, the learning unit 305 provides the safety vision device 20 with the constructed 3D skeleton estimation model 231 and the robot 3D skeleton estimation model 232 composed of the 2D skeleton estimation model 2321 and the joint angle estimation model 2322.

[0136] Hereinafter, the construction of the 2D skeleton estimation model 2321 and the joint angle estimation model 2322 that constitute the robot 3D skeleton estimation model 232 will be described.

[0137] <2D skeleton estimation model 2321>

[0138] The learning unit 305 performs supervised learning using the input data of the frame image of the robot 10 captured by the camera 40 and the training data of the label representing the 2D pose of the center positions of the joint axes J1 to J6 at the time of capturing the frame image, for example, according to the deep learning model used by the known markerless animal tracking tool (for example, DeepLabCut) and the like, and generates the 2D skeleton estimation model 2321 that inputs the frame image of the operator U and the robot 10 captured by the camera 40 and outputs the 2D pose representing the pixel coordinates of the center positions of the joint axes J1 to J6 of the robot 10 in the captured frame image.

[0139] Specifically, the 2D skeleton estimation model 2321 is constructed according to a neural network, that is, a convolutional neural network (CNN: Convolutional Neural Network).

[0140] The convolutional neural network has a configuration of a convolution layer, a pooling layer, a fully connected layer, and an output layer.

[0141] In the convolution layer, a filter of a prescribed parameter is applied to the input frame image in order to perform feature extraction such as edge extraction. The prescribed parameter in the filter corresponds to the weight of the neural network, and learning is performed by repeatedly performing forward propagation and backward propagation.

[0142] In the pooling layer, the image output from the convolution layer is blurred in order to allow the position of the robot 10 to shift. Thus, even if the position of the robot 10 changes, it can be regarded as the same object.

[0143] By combining these convolution layer and the pooling layer, it is possible to extract a feature amount from the frame image.

[0144] In the fully connected layer, the image data from which the feature portion is extracted by the convolution layer and the pooling layer is combined with 1 node, and a value converted by an activation function, that is, a feature map of the certainty is output.

[0145] Figure 11 is a diagram showing an example of the feature map of the joint axes J1 to J6 of the robot 10.

[0146] As shown in Figure 11 , in the feature map of each joint axis J1 to J6, the value of the certainty c i is represented in the range of 0 to 1, and the closer the cell is to the center position of the joint axis, the closer the value is to "1", and the farther the cell is from the center position of the joint axis, the closer the value is to "0".

[0147] In the output layer, with respect to the output from the fully connected layer, the row, the column, and the maximum of the certainty of the cell in which the maximum of the certainty is obtained in the feature map of each joint axis J1 to J6 are output. Further, in the output layer, when the frame image is convolved by 1 / N in the convolution layer, the row and the column of the cell are set to N times, and are set to the pixel coordinates indicating the center position of each joint axis J1 to J6 in the frame image (N is an integer of 1 or more).

[0148] Figure 12 is a diagram showing an example of the comparison of the frame image and the output result of the 2-dimensional skeleton estimation model 2321.

[0149] <Joint Angle Estimation Model 2322>

[0150] The learning unit 305 performs supervised learning using training data of the distance and the slope between the camera 40 and the robot 10, and the input data indicating the 2-dimensional pose of the center positions of the joint axes J1 to J6 normalized as described above, and the label data of the angles of the joint axes J1 to J6 of the robot 10 at the time of capturing the frame image, and generates the joint angle estimation model 2322.

[0151] Further, the learning unit 305 normalizes the 2-dimensional pose of the joint axes J1 to J6 output from the 2-dimensional skeleton estimation model 2321, but the 2-dimensional skeleton estimation model 2321 can be generated in a manner that the normalized 2-dimensional pose is output from the 2-dimensional skeleton estimation model 2321.

[0152] Figure 13 is a diagram showing an example of the joint angle estimation model 2322. Here, as shown in Figure 13As shown, the joint angle estimation model 2322 is exemplified as a multi-layer neural network that takes as an input layer 2-dimensional pose data output from the 2-dimensional skeleton estimation model 2321 and representing positions of the normalized joint axes J1 to J6, and the distance and slope between the camera 40 and the robot 10, and takes as an output layer the angles of the joint axes J1 to J6. Further, the 2-dimensional pose contains the center positions of the normalized joint axes J1 to J6, i.e., coordinates (x i , y i ), and the confidence c i (x i , y i , c i ) that is set to "1" when the confidence output from the 2-dimensional skeleton estimation model 2321 is 0.5 or more, and set to "0" when it is less than 0.5.

[0153] Further, the "slope Rx of the X-axis", the "slope Ry of the Y-axis", and the "slope Rz of the Z-axis" are the rotation angles of the camera 40 with respect to the robot 10 around the X-axis, the Y-axis, and the Z-axis in the world coordinate system, respectively, which are calculated from the 3-dimensional coordinate values of the camera 40 in the world coordinate system and the 3-dimensional coordinate values of the robot origin of the robot 10 in the world coordinate system.

[0154] Further, the learning unit 305 can further perform supervised learning on the robot 3-dimensional skeleton estimation model 232 composed of the 2-dimensional skeleton estimation model 2321 and the joint angle estimation model 2322 when new training data is obtained after the learning completion model composed of the 2-dimensional skeleton estimation model 2321 and the joint angle estimation model 2322 is constructed, and thereby update the robot 3-dimensional skeleton estimation model 232 composed of the 2-dimensional skeleton estimation model 2321 and the joint angle estimation model 2322 that is constructed once.

[0155] Thus, the training data can be automatically obtained from the usual photographing of the robot 10, and thereby the estimation accuracy of the angles of the joint axes J1 to J6 of the robot 10 can be improved on a daily basis.

[0156] The above-described supervised learning can be performed by online learning, batch learning, or mini-batch learning.

[0157] The online learning is a learning method in which supervised learning is performed every time a frame image of the robot 10 is captured and training data is created. The batch learning is a learning method in which a plurality of training data corresponding to repetitions are collected during a period in which frame images of the robot 10 are repeatedly captured and training data are created, and supervised learning is performed using all the collected training data. The mini-batch learning is a learning method in which supervised learning is performed every time training data of a certain degree are accumulated, which is intermediate between the online learning and the batch learning.

[0158] With the above-described machine learning device, the robot 3D skeleton estimation model possessed by the safety vision device 20 can be generated.

[0159] The above describes one embodiment, but the safety vision device 20 is not limited to the above-described embodiment, and includes variations, modifications, and the like within a range capable of achieving the object.

[0160] <Modification Example 1>

[0161] In the above-described embodiment, the safety vision device 20 calculates one region with respect to the worker U, but is not limited thereto. For example, the safety vision device 20 can calculate two regions RU1, RU2 of different sizes with respect to the worker U, like the regions R1, R2 of the robot 10. Further, the length, depth, and height in the region RU1 are set to be larger than those in the region RU2.

[0162] The safety vision device 20 can output, to a robot control device (not shown), an instruction to decelerate the robot 10, for example, when the region RU1 of the worker U overlaps the region R1 of the robot 10, the region RU2 of the worker U overlaps the region R1 of the robot 10, or the region RU1 of the worker U overlaps the region R2 of the robot 10.

[0163] Further, the safety vision device 20 can output, to a robot control device (not shown), an instruction to stop the robot 10 when the smallest region RU2 of the worker U overlaps the region R2 of the robot 10.

[0164] Further, the region of the worker U and the region of the robot 10 are set to one or two, but can be set to three or more regions.

[0165] Thus, the safety vision device 20 can more finely avoid collision between the worker U and the robot 10.

[0166] <Modification Example 2>

[0167] Further, for example, in the above-described embodiment, the safety vision device 20 uses the human 3-dimensional skeleton estimation model 231 and the robot 3-dimensional skeleton estimation model 232 to estimate the 3-dimensional joint data of the worker U and the angles of the joint axes J1 to J6 of the robot 10 from the frame images of the worker U and the robot 10 inputted and the distance and the slope between the camera 40 and the robot 10, but is not limited thereto. For example Figure 6 As shown, the server 50 can also store the human 3-dimensional skeleton estimation model 231 and the robot 3-dimensional skeleton estimation model 232, and share the human 3-dimensional skeleton estimation model 231 and the robot 3-dimensional skeleton estimation model 232 with m safety vision devices 20A(1) to 20A(m) (m is an integer of 2 or more) connected to the server 50 via the network 60. Thus, even if a new robot and a safety vision device are installed, the human 3-dimensional skeleton estimation model 231 and the robot 3-dimensional skeleton estimation model 232 can be applied.

[0168] Further, the robots 10A(1) to 10A(m) respectively correspond to the robots 10 of Figure 1 The safety vision devices 20A(1) to 20A(m) respectively correspond to the safety vision devices 20 of Figure 1

[0169] Further, each function included in the safety vision device 20 in an embodiment can be realized by hardware, software, or a combination thereof, respectively. Here, realization by software means realization by a computer reading in a program and executing.

[0170] Each constituent included in the safety vision device 20 can be realized by hardware including an electronic circuit or the like, software, or a combination thereof. When realized by software, a program constituting the software is installed in a computer. Further, these programs can be distributed to users recorded in a removable medium or can be distributed by being downloaded to a computer of a user via a network. Further, when constituted by hardware, for example, a part or all of the functions of each structural unit included in the above-described device can be constituted by an integrated circuit (IC) such as an ASIC (Application Specific Integrated Circuit), a gate array, a FPGA (Field Programmable Gate Array), a CPLD (Complex Programmable Logic Device), and the like.

[0171] ​The program can be stored and supplied to the computer using various types of non-transitory computer readable media. The non-transitory computer readable media include a variety of types of tangible storage media. Examples of the non-transitory computer readable media include a magnetic recording medium (such as a floppy disk, a tape, and a hard disk drive), an optical magnetic recording medium (such as a magneto-optical disk), a CD-ROM (Read Only Memory), a CD-R, a CD-R / W, and a semiconductor memory (such as a mask ROM, a PROM (Programmable ROM), an EPROM (Erasable PROM), a flash ROM, and a RAM). In addition, the program can be supplied to the computer by various types of transitory computer readable media. Examples of the transitory computer readable media include an electrical signal, an optical signal, and an electromagnetic wave. The transitory computer readable medium can supply the program to the computer via a wired communication path such as an electrical wire, and an optical fiber, or a wireless communication path.

[0172] Further, the steps of the program described in the recording medium include processing performed in a time series in the order, and also include processing not necessarily performed in a time series, and processing performed in parallel or individually.

[0173] In other words, the safety vision device and the safety vision system of the present disclosure can take various embodiments having the following structures.

[0174] (1) The safety vision device 20 of the present disclosure has: a human 3D skeleton estimation model 231 that inputs a 2D image of a human and outputs 3D joint data of 3D coordinate values representing joint positions of the human; a robot 3D skeleton estimation model 232 that inputs a 2D image of a robot 10, and a distance and a slope between a camera 40 that has captured the 2D image of the robot 10 and the robot 10, and outputs angles of a plurality of joint axes J1 to J6 included in the robot 10; an input unit 213 that inputs a 2D image of an operator U and the robot 10 captured by the external camera 40, and a distance and a slope between the external camera 40 and the robot 10; an estimation unit 214 that inputs the 2D image input by the input unit 213, and the distance and the slope between the external camera 40 and the robot 10, to the human 3D skeleton estimation model 231 and the robot 3D skeleton estimation model 232, estimates 3D joint data of 3D coordinate values representing joint positions of the operator U, and angles of the plurality of joint axes J1 to J6 included in the robot 10; and a proximity determination unit 215 that calculates a region representing a range of the operator U and regions R1, R2 representing a range of the robot 10 from the 3D joint data and the angles of the plurality of joint axes J1 to J6, and outputs a deceleration or stop instruction for the robot 10 according to an overlap of the calculated region of the operator U and the regions of the robot 10.

[0175] According to the safety vision device 20, the robot can be decelerated or stopped without using a region sensor when the operator intrudes into the action region of the robot.

[0176] (2) In the safety vision device 20 described in (1), the 2D image can be a frame image captured by the camera 40 at a predetermined frame rate.

[0177] Thus, the safety vision device 20 can continuously track the actions of the operator U and the robot 10.

[0178] (3) In the safety vision device 20 described in (1) or (2), the safety vision device 20 can further have: a notification unit 216 that outputs a warning sound when the proximity determination unit 215 outputs the deceleration or stop instruction.

[0179] Thus, the safety vision device 20 can warn the operator U.

[0180] (4) In the safety vision device 20 described in any one of (1) to (3), the human 3D skeleton estimation model 231 and the robot 3D skeleton estimation model 232 can be provided in a server 50 to which the safety vision device 20 is accessibly connected via a network 60.

[0181] Thus, even if a new robot and the safety vision device are configured, the safety vision device 20 is able to apply the learning completed model.

[0182] (5) The safety vision system 1 of the present disclosure has: a robot 10, a camera 40; and the safety vision device 20 described in any one of (1) to (4).

[0183] The safety vision system 1 is able to obtain the same effects as (1) to (4).

[0184] Explanation of symbols

[0185] 1 Safety vision system

[0186] 10 Robot

[0187] 20 Safety vision device

[0188] 21 Control section

[0189] 211 3-dimensional object recognition section

[0190] 212 Self position estimation section

[0191] 213 Input section

[0192] 214 Estimation section

[0193] 215 Approach determination section

[0194] 216 Notification section

[0195] 22 Communication section

[0196] 23 Storage section

[0197] 231 Human 3-dimensional skeleton estimation model

[0198] 232 Robot 3-dimensional skeleton estimation model

[0199] 2321 2-dimensional skeleton estimation model

[0200] 2322 Joint angle estimation model

[0201] 30 Machine learning device

[0202] 301 Information acquisition section

[0203] 302 2-dimensional pose acquisition section

[0204] 303 Input data acquisition section

[0205] 304 Label acquisition section

[0206] 305 Learning section

[0207] 306 storage section

[0208] 40 camera

[0209] 50 server

[0210] 60 network

Claims

1. A safety vision device, characterized in that, having: a human 3-dimensional skeleton estimation model as a learning completed model, which inputs a 2-dimensional image of a human and outputs 3-dimensional joint data of 3-dimensional coordinate values representing joint position of the human; a 2-dimensional skeleton estimation model as a learning completed model, which inputs a 2-dimensional image of a robot and outputs 2-dimensional pose data representing center position of a joint axis at the time of capturing the 2-dimensional image; a robot 3-dimensional skeleton estimation model as a learning completed model, which inputs a distance and a slope between a camera that captures a 2-dimensional image of the robot and the robot, and 2-dimensional pose data representing center position of the joint axis, and outputs angles of a plurality of joint axes included in the robot; an input unit that inputs a 2-dimensional image of a worker and the robot captured by an external camera, and a distance and a slope between the external camera and the robot; an estimation unit that inputs the 2-dimensional image of the worker input by the input unit to the human 3-dimensional skeleton estimation model, estimates 3-dimensional joint data of 3-dimensional coordinate values representing joint position of the worker, and inputs the 2-dimensional image of the robot input by the input unit to the 2-dimensional skeleton estimation model, and inputs the output 2-dimensional pose data representing center position of the joint axis and the distance and the slope between the external camera and the robot to the robot 3-dimensional skeleton estimation model, and estimates angles of a plurality of joint axes included in the robot; and a proximity determination unit that calculates a region representing a range of the worker and a region representing a range of the robot, and outputs a deceleration or stop instruction for the robot according to an overlap of the calculated region of the worker and the region of the robot, wherein the region representing the range of the worker is calculated by modeling a solid having a length, a depth, and a height set in advance by a skeleton line generated from the 3-dimensional joint data, and the region representing the range of the robot is calculated from a link generated by modeling a solid and the angles of the plurality of joint axes, the solid having a length, a depth, and a height set in advance by a skeleton link generated from the plurality of joint axes.

2. The safety vision device according to claim 1, wherein the 2-dimensional image is a frame image captured by the camera at a prescribed frame rate.

3. The safety vision device according to claim 1 or 2, wherein the safety vision device further has a notification unit that outputs a warning sound when the proximity determination unit outputs the deceleration or stop instruction.

4. The safety vision device according to claim 1 or 2, wherein the human 3-dimensional skeleton estimation model and the robot 3-dimensional skeleton estimation model are provided in a server that is accessibly connected via a network from the safety vision device.

5. A security vision system characterized by, having: a robot; a camera; and the safety vision device according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Robot control device and robot control method

    JP2004243427A

  • Method for observation of a person in an industrial environment

    CN101511550A

  • Man-robot communion safety protection control system based on vision

    CN108527370A

  • Three-dimensional space monitoring device, three-dimensional space monitoring method, and three-dimensional space monitoring program

    CN111372735A