Spatial perception assistance system
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2026-08-13
Smart Images

Figure JP2025019987_13082026_PF_FP_ABST
Abstract
Description
Spatial Grasp Support System
[0001] The present disclosure relates to a spatial grasp support system that enables an operator who operates a machine while viewing an image displayed on an image display device to easily grasp the working space.
[0002] In recent years, a machine operation system in which an operator operates a machine while viewing an image displayed on an image display device has been put into practical use.
[0003] In a machine operation system in which an operator operates a machine while viewing an image displayed on an image display device, images of the machine, workpiece, and surrounding environment are captured by a camera, and the captured images are transmitted to and displayed on an image display device installed around the operator. The operator operates the machine while viewing the image displayed on the image display device. For example, when operating a robot, the operator operates the robot while viewing the image displayed on the image display device, and performs operations such as a pick-up operation of gripping and lifting a workpiece with an end effector and a place operation of placing the workpiece gripped by the end effector at another location.
[0004] In order to perform a pick operation of gripping and lifting a workpiece with an end effector and a place operation of placing the workpiece gripped by the end effector, it is necessary to move the end effector directly above the target point. However, since the image captured by the camera lacks information corresponding to the depth, it is difficult to grasp the current position of the end effector and the position directly above the target point from the image displayed on the image display device and to determine the direction in which the machine should be operated.
[0005] Patent Document 1 discloses a device that displays on an output device a sight coordinate axis including a line segment extending vertically from a control point set on a drive target device to the surface of a workpiece within a virtual space defined by three-dimensional data of a structure. According to the technique disclosed in Patent Document 1, since a line segment extending vertically from a control point set on a drive target device to the surface of a workpiece is visualized, an operator can position the end effector by referring to the sight coordinate axis.
[0006] Japanese Patent No. 6385627
[0007] However, the technology disclosed in Patent Document 1 does not include a technology for displaying support information on the actual image of a camera with distortion. Furthermore, the technology disclosed in Patent Document 1 is based on computer graphics operation aimed at repetitive work by automated operation after teaching, and in order to realize the technology, three-dimensional data of the surrounding environment without defects, including the workpiece and blind spots, is an indispensable element in generating the image to be displayed on the output device. For this reason, it is necessary to capture the actual image of the unpredictable work environment in real time and display it on the image display device, which makes it difficult to prepare such three-dimensional data, and therefore it could not be applied to general-purpose machine operation systems in which the work content may change each time.
[0008] For example, consider a machine operation scenario where the machine moves to unspecified stores and manipulates a wide variety of products on multiple product display shelves in each store. In this case, the machine's own position becomes uncertain due to movement, making it difficult to easily obtain accurate three-dimensional coordinates of the product display shelves as seen from the machine. Furthermore, it is difficult to maintain all three-dimensional data for any given product and to use three-dimensional data identified from that data based on video. Moreover, if a new product is launched or an existing product is renewed, and a product for which no three-dimensional data is available appears, the machine will be unable to respond. In addition, the situation on the product display shelves changes dynamically due to customer access in the stores, but it is not realistic to assume that every shelf in every store is equipped with a three-dimensional scanner to provide the machine with real-time three-dimensional information. Even if a three-dimensional scanner were available, it would be impossible to avoid problems such as blind spots on the back of workpieces or workpieces stacked on lower levels being obscured by workpieces stacked on upper levels. Thus, under general-purpose conditions, it was difficult to prepare the three-dimensional data that is an essential element in the technology disclosed in Patent Document 1.
[0009] This disclosure is made in view of the above, and aims to provide a spatial awareness support system that can be applied to machine operation systems where the work content changes each time, and that can help the operator grasp the real-world position of the end effector.
[0010] To solve the above-mentioned problems and achieve the objective, the spatial awareness support system according to this disclosure comprises a machine equipped with an end effector that performs work on a workpiece, a control device that controls the machine, an operation input device connected to the control device that accepts operator input operations, a camera that photographs the area around the end effector, an image display device that displays an operation image for the operator to refer to when performing input operations, and a depth sensor provided on the machine that measures the distance to an object on the extension of the center of the end effector. The control device includes a display image generation unit that calculates the screen coordinates on the image captured by the camera of a first point indicating the surface of an object on the extension of the center of the end effector based on the camera parameters and lens distortion coefficient of the camera, the distance to an object on the extension of the center of the end effector measured by the depth sensor, and commands given to the machine via input operations or status values fed back from the machine, and generates an operation image by superimposing a marker of a preset shape indicating that it is the surface of an object on the extension of the center of the end effector onto the screen coordinates of the first point on the image captured by the camera.
[0011] Alternatively, in order to solve the above-mentioned problems and achieve the objectives, the spatial awareness support system according to this disclosure comprises a machine equipped with an end effector that performs work on a workpiece, a control device that controls the machine, an operation input device connected to the control device that receives input operations from an operator, a camera that photographs the area around the end effector, and an image display device that displays an operation image that the operator refers to when performing input operations. The control device includes a display image generation unit that generates an operation image by drawing a transparent three-dimensional figure of a predetermined shape, placed on the upper surface of a stage on which a workpiece is placed, onto the image captured by the camera, based on the camera parameters and lens distortion coefficient of the camera and a command given to the machine or a status value fed back from the machine.
[0012] According to this disclosure, the spatial awareness support system can be applied to machine operation systems where the work content changes each time, and it has the effect of providing an operator with the ability to grasp the real-world position of the end effector.
[0013] Figure showing the configuration of the spatial awareness support system according to Embodiment 1 Figure showing the configuration of the control device for the spatial awareness support system according to Embodiment 1 Figure showing a first example of an operation video displayed on a video display device by the spatial awareness support system according to Embodiment 1 Figure showing a second example of an operation video displayed on a video display device by the spatial awareness support system according to Embodiment 1 Figure showing a third example of an operation video displayed on a video display device by the spatial awareness support system according to Embodiment 1 Figure showing an example of marker transitions drawn in the operation video by the spatial awareness support system according to Embodiment 1 Figure Flowchart showing the processing flow of the spatial awareness support system according to Embodiment 1 Flowchart showing the processing flow for measuring the distance from the machine to the top surface of an object directly below the center of the end effector of the spatial awareness support system according to Embodiment 1 Figure showing an example of drawing a marker on the depth camera image of the spatial awareness support system according to Embodiment 1 Figure showing the configuration of the spatial awareness support system according to Embodiment 2 Figure showing the configuration of the control device for the spatial awareness support system according to Embodiment 2 Figure 1 shows a first example of an operation video displayed on a video display device by the spatial awareness support system. Figure 2 shows a flowchart showing the processing flow of the spatial awareness support system according to Embodiment 2. Figure 3 shows a third example of an operation video displayed on a video display device by the spatial awareness support system according to Embodiment 2. Figure 4 shows an example of a problem solved by the spatial awareness support system according to Embodiment 2. Figure 5 shows another example of a problem solved by the spatial awareness support system according to Embodiment 2. Figure 6 shows an example of the spatial awareness support effect on the operator by the operation video displayed on a video display device by the spatial awareness support system according to Embodiment 2. Figure 7 shows a first example of an operation video displayed on a video display device by the spatial awareness support system according to Embodiment 3. Figure 8 shows a second example of an operation video displayed on a video display device by the spatial awareness support system according to Embodiment 3. Figure 9 shows an example of a hardware configuration that realizes the control device of the spatial awareness support system according to Embodiments 1, 2 and 3.
[0014] The spatial awareness support system according to an embodiment will be described in detail below with reference to the drawings.
[0015] Embodiment 1. Figure 1 is a diagram showing the configuration of the spatial awareness support system according to Embodiment 1. The spatial awareness support system 100 according to Embodiment 1 includes a machine 10 that performs work on a workpiece 30 using an end effector 11 attached to the machine end 10a, and a depth sensor 12 that measures the distance to the target and outputs depth information. The spatial awareness support system 100 includes a camera system 40 installed on the side of the work area of the end effector 11, an image display device 50 that displays an operation image which is an image for operating the machine 10, an operation input device 60 that operates the machine 10, and a control device 70 that controls the machine 10.
[0016] The machine 10, camera system 40, video display device 50, operation input device 60, and control device 70 constitute a machine operation system in which the operator operates the machine 10 while referring to the operation video displayed on the video display device 50, which captures images of the workpiece 30 and the surrounding environment in real time and displays them on the video display device 50. The camera system 40 includes a side camera 40a that captures the area around the end effector 11, and a side camera drive device 40b that changes the direction of the field of view of the side camera 40a in accordance with the movement of the end effector 11. The drive angle, which indicates the direction of the field of view of the side camera 40a, is fed back to the camera control unit 71 as the camera drive status.
[0017] In the following description, machine 10 is assumed to be a 6-axis vertical articulated robot 10b, and the end effector 11 is assumed to be a gripping type gripper 11a equipped with multiple fingers 111. However, machine 10 may be something other than a 6-axis vertical articulated robot 10b, such as a SCARA robot, and the end effector 11 may be something other than a gripping type gripper 11a, such as a suction type gripper. Machine 10 changes its posture and moves the end effector 11 by driving actuators installed corresponding to each joint axis according to operations input from the operator via the operation input device 60.
[0018] The depth sensor 12 is provided with a flange in the same shape as the machine end 10a, and an end effector 11 for the machine 10 can be attached to it. That is, the end effector 11 can be attached to the machine end 10a via the depth sensor 12. Therefore, the spatial awareness support system 100 can use an existing end effector 11 for the machine 10 or any general end effector 11 and a corresponding tool changer. The depth sensor 12 is installed with a tilt angle at an offset position from the center of the end effector 11. Here, the center of the end effector 11 is the point that is the center of the tips of the multiple fingers 111 provided on the end effector 11, which is a gripping gripper 11a, on the extension of the sixth axis of the machine 10, which is a six-axis vertical articulated robot 10b. The offset amount and tilt angle are known design values that are set in advance. Furthermore, in the following description, the depth sensor 12 is assumed to be a depth camera 12a attached to the machine end 10a of the machine 10 so as to photograph the area around the end effector 11. However, the depth sensor 12 may be something other than a depth camera 12a.
[0019] Here, the coordinate system of the depth camera 12a is the local coordinate system of the depth camera 12a, with the focal point of the camera lens as the origin, the depth direction along the optical axis as the Z direction, the downward direction of the camera image as the Y direction, and the rightward direction of the camera image as the X direction.
[0020] The depth camera 12a is centered such that a virtual line extending from the center of the end effector 11 along the sixth axis of the machine 10, which is a six-axis vertical articulated robot 10b, is contained within the YZ plane of the depth camera 12a. Therefore, it is guaranteed that the position on the extension of the center of the end effector 11 will be located in the pixels in the central column in the X direction of the depth camera 12a's image, even if the depth camera 12a is installed with offset and tilt.
[0021] Machine 10 performs tasks such as gripping and lifting a workpiece 30 on the stage 80 and moving it to another location. The operator moves the end effector 11 directly above the target point by operating the operation input device 60 while viewing the image captured by the side camera 40a and displayed on the image display device 50. The operator then lowers the end effector 11, which is positioned directly above the target point, to grip the workpiece 30 or release the gripped workpiece 30. When gripping a workpiece 30 on the stage 80, the target point directly above is directly above the workpiece 30. When placing a gripped workpiece 30, the target point directly above is directly above the point where the workpiece 30 will be placed.
[0022] The operation input device 60 can be a pointing device such as a joystick, joypad, motion capture device, puppet, touch panel, or pointing stick, but it may also be a keyboard, an audio input device such as a microphone, or an eye-tracking input device.
[0023] The video display device 50 can be a general display device such as a monitor, tablet terminal, or smartphone terminal that can display still images and moving images.
[0024] Figure 2 shows the configuration of the control device for the spatial awareness support system according to Embodiment 1. The control device 70 includes a camera control unit 71 that controls the camera system 40 according to the operation input from the operator at the operation input device 60, thereby causing the field of view of the side camera 40a to follow the end effector 11. The control device 70 includes a display image generation unit 72 that superimposes an augmented reality image onto the image captured by the side camera 40a based on the image captured by the side camera 40a and depth information output by the depth sensor 12, making it easier to recognize the position of the end effector 11, and displays it on the image display device 50 as an operation image, and a machine control unit 73 that controls the machine 10. The machine control unit 73 drives the end effector 11 by outputting commands to the machine 10, and also receives feedback of machine status and camera drive angle status from the machine 10 and the camera system 40.
[0025] The spatial awareness support system 100 according to Embodiment 1 superimposes an augmented reality image of a point indicating the surface of an object on the extension of the center of the end effector 11 onto the image captured by the side camera 40a. Hereinafter, the augmented reality image of a point indicating the surface of an object on the extension of the center of the end effector 11 will be referred to as a "marker". The augmented reality image of the line connecting the center of the end effector 11 and the marker will be referred to as a "guideline". The augmented reality image of the "marker" and "guideline" combined will be referred to as a "pointer". If the point indicating the surface of an object on the extension of the center of the end effector 11 is designated as the first point, and the second point is predetermined around the end effector 11, the marker is drawn at the first point, and the guideline 52 is drawn as a line segment connecting the first point and the second point. Here, the second point is, for example, the center of the tip of the end effector 11. By making the second point the center of the tip of the end effector 11, it becomes easier to obtain spatial information such as the positional relationship between the end effector 11 and objects in the work environment, including the workpiece 30, and whether or not the workpiece is between the fingers 111 necessary for gripping the workpiece 30. However, the second point is not limited to the center of the tip of the end effector 11.
[0026] Figure 3 shows a first example of an operation video displayed on an image display device by the spatial awareness support system according to Embodiment 1. When the end effector 11 is not gripping the workpiece 30, the marker 51 is drawn on the upper surface of the workpiece 30, which is on the extension of the center of the end effector 11. The guideline 52 is drawn as a line segment connecting the second point, which is the center of the end effector 11, and the marker 51 drawn at the first point.
[0027] Figure 4 shows a second example of an operation video displayed on an image display device by the spatial awareness support system according to Embodiment 1. The end effector 11 is not gripping the workpiece 30, but the workpiece 30 is between the fingers 111, and the upper surface of the workpiece 30 is positioned above the tip of the end effector 11. In this state, the display image generation unit 72 considers the workpiece 30 to be an object on the extension of the center of the end effector 11 and draws a marker 51 on the upper surface of the workpiece 30. The guideline 52 is drawn as a line segment connecting the second point, which is the center of the end effector 11, and the marker 51 drawn at the first point.
[0028] Figure 5 shows a third example of an operation video displayed on a video display device by the spatial awareness support system according to Embodiment 1. When the end effector 11 grips and lifts the workpiece 30, the marker 51 is drawn not on the top surface of the workpiece 30, but on the top surface of an object other than the workpiece 30, on the extension of the center of the end effector 11. The guideline 52 is drawn as a line segment connecting the second point, which is the center of the end effector 11, and the marker 51 drawn at the first point. As a result, the spatial awareness support system 100 according to Embodiment 1 can present the operator with a target point when transporting the workpiece 30.
[0029] Figure 6 shows an example of marker transitions drawn in the operation video by the spatial awareness support system according to Embodiment 1. In state (A), the end effector 11 is located directly above the dish 31, not directly above the workpiece 30, so the marker 51 is drawn on the top surface of the dish 31. When the end effector 11 moves directly above the workpiece 30 to state (B), the marker 51 is drawn on the top surface of the workpiece 30. When the end effector 11 moves further and comes to be directly above a part of the workpiece 30 that is a different color from the part that was directly below the center of the end effector 11 in state (B), to state (C), the color of the marker 51 changes. When the end effector 11 moves further and passes directly above the workpiece 30 to state (D), the marker 51 is drawn again on the top surface of the dish 31 with the same color as in state (A).
[0030] Figure 7 is a flowchart showing the processing flow of the spatial awareness support system according to Embodiment 1. In step S11, the display image generation unit 72 measures the distance from the machine 10 to the upper surface of an object that lies on the extension of the center of the end effector 11.
[0031] The details of the process in step S11 will now be explained. Figure 8 is a flowchart showing the process flow for measuring the distance from the machine to the top surface of an object on the extension of the center of the end effector in the spatial awareness support system according to Embodiment 1. In step S111, the display image generation unit 72 sequentially scans the central column of the depth matrix output by the depth camera 12a and performs deprojection processing using the screen coordinates, the depth value recorded in the pixels of the screen coordinates, and the camera parameters of the depth camera 12a to obtain the depth camera coordinates of the subject captured in the pixel in the tilted and offset depth camera 12a. Here, deprojection processing refers to the calculation of projecting the screen coordinates of the object image on the image onto a three-dimensional coordinate system in real space. In step S112, the display image generation unit 72 performs a process to rotate the coordinate values obtained in step S111 at the depth camera coordinates in the opposite direction by the same magnitude as the tilt angle of the depth camera 12a, converting them into coordinate values of the depth camera 12a that are parallel to the end effector 11 at the installation position of the depth camera 12a.
[0032] In step S113, the display image generation unit 72 translates the origin of the camera coordinates, which were rotated in step S112, toward the center of the end effector 11 by the offset amount of the depth camera 12a. By translating the origin of the camera coordinates by the same amount as the offset amount of the depth camera 12a, the depth camera coordinates after the translation become the same as the depth camera coordinates when the depth camera 12a is installed in the center of the end effector 11 and facing the same direction as the end effector 11. In the depth camera coordinates after this coordinate transformation by rotation and translation, the scanned pixels where both the X and Y components are zero are the pixels that have depth information of an object on the extension of the end effector 11, and the value of the Z component of the transformed depth camera coordinates is the distance from the camera focal plane to the object on the extension of the center of the end effector 11.
[0033] Step S114 is a process for determining the end of the loop from steps S111 to S113. In step S114, the display image generation unit 72 calculates the square root of the sum of the squares of the X and Y components of the depth camera coordinate value after coordinate transformation obtained in step S113 for the scanned pixel, and calculates the distance from the z-axis, which is the coordinate axis in the Z direction, to the depth camera coordinate indicated by the scanned pixel. If the distance from the z-axis to the pixel is less than or equal to a preset tolerance value, the display image generation unit 72 determines that it has found a pixel where both the X and Y components are zero, that is, the depth camera coordinate of an object on the extension line of the end effector 11, where the Z component indicates the distance to an object on the extension line of the end effector 11, and terminates the process. If the distance from the z-axis to the pixel is greater than the preset tolerance value, the display image generation unit 72 returns to step S111 and repeats the process from steps S111 to S114 until a pixel is found whose distance from the z-axis is less than or equal to the preset tolerance value.
[0034] Hereinafter, the depth camera coordinates of an object on the extension of the end effector after coordinate transformation of the search pixels obtained in step S11 will be referred to as the "three-dimensional coordinates of the drawing point" in the depth camera coordinate system. The "three-dimensional coordinates of the drawing point" in the depth camera coordinate system are equivalent depth camera coordinates to the coordinates indicated by the pixel at the focal position when the depth camera 12a is installed without tilt angle and offset, that is, the pixel at the focal position when the depth camera 12a is installed facing directly downwards in the center of the end effector 11. The X and Y components are approximately zero, and the Z component is an end effector local coordinate system indicating the distance to the object on the extension of the end effector 11.
[0035] In step S12, the display image generation unit 72 converts the "three-dimensional coordinates of the drawing points" in the depth camera coordinate system obtained in step S11, when the depth camera 12a is installed in the center of the end effector 11 in the direction of the end effector 11, from the end effector local coordinate system to the world coordinate system using the offset value from the machine end of the depth camera 12a in the direction of the end effector 11 and the world coordinates of the machine end. The world coordinates of the machine end may be determined based on a command value output by the machine control unit 73 to the machine 10, or based on a status value fed back from the machine 10 to the machine control unit 73.
[0036] In step S13, the display image generation unit 72 converts the coordinate values of the "three-dimensional coordinates of the drawing point" in the world coordinate system obtained in step S12 into coordinate values in the side camera local camera coordinate system. For example, the display image generation unit 72 converts the three-dimensional coordinates of the drawing point in the world coordinate system into coordinate values in the side camera local camera coordinate system using the relative position from the machine origin where the side camera is installed and the camera drive angle. The camera system drive angle may be determined based on a command output by the camera control unit 71 to the side camera drive device 40b, or it may be determined by the camera drive status fed back to the camera control unit 71.
[0037] In step S14, the display image generation unit 72 performs a calculation to project the coordinate values of the "three-dimensional coordinates of the drawing point" in the side camera's local camera coordinate system onto the coordinate values in the side camera's screen coordinate system. This process is generally called "projection processing". For example, the display image generation unit 72 converts the coordinate values in the side camera's local camera coordinate system to the coordinate values in the side camera's screen coordinate system based on the camera matrix and lens distortion coefficient of the side camera 40a. The pixels of the side camera's screen coordinates obtained in step S14 are the pixels to be drawn. Here, the "pixels to be drawn" are the pixels corresponding to the first point on the upper surface of an object that lies on the extension of the center of the end effector 11.
[0038] In step S15, a marker 51 centered on the pixel to be drawn, based on the screen coordinates of the side camera obtained in step S14, is drawn on the surface of the image of the object on the extension of the center of the end effector 11, and a guideline 52 connecting the midpoint of the end effector 11 and the marker 51 is drawn. The marker 51 is drawn in a predetermined arbitrary shape. For example, the marker 51 may be a circle, triangle, square, asterisk shape, star shape, letter, or other arbitrary shape of character illustration centered on the pixel to be drawn, or it may be a crosshair shape with an intersection point at the pixel to be drawn. When drawing a crosshair-shaped marker 51, the display image generation unit 72 similarly obtains the coordinate values in screen coordinates not only for one point on the extension of the center of the end effector 11, but also for multiple points that make up the line segment of the crosshair. This makes it possible to draw a crosshair-shaped marker 51 along a curved surface even if the surface on which the marker 51 is drawn is curved.
[0039] When drawing the marker 51, the display image generation unit 72 determines the drawing color of the marker 51 based on the color of the pixel to be drawn. For example, the display image generation unit 72 converts the color of the pixel to be drawn from the RGB color space to the HSV color space and draws a highly visible marker 51 using a color whose hue difference with the color of the pixel to be drawn is greater than a preset value. An example of a color whose hue difference with the color of the pixel to be drawn is greater than a preset value is the complementary color of the color of the pixel to be drawn, but it is not limited to complementary colors. If the pixel to be drawn is an achromatic color with no hue and a complementary color cannot be determined, or if the pixel to be drawn is low in saturation close to achromatic and contrast with the pixel to be drawn cannot be obtained using a complementary color, the display image generation unit 72 may draw the marker 51 with a color whose brightness difference with the pixel to be drawn is greater than or equal to a preset value. Alternatively, the display image generation unit 72 may draw the marker 51 with a preset fixed color regardless of the color of the pixel to be drawn.
[0040] Since the distance from the location where the end effector 11 is installed to the location where the depth camera 12a is installed is a known jig design dimension, the distance to an object on the extension of the center of the end effector 11 can be converted into the machine 10's reference coordinates. Since it is also possible to convert these coordinates into the side camera's local camera coordinate system by translational and rotational coordinate transformations, by converting the surface coordinates of an object on the extension of the center of the end effector 11 into screen coordinates using the camera parameters of the side camera 40a, it is possible to draw the marker 51 as described above on the side camera image.
[0041] Furthermore, if the depth sensor 12 is a depth camera 12a, the RGB image from the depth camera 12a can be used as a close-range overhead view, and the marker 51 can be superimposed and displayed on the video display device 50 as a work image. In other words, the depth camera 12a may be assigned both the function of capturing the area around the end effector 11 and the function of measuring the distance to an object on the extension of the center of the end effector 11. Figure 9 is a diagram showing an example of drawing a marker on the depth camera image of the spatial awareness support system according to Embodiment 1. By drawing the marker 51 on the surface of the workpiece 30 in the RGB image from the depth camera 12a, the operator can easily determine whether or not the end effector 11 is in a position to grip the workpiece 30. The RGB image from the depth camera 12a has high resolution due to the short distance to the workpiece 30, but it cannot obtain a field of view that allows for an overall view like the side camera 40a. However, if a narrow field of view is acceptable, a machine operation system consisting only of a depth camera 12a, omitting the camera system 40, is also practical. In such a configuration, a machine operation system with spatial awareness support functionality can be constructed using video by simply adding a depth sensor 12 unit between the machine end of a direct-view machine operation system and the end effector 11 being used.
[0042] The spatial awareness support system 100 according to Embodiment 1 can display an operation video on the video display device 50 in which a marker 51 indicating the surface of an object on the extension of the center of the end effector 11 is drawn. Therefore, the operator operating the machine 10 can accurately recognize the position directly above the end effector 11. Furthermore, since the spatial awareness support system 100 according to Embodiment 1 does not require the prior preparation of three-dimensional data of the workpiece 30 when drawing the marker 51 on the operation video, it can be immediately applied to existing machine operation systems that capture images of the workpiece 30 and the surrounding environment in real time and display them on the video display device 50.
[0043] Furthermore, the spatial awareness support system 100 according to Embodiment 1 does not require special display devices for displaying three-dimensional images, such as head-mounted displays and 3D displays. Therefore, the spatial awareness support system 100 according to Embodiment 1 has ubiquitous capabilities that can be used in a wide range of operating environments where special display devices are not installed, and it is possible to use a general display device as the image display device 50 to make it easier for the operator to grasp the position of the end effector 11. In addition, there are individual differences in how three-dimensional images are perceived, but since the image displayed on the image display device 50 of the spatial awareness support system 100 according to Embodiment 1 is not a three-dimensional image, there are no individual differences in the degree to which the operator grasps the position of the end effector 11. Also, three-dimensional images can be difficult to see in three dimensions if there is a difference in the visual acuity of both eyes, but since the spatial awareness support system 100 according to Embodiment 1 does not display a three-dimensional image on the image display device 50, even an operator with extremely low visual acuity in one eye can easily recognize the position of the end effector 11. Therefore, even when the spatial awareness support system 100 according to Embodiment 1 is applied to a machine operation system where the machine is operated while directly looking at it, it can achieve barrier-free operation through the introduction of camera images. Furthermore, while three-dimensional images can easily cause fatigue in viewers because they use optical illusions to simulate three-dimensional shapes, the spatial awareness support system 100 according to Embodiment 1 does not display three-dimensional images on the image display device 50. This reduces fatigue for operators who need to operate the machine 10 while looking at the operation images displayed on the image display device 50, enabling them to work for longer periods.
[0044] Furthermore, since a three-dimensional image composed of left and right images handles twice the amount of data as a two-dimensional image, using a three-dimensional image for operation increases the transmission time of the image data and the image processing time, requiring high computational processing power from the image processing device, and also increasing the display delay that worsens the operability of the machine. In contrast, in the spatial awareness support system 100 according to Embodiment 1, since the operation image is not a three-dimensional image, the computational processing power required of the control device 70 that processes the image is small, and the display delay of the image display device 50 is further reduced.
[0045] Also, the spatial grasping support system 100 according to Embodiment 1 does not irradiate the workpiece 30 with a laser beam or the like. Instead, in order to make it easier for the operator to recognize the position of the end effector 11 by means of an extended virtual reality video by software processing, even when the workpiece 30 is being gripped by the end effector 11, the optical path is not blocked by the workpiece as shown in FIG. 5, and the operator can easily grasp the position of the end effector 11. Further, the spatial grasping support system 100 according to Embodiment 1 makes it easier for the operator to grasp the position of the end effector 11 by means of an extended virtual reality video by software processing. Therefore, even when working in a bright working environment by the machine 10 installed outdoors, the operator can easily grasp the position of the end effector 11 without being affected by external light.
[0046] For example, between a beach in the middle of summer and a dimly lit room, there is a difference of 12 in the exposure value showing a light amount twice as much for each increment, that is, a difference in brightness of 2 to the 12th power = 4096 times. Even limited to indoors, there is a difference of 6 in the exposure value depending on the intensity of illumination, that is, a difference in brightness of 2 to the 6th power = 64 times.
[0047] The human eye has a dynamic range of 10 to the 9th power with respect to brightness. In addition to being able to adjust the light amount by about 20 times with a reflection of about 0.2 seconds by the iris, by taking about 20 minutes, the cone cells and rod cells undergo light adaptation or dark adaptation, and the retina itself changes its sensitivity by a factor of 10 to the 6th power.
[0048] The camera obtains proper exposure of subjects such as workpieces and working environments in response to differences in brightness through combinations of ISO sensitivity, which is a standard for photographic films established by the International Organization for Standardization (ISO), ND (Neutral Density) filters, aperture, and shutter speed. For example, a five-step aperture reduction from F2.8 to F16 or a shutter speed change from 1 / 30 second to 1 / 1000 second means a differential adjustment of 5 in the exposure value. Thus, when the camera side makes an exposure value adjustment equivalent to 2 to the 6th power to 2 to the 12th power for differences in external brightness, a change in the emission intensity of the same magnification is required to make the light spot of the light-emitting device appear equally on the camera image. That is, assuming a Class 1 laser with an output of 0.2 mW is made to correspond to a dim indoor environment, an output of about 800 W is required for a 4096-fold case, and an output of about 10 W is required even for a 64-fold case.
[0049] Although it is possible to emit a laser with these outputs, problems arise such as the need to set the emission wavelength in the visible region, the possibility of damaging the workpiece, the device, the working environment, etc. because the output value falls within the range of a laser processing machine, and the need to ensure safety for people, making it virtually impossible to cover the influence of external light by changing the emission intensity. Thus, in a configuration using a light-emitting device, the influence of external light cannot be covered by changing the emission intensity.
[0050] Furthermore, if the surface of the workpiece 30 is mirror-like, it is difficult to display the marker 51 on the surface of the workpiece 30 by laser irradiation because the laser light will be reflected even if it is irradiated onto the workpiece 30. In contrast, the spatial awareness support system 100 according to Embodiment 1 makes it easier for the operator to grasp the position of the end effector 11 through augmented virtual reality images processed by software, so it is possible to display the marker 51 on the surface of the workpiece 30 even if the surface of the workpiece 30 is mirror-like. Also, if the workpiece 30 is transparent, it is difficult to display the marker 51 on the surface of the workpiece 30 by laser irradiation because the laser light will pass through even if it is irradiated onto the workpiece 30. However, the spatial awareness support system 100 according to Embodiment 1 makes it easier for the operator to grasp the position of the end effector 11 through augmented virtual reality images processed by software, so it is possible to display the marker 51 on the surface of the workpiece 30 even if the surface of the workpiece 30 is transparent.
[0051] Furthermore, the spatial awareness support system 100 according to Embodiment 1 does not require a light-emitting device, thus enabling cost reduction and avoiding the risk of light-emitting device failure. In addition, since the spatial awareness support system 100 according to Embodiment 1 does not require the integration of a light-emitting device into the end effector 11, existing end effectors 11 can be used, and design flexibility can be increased. Moreover, since the spatial awareness support system 100 according to Embodiment 1 does not use a light-emitting device, the consideration of pedestrian safety required when using laser light is unnecessary, and operation in places with pedestrian traffic becomes possible.
[0052] Embodiment 2. Figure 10 shows the configuration of the spatial awareness support system according to Embodiment 2. The spatial awareness support system 100 according to Embodiment 2 differs from the spatial awareness support system 100 according to Embodiment 1 in that it does not have a depth camera 12a which is a depth sensor 12. Since the configuration shown in Figure 10 is a general configuration for a machine operation system that uses video, the spatial awareness support system 100 according to Embodiment 2 can be realized without adding new hardware to any existing machine operation system.
[0053] Figure 11 is a diagram showing the configuration of the control device for the spatial awareness support system according to Embodiment 2. Figure 12 is a diagram showing a first example of an operation video displayed on an image display device by the spatial awareness support system according to Embodiment 2. The spatial awareness support system 100 according to Embodiment 2 superimposes a composite reality image, which includes a transparent three-dimensional figure 55 that is set to an arbitrary shape in advance and intersecting lines 56 that indicate the position of the end effector 11 in the movable range, onto the image captured by the side camera 40a. In the following description, the transparent three-dimensional figure 55 is assumed to be a shape set based on the movable range of the end effector 11. The transparent three-dimensional figure 55 may be the same as the shape that indicates the movable range of the end effector 11, or it may be a shape that is slightly larger or slightly smaller than the shape that indicates the movable range of the end effector 11. The intersecting lines 56 are composed of a plurality of line segments whose intersection points are points where the extension of the sixth axis of the machine 10, which is the orientation of the end effector 11, penetrates the surface constituting the transparent three-dimensional figure 55.
[0054] The crossing line 56 includes a first crossing line 56a drawn on the upper surface of the transparent solid figure 55 and a second crossing line 56b drawn on the bottom surface of the transparent solid figure 55. The first crossing line 56a is drawn by two or more intersecting line segments 56e, and the second crossing line 56b is drawn by two or more intersecting line segments 56f. The line segments 56e constituting the first crossing line 56a include two line segments parallel to the two directional axes of the horizontal movement of the end effector 11 by a command given to the machine 10. The second crossing line 56b is drawn at a position where the first crossing line 56a is projected onto the bottom surface of the transparent solid figure 55, parallel to the direction of vertical movement of the end effector 11 by a command given to the machine 10 in response to an up-and-down input operation to the operation input device 60. Here, the first crossing line 56a and the second crossing line 56b are assumed to be cross lines drawn by two intersecting line segments. The transparent solid figure 55 includes a line segment 55a that forms the top surface, a line segment 55b that forms the bottom surface, a line segment 55c that connects the top surface and the bottom surface, and a line segment 56d that connects the intersection points of the first intersecting line 56a and the second intersecting line 56b.
[0055] The intersection point of the first intersecting line 56a is the second point, which is the center of the end effector 11. When the machine 10 operates in a cylindrical coordinate system, the three-dimensional shape is suitable as a sector or cylinder, and when the machine 10 operates in a Cartesian coordinate system, a rectangular parallelepiped is suitable, but is not limited to these. In the following description, we will assume that the machine 10 operates in a Cartesian coordinate system and take the example of applying a rectangular parallelepiped to the transparent three-dimensional figure 55, the base of which is within the range of motion of the end effector 11. When the machine 10 operates in a Cartesian coordinate system, the display image generation unit 72 draws a first intersection line 56a on the top surface of the transparent solid figure 55, with the second point as the intersection point and parallel to the two directional axes of the horizontal movement of the end effector 11 according to the command given to the machine 10. Then, a second intersection line 56b is drawn on the bottom surface of the transparent solid figure 55, projected from the first intersection line 56a and parallel to the direction of the vertical movement of the end effector 11 according to the command given to the machine 10 by an up-and-down input operation to the operation input device 60. The side surface of the transparent solid figure 55 is drawn by a plane or curved surface composed of a set of points that the end effector 11 can reach at the height between the bottom surface of the transparent solid figure 55 and the top surface of the transparent solid figure 55.
[0056] Figure 13 is a flowchart showing the processing flow of the spatial awareness support system according to Embodiment 2. In step S21, the display image generation unit 72 defines the coordinates of the endpoints of the line segments constituting the transparent solid figure 55 in the world coordinate system, which is the coordinate system of the machine. The endpoint coordinates when the horizontal plane shape of the transparent solid figure 55 is the orthogonal movable range of the machine 10 are as follows. The x and y coordinates of the vertices of the transparent solid figure 55 are the same as the horizontal plane coordinates of the machine movable range end. The z coordinate of the vertices of the transparent solid figure 55 is the same as the z coordinate of the second intersecting line 56b on the bottom surface, and is the stage height. The z coordinate of the vertices of the transparent solid figure 55 is the same as the z coordinate of the second point, which is the center of the end effector 11, and the first intersecting line 56a on the top surface, and is the value obtained by subtracting the height offset amount between the machine end and the tip of the end effector from the machine command value or status value indicating the position of the machine end. The y coordinates of the left and right line segments constituting the intersecting line 56 are the y coordinates of the machine command value or status value. The x-coordinates of the line segments before and after the intersection line 56 are the x-coordinates of the machine command value or status value.
[0057] By aligning the bottom surface of the transparent 3D figure 55 with the range of motion of the end effector 11, the operator can more easily grasp the range of motion of the end effector 11. However, from the perspective of spatial awareness alone, the bottom surface of the transparent 3D figure 55 does not necessarily have to coincide with the range of motion of the end effector 11. Furthermore, by aligning the bottom surface of the transparent 3D figure 55 with the top surface of the stage 80, which is the lowest surface that the machine 10 accesses during operation, and aligning the top surface of the transparent 3D figure 55 with the horizontal plane containing the tip of the end effector 11, the shape of the bottom surface of the transparent 3D figure 55 can be made to indicate the reachable range of the end effector 11 on the lowest surface that the machine 10 accesses during operation, and the reachable range of the end effector 11 on the horizontal plane containing the tip of the end effector 11 can be made to the top surface of the transparent 3D figure 55. In this way, the operator can more easily intuitively recognize the height of the end effector 11 from the stage 80.
[0058] In order to align the bottom surface of the transparent three-dimensional figure 55 with the top surface of the stage 80, "stage height," which is information about the relative height from the machine origin to the stage surface, is required. If the machine 10 is stationary, the stage height can be stored in the control device 70 as a value measured in advance. On the other hand, if the machine 10 is movable and the distance to the stage 80 is indeterminate, the stage height must be determined by some method. For example, the stage height can be determined based on the command value when the end effector 11 is lowered until it touches the stage surface. Alternatively, the stage height can be determined by attaching a depth camera 12a to the end effector 11 and using the depth camera 12a. Furthermore, the stage height can be determined by installing a hand camera on the end effector 11 and an active or passive marker on the stage 80, and analyzing the image captured by the hand camera of the active or passive marker. Finally, the stage height can also be determined by pre-displaying the height from the floor on the stage 80 as a numerical value and reading it optically. Alternatively, the control device 70 may be pre-stored with multiple values for stage height, and a value close to the height of the stage 80 from the floor may be selected.
[0059] Furthermore, the display image generation unit 72 calculates the coordinate values in the world coordinate system of the ends of line segments that have the sides of the transparent solid figure 55 as their ends, using the position on the extension of the center of the end effector 11 as the intersection point on the bottom surface of the transparent solid figure 55. For example, if the transparent solid figure 55 is a rectangular prism, the coordinate values of the ends of at least two line segments are calculated, and if the transparent solid figure 55 is a sector or cylinder, the coordinate values of the ends of one circular arc and one radial line segment are calculated at least, but the coordinates of the ends of more line segments than the number exemplified may be calculated. Note that by making the ends of the line segments intersection points with the sides of the transparent solid figure 55, the operator can easily grasp the range of motion of the end effector 11, but the ends of the line segments do not necessarily have to be intersection points with the sides of the transparent solid figure 55.
[0060] In step S22, the display image generation unit 72 converts the coordinate values of the endpoints of the line segments constituting the transparent three-dimensional figure 55 and the intersecting line 56 defined in step S21 from the world coordinate system to the side camera local camera coordinate system, based on the physical distance to the installation position of the side camera 40a relative to the machine installation point, which is the world coordinate origin of the side camera 40a, and the orientation indicated by the drive angle of the side camera 40a. The drive angle of the side camera 40a may be a command value output by the camera control unit 71 of the control device 70, or it may be a status value of the side camera drive device 40b that is fed back from the camera system 40.
[0061] In step S23, the display image generation unit 72 converts the coordinate values of the endpoints from the camera coordinate system to the screen coordinate system. Similar to Embodiment 1, the display image generation unit 72 converts the coordinate values in the side camera's local camera coordinate system to the side camera's screen coordinates based on the camera matrix and lens distortion coefficient of the side camera 40a.
[0062] In step S24, the display image generation unit uses the coordinate values of the screen coordinate system to draw a transparent three-dimensional figure 55 that shows the range of motion of the end effector 11 and an intersecting line 56 that shows the position of the end effector 11 within the range of motion.
[0063] If the target image on which the transparent 3D figure 55 and intersecting lines 56 are superimposed does not have lens distortion, the display image generation unit 72 draws the transparent 3D figure 55 and intersecting lines 56 with line segments that linearly connect the two endpoints. On the other hand, if the target image on which the transparent 3D figure 55 and intersecting lines 56 are superimposed does have lens distortion, it is preferable to draw the transparent 3D figure 55 and intersecting lines 56 with line segments that gradually increment the camera coordinates from one endpoint to the other. In other words, if the target image on which the transparent 3D figure 55 and intersecting lines 56 are superimposed does have lens distortion, it is preferable to superimpose the transparent 3D figure 55 and intersecting lines 56 with curves in the target image on which the transparent 3D figure 55 and intersecting lines 56 are superimposed, as they are straight lines in real space. By doing so, the transparent 3D figure 55 and intersecting lines 56 are drawn with the same distortion as in real space, thereby improving the spatial presentation capability to the operator.
[0064] Furthermore, the camera model used for converting from camera coordinates to screen coordinates may have limitations on the field of view it can handle. For example, a camera model that performs perspective projection transformation normalized by distance will result in division by zero at a point with a field of view of 180° where the distance is zero, i.e., a point coplane with the focal plane. This also increases the camera calibration error for obtaining distortion coefficients and other values for surrounding points, making it likely that the normalized value will deviate significantly from a reasonable value. For this reason, if the camera coordinates of the endpoints of the line segments 55a, 55b, 56e, and 56f that constitute the transparent solid figure 55 and the intersecting lines 56 are located outside a cone with the camera origin of the side camera 40a as its vertex and a preset field of view as its apex angle, the display image generation unit 72 saturates the image by treating the endpoints of the drawing line segments as being located on a cone with a preset field of view as its apex angle, and then converts them to screen coordinates to draw the transparent solid figure 55 and the intersecting lines 56. For example, in the spatial awareness support system 100 according to Embodiment 2, if the camera coordinates of the endpoints of the line segments 55a, 55b, 56e, and 56f that constitute the transparent three-dimensional figure 55 and the intersecting line 56 are located outside a cone with the camera origin of the side camera 40a as its vertex and a vertex angle of 165 degrees, the display image generation unit 72 saturates the endpoints of the drawing line segments as if they were located on the side surface of a cone with a vertex angle of 165 degrees, converts them to screen coordinates, and then draws the transparent three-dimensional figure 55 and the intersecting line 56.
[0065] Furthermore, the display image generation unit 72 draws each of the line segments 56e and 56f that form the intersecting lines 56 and the line segment 56d that connects the midpoints of the intersecting lines 56 with a different line type or color than the line segment 55a that forms the top surface of the transparent three-dimensional figure 55, the line segment 55b that forms the bottom surface of the transparent three-dimensional figure 55, and the line segment 55c that connects the top surface and the bottom surface of the transparent three-dimensional figure 55. For example, the display image generation unit 72 draws the line segment 55a that forms the top surface of the transparent three-dimensional figure 55, the line segment 55b that forms the bottom surface of the transparent three-dimensional figure 55, and the line segment 55c that connects the top surface and the bottom surface of the transparent three-dimensional figure 55 with solid black lines, the line segment 56e that forms the intersecting lines 56 with a dashed red line, the line segment 56f that forms the intersecting lines 56 with a dashed black line, and the line segment 56d that connects the midpoints of the intersecting lines 56 with a solid green line. This makes it easier for the operator to recognize the position of the end effector 11, even if some line segments overlap during system use.
[0066] Figure 14 shows a second example of an operation video displayed on an image display device by the spatial awareness support system according to Embodiment 2. When the end effector 11 is tilted, the range of motion of the end effector 11 changes, and the drawing position and shape of the transparent three-dimensional figure 55 change in accordance with the tilt of the end effector 11. By referring to the operation video after the end effector 11 has been tilted, the operator can recognize spatial information such as the workpiece 30 being located near the center of the range of motion of the end effector 11, the end effector 11 being located slightly to the left and higher than the workpiece 30, and that in order to position the end effector 11 directly above the workpiece 30, the end effector 11 should be moved approximately along the Y-axis in the negative direction by about the width of the workpiece 30, and then moved approximately 1 / 4 of the length of the workpiece 30 in the X direction.
[0067] In the example shown in Figure 14, the end effector 11 is tilted while positioned near the center of the transparent solid figure 55, so the second intersection line 56b is drawn inside the transparent solid figure 55. However, if the end effector 11 is tilted towards the outside of the transparent solid figure 55 at the periphery of the transparent solid figure 55, the second intersection line 56b may also be drawn outside the transparent solid figure 55.
[0068] Figure 15 shows a third example of an operation video displayed on an image display device by the spatial awareness support system according to Embodiment 2. In the machine's orientation shown in Figure 15, the end effector 11 is positioned so as not to reach the end of the machine's movable range when the end effector 11 is not tilted. However, as the end effector 11 tilts, the transparent solid figure 55 of the augmented reality drawn in accordance with the tilt of the end effector 11 changes, and as a result, the first intersection line 56a and the second intersection line 56b are drawn in a T-shape, similar to the end of the movable range. At this time, the control device 70 transmits the same information to the machine control unit 73, and changes the machine's movable range in real time in conjunction with the tilt of the end effector 11. As a result, the machine control unit 73 no longer accepts movement commands to the machine, as shown in the augmented reality video, and buffering between the end effector 11 and the machine itself is automatically avoided.
[0069] Figure 16 shows an example of a problem solved by the spatial awareness support system according to Embodiment 2. When the transparent three-dimensional figure 55 and the intersecting lines 56 are hidden, an operator referring to the operation video may perceive that the end effector 11 is located at point A in Figure 16. Upon such recognition, the operator will move the end effector 11 in the direction indicated by the dashed arrow in Figure 16 in order to move the end effector 11 toward the target 30a, which is the workpiece 30 to be picked up.
[0070] Figure 17 shows another example of a problem solved by the spatial awareness support system according to Embodiment 2. The operation video shown in Figure 17 and the operation video shown in Figure 16 are the same video. When the transparent 3D figure 55 and the intersecting lines 56 are hidden, an operator referring to the operation video may perceive that the end effector 11 is located on a workpiece 30 other than the target 30a, which is the workpiece 30 to be picked up. Upon such perception, the operator will move the end effector 11 in the direction indicated by the dashed arrow in Figure 16 in order to move the end effector 11 toward the target 30a, which is the workpiece 30 to be picked up.
[0071] Even when referring to the same operational video, the recognition of the location of the end effector 11 may differ among operators. While a skilled operator can improve the accuracy of estimating the position of the end effector 11 based on its size in the operational video, accurately recognizing the precise location of the end effector 11 and its direction of movement toward the target position is difficult even for a skilled operator.
[0072] Figure 18 shows an example of the spatial awareness support effect on the operator provided by the operation video displayed on the video display device of the spatial awareness system according to Embodiment 2. The operation video shown in Figure 18 is the same video as the operation videos shown in Figures 16 and 17. By displaying a composite reality video including a transparent 3D figure 55 and intersecting lines 56 as an operation video on the video display device 50, the operator can intuitively recognize the position of the end effector 11 in real space and where it will land when it is lowered straight down. In the example shown in Figure 18, since the video was taken from the diagonal left, the operator can easily recognize that the forward direction of the machine 10 is towards the upper left of the video. The operator can also easily recognize that the correct direction of movement to the target 30a, which is the workpiece 30 to be picked up, is the direction indicated by the dashed arrow in Figure 18.
[0073] Furthermore, the transparent three-dimensional figure 55 and the intersecting lines 56 visualize the "horizontal position" and "height from the stage" of the end effector 11 within its movable range in real space. As a result, an operator referring to the operation video shown in Figure 18 can intuitively recognize that the end effector is "at a sufficient height to move horizontally safely without getting caught on the workpieces on the stage," and that by moving the end effector 11 almost directly to the left by an amount roughly equal to the length of the workpiece 30, it can be moved onto the workpiece 30 which will be the target 30a.
[0074] The spatial awareness support system 100 according to Embodiment 2 superimposes a composite reality image, which includes a transparent three-dimensional figure 55 indicating the movable range of the end effector 11 and intersecting lines 56 indicating the position of the end effector 11 within the movable range, onto the image captured by the side camera 40a. As a result, the horizontal position of the end effector 11 within the movable range of the end effector 11 in real space and the height of the end effector 11 from the stage 80 are visualized in the two-dimensional image based on the line of sight direction, size, position of the intersection of the intersecting lines of the bottom and top surfaces shown by the transparent three-dimensional figure 55, and the position and length of the line segment 56d. This helps the operator understand the position of the end effector 11 and the direction and amount of operation required to achieve the desired task. Furthermore, the spatial awareness support system 100 according to Embodiment 2 does not require the prior preparation of three-dimensional data of the workpiece 30 when drawing transparent three-dimensional figures 55 and intersecting lines 56 on the operation video. Therefore, it can be applied to existing machine operation systems that need to capture images of the workpiece 30 and the surrounding environment in real time in any work environment and display them on the image display device 50, without adding any new hardware.
[0075] Furthermore, the spatial awareness support system 100 according to Embodiment 2 does not require special display devices for displaying three-dimensional images, such as head-mounted displays and 3D displays. Therefore, the spatial awareness support system 100 according to Embodiment 2 has ubiquitous capabilities that can be used in a wide range of operating environments where special display devices are not installed, and allows operators to easily grasp the position, direction of operation, and amount of operation of the end effector 11 by using a general display device as the image display device 50. In addition, there are individual differences in how three-dimensional images are perceived, but since the image displayed on the image display device 50 of the spatial awareness support system 100 according to Embodiment 2 is not a three-dimensional image, there are no individual differences in the degree to which operators grasp the position, direction of operation, and amount of operation of the end effector 11. Also, three-dimensional images can be difficult to see in three dimensions if there is a difference in visual acuity between the two eyes, but since the spatial awareness support system 100 according to Embodiment 2 does not display three-dimensional images on the image display device 50, even operators with extremely low visual acuity in one eye can easily recognize the position, direction of operation, and amount of operation of the end effector 11. Therefore, even when the spatial awareness support system 100 according to Embodiment 2 is applied to a machine operation system where the machine is operated while directly looking at it, barrier-free access can be achieved by introducing camera images. Furthermore, while three-dimensional images can easily cause fatigue in viewers because they use optical illusions to simulate three-dimensional shapes, the spatial awareness support system 100 according to Embodiment 2 does not display three-dimensional images on the image display device 50. This reduces fatigue for operators who need to operate the machine 10 while looking at the operation images displayed on the image display device 50, enabling them to work for longer periods. It should be noted that information on the "direction of operation," which is essential for machine operation, cannot be presented to the operator even using three-dimensional images. However, the spatial awareness support system 100 according to Embodiment 2 offers the unique advantage of being able to present information on the direction of operation to the operator.
[0076] Furthermore, since a three-dimensional image composed of left and right images handles twice the amount of data as a two-dimensional image, using a three-dimensional image for operation increases the transmission time of the image data and the image processing time, requiring high computational processing power from the image processing device, and also increasing the display delay that worsens the operability of the machine. In contrast, in the spatial awareness support system 100 according to Embodiment 2, since the operation image is not a three-dimensional image, the computational processing power required of the control device 70 that processes the image is small, and the display delay of the image display device 50 is further reduced.
[0077] Embodiment 3. The configuration of the spatial awareness support system 100 according to Embodiment 3 is the same as that of the spatial awareness support system 100 according to Embodiment 1. The spatial awareness support system 100 according to Embodiment 3 draws a mixed reality image on the operation image that includes both the pointer described in Embodiment 1 and the transparent three-dimensional figure 55 and intersecting lines 56 described in Embodiment 2.
[0078] Figure 19 shows a first example of an operation video displayed on an image display device by the spatial awareness support system according to Embodiment 3. Since the end effector 11 is not gripping the workpiece 30, a marker 51 is drawn on the upper surface of the stage 80, which is an object on the extension of the center of the end effector 11, in the operation video. In addition, a transparent three-dimensional figure 55 indicating the range of motion of the end effector 11 and an intersecting line 56 indicating the position of the end effector 11 are drawn in the operation video. Therefore, when the operator performs a pick operation in which the end effector 11 grips and lifts the workpiece 30, the end effector 11 can be easily positioned directly above the workpiece 30. Note that the line segment 56d connecting the midpoints of the first intersecting line 56a and the second intersecting line 56b is drawn overlapping with the guideline 52 connecting the center of the end effector 11 and the marker 51. An operator referring to the operation video shown in Figure 19 can recognize, based on the transparent 3D figure 55 and the image of the machine 10 in the operation video, that the left direction in the operation video is the front direction of the machine 10, that the workpiece 30 is located in the fourth quadrant, and that in order to approach the workpiece 30, it is necessary to move the end effector 11 to the right and slightly longer than the length of the workpiece 30.
[0079] Figure 20 shows a second example of an operation video displayed on a video display device by the spatial awareness support system according to Embodiment 3. Since the end effector 11 is gripping the workpiece 30, a marker 51 is drawn on the upper surface of the stage 80, which is an object on the extension of the center of the end effector 11, in the operation video. In addition, a transparent three-dimensional figure 55 indicating the range of motion of the end effector 11 and an intersecting line 56 indicating the position of the end effector 11 are drawn in the operation video. Therefore, when the operator performs a placing operation to place the workpiece 30 gripped by the end effector 11 at the target location, the end effector 11 can be easily positioned directly above the target location. An operator referring to the operation video shown in Figure 20 can recognize, based on the transparent 3D figure 55 and the image of the machine 10 in the operation video, that the left direction in the operation video is the front direction of the machine 10, that the target position for placing the workpiece 30 on the plate 31 is in the fourth quadrant, and that in order to approach the workpiece 30, it is necessary to move the end effector 11 to the right and slightly longer than the length of the workpiece 30.
[0080] As described above, the spatial awareness support system 100 according to Embodiment 3 superimposes a composite reality image onto the image captured by the side camera 40a. This composite reality image visualizes the operator's position in real space, the direction of operation to reach the target object, and the amount of operation, even on a two-dimensional image. This composite reality image is created by a marker 51, which is an augmented reality image showing the surface of an object on the extension of the center of the end effector 11; a transparent three-dimensional figure 55 showing the range of motion of the end effector 11; an intersecting line 56 showing the horizontal position of the end effector 11 within the range of motion; and a line segment 56d showing the height of the end effector 11. This composite reality image visualizes the operator's position in real space, the direction of operation to reach the target object, and the amount of operation to reach the target object, even on a two-dimensional image. Therefore, even in images without depth information, this system can more powerfully support the operator in performing accurate operations. Furthermore, since the spatial awareness support system 100 according to Embodiment 3 does not require the prior preparation of three-dimensional data of the workpiece 30, it can be applied to existing machine operation systems that need to capture images of the workpiece 30 and the surrounding environment in real time during any operation and display them on the image display device 50, simply by inserting a depth sensor 12 unit between the machine end 10a and the end effector 11 in use.
[0081] Furthermore, the spatial awareness support system 100 according to Embodiment 3 does not require special display devices for displaying three-dimensional images, such as head-mounted displays and 3D displays. Therefore, the spatial awareness support system 100 according to Embodiment 3 has ubiquitous capabilities that can be used in a wide range of operating environments where special display devices are not installed, and allows operators to easily grasp the position, direction of operation, and amount of operation of the end effector 11 by using a general display device as the image display device 50. In addition, there are individual differences in how three-dimensional images are perceived, but since the image displayed on the image display device 50 of the spatial awareness support system 100 according to Embodiment 3 is not a three-dimensional image, there are no individual differences in the degree to which operators grasp the position, direction of operation, and amount of operation of the end effector 11. Also, three-dimensional images can be difficult to see in three dimensions if there is a difference in visual acuity between the two eyes, but since the spatial awareness support system 100 according to Embodiment 3 does not display three-dimensional images on the image display device 50, even operators with extremely low visual acuity in one eye can easily recognize the position, direction of operation, and amount of operation of the end effector 11. Therefore, even when the spatial awareness support system 100 according to Embodiment 3 is applied to a machine operation system where the machine is operated while directly looking at it, barrier-free operation can be achieved by introducing camera images. Furthermore, while three-dimensional images can easily cause fatigue in viewers because they use optical illusions to simulate three-dimensional shapes, the spatial awareness support system 100 according to Embodiment 3 does not display three-dimensional images on the image display device 50. This reduces fatigue for operators who need to operate the machine 10 while looking at the operation images displayed on the image display device 50, enabling them to work for longer periods. It should be noted that information on the "direction of operation," which is essential for machine operation, cannot be presented to the operator even using three-dimensional images. However, the spatial awareness support system 100 according to Embodiment 3 offers the unique advantage of being able to present information on the direction of operation to the operator.
[0082] Furthermore, since a three-dimensional image composed of left and right images handles twice the amount of data as a two-dimensional image, using a three-dimensional image for operation increases the transmission time of the image data and the image processing time, requiring high computational processing power from the image processing device, and also increasing the display delay that worsens the operability of the machine. In contrast, in the spatial awareness support system 100 according to Embodiment 3, since the operation image is not a three-dimensional image, the computational processing power required of the control device 70 that processes the image is small, and the display delay of the image display device 50 is further reduced.
[0083] Next, the hardware configuration of the control device 70 will be described. Figure 21 is a diagram showing an example of a hardware configuration for realizing the control device of the spatial awareness support system according to Embodiment 1, Embodiment 2, and Embodiment 3. The control device 70 is realized as a computer system by a processing circuit that includes a processor 91 that performs various processes, a memory 92 which is the main memory, and a storage device 93 that stores information.
[0084] The processor 91 may be an arithmetic unit, microprocessor, microcomputer, CPU (Central Processing Unit), or DSP (Digital Signal Processor). The memory 92 can use volatile semiconductor memory such as RAM (Random Access Memory). The storage device 93 stores a program for rendering augmented reality images onto the control screen to make it easier for the operator to understand the position of the end effector 11. The storage device 93 can use magnetic disks, optical disks, magneto-optical disks, silicon disks, as well as non-volatile semiconductor memory such as ROM (Read Only Memory), flash memory, EPROM (Erasable Programmable Read Only Memory), and EEPROM® (Electrically Erasable Programmable Read Only Memory). The processor 91 reads the program stored in the storage device 93 into the memory 92 and executes it. The processor 91 reads the program stored in the storage device 93 into the memory 92 and executes it, thereby realizing the functions of the control device 70.
[0085] The configurations shown in the above embodiments are merely examples of the content, and can be combined with other known technologies. It is also possible to omit or modify parts of the configuration without departing from the gist of the invention.
[0086] 10 Machine, 10a Machine end, 10b 6-axis vertical articulated robot, 11 End effector, 11a Gripping gripper, 12 Depth sensor, 12a Depth camera, 30 Workpiece, 30a Target, 31 Dish, 40 Camera system, 40a Side camera, 40b Side camera drive unit, 50 Image display device, 51 Marker, 52 Guideline, 55 Transparent 3D figure, 55a, 55b, 55c, 56d, 56e, 56f Line segment, 56 Intersecting line, 56a First intersecting line, 56b Second intersecting line, 60 Operation input device, 70 Control device, 71 Camera control unit, 72 Display image generation unit, 73 Machine control unit, 80 Stage, 91 Processor, 92 Memory, 93 Storage device, 100 Spatial awareness support system, 111 Finger.
Claims
1. A spatial awareness support system comprising: a machine equipped with an end effector for performing work on a workpiece; a control device for controlling the machine; an operation input device connected to the control device for receiving input operations from an operator; a camera for photographing the area around the end effector; an image display device for displaying an operation image for the operator to refer to when performing the input operation; and a depth sensor provided on the machine for measuring the distance to an object on the extension of the center of the end effector, wherein the control device calculates the screen coordinates on the image captured by the camera of a first point indicating the surface of an object on the extension of the center of the end effector based on the camera parameters and lens distortion coefficient of the camera, the distance to an object on the extension of the center of the end effector measured by the depth sensor, and a command given to the machine via the input operation or a status value fed back from the machine, and generates the operation image by superimposing a preset shaped marker indicating that it is the surface of an object on the extension of the center of the end effector onto the screen coordinates of the first point on the image captured by the camera.
2. The spatial awareness support system according to claim 1, characterized in that the marker is circular, triangular, square, asterisk-shaped, star-shaped, letter-shaped, or crosshair-shaped.
3. The spatial awareness support system according to claim 1 or 2, characterized in that the display image generation unit draws the marker on the upper surface of an object directly below the workpiece while the end effector is gripping the workpiece.
4. The spatial awareness support system according to any one of claims 1 to 3, characterized in that the display image generation unit draws the marker with a predetermined fixed color.
5. The spatial awareness support system according to any one of claims 1 to 3, characterized in that the display image generation unit draws the marker in a color in which the hue difference between the marker and the pixels of screen coordinates on the image captured by the camera, which shows the surface of an object on the extension of the center of the end effector, is greater than a preset value.
6. The spatial awareness support system according to claim 4 or 5, characterized in that the display image generation unit draws the marker by setting the brightness difference between the marker and the pixels of screen coordinates on the image captured by the camera, which represent the surface of an object on the extension of the center of the end effector, to be greater than or equal to a preset value.
7. The spatial awareness support system according to any one of claims 1 to 6, characterized in that the camera is a side camera that photographs the end effector from the side.
8. The spatial awareness support system according to any one of claims 1 to 6, characterized in that the camera is a handheld camera installed on the end effector and takes overhead shots of the end effector.
9. The spatial awareness support system according to claim 8, characterized in that the hand camera is a depth camera that also functions as a depth sensor.
10. The spatial awareness support system according to any one of claims 1 to 6, characterized in that the camera includes a side camera that photographs the end effector from the side and a hand camera that is installed on the end effector and photographs the end effector from above.
11. The spatial awareness support system according to claim 10, characterized in that the hand camera is a depth camera that also functions as a depth sensor, the image display device is capable of switching between displaying the image captured by the side camera and the image captured by the hand camera, and the display image generation unit superimposes the marker onto the image captured by the side camera and the image captured by the hand camera.
12. The spatial awareness support system according to any one of claims 1 to 11, characterized in that the display image generation unit calculates the screen coordinates on the image captured by the camera of a second point predetermined near the tip of the end effector using the camera parameters, the lens distortion coefficient, and commands given to the machine, and superimposes a line segment connecting the screen coordinates of the second point and the screen coordinates of the first point onto the image captured by the camera.
13. The spatial awareness support system according to any one of claims 1 to 11, characterized in that the display image generation unit draws a transparent three-dimensional figure of a predetermined shape on the upper surface of the stage on which the workpiece is placed, on the image captured by the camera, based on the camera parameters of the camera, the lens distortion coefficient, and a command given to the machine or a status value fed back from the machine.
14. The spatial awareness support system according to claim 13, characterized in that the shape of the transparent three-dimensional figure is set based on the movable range of the end effector.
15. The spatial awareness support system according to 14, characterized in that the display image generation unit draws a first intersection line on the upper surface of the transparent three-dimensional figure, which includes two line segments parallel to the two directional axes of the horizontal movement of the end effector according to a command given to the machine, with a predetermined second point near the tip of the end effector as the intersection point, and the first intersection line projected parallel to the vertical movement direction of the end effector according to a command given to the machine in response to an up and down input operation to the operation input device, and draws a second intersection line on the bottom surface of the transparent three-dimensional figure.
16. The spatial awareness support system according to 12, characterized in that the display image generation unit draws a transparent three-dimensional figure, which has a shape set based on the movable range of the end effector and is placed on the stage surface on which the workpiece is placed, on the image captured by the camera, based on the camera parameters of the camera, the lens distortion coefficient, and a command given to the machine or a status value fed back from the machine.
17. A spatial awareness support system comprising: a machine equipped with an end effector for performing work on a workpiece; a control device for controlling the machine; an operation input device connected to the control device for receiving input operations from an operator; a camera for photographing the area around the end effector; and an image display device for displaying an operation image for the operator to refer to when performing the input operations, wherein the control device comprises a display image generation unit that generates the operation image by drawing a transparent three-dimensional figure of a predetermined shape, placed on the upper surface of a stage on which the workpiece is placed, onto the image captured by the camera, based on the camera parameters and lens distortion coefficient of the camera and a command given to the machine or a status value fed back from the machine.
18. The spatial awareness support system according to claim 17, characterized in that the shape of the transparent three-dimensional figure is set based on the movable range of the end effector.
19. The spatial awareness support system according to 16 or 17, characterized in that the display image generation unit draws a first intersection line on the upper surface of the transparent three-dimensional figure, which includes two line segments parallel to the two directional axes of the horizontal movement of the end effector according to a command given to the machine, with a second point predetermined around the end effector as the intersection point, and the first intersection line projected parallel to the vertical movement direction of the end effector according to a command given to the machine in response to an up-and-down input operation to the operation input device, and draws a second intersection line on the bottom surface of the transparent three-dimensional figure.
20. The spatial awareness support system according to claim 15 or 19, characterized in that the endpoints of the line segments constituting the first and second intersecting lines are located on the side surface of the transparent three-dimensional figure.
21. The spatial awareness support system according to any one of claims 15, 19, and 20, characterized in that the display image generation unit draws the transparent three-dimensional figure which includes a line segment connecting the intersection point of the first intersecting line and the intersection point of the second intersecting line.
22. The spatial awareness support system according to 21, characterized in that, if the upper surface of an object directly below the end effector is above the upper surface of the transparent three-dimensional figure, the display image generation unit extends the line segment connecting the intersection of the first intersecting lines and the intersection of the second intersecting lines to the upper surface of the object directly below the end effector and draws it.
23. The spatial awareness support system according to any one of claims 13 to 22, characterized in that the transparent three-dimensional figure has a bottom surface that shows the reachable range of the end effector on the lowest surface that the machine accesses during operation, a top surface that shows the reachable range of the end effector on a horizontal plane including the tip of the end effector, and a side surface that is a plane or curved surface composed of a set of points reachable by the end effector at the height between the bottom surface of the transparent three-dimensional figure and the top surface of the transparent three-dimensional figure.
24. The spatial awareness support system according to any one of claims 13 to 23, characterized in that the display image generation unit changes the drawing position and shape of the transparent three-dimensional figure in accordance with the change in the range of motion of the tip of the end effector due to a change in the posture of the end effector.