Spatial awareness support system

The spatial understanding support system addresses the challenge of determining end effector position by superimposing markers and guidelines on real-time camera images, enabling accurate operation in dynamic environments without 3D data, reducing fatigue, and eliminating the need for specialized display devices or light-emitting devices.

JP7814651B1Active Publication Date: 2026-02-16MITSUBISHI ELECTRIC CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025571497
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2025-02-07
Filing Date
2025-06-03
Publication Date
2026-02-16
Estimated Expiration
2045-06-03

Smart Images

  • Figure 0007814651000001
    Figure 0007814651000001
  • Figure 0007814651000002
    Figure 0007814651000002
  • Figure 0007814651000003
    Figure 0007814651000003
Patent Text Reader

Abstract

The spatial understanding support system includes a machine (10) equipped with an end effector (11), a camera that captures images of the area around the end effector (11), an image display device (50) that displays the image captured by the camera, a control device (70) that controls the machine (10), and a depth sensor that measures the distance to an object on an extension of the center of the end effector (11). The control device (70) includes a display image generation unit (72) that calculates the screen coordinates on the image captured by the camera of a first point that indicates the surface of an object on an extension of the center of the end effector (11) based on the camera parameters and lens distortion coefficient of the camera, the distance to the object on an extension of the center of the end effector (11) measured by the depth sensor, and commands given to the machine (10), and superimposes a marker on the screen coordinates of the first point.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a space understanding support system that enables an operator operating a machine to easily understand a work space while viewing an image displayed on a video display device. [Background technology]

[0002] 2. Description of the Related Art In recent years, machine operation systems have been put into practical use in which an operator operates a machine while viewing an image displayed on a video display device.

[0003] In a machine operation system in which an operator operates a machine while viewing an image displayed on a video display device, images of the machine, workpiece, and surrounding environment are captured by a camera, and the captured images are transmitted to and displayed on a video display device installed around the operator. The operator operates the machine while viewing the image displayed on the video display device. For example, when operating a robot, the operator operates the robot while viewing the image displayed on the video display device, and performs tasks such as a pick-up operation in which an end effector grasps and lifts a workpiece, and a place operation in which the workpiece grasped by the end effector is placed in another location.

[0004] In order to perform a pick operation, in which an end effector grasps and lifts a workpiece, and a place operation, in which an end effector places a workpiece grasped by the end effector, it is necessary to move the end effector directly above the target point. However, because the image captured by the camera lacks information corresponding to depth, it is difficult to determine the current position of the end effector and the position directly above the target point from the image displayed on the video display device and to determine the direction in which the machine should be operated.

[0005] Patent Document 1 discloses a device that displays on an output device, within a virtual space defined by three-dimensional data of a structure, a target coordinate axis including a line segment extending vertically from a control point set on the device to be driven to the surface of the workpiece. According to the technology disclosed in Patent Document 1, a line segment extending vertically from a control point set on the device to be driven to the surface of the workpiece is visualized, allowing the operator to position the end effector with reference to the target coordinate axis. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Patent No. 6385627 Summary of the Invention [Problem to be solved by the invention]

[0007] However, the technology disclosed in Patent Document 1 does not include a technology for presenting support information on actual images captured by a distorted camera. Furthermore, the technology disclosed in Patent Document 1 is based on computer graphics operations aimed at performing repetitive tasks through automated driving after instruction, and the configuration for realizing the technology requires complete 3D data of the workpiece and the surrounding environment, including blind spots, to generate images to be displayed on an output device. Therefore, it is difficult to prepare this 3D data, and actual images of the unstable work environment must be captured in real time and displayed on a video display device. This makes it difficult to apply the technology to a general-purpose machine operation system where the work content can change from time to time.

[0008] Consider a machine operating a variety of products displayed on shelves in various stores. Because the machine's position becomes unstable due to movement, it is difficult to accurately obtain the three-dimensional coordinates of the shelves as seen by the machine. It is also difficult to retain all three-dimensional data for any product and use the three-dimensional data identified from the stored data based on video. Furthermore, it becomes impossible to respond to new products that do not have three-dimensional data due to the release of new products or the renewal of existing products. Furthermore, the situation on the shelves changes dynamically depending on customer access. It is unrealistic to assume that every shelf in every store is equipped with a three-dimensional scanner to provide real-time three-dimensional information to the machine. Even if a three-dimensional scanner were available, it would be impossible to avoid problems such as blind spots on the back of the workpieces or the workpieces on the lower shelves being in the shadow of the workpieces on the upper shelves. Thus, under general conditions, it was difficult to obtain three-dimensional data, an essential element of the technology disclosed in Patent Document 1.

[0009] The present disclosure has been made in consideration of the above, and aims to provide a spatial understanding support system that can be applied to machine operation systems in which the work content changes from time to time and can assist an operator in understanding the real-space position of an end effector. [Means for solving the problem]

[0010] To solve the above-mentioned problems and achieve the object, the spatial understanding assistance system according to the present disclosure includes a machine equipped with an end effector that performs an operation on a workpiece, a control device that controls the machine, an operation input device connected to the control device and accepting input operations from an operator, a camera that captures images of the area around the end effector, an image display device that displays an operation image that the operator refers to when performing input operations, and a depth sensor that is provided on the machine and measures the distance to an object on an extension of the center of the end effector. The control device includes a display image generation unit that calculates screen coordinates on the image captured by the camera of a first point that indicates the surface of the object on the extension of the center of the end effector based on camera parameters and lens distortion coefficients of the camera, the distance to the object on the extension of the center of the end effector measured by the depth sensor, and commands given to the machine via input operations or status values ​​fed back from the machine, and generates an operation image by superimposing a marker of a predetermined shape that indicates the surface of the object on the extension of the center of the end effector on the screen coordinates of the first point on the image captured by the camera.

[0011] Alternatively, in order to solve the above-mentioned problems and achieve the object, a spatial understanding assistance system according to the present disclosure includes a machine equipped with an end effector that performs work on a workpiece, a control device that controls the machine, an operation input device connected to the control device and accepting input operations from an operator, a camera that captures images of the periphery of the end effector, and an image display device that displays an operation image that the operator refers to when performing input operations. The control device includes a display image generation unit that generates the operation image by drawing a transparent three-dimensional figure of a predetermined shape that is placed on the top surface of a stage on which the workpiece is placed, on the image captured by the camera, based on the camera parameters and lens distortion coefficient of the camera, and commands given to the machine or status values ​​fed back from the machine. [Effects of the Invention]

[0012] According to the present disclosure, it is possible to obtain a spatial understanding support system that can be applied to a machine operation system in which the work content changes each time, and can assist an operator in understanding the position of the end effector in real space. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 is a diagram showing a configuration of a space recognition support system according to a first embodiment. [Figure 2] FIG. 1 is a diagram showing a configuration of a control device of a space recognition assistance system according to a first embodiment; [Figure 3] FIG. 10 is a diagram showing a first example of an operation image displayed on the image display device by the space recognition assistance system according to the first embodiment; [Figure 4] FIG. 10 is a diagram showing a second example of an operation image displayed on the image display device by the spatial recognition assistance system according to the first embodiment; [Figure 5] FIG. 10 is a diagram showing a third example of an operation image displayed on the image display device by the spatial recognition assistance system according to the first embodiment. [Figure 6] FIG. 10 is a diagram showing an example of marker transitions drawn in an operation image by the spatial recognition support system according to the first embodiment; [Figure 7] 1 is a flowchart showing a processing flow of a space recognition support system according to a first embodiment. [Figure 8] 1 is a flowchart showing a process flow for measuring the distance from the machine of the spatial recognition support system according to the first embodiment to the top surface of an object directly below the center of the end effector. [Figure 9] FIG. 1 is a diagram showing an example of markers drawn on a depth camera image of the space recognition support system according to the first embodiment; [Figure 10] FIG. 10 is a diagram showing the configuration of a space recognition support system according to a second embodiment. [Figure 11] FIG. 10 is a diagram showing the configuration of a control device of a space recognition assistance system according to a second embodiment. [Figure 12] FIG. 10 is a diagram showing a first example of an operation image displayed on the image display device by the space recognition assistance system according to the second embodiment; [Figure 13]10 is a flowchart showing a process flow of a space recognition support system according to a second embodiment. [Figure 14] FIG. 10 is a diagram showing a second example of an operation image displayed on the image display device by the space recognition assistance system according to the second embodiment. [Figure 15] FIG. 10 is a diagram showing a third example of an operation image displayed on the image display device by the space recognition assistance system according to the second embodiment. [Figure 16] FIG. 10 is a diagram showing an example of a problem solved by the space recognition support system according to the second embodiment. [Figure 17] FIG. 10 is a diagram showing another example of a problem solved by the space recognition support system according to the second embodiment. [Figure 18] FIG. 10 is a diagram showing an example of the effect of supporting an operator in understanding a space by an operation image displayed on a video display device by the space understanding system according to the second embodiment. [Figure 19] FIG. 10 is a diagram showing a first example of an operation image displayed on the image display device by the space recognition assistance system according to the third embodiment. [Figure 20] FIG. 10 is a diagram showing a second example of an operation image displayed on the image display device by the spatial recognition assistance system according to the third embodiment. [Figure 21] FIG. 1 is a diagram showing an example of a hardware configuration for realizing a control device of a space recognition support system according to Embodiments 1, 2, and 3. DETAILED DESCRIPTION OF THE INVENTION

[0014] A spatial recognition support system according to an embodiment will be described in detail below with reference to the accompanying drawings.

[0015] Embodiment 1 1 is a diagram showing the configuration of a spatial understanding support system according to embodiment 1. The spatial understanding support system 100 according to embodiment 1 includes a machine 10 that performs work on a workpiece 30 using an end effector 11 attached to a machine end 10a, and a depth sensor 12 that measures the distance to the target and outputs depth information. The spatial understanding support system 100 also includes a camera system 40 installed to the side of the working area of ​​the end effector 11, an image display device 50 that displays an operation image for operating the machine 10, an operation input device 60 that operates the machine 10, and a control device 70 that controls the machine 10.

[0016] The machine 10, camera system 40, video display device 50, operation input device 60, and control device 70 constitute a machine operation system that captures images of the workpiece 30 and the surrounding environment in real time and displays them on the video display device 50, and allows the operator to operate the machine 10 while referring to the operation image displayed on the video display device 50. The camera system 40 includes a side camera 40a that captures images of the area around the end effector 11, and a side camera drive device 40b that changes the orientation of the angle of view of the side camera 40a in accordance with the movement of the end effector 11. The drive angle, which indicates the orientation of the angle of view of the side camera 40a, is fed back to the camera control unit 71 as a camera drive status.

[0017] In the following description, the machine 10 is a six-axis vertical articulated robot 10b, and the end effector 11 is a grasping-type gripper 11a equipped with a plurality of fingers 111. Note that the machine 10 may be a robot other than the six-axis vertical articulated robot 10b, such as a SCARA robot, and the end effector 11 may be a device other than the grasping-type gripper 11a, such as a suction-type gripper. The machine 10 changes its posture and moves the end effector 11 by driving actuators installed corresponding to each joint axis in accordance with operations input by an operator via the operation input device 60.

[0018] The depth sensor 12 has a flange with the same shape as the machine end 10a, allowing the attachment of an end effector 11 for the machine 10. That is, the end effector 11 can be attached to the machine end 10a via the depth sensor 12. Therefore, the spatial recognition assistance system 100 can use an existing end effector 11 for the machine 10 or any general end effector 11 and a corresponding tool changer. The depth sensor 12 is installed at a position offset from the center of the end effector 11 with a tilt angle. Here, the center of the end effector 11 is the center of the tips of the multiple fingers 111 of the end effector 11, which is a grasping gripper 11a, on the extension of the sixth axis of the machine 10, which is a six-axis vertical articulated robot 10b. The offset amount and tilt angle are predetermined known design values. In the following description, the depth sensor 12 is assumed to be a depth camera 12a attached to the machine end 10a of the machine 10 so as to capture images of the periphery of the end effector 11. Note that the depth sensor 12 may be something other than the depth camera 12a.

[0019] Here, the coordinate system of the depth camera 12a is a local coordinate system of the depth camera 12a, with the focal point of the camera lens as the origin, the depth direction along the optical axis as the Z direction, the downward direction of the camera image as the Y direction, and the rightward direction of the camera image as the X direction.

[0020] The depth camera 12a is centered so that an imaginary line extending from the center of the end effector 11 along the sixth axis of the machine 10, which is a six-axis vertical articulated robot 10b, is included in the YZ plane of the depth camera 12a. Therefore, even if the depth camera 12a is installed with an offset and tilt, the position on the extension line of the center of the end effector 11 is guaranteed to exist in the pixel in the center column in the X direction of the image of the depth camera 12a.

[0021] The machine 10 performs tasks such as grasping and lifting the workpiece 30 on the stage 80 and moving it to another location. The operator operates the operation input device 60 while watching the image captured by the side camera 40a and displayed on the image display device 50 to move the end effector 11 to directly above the target point, and then lowers the end effector 11 positioned directly above the target point to grasp the workpiece 30 or release the grasped workpiece 30. Here, when performing the task of grasping the workpiece 30 on the stage 80, directly above the target point is directly above the workpiece 30, and when performing the task of placing the grasped workpiece 30, directly above the target point is directly above the point where the workpiece 30 will be placed.

[0022] The operation input device 60 can be a pointing device such as a joystick, joypad, motion capture, puppet, touch panel, or pointing stick, but it may also be a keyboard, a voice input device such as a microphone, or an eye-gaze input device.

[0023] The video display device 50 may be a general display device such as a monitor, tablet terminal, or smartphone terminal that is capable of displaying still images and moving images.

[0024] 2 is a diagram showing the configuration of a control device of the spatial recognition assistance system according to the first embodiment. The control device 70 includes a camera control unit 71 that controls the camera system 40 in accordance with an operation input by an operator via the operation input device 60, thereby causing the angle of view of the side camera 40a to track the end effector 11. The control device 70 also includes a display image generation unit 72 that superimposes an augmented reality image onto the image captured by the side camera 40a, making it easier to recognize the position of the end effector 11, based on the image captured by the side camera 40a and the depth information output by the depth sensor 12, and displays the superimposed image as an operation image on the image display device 50. The machine control unit 73 also includes a machine control unit 73 that controls the machine 10. The machine control unit 73 outputs commands to the machine 10 to drive the end effector 11, and receives feedback of the machine status and the camera drive angle status from the machine 10 and the camera system 40.

[0025] The spatial understanding assistance system 100 according to the first embodiment superimposes an augmented reality image of a point indicating the surface of an object on an extension of the center of the end effector 11 onto an image captured by the side camera 40a. Hereinafter, the augmented reality image of the point indicating the surface of the object on an extension of the center of the end effector 11 is referred to as a "marker." Furthermore, the augmented reality image of the line connecting the center of the end effector 11 and the marker is referred to as a "guideline." Furthermore, the augmented reality image combining the "marker" and the "guideline" is referred to as a "pointer." If a point indicating the surface of an object on an extension of the center of the end effector 11 is referred to as a first point and a predetermined second point on the periphery of the end effector 11 is referred to as a second point, the marker is drawn at the first point, and the guideline 52 is drawn as a line connecting the first point and the second point. Here, the second point is, for example, the center of the tip of the end effector 11. By setting the second point to the center of the tip of the end effector 11, it becomes easier to obtain spatial information about the positional relationship between the end effector 11 and objects in the working environment, including the workpiece 30, and whether or not the workpiece is between the fingers 111 required to grasp the workpiece 30. However, the second point is not limited to the center of the tip of the end effector 11.

[0026] 3 is a diagram showing a first example of an operation image displayed on the image display device by the spatial recognition support system according to embodiment 1. When the end effector 11 is not gripping the workpiece 30, a marker 51 is drawn on the upper surface of the workpiece 30 on an extension of the center of the end effector 11. A guideline 52 is drawn as a line connecting a second point, which is the center of the end effector 11, and the marker 51 drawn at the first point.

[0027] 4 is a diagram showing a second example of an operation image displayed on the image display device by the spatial recognition support system according to the first embodiment. In this state, the end effector 11 is not gripping the workpiece 30, but the workpiece 30 is between the fingers 111, and the top surface of the workpiece 30 is positioned above the tip of the end effector 11. In this state, the display image generation unit 72 regards the workpiece 30 as an object on an extension of the center of the end effector 11, and draws a marker 51 on the top surface of the workpiece 30. The guide line 52 is drawn as a line connecting the second point, which is the center of the end effector 11, and the marker 51 drawn at the first point.

[0028] 5 is a diagram showing a third example of an operation image displayed on the image display device by the spatial understanding support system according to the first embodiment. When the end effector 11 grips and lifts the workpiece 30, the marker 51 is drawn not on the top surface of the workpiece 30 but on the top surface of an object other than the workpiece 30 on an extension of the center of the end effector 11. The guide line 52 is drawn as a line connecting the second point, which is the center of the end effector 11, and the marker 51 drawn at the first point. This allows the spatial understanding support system 100 according to the first embodiment to present a target point to the operator when transporting the workpiece 30.

[0029] 6 is a diagram showing an example of the transition of a marker drawn in an operation image by the spatial recognition support system according to the first embodiment. In state (A), the end effector 11 is positioned directly above the plate 31, not directly above the workpiece 30, so the marker 51 is drawn on the top surface of the plate 31. When the end effector 11 moves directly above the workpiece 30, entering state (B), the marker 51 is drawn on the top surface of the workpiece 30. When the end effector 11 moves further and comes directly above a part of the workpiece 30 that is a different color from the part that was directly below the center of the end effector 11 in state (B), entering state (C), the color of the marker 51 changes. When the end effector 11 moves further and passes directly above the workpiece 30, entering state (D), the marker 51 is again drawn on the top surface of the plate 31 in the same color as in state (A).

[0030] 7 is a flowchart showing the flow of processing in the spatial understanding support system according to Embodiment 1. In step S11, the display image generation unit 72 measures the distance from the machine 10 to the top surface of an object on an extension of the center of the end effector 11.

[0031] The details of the processing of step S11 will be described. FIG. 8 is a flowchart showing the flow of processing for measuring the distance from the machine of the spatial recognition support system according to embodiment 1 to the top surface of an object on an extension line of the center of the end effector. In step S111, the display image generation unit 72 sequentially scans the central column of the depth matrix output by the depth camera 12a, and performs deprojection processing using the screen coordinates, the depth value recorded in the pixel of the screen coordinates, and the camera parameters of the depth camera 12a to obtain the depth camera coordinates of the object captured in the pixel, which are tilted and offset, in the depth camera 12a. Here, deprojection processing refers to performing a calculation to reverse-project the screen coordinates of the object image on the image onto a three-dimensional coordinate system in real space. In step S112, the display image generation unit 72 performs a process of rotating and moving the coordinate values ​​in the depth camera coordinates obtained in step S111 in the opposite direction by the same amount as the tilt angle of the depth camera 12a, and converts them into coordinate values ​​in the depth camera 12a that is oriented parallel to the end effector 11 at the installation position of the depth camera 12a.

[0032] In step S113, the display image generation unit 72 translates the origin of the camera coordinates that was rotated and moved in step S112 by the offset amount of the depth camera 12a toward the center of the end effector 11. By translating the origin of the camera coordinates by the same amount as the offset amount of the depth camera 12a, the depth camera coordinates after the translation become the same as the depth camera coordinates when the depth camera 12a is installed at the center of the end effector 11 and facing the same direction as the end effector 11. In the depth camera coordinates after the coordinate transformation by this rotation and translation, a scanned pixel whose X-direction component and Y-direction component are both zero is a pixel having depth information of an object on an extension line of the end effector 11, and the value of the Z-direction component of the converted depth camera coordinates is the distance from the camera focal plane to an object on an extension line of the center of the end effector 11.

[0033] Step S114 is a process for determining whether the loop from step S111 to step S113 is terminated. In step S114, the display image generation unit 72 calculates the square root of the sum of the squares of the X-direction component and the Y-direction component of the depth camera coordinate value of the scanned pixel after the coordinate transformation obtained in step S113, and calculates the distance from the z-axis, which is the coordinate axis in the Z direction, to the depth camera coordinate indicated by the scanned pixel. If the distance from the z-axis to the pixel is equal to or less than a preset tolerance, the display image generation unit 72 determines that it has successfully searched for a pixel whose X-direction component and Y-direction component are both zero, i.e., the depth camera coordinate of an object on the end effector extension line whose Z component indicates the distance to the object on the end effector extension line, and terminates the process. If the distance from the z-axis to the pixel is greater than the preset tolerance, the display image generation unit 72 returns to step S111 and repeats the processes from step S111 to S114 until it finds a pixel whose distance from the z-axis is equal to or less than the preset tolerance.

[0034] Hereinafter, the depth camera coordinates of the object on the extension line of the end effector after the coordinate conversion of the search pixel obtained in step S11 are referred to as the "three-dimensional coordinates of the drawing point" in the depth camera coordinate system. The "three-dimensional coordinates of the drawing point" in the depth camera coordinate system are depth camera coordinates equivalent to the coordinates indicated by the pixel at the focal position when the depth camera 12a is installed without a tilt angle or offset, that is, the pixel at the focal position when the depth camera 12a is installed facing directly downward at the center of the end effector 11, and are also an end effector-local coordinate system in which the X and Y components are approximately zero and the Z component indicates the distance to the object on the extension line of the end effector 11.

[0035] In step S12, the display image generation unit 72 converts the "three-dimensional coordinates of the drawing point" in the depth camera coordinate system obtained in step S11 when the depth camera 12a is placed at the center of the end effector 11 in the direction of the end effector 11 from the end effector local coordinate system to the world coordinate system using an offset value of the direction from the machine end of the depth camera 12a to the end effector 11 and the world coordinate of the machine end. The world coordinate of the machine end may be determined based on a command value output by the machine control unit 73 to the machine 10, or may be determined based on a status value fed back from the machine 10 to the machine control unit 73.

[0036] In step S13, the display image generation unit 72 converts the coordinate values ​​of the "three-dimensional coordinates of the drawing point" in the world coordinate system calculated in step S12 into coordinate values ​​in the camera coordinate system local to the side camera. For example, the display image generation unit 72 converts the three-dimensional coordinates of the drawing point in the world coordinate system into coordinate values ​​in the camera coordinate system local to the side camera using the relative position from the mechanical origin where the side camera is installed, the camera drive angle, etc. The camera system drive angle may be determined based on a command output by the camera control unit 71 to the side camera drive device 40b, or may be determined based on a camera drive status fed back to the camera control unit 71.

[0037] In step S14, the display image generation unit 72 performs a calculation to project the coordinate values ​​in the side camera's local camera coordinate system of the "three-dimensional coordinates of the drawing point" onto coordinate values ​​in the side camera's screen coordinate system. This process is generally referred to as "projection processing." For example, the display image generation unit 72 converts the coordinate values ​​in the side camera's local camera coordinate system into coordinate values ​​in the side camera's screen coordinate system based on the camera matrix and lens distortion coefficient of the side camera 40a. The pixel at the side camera's screen coordinates calculated in step S14 is the pixel to be drawn. Here, the "pixel to be drawn" is the pixel that corresponds to the first point on the top surface of the object on an extension of the center of the end effector 11.

[0038] In step S15, a marker 51 centered on the pixel to be drawn is drawn on the surface of the image of the object on an extension of the center of the end effector 11 at the screen coordinates of the side camera calculated in step S14, and a guideline 52 connecting the midpoint of the end effector 11 and the marker 51 is also drawn. The marker 51 is drawn in a predetermined arbitrary shape. For example, the marker 51 may be a circle, triangle, square, asterisk, star, character, or any other arbitrary shaped character illustration centered on the pixel to be drawn, or may be a crosshair shape with an intersection at the pixel to be drawn. When drawing a crosshair-shaped marker 51, the display image generation unit 72 calculates the coordinate values ​​in the screen coordinate system not only of a single point on an extension of the center of the end effector 11, but also of multiple points constituting the line segments of the crosshair. This allows the crosshair-shaped marker 51 to be drawn along a curved surface, even if the surface on which the marker 51 is drawn is curved.

[0039] When drawing the marker 51, the display image generation unit 72 determines the drawing color of the marker 51 based on the color of the pixel to be drawn. For example, the display image generation unit 72 converts the color of the pixel to be drawn from the RGB color space to the HSV color space and draws the marker 51 with high visibility using a color whose hue difference with the color of the pixel to be drawn is greater than a preset value. An example of a color whose hue difference with the color of the pixel to be drawn is greater than a preset value includes, but is not limited to, the complementary color of the pixel to be drawn. Note that if the pixel to be drawn is an achromatic color with no hue and a complementary color cannot be obtained, or if the pixel to be drawn has low saturation close to an achromatic color and therefore cannot provide contrast with the pixel to be drawn using a complementary color, the display image generation unit 72 may draw the marker 51 using a color whose brightness difference with the pixel to be drawn is greater than a preset value. Alternatively, the display image generation unit 72 may draw the marker 51 using a predetermined fixed color regardless of the color of the pixel to be drawn.

[0040] Since the distance from the location where the end effector 11 is installed to the location where the depth camera 12a is installed is a known jig design dimension, the distance to the object on an extension of the center of the end effector 11 can be converted into the reference coordinates of the machine 10. Since these coordinates can also be converted into the side camera's local camera coordinate system by translational and rotational coordinate transformation, if the coordinates of the surface of the object on an extension of the center of the end effector 11 are converted into screen coordinates using the camera parameters of the side camera 40a, it is possible to draw the above-mentioned marker 51 on the side camera image.

[0041] In addition, if the depth sensor 12 is a depth camera 12a, the RGB image of the depth camera 12a can be used as a bird's-eye view image from a short distance, with the marker 51 superimposed and displayed on the image display device 50 as a work image. That is, the depth camera 12a may be configured to capture both the periphery of the end effector 11 and the distance to an object on an extension of the center of the end effector 11. FIG. 9 is a diagram showing an example of a marker drawn on the depth camera image of the spatial recognition support system according to the first embodiment. By drawing the marker 51 on the surface of the workpiece 30 in the RGB image of the depth camera 12a, the operator can easily determine whether the end effector 11 is in a position where it can grasp the workpiece 30. While the RGB image of the depth camera 12a has high resolution due to the short distance to the workpiece 30, it does not have the same field of view as the side camera 40a, which allows for an overview of the entire workpiece 30. However, if the narrow field of view is acceptable, it is also possible to use a machine operation system that is configured with only the depth camera 12a, omitting the camera system 40. In such a configuration, a machine operation system using video with a spatial awareness support function can be constructed by simply inserting the depth sensor 12 unit between the end of the machine in the direct vision machine operation system and the end effector 11 in use.

[0042] The spatial understanding support system 100 according to the first embodiment can display, on the image display device 50, an operation image in which a marker 51 indicating the surface of an object on an extension of the center of the end effector 11 is drawn. This allows the operator operating the machine 10 to accurately recognize the position directly above which the end effector 11 is located. Furthermore, the spatial understanding support system 100 according to the first embodiment does not require the preparation of three-dimensional data of the workpiece 30 in advance in order to draw the marker 51 in the operation image, and therefore can be widely and immediately applied to existing machine operation systems that capture images of the workpiece 30 and the surrounding environment in real time and display them on the image display device 50.

[0043] Furthermore, the spatial understanding support system 100 according to the first embodiment does not require a special display device for displaying three-dimensional images, such as a head-mounted display or a 3D display. Therefore, the spatial understanding support system 100 according to the first embodiment has ubiquitous capabilities that allow it to be used in a wide range of operating environments where no special display devices are installed, and can easily allow the operator to grasp the position of the end effector 11 by using a general display device as the image display device 50. Although three-dimensional images tend to appear stereoscopically different from one another, the spatial understanding support system 100 according to the first embodiment does not display three-dimensional images on the image display device 50, eliminating individual differences in the degree to which the operator can grasp the position of the end effector 11. Furthermore, three-dimensional images tend to appear stereoscopically differently if there is a difference in the visual acuity between the two eyes. However, the spatial understanding support system 100 according to the first embodiment does not display three-dimensional images on the image display device 50, and therefore, even an operator with extremely poor visual acuity in one of his or her eyes can easily recognize the position of the end effector 11. Therefore, even when the spatial understanding support system 100 according to the first embodiment is applied to a machine operation system in which the machine is operated while looking directly at it, it is possible to realize barrier-free access by introducing camera images. Furthermore, because three-dimensional images utilize an optical illusion to display a pseudo three-dimensional shape, the viewer of the image is likely to become fatigued, but the spatial understanding support system 100 according to the first embodiment does not display three-dimensional images on the image display device 50. This reduces fatigue of the operator who must operate the machine 10 while looking at the operation image displayed on the image display device 50, enabling longer work hours.

[0044] Furthermore, since a three-dimensional image composed of left and right images handles twice the amount of data as a two-dimensional image, if the operation image is a three-dimensional image, the transmission time of the image data and the image processing time increase, requiring a high level of processing power for the device processing the image, and increasing the display delay that deteriorates the operability of the machine.In contrast, in the spatial awareness support system 100 according to the first embodiment, the operation image is not a three-dimensional image, so the processing power required for the control device 70 that processes the image is small, and the display delay of the image display device 50 is further reduced.

[0045] Furthermore, the spatial recognition support system 100 according to the first embodiment does not irradiate the workpiece 30 with laser light or the like, but rather makes it easier for the operator to recognize the position of the end effector 11 by using an augmented virtual reality image processed by software. Therefore, even when the workpiece 30 is gripped by the end effector 11, the light path is not blocked by the workpiece as shown in Fig. 5, making it easier for the operator to recognize the position of the end effector 11. Furthermore, the spatial recognition support system 100 according to the first embodiment makes it easier for the operator to recognize the position of the end effector 11 by using an augmented virtual reality image processed by software. Therefore, even when working in a bright work environment with a machine 10 installed outdoors, the operator can easily recognize the position of the end effector 11 without being affected by external light.

[0046] For example, between a midsummer beach and a dimly lit room, the exposure value, which indicates twice the amount of light for each increment of 1, varies by 12, or 2 to the power of 12 = 4096 times the brightness.Even just indoors, the exposure value can vary by 6 depending on the intensity of the lighting, or 2 to the power of 6 = 64 times the brightness.

[0047] The human eye has a dynamic range of brightness of 10 to the power of 9. In addition to the fact that the iris can adjust the amount of light by about 20 times with a reflection of about 0.2 seconds, over a period of about 20 minutes, the cone and rod cells adapt to light or darkness, causing the retina itself to change in sensitivity by as much as 10 to the power of 6.

[0048] Cameras adjust for differences in brightness by combining ISO speed (a standard for photographic film established by the International Organization for Standardization (ISO)) and neutral density (ND) filters with aperture and shutter speed to achieve the correct exposure for the subject, such as the work area or working environment. For example, stopping down the aperture five stops from f / 2.8 to f / 16, or changing the shutter speed from 1 / 30 to 1 / 1000 seconds, means adjusting the exposure value by a factor of five. When the camera adjusts the exposure value by a factor of 2 to 2 to accommodate differences in external brightness, a corresponding change in the light intensity of the light-emitting device is required to ensure that the light spot appears the same on the camera image. In other words, a Class 1 laser with an output of 0.2 mW would require approximately 800 W of output power at 4096x magnification and approximately 10 W at 64x magnification.

[0049] Although it is possible to emit a laser with these outputs, there are problems such as the need to make the emission wavelength in the visible range, and the need to ensure personal safety due to the output value being in the range of a laser processing machine, which could potentially damage the workpiece, equipment, and working environment. Therefore, compensating for the effects of external light by changing the emission intensity is virtually impossible. In this way, with a configuration using a light-emitting device, it is not possible to compensate for the effects of external light by changing the emission intensity.

[0050] Furthermore, if the surface of the workpiece 30 is mirror-finished, even if a laser beam or the like is irradiated onto the workpiece 30, it is reflected, making it difficult to display the marker 51 on the surface of the workpiece 30 by laser irradiation. In contrast, the spatial understanding support system 100 according to the first embodiment uses augmented virtual reality images generated by software processing to make it easier for the operator to grasp the position of the end effector 11, so it is possible to display the marker 51 on the surface of the workpiece 30 even if the surface of the workpiece 30 is mirror-finished. Furthermore, if the workpiece 30 is transparent, even if a laser beam or the like is irradiated onto the workpiece 30, it is transmitted through the workpiece 30, making it difficult to display the marker 51 on the surface of the workpiece 30 by laser irradiation. However, the spatial understanding support system 100 according to the first embodiment uses augmented virtual reality images generated by software processing to make it easier for the operator to grasp the position of the end effector 11, so it is possible to display the marker 51 on the surface of the workpiece 30 even if the surface of the workpiece 30 is transparent.

[0051] Furthermore, the spatial understanding support system 100 according to the first embodiment does not require a light-emitting device, which allows for cost reduction and avoids the risk of the light-emitting device breaking down. Furthermore, the spatial understanding support system 100 according to the first embodiment does not require a light-emitting device to be incorporated into the end effector 11, which allows for the use of existing end effectors 11 and increases the degree of freedom in design. Furthermore, since the spatial understanding support system 100 according to the first embodiment does not use a light-emitting device, there is no need to take into consideration personal safety, which is necessary when using laser light, and operation in places with pedestrian traffic is possible.

[0052] Embodiment 2 Fig. 10 is a diagram showing the configuration of a spatial understanding support system according to embodiment 2. The spatial understanding support system 100 according to embodiment 2 differs from the spatial understanding support system 100 according to embodiment 1 in that it does not include a depth camera 12a, which is the depth sensor 12. The configuration shown in Fig. 10 is a general configuration of a machine operation system using video, and therefore the spatial understanding support system 100 according to embodiment 2 can be realized without adding new hardware to any existing machine operation system.

[0053] FIG. 11 is a diagram illustrating the configuration of a control device of a spatial understanding support system according to a second embodiment. FIG. 12 is a diagram illustrating a first example of an operation image displayed on an image display device by the spatial understanding support system according to the second embodiment. The spatial understanding support system 100 according to the second embodiment superimposes a mixed reality image including a transparent three-dimensional figure 55 preset to an arbitrary shape and an intersection line 56 indicating the position of the end effector 11 within its movable range on an image captured by the side camera 40a. In the following description, it is assumed that the shape of the transparent three-dimensional figure 55 is set based on the movable range of the end effector 11. The transparent three-dimensional figure 55 may be the same as the shape indicating the movable range of the end effector 11, or may be slightly larger or smaller than the shape indicating the movable range of the end effector 11. The intersection line 56 is composed of multiple line segments whose intersection point is the point where an extension of the sixth axis of the machine 10, which indicates the orientation of the end effector 11, passes through the surface constituting the transparent three-dimensional figure 55.

[0054] The intersecting lines 56 include a first intersecting line 56a drawn on the top surface of the transparent three-dimensional figure 55 and a second intersecting line 56b drawn on the bottom surface of the transparent three-dimensional figure 55. The first intersecting line 56a is drawn by two or more intersecting line segments 56e, and the second intersecting line 56b is drawn by two or more intersecting line segments 56f. The line segments 56e constituting the first intersecting line 56a include two line segments parallel to the two directional axes of the horizontal movement of the end effector 11 in response to a command given to the machine 10. The second intersecting line 56b is drawn at a position where the first intersecting line 56a is projected onto the bottom surface of the transparent three-dimensional figure 55, parallel to the vertical movement direction of the end effector 11 in response to a command given to the machine 10 in response to an input operation in the vertical direction on the operation input device 60. Here, the first intersecting line 56a and the second intersecting line 56b are assumed to be crosshairs drawn by two intersecting line segments. The transparent three-dimensional figure 55 includes a line segment 55a forming the top surface, a line segment 55b forming the bottom surface, a line segment 55c connecting the top surface and the bottom surface, and a line segment 56d connecting the intersection points of the first intersection line 56a and the second intersection line 56b.

[0055] The intersection of the first intersection line 56a is a second point, which is the center of the end effector 11. When the machine 10 operates in a cylindrical coordinate system, a prism or a cylinder is suitable as the three-dimensional shape, and when the machine 10 operates in a Cartesian coordinate system, a rectangular parallelepiped is suitable, but is not limited to these. In the following explanation, it is assumed that the machine 10 operates in a Cartesian coordinate system, and an example is taken in which the transparent three-dimensional figure 55 is a rectangular parallelepiped whose bottom surface is the range of motion of the end effector 11. When the machine 10 operates in a Cartesian coordinate system, the display image generation unit 72 draws a first intersection line 56a on the top surface of the transparent three-dimensional figure 55, the first intersection line 56a being the intersection point and parallel to the two directional axes of the horizontal movement of the end effector 11 in response to a command given to the machine 10, and draws a second intersection line 56b projected from the first intersection line 56a on the bottom surface of the transparent three-dimensional figure 55, the second intersection line 56b being parallel to the direction of the up and down movement of the end effector 11 in response to a command given to the machine 10 by an up and down input operation on the operation input device 60. The side surface of the transparent three-dimensional figure 55 is drawn by a plane or a curved surface constituted by a set of points reachable by the end effector 11 at a height between the bottom surface of the transparent three-dimensional figure 55 and the top surface of the transparent three-dimensional figure 55.

[0056] FIG. 13 is a flowchart showing the processing flow of the spatial understanding support system according to the second embodiment. In step S21, the display image generation unit 72 defines the coordinates of the endpoints of the line segments constituting the transparent three-dimensional figure 55 in the world coordinate system, which is the coordinate system of the machine. When the horizontal shape of the transparent three-dimensional figure 55 is defined as the orthogonal movable range of the machine 10, the endpoint coordinates are as follows: The x and y coordinates of the vertices of the transparent three-dimensional figure 55 are the same as the horizontal coordinates of the ends of the machine's movable range. The z coordinates of the vertices of the transparent three-dimensional figure 55 on the bottom surface are the same as the z coordinates of the second intersection line 56b and the first intersection line 56a, which are the center of the end effector 11, on the top surface, and are the machine command value or status value indicating the position of the machine end minus the offset amount in the height direction between the machine end and the tip of the end effector. The y coordinates of the left and right line segments constituting the intersection line 56 are the y coordinates of the machine command value or status value. The x coordinates of the line segments before and after the intersection line 56 are the x coordinates of the machine command value or the status value.

[0057] By aligning the bottom surface of the transparent three-dimensional figure 55 with the movable range of the end effector 11, the operator can easily grasp the movable range of the end effector 11. However, from the perspective of spatial understanding alone, the bottom surface of the transparent three-dimensional figure 55 does not necessarily have to coincide with the movable range of the end effector 11. Furthermore, by aligning the bottom surface of the transparent three-dimensional figure 55 with the top surface of the stage 80, which is the lowest surface accessible by the machine 10 during work, and the top surface of the transparent three-dimensional figure 55 with the horizontal plane including the tip of the end effector 11, the bottom surface of the transparent three-dimensional figure 55 can be defined as a shape indicating the reachable range of the end effector 11 on the lowest surface accessible by the machine 10 during work, and the top surface of the transparent three-dimensional figure 55 can be defined as the reachable range of the end effector 11 on the horizontal plane including the tip of the end effector 11. This makes it easier for the operator to intuitively recognize the height of the end effector 11 from the stage 80.

[0058] In order to align the bottom surface of the transparent three-dimensional figure 55 with the top surface of the stage 80, "stage height," which is information on the relative height of the stage surface from the machine origin, is required. If the machine 10 is stationary, the stage height can be measured in advance and stored in the control device 70. On the other hand, if the machine 10 is mobile and the distance from the stage 80 is indefinite, the stage height must be determined in some way. For example, the stage height can be determined based on the command value when the end effector 11 is lowered until it touches the stage surface. Alternatively, the stage height can be determined using a depth camera 12a attached to the end effector 11. Alternatively, the stage height can be determined by installing a hand camera on the end effector 11 and installing active or passive markers on the stage 80, and analyzing the images of the active or passive markers captured by the hand camera. Alternatively, the stage height can be determined by optically reading the height from the floor, which is numerically displayed on the stage 80 in advance. It is also possible to store a plurality of pieces of stage height information in the control device 70 in advance, and select a value that is closest to the height of the stage 80 from the floor.

[0059] Furthermore, the display image generation unit 72 calculates the coordinate values ​​in the world coordinate system of both ends of a line segment whose both ends are the side surfaces of the transparent three-dimensional figure 55, with the position on the bottom surface of the transparent three-dimensional figure 55 on an extension of the center of the end effector 11 as the intersection point. For example, if the transparent three-dimensional figure 55 is a rectangular parallelepiped, the coordinate values ​​of both ends of at least two line segments are calculated. If the transparent three-dimensional figure 55 is a prism or a cylinder, the coordinate values ​​of both ends of at least one arc and one radial line segment are calculated. However, the coordinates of both ends of more line segments than the number shown in the example may be calculated. Note that by setting the both ends of the line segment as the intersection points with the side surfaces of the transparent three-dimensional figure 55, the operator can easily grasp the movable range of the end effector 11, but the both ends of the line segment do not necessarily have to be the intersection points with the side surfaces of the transparent three-dimensional figure 55.

[0060] In step S22, the display image generation unit 72 converts the coordinate values ​​of both end points of the line segments constituting the transparent three-dimensional figure 55 and the intersection line 56 defined in step S21 from the world coordinate system to a camera coordinate system local to the side camera, based on the physical distance of the installation position of the side camera 40a relative to the machine installation point, which is the world coordinate origin of the side camera 40a, and the attitude indicated by the drive angle of the side camera 40a. The drive angle of the side camera 40a may be a command value output by the camera control unit 71 of the control device 70, or may be a status value of the side camera drive device 40b fed back from the camera system 40.

[0061] In step S23, the display image generation unit 72 converts the coordinate values ​​of the endpoints from the camera coordinate system to the screen coordinate system. As in the first embodiment, the display image generation unit 72 converts the coordinate values ​​in the side camera local camera coordinate system into coordinate values ​​in the screen coordinates of the side camera 40a based on the camera matrix and lens distortion coefficient of the side camera 40a.

[0062] In step S24, the display image generating unit uses the coordinate values ​​of the screen coordinate system to draw a transparent three-dimensional figure 55 indicating the movable range of the end effector 11 and a cross line 56 indicating the position of the end effector 11 within the movable range.

[0063] If the target image on which the transparent three-dimensional figure 55 and the crossing line 56 are superimposed does not have lens distortion, the display image generation unit 72 draws the transparent three-dimensional figure 55 and the crossing line 56 using line segments that linearly connect both endpoints. On the other hand, if the target image on which the transparent three-dimensional figure 55 and the crossing line 56 are superimposed has lens distortion, it is preferable to draw the transparent three-dimensional figure 55 and the crossing line 56 using line segments whose camera coordinates are gradually incremented from one endpoint to the other. In other words, if the target image on which the transparent three-dimensional figure 55 and the crossing line 56 are superimposed has lens distortion, it is preferable to superimpose the transparent three-dimensional figure 55 and the crossing line 56 using curves in the target image on which the transparent three-dimensional figure 55 and the crossing line 56 are superimposed, even though they are straight lines in real space. In this way, the transparent three-dimensional figure 55 and the crossing line 56 are drawn with the same distortion as in real space, thereby improving the spatial presentation ability for the operator.

[0064] Note that the camera model used to convert camera coordinates to screen coordinates may have limitations on the view angle it can support. For example, a "camera model that performs perspective projection conversion normalized by distance" results in zero division at a 180° view angle where the distance is zero, i.e., at a point on the same plane as the focal plane. This results in increased camera calibration errors, such as distortion coefficients, at surrounding points, leading to normalized values ​​that are likely to deviate from reasonable values. Therefore, if the camera coordinates of the endpoints of line segments 55a, 55b, 56e, and 56f constituting the transparent three-dimensional figure 55 and the intersection line 56 are outside a cone whose apex is the camera origin of the side camera 40a and whose vertex is a preset view angle, the display image generation unit 72 performs saturation processing assuming that the endpoints of the drawing line segments are located on a cone whose apex is the preset view angle, and then converts the endpoints to screen coordinates to draw the transparent three-dimensional figure 55 and the intersection line 56. As an example, in the spatial understanding support system 100 according to the second embodiment, when the camera coordinates of the endpoints of the line segments 55a, 55b, 56e, and 56f constituting the transparent three-dimensional figure 55 and the intersection line 56 are located outside a cone having the camera origin of the side camera 40a as its vertex and an apex angle of 165 degrees, the display image generation unit 72 performs saturation processing assuming that the endpoints of the drawing line segments are located on the side surface of the cone having an apex angle of 165 degrees, and then converts them into screen coordinates to draw the transparent three-dimensional figure 55 and the intersection line 56.

[0065] Furthermore, the display image generation unit 72 draws each of the line segments 56e and 56f forming the intersection line 56 and the line segment 56d connecting the midpoints of the intersection line 56 in a line type or color different from the line segment 55a forming the top surface of the transparent three-dimensional figure 55, the line segment 55b forming the bottom surface of the transparent three-dimensional figure 55, and the line segment 55c connecting the top surface and the bottom surface of the transparent three-dimensional figure 55. For example, the display image generation unit 72 draws the line segment 55a forming the top surface of the transparent three-dimensional figure 55, the line segment 55b forming the bottom surface of the transparent three-dimensional figure 55, and the line segment 55c connecting the top surface and the bottom surface of the transparent three-dimensional figure 55 in black solid lines, the line segment 56e forming the intersection line 56 in a red dashed line, the line segment 56f forming the intersection line 56 in a black dashed line, and the line segment 56d connecting the midpoints of the intersection line 56 in a green solid line. This makes it easier for the operator to recognize the position of the end effector 11 even if some line segments are drawn overlapping each other while the system is in use.

[0066] 14 is a diagram showing a second example of an operation image displayed on the image display device by the spatial understanding support system according to the second embodiment. When the movable range of the end effector 11 changes due to tilting of the end effector 11, the rendering position and shape of the transparent three-dimensional figure 55 change in accordance with the tilt of the end effector 11. By referring to the operation image after tilting the end effector 11, the operator can recognize spatial information such as that the workpiece 30 is located approximately in the center of the movable range of the end effector 11, that the end effector 11 is located slightly higher and to the left of the workpiece 30, and that to position the end effector 11 directly above the workpiece 30, the end effector 11 should be moved approximately along the Y-axis in the negative direction by about the width of the workpiece 30, and then moved in the +X direction by about ¼ of the length of the workpiece 30.

[0067] In the example shown in Figure 14, the end effector 11 is tilted while positioned near the center of the transparent three-dimensional figure 55, so the second intersection line 56b is drawn inside the transparent three-dimensional figure 55. However, if the end effector 11 is tilted toward the outside of the transparent three-dimensional figure 55 at the periphery of the transparent three-dimensional figure 55, the second intersection line 56b may be drawn outside the transparent three-dimensional figure 55.

[0068] FIG. 15 is a diagram illustrating a third example of an operation image displayed on the image display device by the spatial understanding support system according to the second embodiment. In the machine posture illustrated in FIG. 15, the end effector 11 is positioned at a position that does not reach the end of the movable range of the machine when the end effector 11 is not tilted. However, as the end effector 11 tilts, the mixed reality transparent three-dimensional figure 55, which is drawn in accordance with the tilt of the end effector 11, changes, and the first intersection line 56a and the second intersection line 56b are drawn in a T-shape, similar to the end of the movable range. At this time, the control device 70 transmits the same information to the machine control unit 73, which changes the movable range of the machine in real time in conjunction with the tilt of the end effector 11. As a result, the machine control unit 73 no longer accepts movement commands to the machine, as shown in the mixed reality image, automatically avoiding collision between the end effector 11 and the machine itself.

[0069] 16 is a diagram showing an example of a problem solved by the spatial understanding support system according to the second embodiment. When the transparent three-dimensional figure 55 and the intersection line 56 are hidden, an operator who refers to the operation video may recognize that the end effector 11 is located at point A in FIG. 16. The operator who recognizes this will perform an operation to move the end effector 11 in the direction indicated by the dashed arrow in FIG. 16 in order to move the end effector 11 toward the target 30a, which is the workpiece 30 that is the target of the pick-up operation.

[0070] FIG. 17 is a diagram illustrating another example of a problem solved by the spatial recognition support system according to the second embodiment. The operation image illustrated in FIG. 17 and the operation image illustrated in FIG. 16 are the same image. If the transparent three-dimensional figure 55 and the intersection line 56 are hidden, an operator viewing the operation image may recognize that the end effector 11 is located on a workpiece 30 other than the target 30a, which is the workpiece 30 to be picked up. The operator, having recognized this, will perform an operation to move the end effector 11 in the direction indicated by the dashed arrow in FIG. 16 in order to move the end effector 11 toward the target 30a, which is the workpiece 30 to be picked up.

[0071] In this way, even when referring to the same operation video, there is a possibility that different operators will have different perceptions of where the end effector 11 is located. A skilled operator can improve the accuracy of estimating the position of the end effector 11 based on the size of the end effector 11 in the operation video, but even a skilled operator will have difficulty accurately recognizing the exact position of the end effector 11 and the direction of movement to the target position.

[0072] FIG. 18 is a diagram illustrating an example of the effect of supporting the operator in spatial understanding by the operation image displayed on the image display device by the spatial understanding system according to the second embodiment. The operation image shown in FIG. 18 is the same as the operation image shown in FIGS. 16 and 17. By displaying a mixed reality image including a transparent three-dimensional figure 55 and an intersection line 56 on the image display device 50 as the operation image, the operator can intuitively recognize the position of the end effector 11 in real space and where the end effector 11 will land if it is lowered straight down. In the example shown in FIG. 18, because the image is captured diagonally from the left, the operator can easily recognize that the forward direction of the machine 10 is the upper left direction of the image. Furthermore, the operator can easily recognize that the correct direction of movement to the target 30a, which is the workpiece 30 to be picked up, is the direction indicated by the dashed arrow in FIG. 18.

[0073] Furthermore, the transparent three-dimensional figure 55 and the intersection line 56 visualize "at what horizontal position" and "at what height from the stage" the end effector 11 is located within its movable range in real space. This allows the operator, referring to the operation video shown in Fig. 18, to intuitively recognize that the end effector is at "a height sufficient to allow safe horizontal movement without getting caught on workpieces on the stage," and that by moving the end effector 11 almost due left by an amount approximately equal to the length of the workpiece 30, it can be moved onto the workpiece 30, which is the target 30a.

[0074] The spatial understanding support system 100 according to the second embodiment superimposes a mixed reality image including a transparent three-dimensional figure 55 indicating the range of motion of the end effector 11 and an intersection line 56 indicating the position of the end effector 11 within the range of motion onto the image captured by the side camera 40a. Therefore, the horizontal position of the end effector 11 within the range of motion of the end effector 11 in real space and the height of the end effector 11 from the stage 80 are also visualized on the two-dimensional image based on the line of sight, size, and position of the intersection of the intersection lines of the bottom and top surfaces indicated by the transparent three-dimensional figure 55, as well as the position and length of the line segment 56d, thereby assisting the operator in understanding the position of the end effector 11 and the operation direction and amount required to perform the desired task. Furthermore, the spatial understanding support system 100 according to the second embodiment does not require the preparation of three-dimensional data of the workpiece 30 in advance in order to draw the transparent three-dimensional figure 55 and the intersecting line 56 in the operation image, and therefore can be applied to an existing machine operation system that needs to capture images of the workpiece 30 and the surrounding environment in any work environment in real time and display them on the image display device 50 without adding any new hardware.

[0075] Furthermore, the spatial understanding support system 100 according to the second embodiment does not require a special display device for displaying three-dimensional images, such as a head-mounted display or a 3D display. Therefore, the spatial understanding support system 100 according to the second embodiment has ubiquitous capabilities that allow it to be used in a wide range of operating environments where no special display device is required. A general display device can be used as the image display device 50 to enable the operator to easily grasp the position, operation direction, and operation amount of the end effector 11. Although three-dimensional images tend to appear stereoscopically different depending on the individual, the spatial understanding support system 100 according to the second embodiment does not display three-dimensional images on the image display device 50. This eliminates individual differences in the degree to which the operator can grasp the position, operation direction, and operation amount of the end effector 11. Furthermore, three-dimensional images are difficult to see stereoscopically if there is a difference in the visual acuity between the two eyes. However, the spatial understanding support system 100 according to the second embodiment does not display three-dimensional images on the image display device 50. Therefore, even an operator with extremely poor visual acuity in one of his or her eyes can easily recognize the position, operation direction, and operation amount of the end effector 11. Therefore, even when the spatial awareness support system 100 according to the second embodiment is applied to a machine operation system in which the operator operates the machine while looking directly at it, it can achieve barrier-free operation by introducing camera images. Furthermore, because three-dimensional images utilize the optical illusion of the human eye to display a pseudo-three-dimensional shape, the viewer of the image can easily become fatigued. However, the spatial awareness support system 100 according to the second embodiment does not display three-dimensional images on the image display device 50. This reduces fatigue of the operator who must operate the machine 10 while viewing the operation image displayed on the image display device 50, enabling longer work hours. Information on the "operation direction," which is essential for machine operation, cannot be presented to the operator even when three-dimensional images are used. However, the spatial awareness support system 100 according to the second embodiment has the unique advantage of being able to present information on the operation direction to the operator.

[0076] Furthermore, because a three-dimensional image composed of left and right images handles twice the amount of data as a two-dimensional image, if the operation image is a three-dimensional image, the transmission time and image processing time of the image data increase, requiring a high level of processing power in the device that processes the image, and increasing display delays that deteriorate the operability of the machine.In contrast, in the spatial awareness support system 100 according to the second embodiment, the operation image is not a three-dimensional image, so the processing power required of the control device 70 that processes the image is small, and the display delay of the image display device 50 is further reduced.

[0077] Embodiment 3 The configuration of the spatial understanding support system 100 according to the third embodiment is the same as that of the spatial understanding support system 100 according to the first embodiment. The spatial understanding support system 100 according to the third embodiment renders, in an operation image, a mixed reality image including both the pointer described in the first embodiment and the transparent three-dimensional figure 55 and the intersection line 56 described in the second embodiment.

[0078] 19 is a diagram showing a first example of an operation image displayed on the image display device by the spatial understanding support system according to the third embodiment. Because the end effector 11 is not gripping the workpiece 30, the operation image depicts a marker 51 on the upper surface of the stage 80, which is an object on an extension of the center of the end effector 11. The operation image also depicts a transparent three-dimensional figure 55 indicating the range of motion of the end effector 11 and an intersection line 56 indicating the position of the end effector 11. This allows the operator to easily position the end effector 11 directly above the workpiece 30 when performing a pick operation in which the end effector 11 grips and lifts the workpiece 30. A line segment 56d connecting the midpoints of the first intersection line 56a and the second intersection line 56b is drawn so as to overlap with a guide line 52 connecting the center of the end effector 11 and the marker 51. An operator looking at the operation image shown in Figure 19 can recognize, based on the transparent three-dimensional figure 55 in the operation image and the image of the machine 10, that the left side of the operation image is the front side of the machine 10, that the workpiece 30 is located in the fourth quadrant, and that in order to approach the workpiece 30, it is necessary to perform an operation to move the end effector 11 toward the front right by a distance slightly longer than the length of the workpiece 30.

[0079] 20 is a diagram showing a second example of an operation image displayed on the image display device by the spatial understanding support system according to the third embodiment. Since the end effector 11 is gripping the workpiece 30, the operation image depicts a marker 51 on the upper surface of the stage 80, which is an object on an extension of the center of the end effector 11. The operation image also depicts a transparent three-dimensional figure 55 indicating the range of motion of the end effector 11 and an intersection line 56 indicating the position of the end effector 11. Therefore, when performing a placing operation to place the workpiece 30 gripped by the end effector 11 at a target point, the operator can easily position the end effector 11 directly above the target point. An operator referring to the operation image shown in Figure 20 can recognize, based on the transparent three-dimensional figure 55 in the operation image and the image of the machine 10, that the left side of the operation image is the front side of the machine 10, that the target position for placing the workpiece 30 on the tray 31 is in the fourth quadrant, and that in order to approach the workpiece 30, it is necessary to perform an operation to move the end effector 11 toward the front right by a distance slightly longer than the length of the workpiece 30.

[0080] In this way, the spatial understanding support system 100 according to the third embodiment superimposes on the image captured by the side camera 40a a mixed reality image that visualizes the operator's own position in real space and the operation direction and operation amount to reach the target on a two-dimensional image using the marker 51, which is an augmented reality image showing the surface of an object on an extension of the center of the end effector 11, the transparent three-dimensional figure 55 showing the movable range of the end effector 11, the intersection line 56 showing the horizontal position of the end effector 11 within the movable range, and the line segment 56d showing the height of the end effector 11. Therefore, the spatial understanding support system 100 according to the third embodiment can more effectively support the operator in performing accurate operations even with images that do not have depth information. Furthermore, since the spatial understanding support system 100 according to the third embodiment does not require the preparation of three-dimensional data of the workpiece 30 in advance, it can be applied to existing machine operation systems that require the operator to capture images of the workpiece 30 and the surrounding environment in real time and display them on the image display device 50 during any operation by simply inserting the depth sensor 12 unit between the machine end 10a and the end effector 11 in use.

[0081] Furthermore, the spatial understanding support system 100 according to the third embodiment does not require a special display device for displaying three-dimensional images, such as a head-mounted display or a 3D display. Therefore, the spatial understanding support system 100 according to the third embodiment has ubiquitous capabilities that allow it to be used in a wide range of operating environments where no special display devices are required. A general display device can be used as the image display device 50 to enable the operator to easily grasp the position, operation direction, and operation amount of the end effector 11. Although three-dimensional images tend to appear stereoscopically different depending on the individual, the spatial understanding support system 100 according to the third embodiment does not display three-dimensional images on the image display device 50. This eliminates individual differences in the degree to which the operator can grasp the position, operation direction, and operation amount of the end effector 11. Furthermore, three-dimensional images are difficult to see stereoscopically if there is a difference in the visual acuity between the two eyes. However, the spatial understanding support system 100 according to the third embodiment does not display three-dimensional images on the image display device 50. Therefore, even an operator with extremely poor visual acuity in one of his or her eyes can easily recognize the position, operation direction, and operation amount of the end effector 11. Therefore, even when the spatial awareness support system 100 according to the third embodiment is applied to a machine operation system in which the operator operates the machine while looking directly at it, it can achieve barrier-free operation by introducing camera images. Furthermore, because three-dimensional images utilize the optical illusion of the human eye to display a pseudo-three-dimensional shape, the viewer of the image can easily become fatigued. However, the spatial awareness support system 100 according to the third embodiment does not display three-dimensional images on the image display device 50. This reduces fatigue of the operator who must operate the machine 10 while viewing the operation image displayed on the image display device 50, enabling longer work hours. Information on the "operation direction," which is essential for machine operation, cannot be presented to the operator even when three-dimensional images are used. However, the spatial awareness support system 100 according to the third embodiment has the unique advantage of being able to present information on the operation direction to the operator.

[0082] Furthermore, because a three-dimensional image composed of left and right images handles twice the amount of data as a two-dimensional image, if the operation image is a three-dimensional image, the transmission time of the image data and the image processing time increase, requiring a high level of processing power for the device that processes the image, and increasing display delays that deteriorate the operability of the machine.In contrast, in the spatial awareness support system 100 according to the third embodiment, the operation image is not a three-dimensional image, so the processing power required for the control device 70 that processes the image is small, and the display delay of the image display device 50 is further reduced.

[0083] Next, a description will be given of the hardware configuration of the control device 70. Fig. 21 is a diagram showing an example of a hardware configuration that realizes the control device of the spatial recognition assistance system according to the first, second and third embodiments. The control device 70 is realized as a computer system by a processing circuit including a processor 91 that executes various processes, a memory 92 that is a main memory, and a storage device 93 that stores information.

[0084] The processor 91 may be a computing unit such as an arithmetic unit, a microprocessor, a microcomputer, a CPU (Central Processing Unit), or a DSP (Digital Signal Processor). The memory 92 may be a volatile semiconductor memory such as a RAM (Random Access Memory). The storage device 93 stores a program for rendering an augmented reality image on an operation image to facilitate the operator's understanding of the position of the end effector 11. The storage device 93 may be a magnetic disk, an optical disk, a magneto-optical disk, a silicon disk, or a non-volatile semiconductor memory such as a ROM (Read Only Memory), a flash memory, an EPROM (Erasable Programmable Read Only Memory), or an EEPROM (Electrically Erasable Programmable Read Only Memory). The processor 91 reads the program stored in the storage device 93 into the memory 92 and executes it. The processor 91 reads the program stored in the storage device 93 into the memory 92 and executes it, thereby realizing the functions of the control device 70.

[0085] The configurations shown in the above embodiments are merely examples of the content, and may be combined with other known technologies, and parts of the configurations may be omitted or modified as long as they do not deviate from the gist of the invention. [Explanation of symbols]

[0086] 10 machine, 10a machine end, 10b 6-axis vertical articulated robot, 11 end effector, 11a grasping gripper, 12 depth sensor, 12a depth camera, 30 workpiece, 30a target, 31 plate, 40 camera system, 40a side camera, 40b side camera drive device, 50 image display device, 51 marker, 52 guideline, 55 transparent three-dimensional figure, 55a, 55b, 55c, 56d, 56e, 56f line segment, 56 intersection line, 56a first intersection line, 56b second intersection line, 60 operation input device, 70 control device, 71 camera control unit, 72 display image generation unit, 73 machine control unit, 80 stage, 91 processor, 92 memory, 93 storage device, 100 spatial understanding support system, 111 finger.

Claims

1. a machine equipped with an end effector that performs work on a workpiece; a control device for controlling the machine; an operation input device connected to the control device and configured to receive input operations from an operator; a camera that photographs the periphery of the end effector; a video display device that displays an operation video to be referred to when the operator performs the input operation; a depth sensor provided on the machine for measuring a distance to an object on an extension line of the center of the end effector; The control device calculates the screen coordinates on the image captured by the camera of a first point indicating the surface of the object on the extension of the center of the end effector based on the camera parameters and lens distortion coefficient of the camera, the distance to the object on the extension of the center of the end effector measured by the depth sensor, and the command given to the machine via the input operation or the status value fed back from the machine, and generates the operation image by superimposing a marker of a predetermined shape indicating that the first point is the surface of the object on the extension of the center of the end effector on the screen coordinates of the first point on the image captured by the camera.

2. The spatial recognition support system according to claim 1 , wherein the marker is in the shape of a circle, a triangle, a rectangle, an asterisk, a star, a letter, or a crosshair.

3. 2. The spatial recognition support system according to claim 1, wherein the display image generation unit draws the marker on an upper surface of an object directly below the workpiece while the end effector is gripping the workpiece.

4. The spatial understanding support system according to claim 1 , wherein the display image generating unit draws the marker in a preset fixed color.

5. The spatial understanding support system according to claim 1, characterized in that the display image generation unit draws the marker in a color whose hue difference with a pixel of a screen coordinate on an image captured by the camera showing the surface of an object on an extension of the center of the end effector is greater than a preset value.

6. The spatial understanding support system according to claim 4, characterized in that the display image generation unit draws the marker by setting a brightness difference between the marker and a pixel at a screen coordinate on an image captured by the camera showing the surface of an object on an extension of the center of the end effector to a preset value or more.

7. The spatial recognition support system according to claim 1 , wherein the camera is a side camera that photographs the end effector from the side.

8. 2. The spatial recognition support system according to claim 1, wherein the camera is a handheld camera that is installed on the end effector and captures an overhead image of the end effector.

9. The spatial recognition support system according to claim 8, wherein the hand camera is a depth camera that also functions as the depth sensor.

10. 2. The spatial understanding support system according to claim 1, wherein the camera includes a side camera that photographs the end effector from the side, and a hand camera that is installed on the end effector and photographs the end effector from above.

11. the hand camera is a depth camera that also serves as the depth sensor, the image display device is capable of switching between displaying the image captured by the side camera and the image captured by the hand camera, 11. The spatial awareness assistance system according to claim 10, wherein the display image generation unit superimposes the marker on the image captured by the side camera and the image captured by the hand camera.

12. The spatial understanding support system described in claim 1, characterized in that the display image generation unit calculates the screen coordinates on the image captured by the camera of a second point predetermined near the tip of the end effector using the camera parameters, the lens distortion coefficient, and instructions given to the machine, and superimposes a line segment connecting the screen coordinates of the second point and the screen coordinates of the first point on the image captured by the camera.

13. The spatial understanding support system according to claim 1, characterized in that the display image generation unit draws a transparent three-dimensional figure of a predetermined shape placed on the upper surface of a stage on which the workpiece is placed on the image captured by the camera based on the camera parameters and the lens distortion coefficient of the camera, and commands given to the machine or status values ​​fed back from the machine.

14. 14. The spatial understanding support system according to claim 13, wherein the shape of the transparent three-dimensional figure is set based on a movable range of the end effector.

15. The display image generation unit a first intersecting line including two line segments parallel to two directional axes of horizontal movement of the end effector according to a command given to the machine, the first intersecting line being an intersecting point at a second point predetermined near the tip of the end effector, on the upper surface of the transparent three-dimensional figure; The spatial understanding support system according to claim 14, characterized in that a second intersection line projected from the first intersection line is drawn on the bottom surface of the transparent three-dimensional figure in parallel to the vertical movement direction of the end effector in response to a command given to the machine in response to an input operation in the vertical direction on the operation input device.

16. The spatial understanding support system according to claim 12, characterized in that the display image generation unit draws, on the image captured by the camera, a transparent three-dimensional figure that is positioned on a stage surface on which the workpiece is placed and has a shape that is set based on the movable range of the end effector, based on the camera parameters and the lens distortion coefficient of the camera, and on commands given to the machine or status values ​​fed back from the machine.

17. a machine equipped with an end effector that performs work on a workpiece; a control device for controlling the machine; an operation input device connected to the control device and configured to receive input operations from an operator; a camera that photographs the periphery of the end effector; a video display device that displays an operation video to be referred to when the operator performs the input operation, The control device is a spatial understanding support system characterized in that it includes a display image generation unit that generates the operation image by drawing a transparent three-dimensional figure of a predetermined shape placed on the top surface of a stage on which the workpiece is placed on the image captured by the camera based on the camera parameters and lens distortion coefficient of the camera, and commands given to the machine or status values ​​fed back from the machine.

18. 18. The spatial understanding support system according to claim 17, wherein the shape of the transparent three-dimensional figure is set based on a movable range of the end effector.

19. The display image generation unit a first intersecting line including two line segments parallel to two directional axes of horizontal movement of the end effector according to a command given to the machine, the first intersecting line being an intersecting point at a second point predetermined on the periphery of the end effector, on the upper surface of the transparent three-dimensional figure; The spatial understanding support system according to claim 16, characterized in that a second intersection line projected from the first intersection line parallel to the vertical movement direction of the end effector in response to a command given to the machine in response to an input operation in the vertical direction on the operation input device is drawn on the bottom surface of the transparent three-dimensional figure.

20. 16. The spatial understanding support system according to claim 15, wherein end points of the line segments constituting the first intersection line and the second intersection line are located on side surfaces of the transparent three-dimensional figure.

21. The spatial understanding support system according to claim 15, wherein the display image generation unit draws the transparent three-dimensional figure including a line segment connecting an intersection of the first intersection line and an intersection of the second intersection line.

22. The spatial understanding support system according to claim 21, characterized in that, when the top surface of the object directly below the end effector is higher than the top surface of the transparent three-dimensional figure, the display image generation unit draws a line segment connecting the intersection of the first intersection line and the intersection of the second intersection line by extending it to the top surface of the object directly below the end effector.

23. The transparent three-dimensional figure is a bottom surface having a shape that indicates a reachable range of the end effector above a lowest surface accessed by the machine in operation; an upper surface having a shape that indicates a reachable range of the end effector on a horizontal plane including the tip of the end effector; The spatial understanding support system according to any one of claims 13 to 22, characterized in that the side surface is a plane or a curved surface formed by a set of points that can be reached by the end effector at a height between the bottom surface of the transparent three-dimensional figure and the top surface of the transparent three-dimensional figure.

24. The spatial understanding support system according to any one of claims 13 to 22, characterized in that the display image generation unit changes the drawing position and shape of the transparent three-dimensional figure in accordance with a change in the range of motion of the tip of the end effector due to a change in the posture of the end effector.

Citation Information

Patent Citations

  • Robot simulation image display system

    JP2010179403A

  • External input device, robot system, control method for robot system, control program, and recording medium

    JP2020075354A

  • Information processing device, method for processing information, robot system, method for manufacturing article using robot system, program, and recording medium

    JP2024048077A

  • Display control device, display control method and display control program

    JP6385627B1

  • Remote control device

    WO2022030047A1