Tactile force presentation system

WO2026167898A1PCT designated stage Publication Date: 2026-08-13MITSUBISHI ELECTRIC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2026-08-13

Smart Images

  • Figure JP2025021163_13082026_PF_FP_ABST
    Figure JP2025021163_13082026_PF_FP_ABST
Patent Text Reader

Abstract

A tactile force presentation system according to the present invention comprises: a machine (10) that comprises an end effector (11) that has a plurality of fingers that perform work on a workpiece; a control device (70) that controls the machine (10); a camera that captures video of the surroundings of the end effector (11); a video display device (50) that displays operation video that is to be referenced when an operator performs input operations; and a means that acquires the force produced at the fingers. The control device (70) comprises a display image generation unit (72) that converts coordinate values for a three-dimensional model of the fingers on a world coordinate system to coordinate values for screen coordinates on the video captured by the camera and renders augmented reality video within the converted three-dimensional model in colors that correspond to the magnitude of the force produced at the fingers to generate the operation video.
Need to check novelty before this filing date? Find Prior Art

Description

Force and Tactile Cueing System

[0001] The present disclosure relates to a force and tactile cueing system that enables an operator who operates a machine while viewing an image displayed on an image display device to easily grasp the force generated in an end effector.

[0002] In recent years, a machine operation system has been put into practical use in which an operator operates a machine while viewing an image displayed on an image display device.

[0003] In a machine operation system in which an operator operates a machine while viewing an image projected on an image display device, an image of the machine, a workpiece, and the surrounding environment is captured by a camera, and the captured image is transmitted to and displayed on an image display device installed around the operator. The operator operates the machine while viewing the image displayed on the image display device. For example, when operating a robot, the operator operates the robot while viewing the image displayed on the image display device, and performs operations such as a pick-up operation of gripping and lifting a workpiece with an end effector and a place operation of placing the workpiece gripped by the end effector at another location. Generally, an end effector called a gripping gripper having a plurality of fingers is used for these operations.

[0004] When gripping a workpiece with an end effector or assembling a workpiece gripped by an end effector to another member, if the force for gripping the workpiece is too weak, the workpiece may fall, and if an excessive force is applied to the workpiece, the workpiece may be damaged. Since the gripping force of the end effector is applied to the workpiece through the fingers, the operator of the machine operation system is required to grasp the force generated in the fingers of the end effector in real time and adjust the gripping force.

[0005] Patent Document 1 discloses a machine operation system in which a light-emitting device that emits light according to force and touch is installed in an end effector. According to the technique disclosed in Patent Document 1, the operator can grasp the force generated in the fingers of the end effector based on the way the light-emitting device emits light in the operation video.

[0006] International Publication No. 2023 / 281648

[0007] However, the technology disclosed in Patent Document 1 has a problem in that, in outdoor environments with high ambient light levels, it is difficult for the operator to recognize the differences in how the light-emitting devices in the control video illuminate, and therefore difficult to grasp the force generated on the end effector's finger.

[0008] For example, between a beach in midsummer and a dimly lit room, there is a difference of 12 in exposure values, where each increment of 1 doubles the amount of light, meaning there is a difference of 2 to the power of 12 = 4096 times in brightness. Even limiting ourselves to indoors, the difference in exposure values ​​depending on the intensity of the lighting can be as much as 6, meaning there is a difference of 2 to the power of 6 = 64 times in brightness.

[0009] The human eye has a dynamic range of 10 to the power of 9 in response to brightness. In addition to the iris being able to adjust the amount of light by about 20 times with a reflection of about 0.2 seconds, the cone and rod cells adapt to light or dark over a period of about 20 minutes, causing the retina itself to undergo a sensitivity change of as much as 10 to the power of 6.

[0010] Cameras achieve proper exposure for subjects such as workpieces and work environments in response to differences in brightness by combining ISO sensitivity and ND (Neutral Density) filters, which are standards for photographic film established by the International Organization for Standardization (ISO), with aperture and shutter speed. For example, stopping down the aperture by 5 stops from F2.8 to F16, or changing the shutter speed from 1 / 30 second to 1 / 1000 second, means adjusting the exposure value by a difference of 5. In this way, when the camera adjusts the exposure value by an amount equivalent to 2 to the power of 6 to 2 to the power of 12 in response to differences in ambient brightness, the same magnification change in light intensity is required so that the light points of the light-emitting device appear the same in the camera image. That is, if a Class 1 laser with an output of 0.2 mW is to be used in a dimly lit room, an output of about 800 W is required at a magnification of 4096, and an output of about 10 W is required even at a magnification of 64.

[0011] While it is technically possible to emit lasers at these output levels, the problem arises of needing to set the emission wavelength to the visible range, and the fact that the output level falls within the range of laser processing machines, potentially damaging workpieces, equipment, and the work environment, as well as the need to ensure human safety. Therefore, it is practically impossible to compensate for the effects of ambient light by changing the emission intensity. Consequently, when using light-emitting devices, the brightness of the work environment in which force and tactile feedback can be provided is limited.

[0012] This disclosure is made in view of the above, and aims to provide a force-tactile feedback system that allows an operator of a machine operation system to intuitively grasp, through vision, the force generated on the finger of an end effector, regardless of the surrounding environment, without the use of additional physical feedback devices.

[0013] To solve the aforementioned problems and achieve the objectives, the force feedback system according to this disclosure comprises a machine equipped with an end effector having multiple fingers for working on a workpiece, a control device for controlling the machine, an operation input device connected to the control device for receiving operator input, a camera for photographing the area around the end effector, an image display device for displaying an operation image for the operator to refer to when performing an input operation, and means for acquiring the force generated on the fingers. The control device includes a display image generation unit that converts the coordinate values ​​of the three-dimensional model of the fingers in the world coordinate system, calculated based on the physical parameters of the fingers and commands given to the machine via input operations or status values ​​fed back from the machine, into screen coordinate values ​​on the image captured by the camera, based on the camera's orientation, the relative position between the camera and the machine, the camera's camera parameters, and the camera's lens distortion coefficient, and generates an operation image by drawing an augmented reality image in a color corresponding to the magnitude of the force generated on the fingers within the three-dimensional model after conversion to screen coordinate values ​​on the image captured by the camera.

[0014] According to this disclosure, it is possible to obtain a force-tactile feedback system that allows an operator of a mechanical operation system to intuitively grasp the force generated on the end effector's finger, regardless of the surrounding environment, through visual means, without the need for additional physical feedback devices.

[0015] A figure showing the configuration of the force-feedback system according to Embodiment 1. A figure showing the configuration of the control device for the force-feedback system according to Embodiment 1. A figure showing a first example of an operation video displayed on a video display device by the force-feedback system according to Embodiment 1. A figure showing a second example of an operation video displayed on a video display device by the force-feedback system according to Embodiment 1. A flowchart showing the operation flow of the force-feedback system according to Embodiment 1. A figure showing an example of a three-dimensional model used for mask generation by the force-feedback system according to Embodiment 1. A figure showing an example of a target image for hue histogram analysis of the force-feedback system according to Embodiment 1. A figure showing an example of hue histogram analysis of the force-feedback system according to Embodiment 1. A figure showing an example of an operation video for the force-feedback system according to Embodiment 1. A figure showing an example of the progression of mask drawing in the force-feedback system according to Embodiment 1. A figure showing an example of a three-dimensional model drawn with colors corresponding to the magnitude of force measured by the force-feedback sensor in the force-feedback system according to Embodiment 1. A figure showing an example of hardware configuration for realizing the control device for the force-feedback system according to Embodiment 1.

[0016] The force-feedback system according to an embodiment will be described in detail below with reference to the drawings.

[0017] Embodiment 1. Figure 1 is a diagram showing the configuration of the force-feedback system according to Embodiment 1. The force-feedback system 100 according to Embodiment 1 comprises a machine 10 that performs work on a workpiece 30 using an end effector 11 attached to the machine end 10a, an image display device 50 that displays an operation image which is an image for operating the machine 10, an operation input device 60 for operating the machine 10, and a control device 70 that controls the machine 10. The force-feedback system 100 also comprises a hand camera 12 that moves with the end effector 11 to photograph the fingers 111 from a certain direction, and a camera system 40 installed to the side of the work area of ​​the end effector 11, as a camera for photographing the area around the machine 10.

[0018] The machine 10, camera system 40, video display device 50, operation input device 60, and control device 70 constitute a machine operation system in which the operator operates the machine 10 while referring to the operation video displayed on the video display device 50, which captures images of the workpiece 30 and the surrounding environment in real time and displays them on the video display device 50. The camera system 40 includes a side camera 40a that captures the area around the end effector 11, and a side camera drive device 40b that changes the direction of the field of view of the side camera 40a in accordance with the movement of the end effector 11. The drive angle, which indicates the direction of the field of view of the side camera 40a as set by the side camera drive device 40b, is fed back to the camera control unit 71 as the camera drive status.

[0019] In the following description, machine 10 is assumed to be a 6-axis vertical articulated robot 10b, and the end effector 11 is assumed to be a gripping type gripper 11a equipped with multiple fingers 111. However, machine 10 may be something other than a 6-axis vertical articulated robot 10b, such as a SCARA robot. Machine 10 changes its posture and moves the end effector 11 by driving actuators installed corresponding to each joint axis according to operations input from the operator via the operation input device 60.

[0020] The hand camera 12 is provided with a flange that has the same shape as the machine end 10a, and an end effector 11 for the machine 10 can be attached to it. That is, the end effector 11 can be attached to the machine end 10a via the hand camera 12. Therefore, a general end effector 11 and quick changer for the machine 10 can be used in the force feedback system 100.

[0021] Here, the coordinate system of the hand camera 12 is the hand camera 12's local coordinate system, with the focal point of the camera lens as the origin, the depth direction along the optical axis as the Z direction, the downward direction of the camera image as the Y direction, and the rightward direction of the camera image as the X direction.

[0022] The hand camera 12 is centered such that a virtual line extending from the center of the end effector 11 along the sixth axis of the machine 10, which is a six-axis vertical articulated robot 10b, is included in the YZ plane of the hand camera 12. A force sensor 112 is installed on the finger 111, which is a means of acquiring the force generated on the finger 111 of the end effector 11. In systems where the drive current of the end effector 11 can be used as a force value, the force sensor 112 can be omitted.

[0023] Machine 10 performs tasks such as gripping and lifting a workpiece 30 on the stage 80 and moving it to another location. The operator operates the operation input device 60 while viewing the image captured by the side camera 40a and displayed on the image display device 50 to move the end effector 11 directly above the target point, and then lowers the end effector 11 positioned directly above the target point to grip the workpiece 30 or release the gripped workpiece 30. Here, when performing the task of gripping a workpiece 30 on the stage 80, the area directly above the target point is directly above the workpiece 30, and when performing the task of placing a gripped workpiece 30, the area directly above the target point is directly above the point where the workpiece 30 will be placed.

[0024] The operation input device 60 can be a pointing device such as a joystick, joypad, motion capture device, puppet, touch panel, or pointing stick, but it may also be a keyboard, an audio input device such as a microphone, or an eye-tracking input device.

[0025] The video display device 50 can be a general display device such as a monitor, tablet terminal, or smartphone terminal that can display still images and moving images.

[0026] Figure 2 shows the configuration of the control device for the force-feedback system according to Embodiment 1. The control device 70 includes a camera control unit 71 that controls the camera system 40 according to the operation input from the operator in the operation input device 60, thereby causing the field of view of the side camera 40a to follow the end effector 11. The control device 70 includes a display image generation unit 72 that superimposes augmented reality images onto the images captured by the side camera 40a or the hand camera 12, showing the force generated on the fingers 111 of the end effector 11 to make it easier for the operator of the machine operation system to grasp the gripping force, and displays it on the image display device 50 as an operation image, and a machine control unit 73 that controls the machine 10. The machine control unit 73 drives the end effector 11 by outputting commands to the machine 10, and also receives feedback of machine status and camera drive angle status from the machine 10 and the camera system 40. The machine control unit 73 also determines the magnitude of the force with which the end effector 11 grips the workpiece 30 based on the output of the force-feedback sensor 112.

[0027] The force-feedback system 100 according to Embodiment 1 superimposes an augmented reality image showing the magnitude of the force generated on the finger 111 of the end effector 11 onto an image captured by a side camera 40a or a hand camera 12. Figure 3 is a diagram showing a first example of an operation video displayed on an image display device by the force-feedback system according to Embodiment 1. Figure 4 is a diagram showing a second example of an operation video displayed on an image display device by the force-feedback system according to Embodiment 1. The first example is an operation video in which an augmented reality image showing the magnitude of the force generated on the finger 111 is superimposed onto an image captured by a side camera 40a. The second example is an operation video in which an augmented reality image showing the magnitude of the force generated on the finger 111 is superimposed onto an image captured by a hand camera 12. A mask 51 is superimposed on the finger 111 portion of the end effector 11 in the operation video, and the mask 51 is drawn in a color corresponding to the magnitude of the force generated when the end effector 11 grasps the workpiece 30. In other words, the mask 51 in the first and second examples is an augmented reality image drawn in a color corresponding to the magnitude of the force generated on the finger 111 of the end effector 11 as measured by the force tactile sensor 112.

[0028] Figure 5 is a flowchart showing the operation flow of the force tactile feedback system according to Embodiment 1. In step S11, the display image generation unit 72 calculates the world coordinates of the camera origin of the side camera 40a based on the relative position and orientation of the side camera 40a. The relative position of the camera is the relative position between the camera origin and the robot origin.

[0029] In step S12, the display image generation unit 72 calculates the coordinate values ​​in the world coordinate system of the vertices of the three-dimensional model 52 of the finger 111 based on the command given to the machine 10 in the world coordinate system. Figure 6 shows an example of a three-dimensional model used for generating a mask in the force-tactile presentation system according to Embodiment 1. The three-dimensional model 52 of the finger 111 is simplified into a rectangular parallelepiped shape that encloses the finger 111 based on the physical parameters of the finger 111, and the shape of the three-dimensional model 52 as seen from the camera is determined by the coordinate values ​​of the eight vertices. Due to this feature, the physical parameters can be dimensions shown in the orthographic drawing or measured values ​​using measuring instruments such as calipers. For this reason, in the force-tactile presentation system 100 according to Embodiment 1, CAD (Computer-Aided Design) data is not required, and the computational load is extremely low compared to when CAD data is used.

[0030] In step S13, the display image generation unit 72 converts the coordinate values ​​of the vertices of the 3D model 52 of the finger 111 in the world coordinate system to the coordinate values ​​in the camera local coordinate system of the side camera 40a or the hand camera 12.

[0031] In step S14, the display image generation unit 72 converts the coordinate values ​​of the three-dimensional model 52 of the finger 111 in the camera local coordinate system to the coordinate values ​​in the screen coordinate system based on the camera parameters and the lens distortion coefficient, and generates a mask 51 of the finger 111.

[0032] In this way, the display image generation unit 72 calculates the screen coordinates of the vertices of the 3D model 52 of the finger 111 of the end effector 11 in the video based on the commands given to the machine 10, the 3D model 52 of the finger 111, and the relative position and orientation of the camera, and generates a mask 51 of the finger 111.

[0033] Converting the world coordinates of a three-dimensional model 52, in which the fingers 111 of the end effector 11 are defined based on the command values ​​of the machine 10, which are world coordinates, into camera coordinates can be achieved by rotational and translational coordinate transformations.

[0034] In order to draw an augmented reality image from the camera coordinates of the 3D model 52 that defines the fingers 111 of the end effector 11, it is necessary to convert the camera local coordinates to 2D screen coordinates. Therefore, the display image generation unit 72 uses a correct distortion function that corrects lens distortion to convert the camera local coordinates to screen coordinates.

[0035] When converting camera local coordinates to screen coordinates, the display image generation unit 72 first performs a perspective projection transformation. In the perspective projection transformation, the left-right coordinate x and up-down coordinate y are normalized by the distance z from the camera origin as shown in equations (1) and (2) below.

[0036]

[0037] Next, the display image generation unit 72 performs distortion correction using distortion parameters, which are camera eigenvalues. For example, if the distortion parameters are k1, k2, p1, p2 [, k3 [, k4, k5, k6 [, s1, s2, s3, s4 [, τx, τy]]]], the display image generation unit 72 performs distortion correction as shown in equation (3) below. 2 Define the following and use x' and y' obtained from the perspective projection transformation to calculate radial distortion, tangential distortion, and thin prism distortion using equations (4) and (5) below, and perform distortion correction. Here, the first term of equations (4) and (5) below is radial distortion, the second and third terms of equations (4) and (5) below are tangential distortion, and the fourth term of equations (4) and (5) below is thin prism distortion.

[0038]

[0039] Next, the display image generation unit 72 performs viewport transformation. The display image generation unit 72 calculates the screen coordinates (u, v) using the distortion-corrected coordinates x'', y'' and the camera matrix. When the camera matrix is ​​represented by the following equation (6), the screen coordinates (u, v) are calculated by the following equations (7) and (8). Here, c x ,c y f is the screen coordinate of the camera origin, and x ,f yis a value obtained by dividing the focal length of each lens by the horizontal pixel pitch and the vertical pixel pitch beside the light receiving element.

[0040]

[0041] Through the above processing, the display image generation unit 72 can convert the camera local coordinates into two-dimensional screen coordinates.

[0042] When creating the mask 51, the display image generation unit 72 performs a histogram analysis of the hue in the pixels of the three-dimensional model 52 of the finger 111 of the end effector 11 to obtain the hue of the most frequent value.

[0043] The display image generation unit 72 generates a mask 51 in the color tracking region 53, which excludes pixels in the hue portion outside the specified range, centered on the most frequent hue, within the three-dimensional model 52 of the finger 111 of the end effector 11. As shown in Figure 6, the pixels in the portion occupied by the finger 111 within the three-dimensional model 52 have the most frequent hue, so the color tracking region 53 corresponds to the portion of the three-dimensional model 52 occupied by the finger 111. In addition to the hue portion outside the specified range, the display image generation unit 72 may also generate a mask 51 which excludes pixels with saturation and brightness outside the specified range. By excluding pixels with low saturation, low brightness, and near maximum brightness, mistracking can be avoided when achromatic objects such as black, white, gray (an intermediate color between these), and silver (a glossy color of these) fall within the tracking hue range depending on the color temperature of the lighting. For example, a white object will have a hue near blue in the white-balanced camera image under the illumination of a white light source. However, white objects with high reflectivity tend to be overexposed compared to subjects of other colors, resulting in a brightness close to the maximum value. Furthermore, even if their hue is blue, their saturation is significantly lower than that of blue objects. Therefore, by excluding approximately 1-2% at the upper end of brightness and 15-20% at the lower end of saturation, the effect of rejecting white subjects that are not to be tracked can be obtained. In addition, if the multiple fingers 111 of the end effector 11 overlap from the camera's line of sight, the overlapping portion is removed from the mask 51 of the finger 111 that is hidden in shadow. As shown in Figure 6, the rejection region 54 where the fingers 111 overlap from the camera's line of sight is removed from the mask 51 of the finger 111 that is hidden in shadow. This prevents the force tactile sensation of the finger 111 that is hidden in the overlapping portion from being superimposed. In this way, the display image generation unit 72 extracts the portion occupied by the finger 111 within the three-dimensional model 52 based on the mode value in the hue histogram analysis of the three-dimensional model 52.

[0044] A process of extracting a hue portion outside a specified range centered on the most frequent hue in pixels within the three-dimensional model 52 of the finger 111 of the end effector 11 will be described. FIG. 7 is a diagram showing an example of a target image for hue histogram analysis of the force and tactile presentation system according to the first embodiment. The black frame in FIG. 7 is the three-dimensional model 52 of the finger 111, and the color to be extracted is assumed to be the portion within the color tracking region 53 indicated by the diagonal hatching in FIG. 7. FIG. 8 is a diagram showing an example of hue histogram analysis of the force and tactile presentation system according to the first embodiment. The solid line in FIG. 8 is the result of single-hue histogram analysis, and the broken line in FIG. 8 is the result of window histogram analysis. The vertical axis indicates the frequency, and the horizontal axis indicates the hue. Here, "window histogram analysis" means performing histogram analysis by reconstructing the total value of the frequencies within a class as the frequency of the central value of the class in each class of the frequency distribution with a preset class width. That is, hue window histogram analysis means performing hue histogram analysis by reconstructing the total value of the hues within a class as the frequency of the hue of the central value of the class in each class of the hue frequency distribution with a preset class width. Also, the preset class width in window histogram analysis is called a "window".

[0045] For example, when the class width is ±10, the frequency of hue 100 in the window histogram is the sum of the frequencies of the single-hue histogram from hue 90 to hue 110, and the frequency of hue 99 in the window histogram is the sum of the frequencies of the single-hue histogram from hue 89 to hue 109. If there is a single peak at hue 50 but almost no frequencies in hues 40 to 49 and hues 51 to 60, the frequency of hue 50 in the window histogram is not much different from that of the single-hue histogram. On the other hand, the frequency of hue 100 in the window histogram when there is a hue gradient near hue 100 may be as high as about 20 times that of the single-hue histogram. Thus, by using window histogram analysis, it becomes possible to extract the hue of a large subject area without being affected by a hue with a small area but a large frequency in the single histogram.

[0046] In Figure 7, the color of the color tracking area 53 within the 3D model 52, shown by hatching, is a gradient, meaning that the color tracking area 53 actually contains multiple colors, not just a single color. On the other hand, the area outside the color tracking area 53, shown by dot hatching in Figure 7, is a single color. Therefore, even though the area outside the color tracking area 53 is smaller than the area inside the color tracking area 53, the single-hue histogram shows the strongest peak for the color outside the color tracking area 53. In Figure 8, the peak of the single-hue histogram for the color outside the color tracking area 53 appears in area A, enclosed by the solid circle on the left.

[0047] The peaks of each color in the color tracking region 53 are lower than the peaks of colors outside the color tracking region 53. However, because the color tracking region 53 has a gradient, the range of hues for which the frequency appears in the single hue histogram is wider in the color tracking region 53.

[0048] Therefore, the display image generation unit 72 performs window histogram analysis on the image targeted for hue histogram analysis. When window histogram analysis is performed, the frequency of the window in the part of the color tracking region 53 that has multiple color peaks, although the peaks of each color are lower than those of the color outside the color tracking region 53, is higher than the frequency of the window in the part of the color tracking region 53 that has only a single hue peak within the window.

[0049] In this way, by performing window histogram analysis, the display image generation unit 72 can extract the hue of the finger 111 portion of the end effector 11 without being affected by strong peaks of a single hue, even though the area is small, due to unevenness in hue caused by factors such as lighting.

[0050] The above process does not need to be performed every frame; for example, it can be performed once per second, or even less frequently. The process of optimizing the tracked color in real time is called "Active Color Tracking" (ACT). Conventional Color Tracking (CCT), which uses fixed colors, requires the "calibration of the tracked hue to match the color temperature of the lighting in each work environment, the camera's exposure, and the white balance," but ACT eliminates this need. Furthermore, ACT technology can immediately address the hue changes shown by fingers in the camera image due to uneven lighting within the work area, which could not be avoided even with calibration in CCT. While the camera's automatic exposure (AE) and automatic white balance (AWB) functions offer high convenience, the dynamic changes in exposure and white balance mean that the previously performed tracked hue calibration in CCT is disrupted. Therefore, it was essential to turn off the AE and AWB functions and manually adjust these camera properties. In systems using CCT, adjusting camera properties and performing tracking hue calibration is time-consuming, and it was difficult to maintain color tracking performance in machines where the end effector moves to a location with significantly different lighting conditions during operation. These problems can be solved with ACT. Furthermore, when supporting switching between multiple viewpoints such as side cameras and hand cameras, systems using CCT required adjusting the camera properties and calibrating the tracking hue for each camera. However, with ACT, there is no need to perform these adjustments or change color tracking parameters each time the camera is switched; simply switching the video automatically adjusts the tracking hue to match the finger color based on the color tone of each camera.

[0051] Generally, extracting tracked hue pixels using color tracking has the problem that the processing load increases in proportion to the area of ​​the search region, making it easy to mistrack objects of the same color other than the target object within the camera's field of view. The latter is particularly difficult to avoid when using a camera with a wide field of view for the sake of workability, and even with a narrow field of view, mistracking is likely to occur if the color of the workpiece or work stage is close to the color of the fingers. Therefore, in laboratories and exhibition halls, it was necessary to create workpiece colors, stage colors, and wall colors with a large hue difference from the finger color, and to install the system so that visitors whose clothing color cannot be specified are not captured in the direction that the camera tracking the movement range of the end effector during demonstration work is facing. In addition, it was necessary to tune the camera's facing range according to the setting of the work content.

[0052] Thus, while systems using CCT can perform stable tracking under specific conditions, they have difficulty adapting to unpredictable work environments where the surrounding environment's colors cannot be limited to a specific color, or where visitors whose clothing colors cannot be specified in the direction the camera is facing are captured in the image. On the other hand, the force-tactile presentation system 100 according to Embodiment 1 performs color tracking by limiting the search area to a small area within the camera image and the subject to be almost entirely occupied by the finger 111, and also uses ACT instead of CCT. This resolves the long-standing, difficult-to-solve problems that color tracking has faced, and enables stable operation in unpredictable environments with a low processing load.

[0053] The area of ​​the finger 111 of the end effector 11 on the control screen is theoretically maximized when the distance between the camera and the finger 111 of the end effector 11 is closest, and minimized when the distance is furthest. When the minimum value of the area of ​​the finger 111 of the end effector 11 on the control screen is defined as AreaMin and the maximum value as AreaMax, by rejecting areas in the color tracking processing range where the area of ​​the extracted region deviates from the range between AreaMin and AreaMax, it is possible to suppress to a certain extent the processing that superimposes augmented reality images that present force feedback on "large surfaces like walls" or "small surfaces like colored wiring" that are the same color as the finger 111 of the end effector 11. This type of processing is called "area filtering".

[0054] In step S15, the display image generation unit 72 draws an augmented reality image in the pixels within the mask 51, converting force feedback into color information. The display image generation unit 72 acquires pressure values ​​as force feedback and normalizes the acquired pressure values ​​between a preset maximum pressure and minimum pressure. The display image generation unit 72 generates a color-inverted image of the mask 51 portion of the operation video and performs alpha blending with a blending ratio corresponding to the normalized pressure value. By alpha blending the color-inverted image with a blending ratio of 0 to 1, a constant scale color change from very light white to pure white is achieved, regardless of the original finger color. Here, since masks 51 are generated independently for the multiple fingers 111 of the end effector 11, each mask 51 corresponding to each finger 111 is set with an independent color according to the force generated. Figure 9 is a diagram showing an example of an operation video of the force feedback presentation system according to Embodiment 1. Because there is a difference in the force applied to the two fingers 111 of the end effector 11, the masks 51 corresponding to each finger 111 are drawn in different colors. In this way, the display image generation unit 72 draws a mask 51, which is an augmented reality image corresponding to each of the multiple fingers 111, using a color corresponding to the magnitude of the force generated individually for each of the multiple fingers 111. Note that while identification of multiple fingers 111 is possible based on the 3D model 52, it cannot be achieved with CCT. Presenting the force sensation of multiple fingers 111 individually provides useful information in applications such as axis insertion.

[0055] Furthermore, the display image generation unit 72 adds a first color border to the mask 51 when the pressure value is below a preset minimum pressure. The display image generation unit 72 adds a second color border to the mask 51 when the pressure value is above a preset maximum pressure. The display image generation unit 72 does not add a border to the mask 51 when the pressure value is above a preset minimum pressure and below a preset maximum pressure. In this way, even if the color of the mask 51 is faint and difficult to perceive immediately after the finger 111 touches the workpiece 30 and a minute force sensation is generated, the operator can reliably recognize that the finger 111 has touched the workpiece 30, and the operator can reliably recognize the moment when the set maximum pressure is exceeded without the color of the mask 51 changing from pure white.

[0056] Figure 10 shows an example of the transition in mask drawing of the force feedback system according to Embodiment 1. It shows how the force applied to the end effector 11 gradually increases from state (A) to state (D). When the finger 111 touches the workpiece 30 and the force applied to the finger 111 is less than the preset minimum pressure, the mask 51 is drawn as an unfilled white frame, as shown in state (A). In Figure 10, the white line used to draw the mask 51 in state (A) is represented by a dashed line. When the force applied to the end effector 11 increases to or exceeds the preset minimum pressure, the white frame disappears, as shown in state (B), and the mask 51 is filled with white, corresponding to the acquired pressure value normalized between the preset maximum and minimum pressures. Since the filling is done by alpha blending of a color-inverted image, the original color of the finger 111 appears to show through more when the pressure is small, and the transparency of the white fill color decreases as the pressure increases. Thus, strictly speaking, mask 51 is not simply filled with white, but rather its whiteness is being altered. Additive mixing of the three primary colors of light, red, green, and blue, results in white. The opposite color of red (RGB = #FF0000) is cyan (RGB = #00FFFF), which is the color obtained by additively mixing blue (RGB = #0000FF) and green (RGB = #00FF00). Additive mixing of red and cyan (#FF0000 + #00FFFF) is equivalent to the additive mixing of red, green, and blue (#FF0000 + #00FF00 + #0000FF), resulting in white (RGB = #FFFFFF). Following this principle, which states that any color becomes white when two color-inverted images are combined in a 1:1 ratio, this process generates a change in whiteness by changing the normalization pressure as the mixing ratio. As shown in state (C), as the force applied to the end effector 11 increases, the mask 51 changes to a color close to pure white, and when the force applied to the finger 111 of the end effector 11 becomes equal to the maximum pressure, the mask 51 is filled with pure white. In this way, the display image generation unit 72 draws the mask 51 by corresponding the intensity of the white color to the magnitude of the force generated on the finger 111.If the force applied to the end effector 11 becomes even stronger and exceeds a preset maximum pressure, the mask 51 corresponding to the finger 111 where the force exceeding the maximum pressure occurred is filled with pure white, and a warning frame is added. The warning frame is drawn in a color other than white, such as red. In Figure 10, the dashed line simulates that in state (D), the warning frame is drawn in a color different from the white of the mask 51.

[0057] In the above description, the augmented reality image was an extracted mask 51 containing the portion of the finger 111 that the finger 111 occupies, rendered in a color corresponding to the force applied to the finger 111 as measured by the force sensor 112. However, the augmented reality image may also be rendered with the entire 3D model 52 in a color corresponding to the magnitude of the force measured by the force sensor 112. Figure 11 shows an example of a 3D model rendered in a color corresponding to the magnitude of the force measured by the force sensor in the force tactile presentation system according to Embodiment 1. When the end effector 11 is displayed as a completely achromatic image under a work light source, it is difficult to extract the portion of the finger 111 from within the 3D model 52 by hue tracking. However, by rendering the entire 3D model 52 in a color corresponding to the magnitude of the force measured by the force sensor 112, it becomes possible to present force tactile sensation to the operator without extracting the portion of the finger 111 within the 3D model 52. Furthermore, if the entire 3D model 52 is to be rendered as an augmented reality image with colors corresponding to the magnitude of the force measured by the force-feedback sensor 112, the 3D model 52 may be a rectangular parallelepiped inscribed within the finger 111. By making the 3D model 52 smaller than the finger 111, it is possible to avoid overlapping the augmented reality image with the workpiece.

[0058] The force-feedback system 100 according to Embodiment 1 can provide force-feedback to the operator without using a light-emitting device, allowing the operator to grasp the force generated on the fingers 111 of the end effector 11 regardless of the surrounding environment. Furthermore, since the operator needs to keep their eyes on the end effector 11 in the operation video during work, displaying force-feedback using numerical values ​​or bar graphs near the end effector 11 would obscure the end effector 11 or the workpiece 30, hindering the work. In order to display force-feedback without obscuring the end effector 11 or the workpiece 30, the positions of the end effector 11 and the workpiece 30 must be known. However, in actual work, the workpiece 30 is not always in the same position each time, and therefore the end effector 11 is not always in the same position when picking or placing. Therefore, in order to display force feedback without obscuring the end effector 11 or workpiece 30, it is necessary to display the force feedback using numerical values ​​or bar graphs at a position away from the end effector 11. However, in this case, eye movement is required to confirm the force feedback, which hinders the work. The force feedback presentation system 100 according to Embodiment 1 generates an operational image by superimposing an augmented reality image showing force feedback onto the portion of the three-dimensional model 52 of the finger 111 that the operator is focusing on during work, from the image captured by the camera. As a result, eye movement is not required when the operator confirms the force feedback. Therefore, the force feedback system 100 according to Embodiment 1 improves workability and the intuitiveness of grasping force feedback, and can reduce fatigue and the occurrence of work errors.

[0059] Furthermore, when visualizing the light emission pattern of a light-emitting device, the sampling theorem dictates that it is necessary to generate an operational image based on video captured at a frame rate of more than twice the maximum frequency of the light-emitting device's on / off switching. This requires high computational processing power in the video processing device, and also increases the likelihood of delays when transmitting the video to the display device. Conversely, since force and tactile sensations are expressed using light emission patterns at or below half the frame rate, the gradation of force and tactile sensation expression is reduced. For example, a frame rate of 30 fps results in a light emission period of 15 Hz, or 67 ms or longer. Moreover, in the case of light emission frequencies of 1 Hz, a drawback arises because the information is constructed along the time axis, meaning that the operator can only recognize the pattern one second after the occurrence of force and tactile sensation. Even by using high transmission bandwidth and processing power to achieve very high frame rates, this problem is not solved. Due to the characteristics of human vision, just as a fluorescent lamp blinking at 50 Hz appears as continuous light, at high light emission frequencies, afterimages cause the light to be perceived not as blinking, but as variations in the intensity of continuous light corresponding to the duty cycle. With the same duty cycle, changes in frequency cannot be recognized, and conversely, the insufficient light output in bright environments, which is already a problem even with continuous illumination, is exacerbated by the duty cycle. In contrast, the force-tactile feedback system 100 according to Embodiment 1 can present force-tactile feedback to the operator without using a light-emitting device, thus reducing the computational processing power required for the image processing device. Furthermore, delays are less likely to occur when transmitting images to the image display device, and since force-tactile information is not constructed in the time axis direction, it is possible to instantly recognize force-tactile feedback in a single frame of video.

[0060] Furthermore, the force-feedback system 100 according to Embodiment 1 does not require a light-emitting device, thus enabling cost reduction and avoiding the risk of light-emitting device failure. In addition, since the force-feedback system 100 according to Embodiment 1 does not require the integration of a light-emitting device into the end effector 11, existing end effectors 11 can be used, and design flexibility can be increased.

[0061] Furthermore, when using light-emitting devices such as liquid crystal panels, it is difficult to determine the tracking hue because the commanded hue of the emitted color does not match the hue on the camera image. For example, when the acquired hue of a camera image taken of a full-color LED (Light-Emitting Diode) emitted with commanded hues from 0 to 180 in a darkroom was measured, the difference between the commanded hue and the acquired hue was a maximum shift of 22, which corresponds to a phase of about 1 / 8 on the hue wheel, and commanded hues with a shift of 5 or more accounted for 1 / 4 of the total. Also, for commanded hues from 130 to 163, the acquired hue remained at 150 and did not change at all. Moreover, when using light-emitting devices, it is required that the light-emitting surface be uniform and that the emitted hue appear to be the same, providing a viewing angle of nearly 180 degrees, but such a configuration is difficult to realize. The force-tactile presentation system 100 according to Embodiment 1 does not use a light-emitting device, so the problems of the viewing angle of light-emitting devices and the mismatch between the commanded hue and the hue on the camera image do not occur.

[0062] Another method exists that uses passive markers made of retroreflective material to detect force and tactile sensation instead of light-emitting devices. However, if the passive markers are obscured by the rotation of the end effector 11, force and tactile sensation cannot be detected. To avoid this problem, it is necessary to provide a structure that supports the markers near the end effector 11, positioned above the rotation mechanism so that the passive markers do not rotate with the rotation of the end effector 11, and so that frame-out due to posture does not occur as much as possible. However, if such a structure is provided, placing the passive markers in a straight line with the camera will obstruct the end effector 11 and the workpiece 30, making it impossible to work. Therefore, when detecting force and tactile sensation using passive markers, it is difficult to find a physical arrangement solution that prevents interference with the main body of the 6-axis vertical articulated robot 10b, does not hinder the rotation of the end effector 11, avoids marker occlusion due to the position and orientation of the end effector, avoids framing out of the camera's field of view when the distance between the camera and the marker is short, and avoids the marker obstructing the work image. For this reason, it is considered virtually impossible to apply the method of detecting force and tactile sensation using passive markers to hand-held camera images with a narrow field of view that provide overhead images from close range. The size of the marker on the image is inversely proportional to the distance between the camera and the marker, and the effect of distance becomes more pronounced with cameras that have a wider field of view, which is advantageous for the work. For this reason, when using a camera with a wide field of view, problems tend to arise in recognizing the marker in response to changes in the distance between the camera and the end effector during work. Thus, there are many hurdles in the recognition of the marker itself when using markers to detect force and tactile sensation. Furthermore, with augmented reality markers capable of estimating posture and distance, there is a problem where the processing load becomes high enough to affect the frame rate, depending on the accuracy. The force-feedback presentation system 100 according to Embodiment 1 does not use markers, thus guaranteeing the completeness of the recognition operation.

[0063] In the force-feedback system 100 according to Embodiment 1, when the finger 111 of the end effector 11 touches the workpiece 30 and force is applied to the finger 111, an augmented reality image showing force-feedback begins to be superimposed on the image captured by the camera, and the outline of the mask 51 disappears. Therefore, based on the disappearance of the outline of the mask 51 in the operation image, even if the color of the mask 51 is faint and difficult to perceive immediately after the finger 111 touches the workpiece 30 and a minute force-feedback is generated, the operator can visually recognize that the finger 111 has touched the workpiece 30. Furthermore, when the force applied to the finger 111 exceeds a preset value, the finger 111 is outlined in the operation video. Therefore, even when the force applied to the finger 111 is close to the set maximum pressure and the change in the mask color is small, that is, when the color of the mask 51 only changes between white just before pure white and pure white, the operator can visually recognize that the force applied to the finger 111 has exceeded the preset value based on the outline applied to the finger 111 in the operation video. In addition, as the force applied to the finger 111 increases, the mask 51 is drawn closer to pure white in the operation video, allowing the operator to intuitively recognize the magnitude of the force applied to the finger 111 on a consistent scale, regardless of the environment, based on the color of the mask 51 in the operation video.

[0064] One approach to presenting force and tactile sensations involves using physical feedback devices worn by operators, and research and development are underway in this area. However, high-performance haptic devices that pursue realism are expensive, and physical feedback devices with moving parts are prone to malfunction. Furthermore, the approach using physical feedback devices worn by operators has two main challenges: if the force returned by the haptic device deviates from the operator's expected sensation, it can cause discomfort and reduce operability; and wearing the physical feedback device for extended periods increases the operator's physical burden. In addition, the approach using physical feedback devices worn by operators requires a special device on the operator's side, limiting machine operation to facilities where the device is installed and compromising ubiquitous accessibility. The force and tactile sensation presentation system 100 according to Embodiment 1 is configured without any additional devices compared to a normal machine operation system. In other words, since it is completed solely with a video display device essential for a machine operation system and does not require additional devices for physical feedback, it does not create problems such as reduced operability or increased physical burden on the operator. It has ubiquitous capabilities that allow machine operation from any location as long as there is a general display device capable of displaying images, including a highly portable smartphone. It should be noted that the five human senses are not processed independently in the brain, and the existence of a cross-modal phenomenon between vision and touch has become clear. As a result, in the force applied by humans, visual information plays a larger role than force and touch from direct physical feedback. Therefore, the force and touch presentation method using visual information in the force and touch presentation system 100 according to Embodiment 1 is rational in terms of human perceptual characteristics.

[0065] Next, the hardware configuration of the control device 70 will be described. Figure 12 is a diagram showing an example of a hardware configuration for realizing the control device of the force tactile feedback system according to Embodiment 1. The control device 70 is realized as a computer system by a processing circuit that includes a processor 91 for executing various processes, a memory 92 which is the main memory, and a storage device 93 for storing information.

[0066] The processor 91 may be an arithmetic unit, microprocessor, microcomputer, CPU (Central Processing Unit), or DSP (Digital Signal Processor). The memory 92 can use volatile semiconductor memory such as RAM (Random Access Memory). The storage device 93 stores a program for rendering augmented reality images onto the control screen to present the force generated on the finger 111 of the end effector 11 to the operator. The storage device 93 can use magnetic disks, optical disks, magneto-optical disks, silicon disks, as well as non-volatile semiconductor memory such as ROM (Read Only Memory), flash memory, EPROM (Erasable Programmable Read Only Memory), and EEPROM® (Electrically Erasable Programmable Read Only Memory). The processor 91 reads the program stored in the storage device 93 into the memory 92 and executes it. The processor 91 reads the program stored in the storage device 93 into the memory 92 and executes it, thereby realizing the functions of the control device 70.

[0067] The configurations shown in the above embodiments are merely examples of the content, and can be combined with other known technologies. It is also possible to omit or modify parts of the configuration without departing from the gist of the invention.

[0068] 10 Machine, 10a Machine end, 10b 6-axis vertical articulated robot, 11 End effector, 11a Gripping gripper, 12 Hand camera, 30 Workpiece, 40 Camera system, 40a Side camera, 40b Side camera drive unit, 50 Image display device, 51 Mask, 52 3D model, 53 Color tracking area, 54 Reject area, 60 Operation input device, 70 Control device, 71 Camera control unit, 72 Display image generation unit, 73 Machine control unit, 80 Stage, 91 Processor, 92 Memory, 93 Storage device, 100 Force tactile feedback system, 111 Finger, 112 Force tactile sensor.

Claims

1. A force-feedback system comprising: a machine equipped with an end effector having a plurality of fingers for performing operations on a workpiece; a control device for controlling the machine; an operation input device connected to the control device for receiving operator input operations; a camera for photographing the area around the end effector; an image display device for displaying an operation image for the operator to refer to when performing the input operations; and means for acquiring the force generated on the fingers, wherein the control device converts the coordinate values ​​of a three-dimensional model of the fingers in a world coordinate system, calculated based on the physical parameters of the fingers and commands given to the machine via the input operations or status values ​​fed back from the machine, into screen coordinate values ​​on the image captured by the camera, based on the camera's orientation, the relative position between the camera and the machine, the camera's camera parameters, and the camera's lens distortion coefficient; and a display image generation unit for generating the operation image by drawing an augmented reality image in a color corresponding to the magnitude of the force generated on the fingers within the three-dimensional model after conversion to screen coordinate values ​​on the image captured by the camera.

2. The force-feedback system according to claim 1, characterized in that the display image generation unit draws the augmented reality image corresponding to each of the plurality of fingers in a color corresponding to the magnitude of the force generated individually in each of the plurality of fingers.

3. The force-feedback presentation system according to claim 1 or 2, characterized in that the display image generation unit renders the augmented reality image by corresponding the intensity of the white color to the magnitude of the force applied to the finger.

4. The force-feedback presentation system according to any one of claims 1 to 3, characterized in that when the force generated in any of the plurality of fingers exceeds a preset maximum pressure, the display image generation unit outlines the augmented reality image corresponding to the finger in which the force exceeding the maximum pressure occurred with a warning display frame.

5. The force-feedback system according to any one of claims 1 to 4, characterized in that the augmented reality image is a mask indicating the portion occupied by the finger within the three-dimensional model.

6. The force-feedback presentation system according to claim 5, characterized in that the display image generation unit extracts the portion occupied by the finger within the three-dimensional model based on the mode value in the histogram analysis of the hue of the three-dimensional model.

7. The force-feedback presentation system according to claim 6, characterized in that the display image generation unit reconfigures the sum of the hue frequencies within each class of the hue frequency distribution within a preset class width as the hue frequency of the central value of the class, and performs a histogram analysis of the hue.

8. The force-feedback system according to any one of claims 1 to 4, characterized in that the augmented reality image is a three-dimensional model whose entire form is rendered in a color corresponding to the magnitude of the force applied to the finger.