Image display system, image display device, and image display method

The video display system addresses the challenge of maintaining accurate virtual-real space superimposition by using a feature point recognition and coordinate processing system to correct video content display in real-time, resulting in efficient and accurate alignment.

JP7691896B2Active Publication Date: 2025-06-12HITACHI LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2021152509
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-17
Publication Date
2025-06-12
Estimated Expiration
2041-09-17

AI Technical Summary

Technical Problem

Existing technologies face challenges in maintaining accurate superimposition of virtual and real spaces, particularly in head-mounted displays, due to delays in video tracking and positional drift caused by user movement.

Method used

A video display system comprising a feature point position recognition unit, a coordinate processing unit, and a video display processing unit, which recognizes feature points in captured images, determines a reference coordinate system, detects camera movement, and corrects the display position of video content in real-time to maintain accurate alignment.

Benefits of technology

The solution enables minimal delay in video tracking with respect to user movement and achieves highly accurate superimposition of virtual and real spaces, improving work efficiency and user comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007691896000001
    Figure 0007691896000001
  • Figure 0007691896000002
    Figure 0007691896000002
  • Figure 0007691896000003
    Figure 0007691896000003
Patent Text Reader

Abstract

To provide a technique that allows an accurate overlap of a real space and a virtual space without much delay of follow-up of a picture to motions of a user.SOLUTION: A picture display system according to one preferred aspect of the present invention includes a feature point position recognition unit, a coordinate processing unit, and a picture display processing unit. The feature point position recognition unit recognizes the position of a feature point of a target on the basis of an image taken by an imaging unit. The coordinate processing unit determines a reference coordinate system on the basis of the position of the feature point, detects a motion of the imaging unit on the basis of sensor information shown by the sensor unit, and corrects the display position of a picture on the basis of the motion of the imaging unit. The picture display unit displays a picture in the corrected display position of a picture when displaying a picture on the basis of the reference coordinate system.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for displaying images, and particularly to a technique for superimposing an image of a virtual space on a real space.

Background Art

[0002] Image display devices and image display systems such as head-mounted displays (HMDs) are known that present information to users by superimposing an image of a virtual space on a real space. In factories and the like, there are cases where work is performed while viewing content such as work processes, but it may be difficult to place a display or the like near the work target. In such a case, by using an image display device such as a head-mounted display, an operator can perform work while referring to work instructions and the like displayed on the image display device, and the work efficiency can be improved.

[0003] Patent Document 1 describes placing a virtual space at a position that is easy for the user to view and comfortably selecting and displaying an image that the user wants to see in a head-mounted display device.

[0004] Patent Document 2 describes a configuration for calculating the three-dimensional position of feature points included in a photographed image based on an image acquired by a camera.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] When a user performs work while viewing content such as work instructions, the positional relationship between the real space and the virtual space is important. If the positional relationship between the real space and the virtual space changes when the user rotates or moves their head, the user may feel discomfort or receive incorrect work instructions.

[0007] In the head-mounted display device described in Patent Document 1, the differential rotation amount of the head with respect to the user's waist is calculated, and the display content displayed in the virtual space is changed according to the differential rotation amount of the head, thereby aligning the positional relationship between the real space and the virtual space. However, since the display position of the content is determined based on the user's movement, the content may drift due to changes over time even if the head is not rotating.

[0008] Also, Patent Document 2 discloses an information processing device that performs self-position estimation and environmental map creation by SLAM (Simultaneous Localization and Mapping). However, since SLAM requires a large amount of time for calculation processing, when the real space is recognized using SLAM and the positional relationship between the real space and the virtual space is aligned, there is a problem that the video change of the video display device cannot follow the rotation of the user's head, in other words, a delay occurs.

[0009] The present invention has been made to solve the above problems, and an object of the present invention is to provide a technology capable of less delay in video tracking with respect to the movement of a user and highly accurate superimposition of the real space and the virtual space.

Means for Solving the Problems

[0010] In a preferred aspect of the present invention, a video display system includes a feature point position recognition unit, a coordinate processing unit, and a video display processing unit. The feature point position recognition unit recognizes the positions of the feature points of an object based on a captured image captured by an imaging unit. The coordinate processing unit determines a reference coordinate system based on the positions of the feature points, detects the movement of the imaging unit based on sensor information indicated by a sensor unit, corrects the display position of a video based on the movement of the imaging unit, and the video display processing unit displays the video at the corrected display position of the video when displaying the video based on the reference coordinate system. This is a video display system characterized by the above.

[0011] In another preferred aspect of the present invention, a video display device for displaying a video includes an imaging unit that captures an object to obtain a captured image, a sensor unit that detects the movement of the imaging unit, a feature point position recognition unit, a coordinate processing unit, a video display processing unit, and a video display unit that displays the video. The feature point position recognition unit recognizes the positions of the feature points of the object based on the captured image. The coordinate processing unit determines a reference coordinate system based on the positions of the feature points and generates correction information for the position of the video displayed by the video display unit based on the movement of the imaging unit. The video display processing unit determines the display position of the video based on the reference coordinate system and the correction information, and the video display unit displays the video at the determined display position. This is a video display device characterized by the above.

[0012] In another preferred aspect of the present invention, a video display method for displaying desired content includes a first step of recognizing the positions of the feature points of an object in an image captured by a camera, a second step of determining a reference coordinate system based on the feature points, a third step of detecting the movement of the camera, and a fourth step of correcting the display position of the content based on the movement of the camera when displaying the content based on the reference coordinate system. This is a video display method that executes the above steps.

Advantages of the Invention

[0013] According to the present invention, it is possible to provide a technology with little delay in video tracking with respect to the movement of a user and enabling highly accurate superimposition of the real space and the virtual space. Problems, configurations, and effects other than those described above will be clarified by the following description of the embodiments.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Mode for Carrying Out the Invention

[0015] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. The following description is for explaining one embodiment of the present invention and does not limit the scope of the present invention. Therefore, those skilled in the art can adopt embodiments in which each of these elements or all of these elements are replaced with equivalents, and these embodiments are also included in the scope of the present invention.

[0016] In the configuration of the embodiments described below, the same reference numerals are commonly used among different drawings for the same part or parts having the same or similar functions, and duplicate descriptions may be omitted.

[0017] When there are a plurality of elements having the same or similar functions, they may be described with the same reference numeral and different subscripts. However, when it is not necessary to distinguish the plurality of elements, the subscripts may be omitted in the description.

[0018] In this specification and the like, notations such as "first", "second", "third", etc. are attached to identify components, and do not necessarily limit the number, order, or content thereof. Also, numbers for identifying components are used for each context, and the numbers used in one context do not necessarily indicate the same configuration in other contexts. Further, it does not prevent a component identified by a certain number from also having the functions of a component identified by another number.

[0019] In the drawings and the like, the positions, sizes, shapes, ranges, etc. of each configuration shown may not represent the actual positions, sizes, shapes, ranges, etc. in order to facilitate understanding of the invention. For this reason, the present invention is not necessarily limited to the positions, sizes, shapes, ranges, etc. disclosed in the drawings and the like.

[0020] Publications, patents, and patent applications cited in this specification form part of the description of this specification as they are.

[0021] Components represented in the singular form in this specification shall include the plural form unless otherwise clearly indicated in the context.

[0022] A typical example of an embodiment is as follows. That is, the video display system includes an imaging unit, a sensor unit, a feature point position recognition unit, a coordinate processing unit, and a video display processing unit. The feature point position recognition unit recognizes the position of the feature points of the object based on the image captured by the imaging unit, and the coordinate processing unit determines a reference coordinate system based on the position of the feature points and corrects the relative displacement of the imaging unit with respect to the reference coordinate system based on the information indicated by the sensor unit to determine a corrected coordinate system or corrected coordinates. The video display processing unit has a configuration to display content at a predetermined position in the corrected coordinate system or to display content at a corrected predetermined position in the reference coordinate system. Here, the relative displacement is a concept including both a displacement in position and a displacement in angle.

Embodiment

[0023] In this embodiment, the video display device 100 is a device having a function of displaying video, such as a head-mounted display, a head-up display, a smartphone, a tablet, or the like.

[0024] FIG. 1 is a diagram showing an example of the functional blocks of the video display device 100 of this embodiment. The video display device 100 includes an imaging unit 101, a sensor unit 102, a feature point position recognition unit 103, a coordinate processing unit 104, a video display processing unit 105, a video display unit 106, and a controller 107.

[0025] The imaging unit 101 captures an image of the surrounding environment of the video display device 100. The sensor unit 102 detects the movement of the video display device 100. Here, the movement is a concept including both displacement of position and change in angle (rotation). The feature point position recognition unit 103 recognizes the position of feature points from the image captured by the imaging unit 101. The coordinate processing unit 104 performs processing of the coordinate system of the video displayed by the video display device 100. The video display processing unit 105 performs display processing of the video displayed by the video display device 100. The video display unit 106 displays the video. The controller 107 comprehensively controls the entire video display device 100. Details of each unit will be described later.

[0026] FIG. 2 is a diagram showing a configuration example of the hardware configuration of the video display device 100 of this embodiment. The video display device 100 includes a camera 111, a sensor 112, a display 113, a CPU (Central Processing Unit) 114, a RAM (Random Access Memory) 115, and a storage medium 116.

[0027] In the configuration example shown in FIG. 2, the video display device 100 includes the camera 111 as the imaging unit 101. The configuration of the camera 111 may be a known configuration described in Patent Document 2 or the like.

[0028] In addition, it is provided with a sensor 112 as the sensor unit 102. The sensor 112 detects the movement of the camera 111, for example, detects the movement of the user's head. The video display unit 106 is provided with a display 113. The configurations of the sensor unit 102 and the display 113 may be known configurations described in Patent Document 1 or the like.

[0029] The display method of the display 113 may be a transmissive display that can also receive visual information from the outside world, or a non-transmissive display that only projects video information. Hereinafter, for the purpose of visually recognizing an image by superimposing it on the real space, it will be described as a transmissive display. The display 113 may be provided with two for both eyes or one for one eye.

[0030] The CPU 114 executes a program stored in the storage medium 116 or the RAM 115. Specifically, when the CPU 114 executes the program, the functions of each unit such as the controller 107 of the video display device 100, the feature point position recognition unit 103, the coordinate processing unit 104, and the video display processing unit 105 are realized.

[0031] The storage medium 116 is a medium for storing the program executed by the CPU 114 and various parameters and data necessary for the execution.

[0032] The RAM 115 is a storage medium for storing the image and various information to be displayed on the display 113. The RAM 115 also functions as a temporary storage area for the programs and data used by the CPU 114. The video display device 100 may be configured to have a plurality of the CPU 114, the RAM 115, and the storage medium 116 respectively.

[0033] Note that the hardware configuration of the video display device 100 is not limited to the configuration shown in FIG. 2. For example, the CPU 114, the storage medium 116, and the RAM 115 may be provided separately from the video display device 100. In that case, the video display device 100 may be realized using a general-purpose computer (e.g., a server computer, a personal computer, a smartphone, etc.).

[0034] For example, the camera 111, the sensor 112, and the display 113 may be mounted on a wearable head-mounted display worn by the user, and the CPU 114, the RAM 115, and the storage medium 116 may be configured by a server at a remote location, and the head-mounted display and the server may be connected by a network.

[0035] Also, a plurality of computers may be connected by a network so that each computer shares the functions of each part of the video display device 100. On the other hand, one or more of the functions of the video display device 100 can also be realized using dedicated hardware.

[0036] FIG. 3 is a diagram showing an example of the video display device 100 of the present embodiment and the user 700. The video display device 100 shown in FIG. 3 is a head-mounted display (also referred to as smart glasses) that can be worn by the user 700 on his or her head.

[0037] FIG. 4 is a diagram showing an example of an implementation form of the video display device 100 of the present embodiment. The video display device 100 includes a camera 111 and a sensor 112. In this example, a gyro sensor is used as the sensor 112. The gyro sensor can detect the movement of the video display device 100.

[0038] The display 113 is built into the lens part of the glasses. In the case of a transmissive display, the user 700 can view the outside world through the display 113 and at the same time view the video projected by the video display unit 106 onto the display 113. The camera 111 captures an image in front of the user, and the camera image almost overlaps with the outside world that the user views. Next, the operation of the video display device 100 of this embodiment will be described in detail.

[0039] FIG. 5 is a diagram showing an example of the relationship between the real space and the coordinate system of this embodiment, which will be described in detail later.

[0040] FIG. 6 is a diagram showing an example of a flowchart related to content display executed by the video display device 100 of this embodiment.

[0041] First, the video display device 100 acquires a camera image (S101). The controller 107 of the video display device 100 sends an imaging command to the imaging unit 101, the imaging unit 101 images according to the command, and the imaging unit 101 transmits the captured image to the controller 107.

[0042] As shown in FIGS. 3 and 4, the camera 111 provided in the video display device 100 is attached to the video display device 100 so that when the user 700 wears the video display device 100 on the head, it can capture an image in front of the user 700. Therefore, the video display device 100 can acquire the situation in the real space in front of the user 700 as an image by acquiring the captured image of the camera 111. That is, an object 131 located in front of the user in the real space is captured in the image captured by the camera 111 (FIG. 5(a)). As the object 131, for example, a distribution board (operation panel) operated by the user 700 is assumed.

[0043] Next, the video display device 100 recognizes the positions of the feature points (S102). The controller 107 of the video display device 100 sends the imaging image and commands received from the imaging unit 101 to the feature point position recognition unit 103. The feature point position recognition unit 103 recognizes the positions of the feature points 132 from the imaging image and transmits the recognition results to the controller 107 (Fig. 5(a)). Although the details of the feature points 132 will be described later, within the imaging image received by the feature point position recognition unit 103 from the controller 107, the characteristic points are the feature points 132. Also, the position of the feature point 132 indicates at which position within the imaging image the feature point 132 is located.

[0044] Next, the video display device 100 determines the reference coordinate system (S103). The controller 107 of the video display device 100 sends the positions of the feature points and commands received from the feature point position recognition unit 103 to the coordinate processing unit 104. The coordinate processing unit 104 determines the reference coordinate system based on the positions of the feature points and transmits the coordinate system information to the controller 107.

[0045] In Fig. 5(a), with the position of the feature point 132 as a reference, the x-axis 133 is set in the downward direction of the paper surface, and the y-axis 134 is set in the rightward direction of the paper surface. However, this embodiment is not limited to this. It is also acceptable to set the origin of the x-axis 133 and y-axis 134 at a position shifted by a predetermined distance in a predetermined direction from the position of the feature point 132, or to set a coordinate system in which the two axes are not orthogonal, or to set a polar coordinate system. Hereinafter, it will be described assuming that, as shown in Fig. 5(a), with the position of the feature point 132 as a reference, the x-axis 133 is set in the downward direction of the paper surface and the y-axis 134 is set in the rightward direction of the paper surface.

[0046] In addition, the coordinate processing unit 104 pre-stores information regarding the positional relationship between the captured image of the imaging unit 101 and the video displayed on the video display unit 106. That is, based on the field of view, direction, and rotation angle of the image captured by the imaging unit 101 and the field of view, direction, and rotation angle of the video displayed on the video display unit 106, information regarding what positions and directions in the video displayed on the video display unit 106 correspond to predetermined positions and directions within the captured image of the imaging unit 101 is pre-stored. Thereby, the coordinates in the captured image and the coordinates in the video can be mutually converted. Such information regarding the positional relationship can be determined in advance based on the specifications of the imaging unit 101 and the video display unit 106. The information for converting the first coordinate system in the captured image to the second coordinate system in the video is sometimes referred to as coordinate conversion information for convenience.

[0047] Based on the information regarding the positional relationship, the coordinate processing unit 104 uses the coordinate conversion information to convert the first reference coordinate system set in the captured image to the second reference coordinate system in the video displayed on the video display unit 106. The coordinate processing unit 104 transmits information regarding the second reference coordinate system in the video displayed on the video display unit 106 to the controller 107. Specifically, for example, it is information regarding the origin and direction of the second reference coordinate system in the video displayed on the video display unit 106.

[0048] In the above description, the coordinate processing unit 104 was described as converting the first reference coordinate system in the captured image to the second reference coordinate system in the video displayed on the video display unit 106 after setting the first reference coordinate system in the captured image. However, the present embodiment is not limited to this. The coordinate processing unit 104 uses the coordinate conversion information to calculate what position in the video displayed on the video display unit 106 corresponds to the position of the feature point 132, and based on this position, the second reference coordinate system may be set in the video displayed on the video display unit 106. Thereby, it is possible to reduce the processing in the coordinate processing unit 104. Since the recognition of the position of the feature point 132 and the determination of the reference coordinate system do not require the use of the entire captured image to perform spatial recognition of the entire real space, it is possible to perform high-speed position recognition and reference coordinate system determination.

[0049] As described above, the video display unit 106 converts the first reference coordinate system set in the captured image into a second reference coordinate system in the video displayed. By positioning the content (object) to be displayed in the captured image based on the converted coordinate system, the positional relationship between the object and the content in the captured image can be determined. As a result, the positional relationship between the object in the real space and the content in the virtual space can be determined. However, this is premised on the fact that the field of view of the camera 111 is fixed.

[0050] Next, the video display device 100 repeatedly determines whether a predetermined time has elapsed until the predetermined time elapses (S104). In other words, it waits until the predetermined time elapses. Specifically, for example, the controller 107 measures the elapsed time since receiving the coordinate system information from the coordinate processing unit 104, determines at predetermined intervals whether the elapsed time exceeds a predetermined time, and repeats until the elapsed time exceeds the predetermined time.

[0051] As another method, the controller 107 has a loop counter, determines whether the loop counter exceeds a predetermined value, increases the value of the loop counter by a predetermined value for each determination, and repeats until the loop counter exceeds the predetermined value.

[0052] As another method, the controller 107 acquires the time, determines at predetermined intervals whether the acquired time exceeds the time when a predetermined time has elapsed from a predetermined time, and repeats until the acquired time exceeds the time when a predetermined time has elapsed from the predetermined time.

[0053] Generally, since the user's head and body move as time passes, the camera 111 moves. At the time of step S103 when the first and second reference coordinate systems are determined, the positional relationship between the object in the captured image and the object in the real space from the camera viewpoint corresponds. That is, the first reference coordinate system defined by the feature points of the object in the captured image is equivalent to the coordinate system defined by the feature points of the object in the real space from the camera viewpoint. Therefore, if the display area is displayed according to the second reference coordinate system determined based on the first reference coordinate system, the content can be displayed at a desired position in the real space.

[0054] However, when the field of view of the camera 111 moves, the positional relationship between the object in the captured image at the time of step S103 and the object in the real space at the current camera viewpoint no longer corresponds. That is, the first reference coordinate system determined in step S103 is no longer equivalent to the coordinate system defined by the feature points of the object in the real space. Therefore, since the positional relationship between the object in the real space and the display area based on the second reference coordinate system is shifted, correction is required.

[0055] For example, in order to make the video follow the movement of the head, the correction is performed, for example, multiple times per second. Specifically, it is about 30 times / second or 60 times / second. The more, the better from the perspective of followability. However, if it cannot be made fast due to the processing load, several times / second may be acceptable depending on the application. Therefore, the predetermined time of S104 (and S110) is determined in consideration of these.

[0056] For the correction of the reference coordinate system, the video display device 100 acquires sensor information (S105). The controller 107 of the video display device 100 sends a command for acquiring sensor information to the sensor unit 102. The sensor unit 102 acquires sensor information according to the command, and the sensor unit 102 transmits the sensed sensor information to the controller 107.

[0057] In an example of the video display device 100 shown in FIG. 4, the video display device 100 includes a gyro sensor as the sensor 112, and the gyro sensor measures the angular velocity of the video display device 100. The gyro sensor transmits the measured angular velocity data to the controller 107.

[0058] In an example of the video display device 100 shown in FIG. 4, it is assumed that the video display device 100 includes a gyro sensor as the sensor 112, but the present embodiment is not limited to this. For example, the video display device 100 may include an acceleration sensor as the sensor 112, and the acceleration sensor measures the acceleration of the video display device 100, and the acceleration sensor transmits the measured acceleration data to the controller 107.

[0059] Next, the video display device 100 recognizes the relative displacement amount and determines the correction coordinate system (S106). The controller 107 of the video display device 100 sends the sensor information and commands received from the sensor unit 102 to the coordinate processing unit 104. The coordinate processing unit 104 recognizes the relative displacement amount 135 and determines the correction coordinate system, and transmits the coordinate system information to the controller 107.

[0060] In one example, the sensor information in S105 is time-series data after S103 (for example, time-series data on how the angular velocity has changed after S103). Alternatively, assuming that the value of the sensor information is substantially constant for a short period of time, the data at the timing of S105 may also be used. Whichever data is used, the relative displacement amount 135 after the time point of S103 can be calculated.

[0061] That is, based on the sensor information received from the controller 107, the coordinate processing unit 104 calculates how much the positional relationship between the feature point 132 in the real space and the video display device 100 has deviated since the reference coordinate system was determined due to the movement of the video display device 100.

[0062] From another perspective of this embodiment, based on the sensor information received from the controller 107, the coordinate processing unit 104 calculates how much the positional relationship between the feature point 132 in the real space and the video display device 100 has deviated since the previous coordinate system information was calculated due to the movement of the video display device 100.

[0063] As shown in FIG. 5(b), it is possible to calculate the relative displacement amount 135 since the reference coordinate system was determined by integrating the displacement amount since the previous coordinate system information was calculated.

[0064] In an example of the video display device 100 shown in FIG. 4, the coordinate processing unit 104 receives angular velocity data from the controller. Now, assuming that the angular velocity is substantially constant, it is assumed that the angular velocity data at the timing of S105 is used. If the time elapsed since the previous coordinate system information was calculated is ΔT and the angular velocity is ω, the amount of angular deviation from the time when the previous coordinate system information was calculated is ΔT×ω. By integrating the amount of angular deviation from the time when the previous coordinate system information was calculated from the time when the reference coordinate system was determined, it is possible to calculate the amount of angular deviation from the time when the first reference coordinate system was determined.

[0065] The coordinate processing unit 104 determines the corrected coordinate system based on the recognized relative deviation amount 135. Specifically, the coordinate processing unit 104 previously stores information regarding the relationship between the relative deviation and the captured image of the imaging unit 101. That is, it previously stores information regarding what distance or angle in what direction in the captured image of the imaging unit 101 corresponds to the relative deviation in the real space. Such information can be determined based on the specifications of the imaging unit 101 and the sensor unit 102. Based on this information and the relative deviation amount, the coordinate processing unit 104 shifts or rotates the first reference coordinate system by an amount corresponding to the relative deviation amount to determine the first corrected coordinate system. Thereby, even when the video display device 100 moves, the first corrected coordinate system becomes a coordinate system that hardly moves in the real space.

[0066] In the example shown in FIG. 5(c), the x'-axis 136 and the y'-axis 137, which are shifted by the relative deviation amount 135 with respect to the first reference coordinate system determined by the x-axis 133 and the y-axis 134, are determined as the first corrected coordinate system.

[0067] The coordinate processing unit 104 uses the coordinate conversion information to convert the first corrected coordinate system set in the captured image into the second corrected coordinate system in the video displayed by the video display unit 106. The coordinate processing unit 104 transmits the second corrected coordinate system information in the video displayed by the video display unit 106 to the controller 107.

[0068] In the above description, the displacement amount from when the previous coordinate system information was calculated was integrated from when the reference coordinate system was determined to calculate the relative displacement amount from when the reference coordinate system was determined. However, this embodiment is not limited to this. The coordinate processing unit 104 may be configured to determine the correction coordinate system from the reference coordinate system or the correction coordinate system when the previous coordinate system information was calculated and the relative displacement amount, with the displacement amount from when the previous coordinate system information was calculated as the relative displacement amount. That is, the difference from the determination of the correction coordinate system in the previous S106 is corrected in the next S106. This eliminates the need for integrating the displacement amount and enables reduction of the processing.

[0069] In the above description, it was described that after the coordinate processing unit 104 sets the first correction coordinate system in the captured image, it converts to the second correction coordinate system in the video displayed by the video display unit 106 according to the coordinate conversion information. However, this embodiment is not limited to this. The coordinate processing unit 104 may use the coordinate conversion information to calculate in which direction and at what distance in the video displayed by the video display unit 106 the relative displacement amount corresponds, and set the second correction coordinate system in the video displayed by the video display unit 106. This enables reduction of the processing in the coordinate processing unit 104.

[0070] As described above, a first reference coordinate system is set based on the feature points in the captured image captured by the camera 111, and a corresponding second reference coordinate system is set in the video displayed by the video display unit 106. The relative displacement amount caused by the position and angular displacement of the camera 111 is calculated based on the signal from the sensor 112. As shown in FIG. 5, the relative displacement amount 135 corresponds to the displacement of an object (typically a reference point) in the captured image. The second reference coordinate system is corrected based on this relative displacement amount. Since the position of the content displayed by the video display unit 106 is determined based on the corrected second reference coordinate system, by correcting the second reference coordinate system to follow the displacement of the object in the captured image to obtain a second corrected coordinate system, the change in the relative positional relationship between the object in the captured image and the content perceived by the user 700 can be suppressed.

[0071] Since the recognition of the relative displacement amount and the determination of the correction coordinate system do not require the spatial recognition of the entire real space, it is possible to quickly recognize the relative displacement amount and determine the correction coordinate system. Note that, by repeatedly determining the reference coordinate system (S103) itself at a high frequency, it is possible to suppress the change in the relative positional relationship between the object and the content in the captured image. However, in this case, there is a problem that the processing load for determining the reference coordinate system increases.

[0072] Next, the video display device 100 determines the display content (S107). The controller 107 of the video display device 100 sends the coordinate system information and the command received from the coordinate processing unit 104 to the video display processing unit 105. The video display processing unit 105 determines the display content to be displayed by the video display device 100 and sends the display content to the controller 107. The determination of the display content will be described later with reference to FIG. 9.

[0073] Based on the coordinate system information received from the controller 107, the video display processing unit 105 determines the position of the content to be displayed with reference to the second correction coordinate system. Even when the video display device 100 moves, since the second correction coordinate system is a coordinate system that hardly moves in the real space, it is possible to perform video display that is hardly affected by the movement of the video display device 100.

[0074] Next, the video display device 100 displays the content (S108). The controller 107 of the video display device 100 sends the display content and the command received from the video display processing unit 105 to the video display unit 106, and the video display unit 106 displays the display content.

[0075] Next, the video display device 100 determines whether there is an end command (S109). As the end command, for example, information input from an input unit (not shown) of the video display device 100 or information obtained by recognizing the information can be used. For example, the video display device 100 includes a microphone as an input unit, the microphone receives a voice command from the user 700, the video display device 100 recognizes the voice command, and when a predetermined command is recognized, it is possible to configure the video display device 100 that determines that there is an end command. If there is an end command, the video display device 100 ends the flow.

[0076] If there is no end command, the video display device 100 repeatedly determines whether a predetermined time has elapsed until the predetermined time elapses (S110). In other words, it waits until the predetermined time elapses. After the predetermined time has elapsed, the video display device 100 repeats the flow from the acquisition of sensor information (S105). Thereby, it is possible to repeatedly determine the correction coordinate system based on the relative displacement amount recognized from the sensor information at a predetermined time interval, and even when the video display device 100 continues to move, it is possible to continuously correct the coordinate system, and it is possible to perform video display that is not substantially affected by the movement of the video display device 100 continuously.

[0077] FIG. 7 is a diagram showing an example of another flowchart related to the content display of the present embodiment. The difference from FIG. 6 will be described. After the video display device 100 repeatedly determines whether a predetermined time has elapsed until the predetermined time elapses in S110, it determines whether a predetermined number of times has elapsed (S111). Specifically, for example, the controller 107 includes a loop counter, and each time a determination is made, the loop counter is incremented by a predetermined value, and it is determined whether the loop counter exceeds a predetermined value.

[0078] In another aspect of this embodiment, the video display device 100 determines whether a predetermined time has elapsed since the previous reference coordinate system was determined. If the predetermined time has elapsed, it is regarded as having elapsed a predetermined number of times. When the video display device 100 determines that the predetermined number of times has elapsed, it repeats from the acquisition of the camera image (S101). When it determines that the predetermined number of times has not elapsed, it repeats from the acquisition of the sensor information (S105).

[0079] Accordingly, the video display device 100 can determine the reference coordinate system using the camera image every predetermined number of times or every predetermined time. For example, even when the camera 111 does not move, signal drift over time or the like may occur. It is possible to improve the position accuracy of the video overlay even for such signal drift. The update of the reference coordinate system can be performed, for example, at a frequency of about once per second.

[0080] FIG. 8 is a diagram showing an example of the feature point 132 according to this embodiment. In FIG. 8, a distribution board 131A is shown as an example of the object 131. As the feature point 132, for example, the vertices 151A to 151D of the object 131 can be used.

[0081] In another aspect of this embodiment, as the feature point 132, for example, the sides 152A to 152D of the object 131 can be used. In another aspect of this embodiment, as the feature point 132, the center 153 of the object 131 or the center of gravity of the object 131 can be used. In another aspect of this embodiment, as the feature point 132, a predetermined position of an object 154 inside the object 131 can be used. For example, the vertex 154A of the object 154 or the side 154B of the object 154 can be used.

[0082] In another aspect of this embodiment, as the feature point 132, a predetermined position 155A of the character string 155 described in the object 131 can be used. As the predetermined position 155A, for example, among the characters at a predetermined character number of the character string 155, the top, bottom, right, left, etc. can be used.

[0083] In another aspect of this embodiment, as the feature point 132, it is possible to use a predetermined symbol 156. The predetermined symbol may be a pre-determined symbol or an unknown symbol. The symbol 156 may be a one-dimensional code, a two-dimensional code, or the like. As the feature point 132, it may be a component of the object 131 that the object 131 is pre-equipped with, or an object attached to the object 131.

[0084] As described above, the feature points can be defined in advance, and the feature point position recognition unit 103 can extract them from the captured image captured by the camera 111 by a known method based on the color (for example, hue, lightness, and saturation) and shape of the feature points.

[0085] FIG. 9 is a diagram showing an example of a method for determining display content by the video display processing unit 105 according to this embodiment. In FIG. 9, for the sake of explanation, the real space and the virtual space as seen from the user 700 are shown together. The object 131 is an object located in the real space. The coordinate processing unit 104 determines the x'-axis 136 and the y'-axis 137 in the captured image as the first correction coordinate system, and converts them into the second correction coordinate system in the video displayed by the video display unit 106.

[0086] Now, as a specific example to help understanding, the object 131 is an object in the real space, such as a distribution board. In this example, the field of view of the camera 111 only needs to be able to recognize the object 131. In this example, the case where content 162, for example, an operation manual for the distribution board, is to be displayed on the right side of the object 131 will be described.

[0087] The video display processing unit 105 determines the position of the content to be displayed based on the second correction coordinate system based on the coordinate system information received from the controller 107. Specifically, the video display device 100 determines the position of the display content as follows, for example.

[0088] The image display processing unit 105 stores the content in the virtual space in a virtual space storage unit (not shown in the figure) in advance. For example, the RAM 115 can be used as the virtual space storage unit. In the virtual space 161 stored in the virtual space storage unit, images of a plurality of contents are positioned and stored at predetermined positions in the virtual space 161 respectively.

[0089] That is, regarding the virtual space, it has a map of which content to display at which location for the object 131. Specifically, in this example, the relative positional relationship between the object 131 and the virtual space 161 is determined in advance, and it is determined that the content 162 is located on the right side of the object 131. The position of the content can be indicated by coordinates according to, for example, the second coordinate system. The size of the virtual space 161 is arbitrary and there is no limit on the size. Regarding the real space, it does not have a map, and it is only necessary to be able to generate the first reference coordinate system based on the camera image.

[0090] In the example of FIG. 9, three contents 162A to 162C are stored in the virtual space 161 with their positions shifted in the left - right direction. The storage format of the virtual space 161 is, for example, an image file such as a bitmap image. In a general example, since the viewing angle that the video display unit 106 can display is small, here, the area that can be displayed by the video display device 100 is only the content located in the display area 163 in the virtual space 161, and the position of the display area 163 is determined by the video display processing unit 105. In the example of FIG. 9, the display area 163 overlaps a part of the content 162A, and the video display device 100 can display a part of the content 162A in the display area 163.

[0091] The position of the display area 163 is determined by the coordinates of the second coordinate system in conjunction with, for example, the movement of the head. In this way, in accordance with the movement of the camera, the display area is cut out from the virtual space to display the video. Regarding the method of viewing the content arranged in the line - of - sight direction by changing the direction of the head, for example, known techniques described in Patent Document 1 etc. may be adopted.

[0092] As described above, in this embodiment, a method of correcting the second reference coordinate system itself is shown. However, the same effect can be achieved by correcting the coordinates of the display area 163 and the coordinates of the content in the second reference coordinate system. In this case, the coordinates of the content and the like may be corrected so that the shift amount is the same as that in the case of correcting the second reference coordinate system and the direction is shifted in the opposite direction.

[0093] After determining the x'-axis 136 and the y'-axis 137 as the first corrected coordinate system in the captured image, the coordinate processing unit 104 calculates in the virtual space 161 the positions and directions corresponding to the x'-axis 136 and the y'-axis 137, and transmits the coordinate system information as the second corrected coordinate system in the virtual space 161 to the controller 107. That is, the coordinate system information represents, for example, the positions and directions in which at least one of the virtual space 161 or the display area 163 corresponds to the second corrected coordinate system.

[0094] The video display processing unit 105 receives the coordinate system information regarding the second corrected coordinate system in the virtual space 161 from the controller 107, and sets the display area 163 at a predetermined position in the second corrected coordinate system.

[0095] The video display processing unit 105 transmits the information within the display area 163 to the controller 107 as display content. Since the second corrected coordinate system is a coordinate system that hardly moves in the real space, the content displayed by the video display device 100 is hardly affected by the movement of the video display device 100. That is, video display is possible such that the user 700 perceives the content as if it were sticking to the real space. In addition, since the content is stored in the virtual space storage unit in advance, it is not necessary to draw a part or the whole of the virtual space every time the video display processing unit 105 determines the display content, and the speed of determining the display content can be increased, that is, high speed can be achieved.

[0096] In a transmissive display that superimposes content in the real space for visual recognition, when the user visually recognizes an object as described above, the display content can be simultaneously displayed. In the above description, since both the reference coordinate system and the correction coordinate system are defined on the object in the captured image, a two-dimensional coordinate system is basically sufficient. The same applies to the coordinate system of the virtual space.

[0097] FIG. 10 is a diagram showing another example of a method for determining display content by the video display processing unit 105 according to the present embodiment. In FIG. 10, for the sake of explanation, the real space and the virtual space as seen from the user 700 are shown together. The video display processing unit 105 previously stores information regarding the coordinates of the content 171 to be displayed by the video display device in the correction coordinate system. For example, the video display processing unit 105 stores information such as displaying the content "operation manual" at a position 5 degrees from the right end of the distribution board. This information can be shown as the coordinates of the second correction coordinate system. When the video display processing unit 105 recognizes that at least a part of the content 171 overlaps the display area 163, it draws the content 171 as the display content and transmits the display content to the controller 107.

[0098] Since the second correction coordinate system is a coordinate system that hardly moves in the real space, the video display device 100 can perform video display such that the user 700 perceives that the content is hardly affected by the movement of the video display device 100, that is, as if the content is sticking to the real space. For example, when the correction coordinate system is attached to the distribution board, the content can be displayed as if it is attached to the distribution board by specifying the coordinates in the correction coordinate system. In addition, since it is not necessary to store the information of the entire virtual space in the virtual space storage unit, it is possible to reduce the usage amount of the storage unit and achieve cost reduction.

[0099] FIG. 11 is a diagram showing an example of the position of display content displayed by the video display device 100 according to the present embodiment. The display position of the display content can be determined in advance by defining the positional relationship with an object in the following method in the first correction coordinate system and converting it to the coordinates of the second correction coordinate system using coordinate conversion information. If the positions of the content 162 in the virtual space 161 and the display area 163 are defined in the second correction coordinate system, the overlapping part will be displayed.

[0100] The eye point 180 indicates the position of the eyes of the user 700 who uses the video display device 100. Also, when the user 700 faces forward and looks straight ahead, a point on the object 131 that the user visually recognizes is defined as the center point 181, and the end of the object 131 is defined as the end 182.

[0101] The video display device 100, for example, displays the display content 184 so as not to overlap with the object 131. Specifically, the display content 184 is displayed such that the end 183 of the display content 184 is on the opposite side of the center point 181 with respect to the end 182 of the object. Thereby, it is possible to prevent the user 700 from being less likely to recognize the object 131 due to the display content 184, and it is possible to improve the visibility of the object 131. From another perspective of the present embodiment, the center point 181 can also be the center of the object 131.

[0102] From another perspective of the present embodiment, the display content is displayed such that the angle θ formed by the line segment connecting the eye point 180 and the center point 181 and the line segment connecting the eye point 180 and the end 183 of the display content is equal to or greater than a predetermined angle. For example, the video display device 100 displays the display content such that the angle θ is 2.5 degrees or more. Thereby, when the user 700 is looking straight ahead, the video display device 100 can display the video while avoiding the discrimination visual field range, so that the visibility of the object 131 by the user 700 can be improved.

[0103] In another aspect of this embodiment, the video display device 100 displays display content such that the angle θ is 15 degrees or more. As a result, the video display device 100 can display a video while avoiding the effective visual field, so that the visibility of the object 131 by the user 700 can be further improved.

[0104] In another aspect of this embodiment, the video display device 100 displays display content such that the angle θ is 30 degrees or more. As a result, the video display device 100 can display a video while avoiding the stable fixation visual field, so that the visibility of the object 131 by the user 700 can be further improved.

[0105] In another aspect of this embodiment, the video display device 100 displays display content such that the angle θ is 50 degrees or more. As a result, the video display device 100 can display a video while avoiding the induced visual field, so that the visibility of the object 131 by the user 700 can be further improved.

[0106] According to this embodiment, since the reference coordinate system is determined based on the captured image of the imaging unit 101 and the corrected coordinate system is determined based on the sensing information from the sensor unit 102, and the display content is determined based on the corrected coordinate system, it is possible to provide the video display device 100 capable of displaying a video with high position accuracy. In addition, since it is possible to recognize the position and relative displacement amount of high-speed feature points, and to determine the reference coordinate system, the corrected coordinate system, and the display content, it is possible to provide the video display device 100 capable of displaying a superimposed video with little delay with respect to the movement of the video display device 100.

Embodiment

[0107] In the second embodiment, the configuration is such that the feature points when determining the reference coordinate system can be changed with time change. In the second embodiment, the differences from the above-described embodiment will be mainly described, and the same components as those in the above-described embodiment are denoted by the same reference numerals, and the descriptions thereof are omitted.

[0108] FIG. 12 is a diagram showing an example of the feature points and the reference coordinate system according to this embodiment.

[0109] FIG. 13 is a diagram showing an example of a flowchart related to content display in this embodiment. Hereinafter, differences from FIG. 7 will be mainly described.

[0110] In the flow shown in FIG. 13, when it is determined in S111 that a predetermined number of times has elapsed, recognition of the position of the feature point (S202) and determination of the reference coordinate system (S203) are repeatedly performed. The times at which recognition of the position of the feature point (S202) and determination of the reference coordinate system (S203) are performed at different times are defined as time T1, time T2, and time T3. Here, time T2 is a time after time T1, and time T3 is a time after time T2.

[0111] Note that although there is naturally a time difference at the times when recognition of the position of the feature point (S202) and determination of the reference coordinate system (S203) are performed, in this embodiment, no particular mention is made of the time difference between recognition of the position of the feature point (S202) and determination of the reference coordinate system (S203), and they are referred to as time T1 etc. as representative times for executing both steps.

[0112] As shown in FIG. 12(a), in S202 at time T1, the feature point position recognition unit 103 recognizes the position of the feature point 200. Also, in S203 at time T1, the coordinate processing unit 104 determines the x-axis 202A and y-axis 203A of the reference coordinate system based on the position of the feature point 200.

[0113] In S202 at time T2, the feature point position recognition unit 103 recognizes the positions of the feature point 200 and the feature point 201, and transmits the recognition results of the positions of the feature point 200 and the feature point 201 different from the feature point 200 to the controller 107. In this example, two feature points are recognized, but three or more may be used.

[0114] As shown in FIG. 12(b), at S203 at time T2, the controller transmits the recognition results of the positions of feature point 200 and feature point 201 to the coordinate processing unit 104. The coordinate processing unit 104 detects the difference 204 in the positions of feature point 200 and feature point 201, and the x-axis 202A and y-axis 203A of the reference coordinate system determined at time T1, or the x'-axis and y'-axis of the corrected coordinate system determined by the coordinate processing unit 104 immediately before time T2, and the x-axis 202B and y-axis 203B of the reference coordinate system with feature point 201 as the reference are substantially coincident. Based on feature point 201, the x-axis 202B and y-axis 203B of the reference coordinate system are determined.

[0115] As shown in FIG. 12(c), at S202 at time T3, the feature point position recognition unit 103 recognizes the position of feature point 201. Also at S203 at time T3, the coordinate processing unit 104 determines the x-axis 202B and y-axis 203B of the reference coordinate system based on the position of feature point 201.

[0116] According to this embodiment, the video display device 100 can change the feature points for determining the reference coordinate system. As a result, even when the video display device 100 moves and the feature points go out of the image range in the image captured by the imaging unit 101, another feature point appearing in the image captured by the imaging unit 101 can be used to determine the reference coordinate system. Therefore, even when the video display device 100 moves significantly, it is possible to provide a video display device 100 capable of high-precision video display.

Embodiment

[0117] In the third embodiment, a configuration is adopted in which the user 700 can indicate feature points by commands. In the third embodiment, the differences from the above-described embodiments are mainly described, and the same components as those in the above-described embodiments are denoted by the same reference numerals, and the descriptions thereof are omitted.

[0118] FIG. 14 is a diagram showing an example of the functional blocks of the video display device 100 according to this embodiment. The video display device 100 includes an input unit 301 and a command recognition unit 302.

[0119] FIG. 15 is a diagram showing an example of the video display device 100 according to the present embodiment. The video display device 100 includes a microphone 310 as an input unit 301.

[0120] FIG. 16 is a diagram showing an example of a flowchart leading to the determination of the reference coordinate system of the present embodiment.

[0121] First, the video display device 100 displays a marker (S301). The controller 107 transmits a command to the video display processing unit 105. The video display processing unit 105 draws a marker on the display content according to the command and transmits the display content to the controller 107.

[0122] FIG. 17 is a diagram showing an example of a state in which a marker is superimposed and displayed in the real space according to the present embodiment. A marker 321 is displayed in a part of the video area 320 displayed by the video display device 100. In FIG. 16, an example of using the symbol "+" as a marker is shown.

[0123] Next, the video display device 100 determines whether there is a command for specifying feature points (S302). The input unit 301 transmits the information input to the input unit to the controller 107. The controller transmits the information input from the input unit to the command recognition unit 302. The command recognition unit 302 detects whether a predetermined command exists in the information received from the controller. The command recognition unit 302 transmits the detection result to the controller 107.

[0124] In the example of the video display device including the microphone 310 as the input unit 301, the command recognition unit 302 detects, for example, whether a predetermined character string uttered by the user 700, such as "feature point set", is included in the information of the sound sensed by the microphone. If the predetermined character string is included, the command recognition unit 302 transmits the recognition result to the controller 107 as having a command for specifying feature points.

[0125] The controller 107 repeats until it receives the recognition result from the command recognition unit 302 that there is a command for specifying feature points.

[0126] Next, the video display device 100 makes the marker invisible (S303). The controller 107 sends a command to the video display processing unit 105. In accordance with the command, the video display processing unit 105 changes its operation so as not to draw the marker on the display content, and transmits the display content to the controller 107.

[0127] Next, the video display device 100 acquires a camera image (S101).

[0128] Next, the video display device 100 recognizes the position of the feature point (S304). The controller 107 transmits a predetermined position 322 of the marker 321 to the feature point position recognition unit 103, and the feature point position recognition unit 103 uses the position 322 of the marker in the camera image as the feature point. Then, it recognizes the position of the feature point within the camera image.

[0129] In the above, the position 322 of the marker is used as the feature point, but the present invention is not limited to this. For example, it is also possible to use the vicinity of the marker as the feature point. For example, it is also possible to use an object existing within a predetermined pixel range centered on the position of the marker in the camera image as the feature point. Thereby, it becomes possible to set the feature point even when there is no object at the position of the marker.

[0130] From another perspective of the present invention, for example, it is also possible to use the location indicated by the marker or its vicinity as the feature point. For example, an arrow is displayed as the marker, and it is also possible to use the tip of the arrow or its vicinity as the feature point. Thereby, it is possible to avoid the marker overlapping an object and improve visibility.

[0131] The above flow is the recognition flow of the position of the feature point of the video display device 100 according to the present embodiment. After the recognition of the position of the feature point (S304), the operations after the execution of S104 in FIG. 7 and the like are executed.

[0132] According to this embodiment, the user 700 can indicate feature points by commands, and it becomes possible to determine a reference coordinate system according to the intention of the user 700. Also, even when it is difficult for the video display device 100 to determine feature points, the user can specify feature points according to their intention, and it is possible to provide the video display device 100 that can stably display video with high position accuracy.

Embodiment

[0133] Example 4 is an example applied to a remote work support system 450 at a factory or a work site where workers are performing assembly, manufacturing, inspection, maintenance, inspection, etc.

[0134] FIG. 18 is a configuration diagram showing an example of a remote work support system 450 that utilizes the video display device 100 according to this embodiment.

[0135] The worker 701 carries the video display device 100 of this embodiment and performs product assembly, inspection, etc. on the factory line. The video display device 100 projects a virtual image in the space as described above and displays content for work support such as work procedure manuals, drawings, checklists, etc. In the embodiment of the present invention, a technology may be installed that can always display a video by determining a spatial position coordinate for a display range in the visual field that does not interfere with the manual work of the worker 701, greatly improving usability. For example, it can be set automatically or by the worker 701 side or the administrator 420 side so as not to display from the center of the line-of-sight direction horizontally ±10° to ±75° and vertically -60° to -90°, ensuring that the content is not displayed at the hands and feet, enabling safe visual field confirmation at the hands and feet, and allowing the inspection work and assembly work that involve moving around to be carried out safely and with confidence.

[0136] The image display device 100 is equipped with communication means (not shown) and is connected to the remote management device 410 of the administrator 420 who is in a remote location via a network 401 such as Wi-Fi (trademark), LTE (trademark), 5G, 6G, wired, or optical, for local or public wireless networks, to share information. The worker 701 owns an image display device 100 such as a tablet, smartphone, HMD, or PC (Personal Computer), which is an edge terminal, and exchanges data through the network 401. Through the network 401, sensing data such as images from cameras 402 around the worker, data from sensors 403, and the status of devices 404 is transmitted to the edge server 407 and the administrator 420. The edge server 407 and the administrator 420 perform situation recognition, judgment, and prediction (especially prediction of dangers, abnormalities, abnormal behaviors, etc.) based on various information. Also, when an abnormal state is recognized in this way, the edge server 407 and the administrator 420 convey work instructions and warnings to the worker 701 through display images, voices, and instruction signals of devices such as the image display device 100.

[0137] Also, when incorporating a work management system (not shown) that allows the worker 701 to inquire about doubts in drawings and work processes and automatically discriminates check mistakes in the work process, the worker 701 or the administrator 420 can display, on the image display device 100 such as an edge terminal, display content corresponding to the on-site situation in real time. Furthermore, content that supports work instructions such as work instructions, text input, and handwritten input can also be displayed from the administrator 420, and the doubts, questions, and concerns of the worker 701 can be resolved by the display content. At this time, the display content is displayed at the absolute coordinate position of the line of sight of the worker 701 described above, or at an absolute or relative coordinate arrangement that does not obstruct the line of sight of the worker 701 due to the content. In this case, since it is displayed at a position that does not obstruct the work area or walking, a more comfortable remote work support system 450 can be provided.

[0138] By superimposing the virtual space video on the real space in this way, a user-friendly video display device 100 such as a head-mounted display and a remote work support system 450 can be provided. In a factory or the like, there are cases where work is performed while viewing content such as work processes. However, it may be difficult to place paper media such as displays, drawings, and work process sheets near the work target, or there may be work that requires both hands at heights. In such cases, by using a video display device 100 such as a head-mounted display and the above-described remote work support system 450, the worker 701 can perform work while referring to work instructions and the like displayed on the video display device 100. By utilizing the above-described remote work support system 450 with excellent usability, safety and work efficiency can be improved.

[0139] The administrator 420 may be located inside the same network 401 as the worker 701, or may be located at another location via the Internet. Also, at least a part of the above-described edge servers 407 may be cloud servers 408. The remote work support system 450 may be provided in the edge server 407, or may be provided in the cloud server 408.

[0140] FIG. 19 is a schematic configuration diagram example of a video transmission system incorporated in the remote work support system 450 of this embodiment. For example, a camera 402 with 4K, 30fps performance is utilized for observing the working environment of the worker 701. USB 3.0 is used for information transmission from the camera to the PC 431. Also, the PC 431 utilizes a Wi-Fi6 router 432 (IEEE 802.11ax, 9.6Gbps) to transmit video information. Further, it has a function of communicating video information, audio, sensor information, etc. to the administrator 420 located on the remote side bidirectionally or unidirectionally with emphasis via the Wi-Fi6 network 433. The Wi-Fi router 434 installed on the remote management side is utilized to connect to the PC 435. For example, the connection from the Wi-Fi router 434 to the PC 435 is set as 1000BASE-T (IEEE 802.3ab) to configure information exchange. Thereby, video information and audio information can be smoothly communicated with a low latency within 200 ms.

[0141] As described above, the remote work support system 450 capable of transmitting 4K video information and audio information at high speed and low latency with at least two or more is realized by the information transmission system in FIG. 19, and at the same time, the spatial display of video content according to the above invention can be realized within 100 ms. Therefore, the total delay amount is 300 ms or less, and a delay time of 300 ms or less, which is the delay time that humans can feel, can be realized. Thereby, a 4K video transmission system with a low latency of 300 ms or less can be constructed, and the remote work support system 450 with improved usability can be realized.

[0142] FIG. 20 is a diagram showing an example of the remote work support system 450. The worker 701 is carrying a video display device 100. The video display device 100 is equipped with a camera (not shown). The camera is shooting in front of the worker 701, and the captured image is transmitted to the remote management device 410 of the administrator 420 via the network 401. The remote management device 410 displays the camera video, and the administrator 420 visually recognizes the displayed video 460.

[0143] The image captured by the camera includes at least the object 131. The display video 460 is displayed on, for example, a tablet that enables touch operations or handwritten input. The administrator 420 recognizes the displayed image and inputs instructions to the worker 701 through touch operations or handwritten input. For example, the worker 701 inputs by indicating a predetermined position where the work should be performed with a circle mark, an arrow, or the like. The input instructions are transmitted to the video display device 100 via the network 401, and the video display device 100 displays the instructions.

[0144] The video captured by the camera may dynamically change due to the movement of the worker 701 or the like. In that case, the remote work support system 450 includes an instruction position recognition unit (not shown). The instruction position recognition unit recognizes the object 131 shown in the video captured by the camera, recognizes the relative position between the position of the instruction input by the user and the object 131, and transmits the position of the instruction to be displayed by the video display device 100 as the relative position with respect to the object 131 via the network 401. The video display device 100 displays the instruction based on the relative position with respect to the object 131. The display process of the video displayed by the video display device 100 is performed by the video display processing unit 105. That is, in this embodiment, the reference coordinate system is determined based on the captured image of the imaging unit 101, and the corrected coordinate system is determined based on the sensing information from the sensor unit 102. To determine the display content based on the corrected coordinate system, the video display device 100 can display the instruction with high position accuracy.

[0145] Note that the position of the instruction displayed by the video display device 100 may be superimposed so as to overlap the object 131, or may be displayed so as not to overlap the object 131. For example, as shown in FIG. 17, the video display device 100 may display the display content 184 so as not to overlap the object 131.

[0146] In the above description, the operator 701 inputs by indicating a predetermined position where the work should be performed with a circle mark, an arrow, etc., but the present invention is not limited to this. For example, the remote management device 410 is provided with a microphone and a voice recognition unit. The administrator 420 issues an instruction to the operator 701 by voice. The microphone collects the voice of the administrator 420, and the collected voice is recognized by the voice recognition unit, and the recognized instruction may be used as an input.

[0147] From another aspect of this embodiment, the administrator 420 inputs an instruction by voice as described above, and also inputs regarding the position where the instruction is to be displayed on a tablet or the like on which the display video 460 is displayed. The video display device 100 may be configured to display an instruction at the display position.

[0148] In the above description, it is assumed that the administrator 420 inputs an instruction by voice, but the present invention is not limited to this. For example, a predetermined instruction may be acquired from a database provided in the edge server 407, the cloud server 408, the remote management device 410, or the remote work support system 450, and the acquired instruction may be used as an input.

Example

[0149] In the above embodiment, it is assumed that the position of the display area 163 is determined according to the direction of the head of the user 700. In this case, the virtual space is fixed to the real space. However, as described in Patent Document 1, the position of the display area 163 can also be determined based on the difference in the movement of the head and waist of the user 700. In this case, when viewed from the user 700, the virtual space always moves in accordance with the direction of the waist.

[0150] According to the above embodiment, with a small processing amount, the delay in video tracking with respect to the movement of the user's head is small, and it is possible to perform highly accurate superimposition of the real space and the virtual space. Therefore, it is possible to reduce energy consumption, reduce carbon emissions, prevent global warming, and contribute to the realization of a sustainable society.

[0151] Note that the present invention is not limited to the above-described embodiments, and various modifications and equivalent configurations within the spirit of the invention are included. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and the present invention is not necessarily limited to those having all the configurations described. Also, a part of the configuration of one embodiment may be replaced with the configuration of another embodiment. Further, the configuration of another embodiment may be added to the configuration of one embodiment. Also, for a part of the configuration of each embodiment, addition, deletion, or replacement with other configurations may be made.

[0152] In addition, each of the above-described configurations, functions, processing units, processing means, etc. may be realized in hardware, for example, by designing a part or all of them with an integrated circuit, or may be realized in software by a processor interpreting and executing a program for realizing each function.

[0153] Information such as programs, tables, files, etc. for realizing each function can be stored in a storage device such as a memory, a hard disk, an SSD (Solid State Drive), or a recording medium such as an IC card, an SD memory card, a DVD.

[0154] Also, the control lines and information lines show those considered necessary for explanation, and do not necessarily show all the control lines and information lines necessary for implementation. In practice, it may be considered that almost all the configurations are interconnected.

Explanation of Reference Numerals

[0155] 100: Image display device, 101: Imaging unit, 102: Sensor unit, 103: Feature point position recognition unit, 104: Coordinate processing unit, 105: Image display processing unit, 106: Image display unit, 107: Controller, 111: Camera, 112: Sensor, 113: Display, 114: CPU, 115: RAM, 116: Storage medium, 121: Gyro sensor, 122: Acceleration sensor, 131: Object, 131A: Distribution board, 132: Feature point, 133: x-axis, 134: y-axis, 135: Relative displacement amount, 136: x'-axis, 137: y'-axis, 151: Vertex, 152: Side, 153: Center, 154: Object, 154A: Vertex, 154B: Side, 155: Character string, 155A: Predetermined position, 156: Symbol, 161: Virtual space, 162: Content, 163: Display area, 171: Content, 180: Eye point, 181: Center point, 182: End, 183: End, 184: Display content, 200: Feature point, 201: Feature point, 202: x-axis, 203: y-axis, 204: Difference in position, 301: Input unit, 302: Command recognition unit, 310: Microphone, 320: Image area, 321: Marker, 322: Position, 401: Network, 402: Camera, 403: Sensor, 404: Device, 407: Edge server, 408: Cloud server, 410: Remote management device, 420: Administrator, 450: Remote work support system, 460: Displayed image, 700: User, 701: Operator

Claims

1. In an image display system, a feature point position recognition unit, a coordinate processing unit, and an image display processing unit, and the feature point position recognition unit recognizes the position of feature points of an object based on a captured image captured by an imaging unit, the coordinate processing unit determines a reference coordinate system based on the position of the feature points, detects the movement of the imaging unit based on sensor information indicated by a sensor unit including at least one of a gyro sensor and an acceleration sensor, and corrects the display position of an image based on the movement of the imaging unit, the image display processing unit displays an image at the corrected display position of the image when displaying the image based on the reference coordinate system. An image display system characterized by this.

2. In the image display system according to Claim 1, the coordinate processing unit converts a first reference coordinate system in the captured image into a second reference coordinate system based on coordinate conversion information regarding the positional relationship between the captured image and the image to be displayed, the image display processing unit displays an image based on the second reference coordinate system. An image display system characterized by this.

3. In the image display system according to Claim 1, the coordinate processing unit corrects the display position of the image based on the movement of the imaging unit by correcting the reference coordinate system to determine a corrected coordinate system, the image display processing unit displays content at a predetermined position in the corrected coordinate system. An image display system characterized by this.

4. In the image display system according to Claim 1, the image display unit for displaying the image, the image display processing unit, the imaging unit, the sensor unit, the feature point position recognition unit, and the coordinate processing unit are provided in a wearable image display device. An image display system characterized by this.

5. In the image display system according to Claim 1, the gyro sensor measures an angular velocity, and the acceleration sensor measures an acceleration. An image display system characterized by this.

6. In the image display system according to Claim 1, the coordinate processing unit corrects the display position of the image based on the movement of the imaging unit after the reference coordinate system is determined. An image display system characterized by this.

7. In the image display system according to Claim 1, the coordinate processing unit determines and updates the reference coordinate system again based on the position of the feature points after a lapse of a predetermined time or more from the time when the reference coordinate system is determined based on the position of the feature points. An image display system characterized by this.

8. In the video display system according to claim 1, when the time T2 has elapsed for a predetermined time or more since the time T1 when the coordinate processing unit determined the reference coordinate system based on the position of the first feature point, the feature point position recognition unit recognizes the position of the first feature point and a second feature point different from the first feature point, the coordinate processing unit determines and updates the reference coordinate system based on the position of the second feature point recognized by the feature point position recognition unit at time T2, the feature point position recognition unit recognizes the position of the second feature point at time T3 when a predetermined time or more has elapsed since time T2, the coordinate processing unit determines and updates the reference coordinate system based on the position of the second feature point recognized by the feature point position recognition unit at time T3. A video display system characterized by this.

9. In the video display system according to claim 1, the feature point is an edge or corner of an object shown in the captured image, or a point where at least one of hue, brightness, and saturation changes by a predetermined amount or more in the captured image, or a predetermined object determined in advance. A video display system characterized by this.

10. In the video display system according to claim 1, an input unit, a video display unit that displays a video superimposed on the captured image, a command recognition unit that recognizes a command input to the input unit, and after displaying a predetermined video superimposed on the captured image, when the command recognition unit recognizes a predetermined command, the feature point position recognition unit uses any one of the location of the predetermined video in the captured image, the vicinity of the location, the location indicated by the video, and the vicinity of the location indicated by the video as a feature point and recognizes the position of the feature point. A video display system characterized by this.

11. In the video display system according to claim 3, the predetermined position is a position that is separated from the front of the user by a predetermined angle or more, or a position that is farther from the feature point when the front of the user is used as a reference. A video display system characterized by this.

12. In a video display device that displays a video, an imaging unit that captures an object and acquires a captured image, a sensor unit that detects the movement of the imaging unit by at least one of a gyro sensor and an acceleration sensor, a feature point position recognition unit, a coordinate processing unit, a video display processing unit, and a video display unit that displays a video, and the feature point position recognition unit recognizes the position of the feature point of the object based on the captured image, The coordinate processing unit determines a reference coordinate system based on the positions of the feature points, and generates correction information for the position of the video displayed by the video display unit based on the movement of the imaging unit. The video display processing unit determines the display position of the video based on the reference coordinate system and the correction information. The video display unit displays the video at the determined display position. A video display device characterized by the above.

13. A video display method for displaying desired content, comprising: a first step of recognizing the positions of feature points of an object in an image captured by a camera; a second step of determining a reference coordinate system based on the feature points; a third step of detecting the movement of the camera by at least one of a gyro sensor and an acceleration sensor; a fourth step of correcting the display position of the content based on the movement of the camera when the content is displayed based on the reference coordinate system. A video display method for performing the above.

14. In the video display method according to Claim 13, the correction includes at least one of a process of correcting the reference coordinate system to a corrected coordinate system and a process of correcting the coordinates of the content. A video display method.

15. In the video display method according to Claim 13, the camera and the sensor for detecting the movement of the camera are provided in a video display device worn by a user, and the content is simultaneously displayed when the user visually recognizes the object. A video display method.

Citation Information

Patent Citations

  • Hybrid registration method and system for target objects of mobile augmented reality (MAR) system

    CN102214000A

  • Information processor, information processing method, and computer program

    JP2008304269A

  • Information processor, information processing method and program

    JP2014199527A

  • Image display device, image display method, and image display program

    JP2015060071A

  • Information device for drawing ar objects based on predictive camera attitude in real time, program and method

    JP2016019199A