Video processing device, video processing method, and video processing program

The image processing device addresses the challenge of inaccurate AR content positioning by detecting moving objects, restoring environmental images, and estimating user gaze to ensure precise AR content display in mixed camera images.

JP7740563B2Active Publication Date: 2025-09-17NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024542503
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-08-24
Publication Date
2025-09-17
Estimated Expiration
2042-08-24

AI Technical Summary

Technical Problem

AR devices struggle to accurately display AR content in the correct position when the camera image contains a mixture of environmental and moving object images, leading to inaccurate superimposition.

Method used

An image processing device that includes an image acquisition unit, moving object detection, environmental image restoration, gaze estimation, and rendering units to calculate and display AR content accurately by complementing the moving object area with environmental images and estimating user gaze based on sensor data.

Benefits of technology

Enables accurate display of AR content even when the camera image includes both environmental and moving object areas, ensuring precise superimposition of AR content on the user's field of view.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007740563000001
    Figure 0007740563000001
  • Figure 0007740563000002
    Figure 0007740563000002
  • Figure 0007740563000003
    Figure 0007740563000003
Patent Text Reader

Abstract

A video processing device according to one embodiment is to be worn by a user and comprises an image acquisition unit that acquires a captured image that has been captured by a camera and includes a mobile body being ridden by the user and the surroundings, a mobile body detection unit that detects the mobile body from the captured image, a surroundings image reconstruction unit that supplements an image of the surroundings at the portion of the detected mobile body when the area of the detected mobile body is smaller than a prescribed threshold value, a line of sight estimation unit that estimates the line of sight of the user on the basis of the captured image, a line of sight movement detection unit that acquires sensor data from a sensor of the video processing device and detects movement of the line of sight of the user from the estimated line of sight on the basis of the sensor data, a rendering unit that calculates the appearance of AR content on the basis of the movement of the estimated line of sight, and an output control unit that performs control to display the calculated AR content.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a video processing device, a video processing method, and a video processing program. [Background technology]

[0002] Users of an Augmented Reality (AR) system can view the real world through a mobile terminal or AR device. In this case, content such as navigation information or 3D data (hereinafter referred to as AR content) is presented as additional information to the real world. In other words, users of the AR system can see the AR content superimposed on the real world, and can use the information in this content. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Explains the mechanisms and principles of AR (Augmented Reality) technology in an easy-to-understand way, and is available online.<URL:https: / / tech-camp.in / note / technology / 18791 / #i-3> [Non-patent document 2] Object Detection, Internet<URL:https: / / ja.wikipedia.org / wiki / %E7%89%A9%E4%BD%93%E6%A4%9C%E5%87%BA> [Non-patent document 3] Badminton Competition x Ultra-Realistic Communication Technology Kirari!, NTT Technical Journal vol.33 No.10 pp.24-29, 2021 [Non-patent document 4] Vehicle speed detection from a single motion blurred image, Internet<URL:https: / / dl.acm.org / doi / abs / 10.1016 / j.imavis.2007.04.004> Summary of the Invention [Problem to be solved by the invention]

[0004] For example, when a user using an AR system is traveling on a moving vehicle, the image displayed by the AR device is a mixture of the camera image showing the environment (the scenery in front of the bicycle) and the moving vehicle (part of the bicycle's body).

[0005] Conventional self-localization processing assumes that the camera image is a reflection of the environment, which makes it difficult to accurately process the image. As a result, AR devices are unable to display AR content in the correct position.

[0006] This invention was made with the above-mentioned circumstances in mind, and its purpose is to provide a technology that enables a mobile terminal or AR device to display AR content in an accurate position even when the camera image contains a mixture of images of the environment and images of a moving object. [Means for solving the problem]

[0007] In order to solve the above problem, one aspect of the present invention is an image processing device worn by a user, comprising: an image acquisition unit that acquires a captured image captured by a camera, the image including a moving object on which the user is riding and the environment; a moving object detection unit that detects the moving object from the captured image; an environmental image restoration unit that, if the area of ​​the detected moving object is smaller than a predetermined threshold, complements an image of the environment in the part of the detected moving object; a gaze estimation unit that estimates the user's gaze based on the captured image; a gaze movement detection unit that acquires sensor data from a sensor provided in the image processing device and detects movement of the user's gaze starting from the estimated gaze based on the sensor data; a rendering unit that calculates how AR content will appear based on the estimated gaze movement; and an output control unit that controls the calculated AR content to be displayed. [Effects of the Invention]

[0008] According to one aspect of the present invention, even if the camera image contains a mixture of areas showing the environment and areas showing a moving object, the mobile terminal or AR device can display the AR content at the correct position, thereby making it possible to accurately present the AR content to the user. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram showing an example of a hardware configuration of a video processing device according to the first embodiment. [Figure 2] FIG. 2 is a block diagram showing the software configuration of the video processing device according to the first embodiment in relation to the hardware configuration shown in FIG. [Figure 3] FIG. 3 is a flowchart showing an example of an operation performed by the video processing device to display AR content at a correct position in a captured image. [Figure 4] FIG. 4 is a diagram showing an example of a captured image. [Figure 5] FIG. 5 is a diagram showing an example in which a moving object is detected in a captured image. [Figure 6] FIG. 6 is a diagram showing an example in which an image of a portion extracted as a moving object is complemented with a surrounding image. [Figure 7] FIG. 7 is a diagram showing an example of setting AR content corresponding to feature points. [Figure 8] FIG. 8 is a diagram showing an example of how the calculated AR content appears. [Figure 9] FIG. 9 is a diagram showing an example of an image captured by a camera in the second embodiment. [Figure 10] FIG. 10 is a block diagram showing the software configuration of the video processing device according to the second embodiment in relation to the hardware configuration shown in FIG. [Figure 11] FIG. 11 is a flowchart showing an example of an operation performed by the video processing device to display AR content at the correct position in a captured image. [Figure 12] FIG. 12 is a diagram showing an example of detecting unevenness on a road surface in a captured image. [Figure 13] FIG. 13 is a diagram showing an example in which unevenness on the road surface is blurred in a captured image. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Hereinafter, elements that are identical or similar to elements already described will be designated by the same or similar reference numerals, and duplicate descriptions will generally be omitted. For example, when there are multiple identical or similar elements, a common reference numeral may be used to describe each element without distinguishing between them, or a subnumber may be used in addition to the common reference numeral to describe each element with distinction between them.

[0011] [First embodiment] (composition) FIG. 1 is a block diagram showing an example of the hardware configuration of a video processing device 1 according to the first embodiment. The image processing device 1 is a computer that analyzes input data and generates and outputs output data. The image processing device 1 may be, for example, an AR device including AR glasses, smart glasses, or other wearable devices. In other words, the image processing device 1 may be a device worn by a user.

[0012] As shown in Fig. 1, the video processing device 1 includes a control unit 10, a program storage unit 20, a data storage unit 30, a communication interface 40, and an input / output interface 50. The control unit 10, the program storage unit 20, the data storage unit 30, the communication interface 40, and the input / output interface 50 are communicatively connected to one another via a bus. Furthermore, the communication interface 40 may be communicatively connected to an external device via a network. Furthermore, the input / output interface 50 is communicatively connected to an input device 2, an output device 3, a camera 4, and an inertial sensor 5.

[0013] The control unit 10 controls the video processing device 1. The control unit 10 includes a hardware processor such as a central processing unit (CPU). For example, the control unit 10 may be an integrated circuit capable of executing various programs.

[0014] The program storage unit 20 may use, as a storage medium, a combination of nonvolatile memory that can be written to and read from at any time, such as an EPROM (Erasable Programmable Read Only Memory), an HDD (Hard Disk Drive), or an SSD (Solid State Drive), and a nonvolatile memory such as a ROM (Read Only Memory). The program storage unit 20 stores programs necessary for executing various processes. That is, the control unit 10 can realize various controls and operations by reading and executing the programs stored in the program storage unit 20.

[0015] The data storage unit 30 is a storage that uses a combination of nonvolatile memory such as a HDD or memory card, which can be written to and read from at any time, and volatile memory such as RAM (Random Access Memory), as a storage medium. The data storage unit 30 is used to store data acquired and generated in the process of the control unit 10 executing programs and performing various processes.

[0016] The communication interface 40 includes one or more wired or wireless communication modules. For example, the communication interface 40 includes a communication module for wired or wireless connection to an external device via a network. The communication interface 40 may also include a wireless communication module for wireless connection to an external device such as a Wi-Fi access point or base station. Furthermore, the communication interface 40 may also include a wireless communication module for wireless connection to an external device using short-range wireless technology. In other words, the communication interface 40 may be any general communication interface as long as it can communicate with an external device under the control of the control unit 10 and send and receive various information including past performance data.

[0017] The input / output interface 50 is connected to the input device 2, the output device 3, the camera 4, the inertial sensor 5, etc. The input / output interface 50 is an interface that enables transmission and reception of information between the input device 2, the output device 3, and the multiple cameras 4 and the inertial sensors 5. The input / output interface 50 may be integrated with the communication interface 40. For example, the video processing device 1 and at least one of the input device 2, the output device 3, the camera 4, and the inertial sensor 5 may be wirelessly connected using short-range wireless technology or the like, and information may be transmitted and received using the short-range wireless technology.

[0018] The input device 2 may include, for example, a keyboard, a pointing device, or the like for the user to input various information including past performance data to the video processing device 1. The input device 2 may also include a reader for reading data to be stored in the program storage unit 20 or the data storage unit 30 from a memory medium such as a USB memory, or a disk device for reading such data from a disk medium.

[0019] The output device 3 includes a display that displays images captured by the camera 4 and AR content. The output device 3 may be integrated with the image processing device 1. For example, if the image processing device 1 is AR glasses or smart glasses, the output device 3 is part of the glasses.

[0020] The camera 4 is capable of capturing images of the environment, such as scenery, and may be a general camera 4 that can be attached to the video processing device 1. Here, the environment refers to a scenery that is generally captured. The camera 4 may be integrated with the video processing device 1. The camera 4 may output the captured image to the control unit 10 of the video processing device 1 via the input / output interface 50.

[0021] The inertial sensor 5 includes, for example, an acceleration sensor, an angular velocity sensor, a geomagnetic sensor, etc. For example, when the video processing device 1 is an AR device, the inertial sensor 5 senses the moving speed and head movement of a user wearing the AR device, and outputs sensor data corresponding to the sensing to the control unit 10.

[0022] FIG. 2 is a block diagram showing the software configuration of the video processing device 1 according to the first embodiment in relation to the hardware configuration shown in FIG. The control unit 10 includes an image acquisition unit 101, a moving object detection unit 102, an environmental image restoration unit 103, a feature point extraction unit 104, a gaze estimation unit 105, a gaze movement estimation unit 106, a sensor data control unit 107, an AR content drawing unit 108, and an output control unit 109.

[0023] The image acquisition unit 101 acquires the captured image captured by the camera 4. The image acquisition unit 101 may store the captured image in the image storage unit 301.

[0024] The moving object detection unit 102 detects a moving object from a captured image. The moving object detection unit 102 detects a moving object MB captured in the captured image. The moving object may be any object such as a motorized bicycle, an electric bicycle, a motorcycle, or a vehicle. Furthermore, the moving object may include only the moving object, or may also include a part of the body of a user riding the moving object, such as an arm. Furthermore, a detection method may use a general technique.

[0025] The environmental image restoration unit 103 complements the video of the part detected as a moving object. The environmental image restoration unit 103 replaces the image of the part detected as a moving object with the complemented image of the environment from the image of the environment of the captured image. Note that the complementing method may use a general technique.

[0026] The feature point extraction unit 104 extracts features within the environment of the captured image as feature points. For example, the feature point extraction unit 104 extracts features in the vicinity of a feature point space stored in a space storage unit 302 (described later).

[0027] The gaze estimation unit 105 estimates the gaze position by matching the feature points extracted by the feature point extraction unit 104 with the feature point space stored in the space storage unit 302. The method for estimating the gaze position will be described in detail later.

[0028] The gaze movement estimation unit 106 estimates gaze movement. The gaze movement estimation unit 106 estimates gaze movement that follows the movement of the user's head by moving the three-dimensional movement measured by the sensor data control unit 107 (described later) to a starting point of the gaze position received from the gaze estimation unit 105. That is, the gaze movement estimation unit 106 estimates the movement of the user's gaze from the gaze position estimated by the gaze estimation unit 105 based on the sensor data.

[0029] The sensor data control unit 107 acquires sensor data from the inertial sensor 5. Then, the sensor data control unit 107 measures the user's head movement, body movement, etc. from the acquired sensor data. For example, the sensor data control unit 107 measures the user's three-dimensional movement (e.g., the user's head movement) based on the sensor data.

[0030] The AR content rendering unit 108 calculates how the AR content will appear. The AR content rendering unit 108 sets the estimated gaze movement included in the gaze movement information received from the gaze estimation unit 105, i.e., the gaze of the destination, in the AR content space corresponding to the feature point space stored in the space storage unit 302, and calculates how the AR content will appear in the set space.

[0031] The output control unit 109 outputs the AR content information. The output control unit 109 controls the output device 3 to render the AR content. For example, the output control unit 109 controls the output device 3 to display the adjusted AR content on AR glasses or the like.

[0032] The data storage unit 30 includes an image storage unit 301 and a spatial storage unit 302 .

[0033] The image storage unit 301 may store the captured image acquired by the image acquisition unit 101. Here, the captured image stored in the image storage unit 301 may include information about the longitude and latitude in the real world where the captured image was captured, which information was acquired by the video processing device 1. Furthermore, the image storage unit 301 may automatically delete the captured image after a predetermined time has elapsed.

[0034] The space storage unit 302 stores a feature point space and an AR content space corresponding to the feature point space. The feature point space may be located at a predetermined position in the captured image, and for example, two positions may be set. Then, the AR content space may be set in the middle of the two feature point spaces.

[0035] (operation) First, a method for displaying AR content on a video processing device 1 (a mobile terminal or an AR device) used by a user in a typical AR system will be described.

[0036] The video processing device 1 extracts feature points from the captured image captured by the camera 4. Furthermore, the video processing device 1 estimates the user's line of sight (position and direction) by comparing data (hereinafter referred to as a feature point space) in which feature points are also extracted from previously captured images (for example, captured images from the previous frame or frames several frames earlier) with the extracted feature points. Here, the feature point space is a space constructed for a "surrounding space" set based on a predetermined usage scene. Therefore, the captured image is also a capture of the surrounding space. Therefore, it is assumed that the positional relationship of both the feature point space and the captured image has been identified based on the "surrounding space."

[0037] The video processing device 1 tracks the movement of the user's head based on the inertial data received from the inertial sensor 5, and follows the above-mentioned user's line of sight in real time.

[0038] Furthermore, the video processing device 1 positions the line of sight of the user, which is being tracked in real time, in the AR content space created in advance, and calculates how the AR content will appear from there.

[0039] Then, the video processing device 1 causes the output device 3 to render the calculated appearance.

[0040] In this way, the AR system generates AR content and displays the AR content on a smartphone or AR glasses, which is the image processing device 1. However, as described above, this method is based on the premise that the camera 4 captures the environment, and therefore, in the case of a captured image that includes both the environment and a moving object MB, the image processing device 1 may fail to display the AR content in the correct position.

[0041] Therefore, the following describes the operation of the video processing device 1 for displaying AR content at the correct position even in a captured image in which the environment and moving objects are mixed.

[0042] FIG. 3 is a flowchart showing an example of the operation of the video processing device 1 to display AR content at the correct position in the captured image. The control unit 10 of the video processing device 1 reads out and executes the program stored in the program storage unit 20, thereby realizing the operation of this flowchart.

[0043] This operation flow is started, for example, when the user inputs an instruction to display AR content, or when a predetermined condition is met and the control unit 10 outputs an instruction to display AR content. Alternatively, this operation flow may be started when the video processing device 1 is started and the camera 4 acquires a captured image. In this operation, the moving object is assumed to be a bicycle.

[0044] In step ST101, the image acquisition unit 101 acquires a captured image captured by the camera 4. The image acquisition unit 101 may store the captured image in the image storage unit 301. The captured image includes an environment and a moving object. Here, the environment may be a general landscape, as described above. Therefore, the environment refers to the part excluding the moving object.

[0045] FIG. 4 is a diagram showing an example of a captured image. In the example of Fig. 4, the image is captured while the user is riding a bicycle, which is a moving object. Therefore, the captured image includes the moving object and the environment. Here, the camera 4 is a camera 4 provided in the AR glasses, which are the image processing device 1, and the captured image is captured by this camera 4.

[0046] For simplicity, the user's arms and the handlebars and wheels of the bicycle are shown as shaded areas in the example of Fig. 4. Furthermore, in the example of Fig. 3, two buildings Bu are shown in the captured image. The user shown in the example of Fig. 4 is facing forward.

[0047] In step ST102, the moving object detection unit 102 detects a moving object MB from the captured image. Here, the moving object detection unit 102 may detect the moving object MB by performing object detection using a general method. For example, the object detection method may be the object detection method disclosed in Non-Patent Document 2. Therefore, a detailed description of the object detection method will be omitted here.

[0048] FIG. 5 is a diagram showing an example when a moving object MB is detected in a photographed image. 5, the bicycle and the user are detected as the moving object MB. That is, the parts detected as the moving object MB include the handlebars and wheels of the bicycle as well as the arms of the user.

[0049] In step ST103, the environment image restoration unit 103 complements the image of the portion detected as the moving object MB. The environment image restoration unit 103 complements the image of the surrounding environment in the image of the portion detected as the moving object MB. In other words, the environment image restoration unit 103 reproduces the surrounding image of the portion detected as the moving object MB. The environment image restoration unit 103 may perform the complementation using a general method. For example, the environment image restoration unit 103 may use the complementation technique disclosed in Non-Patent Document 3. Therefore, a detailed description of the complementation technique will be omitted here.

[0050] FIG. 6 is a diagram showing an example of a case where an image of a portion extracted as a moving object MB is complemented with a peripheral image. As shown in FIG. 6, by complementing the image of the moving object MB with the surrounding image, the environmental image restoration unit 103 can obtain an image that would have been captured if the moving object MB had not been present.

[0051] In step ST104, the gaze estimation unit 105 estimates the gaze of the user. First, the feature point extraction unit 104 extracts feature points from the captured image. The feature points may be, for example, two buildings Bu.

[0052] FIG. 7 is a diagram showing an example of setting AR content corresponding to feature points. As shown in FIG. 7, the space storage unit 302 preliminarily sets a "feature point space" in which two buildings Bu are feature points, and also stores an "AR content space" in which a smiley face, which is AR content, is placed midway between the two buildings Bu in the space. That is, the space storage unit 302 stores the feature point space and an AR content space corresponding to the feature point space. Here, in FIG. 7, the AR content is indicated by the reference symbol ARC. The example in FIG. 7 is just an example, and the space storage unit 302 may store AR content spaces together with various feature point spaces.

[0053] The feature point extraction unit 104 surveys the entire captured image and extracts specific feature points such as edges of objects or corners of objects. Here, the object may be the boundary of an object. For example, in FIG. 7, it is possible to extract the boundary of the building Bu, which is a specific feature point.

[0054] Then, the gaze estimation unit 105 estimates the position of the gaze by matching the feature points extracted by the feature point extraction unit 104 with the feature point space stored in the space storage unit 302. For example, the gaze estimation unit 105 may estimate the gaze using a vision-based AR technology or the like. The vision-based AR technology may be a general technology such as PTAM, SmartAR, or Microsoft Hololens, which is a markerless AR technology. Therefore, a detailed description of the AR technology will be omitted here. The gaze estimation unit 105 outputs the estimated gaze position to the gaze movement estimation unit 106.

[0055] For example, it is not efficient to estimate the gaze position by storing an image corresponding to the captured image in the space storage unit 302 and comparing the images. Therefore, the gaze estimation unit 105 can perform high-speed processing by extracting and simplifying feature points instead of the image itself and comparing them with the feature point space.

[0056] When the user is facing forward, the area of ​​the moving object captured in the captured image is small. Therefore, since the area of ​​the moving object in the captured image is small, the environmental image restoration unit 103 can complement the part detected as the moving object (high image restoration accuracy). As a result, the gaze estimation unit 105 can successfully estimate the gaze.

[0057] In step ST105, the gaze movement estimation unit 106 estimates gaze movement. First, the sensor data control unit 107 acquires sensor data from the inertial sensor 5. Then, the sensor data control unit 107 measures the user's head movement, body movement, etc. from the acquired sensor data. Specifically, for example, the inertial sensor 5 is an inertial measurement unit (IMU), and the sensor data control unit 107 acquires sensor data such as acceleration, angular velocity, and geomagnetism from the inertial sensor 5. Then, the sensor data control unit 107 may measure the user's three-dimensional movement (for example, the user's head movement) based on this data. Then, the sensor data control unit 107 outputs the measurement result to the gaze movement estimation unit 106.

[0058] The gaze movement estimation unit 106 estimates gaze movement that follows the movement of the user's head by moving the three-dimensional movement measured by the sensor data control unit 107 to a starting point of the gaze position received from the gaze estimation unit 105. That is, the gaze movement estimation unit 106 estimates the movement of the user's gaze from a starting point of the gaze position estimated by the gaze estimation unit 105 based on the sensor data. Then, the gaze movement estimation unit 106 outputs gaze movement information including the estimated gaze movement to the AR content rendering unit 108.

[0059] In step ST106, the AR content rendering unit 108 calculates how the AR content will appear. The AR content rendering unit 108 sets the estimated gaze movement included in the gaze movement information received from the gaze estimation unit 105, i.e., the gaze of the destination, in the AR content space corresponding to the feature point space stored in the space storage unit 302, and calculates how the AR content will appear in the set space.

[0060] FIG. 8 is a diagram showing an example of how the calculated AR content appears. As shown in FIG. 8, the AR content rendering unit 108 adjusts the appearance of the AR content based on the movement information of the line of sight movement, and outputs AR content information for rendering the adjusted AR content ARC to the output control unit 109.

[0061] In step ST107, the output control unit 109 outputs the AR content information. The output control unit 109 controls the output device 3 to render the AR content ARC. For example, the output control unit 109 controls the output device 3 to display the adjusted AR content ARC on AR glasses or the like.

[0062] (Operation and effect of the first embodiment) According to the first embodiment, the video processing device 1 can accurately present the AR content ARC to the user even when the image captured by the camera 4 contains a mixture of parts that show the environment and parts that show the moving object MB.

[0063] [Second embodiment] In the second embodiment, for example, when a user riding a moving object such as a bicycle looks down, the moving object occupies most of the image captured by the camera 4.

[0064] FIG. 9 is a diagram showing an example of an image captured by the camera 4 in the second embodiment. When a moving object MB occupies most of the captured image as shown in FIG. 9, the interpolation process may fail even if the process in the first embodiment is performed.

[0065] In the second embodiment, a method will be described that enables AR content to be accurately presented to a user even when the user looks down and a moving object occupies most of the captured image.

[0066] (composition) The hardware configuration of the video processing device 1 in the second embodiment may be the same as the hardware configuration in the first embodiment, so a duplicated description will be omitted here.

[0067] FIG. 10 is a block diagram showing the software configuration of the video processing device 1 according to the second embodiment in relation to the hardware configuration shown in FIG. In the second embodiment, the control unit 10 differs from the control unit 10 in the first embodiment in that it includes a moving object area calculation unit 110 and a road surface image analysis unit 111.

[0068] The moving object area calculation unit 110 calculates the area of ​​the moving object MB. The moving object area calculation unit 110 may calculate the area of ​​the moving object MB, or may calculate the proportion of the moving object MB that occupies the captured image. In addition, the moving object area calculation unit 110 determines whether the calculated area of ​​the moving object MB is equal to or greater than a threshold value.

[0069] The road surface image analysis unit 111 detects vertically elongated portions caused by unevenness in the environment shown in the captured image moving within the interval of the shutter speed of the camera 4. The road surface image analysis unit 111 may use a general method to detect unevenness (e.g., pebbles) shown in the environment of the captured image and detect deformation of the uneven pebbles. Furthermore, the road surface image analysis unit 111 estimates the moving speed of the moving object MB from the degree of deformation. Details of the method for estimating the moving speed will be described later.

[0070] The data storage unit 30 also differs from the first embodiment in that it includes a gaze storage unit 303. The gaze storage unit 303 may store information about the gaze estimated by the gaze estimation unit 105. Note that the stored information about the gaze may be deleted after a certain period of time has elapsed.

[0071] (operation) FIG. 11 is a flowchart showing an example of the operation of the video processing device 1 to display AR content at the correct position in a captured image. The control unit 10 of the video processing device 1 reads out and executes the program stored in the program storage unit 20, thereby realizing the operation of this flowchart.

[0072] This operation flow is started, for example, when the user inputs an instruction to display AR content or when a predetermined condition is satisfied and the control unit 10 outputs an instruction to display AR content. Alternatively, this operation flow may be started when the video processing device 1 is started and the camera 4 acquires a captured image.

[0073] Steps ST201 and ST202 may be the same as steps ST101 and ST102 described with reference to Fig. 3, and therefore a duplicated description will be omitted here. Note that in step ST202, the moving object detection unit 102 may output the captured image and information about the detected moving object MB to the moving object area calculation unit 110.

[0074] In step ST203, the moving object area calculation unit 110 calculates the area of ​​the moving object MB.

[0075] The moving object area calculation unit 110 may calculate the area of ​​the moving object MB, or may calculate the proportion of the moving object MB that occupies the captured image.

[0076] In step ST204, the moving object area calculation unit 110 determines whether the calculated area of ​​the moving object MB is equal to or greater than a threshold. If it is determined that the calculated area of ​​the moving object MB is not equal to or greater than the predetermined threshold, i.e., the area of ​​the moving object MB does not occupy most of the captured image, the process proceeds to step ST205. On the other hand, if it is determined that the calculated area is equal to or greater than the predetermined threshold, i.e., the area of ​​the moving object MB occupies most of the captured image, the process proceeds to step ST207. Note that, when the moving object area calculation unit 110 calculates the ratio, it may determine whether the ratio of the moving object MB is equal to or greater than a threshold. Steps ST205 and ST206 may be similar to steps ST103 and ST104 described with reference to FIG. 3, and therefore a duplicated description will be omitted here. However, the gaze estimation unit 105 may store information about the estimated gaze estimation together with time information in the gaze storage unit 303.

[0077] For example, as shown in FIG. 7, when the user is looking down, the moving object occupies a large portion of the captured image. In such a case, as in the first embodiment, even if the environmental image restoration unit 103 attempts to complement the portion detected as the moving object, the accuracy deteriorates and restoration cannot be performed correctly. As a result, the gaze estimation unit 104 fails to estimate the gaze. Therefore, as will be described below, the process estimates the user's gaze using a method that can estimate the gaze even when the user is looking down.

[0078] In step ST207, the road surface image analysis unit 111 detects portions of the environment shown in the captured image that have been deformed vertically due to movement within the shutter speed interval of the camera 4. The road surface image analysis unit 111 may detect irregularities (e.g., pebbles) in the environment shown in the captured image using a general method, and detect deformation of the pebbles. For example, an irregularity (pebbles) is detected as a rectangle circumscribing the irregularity. The road surface image analysis unit 111 then estimates the degree of deformation from the ratio of the length to the width of this rectangle. The road surface image analysis unit 111 then estimates the moving speed of the moving object MB from the degree of deformation. The road surface image analysis unit 111 estimates the speed of the moving object MB shown in the captured image based on, for example, the intensity of blur, the mounting angle of the camera 4, etc. Here, the speed of the moving object MB may be estimated using a general technique. For example, the method for estimating the speed of a moving body MB may be the method for estimating the speed of a moving body MB as disclosed in Non-Patent Document 4. Therefore, a detailed description of the method for estimating the speed of a moving body MB will be omitted here.

[0079] FIG. 12 is a diagram showing an example of detecting unevenness on a road surface in a captured image. In Fig. 12, unevenness on the road surface is indicated by reference symbol CC. As shown in Fig. 12, the road surface image analysis unit 111 detects unevenness such as pebbles on the road surface from the captured image.

[0080] FIG. 13 is a diagram showing an example in which unevenness on the road surface is blurred in a captured image. 13, the faster the speed of the moving object MB, the more blurred the detected pebble becomes. The road surface image analysis unit 111 may estimate the speed of the moving object MB using the degree of this blur.

[0081] In step ST208, the gaze estimation unit 105 estimates the gaze. The gaze estimation unit 105 acquires information about the gaze estimation stored previously from the current time from the gaze storage unit 303, and estimates the current gaze position based on the result of the gaze estimation and the movement speed.

[0082] Steps ST209 to ST211 may be the same as steps ST105 to ST107 described with reference to FIG. 3, and therefore a duplicated description will be omitted here.

[0083] For example, the AR content space is linked to a coordinate system such as longitude and latitude in the real world. Therefore, even if the gaze estimation method uses the distortion of unevenness in the captured image, it is possible to link and map the movement distance calculated from the speed and time with the distance between two points in the AR content space in the longitude and latitude coordinate system. Therefore, by performing the processes in steps ST209 to ST211, the video processing device 1 can correctly recognize the AR content space and accurately present the AR content ARC to the user.

[0084] (Operation and effect of the second embodiment) According to the second embodiment, the video processing device 1 can accurately present the AR content ARC to the user even when the image captured by the camera 4 contains a larger portion of the moving body MB than the portion showing the environment.

[0085] [Other embodiments] In the above embodiment, an example has been described in which a captured image taken by the camera 4 included in the video processing device 1 is used, but the captured image is not limited to being taken by the camera 4 included in the video processing device 1. For example, it may be an independent camera 4 connected to the video processing device 1. However, it is assumed that the camera 4 is installed in a location where it can capture an image that allows the user's line of sight to be estimated (for example, above the user's head, etc.).

[0086] The techniques described in the above embodiments can be stored as a program (software means) that can be executed by a computer on a storage medium such as a magnetic disk (e.g., a floppy disk, a hard disk, etc.), an optical disk (e.g., a CD-ROM, a DVD, an MO, etc.), or a semiconductor memory (e.g., a ROM, a RAM, a flash memory, etc.), and can also be distributed by transmitting the program via a communication medium. The program stored on the medium also includes a configuration program that configures the software means (including not only executable programs but also tables and data structures) that the computer executes. The computer that implements this device loads the program stored on the storage medium and, in some cases, configures the software means using the configuration program, and executes the above-described processing by controlling the operation of the software means. The term "storage medium" as used herein is not limited to storage media for distribution, but also includes storage media such as magnetic disks and semiconductor memories installed inside the computer or in devices connected via a network.

[0087] In short, this invention is not limited to the above-described embodiments, and various modifications can be made in the implementation stage without departing from the spirit of the invention. Furthermore, the embodiments may be implemented in combination as appropriate as possible, and in such cases, the combined effects can be obtained. Furthermore, the above-described embodiments include inventions at various stages, and various inventions can be extracted by appropriately combining the disclosed multiple constituent elements. [Explanation of symbols]

[0088] 1...Video processing device 2...Input device 3...Output device 4. Camera 5...Inertial sensor 10...Control unit 101...Image acquisition unit 102...Moving object detection unit 103...Environment image restoration unit 104...Feature point extraction unit 105... Gaze estimation section 106... Gaze movement estimation unit 107...Sensor data control unit 108…AR content drawing section 109...Output control section 110...Moving object area calculation unit 111...Road surface image analysis section 20...Program memory section 30...Data storage unit 301...Image storage unit 302…Spatial storage unit 303…Gaze memory unit 40...Communication interface 50...Input / output interface MB: Mobile Bu…Building

Claims

1. A video processing device worn by a user, an image acquisition unit that acquires a captured image including the moving object on which the user is riding and the environment, the captured image being captured by a camera; a moving object detection unit that detects the moving object from the captured image; an environmental image restoration unit that, when the area of ​​the detected moving object is smaller than a predetermined threshold, complements an image of the environment in the part of the detected moving object; a gaze estimation unit that estimates a gaze of the user based on the captured image; a gaze movement detection unit that acquires sensor data from a sensor included in the video processing device and detects a movement of the user's gaze from the estimated gaze point based on the sensor data; a rendering unit that calculates how the AR content appears based on the estimated movement of the line of sight; an output control unit that controls the calculated AR content to be displayed; A video processing device comprising:

2. a moving object area calculation unit that calculates an area of ​​the detected moving object and determines whether the area exceeds a predetermined threshold; The image processing device according to claim 1 , further comprising an image analysis unit that, if the area exceeds a predetermined threshold, detects unevenness in the environment of the captured image, detects deformation of the unevenness, and estimates the speed of the moving object based on the degree of the deformation.

3. The image processing device according to claim 2 , wherein the gaze estimation unit estimates the current gaze of the user based on the speed of the moving object and a previous gaze estimation.

4. a feature point extraction unit that extracts feature points within the environment in the captured image; a storage unit that stores the feature point space; The image processing device according to claim 1 , further comprising: a gaze estimation unit that estimates a gaze based on the extracted feature points and the feature point space.

5. the storage unit further stores an AR content space corresponding to the feature point space; The image processing device according to claim 4 , wherein the rendering unit sets a destination of the line of sight in the AR content space, and calculates how the AR content will appear in the set space.

6. The image processing device according to claim 4 , wherein the feature point extraction unit extracts features at edges or corners of objects in the captured image as feature points.

7. A video processing method executed by a processor of a video processing device worn by a user, comprising: Acquiring a captured image including a moving object on which the user is riding and an environment, the captured image being captured by a camera; Detecting the moving object from the captured image; if the area of ​​the detected moving object is smaller than a predetermined threshold, completing an image of the environment in the part of the detected moving object; estimating a line of sight of the user based on the captured image; acquiring sensor data from a sensor included in the video processing device; Detecting a movement of the user's gaze starting from the estimated gaze point based on the sensor data; Calculating how the AR content appears based on the estimated movement of the line of sight; Controlling the display of the calculated AR content; A video processing method comprising:

8. 1. A video processing program comprising instructions to be executed by a processor of a video processing device worn by a user, the instructions comprising: Acquiring a captured image including a moving object on which the user is riding and an environment, the captured image being captured by a camera; Detecting the moving object from the captured image; if the area of ​​the detected moving object is smaller than a predetermined threshold, completing an image of the environment in the part of the detected moving object; estimating a line of sight of the user based on the captured image; acquiring sensor data from a sensor included in the video processing device; Detecting a movement of the user's gaze starting from the estimated gaze point based on the sensor data; Calculating how the AR content appears based on the estimated movement of the line of sight; Controlling the display of the calculated AR content; A video processing program comprising:

Citation Information

Patent Citations

  • Image processing apparatus and image processing method

    JP2019128824A

  • Information processing device, imaging device, and imaging system

    JP2019145021A

  • Work vehicle display system and work vehicle display method

    JP2022072598A

  • Information processing device, information processing method, and program

    WO2018207426A1

  • Information processing device, information processing method, and program

    WO2019087658A1