Information processing apparatus, information processing method, and program
The information processing device addresses jitter and responsiveness issues in mixed reality and augmented reality by adjusting filter weights based on distance or attributes, enhancing user experience.
Patent Information
- Application Number
- JP2024122146
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2026-02-10
AI Technical Summary
Existing image-based measurement methods for imaging devices in mixed reality and augmented reality suffer from variability in measurement results due to noise, leading to jitter and decreased responsiveness, which degrades the user experience.
An information processing device that adjusts the filter weight based on the distance between the imaging device and virtual content, or attributes like movement speed, to reduce jitter while maintaining responsiveness.
The device effectively reduces jitter perception while minimizing the decrease in responsiveness to displayed virtual content by dynamically adjusting filter weights based on distance, movement speed, or other attributes.
Smart Images

Figure 2026020686000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] Conventionally, image-based measurement of the position and orientation of an imaging device has been used for various purposes. For example, alignment between real space and virtual objects in mixed reality (MR) technology and augmented reality (AR) technology can be cited. The measured position and orientation of the imaging device contain noise, which causes variability in the measurement results. For example, in MR, if the variability is large, the alignment of the virtual object becomes unstable, which manifests as jitter and can degrade the user experience.
[0003] Therefore, a method is known for reducing jitter by filtering (applying a filter to the position and orientation) using past time-series position and orientation measurement results. However, while a filter (time-series filter) can reduce jitter, it also leads to a decrease in responsiveness. In other words, achieving both stability and responsiveness is an issue.
[0004] Non-Patent Document 1 discloses a method for achieving both stability and responsiveness by adjusting the filter weight according to the moving speed of an imaging device, based on the human perceptual characteristic that the faster the moving speed of an imaging device, the lower the perceptual ability, and the slower the moving speed, the higher the perceptual ability. Here, the filter weight is used to adjust the stability and responsiveness. When the filter weight is large, the degree of jitter reduction improves but the responsiveness decreases. On the other hand, when the filter weight is small, the responsiveness increases but the degree of jitter reduction decreases. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Gery Casiez et.al,1£ Filter:A Simple Speed-based Low-pass Filter for Noisy Input in Interactive Systems [Non-patent document 2] Raul Mur-Artal et.al,ORB-SLAM:A Versatile and Accurate Monocular SLAM System.IEEE Transactions on Robotics Summary of the Invention [Problem to be solved by the invention]
[0006] In the method disclosed in Non-Patent Document 1, emphasis is placed on responsiveness, and the weight of the filter is uniformly reduced, which makes it easier for the user to perceive jitter depending on the virtual content to be displayed.
[0007] SUMMARY OF THE INVENTION It is therefore an object of the present invention to provide an information processing device that can reduce jitter while preventing a user from perceiving a decrease in responsiveness to displayed virtual content. [Means for solving the problem]
[0008] An information processing device as one aspect of the present invention includes an acquisition means for acquiring an image from an imaging device, a derivation means for deriving the position and orientation of the imaging device from the image, a storage means for storing information on virtual content to be rendered based on the position and orientation, a determination means for determining a weight of a filter based on the information on the virtual content, and an application means for applying the filter to the position and orientation based on the weight of the filter.
[0009] Other objects and features of the present invention are illustrated in the following examples. [Effects of the Invention]
[0010] According to the present invention, it is possible to provide an information processing device that can reduce jitter while suppressing a user's perception of a decrease in responsiveness to displayed virtual content. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 2 is a hardware configuration diagram of an information processing device in each embodiment. [Figure 2] 1 is a conceptual diagram showing a usage scene of an information processing device according to first and second embodiments. [Figure 3] FIG. 1 is a block diagram of an information processing device according to first and second embodiments. [Figure 4] 1 is a flowchart showing an information processing method in the first and second embodiments. [Figure 5] FIG. 10 is a block diagram of an information processing device as a modified example of the second embodiment. [Figure 6] FIG. 10 is a conceptual diagram showing a user interface as a modified example of the second embodiment. [Figure 7] FIG. 10 is a block diagram of an information processing device according to a third embodiment. [Figure 8] 11 is a flowchart illustrating an information processing method according to a third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0013] Prior to describing each embodiment of the present invention, a hardware configuration capable of realizing an information processing device of each embodiment will be described with reference to Fig. 1. Fig. 1 is a hardware configuration diagram of the information processing device in each embodiment. A CPU (Central Processing Unit) 10 controls each unit connected to a bus 60 via the bus 60. An input I / F (Interface) 40 acquires an input signal in a format that can be processed by the information processing device from an external device (such as an imaging device, a display device, or an operation device). An output I / F (Interface) 50 outputs an output signal to the external device (such as a display device) in a format that can be processed by the external device.
[0014] Programs for realizing the functions of each embodiment are stored in a storage medium such as a read-only memory (ROM) 20. The ROM 20 stores an operating system (OS) and device drivers. A memory such as a random access memory (RAM) 30 temporarily stores these programs. The CPU 10 executes the programs stored in the RAM 30, thereby performing processing according to the flowcharts described below and realizing the functions of each embodiment. However, each embodiment is not limited to this, and instead of software processing using the CPU 10, the functions of each embodiment can also be realized using hardware having a calculation unit or circuit corresponding to the processing of each functional unit. [Example]
[0015] In the first embodiment, an example in which the information processing device of the present invention is applied to a head-mounted display (HMD) will be described. The head-mounted display in this embodiment relates to a technology called augmented reality (AR) or mixed reality (MR).
[0016] FIG. 2 is an image diagram showing a usage scene of the head-mounted display 2 equipped with the information processing device of this embodiment. The head-mounted display 2 displays an image in which virtual content is superimposed on an image. Virtual content 3 is a mug-shaped virtual content displayed by the head-mounted display 2. Virtual content 4 is a virtual content of a user operation panel displayed by the head-mounted display 2. Here, virtual content 3 and virtual content 4 are displayed as if they are floating in the air in front of the head-mounted display 2. Virtual content 5 is a virtual content of a car displayed by the head-mounted display 2. Virtual content 5 is displayed as if it is moving on the floor in front.
[0017] We aim to provide a wearer (user) of the head-mounted display 2 with an experience in which objects (virtual objects) virtually rendered using computer graphics technology appear as if they were actually present. To achieve this, it is necessary to accurately estimate the position and orientation of an imaging device that captures images to be displayed on the head-mounted display 2. The imaging device has a variable position and orientation and captures captured images of a subject. Hereinafter, a three-dimensional coordinate system with the optical center of the imaging device as the origin, the optical axis direction as the Z axis, the horizontal direction of the image as the X axis, and the vertical direction of the image as the Y axis is defined as the imaging device coordinate system or imaging coordinate system. The position and orientation of the imaging device refer to the position and orientation of the imaging coordinate system (e.g., the position of the origin and the direction of the Z axis) relative to a reference coordinate system (world coordinate system) defined in the space (scene) where the image is captured. The position and orientation of the imaging device have six degrees of freedom (three degrees of freedom for position and three degrees of freedom for orientation).
[0018] In this embodiment, the filter weight (strength) is determined (changed) based on the distance between the virtual content and the imaging device. Specifically, the closer the distance, the greater the filter weight, and the farther the distance, the smaller the filter weight. Human perception characteristics dictate that the closer the distance, the larger the display area on the image. For this reason, while jitter is more easily perceived, the amount of movement on the image due to changes in position and posture is smaller, making it less likely that a decrease in responsiveness will be perceived. The filter in this embodiment is designed based on the above-mentioned human perception characteristics.
[0019] 3 is a block diagram of an information processing device 1 in this embodiment. The information processing device 1 includes a storage unit (storage means) 102, an acquisition unit (acquisition means) 103, a derivation unit (derivation means) 104, a determination unit (determination means) 105, and an application unit (application means) 106. The acquisition unit 103 is connected to the imaging device 101, and the application unit 106 and storage unit 102 are connected to a display unit (display means) 107.
[0020] The imaging device 101 is a camera capable of capturing monochrome stereo images, and captures images at regular intervals, although the type of imaging device 101 is not limited to this.
[0021] The storage unit 102 stores a three-dimensional map required for calculating the position and orientation of virtual content for rendering the virtual content, a three-dimensional model, and the position and orientation of the image capture device 101. The three-dimensional map refers to, for example, a group of key frames described in Non-Patent Document 2. The acquisition unit 103 acquires captured images captured by the image capture device 101. The derivation unit 104 calculates (derives) the position and orientation of the image capture device 101 by applying SLAM (Simultaneous Localization and Mapping) to the captured images acquired by the acquisition unit 103.
[0022] The determination unit 105 determines (changes) filter weights to be applied to variables representing six degrees of freedom of the position and orientation, based on the position and orientation of the image capture device 101 calculated by the derivation unit 104. The application unit 106 applies the filter to the position and orientation of the image capture device 101 calculated by the derivation unit 104, based on the filter weights determined by the determination unit 105, and updates the position and orientation of the image capture device 101. In other words, the application unit 106 applies the filter to the position and orientation of the image capture device 101 so as to reduce jitter that occurs when aligning a captured image with virtual content.
[0023] The display unit 107 receives the position and orientation of the virtual content and the three-dimensional model held by the holding unit 102, and the position and orientation of the image capture device 101 calculated by the application unit 106, and renders the virtual content on the screen.
[0024] Note that the processing procedures shown in the flowcharts below are merely examples. For example, it is possible to combine the procedures, combine multiple processes, or subdivide the processes. It is also possible to extract each process individually and have it function as a single functional element, and to use it in combination with processes other than those shown.
[0025] 4 is a flowchart showing a processing procedure (information processing method) in this embodiment. Each step in FIG.
[0026] First, in step S1010, the information processing device 1 performs an initialization process. This puts the information processing device 1 into an operable state. Next, in step S1020, the acquisition unit 103 acquires one frame of a captured image captured by the imaging device 101. Next, in step S1030, the derivation unit 104 derives the position and orientation of the imaging device 101 when the captured image was captured, using the three-dimensional map stored in the storage unit 102 and the captured image acquired in step S1020. Note that various known methods can be used as the derivation method. In this embodiment, the method for deriving the position and orientation of the imaging device disclosed in Non-Patent Document 2 is used. That is, the position and orientation of the imaging device 101 are repeatedly corrected so as to reduce the difference between the image positions of the feature points on the input image calculated based on the three-dimensional positions of the feature points and the derived position and orientation of the imaging device 101, and the image positions of the feature points on the captured image.
[0027] Subsequently, in step S1040, the determination unit 105 determines the filter weight α (filter strength) using the position and orientation of the image capture device derived in step S1030 and the drawing position of the virtual content held by the holding unit 102. The filter used in this embodiment uses the following formula (1) based on Non-Patent Document 1. However, the type of filter is not limited to this.
[0028]
number
[0029] In formula (1), X t represents the position and orientation of the image capture device calculated by the derivation unit 104 at time t (latest time), and X^ t represents the position and orientation of the image capture device at time t after the filter application unit 106 applies the filter. t-1represents the position and orientation of the image capture device 101 at time t-1 after the application unit 106 applies the filter. α represents the weight of the filter, and in equation (1), represents the weight of the position and orientation at time t and the position and orientation at time t-1. As α increases, jitter is reduced, but responsiveness also decreases.
[0030] In this embodiment, the distance d between the image capture device 101 and the virtual content is calculated (acquired) using the position and orientation of the image capture device 101 and the drawing position of the virtual content stored in the storage unit 102. The larger the distance d, the smaller the filter weight α is set, and the smaller the distance d, the larger the filter weight α is set. Specifically, the upper limit value d for the distance is max and the lower limit d min and the upper limit α of the weight α max and the lower limit α min and the weight α is determined using the following equation (2).
[0031]
number
[0032] Note that α≦α min In the case of α=α min , α ≥ α max In the case of α=α max The value of α is determined for each virtual content to be drawn. Note that the upper and lower limits of α are set to prevent the responsiveness from decreasing drastically due to the filter weight being too large.
[0033] A specific example will be described with reference to FIG. 2. It is assumed that the distance between the mug of virtual content 3 and the head-mounted display 2 is 3.0 [m], and the distance between the operation panel of virtual content 4 and the head-mounted display 2 is 0.1 [m]. min =0.3, α max =0.7, d min =1.0[m], d max= 5.0 [m]. In this case, when the filter of formula (2) is applied to virtual content 3, d = 3.0, so α = 1 - (3.0 - 1.0) / (5.0 - 1.0) = 0.5. When the filter weight of virtual content 4 is calculated using formula (2), d = 0.1 ( <d min ) and α=α max =0.7.
[0034] As a result, the filter weight α is determined to be α=0.3 for virtual content 3 and α=0.7 for virtual content 4, with the filter weight being greater for virtual content 4, which is closer to the head-mounted display 2, than for virtual content 3. Therefore, virtual content 4, which is closer and therefore more likely to perceive jitter and less likely to perceive a decrease in responsiveness, has a larger filter weight, which reduces responsiveness but has a stronger jitter suppression effect. On the other hand, virtual content 3, which is farther away and therefore more likely to perceive a decrease in responsiveness and less likely to perceive jitter, has a smaller filter weight, which reduces the jitter suppression effect but can suppress a decrease in responsiveness.
[0035] Next, in step S1050, the information processing device 1 determines whether the calculation of the filter weight α for all virtual content to be rendered is complete. If it is determined that the calculation of the filter weight α for all virtual content is complete, the process proceeds to step S1060. On the other hand, if it is determined that the calculation of the weight α is not complete, the process proceeds to step S1040.
[0036] In step S1060, the application unit 106 applies the filter weight α determined in step S1040 to equation (1) for each virtual content, and calculates X^ representing the position and orientation of the image capture device 101. t Calculate.
[0037] Next, in step S1070, the information processing device 1 determines whether or not to terminate the system. If it is determined that an instruction to terminate has been input by an instruction input means (not shown), the system is terminated. On the other hand, if it is determined that an instruction to terminate has not been input, the process returns to step S1020, and the processes of steps S1020 to S1060 are repeatedly performed each time a captured image is acquired from the imaging device 101.
[0038] In this embodiment, the filter weight is adjusted based on the position and orientation of the image capture device 101 and the distance to the virtual content. As a result, the filter weight is increased for virtual content that is close, making it possible to suppress jitter. On the other hand, the filter weight is decreased for virtual content that is far away, making it possible to suppress a decrease in responsiveness. In other words, by adjusting the filter weight according to the distance from the content to be displayed, it is possible to reduce jitter while suppressing the user's perception of a decrease in responsiveness to the displayed virtual content.
[0039] In this embodiment, the derivation unit 104 derives the position and orientation of the imaging device using SLAM using camera images, which is a method disclosed in Non-Patent Document 2, but the position and orientation estimation method is not limited to this.
[0040] For example, estimation may be performed using visual inertial SLAM, which combines an inertial sensor (Inertial Measurement Unit, IMU) and a camera, or using measurement data from a measuring device other than a camera, such as LiDAR (Light Detection and Ranging).
[0041] The derivation method is not limited to SLAM, and optical markers (e.g., two-dimensional QR codes) (registered trademark) may be placed in the environment, the optical markers may be recognized in an image, and the position and orientation of the imaging device relative to the optical markers may be estimated. Also, the position and orientation of the head-mounted display may be estimated using a camera installed in the environment, such as a motion capture camera.
[0042] In addition, when virtual content is superimposed on a physical object, a three-dimensional model of the target physical object may be stored in advance, and model-based estimation may be performed in which the position and orientation are measured by matching the measured three-dimensional point cloud with the three-dimensional model. [Example]
[0043] In the first embodiment, the determination unit 105 determined (changed) the filter weight based on the distance between the position of the imaging device 101 and the rendering position of the virtual content. In the second embodiment, a method for determining the filter weight based on attribute information of the virtual content will be described. Note that the attribute information of the virtual content in this embodiment is information on the moving speed. However, this embodiment is not limited to this, and other information may be used as long as it expresses the characteristics of the virtual content, such as the size or shape of the virtual content. Furthermore, the moving speed increases when the virtual content itself moves quickly in space, such as in animation, and decreases when the virtual content is positioned in space.
[0044] In this embodiment, the filter weight is increased as the moving speed of the virtual content decreases, and decreased as the moving speed increases. Human perception characteristics are such that the slower the moving speed (closer to a stationary state), the smaller the amount of movement on the image, making it easier to perceive even slight fluctuations due to jitter, but also making it harder to perceive a decrease in responsiveness. The filter in this embodiment is designed based on the above-described human perception characteristics.
[0045] The information processing method of this embodiment is realized by the configuration of the information processing device 1 described with reference to Fig. 3, similarly to the first embodiment. Note that the input and output data of the acquisition unit 103, derivation unit 104, determination unit 105, and application unit 106 are the same as those of the first embodiment, and therefore description thereof will be omitted here. The holding unit 102 is different from that of the first embodiment, and therefore the differences from the first embodiment will be described.
[0046] As in the first embodiment, the storage unit 102 stores the position and orientation of the virtual content for rendering the virtual content, a three-dimensional model, and a three-dimensional map required by the derivation unit 104 to calculate the position and orientation of the image capture device 101. The storage unit 102 also stores the movement speed as attribute information of the virtual content. Note that the movement speed is calculated using the difference in the rendering position of the virtual content at each time.
[0047] Next, a processing flow (information processing method) of the information processing device 1 in this embodiment will be described. The flowchart of this embodiment is the same as that of FIG. 4 described in embodiment 1. In this embodiment, only the processing of step S1040 differs from embodiment 1, and therefore, a description of the processing other than step S1040 will be omitted.
[0048] In step S1040, the determination unit 105 determines a filter weight α (filter strength). The filter is determined using equation (1) as in the first embodiment. The filter weight α in this embodiment is calculated based on the attribute information of the virtual content. Specifically, the filter weight α can be expressed by the following equation (3) using the moving speed v, which is the attribute information of the virtual content stored in the storage unit 102.
[0049]
number
[0050] α max is the upper limit of the weight α, α min is the lower limit of the weight α, which is set in advance. min In the case of α=α min and α ≥ α max In the case of α=α max Also, v max is the upper limit of the movement speed, v minis the lower limit of the movement speed, and is set in advance. The value of weight α is determined for each virtual content to be drawn. The upper and lower limit values of weight α are set in order to prevent an extreme drop in responsiveness caused by making the filter weight too large.
[0051] A specific example will be described with reference to Fig. 2. It is assumed that the moving speed of the operation panel of the virtual content 4 is 0.0 [m / s], and the moving speed of the car of the virtual content 5 is 1.5 [m / s]. min =0.3, α max =0.7, v min =0.0[m / s], v max = 3.0 [m / s]. In this case, when the filter of formula (3) is applied to virtual content 4, v = 0.0 [m / s] (≦ v min ), so α=α max = 0.7, and when the filter of formula (3) is applied to virtual content 5, v = 1.5 [m / s]. Therefore, α = 1 - (1.5 - 0.0) / (3.0 - 0.0) = 0.5.
[0052] As a result, the filter weight α is determined to be α=0.7 for virtual content 4 and α=0.5 for virtual content 5. The filter weight α is small for virtual content that moves quickly in space, such as the car in virtual content 5. On the other hand, the filter weight α is large for virtual content that is fixed in space, such as the operation panel in virtual content 4. Therefore, virtual content 4, which moves slowly and therefore is more likely to perceive jitter and less likely to perceive a decrease in responsiveness, has a large filter weight, which reduces responsiveness but has a strong jitter suppression effect. On the other hand, virtual content 5, which moves quickly and therefore is more likely to perceive a decrease in responsiveness and less likely to perceive jitter, has a small filter weight, which reduces the jitter suppression effect but can suppress a decrease in responsiveness.
[0053] In this embodiment, the filter weight is adjusted based on the attributes of the virtual content. As a result, the filter weight is increased for virtual content with a slow movement speed, thereby suppressing jitter. On the other hand, the filter weight is decreased for virtual content with a fast movement speed, thereby suppressing a decrease in responsiveness. In other words, by adjusting the filter weight according to the attributes of the virtual content to be displayed, it is possible to reduce jitter while minimizing the user's perception of a decrease in responsiveness for the displayed virtual content.
[0054] In this embodiment, the moving speed is used as an attribute of the virtual content, but the present invention is not limited to this and may be, for example, the size or shape of the virtual content. For example, a case where size is used as an attribute of the virtual content will be described. Here, the size of the virtual content refers to the volume of the three-dimensional model of the virtual content, or the area that the three-dimensional model occupies in the image when projected onto the image.
[0055] In this example, the rendering position of the virtual content stored in the storage unit 102 and the position and orientation of the imaging device derived by the derivation unit 104 are used to increase the filter weight α as the area of the virtual content projected onto the image becomes smaller. Conversely, the larger the area, the smaller α is set. This is based on the human perception characteristic that the smaller the area, the more quickly the movement is perceived in the image, and therefore the more likely it is to be perceived as jitter. By adjusting the filter weight according to the attributes (size) of the virtual content to be displayed, it is possible to reduce jitter while minimizing the user's perception of jitter and reduced responsiveness of the displayed virtual content.
[0056] Next, a case where shape is used as an attribute of virtual content will be described. Here, the shape of virtual content represents the degree of unevenness included in the three-dimensional model of the virtual content.
[0057] In this example, normal vectors are acquired from a three-dimensional model of the virtual content stored in the storage unit 102, and the degree of unevenness of the virtual content is calculated from the variation in the angle between adjacent normal vectors. The greater the degree of unevenness, the greater the filter weight α. Conversely, the smaller the degree of unevenness, the smaller α is set. Alternatively, the degree of unevenness may be calculated by calculating the normal vectors of the contours on the image when the three-dimensional model is projected onto the image, and then calculating the degree of unevenness of the virtual content from the variation in the angle between adjacent normal vectors. This is based on the perception characteristic that humans easily perceive changes in the boundaries and contours of objects, and that the greater the degree of unevenness, the more likely they are to be perceived as jitter. By adjusting the filter weights according to the attributes of the virtual content to be displayed, jitter can be reduced while minimizing the user's perception of jitter and reduced responsiveness in the displayed virtual content.
[0058] In addition to these, for example, the filter weight may be increased as the complexity of the texture of the object increases, or the filter weight may be increased as the gradient of brightness of the contour on the image of the virtual content increases.
[0059] Alternatively, the final filter weight α may be determined from a combination of multiple attributes. For example, the filter weight corresponding to attribute n is determined as α n In this case, the final α may be expressed as a weighted average as shown in the following equation (4).
[0060]
number
[0061] where w n represents the weight of attribute n. This allows you to set filter weights based on multiple attribute information.
[0062] Furthermore, in this embodiment, the creator of the virtual content may set in advance a jitter tolerance or a tolerance for a decrease in responsiveness for each piece of virtual content as attribute information of the virtual content. For example, the creator of the virtual content may set a jitter tolerance for each piece of virtual content and store it in advance in the storage unit 102. Based on the jitter tolerance, the determination unit 105 may increase the filter weight α as the jitter tolerance decreases and decrease the filter weight α as the jitter tolerance increases. The tolerance may also be adjusted in accordance with the user's perceptual ability. For example, the user may be shown a video to which pseudo-jitter has been added in advance to measure the user's perceptual ability, and the tolerance may be set to be lower as the perceptual ability increases and higher as the perceptual ability decreases.
[0063] Furthermore, the user wearing the head-mounted display may adjust the tolerance for jitter in virtual content while using the head-mounted display. Specifically, the tolerance for jitter in virtual content that the user perceives may be adjusted to be smaller.
[0064] 5 is a block diagram of an information processing device 1a according to this modification, including an operation unit (operation means) 110 for operating a user interface when the user sets the jitter tolerance. A display unit 107 displays a screen for the user to operate. The operation unit 110 is displayed on the display unit 107 and corresponds to the virtual content 4 in FIG. 2. The virtual content 4 is an operation panel that the user can interact with, allowing the user to arbitrarily set the jitter tolerance of the selected virtual content.
[0065] Fig. 6 is an image diagram of the operation panel when the user sets the jitter tolerance. In the example of Fig. 6, the jitter tolerance is set to 0.3 for virtual content 3, and 0.7 for virtual content 5. The tolerances set in this way are input to determination unit 105, which determines the filter weights based on the tolerances so that the smaller the tolerance, the larger the filter weight α, and the larger the tolerance, the smaller the filter weight α. [Example]
[0066] In the third embodiment, a method for determining (changing) the weight of a filter based on the gaze direction of a user wearing a head-mounted display will be described. The gaze direction of the user is the central field of the user's field of vision, and the area away from the gaze direction is the peripheral field of vision. In this embodiment, the closer the direction is to the gaze direction, the greater the filter weight, and the farther the direction is from the gaze direction, the smaller the filter weight. Human perception has the property that the central field is more likely to capture minute movements than the peripheral field, while the peripheral field is more likely to capture high-speed movements, and the filter in this embodiment is designed based on this property.
[0067] 7 is a block diagram of an information processing device 1b in this embodiment. The acquisition unit 103, derivation unit 104, application unit 106, and storage unit 102 are the same as those in the first embodiment, and therefore detailed description thereof will be omitted. This embodiment differs from the first embodiment in that a gaze imaging device 108 is provided outside the information processing device 1b, and a gaze acquisition unit (gaze acquisition means) 109 is provided inside the information processing device 1b.
[0068] The gaze imaging device 108 captures an image of the eyes of a user wearing a head-mounted display and outputs it to the gaze acquisition unit 109. The gaze acquisition unit 109 acquires the image captured by the gaze imaging device 108, calculates (tracks, estimates, or acquires) the gaze direction of the user using the acquired image, and outputs it to the determination unit 105. The determination unit 105 receives as input the position and orientation of the imaging device calculated by the derivation unit 104 and the gaze direction acquired by the gaze acquisition unit 109, and determines the weight of a filter to be applied to variables representing the six degrees of freedom of the position and orientation.
[0069] 8 is a flowchart showing the process (information processing method) of the information processing device 1b in this embodiment. Note that the processes other than step S2010 and step S1040 in FIG. 8 are the same as those in the first embodiment, and therefore detailed description thereof will be omitted.
[0070] In step S2010, the gaze acquisition unit 109 acquires an image captured by the gaze imaging device 108 and calculates the gaze direction from the acquired image. The gaze direction is calculated using Video-oculography (VOG), which is generally used in head-mounted displays.
[0071] In step S1040, the determination unit 105 calculates the angle θ between the line-of-sight vector calculated in step S2010 and the vector of the direction connecting the position (drawing position) of the virtual content and the position of the imaging device 101. The filter can be obtained using equation (1), as in the first embodiment. The larger the calculated angle, the smaller the filter weight α (filter strength), and the smaller the distance, the larger the weight α. Specifically, the weight α is expressed by the following equation (5) based on the calculated angle θ.
[0072]
number
[0073] α max is the upper limit of the weight α, α min is the lower limit of the weight α, which is set in advance. max , θ min are the upper and lower limit values of the angle θ between the line of sight vector and the vector connecting the position of the virtual content and the position of the image capture device 101, and these are set in advance. The value of the weight α is determined for each virtual content to be rendered. The upper and lower limit values of the weight α are set to prevent an extreme drop in responsiveness caused by making the filter weight too large.
[0074] A specific example will be described with reference to Fig. 2. It is assumed that the angle between the line of sight and the direction connecting the position of the virtual content 3 and the position of the image capture device 101 is 20°, and the angle between the line of sight and the direction connecting the position of the virtual content 4 and the position of the image capture device 101 is 75°. min =0.3, α max =0.7, θ min =0[°], θmax =50[°].
[0075] In this case, when the filter of formula (2) is applied to virtual content 3, θ=10, so α=1-(20-0) / (50-0)=0.6. When the filter of formula (2) is applied to virtual content 4, θ=75(>θ max ) and α=α min =0.3.
[0076] As a result, the filter weight α is determined to be α = 0.6 for virtual content 3 and α = 0.3 for virtual content 4, with a larger filter weight for virtual content 3, which is closer to the line of sight, than for virtual content 4. Therefore, virtual content 3, which is viewed in the central visual field and therefore more likely to perceive jitter and less likely to perceive a decrease in responsiveness, has a larger filter weight, which reduces responsiveness but has a stronger jitter suppression effect. On the other hand, virtual content 4, which is viewed in the peripheral visual field and therefore more likely to perceive a decrease in responsiveness and less likely to perceive jitter, has a smaller filter weight, which reduces the jitter suppression effect but can suppress a decrease in responsiveness.
[0077] In this embodiment, the filter weight is adjusted based on the line of sight of the user wearing the head-mounted display. As a result, the filter weight is increased for virtual content that is close to the line of sight, thereby reducing jitter. On the other hand, the filter weight is decreased for virtual content that is far from the line of sight, thereby reducing a decrease in responsiveness. In other words, the filter weight is adjusted according to the angle with the line of sight. As a result, it is possible to reduce jitter while minimizing the user's perception of a decrease in responsiveness to the displayed virtual content.
[0078] In this embodiment, the filter weight is adjusted based on the distance between the gaze direction vector and the virtual content, but this is not limiting. As a modification of this embodiment, for example, the filter weight may be adjusted based on the distance between the gaze point and the virtual content. When gaze imaging devices are attached to both eyes, the position where the gaze directions of both eyes of the user intersect becomes the position of the gaze point. Therefore, the position of the gaze point can be calculated geometrically. The shorter the calculated distance between the gaze point and the virtual content, the larger the filter weight α, and the larger the distance, the smaller the filter weight α.
[0079] In each embodiment, for example, the operation unit 110 may be configured to be able to set a mode for selecting the type of information of the virtual content used to determine the filter weight α based on a user operation. In this case, the determination unit 105 is configured to be able to change the type of information of the virtual content used to determine the filter weight according to the mode.
[0080] More specifically, the mode includes at least two of a first mode, a second mode, and a third mode. When the first mode is set, the determination unit 105 determines the filter weight α based on the distance between the image capture device 101 and the virtual content (first information). When the second mode is set, the determination unit 105 determines the filter weight α based on attribute information of the virtual content (second information). When the third mode is set, the determination unit 105 determines the filter weight α based on the angle between the user's line of sight and the direction connecting the rendering position of the virtual content and the position of the image capture device 101 (third information).
[0081] In each embodiment, changing the type of information of the virtual content according to the mode may include changing the contribution (sensitivity, influence, priority) of the first information, second information, and third information that contribute to determining the filter weight α. For example, when the first mode is set, the filter weight α may be determined by increasing the contribution of the first information and decreasing the contribution of the second information and the third information (first priority mode). Similarly, when the second mode is set, the contribution of the second information may be higher than the contribution of the first information and the third information (second priority mode). Furthermore, when the third mode is set, the contribution of the third information may be higher than the contribution of the first information and the second information (third priority mode).
[0082] (Other Examples) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0083] According to each embodiment, it is possible to provide an information processing device, an information processing method, and a program that can reduce jitter while suppressing the user's perception of a decrease in responsiveness to displayed virtual content.
[0084] The disclosure of each embodiment includes the following configurations and methods. (Configuration 1) an acquisition means for acquiring a captured image from an imaging device; a derivation means for deriving the position and orientation of the imaging device from the captured image; a storage means for storing information on virtual content to be rendered based on the position and the orientation; determining means for determining a weight of a filter based on the information of the virtual content; applying means for applying the filter to the position and the orientation based on the weight of the filter; An information processing device comprising: (Configuration 2) the information about the virtual content includes information about a drawing position of the virtual content, The determining means acquiring a distance between the imaging device and the virtual content based on the drawing position of the virtual content and the position of the imaging device; determining the weight of the filter so that the weight becomes larger as the distance becomes shorter; 2. The information processing device according to configuration 1, (Configuration 3) the information about the virtual content includes attribute information about the virtual content, the determining means determines the weight of the filter based on attribute information of the virtual content; 2. The information processing device according to configuration 1, (Configuration 4) the attribute information of the virtual content includes information about a moving speed of the virtual content, the determining means determines the weight of the filter so that the weight becomes smaller as the moving speed becomes smaller; 4. The information processing device according to configuration 3, (Configuration 5) the attribute information of the virtual content includes information on the volume of a three-dimensional model of the virtual content or information on the area of the virtual content projected onto the captured image; the determining means determines the weight of the filter so that the weight becomes larger as the volume or the area becomes smaller; 4. The information processing device according to configuration 3, (Configuration 6) the attribute information of the virtual content includes information on the degree of unevenness of the virtual content; the determining means determines the weight of the filter so that the weight increases as the degree of unevenness increases; 4. The information processing device according to configuration 3, (Configuration 7) the storage means stores information on jitter tolerance set in advance for each virtual content as the attribute information of the virtual content; the determining means determines the weight of the filter so that the weight becomes larger as the jitter tolerance becomes smaller; 4. The information processing device according to configuration 3, (Configuration 8) further comprising an operation means for setting the jitter tolerance based on a user operation; 8. The information processing device according to configuration 7, (Configuration 9) further comprising a display means for displaying a screen to be operated by a user; the operation means is displayed on the display means; 9. The information processing device according to configuration 8, (Configuration 10) The device further includes a gaze acquisition means for acquiring a gaze direction of a user, the information about the virtual content includes information about an angle between the line of sight and a direction connecting a drawing position of the virtual content and a position of the imaging device, the determining means determines the weight of the filter based on the angle such that the smaller the angle, the larger the weight of the filter; 2. The information processing device according to claim 1, (Configuration 11) the line-of-sight acquisition means estimates a gaze point at which the line-of-sight directions of both eyes of the user intersect; the determining means determines the weight of the filter based on a distance between the position of the fixation point and the position of the imaging device such that the weight of the filter increases as the distance decreases; 11. The information processing device according to configuration 10, (Configuration 12) further comprising an operation means for setting a mode based on a user's operation; the determining means changes the type of information of the virtual content used to determine the weight of the filter according to the mode; 2. The information processing device according to configuration 1, (Configuration 13) the modes include at least two of a first mode, a second mode, or a third mode; The determining means When set to the first mode, determining the weight of the filter based on a distance between the imaging device and the virtual content; When the second mode is set, the weight of the filter is determined based on attribute information of the virtual content; When the third mode is set, determining the weight of the filter based on an angle between a line of sight of the user and a direction connecting a drawing position of the virtual content and a position of the imaging device; 13. The information processing device according to configuration 12, (Configuration 14) The applying means applies the filter to the position and the attitude so as to reduce jitter that occurs when aligning the captured image with the virtual content. 14. The information processing device according to any one of configurations 1 to 13, (Method 1) an acquisition step of acquiring a captured image from an imaging device; a deriving step of deriving the position and orientation of the imaging device from the captured image; a storing step of storing information about virtual content to be rendered based on the position and the orientation; determining a filter weight based on the information of the virtual content; applying the filter to the position and the pose based on the weights of the filter; 10. An information processing method executed by an information processing device, comprising: (Configuration 15) A program that causes a computer to execute each step of the information processing method described in Method 1.
[0085] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments and various modifications and changes are possible within the scope of the present invention. At least a part of each embodiment may be combined. [Explanation of symbols]
[0086] 1, 1a, 1b Information processing device 3, 4, 5 Virtual Content 101 Imaging device 102 Holding part (holding means) 103 Acquisition unit (acquisition means) 104 Derivation part (derivation means) 105 Determination unit (determination means) 106 Application part (applicable means)
Claims
1. an acquisition means for acquiring a captured image from an imaging device; a derivation means for deriving the position and orientation of the imaging device from the captured image; a storage means for storing information on virtual content to be rendered based on the position and the orientation; determining means for determining a weight of a filter based on the information of the virtual content; applying means for applying the filter to the position and the orientation based on the weight of the filter; An information processing device comprising:
2. the information about the virtual content includes information about a drawing position of the virtual content, The determining means acquiring a distance between the imaging device and the virtual content based on the drawing position of the virtual content and the position of the imaging device; determining the weight of the filter so that the weight becomes larger as the distance becomes shorter; 2. The information processing device according to claim 1,
3. the information about the virtual content includes attribute information about the virtual content, the determining means determines the weight of the filter based on attribute information of the virtual content; 2. The information processing device according to claim 1,
4. the attribute information of the virtual content includes information about a moving speed of the virtual content, the determining means determines the weight of the filter so that the weight becomes smaller as the moving speed becomes smaller; 4. The information processing device according to claim 3,
5. the attribute information of the virtual content includes information on the volume of a three-dimensional model of the virtual content or information on the area of the virtual content projected onto the captured image; the determining means determines the weight of the filter so that the weight increases as the volume or the area decreases; 4. The information processing device according to claim 3,
6. the attribute information of the virtual content includes information on the degree of unevenness of the virtual content; the determining means determines the weight of the filter so that the weight increases as the degree of unevenness increases; 4. The information processing device according to claim 3,
7. the storage means stores information on jitter tolerance set in advance for each virtual content as the attribute information of the virtual content; the determining means determines the weight of the filter so that the weight becomes larger as the jitter tolerance becomes smaller; 4. The information processing device according to claim 3,
8. further comprising an operation means for setting the jitter tolerance based on a user operation; 8. The information processing device according to claim 7,
9. further comprising a display means for displaying a screen to be operated by a user; the operation means is displayed on the display means; 9. The information processing device according to claim 8,
10. The device further includes a gaze acquisition means for acquiring a gaze direction of a user, the information about the virtual content includes information about an angle between the line of sight and a direction connecting a drawing position of the virtual content and a position of the imaging device; the determining means determines the weight of the filter based on the angle such that the smaller the angle, the larger the weight of the filter; 2. The information processing device according to claim 1,
11. the line-of-sight acquisition means estimates a gaze point at which the line-of-sight directions of both eyes of the user intersect; the determining means determines the weight of the filter based on a distance between the position of the fixation point and the position of the imaging device such that the weight of the filter increases as the distance decreases; The information processing device according to claim 10 ,
12. further comprising an operation means for setting a mode based on a user's operation; the determining means changes the type of information of the virtual content used to determine the weight of the filter according to the mode; 2. The information processing device according to claim 1,
13. the modes include at least two of a first mode, a second mode, or a third mode; The determining means When set to the first mode, determining the weight of the filter based on a distance between the imaging device and the virtual content; When the second mode is set, the weight of the filter is determined based on attribute information of the virtual content; When the third mode is set, determining the weight of the filter based on an angle between a line of sight of the user and a direction connecting a drawing position of the virtual content and a position of the imaging device; The information processing device according to claim 12 .
14. The applying means applies the filter to the position and the attitude so as to reduce jitter that occurs when aligning the captured image with the virtual content. The information processing device according to any one of claims 1 to 13,
15. an acquisition step of acquiring a captured image from an imaging device; a deriving step of deriving the position and orientation of the imaging device from the captured image; a storing step of storing information about virtual content to be rendered based on the position and the orientation; determining a filter weight based on the information of the virtual content; applying the filter to the position and the pose based on the weights of the filter; 10. An information processing method executed by an information processing device, comprising:
16. A program causing a computer to execute each step of the information processing method according to claim 15.