Information processing device, method for controlling information processing device, system, and program

By detecting feature points using inertial sensors and camera units within the HMD, user position and orientation estimation is achieved without the need for additional motion sensors. This solves the problem of mismatch between HMD video display and actual position and orientation in vehicles, providing accurate position and orientation acquisition.

CN121925676APending Publication Date: 2026-04-24CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CANON KK
Filing Date
2024-09-20
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies have difficulty distinguishing and acquiring the user's position and orientation from the position and orientation inside the vehicle when the user uses a head-mounted display (HMD) while riding in a vehicle, resulting in a mismatch between the video display and the user's position and orientation. In addition, additional motion sensors are required.

Method used

By acquiring information from inertial sensors within the HMD and images captured by the camera unit, feature points are detected. The estimation unit then distinguishes feature point groups and estimates their position and orientation, enabling the acquisition of user position and orientation without the need for additional motion sensors.

Benefits of technology

When a user is riding a moving vehicle, the system can accurately distinguish and obtain the user's position and orientation, as well as the position and orientation inside the moving vehicle, thus solving the problem of inconsistent video display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121925676A_ABST
    Figure CN121925676A_ABST
Patent Text Reader

Abstract

An information processing apparatus includes: a determination unit configured to determine a feature point based on a detection result of an inertial sensor included in a display apparatus and a feature point detected from a captured image acquired from an imaging unit included in the display apparatus; determining whether the feature points correspond to a first feature point group or a second feature point group different from the first feature point group; and an estimation unit for estimating the position or orientation of the display device on the basis of the first feature point group, or estimating the position or orientation of the display device on the basis of the second feature point group.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an information processing device. Background Technology

[0002] In recent years, mixed reality (MR) and virtual reality (VR) technologies, which use head-mounted displays (HMDs) to allow users to experience spaces different from their real-world environment, have become widely known. These technologies consider various controls that users make while wearing the HMD. The HMD obtains its position and orientation from information from accelerometers and angular velocity sensors embedded within it, thereby enabling control.

[0003] When using an HMD as described above while a user is riding in a vehicle such as a car, the information obtained from the accelerometer and angular velocity sensor includes the user's motion information as well as the vehicle's motion information. This can lead to the following problem: the video image displayed on the video display device does not correspond to the amount and direction of the user's movement inside the vehicle. To solve this problem, PTL 1 describes a technique for determining the video image displayed on the HMD by subtracting the vehicle's motion information obtained from the motion sensor located on the vehicle from the user's motion information obtained from the motion sensor located on the user's body.

[0004] Reference List Patent documents PTL 1: Japanese Patent Application Publication No. 2019-049831 Summary of the Invention

[0005] Technical issues The prior art disclosed in the aforementioned patent documents requires manpower and time to prepare motion sensors separately for installation on vehicles, which are not needed when the user is using an HMD (Head-Down Device) without riding in the vehicle. Therefore, the present invention aims to provide an information processing device that, when the user is riding in a moving body, can distinguish and acquire the user's position and orientation, including the position and orientation of the moving body, as well as the user's position and orientation inside the moving body, without the need for separate motion sensors.

[0006] Problem Solution According to one aspect of the present invention, an information processing apparatus for communicating with a display device is provided, the information processing apparatus comprising: a first acquisition unit for acquiring detection results of an inertial sensor included in the display device; a second acquisition unit for acquiring captured images from a camera unit included in the display device; a detection unit for detecting feature points of a subject in the captured images acquired by the second acquisition unit; a determination unit for determining, based on the detection results of the inertial sensor acquired by the first acquisition unit and the feature points detected by the detection unit, whether the feature points correspond to a first set of feature points or a second set of feature points different from the first set of feature points; and an estimation unit for estimating the position or orientation of the display device based on the first set of feature points, or estimating the position or orientation of the display device based on the second set of feature points.

[0007] Advantages of the invention According to the present invention, when a user is riding in a mobile body, an information processing device is provided that can distinguish and acquire the user's position and orientation, including the position and orientation of the mobile body, as well as the user's position and orientation inside the mobile body, without the need for a separate motion sensor. Attached Figure Description

[0008] Figure 1 This is a diagram illustrating an information processing system according to a first embodiment.

[0009] Figure 2 This is a block diagram based on the first embodiment.

[0010] Figure 3 This is an illustration of a scenario where a user uses a head-mounted display (HMD) 100 while riding in a mobile vehicle.

[0011] Figure 4 This is a diagram illustrating the self-position estimation processing of camera video images.

[0012] Figure 5 This diagram illustrates the calculation and processing of the position and orientation movement data obtained from the orientation sensor unit 204.

[0013] Figure 6 This diagram illustrates the self-position estimation processing and movement calculation processing based on the position and orientation detected by the orientation sensor unit 204 when the moving body is moving and the user is also moving inside the moving body.

[0014] Figure 7 This is a flowchart illustrating the processing procedure of the information processing device. Detailed Implementation

[0015] The embodiments will now be described in detail with reference to the accompanying drawings. It should be noted that the following embodiments do not limit the invention according to the claims. Although multiple features are described in the embodiments, not all of these features are essential to the invention, and multiple features may be combined as appropriate. Furthermore, in the drawings, the same or similar components are denoted by the same reference numerals, and repeated descriptions will be omitted.

[0016] (First embodiment) Reference Figure 1 The information processing system 1 according to the first embodiment is described. The information processing system 1 includes a head-mounted display (HMD) 100, a personal computer (PC) 110, and a controller 120.

[0017] HMD 100 is a head-mounted display device (electronic device) that can be mounted on a user's head. HMD 100 displays composite images obtained by combining images of the area in front of the user captured by HMD 100 with content such as computer graphics (CG) in a form corresponding to the orientation of HMD 100.

[0018] The PC 110 controls the HMD 100. The PC 110 connects via a wired cable such as a Universal Serial Bus (USB) cable or via Bluetooth. ® ) and Wi-Fi ® The PC 110 is wirelessly connected to the HMD 100. The PC 110 combines the captured images with CG to generate a composite image and transmits the composite image to the HMD 100. It should be noted that the PC is described as an example of an information processing device; however, the information processing device is not limited to this. For example, the information processing device may be a smartphone or tablet terminal, and components of the PC 110 may be included in the HMD 100.

[0019] Controller 120 performs various controls on HMD 100. When PC 110 operates in a specific control mode and the user operates controller 120, HMD 100 is controlled based on the user's operation. For example... Figure 1 As shown, the controller 120 can be a wearable ring supported by the user's finger, or a handheld device held by the user. The controller 120 includes physical buttons for performing confirmation or selection operations on the display screen. The controller 120 communicates with the PC 110 via Bluetooth. ® Wireless communication. The controller is not limited to being configured to communicate with the PC 110; alternatively, it can be configured to communicate with the HMD 100.

[0020] The user can change the instruction position corresponding to the movement of the controller on the display through the motion controller 120. The instruction position is sometimes represented by a point, and sometimes can be represented by a virtual ray connecting the point of the instruction position to the controller using a straight line (line segment) or a dashed line. When the physical button is pressed, the determination operation or selection operation of the menu can be performed. As the shape of the controller 120, as described above, it is a ring shape and a hand-held shape; however, the shape of the controller 120 is not limited to this, as long as the controller 120 can be supported by a finger, a hand or an arm. In addition, the button is described as a physical button; however, other components such as a touchpad, a touch screen, a scroll wheel and a trackball can also be used, as long as the operation can be performed through the component, and in addition to pressing the button, another option for operation can be a sliding operation, a flicking operation or a touch operation.

[0021] The controller can also be worn on a finger, a hand and / or an arm.

[0022] The controller can also be attached to an object held by hand, and position information and orientation information about the attachment position can be obtained from the sensor. Examples of such objects include objects similar in shape to tools.

[0023] <Internal Structure of the HMD> will be described with reference to Figure 2 The internal structure of the HMD 100 will be described. The HMD 100 includes an HMD control unit 201, a camera unit 202, an image display unit 203, an orientation sensor unit 204, a non-volatile memory 205, a working memory 206, and a line-of-sight camera unit 207.

[0024] The HMD control unit 201 is a central processing unit (CPU) for controlling the components of the HMD 100. The HMD control unit 201 obtains a composite image (an image obtained by combining a captured image of the space in front of the user captured by the camera unit 202 with CG) from the PC 110, and then displays the composite image on the image display unit 203. Instead of the HMD control unit 201 controlling the entire device, multiple hardware components can share the processing to control the entire device.

[0025] The camera unit 202 includes two cameras (camera devices). The two cameras capture images for combining with images of virtual space to generate position and orientation information, and include a left-eye camera unit and a right-eye camera unit. The left-eye camera unit captures a moving image of real space corresponding to the left eye of the HMD 100 wearer, and outputs images of each frame in the moving image (captured images). The right-eye camera unit captures a moving image of real space corresponding to the right eye of the HMD 100 wearer, and outputs images of each frame in the moving image (captured images). In other words, the captured images acquired by the camera unit 202 are stereoscopic images with parallax substantially consistent with the positions of the left and right eyes of the HMD 100 wearer. Furthermore, distance information about the distance from the two cameras to the object can be obtained through distance measurement using the stereo cameras. In the HMD for a mixed reality (MR) system, the central optical axis of the camera unit's imaging range is preferably configured to substantially coincide with the line of sight of the HMD wearer.

[0026] Both the left-eye and right-eye camera units include an optical system and a camera device. Light entering from the outside passes through the optical system into the camera device, and the camera device outputs an image corresponding to the incident light as the captured image. Images of the object (the area in front of the user) captured by both cameras are output to PC 110 and control unit 201. Instead of capturing images, camera unit 202 can capture and output video images.

[0027] The image display unit 203 displays a composite image. The image display unit 203 includes a liquid crystal panel, an organic electroluminescent (EL) panel, etc. When the user wears the HMD 100, the image display unit 203 is positioned in front of the user's eyes. A device using a transflective mirror can be used as the image display unit 203. In this case, for example, the image display unit 203 can display images using a technology commonly known as augmented reality (AR), making the CG appear as if it were directly superimposed onto the real space visible through the transflective mirror. The image display unit 203 can also display images of a completely virtual space without using captured images by using a technology commonly known as virtual reality (VR).

[0028] Orientation sensor unit 204 is a motion sensor that acquires orientation (and position) information of HMD 100. Orientation sensor unit 204 can acquire orientation information of the user (the user wearing HMD 100) corresponding to the orientation (and position) of HMD 100. Orientation sensor unit 204 includes an inertial measurement unit (IMU) composed of inertial sensors such as accelerometers and angular accelerometers, and a geomagnetic sensor. Orientation sensor unit 204 is used to acquire information about the user's orientation (orientation information), and HMD control unit 201 outputs the detection result of the user's orientation information (orientation information) to PC 110. Orientation information can be acquired from one or more of magnetic sensors (including geomagnetic sensors), ultrasonic sensors, accelerometers, and angular velocity sensors.

[0029] The HMD control unit 201 estimates the position or orientation of the user's hand and finger joints based on images obtained from two cameras by the camera unit 202. Joints include feature points of parts such as finger joints, fingertips, back of the hand (palm), and forearm. Each joint indicates a coordinate position, and orientation can be estimated based on information from multiple joints. Methods for estimating the hand's position and orientation, as well as the position or orientation of the hand's joints, include, for example, known object recognition or pose estimation methods using machine learning via convolutional neural networks. Furthermore, the position information of each hand joint in the depth direction can be obtained by stereo matching using the images from the two cameras obtained by the camera unit 202, calculating the distance from the camera unit 202 to each joint using, for example, triangulation. The estimated coordinate information of each hand joint is output from the control unit 201 to the PC 110.

[0030] The non-volatile memory 205 is an electrically erasable and writable non-volatile memory, and stores programs, etc., to be executed by the control unit 211 described below.

[0031] The working memory 206 is used as a buffer memory for temporarily storing image data captured by the camera unit 202, an image display memory for the image display unit 203, and a working area for the control unit 201.

[0032] The gaze-capture camera unit 207 is a camera that acquires images for detecting the user's gaze, and it is attached to the interior of the HMD to capture images of the user's eyes when the user is wearing the HMD 100. The image of the object (the user's eyes) captured by the camera is output to the control unit 211 of the PC 110 via the HMD control unit 201. The control unit 211 detects the gaze of the user wearing the HMD 100 from the images captured by the gaze-capture camera unit 207 and designates the portion of the image displayed on the image display unit 203 that the user is looking at.

[0033] <Internal Structure of the Controller> Now, refer to Figure 2 to describe the internal structure of the controller 120. The controller 120 includes a controller control unit 221, an operation unit 222, a communication unit 223, and a controller orientation sensor unit 224.

[0034] The controller control unit 221 is a CPU configured to control the components of the controller 120. Instead of controlling the entire device through the controller control unit 221, multiple hardware components can share the processing to control the entire device.

[0035] The operation unit 222 includes buttons. The operation unit 222 detects whether a button is operated and transmits the detection information to the PC 110 through the communication unit 223. The operation unit 222 can also include various types of input formats.

[0036] The communication unit 223 performs Bluetooth (Bluetooth ® ) wireless communication with the PC 110. In the case where multiple controllers are provided, each controller performs Bluetooth (Bluetooth ® ) wireless communication with the PC 110.

[0037] The controller orientation sensor unit 224 includes an inertial measurement unit (IMU) composed of inertial sensors (such as an acceleration sensor and an angular acceleration sensor) and a geomagnetic sensor. The IMU detects changes in the position or orientation of the controller 120. The detected change information of the position and orientation is communicated from the communication unit 223 to the PC 110 through the controller control unit 221.

[0038] The output unit 225 includes a light-emitting diode (LED) light source, a speaker, a vibration element, etc.

[0039] <Internal Structure of the PC> Now, refer to Figure 2 to describe the internal structure of the PC 110. The PC 110 includes a control unit 211, a non-volatile memory 212, a working memory 213, a communication unit 214, and a recording medium 215.

[0040] Control unit 211 is a CPU configured to control each unit of PC 110 according to the input signals and program described below. Instead of control unit 211 controlling the entire device, multiple hardware components can share the processing to control the entire device. Control unit 211 receives images (captured images) acquired by camera unit 202 and orientation information acquired from HMD 100 by orientation sensor unit 204. Control unit 211 performs image processing on the captured images to compensate for aberrations in the optical systems of camera unit 202 and image display unit 203. Furthermore, control unit 211 combines the captured images with optional CG to generate a composite image. Control unit 211 transmits the composite image to HMD control unit 201 of HMD 100.

[0041] Furthermore, the control unit 211 determines the number of controllers included in the captured image. Additionally, the control unit 211 uses information acquired via the communication unit 214 to identify the attachment location of each controller. Then, based on the identification results, the control unit 211 controls each controller to modify its operation based on the input information.

[0042] The control unit 211 controls the position, orientation, and size of the CG in the composite image based on information (distance information and orientation information) acquired by the HMD 100. For example, in the space represented by the composite image, when a virtual object indicated by the CG is positioned near a specific object existing in real space, the control unit 211 increases the size of the virtual object (CG) as the distance between the specific object and the camera unit 202 decreases. Furthermore, for example, the control unit 211 draws the virtual object (CG) according to changes in the position and orientation of the HMD 100, thereby drawing CGs whose position and orientation change along with the HMD 100 and CGs whose position and orientation do not change with the HMD 100, and combines the CGs with the real image. As described above, by controlling the position, orientation, and size of the CG, the control unit 211 can generate a composite image that makes CG objects not positioned in real space appear as if they were positioned in real space.

[0043] In addition, control unit 211 receives information estimated by control unit 201 of HMD 100. The received information is temporarily stored in working memory 213.

[0044] The control unit 211 also receives information about changes in the position or orientation of the controller 120 from the communication unit 223 of the controller 120 via the communication unit 214. The control unit 211 overlays and displays the indicated position corresponding to the change in the position or orientation of the controller 120 onto the composite image. The control unit 211 may also overlay and display the indicated position corresponding to the change in both the position and orientation of the controller 120 onto the composite image.

[0045] Non-volatile memory 212 is an electrically erasable and writable non-volatile memory, and stores information such as programs and CGs to be executed by control unit 211 described below. Control unit 211 can switch the CGs read from non-volatile memory 212 (i.e., CGs used to generate composite images).

[0046] The working memory 213 is used as a buffer memory for temporarily storing image data captured by the camera unit 202 and time-series information about the estimated coordinate positions of each joint of the hand, the image display memory of the image display unit 203, the working area of ​​the control unit 211, etc.

[0047] The estimation of hand joint points can be performed by PC 110. In this case, after the camera unit 202 outputs the captured image to PC 110, the control unit 211 of PC 110 estimates the position or orientation of each hand joint point, processes the image using the information, and outputs the resulting image to HMD 100. Control unit 211 can estimate the position and orientation of each hand joint point, process the image using the information, and output the resulting image to HMD 100.

[0048] In addition to the components of the MR system, for example, the control unit of either the HMD 100 or the PC 110 performs various image processing operations on the captured images obtained by the camera unit 202 and the displayed images displayed on the image display unit 203. However, processing is not the primary objective of this invention, and therefore its description will be omitted.

[0049] (Example of the use of an information processing device) Now, refer to Figure 3 Describe a scenario where a user uses the HMD 100 while traveling on a moving vehicle such as a car or train. Figure 3 The illustration shows an example of a user wearing an HMD 100 on their head. The external scenery appears in window 301, set in a car or train, and changes according to the speed of the moving object as it moves. Figure 3 In this context, building 303 appears as external scenery. Changes in the scenery are presented as camera video images captured by camera unit 202. For example, the moving object might be a convertible without windows, and the scenery might be displayed directly without being seen through the windows. Furthermore, the camera video images captured by camera unit 202 include interior scenes 302 of the moving object, including walls and doors.

[0050] (Self-position estimation using camera video images) Simultaneous Localization and Mapping (SLAM) is a known method for estimating one's own position using camera video images. Publicly available SLAM methods detect the positions of feature points in camera video images and determine their own position in 3D space based on the amount of movement of each feature point between frames. Various publicly available SLAM algorithms are known for estimating one's own position in 3D space based on these feature points.

[0051] Now, refer to Figure 4 Describe the self-position estimation process using camera video images. Figure 4 The illustration shows an example of self-position estimation using camera video images, and is a diagram observed from a third-party perspective. Figure 4 The diagram shows from Figure 3 The state shown indicates that the arrow has moved along the direction indicated by movement amount 400. When viewed from inside the moving body, building 303 appears to have moved along the direction opposite to that indicated by movement amount 400 by the arrow length of movement amount 404. Figure 4 The illustration shows a scenario where the user is not moving inside the moving object.

[0052] When the user is inside the moving body, the camera video image shows a window 301 set on the moving body and an interior scene 302 inside the moving body, including walls and doors.

[0053] When a mobile object moves and the user also moves inside the mobile object, the amount of movement associated with the movement of the mobile object and the user's own movement is detected based on external feature points detected in window 301 set on the mobile object. In this example, since the user has not moved, the building 303 shown by the dashed line has moved relatively, and this is observed from the HMD at the location of building 403. For example, the following situation is detected: Figure 3 The frame shown is Figure 4 Between the frames shown, Figure 4 Feature point 405 of building 403 obtained at the time point shown is from Figure 3 The feature point 305 detected at the time point shown has moved by a movement amount of 404. By using the external feature points detected in the window 301 set on the moving body to estimate its own position, the HMD 100's own position outside the moving body can be estimated.

[0054] For feature points within the interior scene 302, including walls and doors, motion amounts unrelated to the movement of the moving body are detected. When a user moves within the moving body, the user's movement within the moving body is detected. For example, since the user has not moved, feature point 406 of the interior scene 302 is detected as a stationary feature point from the HMD. When using the feature points of the interior scene 302 to estimate its own position, the HMD 100's own position within the moving body can be estimated.

[0055] (Based on the amount of movement of position and orientation detected by orientation sensor unit 204) Now, refer to Figure 5 The description describes the motion calculation process performed using the position and orientation acquired by the orientation sensor unit 204. Figure 5 An example of the amount of movement detected by the orientation sensor unit 204 is illustrated. The orientation sensor unit 204 acquires the acceleration and direction of movement of the HMD 100. The orientation sensor unit 204 calculates the amount of movement 501 of the HMD 100 by accumulating periodically acquired information. The amount of movement 501 of the HMD 100 is a three-dimensional amount of movement. The amount of movement can be represented by three components based on reference coordinates determined for the HMD 100. In addition to the amount of movement, the orientation can also be represented by rotational components, which are divided into three three-dimensional vectors: roll, pitch, and yaw.

[0056] (A method that provides two types of self-position estimation results) Reference Figure 3 , Figure 6 and Figure 7 Describe the processing that provides two types of self-position estimation results.

[0057] Figure 6 This diagram illustrates the self-position estimation processing and movement calculation processing based on the position and orientation detected by the orientation sensor unit 204 when the moving body is moving and the user is also moving inside the moving body. Figure 6 The illustration shows the user at Figure 4 The scene shown depicts movement. Figure 6 The illustration shows a user moving inside the moving body to position HMD 101 in the direction indicated by the movement amount 600, the distance traveled being the length of the arrow. In other words, Figure 6 The illustration shows HMD 100 moving inside the moving body to the position of HMD 101 in the direction indicated by the movement amount 600, and the movement distance is the length of the arrow.

[0058] Figure 7 This is a flowchart illustrating the processing flow that provides two types of self-position estimation results.

[0059] In step S701, the control unit 211 controls the camera unit 202 and stores the video images obtained by capturing images of the user's surroundings in the working memory 213. Recording continues, and the camera video images are repeatedly stored in the working memory 213.

[0060] In step S702, the control unit 211 detects feature points of the camera video images stored in the working memory 213, and stores the position information of the detected feature points for each frame in the working memory 213. Additionally, the control unit 211 compares the position of each feature point across frames to calculate the amount of movement of each feature point between frames. Figure 6 As shown, the external feature point 405 detected in the window 301 set on the moving body has moved by a movement amount 602 due to the movement of the moving body and the movement of the user inside the moving body. The feature point 406 of the internal scene 302 has also moved by a movement amount 603 due to the movement of the user inside the moving body.

[0061] In step S703, the control unit 211 acquires the movement amount 601 of the HMD 100 from the orientation sensor unit 204. For the movement amount 601 of the HMD 100, the movement amount obtained by adding the movement of the moving body to the movement of the user inside the moving body is detected. The movement amount 601 detected solely from the orientation sensor unit 204 is insufficient to detect both the movement of the moving body and the movement of the user inside the moving body; therefore, the movement amount of the user when viewed from outside the moving body is also detected.

[0062] In step S704, the control unit 211 converts the movement amount 501 of the HMD 100 into the movement amount on the camera video image. For example, consider the following method: from the position information obtained by adding the movement amount 501 of the HMD 100 to its own position as three-dimensional position information, select a feature point on the camera video image and calculate the movement amount of the selected feature point.

[0063] like Figure 4 As shown, this assumes the mobile body moves forward while the user-worn HMD remains completely stationary inside the mobile body, and that the amount of movement detected by the sensor unit 204 is positive. Figure 4In this context, feature points 305 and 405 are feature points detected in region 301 corresponding to the area outside the moving body, while feature point 406 is a feature point detected inside the moving body. In this case, it is assumed that the movement amount of the feature point detected in region 301 corresponding to the area outside the moving body, among the feature points in the camera video image, is a negative value with an equal absolute value detected towards sensor unit 204. As described above, the movement direction of the feature points in the camera video image is opposite to the movement direction of HMD 100 detected towards sensor unit 204. Therefore, the sign of the detected feature point movement amount is inverted, thereby converting the movement amount 501 of HMD 100 into the movement amount in the camera video image. Thereafter, the movement amount of HMD 100 in the camera video image, converted from the movement amount 501 of HMD 100, is subtracted from the movement amount of each feature point between frames. Feature point 305 moves to the position of feature point 405, and the length of arrow 404 is moved in the camera video image. The movement amount of arrow 400 is subtracted from the movement amount of arrow 404. When errors are not considered, the absolute value of the movement after subtraction becomes zero in the graph. As described above, feature points whose absolute value of the movement after subtraction is less than a predetermined threshold are identified as first feature points in the vehicle's external region and stored in the working memory 213. At this time, when stationary objects in the vehicle's external region are used as feature points, it is assumed that the result of subtraction is that the movement becomes zero; however, since errors may occur, the predetermined threshold is set to a value that takes into account the error.

[0064] Furthermore, when the user is not moving inside the vehicle, the movement of feature point 406 in the camera video image becomes zero. The movement of arrow 400 is subtracted from the movement of the feature point with zero movement. When error is disregarded, the absolute value of the movement after subtraction is equal to the absolute value of the movement of arrow 400 in the image. As described above, feature points whose absolute value of the subtracted movement is greater than or equal to a predetermined threshold are identified as second feature points in the vehicle interior region and stored in the working memory 213. The predetermined threshold used to determine the first feature point and the predetermined threshold used to determine the second feature point can be set to the same value or different values.

[0065] Suppose that in the vehicle's external region, objects such as buildings that are actually stationary but appear to be moving relative to each other, as well as moving objects such as cars, are detected. In this embodiment, when error is disregarded, among feature points whose absolute values ​​of movement after subtraction are all greater than or equal to a threshold, the absolute value of the movement of each feature point in the vehicle's internal region after subtraction is equal to the movement of the moving object. Therefore, feature points in the vehicle's internal region have values ​​similar to the absolute values ​​of movement after subtraction. As a result of subtraction, if a large number of feature points with similar values ​​are detected, these feature points are identified as a second feature point group and stored in the working memory 213. Whether feature points have similar values ​​can be determined based on whether their similar values ​​are close to each other within a certain error range. Among the feature points identified as the second feature point group, feature points surrounded by feature points detected as first feature points in the vehicle's external region can be determined not to be feature points in the vehicle's internal region and are not identified as part of the second feature point group.

[0066] It should be noted that the sign of the amount of movement toward sensor unit 204 can be inverted, thereby converting the amount of movement on the camera video image (i.e., the amount of movement of feature points detected from the captured image). Taking into account the opposite direction, instead of subtraction, addition can be performed with the signs of the amount of movement toward sensor unit 204 and the amount of movement on the camera video image being opposite to each other.

[0067] In step S705, the control unit 211 determines whether to use the first feature point or the second feature point for self-position estimation. If the first feature point (i.e., the feature point in the external region of the vehicle) is used for self-position estimation, the process proceeds to step S606. If the second feature point (i.e., the feature point in the internal region of the vehicle) is used for self-position estimation, the process proceeds to step S607. The control unit 211 determines the feature point used for self-position estimation based on the fact that the user has selected whether to use the first or second feature point.

[0068] In step S706, the control unit 211 uses the first feature point (i.e., the feature point in the external region of the vehicle) to estimate its own position. After this, the process ends.

[0069] After performing self-position estimation in step S706, the control unit 211 can draw virtual objects based on the estimation results of its own position.

[0070] After estimating its own position in step S706, the estimation result can be transmitted to the application selected by the user.

[0071] In step S707, the control unit 211 estimates its own position based on the second feature point (i.e., the feature point in the vehicle interior region). The process then ends.

[0072] After performing self-position estimation in step S707, the control unit 211 can draw virtual objects based on the estimation results of its own position.

[0073] After estimating its own position in step S707, the estimation result can be transmitted to the application selected by the user.

[0074] While the invention has been described in detail based on preferred embodiments, it is not limited to the specific embodiments, and various forms are included without departing from the spirit of the invention. The above embodiments can be appropriately combined in parts.

[0075] (Other embodiments) The present invention is also achieved by the following process: The process involves providing software (programs) that implement the functions of the above embodiments to a system or device via a network or various storage media, and causing the computer (or control unit, microprocessor unit (MPU), etc.) of the system or device to read and execute the program code. In this case, the program and the storage medium storing the program constitute the present invention.

[0076] As described above, the present invention has been described in detail based on preferred embodiments; however, the present invention is not limited to the specific embodiments, and various forms are included in the present invention without departing from the spirit of the invention. The above embodiments can be appropriately combined in parts.

[0077] Each functional unit in each embodiment (each variant) may or may not be independent hardware. The functions of two or more functional units may be implemented using shared hardware. Each of the multiple functions of a single functional unit may be implemented using independent hardware. Alternatively, each functional unit may be implemented using hardware such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), and a digital signal processor (DSP), or may not be implemented using such hardware. For example, the device may include a processor and a memory (storage medium) storing a control program. Furthermore, the functions of at least some of the functional units included in the device may be implemented by the processor reading the control program from the memory and executing the control program.

[0078] This invention can be implemented by providing a program that implements one or more functions of the above embodiments to a system or device via a network or storage medium, and by causing one or more processors in the computer of the system or device to read and execute the program. Furthermore, this invention can be implemented by a circuit (e.g., an ASIC) that implements one or more functions.

[0079] The disclosure of this embodiment includes the following structures, methods, procedures, and media.

[0080]

Construction 1

[0081]

Construction 2

[0082]

Construction 3

[0083]

Construction 4

[0084]

Construction 5

[0085]

Construction 6

[0086]

Construction 7

[0087]

Construction 8

[0088]

Construction 9

[0089]

Construction 10

[0090]

Construction 11

[0091]

Construction 12

[0092]

Control Methods

[0093]

program

[0094]

system

[0095] This invention is not limited to the embodiments described above, and various changes and modifications can be made without departing from the spirit and scope of the invention. Therefore, in order to make the scope of the invention clear to the public, the claims are hereby appended.

[0096] This application is based on and claims priority to Japanese Patent Application No. 2023-171953, filed on October 3, 2023, the entire contents of which are incorporated herein by reference.

Claims

1. An information processing apparatus for communicating with a display device, the information processing apparatus comprising: The first acquisition unit is used to acquire the detection results of the inertial sensor included in the display device; The second acquisition unit is used to acquire captured images from the camera unit included in the display device; A detection unit is configured to detect feature points of a subject in the captured image obtained by the second acquisition unit. The determining unit is configured to determine, based on the detection result of the inertial sensor acquired by the first acquiring unit and the feature point detected by the detection unit, whether the feature point corresponds to a first feature point group or a second feature point group different from the first feature point group. as well as An estimation unit is configured to estimate the position or orientation of the display device based on the first set of feature points, or to estimate the position or orientation of the display device based on the second set of feature points.

2. The information processing apparatus according to claim 1, wherein, The information processing device is located in the display device.

3. The information processing apparatus according to claim 1 or 2, further comprising a third acquisition unit, the third acquisition unit being used to acquire the movement amount of the feature point. in, The determining unit determines whether the feature point corresponds to the first feature point group or the second feature point group based on the detection result of the inertial sensor obtained by the second acquiring unit and the movement amount of the feature point obtained by the third acquiring unit.

4. The information processing apparatus according to claim 3, wherein, The determining unit determines whether the feature point corresponds to the first feature point group or the second feature point group based on a first value obtained by subtracting the movement amount based on the detection result of the inertial sensor from the movement amount of the feature point.

5. The information processing apparatus according to claim 4, wherein, If the first value is less than a predetermined threshold, the determining unit determines that the feature point corresponds to the first feature point group.

6. The information processing apparatus according to claim 4, wherein, If the first value is greater than or equal to a predetermined threshold and the first values ​​of the plurality of feature points are within a predetermined range, the determining unit determines that the plurality of feature points correspond to the second feature point group.

7. The information processing apparatus according to claim 6, wherein, If the first value is greater than or equal to the predetermined threshold but the feature point is surrounded by feature points corresponding to the first feature point group, the determining unit determines that the feature point does not correspond to the second feature point group.

8. The information processing apparatus according to claim 1, wherein, When the estimation unit estimates the position or orientation of the display device based on the first feature point group, the estimation unit estimates the position or orientation of the display device based on the detection result of the inertial sensor obtained by the first acquisition unit.

9. The information processing apparatus according to claim 1, wherein, The estimation unit estimates the position or orientation of the display device based on the first set of feature points, or based on the second set of feature points, depending on the application selected by the user.

10. The information processing apparatus according to claim 9, further comprising a communication unit for transmitting the estimation result of the estimation unit to the application.

11. The information processing apparatus according to claim 10, in, The estimation unit estimates the position or orientation of the display device based on the first set of feature points, and estimates the position or orientation of the display device based on the second set of feature points. The communication unit transmits the estimation results of the estimation unit based on the first feature point group and the estimation results of the estimation unit based on the second feature point group to the application.

12. The information processing apparatus according to claim 1, further comprising a drawing unit, the drawing unit being used to draw a virtual object based on the estimation result of the estimation unit.

13. A control method for an information processing apparatus communicating with a display device, the method comprising: As a first acquisition step, the detection results of the inertial sensors included in the display device are acquired; As a second acquisition step, a captured image is acquired from the camera unit included in the display device; As a detection step, feature points of the subject are detected from the captured image obtained in the second acquisition step; As a determining step, based on the detection result of the inertial sensor acquired in the first acquisition step and the feature point detected in the detection step, it is determined whether the feature point corresponds to a first feature point group or a second feature point group that is different from the first feature point group. as well as As an estimation step, the position or orientation of the display device is estimated based on the first feature group, or the position or orientation of the display device is estimated based on the second feature point group.

14. A program that causes a computer to act as each unit of the information processing apparatus according to claim 1.

15. An information processing system, the information processing system comprising: Display device; A first acquisition device is used to acquire the detection results of the inertial sensor included in the display device; The second acquisition device is used to acquire captured images from the camera unit included in the display device; A detection device, the detection device being used to detect feature points of a subject in the captured image acquired by the second acquisition device; A determining device is used to determine, based on the detection result of the inertial sensor acquired by the first acquiring device and the feature point detected by the detection device, whether the feature point corresponds to a first feature point group or a second feature point group different from the first feature point group. as well as An estimation device is used to estimate the position or orientation of the display device based on the first set of feature points, or based on the second set of feature points.

Citation Information

Patent Citations

  • Video display control device, video display system and video display control method

    JP2019049831A

  • Compositions and methods for reducing ocular neovascularization

    JP2023171953A