Information processing apparatus, method of controlling the same, and storage medium
The information processing device manages display content and environment maps to prevent exposure of sensitive information and maintain optical consistency in HMD systems, addressing issues of information leakage and obstructive objects in mixed reality imaging.
Patent Information
- Application Number
- JP2024109919
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-08
- Publication Date
- 2026-01-21
AI Technical Summary
In immersive imaging technologies using HMDs, there are issues where confidential information is exposed when the HMD is disconnected, and obstructive objects in the real world affect the optical consistency of mixed reality images.
An information processing device controls a non-head-mounted first display device and a head-mounted second display device, acquiring stop information to manage the display content on the second device when the HMD is stopped, ensuring that sensitive information is not displayed on the first device and adjusting the environment map to maintain optical consistency.
Prevents unintended exposure of sensitive information and improves optical consistency in mixed reality images by managing display content and environment maps effectively.
Smart Images

Figure 2026009780000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to imaging technology using an HMD. [Background technology]
[0002] In recent years, technologies that allow users to enjoy immersive images using head-mounted displays (HMDs) have become widespread. For example, in mixed reality (MR), a type of cross-reality (XR) technology that combines virtual and real worlds, HMD users can perform office tasks such as document creation while viewing a virtual display superimposed on a real-world image (see FIG. 1(a)). Alternatively, before purchasing furniture, users can get a sense of how the furniture will blend in with the surrounding environment by viewing a mixed reality image in which a computer graphic representation of the furniture is superimposed on an image of the surrounding environment (see FIG. 1(b)). Regarding the former use case, for example, Patent Document 1 (Patent Document 1) discloses a technology that extends the screen of a desktop LCD monitor with an HMD and displays confidential information on the HMD screen, thereby preventing others from viewing the confidential information. Regarding the latter use case, Patent Document 2 discloses a technology that maintains an environment mapping for a virtual object expressed in CG and reproduces the reflection of virtual lighting by drawing virtual lighting on the environment map according to the position and orientation of the virtual lighting. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-233963 [Patent Document 2] Japanese Patent Application Laid-Open No. 2007-18173 Summary of the Invention [Problem to be solved by the invention]
[0004] In the former use case, when the HMD is disconnected, the screen viewed on the HMD is aggregated onto an LCD monitor in the real world, making it viewable by others, posing a risk of confidential information being seen by others. In the latter use case, when an obstructive object in the real world is removed from the image by concealment processing, the obstructive object (concealed object) in the environment map affects the CG, reducing the optical consistency of the mixed reality image. [Means for solving the problem]
[0005] The information processing device according to the present disclosure is an information processing device that controls a non-head-mounted first display device and a head-mounted second display device, and is characterized by having an acquisition means that acquires stop information for displaying a virtual screen on the second display device, and a processing means that, when the acquisition means acquires the stop information, determines the display content of the real screen on the first display device when displaying the virtual screen on the second display device stops. [Effects of the Invention]
[0006] According to the present disclosure, in XR technology using an HMD, it is possible to prevent situations from occurring that are not in line with the wearer's intentions. [Brief explanation of the drawings]
[0007] [Figure 1] 1A and 1B are diagrams illustrating the background art. [Figure 2] FIG. 1A is a diagram showing an example of the configuration of a display system, and FIG. 1B is a diagram showing an example of the hardware configuration of an HMD. [Figure 3] 1A is a diagram showing an example of a hardware configuration of an information processing device, and FIG. 1B is a diagram showing an example of a software configuration of the information processing device according to the first embodiment. [Figure 4] 10A and 10B are diagrams showing an example of processing of display content. [Figure 5] 4 is a flowchart showing an operation flow in the information processing device according to the first embodiment. [Figure 6]6 is a flowchart showing details of a display determination process according to the first embodiment. [Figure 7] FIG. 10 is a diagram showing an example of window information in a table format. [Figure 8] 10 is a flowchart showing details of a display determination process according to a first modification of the first embodiment. [Figure 9] FIG. 10 is a diagram showing an example of window information in a table format. [Figure 10] 10 is a flowchart showing details of a display determination process according to a second modification of the first embodiment. [Figure 11] FIG. 10 is a diagram showing an example of form invisibility processing. [Figure 12] 10A and 10B are diagrams illustrating specific examples of situations in which optical consistency in a mixed reality image is reduced. [Figure 13] 10A is a diagram showing an example of the software configuration of an information processing device according to a second embodiment, and FIG. 10B is a diagram showing an example of the software configuration of an information processing device according to a first modified example of the second embodiment. [Figure 14] (a) is a diagram illustrating a real image, and (b) and (c) are diagrams illustrating an environment map. [Figure 15] 10 is a flowchart showing an operation flow in an information processing device according to a second embodiment. [Figure 16] FIG. 10A is a diagram showing an example of an object detection result, and FIG. 10B is a diagram showing an example of a detection result table. [Figure 17] 10 is a flowchart showing details of removal region determination processing according to the second embodiment. [Figure 18] FIG. 1A is a diagram showing an example of an object detection feature region image, FIG. 1B is a diagram showing an example of an image feature region image, and FIG. 1C is a diagram showing an example of a three-dimensional feature region image. [Figure 19] FIG. 10 is a diagram for explaining an example of a situation in Modification 1 of Embodiment 2. [Figure 20] 10 is a flowchart showing the operation flow of an information processing device according to a first modification of the second embodiment. [Figure 21] FIG. 10 is a diagram for explaining an example of a situation in Modification 2 of Embodiment 2. [Figure 22]10A is a diagram showing an example of the software configuration of an information processing device according to a second modification of the second embodiment, and FIG. 10B is a diagram showing an example of the software configuration of an information processing device according to a third modification of the second embodiment. [Figure 23] 10 is a flowchart showing the operation flow of an information processing device according to a second modification of the second embodiment. [Figure 24] 10 is a flowchart showing details of a removal region determination process according to a second modification of the second embodiment. [Figure 25] 10 is a flowchart showing the operation flow of an information processing device according to a third modification of the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0008] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the following embodiments do not limit the present invention, and not all of the combinations of features described in the embodiments are necessarily essential to the solution of the present invention. Note that the same components will be described with the same reference numerals.
[0009] [Embodiment 1] In this embodiment, in a use case where the screen of a display device placed on a desk is extended by an HMD (see Figure 1(a) above), when the HMD display is stopped, the content displayed on the HMD is not displayed on the screen of the display device on the desk.
[0010] <System configuration> Fig. 2(a) is a diagram showing a configuration example of the display system. The display system shown in Fig. 2(a) is composed of an information processing apparatus 1, a first display apparatus 2, and a second display apparatus 3. The first display apparatus 2 is a display apparatus installed on a desk or a table in a state viewable by others, and examples thereof include a liquid crystal monitor, a portable terminal, a projector, a 3D-compatible monitor, etc. Hereinafter, for convenience of explanation, the first display apparatus 2 will be referred to as a "monitor". The second display apparatus 3 is a video see-through type HMD (head-mounted display) worn on a person's head. In the present embodiment, the HMD 3 captures the monitor 2 through an internal RGB camera and displays a composite reality image with an extended display area by superimposing a virtual screen (virtual display) on the captured video. Note that the HMD 3 may be an optical see-through type. Furthermore, instead of the see-through type, a VR HMD that displays a virtual reality image represented only by CG may be used. In the present embodiment, the information processing apparatus 1 will be described as an independent system configuration from the HMD 3, but a configuration such as an integrated HMD system including the information processing apparatus 1 inside the HMD 3 may be adopted. Also, a configuration such as an integrated PC including the information processing apparatus 1 inside the monitor 2 may be adopted.
[0011] <Hardware Configuration of HMD> FIG. 2(b) illustrates an example of the hardware configuration of the HMD 3. The HMD 3 has multiple RGB cameras 201 and an inertial measurement unit (IMU) (not shown) to achieve inside-out position tracking. The IMU detects three-dimensional inertial motion (translational and rotational motion in three orthogonal axes) and is composed of a gyro sensor for detecting rotational motion and an acceleration sensor for detecting translational motion. The HMD 3 also has a distance sensor 202, such as a LiDAR (Light Detection and Ranging) sensor, for acquiring depth information. The HMD 3 also has a left-eye display 203 and a right-eye display 205, each of which is configured using a liquid crystal panel or an organic electroluminescence (EL) panel, for displaying images for the left and right eyes. Furthermore, left-eye eyepieces 204 and right-eye eyepieces 206 are disposed in front of the displays 203 and 205, respectively. The wearer observes enlarged virtual images of the images displayed for the left and right eyes through the lenses 204 and 206. The HMD 3 generates an image for the left eye and an image for the right eye based on a mixed reality image provided by the information processing device 1, and displays the image for the left eye on the left eye display 203 and the image for the right eye on the right eye display 206. At this time, by providing an appropriate parallax between the image for the left eye and the image for the right eye, it is possible to give the wearer an image perception with a sense of depth. Note that the HMD 3 has other components in addition to those described above, but since they are not the main focus of the present invention, their description will be omitted.
[0012] <Hardware configuration of information processing device> An example of the hardware configuration of the information processing device 1 will be described with reference to FIG. 3(a). In FIG. 3(a), a CPU 101 uses a RAM 102 as a work memory, executes programs stored in a ROM 103 and a hard disk drive (HDD) 105, and controls the operation of each block (described later) via a system bus 110. An HDD interface (hereinafter, interface will be referred to as "I / F") 104 connects a secondary storage device such as the HDD 105 or an optical disk drive. The HDD I / F 104 is, for example, an I / F such as a serial ATA (SATA). The CPU 101 can read data from and write data to the HDD 105 via the HDD I / F 104. Furthermore, the CPU 101 can load data stored in the HDD 105 into the RAM 102, and conversely, can save data loaded in the RAM 102 to the HDD 105. The CPU 101 can then execute the data loaded in the RAM 102 as a program. The input I / F 106 connects an input device 107 such as a keyboard, mouse, digital camera, scanner, or acceleration sensor. The input I / F 106 is, for example, a serial bus I / F such as USB or IEEE1394. The CPU 101 can read data from the input device 107 via the input I / F 106. The output I / F 108 connects the monitor 2 and HMD 3. The output I / F 108 is, for example, a video output I / F such as DVI or HDMI (registered trademark). The CPU 101 can send data to the monitor 2 and HMD 3 via the output I / F 108 and display a predetermined video. The information processing device 1 includes other components than those described above, but these are not the focus of the present invention and will not be described here.
[0013] Note that both the information processing device 1 and the HMD 3 may have a system configuration having the hardware configuration shown in Fig. 3(a). In this case, the HMD 3 and the information processing device 1 may exchange information about application windows (hereinafter simply referred to as "windows") through their respective input I / Fs 106 and output I / Fs 108. This allows, for example, the use of both the input function of the HMD 3 and the input of the information processing device 1. If the HMD 3 has an input function based on hand gesture recognition, both hand gesture input by the HMD 3 and input by mouse and keyboard operation by the information processing device 1 are possible.
[0014] <Software configuration of information processing device 1> 3(b) is a block diagram showing an example of the software configuration (logical configuration) of the information processing device 1 according to this embodiment. In FIG. 3(b), the information processing device 1 is composed of an input receiving unit 10 and an image processing unit 11, and the image processing unit 11 has a display determination unit 12 and a processing unit 13.
[0015] The input receiving unit 10 receives various user inputs, including data on images (real images) obtained by capturing a real space with the HMD 3. In this embodiment, one of the user inputs is information to stop display on the HMD 3. The stop information includes, for example, an instruction to stop display on the HMD 3 by the user via a mouse, keyboard, or hand gesture, as well as a detection signal indicating removal of the HMD 3 main body from the head or disconnection from the HMD 3 (for example, cable disconnection). The stop information received by the input receiving unit 10 is output to the display determination unit 12.
[0016] The display determination unit 12 determines what display should be performed on the monitor 2 when the display of the HMD 3 is stopped. The determination result is output to the processing unit 13.
[0017] The processing unit 13 processes the display content of the monitor 2 based on the determination result of the display determination unit 12. (a) and (b) of FIG. 4 are diagrams illustrating an example of the processing process. (a) of FIG. 4 illustrates a case where a window displayed on the HMD 3 is minimized and displayed on the screen of the monitor 2. The upper part of FIG. 4(a) illustrates a state before the display of the HMD 3 is stopped, in which a window 405 is displayed on the real screen 400 of the monitor 2 and a window 404 is displayed on the virtual screen 402. Furthermore, the real icon portion 401 in the real screen 400 displays an icon of an application corresponding to the window 405 belonging to the monitor 2. Similarly, the virtual icon portion 403 in the virtual screen 402 displays an icon of an application corresponding to the window 404 belonging to the HMD 3. The user can hide or redisplay the window corresponding to the icon by clicking on the icon. The lower part of FIG. 4(a) shows the state after the display of the HMD3 has stopped, in which the virtual screen 402 has disappeared and a window 405 has been displayed on the real screen 400 of monitor 2. An icon corresponding to the window 404 that was displayed on the virtual screen 402 has also been added to the real icon section 401. FIG. 4(b) shows an example in which the area of the window that was displayed on monitor 2 and that corresponds to the window that was displayed on HMD3 has been filled in. In the upper part of FIG. 4(b), the display area of window 411 has been expanded by HMD3. That is, the portion of window 411 that overlaps monitor 2 is displayed on monitor 2, and the portion that extends beyond monitor 2 to the left is superimposed on the real image by HMD3 and displayed on the virtual screen 402. The lower part of FIG. 4(b) shows the state after the display of the HMD3 has stopped, in which the virtual screen 402 has disappeared and the area of window 412 on the real screen 400 of monitor 2 that corresponds to the area that was displayed on the virtual screen 402 has been filled in. When performing such a filling process, it is sufficient to refer to the window information and perform a process of filling the area with, for example, gray based on the coordinates and size information of the area displayed on HMD 3. Furthermore, although not shown here, processing may be performed such that the window displayed on HMD 3 is placed behind window 405 displayed on monitor 2 (reversing the front-to-back relationship).
[0018] <Operation flow of information processing device> 5 is a flowchart showing the operation flow of the information processing device 1 according to this embodiment. In the following description, the symbol "S" means step.
[0019] In S501, the input receiving unit 10 acquires information about stopping the display on the HMD 3. As described above, this stop information is obtained, for example, by an operation signal when the user presses a display stop button (not shown) provided on the HMD 3, or by the user inputting a display stop instruction to the information processing device 1 through the input device 107. Note that the display may be stopped upon detecting that the output I / F 108 to the HMD 3 has been disconnected. Alternatively, the display may be stopped upon detecting that the power supply to the HMD 3 has been cut off.
[0020] In S502, the display determination unit 12 determines what kind of display to perform on the monitor 2 in response to the display stop of the HMD 3. Fig. 6 is a flowchart showing the details of the display determination process in this step. Here, a description will be given with reference to the flowchart in Fig. 6.
[0021] In S601, window information (hereinafter referred to as "window information") corresponding to all running applications is obtained. FIG. 7 shows an example of window information in table format, with each row corresponding to one window. The "ID" column contains a unique ID that uniquely identifies each window. The "Application Name" column contains the application name corresponding to the window. The "Coordinate" column contains the x and y coordinate values of the upper left corner of the window. Here, the origin of the coordinates is assumed to be the coordinate of the upper left corner of monitor 2. The "Size" column contains values representing the width and height (width, height) of the window. The "Display Status Flag" column contains a flag value indicating whether the window is displayed or hidden (minimized). In this embodiment, "True" is a value indicating a displayed state, and "False" is a value indicating a hidden state. The "Belonging" column contains a value specifying whether the window belongs to monitor 2, HMD 3, or both. In this embodiment, for convenience, the strings "monitor" and "HMD" are used, but they may also be numeric values, such as "0 = monitor + HMD," "1 = monitor," and "2 = HMD." Note that the items (columns) constituting the table are merely examples; for example, "start point and end point" may be used instead of "size." The origin may also be the upper left corner of both HMD3 and monitor 2. That is, window 404 belonging to HMD3 may have the upper left corner of HMD3 as its origin, and window 405 belonging to monitor 2 may have the upper left corner of monitor 2 as its origin. Furthermore, instead of the upper left corner, the origin may be the center or the lower right corner of the screen. The table may also include information such as the anteroposterior relationship between windows. The anteroposterior relationship determines which window is drawn on the display device when multiple windows are positioned so that they are drawn in the same area. The data format of the window information does not have to be the table format described above; any data format that can grasp the above information about windows corresponding to all running applications is sufficient.
[0022] In S602, a window of interest is determined from among the windows included in the window information acquired in S601, and it is determined whether the window belongs to an HMD. In this embodiment, this determination is made using the value in the "Belonging" column of the table in FIG. 7 described above. If the window belongs to HMD3, S603 is executed next. On the other hand, if the window of interest does not belong to an HMD, this flow is terminated.
[0023] In S603, the affiliation indicated by the window information acquired in S601 for the window of interest is changed from HMD to monitor. In this embodiment, the value in the "affiliation" column of the table in Fig. 7 is changed from "HMD" to "monitor."
[0024] In S604, the display state of the window of interest indicated in the window information acquired in S601 is changed from "display" to "hide." In this embodiment, the value of the "display state flag" column in the table in Fig. 7 is changed from "True" to "False."
[0025] In S605, it is determined whether there are any unprocessed windows. If all windows in the window information have been processed, the process exits this flow. On the other hand, if there are any unprocessed windows, the process returns to S602 and continues. This completes the display determination process according to this embodiment. Returning to the explanation of the flow in FIG. 5.
[0026] In S503, the processing unit 13 processes the display screen of the monitor 2 based on the window information after S502 so that the window of the HMD 3 whose display state has been changed from "display" to "hide" is not displayed on the monitor 2. In this embodiment, if there is an unminimized window among the windows whose value in the "display state" column of the table in Fig. 7 is "False," the window will be minimized (windows that were originally minimized will remain as they are).
[0027] The above is the operational flow of the information processing device 1 when stopping the display on the HMD 3. While the present embodiment has been described using a two-dimensional window as an example of the display screen, a three-dimensional window may also be used. In this case, the "coordinates" included in the window information need to be three-dimensional (x, y, z) rather than two-dimensional (x, y). Furthermore, the "size" does not have to be (width, height) but includes information necessary for three-dimensional display, such as the window orientation and magnification. While the present embodiment has been described using a single monitor 2, two or more monitors may also be used. In this case, the "belonging" included in the window information may contain a character string or the like that uniquely identifies each display device. Although the present embodiment has been described using a single window for each application, multiple windows may also be associated with one application. Furthermore, one application may display a window on both the monitor 2 and the HMD 3. In this case, an icon corresponding to the application may be displayed in both the real icon section 401 and the virtual icon section 403.
[0028] <Variation 1> In the above-described embodiment, the content of the display screen of the HMD 3 is processed based on the display stop information so that it cannot be seen by others, but there may be cases where it is acceptable to display it as is on the monitor 2. In this modified example, an embodiment will be described in which it is up to the user to decide whether or not to display the content of the display screen of the HMD 3 on the monitor 2. FIG. 8 is a flowchart showing the details of the display determination process according to this modified example. Steps that are the same as those in the flowchart of FIG. 6 above will be given the same step numbers, and their explanation will be omitted. The display determination process of this modified example will be described below with reference to the flowchart of FIG. 8.
[0029] In S801, information indicating that the display state of the window of interest displayed on HMD3 has been changed from "display" to "hidden" is saved. For example, a "change flag" column is newly added to the table of window information described above, as shown in FIG. 9, and a flag value indicating whether or not the window has been changed from "display" to "hidden" is saved. This flag value is initialized to "False," which indicates that no change will be made, and when the window has been changed from "display" to "hidden," "True" is substituted, thereby saving the fact that the window of interest has been hidden. This makes it possible to distinguish it from windows that were originally hidden on monitor 2.
[0030] In S802, a user interface (UI) (not shown) is displayed to allow the user to select whether or not to display the window of interest that was displayed on HMD 3 on monitor 2. This UI prompts the user to press, for example, a "Yes" button if the content that was displayed on HMD 3 should also be displayed on monitor 2, or a "No" button if not, and is displayed, for example, as a pop-up on the actual screen of monitor 2. Note that the above-mentioned UI is just an example, and it may also be one that prompts the user to enter a password instead of pressing the "Yes" button, for example.
[0031] In S803, if the user selects via the UI displayed in S902 that the window of interest that was displayed on HMD 3 should also be displayed on monitor 2, then S804 is executed. On the other hand, if the user does not select that the window of interest that was displayed on HMD 3 should also be displayed on monitor 2, this flow is exited.
[0032] In S804, the display state of the window selected by the user is changed from "Hide (False)" to "Display (True)" in the window information so that the window is also displayed on monitor 2. At this time, the change flag is also initialized, thereby preventing malfunctions when the same window is displayed again on HMD 3.
[0033] The above is the content of the display determination process according to this modified example. Note that for windows belonging to HMD3 but currently hidden, the change flag value remains “False” and the window is not subject to redisplay. While an example of displaying a UI for confirming the user's intention has been shown, the user may set whether or not to redisplay the window in advance and store it in RAM 102, etc., and then determine whether or not to execute step S804 by reading the setting instead of steps S802 and S803. While this modified example shows an example of temporarily setting the window to hidden and then redisplaying it, it is not necessary to temporarily set the window to hidden. For example, the above-described UI may be displayed immediately after stopping the display of the window on HMD3, and whether or not to display it on monitor 2 may be controlled according to a user selection. Instead of displaying the UI as a window on monitor 2, a hardware button (not shown) may be provided, and pressing the button may be detected as a user selection.
[0034] <Variation 2> Furthermore, if the window displayed on the HMD 3 is a form for inputting personal information or the like, when the form is also displayed on the monitor 2 in accordance with a user selection, the character string or other portion already entered by the user into the form may be made invisible. FIG. 10 is a flowchart showing details of the display determination process according to this modified example. Steps that are the same as those in the flowchart of FIG. 8 described above will be given the same step numbers, and their explanation will be omitted. The display determination process of this modified example will be described below with reference to the flowchart of FIG. 10.
[0035] In S1001, form information is acquired. For example, in the case of an HTML web page, the form information is acquired by searching for the HTML description of the part corresponding to the form and parsing (syntax analysis) the description found. This form information includes information on whether a character string or the like has been entered, and may also include a form identifier, the form position, the form type, and the like. At this time, the display screen of the HMD 3 may be captured, and a known image recognition technique may be applied to the captured image to estimate the position of the form and whether a character string or the like has been entered. Note that form types that can be viewed by others may be filtered, and form information for those forms may not be acquired.
[0036] In S1002, the next process to be executed is determined based on whether or not a form exists in the window of interest. If the window of interest contains a form, S1003 is executed next; if not, S1004 is executed next. The following steps S1003 to S1005 are executed for each form in the window of interest. Note that a description of S1004 when no form exists will be omitted.
[0037] In S1003, it is determined whether characters or the like have been entered into the form of interest among one or more forms included in the window of interest. If characters or the like have been entered, S1004 is executed next, and if characters or the like have not been entered, S1004 is skipped and S1005 is executed.
[0038] In S1004, the display state of the form of interest is set to invisible. Specifically, for example, a form list is prepared for each window, and for the form of interest, a form identifier is set in association with flag information (form flag) indicating whether to make it invisible. For example, this flag value is set to "True" where "True" indicates visible and "False" indicates invisible, with the initial value being set to "True." As a result, only forms that have already been filled in are set to invisible. At this time, for example, the form may be held in the form of a tuple, which is a data type in which multiple elements are arranged in a fixed order, that is, (form identifier, form flag value).
[0039] In S1005, the next process to be executed is determined depending on whether there are any unprocessed forms in the window of interest. If all forms have been processed, S901 is executed next. On the other hand, if there are any unprocessed forms, the process returns to S1003 and continues. Each step from S801 onwards is the same as in Modification 1, so a description thereof will be omitted. The above is the content of the display determination process according to this modification.
[0040] After that, the process returns to the flow of FIG. 5 and S503 is executed. In S503, the processing unit 13 performs processing to make the form in the window invisible according to the form flag. FIG. 11 is a diagram showing an example of the form invisibility processing in which the form is filled in with a rectangle. The upper part of FIG. 11 shows the state before the display on the virtual display 402 is stopped, and a mailer window 1100 consisting of two forms 1101 and 1102 is displayed. The form 1101 is an empty form into which the user has not yet entered any characters, and the form 1102 is an entered form into which a sentence being created has already been entered. In this state, when the display of the window 1100 on the HMD 3 is stopped, a new window 1110 is displayed on the real screen 400 of the monitor 2. The lower part of FIG. 11 shows the state after the display on the virtual display 402 is stopped, and it can be seen that the empty form 1101 remains displayed as is, and the entered form 1102 is filled in so that the characters, etc., are not visible.
[0041] Note that the invisibility process may be performed by methods other than filling in the form, for example, by scrolling the page so that the completed form becomes invisible. Furthermore, in this modified example, whether to make the form invisible is determined on a form-by-form basis, but this determination may also be made on an even smaller form-by-form basis. For example, the characters in the completed form may be made invisible by replacing them with black circles. In this case, this can be achieved by saving the information on the characters already entered in the form as part of the form information and referencing this in S503. Furthermore, in this modified example, a text input form has been described as an example, but the present invention is not limited to this, and other methods may also be used, such as check boxes, combo boxes, and toggle switches.
[0042] As described above, in this embodiment, when the HMD display is stopped, the window displayed on the monitor is processed so that the window displayed on the HMD is not displayed as is on the monitor. This prevents the window that the user was looking at on the HMD from being displayed as is on the monitor and being seen by others.
[0043] [Embodiment 2] In this embodiment, in a use case where a mixed reality image is viewed through an HMD (see FIG. 1(b) above), an aspect will be described in which the influence on an environment map caused by removing a specific object from a real image is reduced and optical consistency in the mixed reality image is improved. Before going into a detailed description of this embodiment, the problem that this embodiment aims to solve will be described in detail.
[0044] <Checking the assignment> For example, when considering purchasing a new kitchen, a user can check how well the new kitchen will blend with their own room by using a mixed reality image in which a CG image of the new kitchen is superimposed on an image of their own room. One known method for this is to erase the current kitchen (which is currently being installed) from the real-world image before superimposing the CG image of the new kitchen. This technique, known as "obscuring reality," compensates and overwrites the pixel values of the object to be erased (hereinafter referred to as the "obscuring object") in the real-world image that reflects real space with the pixel values of the surrounding area, making the obscuring object appear as if it were not there. In the example above, erasing the current kitchen from the real-world image prevents the current kitchen from appearing to overlay the CG image of the new kitchen. Meanwhile, there is a technique called environment mapping, which expresses reflections in the superimposed CG. Environment mapping stores the surrounding environment reflected in the CG as image data (called an "environment map") and references the environment map when rendering the CG to express reflections in real space in the CG. In environment mapping, the normal vector of the CG surface is used as an axis to determine the reflection direction of the line-of-sight vector from the virtual viewpoint to the surface, and the pixel in the environment map corresponding to the reflection direction is used to represent the CG reflection. In this case, refraction can also be represented by changing the reflection direction according to the refractive index. This improves the optical consistency of the generated mixed reality image. Optical consistency refers to the degree to which the optical representation of CG, such as reflection and shadow, is consistent with reality. High-quality mixed reality images can be generated using technologies such as fading reality and environment maps. However, if the environment map contains fading objects, the fading objects that should have disappeared will affect the CG reflection, reducing the optical consistency between the real space and the CG. This is the challenge of this embodiment.
[0045] <Examples of specific situations where issues arise> A specific example of a situation where optical consistency in a mixed reality image is reduced by performing CG synthesis while an obscuring object is reflected in the environment map will be described.
[0046] <Example 1> The first example situation is one in which the environment map acquisition position and the user viewing position are different. FIG. 12(a) is a diagram illustrating Example 1. A user is standing facing a disappearing object (≒ virtual object) from position A and is viewing a mixed reality (MR) image in which the real-world disappearing object has been erased and a CG (virtual object) has been superimposed on it. The environment map used to render the CG is acquired at the environment map acquisition position (≒ position A). Generally, the environment map is stored as, for example, an equirectangular image. A disappearing object exists in the front direction as viewed from the environment map acquisition position (≒ position A), and the disappearing object is not reflected in the CG in the video viewed by the user from position A. However, when the user moves to position B, the disappearing object reflected in the environment map is used as a reflection in the CG in the video viewed by the user.
[0047] <Example situation 2> The second example situation is when the user is sandwiched between obscuring objects. Figure 12(b) illustrates example situation 2. The user is standing between obscuring objects 1 and 2, facing obscuring object 1 (≒ virtual object 1). The user is viewing a mixed reality (MR) image in which the real obscuring objects 1 and 2 have been erased and CG1 (virtual object 1) has been superimposed on them. The environment map used to render CG1 was acquired at the same position as the user's viewing position. Obscuring object 1 is located directly in front of the environment map acquisition position (≒ position A), and obscuring object 2 is located in the opposite direction. Therefore, obscuring object 2, which is displayed in the environment map, is reflected in CG1 in the user's viewing image. Similarly, when the user views CG2, obscuring object 1 displayed in the environment map is used in the reflection, and obscuring object 1 displayed in the environment map is reflected in CG2 in the viewing image.
[0048] <Functional configuration of information processing device> Fig. 13(a) is a block diagram showing an example of the software configuration (logical configuration) of an information processing device 1 according to this embodiment. In Fig. 13(a), the information processing device 1 is made up of an input receiving unit 10 and an image processing unit 11, and the image processing unit 11 is made up of an object detection unit 14, a feature extraction unit 15, an elimination processing unit 16, an environment map generation unit 17, a removal region determination unit 18, and a removal processing unit 19. Each unit will be described below.
[0049] The input receiving unit 10 receives various user inputs in addition to the real image captured by the HMD 3. As shown in Fig. 14(a), the real image in this embodiment is in a uv coordinate system where the horizontal direction is represented by the u axis and the vertical direction is represented by the v axis, and color values of three RGB channels are held for each pixel. The acquired real image data is output to the object detection unit 14, the elimination processing unit 16, and the environment map generation unit 17.
[0050] The object detection unit 14 detects an occluding object in a real image. This detection is performed by detecting an object specified by the user as a target for occlusion processing among objects appearing in the real image using a known method such as machine learning or pattern matching, and acquiring an area corresponding to the detected occluding object and the type (class) of the occluding object. The area corresponding to the occluding object is, for example, a rectangular area surrounding the occluding object. Information on the rectangular area and class of the detected occluding object is output to the feature extraction unit 15 and the occlusion processing unit 16.
[0051] The feature extraction unit 15 extracts feature information of an obscuring object appearing in a real image. This feature information includes three types of feature: object detection feature, image feature, and three-dimensional feature. The extracted feature information is output to the removal area determination unit 18.
[0052] The environment map generation unit 17 generates an environment map of the real space where the user uses the HMD 3. As shown in FIG. 14(b), the direction at the acquisition position of the environment map can be expressed using polar coordinates (θ, Φ), and the environment map is an image obtained by quantizing these polar coordinates and mapping the light at each polar coordinate. FIG. 14(c) shows an environment map in equirectangular projection. Note that the format of the environment map does not necessarily have to be equirectangular projection, and it can also be a cube map format, a spherical map, or the like. The generated environment map is output to the removal area determination unit 18.
[0053] The removal area determination unit 18 determines a removal area, which is an area to be subjected to removal processing for removing an obscuring object from the environment map. Information on the determined removal area is output to the removal processing unit 19.
[0054] The removal processing unit 19 performs a complementation process on the removal area determined by the removal area determination unit 18, and removes the obscuring object from the environment map.
[0055] <Operation flow of information processing device> Fig. 15 is a flowchart showing the operation flow of the information processing device 1 according to this embodiment. A series of processes shown in the flowchart of Fig. 15 is executed on a frame-by-frame basis. In the following description, the symbol "S" means step.
[0056] In S1501, the input receiving unit 10 acquires real image data from the HMD 3. As described above, the real image acquired here is an image captured by the RGB camera 201.
[0057] In S1502, the input receiving unit 10 receives a user input specifying an obscuring object among the objects appearing in the real image acquired in S1501. The obscuring object may be specified, for example, by a hand gesture in which the user points to the obscuring object using their finger. In the case of a hand gesture, the RGB camera 201 of the HMD 3 detects the user's finger using, for example, known machine learning, and acquires the image coordinates of the fingertip, thereby accepting the specification of the obscuring object. The coordinate information of the obscuring object previously specified and saved by the user may be read from the HDD 105 and acquired. Alternatively, for example, the UV coordinates of the object pointed at by the user may be acquired, or the user's specification may be accepted via the input device 107, such as a mouse.
[0058] In S1503, the environment map generation unit 17 generates an environment map. For example, the environment map is generated by reading from the HDD 105 data of a real image (360-degree image) previously captured and saved using a 360-degree camera, and converting the 360-degree image into an equirectangular image. Here, if the environment map is captured in advance using a 360-degree camera, the position where the environment map was acquired may differ from the position of the HMD 3, resulting in an inappropriate positional relationship between the reflections. To alleviate this situation, a known warping technique may be used to correct the positional deviation of the environment map. Alternatively, a 360-degree camera may be installed above the HMD 3, or multiple cameras may be installed around the HMD 3, and the images may be combined to generate an environment map at the user's position in real time. Note that the method for generating the environment map is not limited to using a 360-degree camera. For example, the environment map may be generated by extrapolating the area outside the angle of view of a camera with an angle of view less than 360 degrees using machine learning, such as deep learning, to interpolate the area outside the angle of view (prerequisite for Modification 1, described later). Furthermore, the visible range of an environment map generated using video from a 360-degree camera may be successively updated (prerequisite for Modification 2 described below). In this case, the orientation of the HMD 3 at startup is set to (θ, Φ) = (0, 0), and the environmental map is updated for the angle of view of the RGB camera 201. When the orientation of the HMD 3 changes, the orientation (θt, Φt) is acquired by a gyro sensor (not shown), and the environmental map for the angle of view of the RGB camera 201 is updated around the orientation (θt, Φt). This process is repeated. At this time, the orientation of the HMD may be estimated using a known SLAM technique instead of a gyro sensor. Alternatively, the initial environment map may be generated using a hybrid method in which areas outside the angle of view of the initial environment map are extrapolated by machine learning, and areas that have been captured by the RGB camera 201 at least once are updated using the captured image.
[0059] In S1504, the object detection unit 14 performs a process of detecting the obscuring object designated by the user for the real image acquired in S1501. The specific procedure is as follows. First, image analysis using known machine learning is performed on the real image to detect objects contained in the real image. FIG. 16(a) is an example of the detection result, showing a rectangular area surrounding an object shown in the real image. Based on the detection result thus obtained, the ID, class, likelihood indicating the class likelihood, and vertex coordinates of the rectangular area are calculated for each object (see the table in FIG. 16(b)). Then, among the detected objects, the object having coordinates closest to the coordinates designated by the user is identified as the obscuring object, and its class and vertex coordinates (u m ,v m ) {m=0,1,2,3} is output as the detection result. Note that the shape of the area representing the detected object does not necessarily have to be rectangular. For example, a known semantic area segmentation technique may be applied to detect a pixel area including coordinates specified by the user as the area of the obscuring object, and information about the pixel area may be output as the detection result.
[0060] In S1505, the feature extraction unit 15 extracts an object detection feature, an image feature, and a three-dimensional feature as feature information of the occluding / disappearing object based on the detection result of S1504. A specific extraction method is as follows.
[0061] <<Extraction of object detection features>> The object detection feature acquires the class of the occluding object included in the above detection result.
[0062] <Image feature extraction> The image feature refers to color information of the obscuring object in the real image. First, a rectangular region of the obscuring object is trimmed from the real image acquired in S1501, and pixel values of the trimmed region are obtained. The rectangular region of the obscuring object may also include its background. In this case, the background may be separated using a known background separation technique, and the average RGB values of only the pixels of the obscuring object in the foreground may be calculated and used as the image feature. The object detection unit 14 may also perform trimming and include the trimmed image in the detection result. In this case, trimming in this step is unnecessary. The value calculated as the image feature is not limited to the average RGB value, but may be other statistical values such as the median. It may also be the pixel value of the center coordinate of the obscuring object region in the real image. It does not necessarily have to be a single value. For example, the colors constituting the obscuring object may be clustered using a known clustering technique such as the X-means algorithm, and a list of the average RGB values for each cluster may be used as the image feature.
[0063] <Extraction of 3D features> The three-dimensional feature amount means the vertex coordinates in the polar coordinate system of the rectangular area of the obscuring object in the real image. Therefore, the vertex coordinates of the rectangular area in the uv coordinate system are converted into a polar coordinate system with the HMD 3 as the origin to obtain the vertex coordinates. First, the image coordinates are converted into a Cartesian coordinate system using the internal parameter K of the RGB camera 201 of the HMD. Here, the internal parameter K is calculated based on the principal point (C x ,C y ) and focal length (f x ,f y The internal parameter K is calculated in advance by a known camera calibration technique and is read from a storage unit such as the HDD 105.
[0064]
number
[0065]
number
[0066] θ=Arccos(z / √(x 2 +y 2 +z 2 )) ...Equation (3) Φ=sgn(y)Arccos(x / √(x 2 +y 2 ))...Equation (4) In this way, the transformation process according to the above formulas (1) to (4) is performed on the vertex coordinates (u m ,v m ) {m=0,1,2,3}, the vertex coordinates (θ m ,Φ m ){m=0,1,2,3} is obtained.
[0067] In S1506, the occlusion processing unit 16 performs occlusion processing on the real image acquired in S1501 to remove the occlusion object from the real image. In this case, for example, a known image completion technique using machine learning is applied to remove the occlusion object from the real image.
[0068] In S1507, the removal area determination unit 18 calculates an area corresponding to the obscuring object from the environment map and determines it as a removal area. Details of the removal area determination process will be described later.
[0069] In S1508, the removal processing unit 19 removes the occluding object from the environment map generated in S1503 based on the removal area determined in S1507. In this case, too, the occluding object is removed from the environment map by applying, for example, a known image completion technique using machine learning to the removal area.
[0070] The above is the content of the operation flow of the information processing device 1 according to this embodiment. Note that although the real image has been described as having three RGB channels in this embodiment, it may be, for example, a five-channel or one-channel black and white image.
[0071] <Details of removal area determination process> 17 is a flowchart showing the details of the removal region determination process in S1507. The following will be explained with reference to the flowchart in FIG.
[0072] In S1701, a removal candidate region is identified based on the object detection features. Specifically, an object corresponding to the class of the occluding object identified as the object detection feature is found in the environment map, and the region of the object is designated as the removal candidate region. First, the object detection technology described above is applied to the environment map to obtain a detection result similar to that of S1504. However, while vertex coordinates are expressed in the UV coordinate system in S1504, this step uses a polar coordinate system. Next, based on the obtained detection result, a region in the environment map where an object of the same class as the occluding object exists is identified as a removal candidate region based on the object detection features. The removal candidate region based on the object detection features thus identified is stored as an object detection feature region image with the same number of pixels as the environment map. Figure 18(a) is an example of an object detection feature region image, with a dashed line indicating a rectangular region in which an object of the same class as the occluding object has been detected. Each rectangular region is assigned an ID, and the pixels that make up each region retain the ID value of the region to which they belong. In the example of Figure 18(a), two objects of the same class as the occluding object are detected, with each pixel in rectangular region 1801 holding an ID of 1, each pixel in rectangular region 1802 holding an ID of 2, and the other regions holding an ID of 0. Note that in this embodiment, an object detection feature region image is generated in which the detected objects are represented by rectangular regions, but this is not limiting. For example, background separation may be applied to the rectangular regions of the environment map, and IDs may be assigned only to the foreground portions and the corresponding pixels in the object detection feature region image, thereby generating an object detection feature region image that represents the shapes of the objects in more detail.
[0073] In S1702, removal candidate regions are identified based on image features. Specifically, regions with colors similar to the average RGB values extracted in S1505 are considered to correspond to obscuring objects, and these regions are designated as removal candidate regions. A similar-color region may be one with a color difference ΔE of less than a predetermined value. If the environment map acquisition position differs from the position of the HMD 3, the lighting environment and the viewpoint from which the obscuring object is photographed change, easily increasing the color difference. To prevent obscuring objects from affecting the CG environment mapping at the color name level, it is desirable to designate a similar-color region as, for example, a region with a color difference of ΔE20 or less. ΔE20 refers to the degree of color difference that allows the color name to match. Note that if the image features are a list of average RGB values for each cluster, for example, if at least one of the average RGB values in the list is less than ΔE20, the region is designated as a obscuring object, and then connected components are calculated. The removal candidate regions identified based on image features are stored as image feature region images with the same number of pixels as the environment map. FIG. 18(b) is an example of an image feature region image, in which connected components of pixels with ΔE less than 20 are represented as black pixel blocks. Here, a connected component is a combination of adjacent pixels with ΔE less than 20. An ID is assigned to each black pixel block, and the pixels that make up each region hold the ID value of the region to which they belong. In the example of FIG. 18(b), three connected components (black pixel blocks) are obtained, with each pixel in black pixel block 1811 holding ID=1, each pixel in black pixel block 1812 holding ID=2, and each pixel in black pixel block 1813 holding ID=3, and the other regions holding ID=0.
[0074] In S1703, removal candidate regions are identified based on the three-dimensional feature amount. Specifically, the vertex coordinates (θ m ,Φ m){m=0,1,2,3}, the corresponding area in the environment map is designated as a removal candidate area. The removal candidate area based on the three-dimensional feature amount thus identified is stored as a three-dimensional feature area image with the same number of pixels as the environment map. Figure 18(c) is an example of a three-dimensional feature area image, with the area of vertex coordinates (θm, Φm){m=0,1,2,3} in the polar coordinate system indicated by dashed lines. In the example of Figure 18(c), each pixel in area 1821 has ID=1, and other areas have ID=0.
[0075] In S1704, the removal area is determined by integrating the three types of removal candidate areas identified in S1701 to S1703 (i.e., the object detection feature area image, image feature area image, and 3D feature area image). If the removal area is determined based solely on the object detection feature, when an object of the same class as the occluding object is detected, that object may also be determined as the removal area (see FIG. 18(a) above). Furthermore, if the removal area is determined based solely on the image feature area image, objects other than the occluding object that have similar image feature values may also be determined as the removal area (see FIG. 18(c) above). To prevent such errors, an AND area is first determined for the three types of removal candidate areas. In other words, the area where areas other than ID=0 overlap in the object detection feature area image, image feature area image, and 3D feature area image is determined. By performing AND (logical product), a removal area based on multiple feature values can be obtained, making it possible to minimize the inclusion of objects other than the occluding object in the removal area. However, simply performing AND tends to result in a smaller area than the actual area of the occluding object. This is because when an obscuring object is composed of multiple colors, only one of the colors may become the image feature region, resulting in a smaller area than the original area of the obscuring object. Furthermore, the viewpoint from which the obscuring object is viewed differs between the position where the real image is acquired and the position where the environment map is acquired, and the shape of the object may change depending on the viewpoint. Therefore, the area represented by the 3D feature region image may also be smaller than the original area of the obscuring object. To prevent the removal region from becoming too small, the OR (logical sum) of the pixel value regions of the object detection feature region image, image feature region image, and 3D feature region image, including the AND region, is taken. For example, suppose the OR region has a pixel value of 1 in the object detection feature region image, a pixel value of 2 in the image feature region image, and a pixel value of 1 in the 3D feature region image. In this case, the removal region is determined by adding together the region with a pixel value of 1 in the object detection feature region image, the region with a pixel value of 2 in the image feature region image, and the region with a pixel value of 1 in the 3D feature region image.
[0076] The above is the details of the removal area determination process according to this embodiment. Note that in this embodiment, object detection features, image features, and three-dimensional features are extracted as feature amounts, and removal candidate areas for occluding objects are obtained from each feature amount and integrated to determine the removal area. However, integrating three slow-moving candidate areas is not an essential configuration, and for example, integration may be performed using two of these removal candidate areas.
[0077] <Variation 1> When extrapolating an environment map of a range corresponding to the outside of the angle of view of a camera with an angle of view of less than 360 degrees, occluding objects on the environment map that is the source of the extrapolation may have an effect. Therefore, in order to reduce this effect, a method of removing occluding objects from the source environment map before performing the extrapolation will be described as Variation 1.
[0078] <Examples of specific situations where issues arise> FIG. 19 is a diagram illustrating an example of a situation in this modified example. First, an environment map (hereinafter referred to as a "limited environment map") of the limited range indicated by the solid arc in FIG. 19 is generated using video from a camera with a field of view narrower than 360 degrees from the environment map acquisition position directly facing the obscuring object. Then, for the range outside the field of view of the camera that is not covered by the limited environment map, extrapolation is performed, for example, using deep learning, to generate an environment map (hereinafter referred to as an "extrapolated environment map") of the range indicated by the two-dot chain arc in FIG. 19. In this case, if the obscuring object reflected in the limited environment map indicated by the solid line is extrapolated as is, for example, the color of the obscuring object may bleed into the extrapolated portion. In such a case, for example, when a user observes the CG from the side, the color bleed reflected in the extrapolated environment map behind the user will be reflected in the CG in the viewed video.
[0079] <Functional configuration of information processing device> FIG. 13(b) is a block diagram showing an example of the software configuration (logical configuration) of the information processing device 1 according to this modification. This modification differs from the block diagram of FIG. 13(a) in that an extrapolation unit 20 follows the removal processing unit 19, and information from the input receiving unit 10 is also provided to the extrapolation unit 20. The input receiving unit 10 outputs to the extrapolation unit 20 angle-of-view information indicating the angle-of-view (angle-of-view smaller than 360 degrees) of the camera that captures the video used by the environment map generation unit 17 when generating the environment map. The removal processing unit 19 removes obscuring objects from the limited environment map generated by the environment map generation unit 17 within a range corresponding to an angle-of-view smaller than 360 degrees, and outputs the removed limited environment map to the extrapolation unit 20. The extrapolation unit 20 extrapolates a range corresponding to outside the angle-of-view of the camera to the limited environment map.
[0080] <Operation flow of information processing device> Fig. 20 is a flowchart showing the operation flow of the information processing device 1 according to this modification. Steps that have the same content as those in the flowchart of Fig. 15 are given the same step numbers, and their explanations will be omitted. Details of this modification will be described below with reference to the flowchart of Fig. 20.
[0081] In S2001, which replaces S1503, the environment map generation unit 17 generates a limited environment map of a range corresponding to an angle of view smaller than 360 degrees. In S2002, which replaces S1508, the removal processing unit 19 removes obscuring objects reflected in the limited environment map. Then, in S2003, the extrapolation unit 20 extrapolates the area outside the range of the limited environment map. That is, the environment map of the range corresponding to the outside of the angle of view indicated by the input angle of view information is inferred (complemented) by, for example, well-known deep learning. The above is the operational flow of the information processing device 1 according to this modification.
[0082] In this modified example, by removing occluding objects reflected in a partial environment map obtained in advance before performing the extrapolation process, it is possible to reduce the adverse effects that occluding objects reflected in the environment map have on the extrapolation process.
[0083] <Variation 2> When the range of the environment map captured by the camera is updated sequentially, a certain occluding object may appear multiple times on the environment map. A method for removing unnecessary occluding objects from the environment map that have appeared multiple times will be described as Variation 2.
[0084] <Examples of specific situations where issues arise> FIG. 21 is a diagram illustrating an example of a situation in this modified example. A 360-degree camera is currently capturing an environment map (shown as a solid line) covering the entire area. Then, a photograph is taken at environment map update position 1, where the occluding object is viewed from the front. The environment map is first updated based on the captured image for the range corresponding to the angle of view indicated by the dashed line. At this time, the occluding object appears in that range of the environment map. Then, the user moves to environment map update position 2, where the occluding object is viewed to the left, and another photograph is taken. The environment map is updated based on the captured image for the range corresponding to the angle of view indicated by the dashed line. Again, the occluding object appears in that range of the environment map. As a result, multiple occluding objects remain visible on the environment map until the next update is performed for the same range based on an image that does not capture the occluding object.
[0085] <Functional configuration of information processing device> Fig. 22(a) is a block diagram showing an example of the software configuration (logical configuration) of the information processing device 1 according to this modification. The difference from the block diagram of Fig. 13(a) is that a posture estimation unit 21 is added, and information from the input reception unit 10 is also provided to the posture estimation unit 21. The input reception unit 10 of this modification further receives input of a depth map and outputs the depth map to the posture estimation unit 21. Then, the posture estimation unit 21 estimates the posture of the HMD 3 based on the depth map. Posture information as the estimation result is output to the feature extraction unit 15 and the environment map generation unit 17.
[0086] <Operation flow of information processing device> Fig. 23 is a flowchart showing the flow of processing in the information processing device 1. Steps that have the same content as those in the flowchart of Fig. 15 above will have the same step numbers, and detailed explanations thereof will be omitted. Details of this modified example will be explained below with reference to the flowchart of Fig. 23.
[0087] In S2301, the input receiving unit 10 acquires a depth map obtained using the distance sensor 202, etc. The depth map is a map with the same number of pixels as the real image, which indicates the distance from the camera for each pixel of the image, and the depth is stored in the pixel corresponding to each pixel of the real image captured by the RGB camera 201.
[0088] In S2302, the posture estimation unit 21 estimates the posture of the HMD 3 based on the depth map acquired in S2301. For example, a known SLAM (Simultaneous Localization and Mapping) technique is applied to this posture estimation, and a rotation matrix R and a translation matrix T that transform the initial posture of the HMD 3 into the current posture are obtained as posture information. Here, the rotation matrix R and the translation matrix T are called external parameters, and in this embodiment, they are described as the initial posture when the HMD 3 is powered on.
[0089] In S2303, the environment map generation unit 17 generates / updates the environment map. That is, the environment map generation unit 17 performs processing to update the environment map within the current angle of view of the RGB camera 201 as needed. In this update, RGB values and posture information of the HMD 3 are stored in each pixel of the environment map.
[0090] In S2304, the object detection unit 14 performs a process of detecting an occluding / disappearing object in the real image, and calculates the class and vertex coordinates (u m ,v m ){m=0,1,2,3}.
[0091] In S2305, the feature extraction unit 15 extracts object detection features, image features, and three-dimensional features as feature information of the occluding or disappearing object based on the detection result of S2304. The methods for extracting the object detection features and image features are the same as those in the above-described embodiment, so a description thereof will be omitted, and only the extraction of the three-dimensional features will be described here.
[0092] The feature extraction unit 15 of this modified example extracts the position (x w , y w ,z w ) is extracted. To do this, first, the vertex coordinates (u m ,v m ) {m=0,1,2,3}, the center coordinates of the obscuring object in the image coordinate system (u ave ,v ave ) is calculated from the depth map obtained in S2301. ave ,v ave Then, using the following equation (5), the center coordinates of the obscuring object in the image coordinate system are calculated as the coordinates (x w , y w ,z w )
[0093]
number
[0094] The steps from S1506 onwards are as explained in the flow of Fig. 15 above. That is, the obscuring process is executed on the real image acquired in S1501 (S1506), and the removal area determination process is executed (S1507). The details of this removal area determination process will be described later. Then, based on the determined removal area, the obscuring object is removed from the environment map generated and updated in S2303 (S1508). This concludes the operation flow of the information processing device 1 according to this modified example.
[0095] <Details of removal area determination process> 24 is a flowchart showing details of the removal region determination process in S1507 according to this modified example. The following will be described with reference to the flowchart in FIG.
[0096] In S2401, similar to S1701 described above, removal candidate regions are identified based on object detection features. In addition, for the identified removal candidate regions, the likelihood representing the object-likelihood included in the table obtained in S2304 (see FIG. 16(b) described above) is linked to the ID and stored as an "object detection score." In the example of FIG. 18(a) described above, the object detection score corresponding to the region with ID=1 and the object detection score corresponding to the region with ID=2 are stored.
[0097] In S2402, similar to S1702 described above, removal candidate regions are identified based on image features. In addition, for the identified removal candidate regions, the average ΔE between the region corresponding to the ID and the occluding object is stored as an "image feature score." In the example of FIG. 18(b) described above, if multiple removal candidate regions with similar image features are identified for the occluding objects corresponding to IDs 1, 2, and 3, the average value of the color difference ΔE for each pixel in each removal candidate region is stored as an "image feature score."
[0098] In S2403, a score map of three-dimensional feature amounts is initialized. This score map has the same number of pixels and the same shape as the environment map, and stores in each pixel a score of the three-dimensional feature corresponding to the pixel position in the environment map (hereinafter referred to as a "three-dimensional feature score"). In this step, each pixel in this score map is initialized to "0." The method for calculating the three-dimensional feature score will be described later.
[0099] In S2402, a pixel of interest (θ I , Φ i The next process to be executed is determined depending on whether or not the obscuring object is reflected in the image. Here, the coordinates (x w , y w , z w) is the pixel of interest (θ I , Φ i If the pixel of interest is within the angle of view of the image captured in the orientation of the HMD 3 at the time when the coordinates (x w , y w , z w ) into the coordinates (u i , v i ) to make this determination.
[0100]
number
[0101] In S2405, the pixel of interest (θ i , Φ i ) is calculated. The three-dimensional feature score indicates how much the position of the HMD 3 has moved from its initial position when the pixel value of the pixel of interest in the environment map is updated. The more the orientation of the HMD 3 has changed since the environment map was updated, the more the viewpoint from which the occluding object is viewed will change. Unless the object is symmetrical about the origin, its shape is likely to change depending on the angle from which it is viewed, and the reliability of the three-dimensional feature values tends to be low. Therefore, the score is determined by how much it has moved from its initial position. The three-dimensional feature score can be calculated, for example, using the following equation (7):
[0102]
number
[0103] In S2406, the three-dimensional feature score obtained in S2405 is used to calculate the target pixel (θ I , Φ i ) position and the score map is updated.
[0104] In S2407, it is determined whether all pixels in the environment map have been processed, and if there are unprocessed pixels, the process returns to S2404 and continues. On the other hand, if all pixels have been processed, S2408 is executed next.
[0105] In S2408, a process is performed to connect adjacent pixels having a pixel value other than "0" in the score map obtained by the processes up to this point. An ID is then assigned to each of the obtained connected components.
[0106] In S2409, removal candidate regions based on the three-dimensional feature values are identified based on the three-dimensional feature scores for each region corresponding to the connected components. Specifically, a three-dimensional feature region image is generated in which the pixels that make up the regions corresponding to the connected components are assigned ID values and the pixels that make up other regions are assigned "0." This three-dimensional feature region image is similar to the above-mentioned Figure 18(c), but in this modified example, it may include multiple three-dimensional feature regions. Then, the average three-dimensional feature score is calculated for each region that corresponds to the connected component (i.e., for each ID).
[0107] In S2410, the area obtained by integrating the three types of removal candidate areas identified in S2401, S2402, and S2409 (i.e., the object detection feature area image, the image feature area image, and the 3D feature area image) is determined as the removal area. In this modified example, integration is performed taking into account the possibility that multiple occluding objects may be captured in the environment map. Specifically, areas where areas with pixel values other than "0" in the three types of feature area images overlap (AND areas) are first found. Then, for each of the obtained AND areas, the scores (object detection score, image feature score, and 3D feature score) of the object detection area, image feature area, and 3D feature area that include that AND area are added together. If the score after the addition is below a threshold, the AND area is not included in the removal area. Then, as in S1704, the removal area is determined by taking the logical OR (logical sum) of the pixel value areas in the object detection feature area image, image feature area image, and 3D feature area image that include an AND area whose score after the addition is equal to or greater than the threshold.
[0108] The above is the content of the removal area determination process according to this modified example. Note that in S2409, instead of threshold processing, regions with high scores corresponding to the number of 3D feature regions in the environment map where occluding objects are displayed may be integrated and determined as removal areas. In this case, first, the number of regions corresponding to the connected components is set to N. Next, an AND area consisting only of image feature regions and object detection feature regions is obtained. Then, the scores of one or more obtained AND areas are added together. In other words, the image feature score and object detection score corresponding to the IDs constituting each AND area are added together. Then, among the one or more AND areas, the scores with the highest combined scores are excluded from the removal area. Finally, an OR is taken to determine the removal area. In this way, the number of occluding objects displayed in the environment map may be estimated based on 3D feature amounts, and removal areas for that number may be calculated based on other feature amounts.
[0109] As described above, according to this modified example, when the range captured by the camera on the environment map is updated sequentially, even if a certain occluding object appears multiple times on the environment map, it is possible to remove unnecessary occluding objects.
[0110] <Variation 3> For example, devices with insufficient memory capacity, such as mobile terminals, may only be able to generate an environment map with a resolution of approximately 1 pixel per degree (ppd). In this case, objects cannot be sufficiently resolved or detected, making it impossible to properly remove occluding objects from the environment map. In such cases, higher optical consistency can be achieved by not performing the removal process. Furthermore, for virtual objects made of materials that exhibit some degree of internal scattering, a low-resolution environment map may be sufficient. However, even in such cases, if the color of the occluding object is unique in the real image, that color may be used for CG reflections, reducing optical consistency. One possible solution to this problem is to remove the color of the occluding object from the environment map. However, removing the color of all occluding objects from the environment map would have the adverse effect of removing the color of any other objects of similar color. Therefore, a third variation on the method for determining whether it is appropriate to remove an occluding object from the environment map is described.
[0111] <Functional configuration of information processing device> Fig. 22(b) is a block diagram showing an example of the software configuration (logical configuration) of the information processing device 1 according to this modification. What differs from the block diagram of Fig. 13(a) is that a removal determination unit 22 is added, and information from the input receiving unit 10 is also provided to the removal determination unit 22. The input receiving unit 10 of this modification outputs real image data to the removal determination unit 22. In addition, the feature extraction unit 15 outputs the extracted feature amount to the removal determination unit 22. Then, the removal determination unit 22 determines whether to remove the occluding object from the environment map, and outputs the determination result to the removal area determination unit 18.
[0112] <Operation flow of information processing device> Fig. 25 is a flowchart showing the flow of processing in the information processing device 1. Steps that have the same content as the flowchart in Fig. 15 above will have the same step numbers, and detailed explanations thereof will be omitted. Details of this modified example will be explained below with reference to the flowchart in Fig. 25.
[0113] In S2501, the removal determination unit 22 determines whether or not to remove the disappearing object from the environment map. Specifically, it determines whether or not there is an area in the real image other than the area corresponding to the disappearing object, which has a similar color and a color difference less than a threshold value with respect to the RGB value of the area corresponding to the disappearing object (for example, the average RGB value of the area corresponding to the disappearing object extracted in 1505). In this case, for example, ΔE20 is used as the threshold. As a result of the determination, if there is no other object of a similar color to the disappearing object in the real image, it determines that the disappearing object should be removed, and proceeds to execution of S1507. On the other hand, if there is another object of a similar color to the disappearing object in the real image, it determines that the disappearing object should not be removed, skips S1507 and S1508, and ends this flow.
[0114] The above is the operation flow of the information processing device 1 according to this modification. As a result, by not performing the removal process when an occluding object cannot be appropriately removed from the environment map, it is possible to improve optical consistency.
[0115] As described above, according to this embodiment including the various modified examples, when performing the obscuration process on the real image, the environment map is also appropriately treated, thereby improving the optical consistency of the CG in the mixed reality image.
[0116] <Other embodiments> The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0117] The present disclosure also includes the following configurations and methods.
[0118] [Configuration 1] An information processing device that controls a non-head-mounted first display device and a head-mounted second display device, an acquisition means for acquiring information on stopping display of the virtual screen on the second display device; a processing means for determining, when the acquisition means acquires the stop information, a display content of the real screen of the first display device at the time when the display of the virtual screen of the second display device is stopped; An information processing device comprising:
[0119] [Configuration 2] 2. The information processing device according to configuration 1, wherein the second display device is a device that displays a mixed reality image in which the virtual screen is superimposed on an image captured of a real space.
[0120] [Configuration 3] 3. The information processing device according to configuration 2, wherein the processing means determines that the content displayed on the virtual screen is not included in the display content on the real screen of the first display device.
[0121] [Configuration 4] The processing means When the acquisition means acquires the stop information, a user interface is displayed on the real screen, allowing a user to select whether or not to include the content displayed on the virtual screen in the display content of the real screen; determining whether or not to include the content displayed on the virtual screen in the content displayed on the real screen based on a user input via the user interface; 3. The information processing device according to configuration 2.
[0122] [Configuration 5] 5. The information processing device according to any one of configurations 2 to 4, wherein, when it is determined that the content displayed on the virtual screen should not be included in the display content of the real screen, the processing means processes the real screen so that the content displayed on the virtual screen cannot be confirmed.
[0123] [Configuration 6] 6. The information processing device according to configuration 5, wherein the processing is a process of minimizing a window that was displayed on the virtual screen on the real screen.
[0124] [Configuration 7] The information processing device described in configuration 5, characterized in that the processing is a process of hiding, on the actual screen of the first display device, a filled-out form among forms in a window displayed on the virtual screen of the second display device.
[0125] [Configuration 8] The information processing device according to configuration 5, wherein the processing is a process of hiding an area displayed on the virtual screen of a window that is displayed across the real screen of the first display device and the virtual screen of the second display device in the mixed reality image, on the real screen.
[0126] [Configuration 9] An information processing device that controls a head-mounted display device that displays a mixed reality image on which a virtual object is superimposed, a concealment means for removing a specific object from an image of a real space for the mixed reality image; a removal means for removing the specific object from an environment map used when rendering the virtual object; An information processing device comprising:
[0127] [Configuration 10] 10. The information processing device according to configuration 9, further comprising: a generating means for generating the environment map from an image obtained by capturing the real space.
[0128] [Configuration 11] a determination means for determining an elimination area in the environment map in which the specific object is reflected, based on feature information of the specific object in an image captured of the real space; the removal means removes the specific object based on the determined removal area. 11. The information processing device according to configuration 10.
[0129] [Configuration 12] a detection means for detecting an area in which the specific object is captured from the captured image of the real space; extraction means for extracting the feature information based on the detection result by the detection means; 12. The information processing device according to configuration 11, comprising:
[0130] [Configuration 13] The information processing device according to configuration 12, characterized in that the feature information is at least one of an object detection feature representing a type of the detected specific object, an image feature representing color information of the detected specific object, and a three-dimensional feature representing vertex coordinates in a polar coordinate system of an area corresponding to the detected specific object.
[0131] [Configuration 14] The information processing device described in configuration 13, characterized in that when the determination means makes the determination based on two or more features among the object detection features, the image features, and the three-dimensional features, it determines the area obtained by taking a logical product and / or logical sum of the areas obtained for each feature as the removal area.
[0132] [Configuration 15] an extrapolation means for extrapolating a range outside the angle of view of a camera having an angle of view of less than 360 degrees, the range of the environment map being generated by the generation means and corresponding to the angle of view, by interpolation using machine learning; the extrapolation means performs the extrapolation on the environment map after the specific object has been removed; 15. The information processing device according to configuration 14.
[0133] [Configuration 16] the generating means updates the range of the generated environment map that is captured in an image obtained by a camera that captures the real space; The information processing device according to configuration 14, wherein the determining means determines the removal area based on the camera posture information when the image at the time of the update was captured and the three-dimensional feature obtained by the extracting means.
[0134] [Configuration 17] a determining means for determining whether to remove the particular object from the environment map; 17. The information processing device according to any one of configurations 9 to 16, wherein the removal means removes the specific object from the environment map when the determination means determines that the specific object should be removed.
[0135] [Configuration 18] 18. The information processing device according to configuration 17, wherein the determination means determines not to remove an object if the image captured of the real space contains an object whose color difference with the specific object is less than a threshold value.
[0136] [Configuration 19] 19. The information processing device according to configuration 18, wherein the threshold value is ΔE20.
[0137] [Method 1] A control method for an information processing device that controls a non-head-mounted first display device and a head-mounted second display device, an acquisition step of acquiring information indicating whether or not the display of the virtual screen on the second display device has stopped; a processing step of determining, when the stop information is acquired in the acquisition step, a display content of the real screen of the first display device when the display of the virtual screen of the second display device is stopped; A control method comprising:
[0138] [Method 2] A control method for an information processing device that controls a head-mounted display device that displays a mixed reality image on which a virtual object is superimposed, comprising: an obscuration step of removing a specific object from an image of a real space for the mixed reality image; a removal step of removing the particular object from an environment map used in rendering the virtual object; A control method comprising:
[0139] [Configuration 20] 20. A program that causes a computer to function as the information processing device according to any one of configurations 1 to 19.
Claims
1. An information processing device that controls a non-head-mounted first display device and a head-mounted second display device, an acquisition means for acquiring information on stopping display of the virtual screen on the second display device; a processing means for determining, when the acquisition means acquires the stop information, a display content of the real screen of the first display device at the time when the display of the virtual screen of the second display device is stopped; An information processing device comprising:
2. 2. The information processing apparatus according to claim 1, wherein the second display device is a device that displays a mixed reality image in which the virtual screen is superimposed on an image captured of a real space.
3. 3. The information processing apparatus according to claim 2, wherein said processing means determines that the content displayed on said virtual screen is not to be included in the display content on the real screen of said first display device.
4. The processing means When the acquisition means acquires the stop information, a user interface is displayed on the real screen, allowing a user to select whether or not to include the content displayed on the virtual screen in the display content of the real screen; determining whether or not to include the content displayed on the virtual screen in the content displayed on the real screen based on a user input via the user interface; 3. The information processing apparatus according to claim 2, wherein:
5. 3. The information processing device according to claim 2, wherein, when it is determined that the content displayed on the virtual screen is not to be included in the display content of the real screen, the processing means processes the real screen so that the content displayed on the virtual screen cannot be confirmed.
6. 6. The information processing apparatus according to claim 5, wherein the processing is a process of minimizing a window displayed on the virtual screen on the real screen.
7. 6. The information processing device according to claim 5, wherein the processing is a process of hiding, on the actual screen of the first display device, an inputted form among forms in a window displayed on the virtual screen of the second display device.
8. The information processing device according to claim 5, characterized in that the processing is a process of hiding an area that was displayed on the virtual screen of a window that was displayed across the real screen of the first display device and the virtual screen of the second display device in the mixed reality image, on the real screen.
9. An information processing device that controls a head-mounted display device that displays a mixed reality image on which a virtual object is superimposed, a concealment means for removing a specific object from an image of a real space for the mixed reality image; a removal means for removing the specific object from an environment map used when rendering the virtual object; An information processing device comprising:
10. 10. The information processing apparatus according to claim 9, further comprising: a generating unit that generates the environment map from an image obtained by capturing the real space.
11. a determination means for determining an elimination area in the environment map in which the specific object is reflected, based on feature information of the specific object in an image captured of the real space; the removal means removes the specific object based on the determined removal area.
11. The information processing apparatus according to claim 10,
12. a detection means for detecting an area in which the specific object is captured from the captured image of the real space; extraction means for extracting the feature information based on the detection result by the detection means; 12. The information processing apparatus according to claim 11, further comprising:
13. 13. The information processing device according to claim 12, wherein the feature information is at least one of an object detection feature representing a type of the detected specific object, an image feature representing color information of the detected specific object, and a three-dimensional feature representing vertex coordinates in a polar coordinate system of an area corresponding to the detected specific object.
14. The information processing device according to claim 13, characterized in that when the determination means makes the determination based on two or more feature amounts among the object detection feature amount, the image feature amount, and the three-dimensional feature amount, the determination means determines the area obtained by taking a logical product and / or a logical sum of the areas obtained for each feature amount as the removal area.
15. an extrapolation means for extrapolating a range outside the angle of view of a camera having an angle of view of less than 360 degrees, the range of the environment map being generated by the generation means and corresponding to the angle of view, by interpolation using machine learning; the extrapolation means performs the extrapolation on the environment map after the specific object has been removed; 15. The information processing apparatus according to claim 14,
16. the generating means updates the range of the generated environment map that is captured in an image obtained by a camera that captures the real space; 15. The information processing apparatus according to claim 14, wherein the determining means determines the removal area based on the orientation information of a camera when the image at the time of the update was captured and the three-dimensional feature amount obtained by the extracting means.
17. a determining means for determining whether to remove the particular object from the environment map; The information processing apparatus according to claim 9 , wherein the removal means removes the specific object from the environment map when the determination means determines that the specific object should be removed.
18. 18. The information processing apparatus according to claim 17, wherein the determining means determines not to remove the specific object when an object having a color difference less than a threshold value with respect to the specific object is captured in the image of the real space.
19. 19. The information processing apparatus according to claim 18, wherein the threshold value is ΔE20.
20. A control method for an information processing device that controls a non-head-mounted first display device and a head-mounted second display device, an acquisition step of acquiring information indicating whether or not the display of the virtual screen on the second display device has stopped; a processing step of determining, when the stop information is acquired in the acquisition step, a display content of the real screen of the first display device when the display of the virtual screen of the second display device is stopped; A control method comprising:
21. A control method for an information processing device that controls a head-mounted display device that displays a mixed reality image on which a virtual object is superimposed, comprising: an obscuration step of removing a specific object from an image of a real space for the mixed reality image; a removal step of removing the particular object from an environment map used in rendering the virtual object; A control method comprising:
22. A program that causes a computer to execute the control method of claim 20 or 21.
Citation Information
Patent Citations
Image processing method and image processor
JP2007018173A
Projection system, projection device, projection method, and control program
JP2012233963A