Information processing apparatus, information processing method, and computer-readable recording medium
By acquiring the viewpoint and focus of attention of the subject, determining the virtual intersection point, and controlling the illumination direction of the illumination device, the problem of guiding the gaze on the screen is solved, achieving a natural gaze guidance effect.
Patent Information
- Application Number
- CN202480049747.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-08
- Filing Date
- 2024-06-24
- Publication Date
- 2026-02-27
AI Technical Summary
Existing technologies struggle to effectively guide the subject's gaze when using screens for chroma key synthesis, leading to a decline in subject image extraction performance.
By obtaining the viewpoint and focus of attention of the subject, the virtual intersection of the virtual line and the screen is determined, and the illumination direction of the illumination device is controlled so that the light spot corresponds to the virtual intersection in real space, thus guiding the subject's line of sight.
Without compromising the performance of subject image extraction, it effectively guides the subject's gaze, providing a natural viewing experience.
Smart Images

Figure CN121587018A_ABST
Abstract
Description
Technical Field
[0001] This technology relates to information processing apparatus, information processing method, and computer-readable recording medium for, for example, guiding the line of sight of a subject. Background Technology
[0002] Patent Document 1 discloses a display imaging device in which a transmissive display is arranged between a camera device and a user, and a video calling system using the device. In this system, a camera device located behind the display captures an image of a user looking at the display, which shows a partner at the other end of the line. This allows a video of the user facing forward to be displayed to the partner at the other end of the line. This allows the user's line of sight to be aligned with the partner's line of sight (e.g., paragraphs
[0012] ,
[0060] , and
[0061] in the specification of Patent Document 1, and...). Figure 1 and Figure 7 ).
[0003] Citation List
[0004] Patent documents
[0005] Patent Document 1: WO 2019 / 207922 Summary of the Invention
[0006] Technical issues
[0007] In recent years, techniques for combining live-action video of a subject to recreate the subject in, for example, virtual space have attracted attention. When capturing video for this combination, a screen for chroma keying is used. Essentially, such screens have limited features, and sometimes the subject is unaware of where to focus their attention during image capture. For example, a monitor or similar device is used to guide the subject's gaze; the monitor or similar device appears in the subject's video, which can lead to difficulties in accurately extracting the subject's image.
[0008] In view of the above, the purpose of this technology is to provide an information processing apparatus, an information processing method, and a computer-readable recording medium that enables the guidance of the subject's line of sight without compromising performance when extracting images of the subject.
[0009] Solution to the problem
[0010] To achieve the above objectives, the information processing apparatus according to embodiments of the present technology includes an acquisition unit, a determination unit, and a control unit. The acquisition unit acquires: the viewpoint position of the subject capturing its image using a screen provided in real space for chroma key synthesis, the desired focus point of attention for the subject, and virtual spatial information indicating the position of the screen in a virtual three-dimensional space. The determination unit determines the virtual intersection point of a virtual line with the screen indicated by the virtual spatial information, the virtual line connecting the viewpoint position and the focus point of attention of the subject in the virtual three-dimensional space. The control unit controls the illumination direction of illumination performed by an illumination device provided in real space, pointing to a guide point on the screen corresponding to the virtual intersection point in real space.
[0011] In this information processing device, a virtual intersection point is determined, at which a virtual line connecting the viewpoint position of the subject and the desired focus point of attention of the subject intersects with the screen used for chroma key synthesis. Then, the illumination device points to a guide point located on the actual screen corresponding to the virtual intersection point. Pointing to the guide point enables the presentation of the desired direction of the subject's gaze. This allows the subject's line of sight to be guided without compromising the performance of image extraction.
[0012] The point of focus can be the viewpoint of a viewer viewing a display device on which an image of the subject is extracted from an image of the subject captured against a screen background.
[0013] The acquisition unit can acquire the viewpoint position of the subject based on the image of the subject, which is captured by multiple camera devices set in real space such that the screen is the background within the image capture range.
[0014] The acquisition unit can generate virtual space information based on images captured by multiple camera devices.
[0015] The acquisition unit can use multiple camera devices to capture images of the light spot, which are then projected onto a screen by an illumination device; and can calculate the three-dimensional coordinates of the light spot based on the images captured using the multiple camera devices.
[0016] The control unit can calculate the irradiation direction based on the position of the guide point and the position of the irradiation device in real space.
[0017] The control unit can calculate the rotation angle of the irradiation direction relative to the reference irradiation direction, which is used as a reference for the irradiation direction in the irradiation device.
[0018] The control unit can calibrate the position of the irradiation device and the reference irradiation direction based on the rotation angle of the irradiation direction when the irradiation device projects the light spot onto each of the multiple camera devices, and based on the installation position of the multiple camera devices.
[0019] The illumination device may include position detection markers. In this case, the control unit can calculate the position of the illumination device and the reference illumination direction based on images of the position detection markers of the illumination device captured by multiple camera devices.
[0020] The control unit can determine whether the irradiation path connecting the position of the irradiation device and the guide point in the real space passes through a non-irradiation area that should be avoided from pointing to the guide point, and when it is determined that the irradiation path passes through a non-irradiation area, the control unit can stop pointing to the guide point.
[0021] The illumination device can be a laser pointer that can rotate in both the elevation and azimuth directions.
[0022] The illumination device can be a laser beam scanning projector that uses a movable mirror to scan a laser beam. In this case, the control unit can control the projector so that a pattern including at least one of text or images is projected onto the guide point.
[0023] The screen can be a fully surround screen that encloses the entire periphery of the image capture space surrounding the subject. In this case, the control unit can control the illumination direction performed by the illumination device, such that it points to a guide point in an area on the fully surround screen that is located in front of the subject when viewed from the subject.
[0024] The location of focus of attention can be the image capture location performed by a virtual camera device set up in a virtual three-dimensional space.
[0025] The focus of attention can be the location of the target of attention that the subject is expected to look at within the background, which is combined with the subject image extracted from the captured image of the subject.
[0026] The information processing method according to embodiments of this technology is an information processing method executed by a computer system. This information processing method includes: acquiring the viewpoint position of a subject capturing an image using a screen disposed in real space for chroma keying, the desired focus point of attention for the subject, and virtual spatial information indicating the position of the screen in a virtual three-dimensional space; determining a virtual intersection point between a virtual line and the screen indicated by the virtual spatial information, the virtual line connecting the viewpoint position and the focus point of attention of the subject in the virtual three-dimensional space; and controlling the illumination direction of illumination performed by an illumination device disposed in real space, such that it points to a guide point on the screen corresponding to the virtual intersection point in real space.
[0027] A computer-readable recording medium according to embodiments of the present technology contains a program that causes a computer system to perform processing, the processing including: acquiring the viewpoint position of a subject capturing an image using a screen for chroma keying disposed in real space, the desired focus of attention of the subject, and virtual spatial information indicating the position of the screen in a virtual three-dimensional space; determining a virtual intersection point between a virtual line and the screen indicated by the virtual spatial information, the virtual line connecting the viewpoint position and the focus of attention of the subject in the virtual three-dimensional space; and controlling the illumination direction of illumination performed by an illumination device disposed in real space, such that it points to a guide point on the screen in real space corresponding to the virtual intersection point. Attached Figure Description
[0028] [ Figure 1 ] Figure 1 An example of the configuration of a telepresence system according to this embodiment is illustrated schematically.
[0029] [ Figure 2 ] Figure 2 An example configuration of an image capture system is illustrated schematically.
[0030] [ Figure 3 ] Figure 3 An example of the configuration of the irradiation device is shown schematically.
[0031] [ Figure 4 ] Figure 4 This is a block diagram illustrating an example of the functional configuration of the gaze guidance controller.
[0032] [ Figure 5 ] Figure 5 This is a schematic diagram used to describe the basic operation of a line-of-sight guidance system.
[0033] [ Figure 6 ] Figure 6 This is a flowchart illustrating an example of how a line-of-sight guidance system operates.
[0034] [ Figure 7 ] Figure 7 This is a flowchart illustrating an example of three-dimensional spatial measurement processing.
[0035] [ Figure 8 ] Figure 8 This is a schematic diagram illustrating an example of three-dimensional spatial measurement processing.
[0036] [ Figure 9 ] Figure 9 This is a flowchart of an example of subject viewpoint measurement processing.
[0037] [ Figure 10 ] Figure 10This is a schematic diagram used to illustrate an example of coordinate transformation processing.
[0038] [ Figure 11 ] Figure 11 This is a schematic diagram illustrating an example of the pointing operation of an irradiation device.
[0039] [ Figure 12 ] Figure 12 This is a schematic diagram illustrating another example of the pointing operation of the irradiation device.
[0040] [ Figure 13 ] Figure 13 This is a schematic diagram illustrating a method for calibrating an irradiation device.
[0041] [ Figure 14 ] Figure 14 This is a schematic diagram illustrating another method for calibrating an irradiation device.
[0042] [ Figure 15 ] Figure 15 Another example of the configuration of the irradiation device is shown schematically.
[0043] [ Figure 16 ] Figure 16 This is a schematic diagram illustrating another example of an application of a line-of-sight guidance system.
[0044] [ Figure 17 ] Figure 17 This is a schematic diagram illustrating another example of an application of a line-of-sight guidance system.
[0045] [ Figure 18 ] Figure 18 This is a schematic diagram illustrating another example of an application of a line-of-sight guidance system. Detailed Implementation
[0046] Embodiments according to this technology will now be described below with reference to the accompanying drawings.
[0047] [Overview of Telepresence Systems]
[0048] Figure 1 An example configuration of a telepresence system according to this embodiment is illustrated schematically. The telepresence system 100 is a system that displays a video of user 1 in real time at a viewing side position, the video being captured at an image capture side position. In other words, the telepresence system 100 functions as a real-time streaming system that provides a user 2 located at the viewing side position with the live video of user 1 captured at the image capture side position.
[0049] In the following text, the user 1 who captures the image at the image capture side position is referred to as Subject 1, and the user 2 who watches the video at the viewing side position is referred to as Viewer 2. Furthermore, the video displayed at the viewing side position in which Subject 1 appears is referred to as Content Video 5. In this embodiment, Content Video 5 is a real-time video that allows the state of Subject 1 to be displayed in real time.
[0050] Figure 1 The lower section schematically illustrates an example of the configuration of the telepresence system 100 at the image capture side location, and Figure 1 The upper section schematically illustrates an example of the configuration of the telepresence system 100 at the viewing side position. The telepresence system 100 includes an image capture system 10, a display system 20, and an eye-guiding system 30.
[0051] Image capture system 10 is a system positioned at the image capture side position that captures images of subject 1 to generate content data of subject 1. The content data is data used to construct the content video 5 displayed at the viewing side position. Display system 20 is a system positioned at the viewing side position that displays the content video 5 based on the content data generated by image capture system 10.
[0052] The gaze guidance system 30 is a system positioned at the image capture side and guiding the gaze L1 of the subject 1 during image capture. The gaze L1 of the subject 1 is controlled such that the gaze L1 of the subject 1 displayed in the content video 5 is directed toward the viewpoint P2 of the viewer 2. When the gaze L1 of the subject 1 is guided as described above, it allows the viewer 2's gaze to align with the subject 1's gaze. This provides the viewer 2 with a natural viewing experience where the viewer 2's gaze aligns with the subject 1's gaze.
[0053] Image capture system
[0054] like Figure 1 As shown in the lower part, the image capture system 10 includes a green screen 11, an image capture unit 12, and an image capture side controller 14. The green screen 11 and the image capture unit 12 are mechanical devices used for image capture in real space. Here, real space refers to three-dimensional space encompassed in the real world. The image capture studio using the mechanical devices is created at the image capture side location (real space). The image capture side controller 14 is, for example, a computer located at the image capture side location, and can, for example, be a server device located at a location different from the image capture studio location, as the image capture side controller 14.
[0055] Green Screen 11 is a screen set up in real space and used for chroma keying. Green Screen 11 includes a green surface (screen surface), also referred to as a green background, and serves as the background for capturing the image used for chroma keying. In addition to Green Screen 11, other image capture screens suitable for chroma keying, such as blue screens, can also be used appropriately. It should be noted that in Figure 1 The illustration of the green screen 11, which is located behind the subject 1 and is used for chroma key synthesis, is omitted.
[0056] For example, chroma key synthesis is a method that involves separating regions with uniform color from a captured image and combining another video with the separated regions. For instance, this method enables the separation of a green background from an image of a subject 1 captured against a green screen 11, and the extraction of the subject 1 corresponding to the foreground. As described above, the image obtained by extracting the subject 1 corresponding to the foreground from the region corresponding to the background is called the subject image 3.
[0057] The image capture unit 12 includes a plurality of camera devices 13, each for capturing an image of the subject 1. For example, any digital camera device including an image sensor such as a complementary metal-oxide-semiconductor (CMOS) sensor or a charge-coupled device (CCD) sensor is used as each of the plurality of camera devices 13.
[0058] Each of the plurality of camera devices 13 captures an image in which the subject 1 appears. Note that examples of images in this disclosure include moving images (videos) with time-varying characteristics and still images without time-varying characteristics. The following description primarily focuses on the case where the plurality of camera devices 13 capture video of the subject 1. However, this technique can also be applied to the case of capturing still images of the subject 1.
[0059] The camera devices 13 of the plurality of camera devices 13 are arranged in real space such that the green screen 11 is the background within the image capture range. Here, the image capture range refers to, for example, the range within the viewing angle of the camera device 13. The green screen 11 is the background of the images of the subject 1 captured by the plurality of camera devices 13. Therefore, each of the images is an image from which the subject image 3 can be accurately extracted. Furthermore, the camera devices 13 of the plurality of camera devices 13 are arranged to capture images of the subject 1 from different directions. This makes it possible to capture the subject image 3 of the subject viewed from various directions.
[0060] The image capture side controller 14 extracts the subject image 3 from the image of the subject 1 captured by the image capture unit 12 (multiple camera devices 13), and uses the subject image 3 to generate content data. For example, the subject image 3 can be used to create a 3D model 6 of the subject 1. In this case, the data of the 3D model 6 is the content data. Furthermore, an image in which only the subject image 3 appears (extracted image) or an image obtained by combining any background with the subject image 3 (combined image) can be generated as content data. The content data generated by the image capture side controller 14 is sent to the viewing side location via a designated network using a protocol such as User Datagram Protocol (UDP).
[0061] [Display System]
[0062] like Figure 1 As shown in the upper part, the display system 20 includes a display device 21, a viewpoint sensor 22, and a viewing-side controller 23. The display device 21 and the viewpoint sensor 22 are mechanical devices for display in real space. The viewing-side controller 23 is, for example, a computer located at a viewing-side position, and can, for example, a server device located in a different location can be used as the viewing-side controller 23.
[0063] Display device 21 is a device for displaying content video 5 thereon. As described above, content video 5 is a video generated using subject image 3, and the video includes subject image 3. In other words, display device 21 can also be a device for displaying subject image 3 extracted from an image of subject 1, which was captured against a green screen 11.
[0064] For example, a display device comprising multiple light-emitting diodes (LEDs) arranged on the display surface (so-called LED wall) of the display device is used as display device 21. For example, Figure 1 The display device 21 shown is provided with a display surface including a floor surface and a vertical surface arranged perpendicular to the floor surface, wherein LEDs are used on both the floor surface and the vertical surface. For example, combining the image display on the vertical surface with the image display on the floor surface enables a display, for example, to make the viewer perceive that the subject 1 jumps out from the vertical surface or the floor surface.
[0065] Note that the configuration of display device 21 is not limited, and for example, a display including display elements such as a liquid crystal panel or an organic EL panel can be used as display device 21. Furthermore, display device 21 can be a display that displays the same video to both eyes of viewer 2, or it can be a display capable of performing stereoscopic display, which allows viewer 2 to perceive a stereoscopic image by independently displaying images with parallax between their right and left eyes. Additionally, a head-mounted display (HMD) worn on the head of viewer 2 can be used as display device 21.
[0066] The viewpoint sensor 22 is a sensor that detects the viewpoint P2 of viewer 2. Specifically, the viewpoint sensor 22 is a sensor capable of detecting the three-dimensional position of viewpoint P2 of viewer 2 in real space. The processing of viewpoint detection P2 can be performed by a computing device included in the viewpoint sensor 22 or by the view-side controller 23, which will be described later.
[0067] exist Figure 1 In the example shown, it is assumed that viewer 2 is watching display device 21 while wearing glasses 26 that include markers 25 for position detection. In this case, the position of glasses 26 is detected as viewer 2's viewpoint P2. A reflective marker that reflects infrared light is used as marker 25 set on glasses 26. In this case, for example, an infrared camera that illuminates infrared light into the image capture range to detect the reflected infrared light is used as viewpoint sensor 22. For example, marker 25 is detected in the video of viewer 2 captured using viewpoint sensor 22, and the position and posture of glasses 26 are calculated based on the positional relationship between markers 25. This makes it possible to detect not only the position of viewer 2's viewpoint P2, but also the orientation of viewer 2's gaze L2.
[0068] The type of viewpoint sensor 22 is not limited, and for example, a sensor capable of detecting depth, such as a ToF camera or a stereo camera, can be used as the viewpoint sensor 22. In this case, the process of detecting the viewer 2's face in the image captured by the camera and calculating the three-dimensional position of the viewer 2's viewpoint P2 is performed. Furthermore, any sensor capable of detecting the three-dimensional position of viewpoint P2 can be used as the viewpoint sensor 22. Additionally, when using an HMD as the display device 21, the three-dimensional position of the HMD at the viewing side position can be calculated as the viewpoint P2 of the viewer 2.
[0069] The viewing-side controller 23 receives content data generated at the image capture-side position via a designated network. Furthermore, the viewing-side controller 23 generates a content video 5 to be displayed on the display device 21 based on the content data, and outputs the generated content video 5 to the display device 21. In this embodiment, the content video 5 is generated based on the viewpoint P2 of the viewer 2 detected using the viewpoint sensor 22. This enables the generation of a content video 5 in which, for example, the presentation of the subject image 3 changes according to the movement of the viewpoint P2 of the viewer 2.
[0070] [Specific configuration of the image capture system]
[0071] Figure 2 An example of the configuration of the image capture system 10 is shown schematically. In the image capture system 10 (image capture studio), in order to capture an image of the subject 1, a green screen 11 and a plurality of camera devices 13 included in the image capture unit 12 are arranged around the image capture space 16. For example, a relatively large space is set as the image capture space 16 to allow the subject 1 to move freely.
[0072] Green screen 11 is an all-around screen that surrounds the entire periphery of the image capture space 16 in which the image of the subject 1 is captured. Therefore, Figure 2 The image capture studio shown is an image capture studio with an all-round green background. Using an all-round screen allows the green screen 11 to be used as a background even when capturing images of subject 1 from, for example, all directions. Furthermore, the shape of the green screen 11 corresponds to the shape of the image capture space 16 without change. Note that it is sufficient if the all-round screen is located behind and in front of subject 1, and the all-round screen does not necessarily have to surround the entire circumference of subject 1 to cover the full 360-degree circle. For example, the green screen 11 in the form of an all-round screen may include a non-image capture area that includes an entrance to the image capture space 16.
[0073] exist Figure 2 In the example shown, a green screen 11 with a cylindrical shape is formed. In this case, the cylindrical sidewalls and floor surface are green. Furthermore, the cylindrical space surrounded by the green screen 11 corresponds to the image capture space 16. Note, for example, that the lighting equipment is appropriately arranged above the image capture space 16 (ceiling portion). Furthermore, the shape of the green screen 11 is not limited, and for example, a green screen 11 with a cuboid shape or a polygonal prism shape can be used.
[0074] Multiple camera devices 13 are arranged to surround the image capture space 16. In other words, images of the subject 1 are captured using multiple camera devices 13 arranged to surround the image capture space 16. For example, each camera device 13 is arranged along the sidewall of the green screen 11 and oriented toward the central portion of the image capture space 16. In this embodiment, the multiple camera devices 13 (image capture unit 12) are a full-around camera device that captures images of the subject 1 from all directions. This makes it possible to capture images of the subject 1 from all directions without blind spots. The multiple camera devices 13 can be embedded in the green screen 11, or they can be arranged inside the green screen 11 using designated fixing devices.
[0075] Furthermore, it is assumed that the intrinsic parameters and extrinsic parameters of each of the multiple imaging devices 13 are known in advance. Examples of intrinsic parameters include parameters such as the focal position of the imaging device 13, the image center, and lens characteristics. Examples of extrinsic parameters include parameters indicating, for example, the position (image capture position) and orientation (image capture orientation) of the imaging device 13. These parameters are recorded in the form of data.
[0076] In this embodiment, a configuration thereof is used. Figure 2 The image capture studio shown uses a fully encircling screen (green screen 11) and a fully encircling camera setup (multiple camera setups 13) to capture real-world volumetric motion images. For example, real-world volumetric technology is a technique used to create a real-world 3D model 6 of the subject 1 based on captured images of the subject 1. Therefore, real-world volumetric motion images are, for example, videos that enable the display of a 3D model 6 performing the same motion as the subject 1 at any scale and any viewpoint.
[0077] When capturing real-world volumetric motion images, the camera devices 13 of the multiple camera devices 13 first capture images of the subject 1 from different directions against a green screen 11. Next, the image capture-side controller 14 extracts the subject image 3 obtained by separating the background from the images captured by each camera device 13. Then, the image capture-side controller 14 uses the data from the subject image 3 to create three-dimensional data of the subject 1, thereby creating a 3D model 6 of the subject 1.
[0078] Examples of data for 3D model 6 include: shape data indicating the model's three-dimensional shape, skeletal structure data for the character model, and texture data indicating surface information. Additionally, examples of data for 3D model 6 may include data about the background of 3D model 6. These data items are sent as content data items to the viewing-side controller 23.
[0079] At the viewing side position, the viewing side controller 23 generates content video 5 (real-time video) based on content data. Content video 5 is, for example, a video obtained by rendering a 3D model 6 at any scale and any viewpoint. In this embodiment, the 3D model 6 is rendered according to the position of viewpoint P2 of viewer 2 and the orientation of viewer 2's line of sight L2. This makes it possible to display video with 6 degrees of freedom (6-DoF). The use of the above-described real-scene volumetric technique enables video representation with high degrees of freedom.
[0080] [Eye Guidance System]
[0081] like Figure 1 As shown in the lower part, the gaze guidance system 30 includes an illumination device 31 and a gaze guidance controller 32. The illumination device 31 is a device used by being arranged in a real space. Specifically, at the image capture side position, the illumination device 31 is suitably arranged in a position where it can project a light spot 33 onto a green screen 11, the light spot 33 being used to guide the gaze L1 of the subject 1 to the green screen 11. The gaze guidance controller 32 is, for example, a computer located at the image capture side position. For example, a server device located in a location different from the image capture studio can be used as the gaze guidance controller 32.
[0082] Figure 3 An example of the configuration of the illumination device 31 is shown schematically. The illumination device 31 is a device that illuminates a laser beam LB to project a light spot 33 onto a green screen 11. The illumination device 31 includes an azimuth rotation unit 34, an elevation rotation unit 35, and a laser pointer 36, and is a movable laser pointer that uses the azimuth rotation unit 34 and the elevation rotation unit 35 to control the orientation of the laser pointer 36.
[0083] The azimuth rotation unit 34 is a rotating platform disposed on the lower part of the illumination device 31, and rotates about the rotation axis C1. For example, the azimuth rotation unit 34 is arranged such that the rotation axis C1 is parallel to the orthogonal direction. The elevation rotation unit 35 is a rotating mechanism mounted on the azimuth rotation unit 34, and rotates about the rotation axis C2. For example, the elevation rotation unit 35 is arranged on the azimuth rotation unit 34 such that the rotation axis C2 is parallel to the horizontal direction. For example, a stepper motor is used to drive the azimuth rotation unit 34 and the elevation rotation unit 35.
[0084] The laser pointer 36 is a light-emitting element that illuminates the laser beam LB. When the laser beam LB, illuminated by the laser pointer 36, hits an object, a light spot 33 corresponding to the bright spot of the laser beam LB is formed at the location on the object hit by the laser beam LB. This allows the light spot 33 to be projected onto the green screen 11. The light spot 33 is a marker indicating the direction in which the subject 1's line of sight L1 is directed, and serves as line-of-sight guidance information for the subject 1. As described above, the illumination device 31 serves as a presentation device that presents the subject 1 with the direction in which the line of sight L1 is directed (line-of-sight guidance information).
[0085] A light-emitting element, such as a laser diode, is used as the laser pointer 36. Furthermore, the color and intensity of the laser beam LB are appropriately set so that the subject 1 can easily perceive the light spot 33 when the laser beam LB illuminates the green screen 11. Moreover, the specific configuration of the laser pointer 36 is not limited.
[0086] The laser pointer 36 is mounted on the elevation rotation unit 35, causing the laser pointer 36 to rotate together with the elevation rotation unit 35. For example, when the elevation rotation unit 35 is driven, the illumination direction (the illumination direction of the laser beam LB) executed by the laser pointer 36 rotates in the elevation direction. This corresponds to the angle control executed in the tilt direction to move the position of the spot 33 up or down. Furthermore, when the azimuth rotation unit 34 is driven, the laser pointer 36 rotates together with the elevation rotation unit 35. Therefore, the illumination direction executed by the laser pointer 36 rotates in the azimuth direction. This corresponds to the angle control executed in the panning direction to move the position of the spot 33 to the right or left.
[0087] As described above, the irradiation device 31 is a laser pointer 36 that is capable of rotating in the elevation direction and the azimuth direction. Therefore, for example, by controlling the rotation angle (i.e., tilt angle and panning angle) in the elevation direction and the azimuth direction, the irradiation direction performed by the laser pointer 36 can be set to the desired direction.
[0088] Note that in Figure 3 In this configuration, the rotation axis C1 of the azimuth rotation unit 34 intersects the rotation axis C2 of the elevation rotation unit 35. Therefore, the intersection point of rotation axes C1 and C2 is a point that will not move even when driving the azimuth rotation unit 34 and the elevation rotation unit 35, and this intersection point can be used, for example, as the origin when controlling the laser pointer 36. Note that the configuration of the irradiation device 31 is not limited to this, and for example, rotation axes C1 and C2 do not necessarily have to intersect. In this case, the rotation angles in the elevation and azimuth directions can be appropriately calculated, and the irradiation direction can be aligned with the desired direction according to the structure of the irradiation device 31.
[0089] Figure 4 This is a block diagram illustrating an example of the functional configuration of the gaze guidance controller 32. The gaze guidance controller 32 is a device that controls the operation of the illumination device 31 to guide the gaze L1 of the subject 1, and is configured, for example, using a personal computer (PC). The gaze guidance controller 32 includes a storage device 40 and a control device 41.
[0090] Storage device 40 is a non-volatile storage device, such as a solid-state drive (SSD) or hard disk drive (HDD). Storage device 40 stores a control program therein. The control program is, for example, a program that controls the overall operation of the gaze guidance controller 32. In this embodiment, storage device 40 corresponds to a computer-readable recording medium in which the program is recorded. Furthermore, the control program corresponds to a program recorded on the recording medium.
[0091] The control device 41 is a computing device that controls the operation of the gaze guidance controller 32. The control device 41 is configured with hardware required by a computer, such as a CPU and memory (RAM and ROM). Various processes are performed by loading a control program stored in the storage device 40 into the RAM and executing the program via the CPU. In this embodiment, the control device 41 corresponds to an information processing device.
[0092] For example, a programmable logic device (PLD) such as a field-programmable gate array (FPGA) or other devices such as an application-specific integrated circuit (ASIC) can be used as the control device 41. Furthermore, for example, a graphics processing unit (GPU) can be used as the control device 41.
[0093] In this embodiment, the CPU of the control device 41, which executes the program (control program) according to this embodiment, implements the attention focus position acquisition unit 42, the image capture processing unit 43, and the gaze guidance processing unit 44 as functional blocks. Then, the information processing method according to this embodiment is executed through these functional blocks. Note that dedicated hardware such as an integrated circuit (IC) can be appropriately used to implement each functional block.
[0094] The attention focus position acquisition unit 42 acquires the attention focus position Pin that the subject 1 is expected to look at. For example, the attention focus position Pin is the position where the subject 1's line of sight L1 is guided (i.e., the position where the subject 1's line of sight is expected to be guided). Note that the attention focus position Pin is the position where the subject 1's line of sight L1 is guided by the subject 1 looking at the light spot 33 projected by the illumination device 31, and it is a position different from the position of the projection point of the light spot 33 projected thereon.
[0095] In this embodiment, the attention focus position Pin is the viewpoint position of the viewer 2 who is viewing the display device 21 positioned on the viewing side. The viewpoint position of the viewer 2 is, for example, determined by a reference... Figure 1 The viewing-side controller 23 and the viewpoint sensor 22 described herein detect the position of viewpoint P2 of viewer 2. Hereinafter, the viewpoint position of viewer 2 may be referred to as viewpoint position P2, using the same reference numerals as viewpoint P2 of viewer 2. For example, viewpoint position P2 of viewer 2 may be detected as a coordinate position in real space set at the viewing-side position. In this case, the attention focus position acquisition unit 42 reads coordinate information related to the coordinates of viewpoint position P2 of viewer 2 via a designated network. Note that the attention focus position acquisition unit 42 may read the detection result performed by the viewpoint sensor 22 to detect viewpoint position P2 of viewer 2.
[0096] The image capture processing unit 43 processes the image captured by the image capture unit 12 (multiple camera devices 13) positioned at the image capture side location to obtain information for guiding the gaze L1 of the subject 1. Specifically, the image capture processing unit 43 obtains the viewpoint position of the subject 1 based on the image of the subject 1 captured by the multiple camera devices 13 positioned in real space such that the green screen 11 is the background within the image capture range. The viewpoint position of the subject 1 is the position of the viewpoint P1 of the subject 1. Hereinafter, the viewpoint position of the subject 1 may be referred to as viewpoint position P1, using the same reference numerals as the viewpoint P1 of the subject 1. For example, the viewpoint position P1 of the subject 1 is detected as a coordinate position in real space positioned at the image capture side location. Reference will be made later to, for example... Figure 9 This describes a method for calculating the viewpoint position P1 of subject 1.
[0097] Furthermore, the image capture processing unit 43 acquires virtual spatial information indicating the position of the green screen 11 in a virtual three-dimensional space based on images captured by the multiple camera devices 13. Here, the virtual three-dimensional space is a three-dimensional space created, for example, by a computer. For example, the virtual spatial information is data obtained by mapping the position of the green screen 11 and reproducing the green screen 11 in the virtual three-dimensional space. The form of the virtual spatial information is not limited, and the virtual spatial information can be, for example, point group data representing individual points on the green screen 11 or polygon data constituting the screen surface. Reference will be made later to examples... Figure 7 and Figure 8 This describes the methods used to generate virtual space information.
[0098] As described above, in the gaze guidance controller 32, the attention focus position acquisition unit 42 acquires the attention focus position Pin (here, the viewpoint position P2 of the viewer 2) that the subject 1 is expected to look at, and the image capture processing unit 43 acquires the viewpoint position P1 of the subject 1 and virtual space information. In this embodiment, the acquisition unit is implemented by the attention focus position acquisition unit 42 and the image capture processing unit 43. This acquisition unit acquires the viewpoint position of the subject that captures its image using a screen set in real space for chroma key synthesis, the attention focus position that the subject is expected to look at, and virtual space information indicating the position of the screen in virtual three-dimensional space.
[0099] The gaze guidance processing unit 44 uses the attention focus position Pin (viewer 2's viewpoint position P2) acquired by the attention focus position acquisition unit 42 and the subject 1's viewpoint position P1 acquired by the image capture processing unit 43 to calculate the projection point Pproj of the light spot 33 projected by the illumination device 31. Here, the projection point Pproj is the guide point to which the subject 1's gaze L1 is actually guided at the image capture side position. In the following text, the position of the projection point Pproj can be referred to as the projection position Pproj. The virtual space information (data obtained by reproducing the green screen 11 in a virtual three-dimensional space) acquired by the image capture processing unit 43 is used in the processing to calculate the projection point Pproj.
[0100] Figure 5 This is a schematic diagram used to describe the basic operation of the line-of-sight guidance system 30. Figure 5 A schematically illustrates a virtual three-dimensional space 18, and Figure 5 Figure B schematically illustrates the real space 17 corresponding to the image capture side position. Hereinafter, the virtual three-dimensional space 18 will be simply referred to as virtual space 18.
[0101] The position of real space 17 is represented by three-dimensional coordinates set for the real world. The axes of the Cartesian coordinates representing the three-dimensional position of real space 17 are called the X-axis, Y-axis, and Z-axis, and the origin of the XYZ coordinate system is denoted by O. The X-axis and Y-axis are axes that are, for example, orthogonal to each other in the horizontal plane, while the Z-axis is an axis that represents an orthogonal direction to the horizontal plane.
[0102] Furthermore, the position of virtual space 18 is represented by a virtual three-dimensional coordinate system. The axes of the Cartesian coordinate system representing the three-dimensional position of virtual space 18 are called the X' axis, Y' axis, and Z' axis, and the origin of the X'Y'Z' coordinate system is denoted by O'. The X' and Y' axes are, for example, axes orthogonal to each other in the horizontal plane of virtual space 18, while the Z' axis is an axis representing an orthogonal direction orthogonal to the horizontal plane of virtual space 18.
[0103] A set of coordinates represented using the XYZ coordinate system of real space 17 and a set of coordinates represented using the X'Y'Z' coordinate system of virtual space 18 correspond one-to-one with each other and can be transformed into one another. For example, Figure 5 The coordinates of the viewpoint position P1 of the subject 1 in B are represented using the XYZ coordinate system set for the real space 17 at the image capture side position. Figure 5 The viewpoint position P1' in A is obtained by transforming the viewpoint position P1 in real space 17 to virtual space 18. The coordinates of the viewpoint position P1' are represented using the X'Y'Z' coordinate system set for virtual space 18.
[0104] On the other hand, the viewpoint position P2 of viewer 2 is represented using an XYZ coordinate system set for the real space 17 at the viewing side position. The XYZ coordinate system at the viewing side position is independent of the XYZ coordinate system at the image capture side position. Figure 5 The coordinate system shown in B is used. However, the viewpoint position P2 can also be uniquely transformed to its position in virtual space 18. Figure 5 The viewpoint position P2' in A is obtained by transforming the viewpoint position P2 of viewer 2 in real space 17 to virtual space 18. Similar to the viewpoint position P1', the coordinates of viewpoint position P2' are represented using an X'Y'Z' coordinate system set for virtual space 18. In this embodiment, viewpoint position P2' corresponds to the attention focus position Pin' in virtual space 18.
[0105] In addition, a virtual green screen 11 (hereinafter referred to as virtual screen 19) is arranged in the virtual space 18 as information for generating the virtual space. The virtual screen 19 is obtained by reproducing the green screen 11 in the real space 17 in the virtual space 18, and the position of each point on the virtual screen 19 is represented using the X'Y'Z' coordinate system.
[0106] As described above, in the virtual space 18, the viewpoint position P1 of the subject 1 at the image capture side position and the position of the green screen 11 are reproduced as viewpoint position P1' and the position of the virtual screen 19. Furthermore, in the virtual space 18, the viewpoint position P2 of the viewer 2 at the viewing side position is reproduced as viewpoint position P2'. The use of the virtual space 18 allows the subject 1, the viewer 2, and the green screen 11 to be virtually arranged in the same space.
[0107] Note that the method used to set the X'Y'Z' coordinate system is unrestricted. For example, it is possible to set the X'Y'Z' coordinate system of virtual space 18 to be consistent with the X'Y'Z coordinate system of real space 17 at the image capture side position. In this case, the origin O' is set to have coordinate values equal to those of the origin O, and the X', Y', and Z' axes are set in the same directions as the X, Y, and Z axes, respectively. This allows the coordinates in virtual space 18 to be used as coordinates in real space 17 without change. Furthermore, for example, the space in which the 3D model 6 of the subject 1 is rendered can be used as virtual space 18, or virtual space 18 can be set as a space independent of the space in which the 3D model 6 is rendered. Furthermore, the coordinate system of virtual space 18 can be arbitrarily set.
[0108] In this embodiment, the gaze guidance processing unit 44 determines the virtual intersection point Pc' between the virtual line L' and the green screen 11 (virtual screen 19) indicated by the virtual space information. The virtual line L' connects the viewpoint position P1' of the subject 1 and the attention focus position Pin' (viewpoint position P2' of the viewer 2) in the virtual space 18. This is the process of calculating the position coordinates of the virtual intersection point Pc' in the virtual space 18.
[0109] like Figure 5 As shown in A, the virtual line L' is a straight line passing through P1' and P2'. In this case, the virtual line L' is the line through which the subject 1 and the viewer 2 make eye contact with each other in the virtual space 18, and can be a virtual eye contact line. The virtual intersection point Pc' is a point in such a virtual line L'. Furthermore, the virtual intersection point Pc' is a point on the virtual screen 19. Therefore, in the real space 17, the point corresponding to the virtual intersection point Pc' (the point obtained by transforming the virtual intersection point Pc' into the real space 17) is a point on the green screen 11, and is a point on which the light spot 33 can be projected.
[0110] For example, suppose that in real space 17, subject 1 is looking at the point corresponding to the virtual intersection point Pc'. This state is the state in the content video 5 displayed at the viewing side position where subject 1 is looking at the viewpoint P2 of viewer 2. In other words, when subject 1 is looking at the virtual intersection point Pc' which is transformed into a point in real space 17, it makes it possible for the line of sight of subject 1 to be aligned with the line of sight of viewer 2 at the viewing side position.
[0111] Therefore, in real space 17, the point corresponding to the virtual intersection point Pc' is set as the guiding point to which the line of sight L1 of the subject 1 is guided, that is, the projection point Pproj on which the light spot 33 is projected by the illumination device 31. The position of the projection point Pproj is calculated by transforming the position coordinates of the virtual intersection point Pc' to the position coordinates in real space 17. As described above, the projection point Pproj can be a point obtained by performing a coordinate transformation on the attention focus position Pin (viewer 2's viewpoint position P2) that the subject 1 is usually expected to look at, the coordinate transformation being performed according to the shape of a portion of the green screen 11 surrounding the subject 1.
[0112] Furthermore, the gaze guidance processing unit 44 controls the illumination direction of the illumination performed by the illumination device 31 installed in the real space 17, directing it towards the projection point Pproj (guide point) located on the green screen 11 in the real space 17 and corresponding to the virtual intersection point Pc'. Therefore, the illumination device 31 irradiates a laser beam LB towards the projection point Pproj, and the light spot 33 is projected onto the projection point Pproj. For example, when the subject 1 moves his / her gaze L1 to look at the light spot 33, this allows the display of a content video 5 at the viewing side position that aligns the gaze of the subject 1 with that of the viewer 2. In this embodiment, the gaze guidance processing unit 44 serves as both a determination unit and a control unit.
[0113] As described above, the green screen 11 is a fully encircling screen. Therefore, for example, when capturing an image of the subject 1 from the front, the screen exists not only in the area corresponding to the background but also in the area in front of the subject 1 when viewed from the front. The area in front of the subject 1 when viewed from the front is, for example, the area that the subject 1 can see when looking forward.
[0114] In this embodiment, the light spot 33 is projected onto the portion of the screen located in front of the subject 1. In other words, the gaze guidance processing unit 44 controls the illumination direction performed by the illumination device 31, directing it towards the projection point Pproj (guide point) in an area on the green screen 11 that serves as a fully encircling screen, located in front of the subject 1 when viewed from above. This allows the subject 1 to easily locate the projection point Pproj. This enables the subject 1's line of sight to naturally align with the viewer 2's line of sight without causing the subject 1 to move unnaturally.
[0115] For example, in the real space 17 (image capture space 16 surrounded by green screen 11) at the image capture side position, a front direction and a rear direction corresponding to the direction in the real space 17 at the viewing side position are pre-set. In this case, the front direction in the image capture space 16 corresponds to, for example, the front direction when viewed from the display device 21 at the viewing side position. In such a positional relationship, the subject 1 performs with the viewer 2 located in the front direction of the image capture space 16 and, in most cases, facing the front direction of the image capture space 16. Furthermore, when viewed from the subject 1, the viewpoint position P2' of the viewer 2 in the virtual space 18 is arranged in front. Therefore, the projection point Pproj is a point in the following area on the green screen 11 that is in front when viewed from the subject 1, and the light spot 33 is projected onto the area in front when viewed from the subject 1.
[0116] Note that viewer 2 can move freely from the viewing side position. Therefore, it is conceivable that, depending on the positional relationship between subject 1 and viewer 2, viewer 2's viewpoint position P2' in virtual space 18 can be located behind when viewed from subject 1. In this case, the point in image capture space 16 corresponding to the virtual intersection Pc' is a point in the following area on green screen 11, located behind when viewed from subject 1. In this case, projection point Pproj can be controlled so that the subject 1's line of sight is guided to the point corresponding to the virtual intersection Pc'. For example, independently of the virtual intersection Pc', projection point Pproj can be set in the area in front when viewed from subject 1, and projection point Pproj can be moved so that it eventually coincides with the point corresponding to the virtual intersection Pc'. This allows the subject 1's line of sight to be naturally guided, aligning with the viewer 2's line of sight.
[0117] In the example above, the positional relationship in real space 17 between the image capture side position and the viewing side position is fixed in advance. However, it is not limited to this; the positional relationship in real space 17 can be adjusted in virtual space 18 such that the projection point Pproj corresponding to the virtual intersection point Pc' is a point in the region in front of the subject 1 when viewed. In this case, a process is performed that rotates or translates at least one of the coordinate systems at the image capture side position or the viewing side position, such that, for example, in virtual space 18, the viewpoint position P2' of the viewer 2 is positioned in front of the subject 1 when viewed. This allows the subject 1 to perform with a high degree of freedom.
[0118] As described above, in the gaze guidance system 30, in order to guide the gaze of the subject 1 when performing image capture with a full-circle green background using a green screen 11, a coordinate transformation is performed on the viewpoint position P2 of the viewer 2, which is the attention focus position Pin, to obtain a projection point Pproj. This coordinate transformation is performed in the virtual space 18 using the viewpoint position P2 of the viewer 2, the viewpoint position P1 of the subject 1, and the shape of the green screen 11 (virtual space information). When the illumination device 31 projects the light spot 33 onto the projection point Pproj calculated as described above, this allows the gaze L1 of the subject 1 to be guided, for example, when capturing a real-world volumetric motion image, so that the subject 1 gazes at any three-dimensional coordinate.
[0119] Furthermore, for example, the bright spot 33 of the laser beam LB has a diameter of several centimeters or less, or typically one centimeter or less. Therefore, even when the bright spot 33 appears in the image of the subject 1, the performance of extracting the subject image 3 is hardly diminished. This makes it possible to guide the line of sight of the subject 1 without reducing the performance of extracting the subject image 3.
[0120] [Operation of the Eye Guidance System]
[0121] Figure 6 This is a flowchart illustrating an example of the operation of the line-of-sight guidance system 30. Figure 6 The process shown is initiated together with the start of the gaze guidance system 30. First, a three-dimensional spatial measurement process (step 101) is performed at the image capture side position to measure the image capture space 16. The three-dimensional spatial measurement is, for example, a process performed before capturing an image of the subject 1. Here, the image capture processing unit 43 generates virtual spatial information indicating the position of the green screen 11 in the virtual space 18 based on images captured by the multiple camera devices 13. Note that the virtual spatial information is appropriately read when the three-dimensional spatial measurement process is performed in advance and when the virtual spatial information related to the green screen 11 is recorded in, for example, the storage device 40.
[0122] Next, subject viewpoint measurement processing (step 102) is performed at the image capture side position to measure the viewpoint position P1 of the subject 1. Here, the capture image processing unit 43 calculates the position coordinates of the viewpoint position P1 of the subject 1 in the real space 17 based on the images captured by the multiple camera devices 13.
[0123] Next, attention focus position acquisition processing (step 103) is performed to obtain the position coordinates of the attention focus position Pin. In this embodiment, the viewpoint position P2 of the viewer 2 measured at the viewing side position is obtained as the attention focus position Pin. Here, the attention focus position acquisition unit 42 reads the position coordinates of the viewpoint position P2 of the viewer 2 in the real space 17 from the viewing side controller 23.
[0124] Next, a coordinate transformation process (step 104) is performed to transform the attention focus position Pin (viewer 2's viewpoint position P2) into a projection point Pproj. Here, the gaze guidance processing unit 44 transforms the viewpoint position P2 of viewer 2 into a projection point Pproj based on the viewpoint position P1 of subject 1 and virtual space information about the green screen 11. The viewpoint position P2 of viewer 2 is measured at the viewing side position, and the projection point Pproj is a point in the real space 17 at the image capture side position.
[0125] Next, the gaze guidance information presentation process is performed, which includes controlling the illumination device 31 to project a light spot 33 corresponding to the gaze guidance information onto the projection point Pproj (step 105). This process is to present a marker (light spot 33) (to which the gaze L1 is directed) to the subject 1 (whose image is being captured). Here, the gaze guidance processing unit 44 calculates the illumination direction in which the projection point Pproj can be illuminated using a laser beam LB. The illumination device 31 is controlled to be oriented toward the illumination direction, and the light spot 33 is projected onto the projection point Pproj.
[0126] When the projected light spot 33 is reached, it is determined whether to terminate image capture of subject 1 (step 106). For example, if image capture of subject 1 continues (No in step 106), step 102 and subsequent processing are executed again. Furthermore, if image capture is terminated due to, for example, the end of subject 1's performance (Yes in step 106), the operation of the gaze guidance system 30 is terminated. The following detailed description refers to... Figure 6 Details of the operations performed in each step of the described steps.
[0127] [3D Spatial Measurement Processing]
[0128] Here, for Figure 6 The specific method of the three-dimensional spatial measurement processing performed in step 101 will be described. As described above, in order to properly present the light spot 33 (line-of-sight guidance information) according to the viewpoint position P1 of the subject 1, it is necessary to obtain three-dimensional position information about the green screen 11. In this embodiment, three-dimensional spatial measurement processing using multiple camera devices 13 and illumination device 31 is performed. The multiple camera devices 13 are also used to capture real-world volumetric motion images and are mounted to surround the image capture space 16.
[0129] Figure 7 This is a flowchart illustrating an example of three-dimensional spatial measurement processing. Figure 8This is a schematic diagram illustrating an example of three-dimensional spatial measurement processing. In the three-dimensional spatial measurement processing, the positions of multiple measurement points 37 on the green screen 11 are measured as position coordinates in the real space 17. Virtual space information indicates the positions of the multiple measurement points 37 in the virtual space 18. In the following description, it is assumed that the internal and external parameters of multiple camera devices 13 are pre-calibrated, and the values of each parameter can be referenced in the form of data.
[0130] First, the illumination of the laser beam LB performed by the illumination device 31 is turned off (step 201), and images are captured by all of the multiple imaging devices 13 (step 202). Here, the captured images are used as background images in which the light spot 33 does not appear. Next, the illumination of the laser beam LB performed by the illumination device 31 is turned on (step 203), and the illumination device 31 is controlled to move the light spot 33 (step 204).
[0131] Figure 8 Multiple measurement points 37 are schematically shown on the green screen 11 to measure their positions. Here, the position of the illumination device 31 in the real space 17 is referred to as position L. The measurement point 37 is the point on which the light spot 33 is projected by the illumination device 31. For example, the operation of micro-driving the illumination device 31 in the elevation and azimuth directions is repeatedly performed to form multiple measurement points 37. In the three-dimensional space measurement process, the light spot 33 is sequentially projected onto the measurement point 37 among the multiple measurement points 37, and the position coordinates in the real space 17 are calculated for each measurement point 37.
[0132] As the light spot 33 moves, a light spot image of the light spot 33 is captured by multiple camera devices 13 (step 205). For example, the light spot image is an image captured by a camera device 13 whose viewing angle includes the projection position of the light spot 33. When the light spot image is captured, the light spot position is extracted (step 206). Here, the light spot position in the image is calculated based on the difference between the light spot image and the background image captured by the camera device 13 that captured the light spot image.
[0133] Next, it is determined whether one of the multiple camera devices 13 has acquired the position of the light spot (step 207). At least two camera devices 13 are required to acquire the position of the light spot. Here, it is determined whether the number of camera devices 13 that have acquired the position of the light spot successfully exceeds a threshold of 2 or more. If the number of camera devices 13 does not exceed the threshold (No in step 207), the process returns to step 204, and the light spot 33 is moved to the next measurement point 37. Alternatively, if the number of camera devices 13 exceeds the threshold (Yes in step 207), the three-dimensional position of the light spot 33 is calculated (step 208).
[0134] Using, for example, the position of the light spot 33 in a planar position in the image acquired by one of the multiple camera devices 13, and the internal and external parameters of each of the camera devices 13, the three-dimensional position of the light spot 33 is calculated by triangulation. The three-dimensional position of the light spot 33 calculated here is stored as the three-dimensional coordinates of a measurement point 37.
[0135] When the three-dimensional position of the light spot 33 is calculated, it is determined whether sufficient measurement points 37 have been acquired (step 209). For example, a threshold is used to determine whether measurement points 37 have been acquired at a sufficient density or whether a sufficient number of measurement points 37 have been acquired. If it is determined that sufficient measurement points 37 have not been acquired (No in step 209), the process returns to step 204, and the illumination device 31 is driven to move the light spot 33 a very small distance. Then, the process of measuring the position of the next measurement point 37 continues. This allows the shape of the surface of the green screen 11 to be measured. Note that if it is determined that sufficient measurement points 37 have been acquired (Yes in step 209), the three-dimensional spatial measurement process terminates.
[0136] In this embodiment, as described above, images of the light spot 33 projected onto the green screen 11 by the illumination device 31 are captured by multiple camera devices 13, and the three-dimensional coordinates of the light spot 33 are calculated based on the images captured by the multiple camera devices 13. This allows data representing the three-dimensional position of the green screen 11 to be acquired without using, for example, a dedicated sensor for measuring the shape of the green screen 11. Compared to, for example, using a dedicated sensor, this allows virtual spatial information to be provided at a lower cost.
[0137] [Subject viewpoint measurement processing]
[0138] Here, for Figure 6 The specific method of subject viewpoint measurement processing performed in step 102 will be described. In order to control the illumination device 31 to properly present the light spot 33 (line guidance information) to the subject 1, it is necessary to properly obtain the viewpoint position P1 of the subject 1. In this embodiment, as in the case of three-dimensional spatial measurement processing, multiple camera devices 13 and illumination devices 31 are used to perform subject viewpoint measurement processing.
[0139] Figure 9This is a flowchart illustrating an example of subject viewpoint measurement processing. In subject viewpoint measurement processing, the viewpoint position P1 of subject 1 is measured as position coordinates in real space 17. First, images of subject 1 are captured by multiple camera devices 13 (step 301). Next, the viewpoint position of subject 1 is estimated in the images captured by the multiple camera devices 13, in which subject 1 appears (step 302). Here, the estimated viewpoint position is the planar position of the viewpoint P1 of subject 1 in the image. For example, methods used to estimate human pose or detect human faces can be used as methods for estimating viewpoint positions in images.
[0140] Next, it is determined whether one of the multiple camera devices 13 has acquired the viewpoint position of the subject 1 in the image (step 303). At least two camera devices 13 are required to acquire the viewpoint position in the image. Here, it is determined whether the number of camera devices 13 that have acquired the viewpoint position in the image successfully exceeds a threshold of 2 or greater. When the number of camera devices 13 does not exceed the threshold (No in step 303), the process returns to step 301, and image capture is performed again by the multiple camera devices 13. Furthermore, if the number of camera devices 13 exceeds the threshold (Yes in step 303), the three-dimensional position (viewpoint position P1) of the viewpoint P1 of the subject 1 is calculated (step 304).
[0141] like Figure 7 As in step 208, the viewpoint position P1 of the subject 1 is calculated. In other words, the viewpoint position P1 is calculated by triangulation using, for example, the viewpoint positions in images acquired by the multiple imaging devices 13, and the internal and external parameters of each of the imaging devices 13. When the viewpoint position P1 of the subject 1 is calculated, the subject viewpoint measurement process ends.
[0142] [Coordinate Transformation Processing]
[0143] Figure 10 This is a schematic diagram illustrating an example of coordinate transformation processing. Here, for... Figure 6 The specific method of coordinate transformation processing performed in step 104 will be described. Figure 10 The arrangement of the viewpoint position P1' of the subject 1, the focus position Pin' (viewpoint position P2' of the viewer 2), and the virtual screen 19 in the virtual space 18 is illustrated schematically. Here, the surface forming the virtual screen 19 is referred to as the forming surface S'.
[0144] The viewpoint position P1' of subject 1 is obtained by transforming the coordinates of the viewpoint position P1 of subject 1 in real space 17 to the coordinates in virtual space 18. The attention focus position Pin' (viewer 2's viewpoint position P2) is obtained by transforming the coordinates of the attention focus position Pin (viewer 2's viewpoint position P2) in real space 17 to the coordinates in virtual space 18. The virtual screen 19 corresponds to the data (virtual space information) obtained by transforming the shape data of the green screen 11 obtained by three-dimensional spatial measurement processing to coordinates in virtual space 18.
[0145] Here, the virtual intersection point Pc' of the virtual line L' and the virtual screen 19 is calculated. The virtual line L' connects the viewpoint position P1' of the subject 1 and the attention focus position Pin'. The formula that the virtual intersection point Pc' satisfies is expressed as indicated below.
[0146] Pc'=(1-t)P1'+tPin' ...(1)
[0147] Here, the virtual intersection point Pc' is a point on the forming surface S' of the virtual screen 19. Furthermore, t is a positive real number (t≥0). For example, the right side of equation (1) represents the position of the end of the vector (line of sight) extending parallel to the virtual line L' and originating at the viewpoint position P1' of the subject 1. Furthermore, t is a parameter representing the length of the line of sight vector (distance from the viewpoint position P1').
[0148] The coordinate transformation process includes, for example, obtaining t that satisfies formula (1) and calculating the coordinates of the virtual intersection point Pc'. For example, t can be changed to determine the coordinates (Pc') that satisfy formula (1). Furthermore, for example, if the forming surface S' of the virtual screen 19 is a parametrically represented surface, t that satisfies formula (1) can be calculated analytically. Moreover, the method used to calculate the virtual intersection point Pc' is not limited, and for example, a method for calculating the intersection point based on a data format, for example, indicating virtual spatial information of the forming surface S', can be appropriately used.
[0149] The virtual intersection point Pc' is transformed into the real space 17 to calculate the projection point Pproj. Specifically, the position coordinates of the virtual intersection point Pc' in the virtual space 18 (position coordinates in the X'Y'Z' coordinate system) are transformed into the position coordinates in the real space 17 (position coordinates in the XYZ coordinate system). The position coordinates obtained through the transformation are the position coordinates of the projection point Pproj.
[0150] [Gaze-guided information presentation processing]
[0151] Figure 11 This is a schematic diagram illustrating an example of the pointing operation of the irradiation device. Here, for... Figure 6The specific method for the gaze guidance information presentation processing performed in step 105 is described below. Figure 11 The arrangement of the illumination device 31 at the image capture side position in the real space 17 at position L and the projection point Pproj on the green screen 11 is schematically shown.
[0152] As described above, the gaze guidance information presentation processing is the process of controlling the irradiation device 31 to project the light spot 33 onto the projection point Pproj. In this embodiment, the irradiation direction performed by the irradiation device 31 is calculated based on the position of the projection point Pproj (guidance point) in the real space 17 and the position L of the irradiation device 31. In other words, the irradiation direction is calculated in which the light spot 33 can be projected from the position L of the irradiation device 31 onto the projection point Pproj. Furthermore, the line connecting the position L of the irradiation device 31 and the projection point Pproj corresponds to the irradiation path 38 of the laser beam LB.
[0153] Specifically, the rotation angles (pan) of the irradiation device 31 in the azimuth direction and the rotation angles (tilt) of the irradiation device 31 in the elevation direction are calculated respectively, serving as parameters for controlling the irradiation direction performed by the irradiation device 31. Furthermore, these rotation angles are relative to a reference irradiation direction Dir, which serves as the reference for the irradiation direction. Therefore, the reference irradiation direction Dir can be the initial direction when angle pan = 0 and angle tilt = 0. It is assumed that the position L of the irradiation device 31 and the orientation of the reference irradiation direction Dir are known in advance. The rotation angles (pan) in the azimuth direction and the rotation angles (tilt) in the elevation direction are calculated according to the formulas shown below.
[0154] L+t·Rpan·Rtilt·Dir=Pproj... (2)
[0155] Here, Rpan is a rotation matrix used to rotate any three-dimensional vector through the angle pan in the azimuth direction. Furthermore, Rtilt is a rotation matrix used to rotate any three-dimensional vector through the angle tilt in the elevation direction. t is a positive real number (t≥0). Equation (2) is a simultaneous equation used to calculate the angle pan, angle tilt, and t such that the end of the vector extending from position L of the irradiation device 31 is aligned with Proj.
[0156] In this embodiment, as described above, the rotation angles (pan and tilt) of the irradiation direction relative to the reference irradiation direction Dir are calculated, which is used as a reference for the irradiation direction in the irradiation device 31. The calculated rotation angles pan and tilt are sent to the irradiation device 31 as control angles. In the irradiation device 31, the azimuth rotation unit 34 is driven so that the angular position coincides with the position of the rotation angle pan. Similarly, the elevation rotation unit 35 is driven so that the angular position coincides with the position of the rotation angle tilt. Therefore, the irradiation direction of the laser beam LB is oriented toward the projection point Pproj, which enables the light spot 33 to be projected onto the projection point Pproj.
[0157] Figure 12 This is a schematic diagram illustrating another example of the pointing operation of the illumination device 31. Here, the following control example is described: controlling the illumination of the illumination device 31 (laser pointer 36) so that the light spot 33 is not displayed on a target on which it is not desired to project a light spot 33.
[0158] For example, when the laser beam LB directly illuminates an image capture device such as the camera device 13, a spot 33 of the laser beam LB appears in the captured image. Furthermore, although the intensity of the laser beam LB is sufficiently reduced, it is dangerous to illuminate the spot 33 onto a person, including the subject 1 in the image capture space 16. As described above, there is a target in the image capture space 16 that is intended to prevent the laser beam LB from illuminating.
[0159] In such cases, reference can be used. Figure 13 The described laser beam LB's illumination path 38 (illumination direction) is used to take preventative measures against illumination. Specifically, the line-of-sight guidance processing unit 44 determines whether the illumination path 38, which connects the projection point Pproj (guide point) in the real space 17 to the position L of the illumination device 31, passes through a non-illumination area 46, in which it avoids pointing towards the projection point Pproj. Furthermore, if it is determined that the illumination path 38 passes through the non-illumination area 46, it stops pointing towards the projection point Pproj.
[0160] Here, the non-illuminated area 46 (where the projection point Pproj is to be avoided) is the area where a target (non-illuminated target 47) is to be prevented from being illuminated by the laser beam LB. Examples of targets include image capturing devices such as the image capturing device 13 and people in the image capturing space 16. Alternatively, an area located at a certain distance from the non-illuminated target 47 can be designated as the non-illuminated area 46.
[0161] exist Figure 12In the example shown, a rectangular object is schematically illustrated as an example of an unlit target 47. For example, suppose a stationary object is positioned at position T in real space 17, where the unlit target 47 is a stationary object. Furthermore, the region located at a specified distance from position T of the unlit target 47 (the shaded area in the figure) is designated as the unlit region 46. Note that the unlit region 46 is typically defined as the entire area where the unlit target 47 is located.
[0162] When the projection point Pproj is calculated, it is determined whether the illumination path 38 passes through the non-illuminated area 46. For example, the distance from the position T of the non-illuminated target 47 to the illumination path 38 is calculated using the position of the projection point Pproj and the position L of the illumination device 31. When this distance is less than the specified distance used to define the non-illuminated area 46, it is determined that the illumination path 38 passes through the non-illuminated area 46. In this case, since there is a possibility that the laser beam LB will illuminate the non-illuminated target 47, the laser pointer 36 is turned off and pointing to the projection point Pproj is stopped.
[0163] As described above, when the distance between the non-irradiated target 47 and the irradiation path 38 is the distance at which the non-irradiated target 47 may be irradiated by the laser beam LB, the laser pointer 36 can be temporarily turned off to prevent the laser beam LB from irradiating the non-irradiated target 47.
[0164] Note that it has already been used. Figure 12 The example shown depicts a case where the non-illuminated target 47 is a stationary object in the image capture space 16. However, similar measures can be taken for subjects 1, for example, whose positions in the image capture space 16 are dynamically changed. Note that, for example, the position of subject 1 can be detected by triangulation using images captured by multiple camera devices 13. Furthermore, the method used to detect the position of subject 1 is not limited.
[0165] [Calibration of the irradiation device]
[0166] Figure 13 This is a schematic diagram illustrating a method for calibrating the irradiation device 31. The case where the position L of the irradiation device 31 and the reference irradiation direction Dir are pre-calibrated and recorded as data has already been described above. For example, in the case of newly installing or moving the irradiation device 31, a new calibration is performed on the position L of the irradiation device 31 and the reference irradiation direction Dir. Here, a method for calibrating the position L of the irradiation device 31 and the reference irradiation direction Dir is described.
[0167] Reference Figure 13This describes the operation of the illumination device 31 and the method performed by the physical positions of the multiple calibration camera devices 13. This process is performed using coordinates in real space 17 at the image capture side position. Specifically, the position L of the illumination device 31 and the reference illumination direction Dir are calibrated based on the rotation angles (pan and tilt) of the illumination direction when the illumination device 31 projects the light spot 33 onto each of the multiple camera devices 13, and based on the mounting positions of the multiple camera devices 13.
[0168] In the following text, the multiple camera devices 13 arranged around the image capture space 16 will be referred to as cams. i i is an index set for each camera device 13. Furthermore, the mounting location of each camera device 13 is referred to as Pcam. i Pcam is obtained by calibrating the position (external parameter) of the camera device 13 in the coordinate system of real space 17. i .
[0169] For example, suppose the rotation angle (pan and tilt) of the illumination device 31 is controlled to illuminate the i-th camera device 13 (cam) with light spot 33. i In this case, the position L of the illumination device 31, the reference illumination direction Dir, and the i-th camera device 13 (cam) are represented by the following formula. i Pcam installation location i The relationship between them.
[0170] L+t i ·Rpan i ·Rtilt i Dir=Pcam i ... (3)
[0171] Here, t i It is a positive real number (t) i ≥0). Rpan i It is the rotation angle pan in the azimuth direction when the light spot 33 is being projected onto the i-th camera device 13. i The rotation matrix is represented by Rtilt. i The rotation angle tilt in the elevation direction when the light spot 33 is being projected onto the i-th camera device 13 is determined by the rotation angle tilt. i The rotation matrix is represented.
[0172] Therefore, in order to calculate the position L of the irradiation device 31 and the reference irradiation direction Dir, a minimization problem expressed using the formula shown below is solved.
[0173] min{Σ|Pcam i -(L+ti ·Rpan i ·Rtilt i ·Dir)|} ... (4)
[0174] To obtain the installation position of each camera device 13, Pcam i External parameters are sufficient. Furthermore, Rpan is recorded each time illumination is applied to the camera device 13. i and Rtilt i Therefore, when the laser beam LB is irradiated onto multiple camera devices 13 and the rotation angle data obtained during irradiation is used, it becomes possible to solve the minimization problem of equation (4). This makes it possible to estimate the position L of the irradiation device 31 and the reference irradiation direction Dir.
[0175] The arrangement of the multiple camera devices 13 used to perform image capture on the subject 1 can be used in this method without change. This makes it possible to calibrate the illumination device 31 at low cost. Note that reference has been made to... Figure 13 The use of camera device 13 (cam) is described i Pcam installation location i The method of execution. However, for example, when a point whose position coordinates are known is located on the surface of the green screen 11 or in the image capture space 16, the position L of the illumination device 31 and the reference illumination direction Dir can also be estimated by performing illumination to that point.
[0176] Figure 14 This is a schematic diagram illustrating another method for calibrating an irradiation device. Here, a method is described for calibrating the position L of the irradiation device 31 and the reference irradiation direction Dir using a position detection mark 39. In this method, as... Figure 14 As shown, a position detection mark 39 is set on the irradiation device 31.
[0177] The position detection mark 39 is a mark used to detect the position and orientation of a target. Here, a planar mark is used, which uses a pattern to represent the position in a plane. An AR mark, such as an ArUco mark, can be used as such a mark. The position detection mark 39 is disposed, for example, in a part that is driven together with the laser pointer 36 of the illumination device 31 (i.e., the part whose position or orientation is changed together with the illumination direction).
[0178] When calibrating the position L and reference illumination direction Dir of the illumination device 31, the position L and reference illumination direction Dir of the illumination device 31 are calculated based on images captured by multiple camera devices 13 of position detection marks 39 set on the illumination device 31. For example, the position detection marks 39 set on the illumination device 31 are detected in the images captured by multiple camera devices 13, and the position and orientation of the position detection marks 39 are estimated. Based on the estimation results, the position L and reference illumination direction Dir of the illumination device 31 are estimated. As described above, using the position detection marks 39 makes it possible, for example, to easily calibrate the position L and reference illumination direction Dir of the illumination device 31 without operating the illumination device 31.
[0179] As described above, in the control device 41 according to this embodiment, a virtual intersection point Pc' is determined. At this virtual intersection point Pc', a virtual line L' connecting the viewpoint position P1' of the subject 1 and the attention focus position Pin' that the subject 1 is expected to gaze at intersects with the green screen 11 (virtual screen 19) used for chroma key synthesis in the virtual space 18. Then, the illumination device 31 points to a projection point Pproj (guide point) located on the actual green screen 11 and corresponding to the virtual intersection point Pc'. The direction of the projection point Pproj enables the subject 1 to be presented with the direction that the subject 1 is expected to gaze at. This allows the gaze of the subject 1 to be guided without degrading the performance of extracting the subject image 3.
[0180] The method for performing image capture using a green background is used as a method for extracting the foreground from the captured image. For example, when performing image capture around a subject 1 (as in the case of volumetric image capture), there is a need for an image capture environment with a full-circle green background around the subject 1.
[0181] For example, when it is desired to guide the gaze L1 of subject 1 in a specific direction, the features of an image capture environment such as a full-circle green background for guiding the gaze L1 of subject 1 are poor, and there is a possibility that subject 1 will not know where to focus its attention during image capture. On the other hand, for example, if a display is introduced to guide the gaze L1 of subject 1, the accuracy of extracting subject image 3 may be reduced.
[0182] In this embodiment, the light spot 33 is projected onto the green screen 11 using an illumination device 31 (movable laser pointer) that illuminates the laser beam LB. Therefore, information for guiding the subject 1's gaze can be presented with minimal area. This allows for sufficient maintenance of the performance in extracting the subject image 3.
[0183] Furthermore, in this embodiment, a coordinate transformation is performed to transform the desired focus point Pin (viewer 2's viewpoint P2) to be presented to subject 1 into a projection point Pproj via virtual space 18. The viewpoint P1 of subject 1 and virtual space information indicating the shape of the green screen 11 are used for this coordinate transformation. This allows for the calculation with high accuracy of the projection point Pproj onto which the light spot 33 will be projected in real space 17, and thus appropriately and accurately guides the subject 1's line of sight L1.
[0184] Furthermore, in this embodiment, the gaze guidance system 30 is used for real-time streaming of real-world volumetric motion images. For example, it is difficult for the subject 1, whose image is captured using a green screen 11, to see what the viewer 2 is doing. On the other hand, the gaze guidance system 30 makes it easy to show the subject 1 where the viewer 2's viewpoint P2 is located. Therefore, the content video 5 viewed by the viewer 2 is a video in which the subject 1 is looking at the viewer 2. This allows the viewer 2 to feel as if the viewer 2 is making eye contact with the subject 1. This enables interactive communication with a sense of realism.
[0185] <Other Implementation Methods>
[0186] This technology is not limited to the above-described embodiments, and various other embodiments can be implemented.
[0187] In the above embodiment, a movable laser pointer that presents a single light spot 33 is used as the illumination device 31. The configuration of the illumination device 31 is not limited, and for example, a device that can project a pattern other than the single light spot 33 as line-of-sight guidance information can be used as the illumination device 31.
[0188] Figure 15 Another example of the configuration of the irradiation device is shown schematically. Figure 15 The illumination device 51 shown is a laser beam scanning projector that uses a movable mirror 53 to scan the laser beam LB. The illumination device 51 includes a light source 52 that emits the laser beam LB and a movable mirror 53 that reflects the laser beam LB emitted by the light source 52.
[0189] The movable mirror 53 is, for example, a MEMS mirror using microelectromechanical systems (MEMS) components. For instance, the movable mirror 53 is configured to tilt by rotating about two tilt axes orthogonal to each other on the mirror surface. Furthermore, the tilt angle of the movable mirror 53 is electrically controlled. This allows for scanning the spot 33 of the laser beam LB and rendering any pattern.
[0190] As described above, when using the laser beam scanning irradiation device 51, the irradiation device 51 can be controlled to project a pattern 55, including at least one of text or an image, onto the projection point Pproj (guide point). For example, in Figure 15 In the example shown, a pattern 55 centered on the projection point Pproj includes textual information, such as the name of viewer 2. Additionally, textual information related to speech or elapsed time can be used. Furthermore, image patterns such as signs or icons representing directions can be projected. This allows for the appropriate guidance of the gaze L1 while simultaneously informing the subject 1 of various information during image capture.
[0191] In the above embodiments, a configuration using a single irradiation device has been described. However, multiple irradiation devices can be used. Using multiple irradiation devices makes it possible to shorten the irradiation distance performed by each irradiation device. This makes it possible, for example, to reduce the irradiation of the laser beam LB onto, for example, a non-irradiated target.
[0192] Furthermore, when using multiple illumination devices, the illumination range of each device can be pre-set according to the shape of the green screen, ensuring that line-of-sight guidance information items such as light spots do not overlap. Additionally, for example, if a non-illuminating target is located in the illumination path of an illumination device, other illumination devices can be used to project light spots.
[0193] The following example has been described in the above embodiments, in which, assuming that the real-time video of subject 1 is provided to viewer 2, the attention focus position Pin that subject 1 is expected to gaze at is set as the viewpoint position P2 of viewer 2. The attention focus position Pin can be set arbitrarily. For example, a three-dimensional position pre-designed in virtual space 18 can be used as the attention focus position Pin.
[0194] Figure 16 This is a schematic diagram illustrating another example of an application of a line-of-sight guidance system. Figure 16 The attention focus position Pin' in virtual space 18 and the viewpoint position P1' of the subject 1 are schematically shown. Here, the attention focus position Pin' is the image capture position Pcam' of image capture performed by the virtual camera device 57 set in virtual three-dimensional space 18 (virtual space 18).
[0195] The virtual camera device 57 is a virtual camera device that captures images of an object (e.g., a 3D model of subject 1) as a target in virtual space 18, and corresponds, for example, to the image capture viewpoint used when rendering the target. When the camera work (movement trajectory 58) of the virtual camera device 57 is predetermined, the image capture position Pcam' of the image capture performed by the virtual camera device 57 is three-dimensional data that changes sequentially over time along the movement trajectory 58.
[0196] In this configuration, while capturing an image of subject 1, a light spot 33 is projected onto a projection point Pproj, which is calculated to orient the line of sight L1 toward the image capture position Pcam'. Therefore, subject 1, capturing its image in real space 17, follows the light spot 33 with its eyes while observing the virtual camera device 57 moving along a movement trajectory 58 in virtual space 18. This allows subject 1 to move in real space 17 while simultaneously experiencing image capture as if following the movement of the camera device.
[0197] Figure 17 This is a schematic diagram illustrating another example of an application of a gaze guidance system. Figure 17 The text describes gaze guidance for multiple subjects 1 when capturing volumetric motion images of a real scene. When capturing images of multiple subjects 1 simultaneously in volumetric image capture, the accuracy of extracting the subject image 3 for each subject 1 may decrease, which could lead to a decrease in the accuracy of the 3D model 6. Therefore, it is possible to capture images of multiple subjects 1 individually. However, when performing individual image capture, the uniformity of the scene may be lost due to variations, for example, in the orientation of the gaze of each subject 1.
[0198] Therefore, an attention-focusing pin is also presented to perform gaze guidance when performing individual image capture. Figure 17 A schematically illustrates the focus point Pin in virtual space 18. a 'Viewpoint position P1 of subject 1a a Here, the focus of attention is Pin. a 'Image capture position Pcam' is the image capture position performed by the virtual camera device 57 that moves along the pre-set movement trajectory 58.
[0199] also, Figure 17 B schematically illustrates the focus point Pin in virtual space 18. b 'Viewpoint position P1 of subject 1b b For example, the point of focus of attention for subject 1b is Pin. b 'Set at the point of focus of attention of the subject 1a, Pina At the same location. In other words, the location of attention, Pin. b 'Set to be along with Figure 17 The movement trajectory 58 shown in Figure A is similar to the movement trajectory 58 of the virtual camera device 57, which performs image capture at the image capture position Pcam'. This allows for the presentation of an attention-focusing position Pin to each subject 1 when performing individual image capture. This enables the performance of uniform image capture when integrating 3D models 6 of multiple people.
[0200] Note that the attention focus pins presented to multiple subjects 1 do not necessarily have to be at the same location. For example, different attention focus pins or a single attention focus pin can be presented to each subject 1 depending on the image capture scenario. This allows for various arrangements, for example, when performing volumetric image capture on multiple subjects 1.
[0201] Figure 18 This is a schematic diagram illustrating another example of an application of a line-of-sight guidance system. Figure 18 A schematically illustrates the image capture space 16 at the image capture side position. Figure 18 B schematically illustrates the background 60 used for composition. The background 60 used for composition can be a real-world image or a virtual constructed image, such as a 3D CG image. Here, the attention focus position Pin that the subject 1 is expected to gaze at is the position of the attention focus target 61 that the subject 1 is expected to gaze at in the background 60, which is combined with the subject image 3 extracted from the captured image of the subject 1.
[0202] For example, Figure 18 Image B shows a subject image 3 with a gray silhouette and combined with a background 60. For example, the traffic light located in the upper left of the background 60 in the image is the focus of attention 61. Therefore, the focus point Pin of subject 1 is set at the position of the focus point 61 (the traffic light in the upper left). Figure 18 As shown in A, in the image capture space 16, a projection point Pproj is calculated to guide the gaze L1 of the subject 1 to the attention focus position Pin, and the light spot 33 is projected onto the projection point Pproj. Note that in Figure 18 In A, when viewed from subject 1, the point of focus, Pin, is set in front of the green screen 11.
[0203] As described above, the position of the object (attention focus target 61) in the background 60 corresponding to the combined target can also be set as the attention focus position Pin. As described above, when the subject image 3 extracted from the image captured by guiding the gaze L1 of the subject 1 is combined with the background 60, the subject image 3 is an image in which the subject 1 is looking at the object that serves as the attention focus target 61. This makes it possible to provide various arrangements consistent with the background 60.
[0204] The above primarily describes the use of multiple camera devices to capture images of a subject. However, the number of camera devices is not limited, and this technique can also be applied when image capture is performed, for example, using a single camera device. Furthermore, the position of the camera device used to capture images of the subject does not necessarily have to be fixed. Note that in the case of using a single camera device, a dedicated sensor capable of performing three-dimensional measurements (e.g., a range sensor or a ToF camera device) can be used to perform three-dimensional measurements on a green screen, etc.
[0205] So far, the descriptions have primarily focused on capturing real-world volumetric motion images and generating 3D models of the subject. The methods or types of video used to capture the subject image are not limited. For example, when performing image capture by extracting the subject image from an image captured using a green screen and combining the subject image with the background without alteration, applying this technique to such image capture also enables the proper guidance of the subject's gaze without compromising the performance of the extracted subject image.
[0206] The above has described an example of a control device corresponding to the processor of the gaze guidance controller executing an information processing method according to the present technology. However, it is not limited thereto; the information processing method and program according to the present technology can be executed by a cooperating gaze guidance controller and another computer, and the information processing device according to the present technology can be implemented by a cooperating gaze guidance controller and another computer, the other computer being able to communicate with the gaze guidance controller via, for example, a network.
[0207] In other words, the information processing methods and programs according to this technology can be executed not only in computer systems including a single computer, but also in computer systems in which multiple computers operate collaboratively. Note that in this disclosure, a system refers to a group of components (e.g., devices and modules (components)), and it is not important whether all components are housed in a single housing. Therefore, both multiple devices housed in separate housings and interconnected via a network, and a single device in which multiple modules are housed in a single housing, are systems.
[0208] The information processing methods and programs according to this technology executed by a computer system include, for example, the following two cases: the acquisition of the viewpoint position of the subject, the position of attention focus, and virtual spatial information about the screen, the determination of virtual intersections, and the control of the illumination direction of illumination performed by the illumination device, all performed by a single computer; and the cases where each process is executed by different computers. Furthermore, the execution of each process by a designated computer includes: having another computer perform part or all of the process and obtaining its results.
[0209] In other words, the information processing methods and procedures according to this technology are also applicable to cloud computing configurations, in which multiple devices share and collaborate to process a single function via a network.
[0210] At least two features of the present technology described above can also be combined. In other words, various features described in different embodiments can be combined arbitrarily, regardless of the embodiment. Furthermore, the various effects described above are not limiting but merely illustrative and may provide other effects.
[0211] In this disclosure, expressions such as “same,” “equal,” and “orthogonal” conceptually include expressions such as “substantially the same,” “substantially equal,” and “substantially orthogonal.” For example, expressions such as “same,” “equal,” and “orthogonal” also include states within a specified range (e.g., a range of + / -10%), wherein expressions such as “completely identical,” “completely equal,” and “completely orthogonal” are used as references.
[0212] Note that this technology can also be configured as follows.
[0213] (1) An information processing device, comprising:
[0214] The acquisition unit acquires:
[0215] The viewpoint position of the subject, the image of which is captured using a screen set in real space for chroma key synthesis.
[0216] The desired focus of attention for the subject, and
[0217] Virtual spatial information indicating the position of the screen in virtual three-dimensional space;
[0218] A determining unit determines the virtual intersection point of a virtual line and the screen indicated by the virtual space information, the virtual line connecting the viewpoint position of the subject and the focus position in the virtual three-dimensional space; and
[0219] The control unit controls the irradiation direction of the irradiation performed by the irradiation device installed in the real space, so that it points to the guide point on the screen corresponding to the virtual intersection point in the real space.
[0220] (2) The information processing apparatus according to (1), wherein,
[0221] The focus of attention is the viewpoint of the viewer viewing the display device, on which the subject image is extracted from an image of the subject captured against the background of the screen.
[0222] (3) The information processing apparatus according to (1) or (2), wherein,
[0223] The acquisition unit acquires the viewpoint position of the subject based on the image of the subject, the image of which is captured by a plurality of camera devices positioned in the real space such that the screen is the background within the image capture range.
[0224] (4) The information processing apparatus according to (3), wherein,
[0225] The acquisition unit generates the virtual space information based on images captured by the plurality of camera devices.
[0226] (5) The information processing apparatus according to (3) or (4), wherein,
[0227] The acquisition unit
[0228] Images of light spots are captured using the plurality of camera devices, and these light spots are projected onto the screen by the illumination device.
[0229] The three-dimensional coordinates of the light spot are calculated based on the images captured using the plurality of camera devices.
[0230] (6) The information processing apparatus according to at least one of (3) to (5), wherein,
[0231] The control unit calculates the irradiation direction based on the position of the guide point in the real space and the position of the irradiation device.
[0232] (7) The information processing apparatus according to (6), wherein,
[0233] The control unit calculates the rotation angle of the irradiation direction relative to the reference irradiation direction, which is used as a reference for the irradiation direction in the irradiation device.
[0234] (8) The information processing apparatus according to (7), wherein,
[0235] The control unit calibrates the position of the irradiation device and the reference irradiation direction based on the rotation angle of the irradiation direction when the irradiation device projects a light spot onto each of the plurality of camera devices, and based on the installation position of the plurality of camera devices.
[0236] (9) The information processing apparatus according to (7), wherein,
[0237] The irradiation device includes a position detection marker, and
[0238] The control unit calculates the position of the irradiation device and the reference irradiation direction based on images of the position detection markers of the irradiation device captured by the plurality of camera devices.
[0239] (10) The information processing apparatus according to at least one of (6) to (9), wherein,
[0240] The control unit determines whether the irradiation path connecting the position of the irradiation device and the guide point in the real space passes through a non-irradiation area, and avoids pointing towards the guide point in the non-irradiation area.
[0241] When the irradiation path is determined to pass through the non-irradiated area, the control unit stops pointing to the guide point.
[0242] (11) The information processing apparatus according to at least one of (1) to (10), wherein,
[0243] The illumination device is a laser pointer that can rotate in both the elevation and azimuth directions.
[0244] (12) The information processing apparatus according to at least one of (1) to (10), wherein,
[0245] The irradiation device is a laser beam scanning projector that uses a movable mirror to scan a laser beam, and
[0246] The control unit controls the projector so that a pattern including at least one of text or image is projected onto the guide point.
[0247] (13) The information processing apparatus according to at least one of (1) to (12), wherein,
[0248] The screen is a fully surround screen that encloses the entire periphery of the image capture space for capturing images of the subject, and
[0249] The control unit controls the direction of illumination performed by the illumination device, such that it points to the guide point in the area of the fully surrounding screen, the area being in front when viewed from the subject.
[0250] (14) The information processing apparatus according to at least one of (1) to (13), wherein,
[0251] The focus of attention is the image capture position performed by a virtual camera device set in the virtual three-dimensional space.
[0252] (15) The information processing apparatus according to at least one of (1) to (14), wherein,
[0253] The focus of attention is the location of the attention target that the subject is expected to look at within the background, which is combined with the subject image extracted from the captured image of the subject.
[0254] (16) An information processing method executed by a computer system, the information processing method comprising:
[0255] Acquire the viewpoint position of the subject capturing its image using a screen set in real space for chroma key synthesis, the desired focus of attention of the subject, and virtual spatial information indicating the position of the screen in virtual three-dimensional space;
[0256] Determine the virtual intersection point of the virtual line with the virtual point on the screen indicated by the virtual space information, wherein the virtual line connects the viewpoint position of the subject and the focus of attention in the virtual three-dimensional space; and
[0257] The direction of illumination performed by the illumination device set in the real space is controlled so that it points to a guide point on the screen corresponding to the virtual intersection point in the real space.
[0258] (17) A computer-readable recording medium having a program recorded thereon that causes processing to be executed, the processing comprising:
[0259] Acquire the viewpoint position of the subject capturing its image using a screen set in real space for chroma key synthesis, the desired focus of attention of the subject, and virtual spatial information indicating the position of the screen in virtual three-dimensional space;
[0260] Determine the virtual intersection point of the virtual line with the virtual point on the screen indicated by the virtual space information, wherein the virtual line connects the viewpoint position of the subject and the focus of attention in the virtual three-dimensional space; and
[0261] The direction of illumination performed by the illumination device set in the real space is controlled so that it points to a guide point on the screen corresponding to the virtual intersection point in the real space.
[0262] List of reference numerals
[0263] 1. Subject
[0264] 2 viewers
[0265] 11 Green Screen
[0266] 13 Camera devices
[0267] 17 Real Space
[0268] 18 Virtual Space
[0269] 31 Irradiation device
[0270] 32 Eye-Guiding Controller
[0271] 41 Control device
[0272] 42. Attention-Focusing Location Acquisition Section
[0273] 43 Image Capture Processing Unit
[0274] 44. Eye-Guidance Processing Department
[0275] 100 Remote Presentation System
Claims
1. An information processing apparatus, comprising: The acquisition unit acquires: The viewpoint position of the subject, the image of which is captured using a screen set in real space for chroma key synthesis. The desired focus of attention for the subject, and Virtual spatial information indicating the position of the screen in virtual three-dimensional space; A determining unit determines the virtual intersection point of the virtual line and the screen indicated by the virtual space information, wherein the virtual line connects the viewpoint position of the subject and the attention focus position in the virtual three-dimensional space; as well as The control unit controls the irradiation direction of the irradiation performed by the irradiation device installed in the real space, so that it points to the guide point on the screen corresponding to the virtual intersection point in the real space.
2. The information processing apparatus according to claim 1, wherein, The focus of attention is the viewpoint of the viewer viewing the display device, on which the subject image is extracted from an image of the subject captured against the background of the screen.
3. The information processing apparatus according to claim 1, wherein, The acquisition unit acquires the viewpoint position of the subject based on the image of the subject, the image of which is captured by a plurality of camera devices positioned in the real space such that the screen is the background within the image capture range.
4. The information processing apparatus according to claim 3, wherein, The acquisition unit generates the virtual space information based on images captured by the plurality of camera devices.
5. The information processing apparatus according to claim 3, wherein, The acquisition unit Images of light spots are captured using the plurality of camera devices, and these light spots are projected onto the screen by the illumination device. The three-dimensional coordinates of the light spot are calculated based on the images captured using the plurality of camera devices.
6. The information processing apparatus according to claim 3, wherein, The control unit calculates the irradiation direction based on the position of the guide point in the real space and the position of the irradiation device.
7. The information processing apparatus according to claim 6, wherein, The control unit calculates the rotation angle of the irradiation direction relative to the reference irradiation direction, which is used as a reference for the irradiation direction in the irradiation device.
8. The information processing apparatus according to claim 7, wherein, The control unit calibrates the position of the irradiation device and the reference irradiation direction based on the rotation angle of the irradiation direction when the irradiation device projects a light spot onto each of the plurality of camera devices, and based on the installation position of the plurality of camera devices.
9. The information processing apparatus according to claim 7, wherein, The irradiation device includes a position detection marker, and The control unit calculates the position of the irradiation device and the reference irradiation direction based on images of the position detection markers of the irradiation device captured by the plurality of camera devices.
10. The information processing apparatus according to claim 6, wherein, The control unit determines whether the irradiation path connecting the position of the irradiation device and the guide point in the real space passes through a non-irradiation area, and avoids pointing towards the guide point in the non-irradiation area. When the irradiation path is determined to pass through the non-irradiated area, the control unit stops pointing to the guide point.
11. The information processing apparatus according to claim 1, wherein, The illumination device is a laser pointer that can rotate in both the elevation and azimuth directions.
12. The information processing apparatus according to claim 1, wherein, The irradiation device is a laser beam scanning projector that uses a movable mirror to scan a laser beam, and The control unit controls the projector so that a pattern including at least one of text or image is projected onto the guide point.
13. The information processing apparatus according to claim 1, wherein, The screen is a fully surround screen that encloses the entire periphery of the image capture space for capturing images of the subject, and The control unit controls the direction of illumination performed by the illumination device, such that it points to the guide point in the area of the fully surrounding screen, the area being in front when viewed from the subject.
14. The information processing apparatus according to claim 1, wherein, The focus of attention is the image capture position performed by a virtual camera device set in the virtual three-dimensional space.
15. The information processing apparatus according to claim 1, wherein, The focus of attention is the location of the attention target that the subject is expected to look at within the background, which is combined with the subject image extracted from the captured image of the subject.
16. An information processing method executed by a computer system, the information processing method comprising: Acquire the viewpoint position of the subject capturing its image using a screen set in real space for chroma key synthesis, the desired focus of attention of the subject, and virtual spatial information indicating the position of the screen in virtual three-dimensional space; Determine the virtual intersection point between the virtual line and the screen indicated by the virtual space information, wherein the virtual line connects the viewpoint position of the subject and the focus of attention in the virtual three-dimensional space; as well as The direction of illumination performed by the illumination device set in the real space is controlled so that it points to a guide point on the screen corresponding to the virtual intersection point in the real space.
17. A computer-readable recording medium having a program recorded thereon that causes processing to be executed, the processing comprising: Acquire the viewpoint position of the subject capturing its image using a screen set in real space for chroma key synthesis, the desired focus of attention of the subject, and virtual spatial information indicating the position of the screen in virtual three-dimensional space; Determine the virtual intersection point between the virtual line and the screen indicated by the virtual space information, wherein the virtual line connects the viewpoint position of the subject and the focus of attention in the virtual three-dimensional space; as well as The direction of illumination performed by the illumination device set in the real space is controlled so that it points to a guide point on the screen corresponding to the virtual intersection point in the real space.
Citation Information
Patent Citations
Display imaging device
WO2019207922A1