Information Processing Apparatus, Information Processing Method, and Program

The information processing apparatus enhances stereoscopic display visibility by guiding user focus through attention-attracting regions in virtual space, addressing fusion difficulties in conventional naked-eye stereoscopic displays.

JP7708178B2Active Publication Date: 2025-07-15SONY GROUP CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023514343
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-04-12
Filing Date
2022-01-28
Publication Date
2025-07-15
Estimated Expiration
2042-01-28

AI Technical Summary

Technical Problem

Conventional naked-eye stereoscopic displays face difficulties in image fusion due to unclear correspondence between left and right images, especially when representing large depths or continuous arrangements of virtual objects, leading to decreased visibility and potential visual fatigue.

Method used

An information processing apparatus that generates and adjusts viewpoint images by detecting an attention-attracting area in a virtual space, creating a control map for eye-catching degree distribution, and correcting image characteristics to guide user focus, thereby enhancing visibility and promoting image fusion.

Benefits of technology

The solution promotes image fusion and improves visibility by making attention-attracting regions more prominent, allowing for clearer stereoscopic perception and reducing visual fatigue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708178000001
    Figure 0007708178000001
  • Figure 0007708178000002
    Figure 0007708178000002
  • Figure 0007708178000003
    Figure 0007708178000003
Patent Text Reader

Abstract

This information processing device (10) comprises a display generation unit (53), a visual attraction area detection unit (51), a map generation unit (52), an image correction unit (54), and a display control unit (55). The display generation unit (53) generates a plurality of viewpoint images for display as a three-dimensional image. The visual attraction area detection unit (51) detects a visual attraction area in a virtual space that should attract the visual attention of a user. The map generation unit (52) generates, for each viewpoint image, a control map that indicates the distribution of the degree of visual attraction in the viewpoint image, on the basis of the distance from the visual attraction area. The image correction unit (54) adjusts the degree of visual attraction of the viewpoint images on the basis of the control map. The display control unit (55) causes a three-dimensional image to be displayed in the virtual space using the plurality of viewpoint images for which the degree of visual attraction has been adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] A naked-eye stereoscopic display that performs stereoscopic display using binocular parallax is known. Left-eye and right-eye viewpoint images are supplied to the left eye and the right eye of an observer. As a result, a display is realized as if a virtual object exists in front of the observer's eyes.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The images projected onto the left and right retinas are fused in the observer's brain and recognized as one stereoscopic image. Such a brain function is called fusion. When the correspondence between the left and right images is clear, fusion is likely to occur. However, when the represented depth is large or the same virtual objects are arranged continuously, the image range to be recognized as one stereoscopic image becomes unclear and fusion becomes difficult. In conventional naked-eye stereoscopic displays, no consideration is given to ease of fusion. Therefore, depending on the display content, fusion may rarely become difficult and visibility may decrease.

[0005] Therefore, the present disclosure proposes an information processing apparatus, an information processing method, and a program capable of realizing a stereoscopic display that is easy to fuse.

Means for Solving the Problems

[0006] According to the present disclosure, there is provided an information processing apparatus including: a display generation unit that generates a plurality of viewpoint images for displaying as a stereoscopic image; an attention attracting area detection unit that detects an attention attracting area in a virtual space to attract a user's visual attention; a map generation unit that generates a control map indicating a distribution of attention degrees in the viewpoint image for each viewpoint image based on a distance from the attention attracting area; an image correction unit that adjusts the attention degree of the viewpoint image based on the control map; and a display control unit that displays the stereoscopic image in the virtual space using the plurality of viewpoint images with the adjusted attention degrees. Further, according to the present disclosure, there are provided an information processing method in which the information processing of the information processing apparatus is executed by a computer, and a program for causing a computer to realize the information processing of the information processing apparatus.

Brief Description of Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Modes for Carrying Out the Invention

[0008] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same parts are denoted by the same reference numerals, and redundant descriptions are omitted.

[0009] Note that the description will be made in the following order. [1. Configuration of the display system] [2. Specific example of a stereoscopic image] [3. Information processing method] [3-1. Generation of a distance map] [3-2. Setting of the spatial distribution of the degree of eye-catchingness] [3-3. Generation and correction processing of a control map] [3-4. Variation of the spatial distribution of the degree of eye-catchingness] [3-5. Specific example of correction signal processing] [4. Modified example] [5. Example of hardware configuration] [6. Effects]

[0010] [1. Configuration of the display system] FIG. 1 is a diagram showing an example of a display system 1.

[0011] The display system 1 has a display 21 with a screen SCR inclined at an angle θ with respect to the horizontal plane BT. The angle θ is, for example, 45 degrees. Hereinafter, the direction parallel to the lower side of the screen SCR is defined as the x direction. The direction in the horizontal plane BT orthogonal to the lower side is defined as the z direction. The direction (vertical direction) orthogonal to the x direction and the z direction is defined as the y direction.

[0012] In the example of FIG. 1, the size of the screen SCR in the x direction is W, the size in the z direction is D, and the size in the y direction is H. A rectangular parallelepiped space with sizes in the x direction, y direction, and z direction of W, 2×D, and H becomes the virtual space VS. A plurality of viewpoint images VI (see FIG. 5) displayed on the screen SCR are presented as stereoscopic images in the virtual space VS. Hereinafter, the vertical plane of the virtual space VS on the front side of the screen SCR as viewed by the user (observer) is referred to as the front surface FT, and the vertical plane of the virtual space VS on the rear side of the screen SCR is referred to as the rear surface RE.

[0013] In the binocular disparity-based display system 1, the viewpoint images VI projected onto the user's left and right eyes are fused and recognized as a single stereoscopic image. However, when the virtual object VOB (see FIG. 3) suddenly jumps out greatly from the screen SCR or is arranged at a deep depth position, or when the same virtual object VOB is continuously arranged, there are cases where the image range to be recognized as a single stereoscopic image becomes unclear. In such cases, rarely, fusion becomes difficult and visibility decreases.

[0014] When viewing a real object, by adjusting the focus of the eyes, the position of the real object in the depth direction can be explored. Therefore, based on the position in the depth direction, the correspondence relationship between the left and right images can be recognized. However, in a naked-eye stereoscopic display, what is actually observed is the stereoscopic illusion image (viewpoint image VI) displayed on the screen SCR. The focal position of the illusion image is fixed on the screen SCR, which serves as the light source. Therefore, it is not possible to explore the depth of the virtual object VOB by adjusting the focus of the eyes. This makes fusion more difficult.

[0015] In the present disclosure, in order to solve such problems, a method is proposed in which a specific region in the virtual space VS is set as an attention - attracting region RA (see FIG. 8), and correction signal processing is performed so that the attention - attracting region RA becomes prominent compared to other regions. By guiding the user's line of sight ES (see FIG. 8) to the attention - attracting region RA, the image range to be recognized as a single stereoscopic image can be easily specified. As a result, fusion is promoted. This will be described in detail below.

[0016] FIG. 2 is a diagram showing the functional configuration of the display system 1.

[0017] The display system 1 includes a processing unit 10, an information presentation unit 20, an information input unit 30, and a sensor unit 40. The processing unit 10 is an information processing device that processes various information. The processing unit 10 controls the information presentation unit 20 based on the sensor information acquired from the sensor unit 40 and the user input information acquired from the information input unit 30.

[0018] The sensor unit 40 includes a plurality of sensors for sensing the external world. The plurality of sensors include, for example, a visible - light camera 41, a distance - measuring sensor 42, a line - of - sight detection sensor 43, etc. The visible - light camera 41 captures a visible - light image of the external world. The distance - measuring sensor 42 detects the distance of a real object existing in the external world using, for example, the flight time of laser light. The line - of - sight detection sensor 43 detects the user's line of sight ES directed to the display 21 using a known eye - tracking technique.

[0019] The information presentation unit 20 presents various information such as video information, acoustic information, and tactile information to the user. The information presentation unit 20 includes, for example, a display 21, a speaker 22, and a haptics device 23. A known display such as an LCD (Liquid Crystal Display) or an OLED (Organic Light - Emitting Diode) is used for the display 21. A known speaker capable of outputting sound, etc., is used for the speaker 22. A known haptics device capable of presenting tactile information related to the display information by, for example, ultrasonic waves, is used for the haptics device 23.

[0020] The information input unit 30 has a plurality of input devices capable of inputting various types of information by a user's input operation. The plurality of input devices include, for example, a touch panel 31, a keyboard 32, a mouse 33, and a microphone 34, etc.

[0021] The processing unit 10 has, for example, a data processing unit 11, an I / F unit 12, a gaze recognition unit 13, a distance recognition unit 14, a user recognition unit 15, an attention - attracting information recognition unit 16, a virtual space recognition unit 17, a timer 18, and a storage unit 19. The processing unit 10 acquires the sensor information detected by the sensor unit 40 and the user input information input from the information input unit 30 via the I / F unit 12.

[0022] Based on the information detected by the gaze detection sensor 43, the gaze recognition unit 13 generates gaze information of the user who directs the gaze ES to the display 21. The gaze information includes information regarding the position of the user's eyes (viewpoint VP: refer to FIG. 8) and the direction of the gaze ES. A known eye - tracking technology is used for the gaze recognition process.

[0023] Based on the information detected by the distance measurement sensor 42, the distance recognition unit 14 generates distance information of a real object existing in the external world. The distance information includes, for example, information on the distance between the real object and the display 21.

[0024] The user recognition unit 15 extracts an image of the user who directs the gaze ES to the display 21 from the visible - light image captured by the visible - light camera 41. The user recognition unit 15 generates operation information of the user based on the extracted user image. The operation information includes, for example, information regarding the situation and gestures of the work being performed by the user while looking at the display 21.

[0025] Based on the user input information, the sensor information, and the content data CT, the attention - attracting information recognition unit 16 generates attention - attracting information of the virtual space VS. The attention - attracting information includes information regarding objects or places (attention - attracting regions RA) that should attract the user's visual attention. The attention - attracting information is used to identify the position information of the key attention - attracting region RA.

[0026] For example, in a general display form where the main object is placed on the front side (the side closer to the user), information for identifying the object on the front side is generated as eye-catching information. When the content data CT includes information on the eye-catching position (object or location) specified by the content producer, the information on the eye-catching position extracted from the content data CT is generated as eye-catching information. When the user continuously gazes at a specific position, information regarding the user's gazing position is generated as eye-catching information. For example, when it is detected from the sensor information that the user inserts a finger or a pen into the virtual space VS and is performing some operation such as drawing or shaping on the virtual object VOB, the position information of the operation location (gazing position) is generated as eye-catching information.

[0027] The virtual space recognition unit 17 generates virtual space information regarding the virtual space VS. The virtual space information includes, for example, the angle θ of the screen SCR, as well as information regarding the position and size of the virtual space VS.

[0028] The data processing unit 11 drives the information presentation unit 20 and the sensor unit 40 synchronously based on the timing signal generated by the timer 18. The data processing unit 11 controls the information presentation unit 20 to display a stereoscopic image with the eye-catching degree (visibility) adjusted according to the distance from the eye-catching region in the virtual space VS. The data processing unit 11 includes, for example, an eye-catching region detection unit 51, a map generation unit 52, a display generation unit 53, an image correction unit 54, and a display control unit 55.

[0029] The display generation unit 53 generates a plurality of viewpoint images VI for display as a stereoscopic image. The viewpoint image VI means a two-dimensional image seen from one viewpoint VP. The plurality of viewpoint images VI include a left-eye image seen from the user's left eye and a right-eye image seen from the user's right eye.

[0030] For example, the display generation unit 53 detects the position and size of the virtual space VS based on the virtual space information acquired from the virtual space recognition unit 17. The display generation unit 53 detects the positions of the user's left and right eyes (viewpoint VP) based on the gaze information acquired from the gaze recognition unit 13. The display generation unit 53 extracts 3D data from the content data CT, and renders the extracted 3D data based on the user's viewpoint to generate a viewpoint image VI.

[0031] The eye-catching area detection unit 51 detects an eye-catching area RA of the virtual space VS that should attract the user's visual attention based on the eye-catching information acquired from the eye-catching information recognition unit 16. The eye-catching area RA is, for example, a specific virtual object VOB presented in the virtual space VS, or a local area within the virtual space VS that includes a specific virtual object VOB. The eye-catching area RA is detected based on, for example, user input information, the user's fixation position, or the eye-catching position extracted from the content data CT. The user's fixation position is detected based on, for example, the user's motion information acquired from the user recognition unit 15.

[0032] The map generation unit 52 generates a control map CM (see FIG. 11) for each viewpoint image VI based on the distance from the eye-catching area RA. The control map CM shows the distribution of the eye-catching degree in the viewpoint image VI. In the virtual space VS, a spatial distribution of the eye-catching degree is set such that the eye-catching degree is the largest in the eye-catching area RA.

[0033] For example, in the control map CM, a distribution of the eye-catching degree is defined such that the eye-catching degree decreases as the distance from the eye-catching area RA increases. The eye-catching degree is calculated based on the distance from the eye-catching area RA. The reference distance may be the distance in the depth direction or the distance in a direction orthogonal to the depth direction. The depth direction may be the user's line of sight direction or the z direction. The distance from the eye-catching area RA is calculated based on the distance information acquired from the distance recognition unit 14.

[0034] The image correction unit 54 adjusts the conspicuousness of the viewpoint image VI based on the control map CM. For example, the image correction unit 54 adjusts the conspicuousness of the viewpoint image VI by adjusting the frequency characteristics, brightness, saturation, contrast, transparency, or hue of the viewpoint image VI for each pixel.

[0035] For example, the image correction unit 54 maximizes characteristics such as frequency characteristics, brightness, saturation, contrast, and transparency in the conspicuous region RA. Thereby, the conspicuous region RA is manifested and the conspicuousness is increased. When a plurality of virtual objects VOB having the same hue are presented in the virtual space VS, the image correction unit 54 can make the hue of the virtual object VOB presented in the conspicuous region RA different from the hues of the virtual objects VOB in other regions. The virtual object VOB with the adjusted hue is identified as a heterogeneous virtual object VOB. Therefore, the conspicuousness of the virtual object VOB in the conspicuous region RA is increased.

[0036] By the above-described image processing, the conspicuousness is adjusted for each region. Therefore, even if homogeneous edges and textures are continuously arranged, they are likely to be perceived as different edges and textures. Thereby, fusion is promoted and a display with high visibility is realized. Since the visibility is improved, a large depth expression becomes possible and the perceived stereoscopic effect is also improved. Furthermore, the difficulty of fusion causes visual fatigue, but by eliminating this, a reduction in visual fatigue can also be expected.

[0037] The display control unit 55 displays a stereoscopic image in the virtual space VS using a plurality of viewpoint images VI with the adjusted conspicuousness.

[0038] Information regarding settings, conditions, and criteria used in various calculations is included in the setting information STI. The content data CT, setting information STI, and program PG used in the above-described processing are stored in the storage unit 19. The program PG is a program that causes a computer to execute information processing according to the present embodiment. The processing unit 10 performs various processes according to the program PG stored in the storage unit 19. The storage unit 19 may be used as a work area for temporarily storing the processing results of the processing unit 10. The storage unit 19 includes, for example, any non-transitory storage medium such as a semiconductor storage medium and a magnetic storage medium. The storage unit 19 includes, for example, an optical disk, a magneto-optical disk, or a flash memory. The program PG is stored, for example, in a non-transitory storage medium readable by a computer.

[0039] The processing unit 10 is, for example, a computer composed of a processor and a memory. The memory of the processing unit 10 includes a RAM (Random Access Memory) and a ROM (Read Only Memory). By executing the program PG, the processing unit 10 functions as a data processing unit 11, an I / F unit 12, a gaze recognition unit 13, a distance recognition unit 14, a user recognition unit 15, an eye-catching information recognition unit 16, a virtual space recognition unit 17, a timer 18, an eye-catching area detection unit 51, a map generation unit 52, a display generation unit 53, an image correction unit 54, and a display control unit 55.

[0040] [2. Specific Examples of Stereoscopic Images] FIG. 3 and FIG. 4 are diagrams showing an example of a stereoscopic image presented in the virtual space VS. FIG. 5 is a diagram showing a viewpoint image VI of the stereoscopic image viewed from a specific viewpoint VP.

[0041] The content data CT includes information on the 3D model of the stereoscopic image. By rendering the 3D model based on the information of the viewing point VP, a viewpoint image VI viewed from an arbitrary viewing point VP is generated. In the examples of FIGS. 3 to 5, a plurality of cube-shaped virtual objects VOB are periodically arranged in the x, y, and z directions. The plurality of virtual objects VOB are widely distributed in the z direction from the front side to the rear side of the screen SCR. The viewpoint image VI includes, for example, a plurality of virtual objects VOB and their shadow images SH.

[0042] The user observes the stereoscopic image from the front side FT of the virtual space VS. Depending on the observation position, virtual objects VOB having the same edges and textures in all directions are observed. Therefore, it is difficult to distinguish each virtual object VOB, and the display becomes very difficult to fuse. To solve this problem, in the present disclosure, a specific space region is made prominent, and the user's line of sight ES is guided to this space region to promote fusion. Hereinafter, an example of information processing will be described.

[0043] [3. Information processing method] [3-1. Generation of distance map] FIG. 6 is a diagram for explaining the signal processing flow. FIG. 7 is a diagram showing the distance map DM of the viewpoint image VI.

[0044] The map generation unit 52 generates a distance map DM for each viewpoint image VI based on the three-dimensional coordinate information of the stereoscopic image. The distance map DM shows the distribution of the distance from the viewing point VP to the surface of the virtual object VOB. The distance map DM defines the distance from the viewing point VP for each pixel. For each pixel of the distance map DM, for example, the distance to the position (for example, the front surface FT) in the virtual space VS closest to the user is set to 0, and the distance to the position (for example, the rear surface RE) in the virtual space VS farthest from the user is set to 1, and the normalized distance value is defined as the pixel value.

[0045] [3-2. Setting of spatial distribution of eye-catching degree] FIGS. 8 to 10 are diagrams showing an example of the spatial distribution AD of the eye-catching degree.

[0046] The eye-catching area detection unit 51 generates position information of the eye-catching area RA based on the eye-catching information. The eye-catching area detection unit 51 supplies the position information of the eye-catching area RA to the map generation unit 52 as a control key. The map generation unit 52 determines the spatial distribution AD of the eye-catching degree in the virtual space VS with reference to the position of the eye-catching area RA. In the example of FIG. 8, the position on the virtual space VS closest to the user's viewpoint VP is determined as the eye-catching area RA. The eye-catching area RA is defined as a planar area orthogonal to the line of sight ES. The eye-catching degree decreases as the distance from the viewpoint VP increases.

[0047] In the example of FIG. 9, the front surface FT of the virtual space VS is determined as the eye-catching area RA. A spatial distribution AD of the eye-catching degree is set such that the eye-catching degree decreases as the distance from the front surface FT increases. The eye-catching degree of the back surface RE is the smallest. In the example of FIG. 10, the area within the virtual space VS where the normalized distance from the viewpoint VP is DS A is determined as the eye-catching area RA. The map generation unit 52 determines a correction value of the image according to the eye-catching degree as a control value CV. The map generation unit 52 determines a control curve CCV that defines the relationship between the distance DS and the control value CV based on the spatial distribution AD of the eye-catching degree.

[0048] [3-3. Generation and Correction Processing of Control Map] FIG. 11 is a diagram showing an example of the control map CM.

[0049] The map generation unit 52 generates the control map CM based on the distance map DM and the spatial distribution AD. For example, the map generation unit 52 generates the control map CM for each viewpoint image VI by applying the control curve CCV to the distance map DM. The control map CM shows the distribution of the control value CV of the corresponding viewpoint image VI. The control map CM defines the control value CV of each pixel of the viewpoint image VI. The image correction unit 54 generates a control signal for correction signal processing based on the control map CM. The image correction unit 54 corrects the viewpoint image VI using the control signal.

[0050] [3-4. Variations in Spatial Distribution of Eye-catching Degree] Figures 12 to 14 are diagrams showing variations in the spatial distribution AD of conspicuousness.

[0051] In the example of Figure 12, the conspicuous area RA is set at the center (indicated by the symbol "a") of the virtual space VS as viewed from the direction of the user's line of sight ES (line-of-sight direction ESD). The planar area at the center of the virtual space VS orthogonal to the line-of-sight direction ESD is the conspicuous area RA. The conspicuousness is the greatest in the conspicuous area RA, and on the front side (the side closer to the user) and the rear side (the side farther from the user) of the conspicuous area RA, the conspicuousness gradually decreases according to the distance from the conspicuous area RA.

[0052] In the first example from the left in Figure 13, the conspicuous area RA is set at the end on the front surface FT side of the virtual space VS as viewed from the line-of-sight direction ESD. The planar area passing through the upper side of the front surface FT and orthogonal to the line-of-sight direction ESD is the conspicuous area RA. The conspicuousness is the greatest in the conspicuous area RA, and from the conspicuous area RA to the center of the virtual space VS, the conspicuousness gradually decreases according to the distance from the conspicuous area RA. On the rear side of the center of the virtual space VS, the conspicuousness does not change.

[0053] In the second example from the left in Figure 13, the conspicuous area RA is set on the front side of the center of the virtual space VS as viewed from the line-of-sight direction ESD. The planar area orthogonal to the line-of-sight direction ESD is the conspicuous area RA. On the front side and the rear side of the conspicuous area RA, the conspicuousness gradually decreases according to the distance from the conspicuous area RA.

[0054] In the third example from the left in Figure 13, the conspicuous area RA is set at the center of the virtual space VS as viewed from the line-of-sight direction ESD. The spatial distribution AD of conspicuousness is the same as in the example of Figure 12.

[0055] In the fourth example from the left in Figure 13, the conspicuous area RA is set on the rear side of the center of the virtual space VS as viewed from the line-of-sight direction ESD. The planar area orthogonal to the line-of-sight direction ESD is the conspicuous area RA. On the front side and the rear side of the conspicuous area RA, the conspicuousness gradually decreases according to the distance from the conspicuous area RA.

[0056] In the fifth example from the left in FIG. 13, the attracting region RA is set at the end on the back surface RE side of the virtual space VS as seen from the line-of-sight direction ESD. A planar region orthogonal to the line-of-sight direction ESD passing through the lower side of the back surface RE is the attracting region RA. The attracting degree is the largest in the attracting region RA, and from the attracting region RA to the central part of the virtual space VS, the attracting degree gradually decreases according to the distance from the attracting region RA. On the far side from the central part of the virtual space VS, the attracting degree does not change.

[0057] In the example of FIG. 14, the attracting region RA is set at the central part of the virtual space VS as seen from the z direction. A planar region at the central part of the virtual space VS orthogonal to the z direction is the attracting region RA. The attracting degree is the largest in the attracting region RA, and on the front side and the back side of the attracting region RA, the attracting degree gradually decreases according to the distance from the attracting region RA.

[0058] [3-5. Specific Examples of Correction Signal Processing] FIGS. 15 to 17 are diagrams showing specific examples of correction signal processing. The left side of each figure shows the viewpoint image VI before correction, and the right side shows the viewpoint image VI (corrected image VIC) after correction.

[0059] The correction signal processing is a process of adjusting the attracting degree of the viewpoint image VI for each pixel. The correction signal processing aims to make the region to be attracted more prominent and the other regions less prominent, or to make it easier to distinguish and recognize each of a large number of virtual objects VOB of the same type. In the correction signal processing, for example, the following multiple processes are performed alone or in combination.

[0060] In the example of FIG. 15, according to the control map CM, processing is performed to increase the frequency characteristics of the region with a high control value CV and decrease those of other parts. By this processing, the sharpness of the main region that becomes the attention - attracting region RA and the virtual object VOB increases, and the visibility of edges and textures also becomes higher than others. As a result, a display that is easy to fuse is obtained. In the example of FIG. 15, the attention - attracting degree of the foremost side as seen from the line - of - sight direction ESD is set to be high. In the corrected image VIC, the sharpness of the foremost side is high, and the sharpness decreases as going to the rear side. By attracting attention to the front side and creating a difference from the rear side at the same time, it becomes easier to fuse.

[0061] In the example of FIG. 16, according to the control map CM, processing is performed to brighten the region with a high control value CV and darken other parts. By this processing, the brightness of the main region that becomes the attention - attracting region RA and the virtual object VOB becomes higher than others, and it becomes more prominent. As a result, the visibility of the main region that becomes the attention - attracting region RA and the virtual object VOB increases, and a display that is easy to fuse is obtained. In the example of FIG. 16, the attention - attracting degree of the foremost side as seen from the line - of - sight direction ESD is set to be high. In the corrected image VIC, the foremost side is bright, and it becomes darker as going to the rear side. By attracting attention to the front side and creating a difference from the rear side at the same time, it becomes easier to fuse.

[0062] In the example of FIG. 17, according to the control map CM, processing is performed to increase the saturation of the region with a high control value CV and decrease the saturation of other parts. By this processing, the saturation of the main region that becomes the attention - attracting region RA and the virtual object VOB becomes more vivid than others. As a result, the visibility of the main region that becomes the attention - attracting region RA and the virtual object VOB increases, and a display that is easy to fuse is obtained. In the example of FIG. 17, the attention - attracting degree of the foremost side as seen from the line - of - sight direction ESD is set to be high. In the corrected image VIC, the foremost side is vivid, and it becomes duller as going to the rear side. By attracting attention to the front side and creating a difference from the rear side at the same time, it becomes easier to fuse.

[0063] Note that the correction signal processing is not limited to the above. For example, according to the control map CM, a process may be performed to increase the local contrast in the region with a high control value CV and decrease the contrast in other parts. The local contrast means the contrast within the virtual object VOB existing in a local space. By this process, the main region that becomes the attention region RA and the texture of the virtual object VOB are displayed more vividly than others. As a result, the visibility of the main region that becomes the attention region RA and the virtual object VOB is enhanced, and a display that is easy to fuse is obtained.

[0064] According to the control map CM, a process may be performed to lower the transparency in the region with a high control value CV and increase the transparency in other parts. By this process, the main region that becomes the attention region RA and the virtual object VOB are more likely to stand out. As a result, the visibility of the main region that becomes the attention region RA and the virtual object VOB is enhanced, and a display that is easy to fuse is obtained.

[0065] When a plurality of homogeneous virtual objects VOB are presented in the virtual space VS, according to the control map CM, a process may be performed to differentiate the virtual objects VOB for each region by changing the color phase of the region with a high control value CV and its region. By this process, each individual virtual object VOB becomes easier to distinguish. As a result, the visibility of the main virtual object VOB that becomes the attention region RA is enhanced, and a display that is easy to fuse is obtained.

[0066] [4. Modification Example] FIG. 18 is a diagram showing a modification example of the correction signal processing.

[0067] All of the above-described correction signal processing is implemented as post-processing applied to the viewpoint image VI. However, a similar display can also be realized by controlling the settings of the individual virtual objects VOB to be drawn, such as their materials, according to the position. For example, the image correction unit 54 extracts a plurality of virtual objects VOB from the content data CT. The image correction unit 54 adjusts the attention degree of the virtual object VOB based on the alpha value according to the distance between the virtual object VOB and the attention region RA for each virtual object VOB.

[0068] In the example of FIG. 18, the front side as viewed from the z - direction is defined as the main region that becomes the attention - attracting region RA. For the virtual object VOB on the rear side, the alpha value of the material is lowered, resulting in a highly transparent display. By this process, the virtual object VOB on the front side that becomes the attention - attracting region RA becomes more prominent. As a result, the visibility of the virtual object VOB on the front side that becomes the attention - attracting region RA is enhanced, and a display that is easy to fuse is obtained.

[0069] [5. Hardware Configuration Example] FIG. 19 is a diagram showing an example of the hardware configuration of the display system 1.

[0070] The display system 1 includes a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, a RAM (Random Access Memory) 903, and a host bus 904a. Further, the display system 1 includes a bridge 904, an external bus 904b, an interface 905, an input device 906, an output device 907, a storage device 908, a drive 909, a connection port 911, a communication device 913, and a sensor 915. The display system 1 may have a processing circuit such as a DSP or an ASIC instead of or together with the CPU 901.

[0071] The CPU 901 functions as an arithmetic processing unit and a control unit, and controls the overall operation within the display system 1 according to various programs. Also, the CPU 901 may be a microprocessor. The ROM 902 stores programs, arithmetic parameters, etc. used by the CPU 901. The RAM 903 temporarily stores programs used in the execution of the CPU 901 and parameters that change appropriately during the execution. The CPU 901 can form, for example, a data processing unit 11, a gaze recognition unit 13, a distance recognition unit 14, a user recognition unit 15, an attention - attracting information recognition unit 16, and a virtual space recognition unit 17.

[0072] The CPU 901, ROM 902, and RAM 903 are interconnected by a host bus 904a including a CPU bus or the like. The host bus 904a is connected to an external bus 904b such as a PCI (Peripheral Component Interconnect / Interface) bus via a bridge 904. Note that it is not always necessary to separately configure the host bus 904a, the bridge 904, and the external bus 904b, and these functions may be implemented in one bus.

[0073] The input device 906 is realized by a device into which information is input by a user, such as a mouse, a keyboard, a touch panel, buttons, a microphone, switches, and levers. Further, the input device 906 may be, for example, a remote control device using infrared rays or other radio waves, or may be an external connection device such as a mobile phone or a PDA corresponding to the operation of the display system 1. Furthermore, the input device 906 may include, for example, an input control circuit that generates an input signal based on information input by the user using the above input means and outputs the input signal to the CPU 901. A user of the display system 1 can input various data to the display system 1 or instruct a processing operation by operating the input device 906. The input device 906 can form, for example, the information input unit 30.

[0074] The output device 907 is formed of a device capable of visually or auditorily notifying the user of the acquired information. Such devices include display devices such as CRT display devices, liquid crystal display devices, plasma display devices, EL display devices, and lamps, audio output devices such as speakers and headphones, and printer devices. The output device 907 outputs, for example, the results obtained by various processes performed by the display system 1. Specifically, the display device visually displays the results obtained by various processes performed by the display system 1 in various forms such as text, images, tables, and graphs. On the other hand, the audio output device converts an audio signal composed of reproduced audio data, acoustic data, etc. into an analog signal and outputs it auditorily. The output device 907 can form, for example, the information presentation unit 20.

[0075] The storage device 908 is a data storage device formed as an example of the storage unit of the display system 1. The storage device 908 is realized, for example, by a magnetic storage unit device such as an HDD, a semiconductor storage device, an optical storage device, or a magneto-optical storage device. The storage device 908 may include a storage medium, a recording device for recording data on the storage medium, a reading device for reading data from the storage medium, and a deleting device for deleting data recorded on the storage medium. This storage device 908 stores programs executed by the CPU 901, various data, and various data acquired from the outside. The above storage device 908 can form, for example, the storage unit 19.

[0076] The drive 909 is a reader / writer for a storage medium and is built into or externally attached to the display system 1. The drive 909 reads the information recorded on a removable storage medium such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory that is mounted, and outputs it to the RAM 903. The drive 909 can also write information to the removable storage medium.

[0077] The connection port 911 is an interface for connecting to an external device, and is a connection port for an external device capable of data transmission, for example, by USB (Universal Serial Bus).

[0078] The communication device 913 is a communication interface formed by, for example, a communication device for connecting to the network 920. The communication device 913 is, for example, a communication card for wired or wireless LAN (Local Area Network), LTE (Long Term Evolution), Bluetooth (registered trademark), or WUSB (Wireless USB). Further, the communication device 913 may be a router for optical communication, a router for ADSL (Asymmetric Digital Subscriber Line), or a modem for various types of communication. This communication device 913 can transmit and receive signals, etc. in accordance with a predetermined protocol such as TCP / IP, for example, between the Internet and other communication devices.

[0079] The sensor 915 is various sensors such as, for example, an acceleration sensor, a gyro sensor, a geomagnetic sensor, an optical sensor, a sound sensor, a distance measuring sensor, a force sensor, etc. The sensor 915 acquires information regarding the state of the display system 1 itself, such as the posture and moving speed of the display system 1, and information regarding the surrounding environment of the display system 1, such as the brightness and noise around the display system 1. Further, the sensor 915 may include a GPS sensor that receives a GPS signal and measures the latitude, longitude, and altitude of the device. The sensor 915 can form, for example, the sensor unit 40.

[0080] Note that the network 920 is a wired or wireless transmission path for information transmitted from devices connected to the network 920. For example, the network 920 may include public line networks such as the Internet, telephone line networks, satellite communication networks, and various LANs (Local Area Networks) including Ethernet (registered trademark), WANs (Wide Area Networks), etc. Further, the network 920 may include a dedicated line network such as an IP-VPN (Internet Protocol-Virtual Private Network).

[0081] [6. Effects] The processing unit 10 includes a display generation unit 53, an eye-catching area detection unit 51, a map generation unit 52, an image correction unit 54, and a display control unit 55. The display generation unit 53 generates a plurality of viewpoint images VI for displaying as a stereoscopic image. The eye-catching area detection unit 51 detects an eye-catching area RA of the virtual space VS that should attract the user's visual attention. The map generation unit 52 generates a control map CM indicating the distribution of the eye-catching degree in the viewpoint image VI for each viewpoint image VI based on the distance from the eye-catching area RA. The image correction unit 54 adjusts the eye-catching degree of the viewpoint image VI based on the control map CM. The display control unit 55 displays a stereoscopic image in the virtual space VS using the plurality of viewpoint images VI with the adjusted eye-catching degree. The information processing method of the present embodiment is such that the processing of the above-described processing unit 10 is executed by a computer. The program of the present embodiment causes a computer to realize the processing of the above-described processing unit 10.

[0082] According to this configuration, it is possible to prompt the observer to gaze at the image area to be recognized as one stereoscopic image. Therefore, fusion is likely to occur.

[0083] In the control map CM, a distribution of the eye-catching degree is defined such that the eye-catching degree decreases as the distance from the eye-catching area RA increases.

[0084] According to this configuration, the eye-catching area RA becomes more prominent than other areas. Therefore, fusion is promoted.

[0085] The map generation unit 52 generates a distance map DM of each viewpoint image VI based on the three-dimensional coordinate information of the stereoscopic image. The map generation unit 52 determines the spatial distribution AD of the saliency degree of the virtual space VS based on the position of the salient region RA. The map generation unit 52 generates a control map CM based on the distance map DM and the spatial distribution AD of the saliency degree.

[0086] According to this configuration, the control map CM can be easily generated based on the three-dimensional coordinate information of the stereoscopic image.

[0087] The image correction unit 54 adjusts the saliency degree of the viewpoint image VI by adjusting the frequency characteristics, brightness, chroma, contrast, transparency, or hue of the viewpoint image VI for each pixel.

[0088] According to this configuration, it becomes easier to prompt the observer to fixate on the salient region RA.

[0089] The image correction unit 54 extracts a plurality of virtual objects VOB from the content data CT. The image correction unit 54 adjusts the saliency degree of the virtual object VOB based on an alpha value corresponding to the distance between the virtual object VOB and the salient region RA for each virtual object VOB.

[0090] According to this configuration, it becomes easier to prompt the observer to fixate on the salient region RA.

[0091] The salient region detection unit 51 detects the salient region RA based on user input information, the user's fixation position, or the salient position extracted from the content data CT.

[0092] According to this configuration, the salient region RA is appropriately set.

[0093] Note that the effects described in this specification are merely examples and are not limiting, and there may be other effects.

[0094] [Appendix] Note that this technology can also adopt the following configuration. (1) A display generation unit that generates a plurality of viewpoint images for display as a stereoscopic image, An eye-catching area detection unit that detects an eye-catching area in a virtual space that should attract the user's visual attention, A map generation unit that generates a control map showing the distribution of eye-catching degrees in the viewpoint image for each viewpoint image based on the distance from the eye-catching area, An image correction unit that adjusts the eye-catching degree of the viewpoint image based on the control map, A display control unit that displays the stereoscopic image in the virtual space using the plurality of viewpoint images with the adjusted eye-catching degree, An information processing apparatus having the above. (2) In the control map, a distribution of the eye-catching degree is defined such that the eye-catching degree decreases as the distance from the eye-catching area increases, The information processing apparatus according to the above (1). (3) The map generation unit generates a distance map for each viewpoint image based on the three-dimensional coordinate information of the stereoscopic image, determines the spatial distribution of the eye-catching degree in the virtual space with reference to the position of the eye-catching area, and generates the control map based on the distance map and the spatial distribution of the eye-catching degree, The information processing apparatus according to the above (1) or (2). (4) The image correction unit adjusts the eye-catching degree of the viewpoint image by adjusting the frequency characteristics, brightness, saturation, contrast, transparency, or hue of the viewpoint image for each pixel, The information processing apparatus according to any one of the above (1) to (3). (5) The image correction unit extracts a plurality of virtual objects from the content data, and adjusts the eye-catching degree of the virtual objects based on an alpha value corresponding to the distance between the virtual object and the eye-catching area for each virtual object, The information processing apparatus according to any one of the above (1) to (3). (6) The eye-catching area detection unit detects the eye-catching area based on user input information, the user's fixation position, or an eye-catching position extracted from the content data, The information processing apparatus according to any one of (1) to (5) above. (7) Generate a plurality of viewpoint images for displaying as a stereoscopic image, Detect an eye-catching area in the virtual space that should attract the user's visual attention, Based on the distance from the eye-catching area, generate a control map for each viewpoint image that shows the distribution of the eye-catching degree in the viewpoint image, Adjust the eye-catching degree of the viewpoint image based on the control map, Display the stereoscopic image in the virtual space using the plurality of viewpoint images with the adjusted eye-catching degree. An information processing method executed by a computer, comprising the above. (8) Generate a plurality of viewpoint images for displaying as a stereoscopic image, Detect an eye-catching area in the virtual space that should attract the user's visual attention, Based on the distance from the eye-catching area, generate a control map for each viewpoint image that shows the distribution of the eye-catching degree in the viewpoint image, Adjust the eye-catching degree of the viewpoint image based on the control map, Display the stereoscopic image in the virtual space using the plurality of viewpoint images with the adjusted eye-catching degree. A program for causing a computer to realize the above.

Explanation of Signs

[0095] 10 Processing unit (information processing apparatus) 51 Eye-catching area detection unit 52 Map generation unit 53 Display generation unit 54 Image correction unit 55 Display control unit AD Spatial distribution of eye-catching degree CM Control map CT Content data DM Distance map PG Program RA Eye-catching area VI Viewpoint image VOB Virtual Object VS Virtual Space

Claims

1. A display generation unit that generates a plurality of viewpoint images for displaying as a stereoscopic image; An eye-catching area detection unit that detects an eye-catching area in a virtual space that should attract the user's visual attention; A map generation unit that generates a control map showing the distribution of the eye-catching degree in the viewpoint image for each viewpoint image based on the distance from the eye-catching area; An image correction unit that adjusts the eye-catching degree of the viewpoint image based on the control map; A display control unit that displays the stereoscopic image in the virtual space using the plurality of viewpoint images with the adjusted eye-catching degree; An information processing apparatus having the above.

2. In the control map, a distribution of the eye-catching degree is defined such that the eye-catching degree decreases as the distance from the eye-catching area increases. The information processing apparatus according to Claim 1.

3. The map generation unit generates a distance map for each viewpoint image based on the three-dimensional coordinate information of the stereoscopic image, determines the spatial distribution of the eye-catching degree in the virtual space with reference to the position of the eye-catching area, and generates the control map based on the distance map and the spatial distribution of the eye-catching degree. The information processing apparatus according to Claim 1.

4. The image correction unit adjusts the eye-catching degree of the viewpoint image by adjusting the frequency characteristics, brightness, saturation, contrast, transparency, or hue of the viewpoint image for each pixel. The information processing apparatus according to Claim 1.

5. The image correction unit extracts a plurality of virtual objects from the content data, and adjusts the eye-catching degree of each virtual object based on an alpha value corresponding to the distance between the virtual object and the eye-catching area. The information processing apparatus according to Claim 1.

6. The eye-catching area detection unit detects the eye-catching area based on user input information, the user's fixation position, or an eye-catching position extracted from the content data. The information processing apparatus according to Claim 1.

7. Generating a plurality of viewpoint images for displaying as a stereoscopic image; Detecting an eye-catching area in a virtual space that should attract the user's visual attention; Generating a control map showing the distribution of the eye-catching degree in the viewpoint image for each viewpoint image based on the distance from the eye-catching area; Adjusting the eye-catching degree of the viewpoint image based on the control map; Displaying the stereoscopic image in the virtual space using the plurality of viewpoint images with the adjusted eye-catching degree. An information processing method executed by a computer, comprising the above.

8. Generating a plurality of viewpoint images for displaying as a stereoscopic image; Detect an eye-catching area of a virtual space that should attract the user's visual attention, Based on the distance from the eye-catching area, generate a control map showing the distribution of eye-catching degrees in the viewpoint image for each viewpoint image, Adjust the eye-catching degree of the viewpoint image based on the control map, Display the stereoscopic image in the virtual space using the plurality of viewpoint images with the adjusted eye-catching degree, A program that causes a computer to realize the above.

Citation Information

Patent Citations

  • Image processing device and method, program, and recording medium

    JP2013162330A

  • Stereoscopic vision image processing apparatus, stereoscopic vision image processing method, and program

    JP2015076776A

  • Image processing apparatus and method

    US20130106844A1

  • Light field display control methods and apparatuses, light field display devices

    US20170372683A1

  • Information processing device, information processing method, and program

    WO2018116580A1