Information processing device and information processing method

The server-side rendering system addresses response delays and processing capacity issues in VR by generating two-dimensional video data and saliency maps for out-of-field objects, ensuring high-quality 6DoF experiences across various devices.

JP7740333B2Active Publication Date: 2025-09-17SONY GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023523964
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-27
Filing Date
2022-01-17
Publication Date
2025-09-17
Estimated Expiration
2042-01-17

AI Technical Summary

Technical Problem

Existing technologies face challenges in delivering high-quality virtual reality (VR) videos due to response delays and the need for high processing capacity at the client device, especially in 6DoF content distribution, which requires rendering large amounts of three-dimensional data.

Method used

A server-side rendering system that generates two-dimensional video data based on user field of view information, estimates recognition positions of out-of-field objects, and generates saliency maps for these areas, using predicted head motion and saliency maps to reduce data transmission and processing load on the client device.

Benefits of technology

This approach reduces response delays and allows for high-quality VR video delivery with reduced data transmission, enabling immersive 6DoF experiences even on devices with lower processing power by offloading rendering to a server.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007740333000001
    Figure 0007740333000001
  • Figure 0007740333000002
    Figure 0007740333000002
  • Figure 0007740333000003
    Figure 0007740333000003
Patent Text Reader

Abstract

An information processing device according to one embodiment of the present technology comprises a rendering unit, an estimation unit, and a generation unit. The rendering unit executes a rendering process on three-dimensional spatial data constituting a virtual space, on the basis of visual field information about the visual field of a user, thereby generating two-dimensional video data corresponding to the visual field of the user. The estimation unit estimates a recognition position at which the user recognizes a recognition target object recognized by the user, in an outside-the-visual-field area not included in the visual field of the user in the virtual space. The generation unit generates a saliency map expressing saliency in the outside-the-visual-field area, on the basis of the estimated recognition position of the recognition target object in the outside-the-visual-field area.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present technology relates to an information processing device and an information processing method that can be applied to the distribution of VR (Virtual Reality) images, etc. [Background technology]

[0002] In recent years, panoramic images that can be viewed in all directions, captured by panoramic cameras, etc., have begun to be distributed as VR images. Furthermore, technology for distributing 6DoF (Degree of Freedom) images (also called 6DoF content) that allow viewers (users) to look around in all directions (freely select the line of sight) and move freely in three-dimensional space (freely select the viewpoint) has been developed. Such 6DoF content dynamically reproduces a three-dimensional space using one or more three-dimensional objects according to the viewer's viewpoint position, line of sight direction, and viewing angle (viewing range) at each time. In such video distribution, it is necessary to dynamically adjust (render) the video data presented to the viewer according to the viewer's visual field. For example, one example of such a technology is disclosed in Patent Document 1.

[0003] Non-Patent Document 1 describes a process for estimating a saliency map for a spherical image. In this estimation process, planar images from various camera directions are extracted from a spherical image, and a saliency map for each planar image is estimated using a saliency map estimation model for planar images. The saliency maps for each planar image are integrated, and a horizon bias is applied to the horizontal line direction at the center of the image, to estimate a saliency map for the spherical image. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Special Publication No. 2007-520925 [Non-patent literature]

[0005] [Non-Patent Document 1] Takao Yamanaka, "Technology for Estimating Prominent Locations in Images: Saliency Map Estimation Using Deep Learning," [online], September 15, 2020, Reiwa 2 New Technology Briefing, Internet<URL:https: / / shingi.jst.go.jp / var / rev0 / 0001 / 1222 / 02_sophia_yamanaka.pdf> Summary of the Invention [Problem to be solved by the invention]

[0006] It is expected that the distribution of virtual images (virtual video) such as VR video will become widespread, and there is a demand for technology that enables the distribution of high-quality virtual video.

[0007] In view of the above circumstances, an object of the present technology is to provide an information processing device and an information processing method that are capable of realizing the delivery of high-quality virtual video. [Means for solving the problem]

[0008] In order to achieve the above object, an information processing device according to an embodiment of the present technology includes a rendering unit, an estimation unit, and a generation unit. The rendering unit generates two-dimensional video data according to the user's field of view by performing a rendering process on three-dimensional space data that constitutes a virtual space based on field of view information relating to the user's field of view. The estimation unit estimates a recognition position, recognized by the user, of a recognition target object recognized by the user in an out-of-field area of ​​the virtual space that is not included in the user's field of view. The generation unit generates a saliency map representing saliency in the out-of-field area based on the estimated recognition position of the object to be recognized in the out-of-field area.

[0009] In this information processing device, the recognition position of a recognition target object in an area outside the field of view is estimated. A saliency map for the area outside the field of view is generated based on the estimated recognition position. This makes it possible to generate a highly accurate saliency map for the area outside the field of view, and to realize the delivery of high-quality virtual video using the saliency map.

[0010] The estimation unit may set, as the recognition target object, an object that has been a rendering target up to the current time.

[0011] The two-dimensional video data may be configured from a plurality of frame images that are consecutive in time series. In this case, the estimation unit may estimate the recognition position of the recognition target object that is not included in the frame image at the current time based on a position in the virtual space that corresponds to a position of the recognition target object in a most recent past frame image that includes the recognition target object.

[0012] The estimation unit may estimate, as the recognition position, a position in the virtual space corresponding to a position of the object to be recognized in the most recent frame image.

[0013] The estimation unit may estimate, as the recognition position, a position shifted along a movement direction of the object to be recognized from a position in the virtual space corresponding to the position of the object to be recognized in the most recent frame image.

[0014] When the estimation unit determines that the user has recognized a sound emitted by the object to be recognized that is not included in the frame image at the current time, the estimation unit may estimate the position where the sound is generated in the virtual space as the recognition position.

[0015] The three-dimensional space data may include three-dimensional space description data that defines a configuration of the virtual space and three-dimensional object data that defines three-dimensional objects in the virtual space. In this case, the three-dimensional space description data may include role information that indicates a role of the recognition target object and fixed position information that indicates a fixed position associated with the role. Furthermore, for the recognition target object to which predetermined role information that is not included in a frame image at the current time is set, if the recognition target object to which the same role information is set has been rendered up to the current time, the estimation unit may estimate the fixed position associated with the role as the recognition position.

[0016] The estimation unit may estimate the fixed position related to the role as the recognition position if the object to be recognized, to which the same role information is set, has been rendered at the fixed position related to the role up to the current time.

[0017] The three-dimensional space data may include three-dimensional space description data that defines a configuration of the virtual space and three-dimensional object data that defines three-dimensional objects in the virtual space. In this case, the three-dimensional space description data may include role information that indicates a role of the recognition target object and fixed position information that indicates a fixed position associated with the role. Furthermore, when the estimation unit determines that the user has recognized a sound emitted by a recognition target object that has predetermined role information set and is not included in a frame image at the current time, the estimation unit may estimate the fixed position associated with the role as the recognition position.

[0018] When the estimation unit determines that the user has recognized a sound made by the object to be recognized, which has the same role information set up to the current time, while the object is in the fixed position related to the role, the estimation unit may estimate the fixed position related to the role as the recognition position.

[0019] The estimation unit may estimate the recognition position based on a position of the object to be recognized in the two-dimensional video data in which the object to be recognized is rendered.

[0020] The generating unit may generate the saliency map such that saliency based on bottom-up attention in the out-of-field region is zero.

[0021] The generation unit may generate the saliency map representing saliency based on top-down attention in the out-of-field region, based on the recognition position of the object to be recognized in the out-of-field region.

[0022] The generating unit may generate the saliency map for the out-of-field area and a saliency map representing saliency of the two-dimensional video data.

[0023] The information processing device may further include a prediction unit that generates the future view information as predicted view information based on the saliency map. In this case, the rendering unit may generate the two-dimensional video data based on the predicted view information.

[0024] The field of view information may include at least one of a position of a viewpoint, a direction of a line of sight, a rotation angle of a line of sight, a position of the user's head, or a rotation angle of the user's head.

[0025] The field of view information may include a head rotation angle of the user. In this case, the prediction unit may predict a future head rotation angle of the user based on the saliency map.

[0026] The two-dimensional video data may be composed of a plurality of frame images that are successive in time series. In this case, the rendering unit may generate frame images based on the predicted view information and output the frame images as predicted frame images.

[0027] An information processing method according to one embodiment of the present technology is an information processing method executed by a computer system, and includes generating two-dimensional video data corresponding to a user's field of view by performing a rendering process on three-dimensional spatial data constituting a virtual space based on field of view information relating to the user's field of view. A recognition position of a recognition target object recognized by the user in an out-of-field area of ​​the virtual space that is not included in the user's field of view is estimated. A saliency map representing saliency in the out-of-field area is generated based on the estimated recognition position of the object to be recognized in the out-of-field area. [Brief explanation of the drawings]

[0028] [Figure 1] FIG. 1 is a schematic diagram illustrating an example of the basic configuration of a server-side rendering system. [Figure 2] FIG. 10 is a schematic diagram illustrating an example of a virtual video that can be viewed by a user. [Figure 3] FIG. 10 is a schematic diagram for explaining a rendering process. [Figure 4] FIG. 1 is a schematic diagram illustrating an example of the configuration of a server-side rendering system. [Figure 5] FIG. 2 is a schematic diagram for explaining a recognition target object and a recognition position. [Figure 6] 10 is a flowchart illustrating an example of generating a rendering video. [Figure 7] FIG. 7 is a diagram for explaining the flowchart shown in FIG. 6, and is a schematic diagram showing the timing of obtaining and generating each piece of information. [Figure 8] FIG. 1 is a schematic diagram illustrating an example of generating a saliency map based on bottom-up attention. [Figure 9] FIG. 10 is a schematic diagram for explaining a problem with the panoramic saliency map of the comparative example. [Figure 10] FIG. 10 is a schematic diagram showing an example of information described in a scene description file used as scene description information in the second embodiment. [Figure 11]FIG. 11 is a schematic diagram showing an example of information described in a scene description file used as scene description information in the third embodiment. [Figure 12] FIG. 11 is a schematic diagram showing an example of information described in a scene description file used as scene description information in the third embodiment. [Figure 13] FIG. 11 is a schematic diagram for explaining estimation of a recognition position of a recognition target object in the third embodiment. [Figure 14] FIG. 11 is a schematic diagram for explaining estimation of a recognition position of a recognition target object in the third embodiment. [Figure 15] 10 is a flowchart illustrating an example of estimating the recognition position of a recognition target object. [Figure 16] 10 is a flowchart illustrating an example of generating a panoramic saliency map. [Figure 17] FIG. 10 is a schematic diagram illustrating an example of a panoramic saliency map. [Figure 18] FIG. 2 is a block diagram showing an example of the hardware configuration of a computer (information processing device) that can realize the server device and the client device. DETAILED DESCRIPTION OF THE INVENTION

[0029] Hereinafter, embodiments of the present technology will be described with reference to the drawings.

[0030] [Server-side rendering system] As one embodiment of the present technology, a server-side rendering system is configured. First, an example of the basic configuration and basic operation of the server-side rendering system will be described with reference to Figs. FIG. 1 is a schematic diagram showing an example of the basic configuration of a server-side rendering system. FIG. 2 is a schematic diagram illustrating an example of a virtual video that can be viewed by a user. FIG. 3 is a schematic diagram for explaining the rendering process. The server-side rendering system can also be called a server-rendering type media distribution system.

[0031] As shown in FIG. 1, the server-side rendering system 1 includes an HMD (Head Mounted Display) 2, a client device 3, and a server device 4. The HMD 2 is a device used to display a virtual image to the user 5. The HMD 2 is worn on the head of the user 5 when in use. For example, when VR video is distributed as virtual video, an immersive HMD 2 configured to cover the field of view of the user 5 is used. When an AR (Augmented Reality) image is distributed as the virtual image, AR glasses or the like are used as the HMD 2. A device other than the HMD 2 may be used as a device for providing a virtual image to the user 5. For example, the virtual image may be displayed on a display provided on a television, a smartphone, a tablet terminal, a PC (Personal Computer), or the like.

[0032] As shown in FIG. 2, in this embodiment, 6DoF images are provided as VR images to a user 5 wearing an immersive HMD 2. Within the virtual space S, which is a three-dimensional space, the user 5 can view images in a 360° range around them, including front, back, left, right, and up and down. For example, within the virtual space S, the user 5 can freely change the position of the viewpoint, the direction of the line of sight, etc., to freely change his or her field of view (field of view) 7. In response to this change in the field of view 7 of the user 5, the image 8 displayed to the user 5 is switched. By performing actions such as turning the face, tilting the face, and looking back, the user 5 can view the surroundings within the virtual space S with a sensation similar to that of the real world. In this way, the server-side rendering system 1 according to this embodiment can deliver photorealistic free viewpoint video, making it possible to provide a viewing experience from any viewpoint position.

[0033] In this embodiment, the HMD 2 acquires visual field information. The visual field information is information relating to the visual field 7 of the user 5. Specifically, the visual field information includes any information that can identify the visual field 7 of the user 5 in the virtual space S. For example, the visual field information includes the position of the viewpoint, the direction of the line of sight, the rotation angle of the line of sight, etc. Furthermore, the visual field information includes the position of the head of the user 5, the rotation angle of the head of the user 5, etc. The rotation angle of the line of sight can be defined, for example, by the rotation angle about an axis extending in the line of sight direction. The rotation angle of the head of user 5 can be defined by the roll angle, pitch angle, and yaw angle when three mutually perpendicular axes set on the head are defined as the roll axis, pitch axis, and yaw axis. For example, the axis extending in the direction of the face is defined as the roll axis. When the face of the user 5 is viewed from the front, the axis extending in the left-right direction is defined as the pitch axis, and the axis extending in the up-down direction is defined as the yaw axis. The roll angle, pitch angle, and yaw angle relative to these roll axis, pitch axis, and yaw axis are calculated as the rotation angle of the head. Note that the direction of the roll axis can also be used as the line of sight direction. Any other information may be used that can identify the field of view of the user 5. As the field of view information, one of the above-exemplified pieces of information may be used, or a combination of a plurality of pieces of information may be used.

[0034] The method for acquiring the visual field information is not limited. For example, the visual field information can be acquired based on the detection result (sensing result) by a sensor device (including a camera) provided in the HMD 2. For example, the HMD 2 is provided with a camera or distance measurement sensor that detects the area around the user 5, an inward-facing camera that can capture images of the left and right eyes of the user 5, etc. The HMD 2 is also provided with an IMU (Inertial Measurement Unit) sensor and a GPS. For example, position information of the HMD 2 acquired by GPS can be used as the viewpoint position of the user 5 or the position of the head of the user 5. Of course, the positions of the left and right eyes of the user 5 may be calculated in more detail. It is also possible to detect the line of sight direction from the captured images of the left and right eyes of the user 5. It is also possible to detect the rotation angle of the line of sight and the rotation angle of the head of the user 5 from the detection results of the IMU.

[0035] Furthermore, the self-position of the user 5 (HMD 2) may be estimated based on the detection result of a sensor device provided in the HMD 2. For example, the self-position estimation can calculate position information of the HMD 2 and posture information such as the direction in which the HMD 2 is facing. From the position information and posture information, it is possible to acquire field of view information. The algorithm for estimating the self-position of the HMD 2 is not limited, and any algorithm such as SLAM (Simultaneous Localization and Mapping) may be used. Furthermore, head tracking for detecting the movement of the head of the user 5 and eye tracking for detecting the movement of the eyes of the user 5 to the left and right may be performed.

[0036] Any other device or algorithm may be used to acquire the visual field information. For example, when a smartphone or the like is used as a device that displays a virtual image to the user 5, an image of the face (head) of the user 5 may be captured, and the visual field information may be acquired based on the captured image. Alternatively, a device including a camera, an IMU, etc. may be worn on the user 5's head or around the eyes. To generate visual field information, any machine learning algorithm using, for example, a DNN (Deep Neural Network) may be used. For example, by using AI (artificial intelligence) that performs deep learning, it is possible to improve the accuracy of generating visual field information. It should be noted that the application of machine learning algorithms may be performed on any process within the present disclosure.

[0037] The HMD 2 and the client device 3 are connected to each other so that they can communicate with each other. The communication method for connecting the two devices to each other so that they can communicate with each other is not limited, and any communication technology may be used. For example, wireless network communication such as WiFi, short-range wireless communication such as Bluetooth (registered trademark), etc. may be used. The HMD 2 transmits the visual field information to the client device 3. The HMD 2 and the client device 3 may be integrated. That is, the HMD 2 may be equipped with the functions of the client device 3.

[0038] The client device 3 and the server device 4 each have hardware necessary for configuring a computer, such as a CPU, a ROM, a RAM, and an HDD (see FIG. 18). The CPU loads a program according to the present technology, which is pre-recorded in the ROM or the like, into the RAM and executes the program, thereby executing the information processing method according to the present technology. For example, any computer such as a PC (Personal Computer) can be used to realize the client device 3 and the server device 4. Of course, hardware such as FPGA and ASIC may also be used. Of course, the client device 3 and the server device 4 are not limited to having the same configuration.

[0039] The client device 3 and the server device 4 are connected via a network 9 so as to be able to communicate with each other. The network 9 is constructed, for example, by the Internet, a wide area communication network, etc. Alternatively, any WAN (Wide Area Network) or LAN (Local Area Network) may be used, and the protocol for constructing the network 9 is not limited.

[0040] The client device 3 receives the visual field information transmitted from the HMD 2. The client device 3 also transmits the visual field information to the server device 4 via the network 9.

[0041] The server device 4 receives the visual field information transmitted from the client device 3. Furthermore, the server device 4 generates two-dimensional video data (rendered video) corresponding to the visual field 7 of the user 5 by performing a rendering process on the three-dimensional space data that constitutes the virtual space S based on the visual field information. The server device 4 corresponds to an embodiment of an information processing device according to the present technology. The server device 4 executes an embodiment of an information processing method according to the present technology.

[0042] As shown in FIG. 3, the three-dimensional space data includes scene description information and three-dimensional object data. The scene description information corresponds to three-dimensional space description data that defines the configuration of the virtual space S (three-dimensional space). The scene description information includes various metadata for reproducing each scene of the 6DoF content. The three-dimensional object data is data that defines three-dimensional objects in the virtual space S (three-dimensional space). In other words, it is data for each object that makes up each scene of the 6DoF content. For example, data of three-dimensional objects such as people and animals, or data of three-dimensional objects such as buildings and trees, or data of three-dimensional objects such as the sky and sea that make up the background, etc. Multiple types of objects may be collectively configured as a single three-dimensional object, and the data for that object may be stored. Three-dimensional object data may consist of mesh data that can be expressed as shape data of a polyhedron and texture data that is data to be applied to the surfaces of the mesh data, or may consist of a set of multiple points (point cloud).

[0043] As shown in FIG. 3, the server device 4 reproduces a virtual space S that constitutes each scene by arranging three-dimensional objects in a three-dimensional space based on the scene description information. 3, an XYZ coordinate system is set in the virtual space S, and a three-dimensional object is placed at a position defined by coordinate values. The coordinate values ​​correspond to position information in the virtual space S, and can also be called world coordinates. There are no limitations on the method for setting the XYZ coordinate system in the virtual space S, and any setting method may be adopted. Using the reproduced virtual space S as a reference, an image seen by the user 5 is cut out (rendering process), and a rendered image, which is a two-dimensional image viewed by the user 5, is generated. The server device 4 encodes the generated rendering image and transmits it to the client device 3 via the network 9 . The rendered image according to the user's field of view 7 can also be said to be an image of a viewport (display area) according to the user's field of view 7.

[0044] The client device 3 decodes the encoded rendering video transmitted from the server device 4. The client device 3 also transmits the decoded rendering video to the HMD 2. 2, the rendered image is played back by the HMD 2 and displayed to the user 5. Hereinafter, the image 8 displayed to the user 5 by the HMD 2 may be referred to as the rendered image 8.

[0045] [Advantages of a server-side rendering system] Another example of a 6DoF video delivery system, such as the one shown in FIG. 2, is a client-side rendering system. In a client-side rendering system, a client device 3 performs rendering processing on three-dimensional spatial data based on visual field information to generate two-dimensional video data (rendered video 8). The client-side rendering system can also be called a client-rendering type media distribution system. In a client-side rendering system, it is necessary to distribute three-dimensional space data (three-dimensional space description data and three-dimensional object data) from the server device 4 to the client device 3. The 3D object data is composed of mesh data or point cloud data. Therefore, the amount of data delivered from the server device 4 to the client device 3 becomes enormous. Furthermore, in order to perform the rendering process, the client device 3 is required to have a fairly high processing capacity.

[0046] In contrast, in the server-side rendering system 1 according to this embodiment, the rendered image 8 after rendering is delivered to the client device 3. This makes it possible to sufficiently reduce the amount of data delivered. In other words, with a small amount of data delivered, it becomes possible to allow the user 5 to experience 6DoF images of a large space composed of a huge amount of 3D object data. In addition, it becomes possible to offload the processing load on the client device 3 side to the server device 4 side, and even if a client device 3 with low processing power is used, it becomes possible for the user 5 to experience 6DoF video.

[0047] [Response delay issue] In the server-side rendering system 1, the view information of the user 5 and the rendered image 8 are transmitted and received via a network 9. Therefore, there is a possibility that a response delay may occur in the display of the rendered image 8 according to the movement of the viewpoint, etc. For example, the user 5 changes the field of view 7 by moving his / her head. The field of view information is acquired by the HMD 2 and transmitted to the client device 3. The client device 3 transmits the received field of view information to the server device 4 via the network 9. The server device 4 performs rendering processing on the three-dimensional space data based on the received visual field information of the user 5, and generates a rendered image 8. The generated rendered image 8 is encoded and transmitted to the client device 3 via the network 9. The client device 3 decodes the received rendered image 8 and transmits it to the HMD 2. The HMD 2 displays the received rendered image 8 to the user 5. The server-side rendering system 1 is configured to execute such a processing flow in real time in response to changes in the field of view of the user 5. In this case, there is a possibility that a delay occurs between when the user 5 changes the field of view and when that change is reflected in the image on the HMD 2, which is called a response delay. This response delay can also be expressed as Motion-to-Photon Latency (T_m2p). It is desirable to keep this response delay to 20 msec or less, which is the limit of human perception.

[0048] This technology is extremely effective in solving the above-mentioned problem of response delay. An embodiment of a server-side rendering system 1 to which this technology is applied will be described in detail below. In the following embodiment, a case where head motion information is used as visual field information of the user 5 will be exemplified. The Head Motion information includes Position information (X, Y, Z) that represents the positional movement of the head of the user 5, and Orientation information (yaw, pitch, roll) that represents the rotational movement of the head of the user 5. The position information (X, Y, Z) corresponds to position information in the virtual space S, and is defined by coordinate values ​​(world coordinates) of the XYZ coordinate system set in the virtual space S. The orientation information (yaw, pitch, roll) is defined by the roll angle, pitch angle, and yaw angle relative to a roll axis, pitch axis, and yaw axis that are set on the head of the user 5 and are perpendicular to each other. Of course, application of this technology is not limited to cases where head motion information (X, Y, Z, yaw, pitch, roll) is used as visual field information of the user 5. This technology is also applicable to cases where other information is used as visual field information.

[0049] In the following embodiment, the server-side rendering system 1 acquires the visual field information of the user 5 in real time, and displays the rendered video to the user 5. In the following description, the time when the visual field information of the user 5 is acquired by the server-side rendering system 1 is referred to as the "current time." In other words, the time when the visual field information of the user 5 is acquired by the HMD 2 is referred to as the "current time." As described above, there is a possibility that a response delay (T_m2p time) may occur between the time when the field of view information acquired at the "current time" is transmitted to the server device 4, the time when the rendered image 8 is generated, and the time when the image is displayed by the HMD 2. By applying this technology, it is possible to sufficiently mitigate the problem of response delays from the "current time," enabling the delivery of high-quality virtual video.

[0050] FIG. 4 is a schematic diagram showing an example of the configuration of a server-side rendering system 1 according to an embodiment of the present technology. The server-side rendering system 1 shown in FIG. 4 includes an HMD 2, a client device 3, and a server device 4. The HMD 2 is capable of acquiring in real time visual field information (head motion information) of the user 5. As described above, the time when the head motion information is acquired by the HMD 2 is the current time. The HMD 2 acquires head motion information at a predetermined frame rate and transmits it to the client device 3. Therefore, "head motion information at the current time" is repeatedly transmitted to the client device 3 at the predetermined frame rate. Similarly, the client device 3 also repeatedly transmits "head motion information at the current time" to the server device 4 at a predetermined frame rate.

[0051] The frame rate for acquiring head motion information (number of times head motion information is acquired per second) is set to be synchronized with the frame rate of the rendered video 8, for example. For example, the rendered video 8 is composed of a plurality of frame images that are successive in time series. Each frame image is generated at a predetermined frame rate. The frame rate for acquiring head motion information is set to be synchronized with the frame rate of the rendered video 8. Of course, this is not a limitation. As described above, AR glasses or a display may be used as a device for displaying a virtual image to the user 5.

[0052] The server device 4 includes a data input unit 11, a head motion information recording unit 12, a prediction unit 13, a rendering unit 14, an encoding unit 15, and a communication unit 16. The server device 4 also includes a saliency map generation unit 17, a saliency map recording unit 18, and a recognition position estimation unit 19. These functional blocks are realized by, for example, a CPU executing a program according to the present technology, and the information processing method according to the present embodiment is executed. Note that dedicated hardware such as an IC (integrated circuit) may be used as appropriate to realize each functional block.

[0053] The data input unit 11 reads out three-dimensional space data (scene description information and three-dimensional object data) and outputs it to the rendering unit . The three-dimensional spatial data is stored, for example, in a storage unit 68 (see FIG. 18) in the server device 4. Alternatively, the three-dimensional spatial data may be managed by a content server or the like that is communicably connected to the server device 4. In this case, the data input unit 11 acquires the three-dimensional spatial data by accessing the content server.

[0054] The communication unit 16 is a module for executing network communication, short-distance wireless communication, etc. with other devices. For example, a wireless LAN module such as WiFi, or a communication module such as Bluetooth (registered trademark) is provided. In this embodiment, the communication unit 16 realizes communication with the client device 3 via the network 9.

[0055] The Head Motion information recording unit 12 records the visual field information (Head Motion information) received from the client device 3 via the communication unit 16 in the storage unit 68 (see FIG. 18 ). For example, a buffer or the like for recording the visual field information (Head Motion information) may be configured. The "Head Motion information at the current time" transmitted at a predetermined frame rate is accumulated and stored in the storage unit 68.

[0056] The prediction unit 13 generates future visual field information as predicted visual field information based on a panoramic saliency map (panoramic saliency map). In this embodiment, future head motion information of the user 5 is predicted and generated as predicted head motion information. The predicted head motion information includes future position information (X, Y, Z) and future orientation information (yaw, pitch, roll). That is, in this embodiment, the head position and head rotation angle are predicted based on the omnidirectional saliency map.

[0057] The rendering unit 14 executes the rendering process exemplified in Fig. 3. That is, by executing the rendering process on the three-dimensional space data based on the visual field information relating to the visual field of the user 5, a rendered image 8 according to the visual field 7 of the user 5 is generated. In this embodiment, the rendering unit 14 generates frame images that make up the rendered video 8 based on the predicted field of view information (predicted head motion information) generated by the prediction unit 13. Hereinafter, the frame images generated based on the predicted head motion information will be referred to as predicted frame images 20. The rendering unit 14 is configured by, for example, a reproduction unit that reproduces the virtual space S, a renderer, a parameter setting unit that sets rendering parameters, etc. The rendering parameters include a resolution map that indicates the resolution for each region. Alternatively, the rendering unit 14 may have any other configuration.

[0058] The encoding unit 15 performs encoding (compression coding) on ​​the rendered video 8 (predicted frame image 20) to generate distribution data. The distribution data is transmitted to the client device 3 via the communication unit 16. For example, the encoding process is performed in real time on each region of the rendered video 8 (predicted frame image 20) based on a QP map (quantization parameter). More specifically, in this embodiment, the encoding unit 15 can suppress image quality degradation due to compression of points of interest and important areas within the predicted frame image 20 by switching the quantization precision (QP: Quantization Parameter) for each area within the predicted frame image 20. In this way, it is possible to suppress an increase in the load of distribution data and processing while maintaining sufficient video quality for areas important to user 5. Note that the QP value here is a value indicating the quantization step when using lossy compression efficiency, and a high QP value reduces the amount of coding and increases compression efficiency, resulting in greater image quality degradation due to compression, while a low QP value increases the amount of coding and decreases compression efficiency, making it possible to suppress image quality degradation due to compression. Any other compression encoding technology may be used. The encoding unit 15 is configured by, for example, an encoder, a parameter setting unit that sets encoding parameters, etc. The encoding parameters include the above-mentioned QP map, etc. For example, the QP map is generated based on the resolution map set by the parameter setting unit of the rendering unit 14. Alternatively, the encoding unit 15 may have any other configuration.

[0059] The saliency map generator 17 generates a panoramic saliency map. The saliency map generation unit 17 generates not only a saliency map for the field of view that represents the saliency of the rendered image (two-dimensional image data) 8 viewed by the user 5, but also a saliency map for the out-of-field area that is not included in the field of view 7 of the user 5. The visual field saliency map is information that quantitatively represents how likely each pixel in the rendered image 8 is to attract attention, based on the mechanism of human visual attention. The saliency map for the out-of-field area can be generated as a map in which the saliency of each pixel is calculated when the out-of-field area of ​​the user 5 is represented as a 2D image. The generation of the omnidirectional saliency map will be described in detail later. The saliency map is also called a saliency map. The omnidirectional saliency map can also be called a celestial sphere.

[0060] The saliency map recording unit 18 records the panoramic saliency map generated by the saliency map generating unit 17 in the storage unit 68 (see FIG. 18 ). For example, a buffer or the like for recording the panoramic saliency map may be configured.

[0061] The recognition position estimation unit 19 estimates the recognition position of the recognition target object recognized by the user 5 in the out-of-field area of ​​the virtual space S that is not included in the field of view 7 of the user 5. The object to be recognized is an object that is assumed to be recognized by the user 5 among the objects in the 6DoF content that the user 5 is viewing. For example, an object that the user 5 has viewed up to the current time is set as a recognition target object. In other words, an object that has been a rendering target up to the current time is set as a recognition target object. Alternatively, an object that has not been viewed up to the current time but whose sound is recognized by the user may be set as a recognition target object. For example, if the volume of the sound emitted by the object is greater than a predetermined reference value (threshold value), the object that emitted the sound may be recognized by the user 5 and set as a recognition target object. In this case, the type or content of the sound may be used as a condition for whether or not to set the object as a recognition target object.

[0062] The recognition position is defined by XYZ coordinate values ​​(world coordinates) in the virtual space S shown in Figure 3. The position where the user 5 would currently recognize the object to be recognized as being located in his or her mind is estimated as the recognition position. Therefore, the recognition position does not necessarily match the position where the object to be recognized actually exists in the 6DoF content. The recognized position can also be said to be a grasped position that the user 5 is aware of.

[0063] FIG. 5 is a schematic diagram for explaining a recognition target object and a recognition position. 5, for ease of understanding, the panoramic image SP of the virtual space S is expressed as an image that is long in the horizontal direction. In practice, the panoramic image SP is often saved as an equirectangular image. In the following description, the predicted user's field of view 7 will be simply referred to as the user's field of view 7. The predicted frame image 20 will be described as the frame image 20.

[0064] 5, three people P1 to P3 and a blinking lighting device L are placed in a virtual space S. In addition, tree, grass, road, and building objects are also placed. At the timing shown in FIG. 5A, the field of view 7 of the user 5 is first directed toward the person P1 on the right. Therefore, a frame image 20 including the person P1 is rendered. At this point, the person P1 is set as an object to be recognized. Then, the position (world coordinates) in the virtual space S where the person P1 is located is estimated as the recognition position. In this way, it is possible to estimate the recognition position based on the position of the object to be recognized within the two-dimensional video data in which the object to be recognized is rendered.

[0065] 5A, the persons P2 and P3 on the left and the blinking lighting device L object on the right are located in the out-of-field area 21 that is not included in the field of view 7 of the user 5. Therefore, the persons P2 and P3 and the lighting device L are not rendered. As a result, although the persons P2 and P3 and the lighting device L are located in the virtual space S, they are determined not to be recognized by the user 5, and are not set as objects to be recognized.

[0066] At the timing shown in FIG. 5B, it is assumed that the field of view 7 of the user 5 is turned to the left. The people P2 and P3 are included in the field of view 7 of the user 5 and are rendered. As a result, the people P2 and P3 are set as objects to be recognized. Then, the positions (world coordinates) in the virtual space S where the people P2 and P3 are located are estimated as recognition positions. Although the person P1 is out of the field of view 7 of the user 5 and is positioned in the out-of-field area 21, the setting as an object to be recognized is maintained because the person P1 is an object that has been rendered in the past. Since the lighting device L is not rendered, it is not set as a recognition target object.

[0067] 5C, it is assumed that the person P1 moves in the out-of-visual-field area 21 toward the lighting device L. In this case, the position of the person P1 moves in the virtual space S. Here, it is assumed that the user 5 is not aware of the movement of the person P1 in the out-of-visual-field area 21. In this case, the perceived position of the person P1 in the user 5's mind does not match the actual position of the person P1. The recognition position estimation unit 19 can estimate the recognition position in the brain of the user 5 with high accuracy in response to the movement of the recognition target object in the out-of-field area 21 as shown in FIG. 5C.

[0068] In this embodiment, the rendering unit 14 functions as an embodiment of a rendering unit according to the present technology. The recognition position estimation unit 19 functions as an embodiment of an estimation unit according to the present technology. The saliency map generator 17 functions as an embodiment of a generator according to the present technology. The prediction unit 13 functions as an embodiment of a prediction unit according to the present technology.

[0069] The client device 3 includes a communication unit 23, a decoding unit 24, and a rendering unit 25. These functional blocks are realized by, for example, a CPU executing a program according to the present technology, and the information processing method according to the present embodiment is executed. Note that dedicated hardware such as an IC (integrated circuit) may be used as appropriate to realize each functional block.

[0070] The communication unit 23 is a module for executing network communication, short-distance wireless communication, etc. with other devices. For example, a wireless LAN module such as WiFi, or a communication module such as Bluetooth (registered trademark) is provided. The decoding unit 24 executes a decoding process on the distribution data, thereby decoding the encoded rendering video 8 (predicted frame image 20). The rendering unit 25 performs rendering processing so that the decoded rendered video 8 (predicted frame image 20) can be displayed by the HMD 2.

[0071] [Prediction accuracy of head motion information] For example, the server device 4 receives the "Head Motion information at the current time" and generates predicted Head Motion information for the future corresponding to the response delay (T_m2p time). Then, a predicted frame image 20 is generated based on the predicted Head Motion information and displayed to the user 5 by the HMD 2. If predicted head motion information can be generated with extremely high accuracy, it will be possible to display a rendered image 8 according to the user's 5 field of view 7 at a time corresponding to the response delay (T_m2p time) from the "current time," and the problem of response delay can be sufficiently suppressed.

[0072] The present inventors have conducted extensive research into head motion prediction in order to improve the accuracy of predicted head motion information. First, the prediction error of head motion prediction tends to increase as the frequency of the head movement signal (sensing result) increases. Due to the characteristics of the human body, rapid changes in rotational movement (high-frequency movement) are possible, but high-frequency movements with sudden changes in position such as forward / backward, up / down, and left / right tend to be difficult to make. Therefore, of these two types of motion, the prediction error for positional movements (X, Y, Z) is low and has very little impact on viewing. On the other hand, the prediction error for rotational movements (yaw, pitch, roll) tends to be large and can easily affect viewing. In other words, improving the prediction accuracy for rotational movements (yaw, pitch, roll) is extremely important.

[0073] The inventors have focused on a panoramic saliency map to improve the accuracy of head motion prediction, especially for rotational movements (yaw, pitch, roll). By generating a high-accuracy panoramic saliency map and using it for head motion prediction, it becomes possible to perform prediction accuracy for rotational movements (yaw, pitch, roll) with very high accuracy.

[0074] [2D video data (rendered video) generation] An example of the operation of generating a rendering video using the panoramic saliency map by the server device 4 will be described. FIG. 6 is a flowchart showing an example of generating a rendered video. FIG. 7 is a diagram for explaining the flowchart shown in FIG. 6, and is a schematic diagram showing the timing of obtaining head motion information, generating predicted head motion information, generating a predicted frame image 20, and generating a panoramic saliency map. In this embodiment, for ease of explanation, it is assumed that visual field information is acquired from the client device 3 at a predetermined frame rate, and predicted head motion information, predicted frame images 20, and a panoramic saliency map are generated at the same frame rate. Of course, the present invention is not limited to such processing. The numbered boxes in Fig. 7 indicate the frames of each process. Fig. 7 shows a schematic diagram of the first frame, from which the process starts, to the 25th frame. In addition, for each frame, a square shape indicates that the data shown on the left side was acquired / generated. The number inside the square shape indicates which frame the data corresponds to.

[0075] First, it is set how far into the future from the "current time" the predicted head motion information is to be generated. In this embodiment, the communication unit 16 measures the network delay with the client device 3 and identifies the predicted time of the target (step 101). That is, the response delay (T_m2p time) is measured, and the T_m2p time is identified as the predicted time. In this embodiment, head motion information for a frame that is a predetermined number of frames in the future than the frame corresponding to the "current time" is predicted and generated as predicted head motion information. The predetermined number of frames is set to the number of frames equivalent to the predicted time T_m2p. For example, in this embodiment, it is assumed that head motion information for five frames ahead is predicted. For example, when "head motion information at the current time" is acquired in the tenth frame, head motion information for the fifteenth frame, which is five frames ahead, is predicted and generated as predicted head motion information. Of course, the specific number of frames is not limited and may be set arbitrarily.

[0076] The communication unit 16 acquires head motion information from the client device 3 (step 102). As shown in Fig. 7, head motion information is acquired at a predetermined frame rate starting from the first frame. The head motion information acquired for each frame is used as is as data corresponding to that frame.

[0077] The prediction unit 13 determines whether or not the amount of head motion information necessary for predicting head motion information has been accumulated (step 103). In this embodiment, it is assumed that 10 frames of head motion information are required to predict head motion information, but the number of frames is not limited to a specific number and may be set arbitrarily. For example, for frames 1 to 9, the amount of head motion information required to predict head motion information has not been accumulated, so the answer in step 103 is No and the process returns to step 102. Therefore, generation of the rendered video 8 (predicted frame image 20) is not executed until the 10th frame. When the head motion information for the tenth frame is acquired, it is determined that the amount of head motion information required for predicting head motion information has been accumulated, and the answer to step 103 is Yes, and the process proceeds to step 104 .

[0078] In step 104, the prediction unit 13 determines whether or not a panoramic saliency map corresponding to the "head motion information at the current time" acquired in step 102 has already been generated. In this embodiment, predicted visual field information (predicted head motion information) is generated using as input historical information of visual field information (head motion information) up to the current time and a panoramic saliency map corresponding to the current time. The omnidirectional saliency map corresponding to the current time is a omnidirectional saliency map generated in the past. Specifically, the omnidirectional saliency map includes a saliency map for the visual field that indicates the saliency of the predicted frame image 20 generated based on predicted visual field information (predicted head motion information) predicted in the past, and a saliency map for an out-of-field area that is not included in the visual field 7 (predicted visual field) of the user 5 based on predicted visual field information (predicted head motion information) predicted in the past.

[0079] In the example shown in FIG. 7, the omnidirectional saliency map corresponding to the "head motion information at the current time" refers to the omnidirectional saliency map corresponding to the frame in which the "head motion information at the current time" is acquired. That is, when the number inside the square shape indicating the Head Motion information and the number inside the square shape indicating the omnidirectional saliency map are equal, a pair of corresponding "Head Motion information at the current time" and omnidirectional saliency map is formed.

[0080] For example, when Head Motion information for the 10th frame is acquired, the frame corresponding to the current time is the 10th frame. In step 104, it is determined whether a panoramic saliency map corresponding to the 10th frame (a panoramic saliency map represented by a square with the number 10 written inside) has been generated. 7, up to the tenth frame, predicted head motion information has not been generated, and predicted frame image 20 has not been generated either. Therefore, since a panoramic saliency map has not been generated, step 104 becomes No, and the process proceeds to step 105.

[0081] In step 105, the prediction unit 13 generates predicted visual field information (predicted head motion information) based on history information of visual field information (head motion information) up to the current time. In this way, when a panoramic saliency map for a frame corresponding to the current time has not been generated, predicted head motion information may be generated based only on the history information of head motion information up to the current time. In this embodiment, at frame 10, predicted head motion information for the future, that is, five frames ahead, is generated based on the history information of head motion information from frame 1 to frame 10. Therefore, as shown in Fig. 7, at the 10th frame, predicted head motion information corresponding to frame 15, five frames ahead, is generated (predicted head motion information represented by a square with the number 15 written inside). There are no particular limitations on the specific algorithm for generating predicted head motion information based on the history information of head motion information up to the current time, and any algorithm may be used. For example, any machine learning algorithm may be used.

[0082] 3 is executed by the rendering unit 14 based on the predicted head motion information, and a rendered video 8 (predicted frame image 20) is generated (step 106). In this embodiment, the predicted frame image 20 corresponding to 15 frames is generated based on the predicted head motion information for the next five frames.

[0083] The recognition position estimation unit 19 estimates the recognition position of the recognition target object in the out-of-field area (step 107). In this embodiment, the recognition position of the recognition target object recognized by the user 5 is estimated in the out-of-field area that is not included in the field of view 7 (predicted field of view) of the user 5 based on the predicted field of view information (predicted head motion information).

[0084] The saliency map generator 17 generates a panoramic saliency map corresponding to the 15 frames based on the predicted frame image 20 and the estimated recognition positions (step 108). The generated panoramic saliency map is recorded and held by the saliency map recording unit 18. As illustrated in Fig. 7, for the 10th frame, a panoramic saliency map corresponding to the 15th frame is recorded.

[0085] 7, in this embodiment, for the frame corresponding to the "current time", a frame image five frames ahead is generated as a predicted frame image 20. Also, for the frame corresponding to the "current time", a panoramic saliency map for five frames ahead is generated. In the present disclosure, a frame image generated at the "current time," i.e., a frame image generated in a frame corresponding to the "current time," is referred to as the "frame image at the current time." Therefore, in this embodiment, the future predicted frame image 20 generated in a frame corresponding to the "current time" corresponds to the "frame image at the current time." On the other hand, the "frame image corresponding to the current time (predicted frame image)" corresponds to a frame image (predicted frame image) generated five frames in the past.

[0086] The encoding unit 15 encodes the predicted frame image 20. The communication unit 16 then transmits the encoded predicted frame image 20 to the client device 3 (step 109). The predicted frame image 20 generated for the tenth frame is transmitted to the HMD 2 via the client device 3 as the first frame of the 6DoF video content, and is displayed to the user 5. This starts the distribution of a virtual video with the effects of response delays sufficiently suppressed. The rendering unit 14 determines whether or not processing has been completed for all frame images (step 110). Here, it is assumed that processing is performed up to frame 25, as shown in the example of FIG. Therefore, step 110 is negative and the process returns to step 102.

[0087] From frame 11 to frame 14 shown in FIG. 7, step 104 is answered as No, and the processing flow proceeds from step 105 to step 106. At frame 15, there exists a omnidirectional saliency map corresponding to frame 15 that was generated in the past frame 10 as the omnidirectional saliency map corresponding to the acquired "Head Motion information at the current time." Therefore, step 104 is Yes, and the process proceeds to step 111.

[0088] In step 111, future head motion information is predicted using the history information of visual field information (head motion information) up to the current time and the omnidirectional saliency map corresponding to the current time as input, and is generated as predicted head motion information. There are no particular limitations on the specific algorithm for generating predicted head motion information using the history information of head motion information and the omnidirectional saliency map as input, and any algorithm may be used. For example, any machine learning algorithm may be used. From then on, up to frame 25, step 104 returns Yes and the panoramic saliency map is used to generate highly accurate predicted Head Motion information. If the processing for all frame images is complete, step 109 returns Yes, and the video generation and distribution processing ends.

[0089] [Considerations on generating omnidirectional saliency maps] The inventors have given extensive consideration to the generation of a panoramic saliency map. Human visual attention can be divided into two types: exogenous attention (bottom-up attention) triggered by visual stimuli before an object is recognized, and endogenous attention (top-down attention) triggered by interest in the object after object recognition. The keyword "saliency" is used in both bottom-up and top-down attention.

[0090] An example of generating a saliency map based on bottom-up attention is to extract from the input video (2D image) each feature amount that attracts exogenous attention (bottom-up attention) through visual stimuli before a person recognizes an object, such as brightness, color, direction, direction of motion, and depth.The final saliency map is generated by calculating each feature map and integrating them so that high saliency is assigned to areas where the value of each feature amount is significantly different from the surroundings.

[0091] FIG. 8 is a schematic diagram showing an example of bottom-up attention-based saliency map generation that can be performed by the present system. In the example shown in FIG. 8, a predicted frame image 20 is input as an input frame. A feature extraction process is performed on the predicted frame image 20, and the features of brightness, color, and direction that attract bottom-up attention are extracted. Note that the predicted frame image 20 of the previous frame, etc., may be used for feature extraction. For each feature amount of brightness, color, and direction, a feature image is generated in which the feature amount is converted into brightness, and a Gaussian pyramid of the feature image is generated. Furthermore, a depth map and a motion vector map image are acquired as parameters (rendering information) related to the rendering process from a renderer constituting the rendering unit 14. The depth map is data including distance information (depth information) to an object to be rendered. The motion vector map image is data including motion information of the object to be rendered. The depth map image is used as a depth feature image to generate a Gaussian pyramid, and the motion vector map image is used as a motion direction feature image to generate a Gaussian pyramid. Center-surround subtraction is performed on each Gaussian pyramid of features, generating feature maps for brightness, color, orientation, motion direction, and depth. These feature maps are then combined to generate a bottom-up attention-based saliency map 22. The specific algorithms for the feature extraction process, Gaussian pyramid generation process, center-surround difference process, and feature map integration process for each feature are not limited. For example, each process can be realized using well-known techniques.

[0092] The depth map image obtained from the renderer is not a depth value estimated by performing 2D image analysis or the like on the predicted frame image 20, but an accurate value obtained in the rendering process. Therefore, by receiving this depth map image directly from the renderer and using it as "depth" feature information in generating the saliency map 22, it becomes possible to generate a more accurate saliency map 22 with higher precision. The motion vector map image obtained from the renderer is not an estimated value obtained by performing 2D image analysis or the like on the predicted frame image 20, but an accurate value obtained in the rendering process. Therefore, by receiving this depth map image directly from the renderer and using it as feature information on the "motion direction" to generate the saliency map 22, it becomes possible to generate a more accurate saliency map with higher precision. It should be noted that other feature quantities such as "brightness" and "color" can also be calculated in the rendering process and used as rendering information. The algorithm for generating a saliency map based on bottom-up attention is not limited, and any other algorithm may be used, for example, any machine learning algorithm. For example, feature information on "depth" and "direction of movement" may be acquired by performing 2D image analysis on the predicted frame image 20, and this information may be used.

[0093] Top-down attention is directed after object recognition based on its meaning, so salience is given to the object. For example, objects that generally attract people's attention, such as human faces, are detected from an image and marked with salience. In addition, salience is assigned to the display area of ​​an object based on the type of object (idol, baseball player, vehicle, etc.), the importance of the object in the 6DoF content, the user's preference for the object (degree of interest or preference, etc.), etc. The algorithm for generating the saliency map based on top-down attention is not limited, and any algorithm may be used, for example, any machine learning algorithm.

[0094] It is also possible to generate a saliency map that includes both bottom-up and top-down attention saliency. For example, by capturing actual human gaze data, creating training data based on that data, and training the model, it is possible to build a machine learning model that can generate a saliency map that includes both bottom-up and top-down attention.

[0095] [Comparative Example of a Method for Generating a Panoramic Saliency Map] Here, a method for generating a panoramic saliency map will be described as a comparative example. A 2D image is rendered for each viewport while changing the viewport at regular intervals across the entire virtual space S (hereinafter, this rendered 2D image will be referred to as a viewport image). Then, for each viewport image, a saliency map based on bottom-up attention or top-down attention as described above is generated and integrated. The panoramic saliency map of this comparative example is information that represents the saliency of all viewport images for the user 5. In other words, it is information that represents the saliency of all areas in the virtual area S when they are within the field of view of the user 5.

[0096] Therefore, when the panoramic saliency map of this comparative example is used, there is a high possibility that the following problems will occur. Fig. 9 is a schematic diagram for explaining a problem with the panoramic saliency map 30 of the comparative example. In Fig. 9, the saliency based on the top-down attention generated by the persons P1 to P3 and the saliency based on the bottom-up attention generated by the lighting device L are schematically illustrated as white regions.

[0097] (1) The saliency resulting from bottom-up attention is not seen outside the field of view of user 5, and therefore there is no visual stimulation and it is essentially zero. The omnidirectional saliency map 30 of the comparative example is created on the assumption that all areas in virtual space S are within the field of view. Therefore, if the omnidirectional saliency map 30 of the comparative example is used as is to predict visual field information, saliency that is outside the field of view and cannot be recognized by user 5 will have a negative impact on the prediction of visual field information. For example, at the timing shown in Fig. 5A, the blinking lighting device L on the right side is not recognized by the user 5. In the omnidirectional saliency map 30 of the comparative example, as shown in Fig. 9A, saliency based on bottom-up attention is assigned to the pixel region of the blinking lighting device L. As a result, saliency that does not exist in the brain of the user 5 is generated, which adversely affects the prediction of visual field information.

[0098] (2) Even in the case of saliency resulting from top-down attention, the saliency of objects that are outside the field of view of the user 5 and have not yet been captured (visually recognized) in the viewport is essentially zero because the user 5 is not even aware of their existence. When the omnidirectional saliency map 30 of the comparative example is used, saliency that is outside the field of view and cannot be recognized by the user 5 adversely affects the prediction of visual field information. For example, at the timing shown in Fig. 5A, the people P2 and P3 on the left side are not recognized by the user 5. In the omnidirectional saliency map 30 of the comparative example, as shown in Fig. 9A, the people P2 and P3 are assigned saliency based on top-down attention. As a result, saliency that does not exist in the brain of the user 5 is generated, which adversely affects the prediction of visual field information.

[0099] (3) Furthermore, in the case of salience resulting from top-down attention, an object previously viewed may be moving in the out-of-field area 21. In this case, the user 5 will not notice the object moving. Therefore, the user will likely recognize that the object is in a position that corresponds to the situation he or she perceived when he or she last viewed the object. For example, if an object was stationary, we would likely recognize it as being in the position where we last saw it, and if the object was moving when we last saw it, we would likely recognize it as being some distance along the direction of movement from where we last saw it. The omnidirectional saliency map 30 of the comparative example is based on the assumption that moving objects and objects after they have moved are also within the field of view, and therefore saliency based on top-down attention is assigned to moving objects and objects after they have moved. For example, if an object was stationary when last seen, i.e., if the user is not aware that the object has moved since then, salience will occur in a location completely different from the location of recognition (grasping location) in the brain. If the object was moving when last seen, there is no problem if the perceived position (perceived position) in the brain matches the actual position of the object. On the other hand, it is not necessarily likely that the position the user expects to be and the actual position of the object will match. Therefore, there is a high possibility that saliency will occur in a position different from the perceived position (perceived position) in the brain. For example, at the timing shown in Fig. 5C, user 5 does not recognize the movement of person P1. In the omnidirectional saliency map 30 of the comparative example, as shown in Fig. 9B, saliency based on top-down attention is assigned to person P1 near lighting device L1. As a result, saliency that does not exist in the brain of user 5 is generated, which adversely affects the prediction of visual field information. As described above, the omnidirectional saliency map 30 of the comparative example does not take into consideration saliency outside the viewport (outside the field of view). Therefore, if the omnidirectional saliency map 30 of the comparative example is used as is for predicting Head Orientation, which is the next direction to which the viewport will be directed from the current viewport, the prediction accuracy may actually decrease or may be useless, thereby failing to achieve the purpose of improving prediction accuracy. To achieve high prediction accuracy, it is important to accurately reflect the actual attention state of the user 5 with respect to both inside and outside the visual field.

[0100] [Estimation of the recognition position of the object to be recognized] In this embodiment, the recognition target object is set and the recognition target position of the recognition target object is estimated by the recognition position estimation unit 19. These processes are executed for each frame. Hereinafter, several examples of estimating the recognition position of a recognition target object will be described. Example 1 The object that is the subject of rendering by the rendering unit 14 is set as the object to be recognized. For example, at the timing shown in Fig. 5A, a person P1 is set as the object to be recognized, and at the timing shown in Fig. 5B, persons P1 and P2 are set as the objects to be recognized.

[0101] In each frame, each time a "current time frame image" (a future predicted frame image 20 generated at the current time) is generated, the recognition position of each set object to be recognized is estimated. For a recognition target object included in the "frame image at the current time", the recognition position is estimated based on the position in the virtual space S that corresponds to the position of the recognition target object in the frame image. For example, at the timing shown in Figure 5A, person P1 is included in frame image (predicted frame image) 20, so the recognized position of person P1 is estimated based on the position in virtual space S that corresponds to the position of person P1 in frame image 20.

[0102] The position in the virtual space S corresponding to the position of the person P1 in the frame image 20 is the position of the person P1 placed in the virtual space S when rendering the frame image 20. The position of the person P1 is estimated as the recognized position recognized by the user 5. If a state in which the person P1 is moving has been rendered, that is, if the user 5 recognizes that the person P1 is moving, a position in the virtual space S corresponding to the position of the person P1 in the frame image 20, that is, a position shifted along the direction of movement from the position of the person P1 placed in the virtual space S, may be estimated as the recognized position recognized by the user 5. The amount of shift may be set appropriately to an amount that is thought to be the amount by which the user 5 would predict the destination. Here, it is assumed that the position in the virtual space S corresponding to the position of the person P1 in the frame image 20 is directly estimated as the recognized position.

[0103] For example, at the timing shown in FIG. 5B, the positions in the virtual space S corresponding to the positions of the people P2 and P3 in the frame image 20 are estimated as the recognized positions recognized by the user 5 with respect to the people P2 and P3.

[0104] For a recognition target object that is not included in the "frame image at the current time," the recognition position is estimated based on the position in the virtual space S that corresponds to the position of the recognition target object in the most recent past frame image 20 that includes the recognition target object. In this embodiment, the position in the virtual space S that corresponds to the position of the recognition target object in the most recent frame image 20 is directly estimated as the recognition position. The most recent past frame image can also be said to be the last frame image rendered.

[0105] For example, at the timing shown in Fig. 5B, person P1 is not included in frame image 20. For person P1 not included in frame image 20, the position in virtual space S corresponding to the position of person P1 in frame image 20 at the timing shown in Fig. 5A is estimated as the recognized position. At the timing shown in Fig. 5C, in the out-of-field area 21, person P1 is moving toward the lighting device L. In this embodiment, the position in the virtual space S corresponding to the position of person P1 in the most recent frame image 20, i.e., the frame image 20 at the timing shown in Fig. 5A, is maintained as the recognized position. This makes it possible to prevent the occurrence of saliency that does not exist in the brain of the user 5, and makes it possible to predict visual field information with high accuracy.

[0106] For example, every time a recognition target object is rendered, the position in virtual space S corresponding to the position of the recognition target object in the rendered frame image 20 is updated and stored as the recognition position. As a result, for a recognition target object that is not included in the frame image 20 at the current time, the position in virtual space S corresponding to the position of the recognition target object in the most recent frame image 20 is held as the recognition position. In this way, the recognition position may be updated for each frame.

[0107] Example 2 Even if the user 5 moves an object out of his or her field of vision, he or she may be able to hear sounds, such as footsteps or voices, emitted from the object and thereby determine the object's location. This second embodiment is designed with such a situation in mind. FIG. 10 is a schematic diagram showing an example of information described in a scene description file used as scene description information in this embodiment. In this embodiment, when 6DoF content is generated, audio data such as footsteps and voices emitted from the objects is generated in association with each object information described in the scene description file. Based on this information, the renderer will render the associated audio data so that it plays from the position of each object.

[0108] In the example shown in FIG. 10, the following information is stored as object information: Name: Name of the object Position: Object position Url: Address of 3D object data Audio...The name of the audio data for the sound emitted from the object In the example shown in FIG. 10, the following information is stored as audio data information linked to object information: Name: Name of the audio data Url: Address of the audio data

[0109] 10, in a hide-and-seek scene, audio data information for "Voice and Footsteps 1" and "Voice and Footsteps 2" is linked to video object information for "Hiding Person 1" and "Hiding Person 2." Examples of the audio data for "Voice and Footsteps 1" and "Voice and Footsteps 2" include replies such as "Not yet" or "That's enough" to the Oni's call, "Are you ready?", as well as the footsteps of someone running away from the Oni.

[0110] The recognition position estimation unit 19 receives from the renderer information on the presence or absence of audio data associated with each object, current volume information, etc. Then, it determines whether the user 5 is currently hearing and recognizing the sound emitted from the object based on whether the volume level exceeds a reference value (threshold) that is a volume level used as a reference for determining whether the user 5 can hear and recognize the sound. The reference volume level may be set arbitrarily. If there is associated audio data and the volume level exceeds a reference level, the current position of the object in virtual space S, i.e., the position where the sound originates in virtual space S, is estimated as the recognized position recognized by user 5. That is, in this embodiment, when it is determined that the user 5 recognizes the sound emitted by the object to be recognized that is not included in the "frame image at the current time", the position where the sound is generated in the virtual space S is estimated as the recognized position.

[0111] For example, in a scene configured based on the scene description file shown in Fig. 10, the recognized positions of people playing hide-and-seek are estimated based on the voices of the people playing hide-and-seek. For example, suppose a voice saying "It's not yet here" is uttered from "Hiding Person 1," who is not included in the frame image. If the volume of the voice exceeds a reference value, it is determined that the voice was heard by user 5, and the current position of "Hiding Person 1" in virtual space S is estimated as the recognized position. This makes it possible to prevent the occurrence of saliency that does not exist in the brain of the user 5, and makes it possible to predict visual field information with high accuracy. When the sound emitted from the object to be recognized is interrupted and can no longer be heard, the position where the sound was last heard, that is, the position where the sound was most recently heard in the past, is maintained as the recognition position.

[0112] Example 3 If the user 5 has previously viewed a scene similar to the scene currently being viewed, the user 5 may grasp the position of an object by associating it with his or her memory of that scene. This embodiment 3 is devised assuming such a case, and the recognized position in the brain of the user 5 is estimated based on past viewing information of a scene similar to the current scene. 11 and 12 are schematic diagrams showing an example of information described in a scene description file used as scene description information in this embodiment. In this embodiment, when generating 6DoF content, the object information described in the scene description file stores the role information of each object in the current scene and its fixed position information (world coordinates) when playing that role. That is, the scene description file stores role information that indicates the role of the object to be recognized and a fixed position (world coordinates) related to the role.

[0113] In the examples shown in FIGS. 11 and 12, the following information is stored as object information: Name: Name of the object Position: Object position Url: Address of 3D object data Role...Role information FixedPos...fixed position information for roles

[0114] In the examples shown in Figures 11 and 12, role information and regular position information for players "Ada Ao" and "Bgawa Bsuke" during offense and defense are stored in a baseball scene. Figure 11 is a scene description file during offense, and Figure 12 is a scene description file during defense. Every time the offense and defense switch positions, the attacking scene description file and the defensive scene description file are updated. Also, if the role of an object changes between offense and defense, the attacking and defensive scene description files are updated accordingly. For example, when the scene description file is updated from Figure 11 to Figure 12 in response to a change of offense and defense, player "A-da A-san" changes from the role of "next batter" to the role of "first baseman," and player "B-gawa B-suke" changes from the role of "batter" to the role of "pitcher." The FixedPos of the "batter," "next batter," "pitcher," and "first baseman" are assumed to be positions in world coordinates for the following positions. "Batter"...the batter's box "Next batter"...Next batter "Pitcher"...pitcher's mound "First baseman"...first base position These FixedPos indicate the general position in that role, and are different from Position, which indicates the actual position of the object. Based on this information, the recognition position estimation unit 19 determines whether the user 5 has previously seen a scene similar to the current scene, and estimates the recognition position.

[0115] 13 and 14 are schematic diagrams for explaining estimation of the recognition position of the recognition target object in the third embodiment. 13 and 14, a baseball stadium is configured as the virtual space S, and a player "Ada A-san" 32 and a player "Bgawa B-suke" 33 are also placed in the virtual space S. The direction of the line of sight of the eyes in the figure corresponds to the field of view (predicted field of view) 7 of the user 5, and a frame image (predicted frame image) 20 of the area of ​​the field of view 7 is generated.

[0116] In Fig. 13A, first, user 5 watches a scene in which player "A-da A-husband" 32 plays first base. That is, user 5 understands that player "A-da A-husband" 32 is at first base in the defensive scene. Of course, player "A-da A-husband" 32 is set as a recognition target object. After that, as the offense and defense switch, player "A-da A-san" 32 becomes the "next batter" and moves to the next batter's circle. As shown in Fig. 13B, user 5 follows player "A-da A-san" 32 with his eyes as he moves to the next batter's circle. 14A, the user 5 turns his field of view 7 toward the player "Bgawa Bsuke" 33 who is the "batter" and watches the batting. At this time, the player "Ada Ao" 32, who is the object to be recognized, is outside the field of view and is not included in the frame image 20. After that, the offense and defense switch, and player "A-da A-san" 32 moves to the first base position as the "first baseman." Meanwhile, as shown in FIG. 14B, user 5 turns his field of view 7 toward the spectator seats and watches the spectators in the cheering section. Player "A-da A-san" 32 moves to first base in an area outside the field of view of user 5, and user 5 does not watch this movement. User 5 can understand that, due to the change of offense and defense, player "A-da A-husband" 32 is now playing defense. Furthermore, based on the past viewing experience of FIG. 13A, user 5 knows that player "A-da A-husband" 32 is at first base when playing defense. Therefore, it can be assumed that user 5 associates player "A-da A-husband" 32 with the position of first base in his mind, just as when viewing FIG. 13A, and updates his mental position. In accordance with this update in the brain, the perceived position estimation unit 19 estimates the perceived position of the player "A-da A-husband" 32 to the "position of first base," which is a fixed position associated with the role of "first baseman." That is, based on the fact that the user 5 is viewing a scene in which the role of the player "A-da A-husband" 32 is "first baseman" in the viewing in Fig. 13A, just as in the present (the player "A-da A-husband" 32 is rendered in the frame image 20), and based on the scene update from the viewing in Fig. 14A to the viewing in Fig. 14D, the role of the player "A-da A-husband" 32 has again become "first baseman," the perceived position estimation unit 19 estimates the perceived position of the player "A-da A-husband" 32 in the user 5's brain to the "position of first base." In this way, in this third embodiment, role information of the object to be recognized in the scene at the time of past viewing (rendering) is retained, and if the object to be recognized is currently out of the field of view, the role of the object is updated in a scene update, and if there is viewing experience at the time of that role, the recognition position is estimated to be the fixed position of that role.

[0117] In the view in FIG. 14B, the player "A-ta A-san" 32 corresponds to a recognition target object to which predetermined role information ("first baseman") that is not included in the "frame image at the current time" is set. Then, if the player "A-ta A-fu" 32 with the same role information ("first baseman") has been rendered up to the current time, the fixed position associated with the role ("first base position") is estimated as the recognized position. This makes it possible to prevent the occurrence of saliency that does not exist in the user's brain, and makes it possible to predict visual field information with high accuracy.

[0118] The scene in which the role of the player "A-da A-san" 32 is "first baseman" may be viewed in another past baseball game. That is, even if a scene in which the role of the player "A-da A-san" 32 is "first baseman" is viewed not only in the game currently being watched but also in another game watched in the past, the "position of first base" may be estimated as the recognized position. That is, as long as a recognition target object to which the same role information is set has been rendered up to the current time, it is possible to estimate a fixed position related to the role as the recognition position.

[0119] In the past, when a state in which the player "A-ta A-fu" 32 was in his regular position "first base position" was rendered, the "first base position" may be estimated as the recognized position. In other words, if a recognition target object (player "A-ta A-fu" 32) with the same role information ("first baseman") has been rendered in a fixed position related to the role ("first base position") up to the current time, the fixed position related to the role ("first base position") may be estimated as the recognition position. This makes it possible to estimate the recognition position when the user is sure that the player "A-ta A-fu" 32 is at the "first base position" during defense. On the other hand, since the recognition target object with role information set is often at its "regular position" in most cases, it is highly likely that it will be rendered as being at its "regular position."

[0120] It is also possible to integrate the processes of (Example 1) to (Example 3) to estimate the recognition position of the recognition target object. For example, the processes are executed with priority in the order of (Example 1), (Example 2), and (Example 3). First, the information with the highest priority is determined to be visual information from the eyes, and (Example 1) is executed. That is, the position where the user 5 is thought to have recognized the recognition target object with his / her eyes is estimated as the recognition position. When the user 5 is visually recognizing the recognition target object, the position seen with the eyes becomes the recognition position of the recognition target object. The recognition position estimation unit 19 determines whether the user 5 is visually recognizing the recognition target object based on whether the recognition target object has been rendered in a frame image (predicted frame image) 20, and estimates the position of the recognition target object in the virtual space S at the time of rendering as the recognition position. In this case, since no other information is required to estimate the recognition position, the sound information used in (Example 2) and the role information and fixed position information used in (Example 3) are not acquired. When the object to be recognized moves out of the field of view 7 (i.e., is no longer rendered), the sound information used in (Example 2) and the role information and fixed position information used in (Example 3) are acquired. When this information is not available, the last viewed position (the position at the time of last rendering) is maintained as the recognition position.

[0121] The next highest priority information is determined to be sound information emitted from the object to be recognized, and (Example 2) is executed. That is, the recognition position is estimated based on the presence or absence of audio data linked to the object to be recognized and the current generation status information of the audio data (generation position, volume information, etc.). When the object to be recognized is out of the field of view and there is no visual information, the position where the object is thought to have been recognized by sound is estimated as the recognized position.

[0122] As the next highest priority information, the role information and the fixed position information are acquired, and (Example 3) is executed. That is, the recognition position is estimated based on the information on whether the user 5 has had a past viewing experience in a scene similar to the current scene and the fixed position information of the object to be recognized in that scene. When neither visual information nor sound information is acquired, a position that is thought to be associated with a position based on past viewing experiences of a scene similar to the current one is estimated as the recognized position. If there is no visual information, sound information, role information, or fixed position information, user 5 is unaware of the existence of the object, and therefore pays zero attention to the object (there is no setting of a recognition target object, and there is no recognition position).

[0123] Fig. 15 is a flowchart showing an example of estimating the recognition position of a recognition target object. The process shown in Fig. 15 can be said to be an example of a process that combines the processes of (Example 1) to (Example 3). Rendering information and scene information for all objects to be recognized in a scene are obtained (step 201). The rendering information includes any information related to the rendering of the recognition target object. Here, the rendering information includes rendering history information of the recognition target object up to the current time, position information of the recognition target object within the frame image 20, etc. The scene information includes scene description information relating to the object to be recognized, such as history information of the scene description information up to the current time. In step 301, an object that is rendered for the first time in the "frame image at the current time" is also set as an object to be recognized, and rendering information and frame information are acquired.

[0124] It is determined whether there are any unprocessed objects to be recognized (step 202). If there are any unprocessed objects to be recognized (Yes in step 202), one of the unprocessed objects to be recognized is selected, and the processing from step 203 onwards is executed.

[0125] It is determined whether the selected object to be recognized is included in the "frame image at the current time" (i.e., whether it has been rendered) (step 203). If the object to be recognized is included in the "frame image at the current time" (Yes in step 203), the current role information of the object to be recognized is added to the role list (step 204). The role list is a list into which role information is input when a recognition target object with role information set thereto has been viewed (i.e., has been rendered) up to the current time. If role information has not been set for the recognition target object, it is not added to the role list. The role list can also be called a role viewed list. The recognition position is estimated to be the current position of the object to be recognized (step 205). Here, the current position of the object to be recognized corresponds to the position in the virtual space S that corresponds to the position of the object to be recognized in the "frame image at the current time". In step 205, the initial recognition position of an object that is rendered for the first time in the "frame image at the current time" is estimated. The recognition position of the object to be recognized, whose recognition position has been estimated up to the current time, is updated. Of course, there may be cases where the result is the same as the recognition position estimated in the past. The estimation of the recognition position of this object to be recognized is now complete, and the process returns to step 202.

[0126] If the object to be recognized is not included in the "frame image at the current time" (No in step 203), it is determined whether audio data is associated with the object to be recognized and whether the volume at the time of current rendering exceeds the reference value (step 206). If step 206 is positive (Yes in step 206), the current role (role information) of the object to be recognized is added to the role list (step 204). In this way, in this embodiment, even if it is determined that the user 5 has recognized a sound emitted by a recognition target object to which the same role information is set up until the current time, the role is added to the role list in the same manner as when a recognition target object to which the same role information is set is rendered. The recognition position is estimated to be the current position of the recognition target object (step 205). Here, the current position of the recognition target object corresponds to the generation position of the sound emitted from the recognition target object in the virtual space S. The recognition position of the recognition target object whose estimated position has been estimated up to the current time is updated. The estimation of the recognition position of this object to be recognized is now complete, and the process returns to step 202.

[0127] If step 206 is negative (No in step 206), it is determined whether the current role of the object to be recognized has changed since the last time the recognition position was updated (step 207). If the current role has not changed since the last time the recognized position was updated (No in step 207), the recognized position is not updated (i.e., the recognized position is unchanged), and the process returns to step 202.

[0128] If the current role has changed since the last time the recognition position was updated (Yes in step 207), it is determined whether the user 5 has previously viewed a scene in which the object to be recognized has the current role (step 208). The determination in step 208 is made by referring to whether the current role information of the object to be recognized has been entered in the role list. If the current role information of the object to be recognized has been entered in the role list, the result in step 208 is affirmative. If the current role information of the object to be recognized has not been entered in the role list, the result in step 208 is negative.

[0129] If the user 5 has not previously viewed a scene in which the object to be recognized plays the current role (No in step 208), the recognition position is not updated (i.e., the recognition position remains unchanged).Then, the process returns to step 202. If the user 5 has previously viewed a scene in which the object to be recognized currently plays a role (Yes in step 208), the recognition position is estimated to be the position of the object to be recognized currently (step 209). Here, the current position of the object to be recognized corresponds to a fixed position related to the role. The recognition position of the object to be recognized, whose estimated position has been estimated in the past frame up to the current time, is updated. The estimation of the recognition position of this object to be recognized is now complete, and the process returns to step 202.

[0130] If there are no unprocessed recognition target objects (No in step 202), the process of estimating the recognition positions of all recognition target objects ends (step 210). As shown in FIG. 15, by using visual information, current audio information, and past audiovisual information, it is possible to estimate the recognition position of the object to be recognized with high accuracy.

[0131] In the estimation example shown in FIG. 15, when it is determined that the user 5 has recognized a sound emitted by a recognition target object set with predetermined role information that is not included in the “frame image at the current time” up until the current time, a fixed position related to the role is estimated as the recognition position. Alternatively, if it is determined that the user 5 has recognized a sound made by an object to be recognized that has the same role information set up until the current time and that is in a fixed position related to the role, the fixed position related to the role may be estimated as the recognized position.

[0132] [Generating a panoramic saliency map] The generation of the panoramic saliency map by the saliency map generation unit 17 will be described. FIG. 16 is a flowchart showing an example of generating a panoramic saliency map. First, in step 301, a saliency map for the visual field is generated based on the "frame image at the current time." In this embodiment, a saliency map for the visual field that includes both saliency based on bottom-up attention and saliency based on top-down attention is generated. For example, saliency based on bottom-up attention and saliency based on top-down attention are detected separately, and then these are added together to generate a saliency map for the field of view corresponding to the final viewport image. Also in the same step 301, a panoramic saliency map of the sky (in this embodiment, an equirectangular image in which all pixels have a value of zero) is prepared, and the saliency map for the field of view is pasted at the location corresponding to the viewport. This makes it possible to generate a panoramic saliency map that does not generate saliency based on bottom-up attention in areas outside the visual field. That is, it is possible to generate a panoramic saliency map in which saliency based on bottom-up attention is zero in areas outside the visual field. As a result, it is possible to prevent unnecessary saliency from occurring, and it is possible to solve the above-mentioned problem point (1). Note that any other method may be used to avoid generating saliency based on bottom-up attention in the out-of-field region. For example, a method may be adopted in which a saliency map for the entire sky (including both saliency based on bottom-up attention and saliency based on top-down attention) is generated once, and then parts of the out-of-field region are masked. On the other hand, according to the method of this embodiment, in which a saliency map for the entire field of view is pasted into the field of view area of ​​the sky omnidirectional saliency map, it is possible to reduce the processing load and shorten the processing time.

[0133] The recognition positions of all objects to be recognized in the out-of-view area estimated by the recognition position estimation unit 19 are acquired (step 302). It is determined whether or not there is an object to be recognized in an unprocessed out-of-field area (hereinafter referred to as an out-of-field object) (step 303). If there are unprocessed out-of-view objects (Yes in step 303), one unprocessed out-of-view object is selected and the process in step 304 is executed.

[0134] In step 304, the position of the out-of-field object in the omnidirectional saliency map (position on the 2D map) is calculated based on the recognized position of the out-of-field object in the virtual space S. The top-down based saliency generated by the out-of-field object is placed at the calculated position. An out-of-view object is an object that has been rendered in the past, and therefore has had its saliency detected in the past using top-down attention in step 301. For example, the saliency value of each pixel along the shape of the object is detected. In this embodiment, the saliency based on the top-down attention for the object to be recognized detected in step 301 is retained. Then, in step 304, the retained saliency (shape and value) based on the top-down attention is reused and arranged on the omnidirectional saliency map. The top-down attention-based saliency placement of this out-of-view object based on its perceived location is now complete, and we return to step 303.

[0135] In addition, in step 304, a method may be adopted in which, after generating a saliency map (top-down attention) for the entire sky, the position at which saliency occurs in the out-of-field area is adjusted to match the estimated recognition position. On the other hand, by reusing the saliency based on the top-down attention detected in step 101, as in this embodiment, rendering processing can be performed only on the viewport, reducing the processing load and the processing time.

[0136] If there are no unprocessed out-of-field objects (No in step 302), the process of generating an omnidirectional saliency map in which saliency based on top-down attention of out-of-field objects is generated based on the recognized position ends.

[0137] Salience based on top-down attention is generated at the recognition position recognized in the brain of the user 5. That is, a saliency map representing saliency based on top-down attention in the out-of-field area is generated based on the recognition position of the object to be recognized in the out-of-field area. This makes it possible to prevent the generation of unnecessary saliency from locations different from the recognition location in the brain, thereby solving the above-mentioned problems (2) and (3).

[0138] Fig. 17 is a schematic diagram showing an example of a panoramic saliency map generated by this embodiment. Fig. 17A is an example of a panoramic saliency map 35 generated at the timing shown in Fig. 5A. Fig. 17B is an example of a panoramic saliency map 35 generated at the timing shown in Fig. 5C. 17A, in the omnidirectional saliency map 35 generated at the timing shown in FIG. 5A, saliency based on the top-down attention of only the person P1 recognized by the user 5 is generated. No saliency based on the bottom-up attention of the lighting device L not recognized by the user 5 is generated. Furthermore, no saliency based on the top-down attention of the persons P1 and P2 not recognized by the user 5 is generated. 17B, ​​for user 5 who is not aware of the movement of person P1, salience based on the top-down attention of person P1 occurs at the position of person P1 before the movement. Furthermore, salience based on the bottom-up attention of lighting device L, which is not recognized by user 5, does not occur. In this way, the panoramic saliency map 35 generated by this embodiment avoids the occurrence of saliency that does not exist in the brain of the user 5, resulting in a highly accurate panoramic saliency map.

[0139] As described above, the server-side rendering system 1 according to this embodiment estimates the recognition position of the recognition target object in the out-of-field area 21. Based on the estimated recognition position, a omnidirectional saliency map 35 including a saliency map in the out-of-field area 21 is generated. This makes it possible to reflect the attention of the user 5 to the outside of the field of view according to the viewing situation at that time in the omnidirectional saliency map 35. As a result, it becomes possible to generate a highly accurate omnidirectional saliency map 35. Since a highly accurate and precise omnidirectional saliency map 35 is generated, it becomes possible to generate predicted head motion information (particularly orientation information) with extremely high accuracy, and it becomes possible to sufficiently suppress the problem of response delay (T_m2p time). In other words, it becomes possible to realize the delivery of high-quality virtual video using the omnidirectional saliency map 35. The high-precision panoramic saliency map 35 generated in this embodiment can also be used for other purposes.

[0140] <Other embodiments> The present technology is not limited to the above-described embodiments, and various other embodiments can be realized.

[0141] The above example shows a case where 6DoF video is distributed as a virtual image. However, this technology is not limited to this and can also be applied to cases where 3DoF video, 2D video, etc. are distributed. Furthermore, AR video, etc., may be distributed as a virtual image instead of VR video. This technology can also be applied to stereo images (for example, right-eye and left-eye images) for viewing 3D images. The present technology can be applied to content that displays any virtual space where an out-of-field area may occur. Furthermore, the saliency map of the out-of-field area is not limited to the saliency map of the entire virtual space, but may also be generated for a portion of the virtual space that becomes an out-of-field area.

[0142] FIG. 18 is a block diagram showing an example of the hardware configuration of a computer (information processing device) 60 that can realize the server device 4 and the client device 3. The computer 60 includes a CPU 61, a ROM (Read Only Memory) 62, a RAM 63, an input / output interface 65, and a bus 64 that interconnects these components. The input / output interface 65 is connected to a display unit 66, an input unit 67, a storage unit 68, a communication unit 69, a drive unit 70, and the like. The display unit 66 is a display device using, for example, a liquid crystal display, an electroluminescent display, etc. The input unit 67 is, for example, a keyboard, a pointing device, a touch panel, or other operating device. When the input unit 67 includes a touch panel, the touch panel can be integrated with the display unit 66. The storage unit 68 is a non-volatile storage device such as a HDD, flash memory, or other solid-state memory. The drive unit 70 is a device capable of driving a removable storage medium 71 such as an optical storage medium or magnetic recording tape. The communication unit 69 is a modem, router, or other communication device that can be connected to a LAN, WAN, or the like and that communicates with other devices. The communication unit 69 may communicate either wired or wirelessly. The communication unit 69 is often used separately from the computer 60. Information processing by the computer 60 having the above-described hardware configuration is realized by cooperation between software stored in the storage unit 68, the ROM 62, etc. and the hardware resources of the computer 60. Specifically, the information processing method according to the present technology is realized by loading a program constituting the software stored in the ROM 62, etc., into the RAM 63 and executing the program. The program is installed in the computer 60 via, for example, a recording medium 61. Alternatively, the program may be installed in the computer 60 via a global network or the like. Any other computer-readable non-transitory storage medium may be used.

[0143] An information processing device according to the present technology may be constructed by a plurality of computers connected to each other so as to be able to communicate via a network or the like working together to execute an information processing method and program according to the present technology. That is, the information processing method and program according to the present technology can be executed not only in a computer system configured by a single computer, but also in a computer system in which multiple computers operate in conjunction with each other. In this disclosure, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all the components are in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems. The information processing method and program execution according to the present technology by a computer system include both cases where, for example, obtaining visual field information, performing rendering processing, setting a recognition target object, estimating a recognition position, generating a panoramic saliency map, etc. are performed by a single computer, and cases where each process is performed by a different computer. Furthermore, the execution of each process by a specific computer includes having another computer execute part or all of the process and obtaining the results. In other words, the information processing method and program according to the present technology can also be applied to a cloud computing configuration in which a single function is shared and processed jointly by multiple devices via a network.

[0144] The configurations and processing flows of the server-side rendering system, HMD, server device, client device, etc. described with reference to the drawings are merely one embodiment and can be modified as desired without departing from the spirit of the present technology. In other words, any other configurations, algorithms, etc. for implementing the present technology may be adopted.

[0145] In this disclosure, to facilitate understanding of the explanation, words such as "approximately," "almost," and "roughly" are used as appropriate. However, there is no clear difference between using and not using words such as "approximately," "almost," and "roughly." That is, in the present disclosure, concepts that define shape, size, positional relationship, state, etc., such as "center," "central," "uniform," "equal," "same," "orthogonal," "parallel," "symmetrical," "extended," "axial direction," "cylindrical," "cylindrical," "ring-shaped," and "annular," are concepts that include "substantially center," "substantially central," "substantially uniform," "substantially equal," "substantially the same," "substantially orthogonal," "substantially parallel," "substantially symmetrical," "substantially extended," "substantially axial direction," "substantially cylindrical," "substantially cylindrical," "substantially ring-shaped," "substantially annular," and the like. For example, this also includes states that fall within a specified range (for example, a range of ±10%) based on criteria such as "perfectly centered," "perfectly central," "perfectly uniform," "perfectly equal," "perfectly the same," "perfectly perpendicular," "perfectly parallel," "perfectly symmetrical," "perfectly extended," "perfectly axial," "perfectly cylindrical," "perfectly cylindrical," "perfectly ring-shaped," and "perfectly annular." Therefore, even if the words "roughly," "almost," "approximately," etc. are not added, it may include concepts that can be expressed by adding "roughly," "almost," "approximately," etc. Conversely, a state expressed by adding "roughly," "almost," "approximately," etc. does not necessarily exclude a complete state.

[0146] In this disclosure, expressions using "more than," such as "greater than A" and "smaller than A," are expressions that comprehensively include both concepts that include equivalent to A and concepts that do not include equivalent to A. For example, "greater than A" is not limited to cases that do not include equivalent to A, but also includes "A or greater." Furthermore, "smaller than A" is not limited to "less than A," but also includes "A or less." When implementing the present technology, specific settings and the like may be appropriately adopted from the concepts included in "greater than A" and "smaller than A" so as to achieve the effects described above.

[0147] It is also possible to combine at least two of the features of the present technology described above. That is, the various features described in each embodiment may be arbitrarily combined without distinction between the embodiments. Furthermore, the various effects described above are merely examples and are not limiting, and other effects may also be achieved.

[0148] The present technology can also be configured as follows. (1) a rendering unit that generates two-dimensional video data according to the user's field of view by performing a rendering process on three-dimensional space data that constitutes a virtual space based on field of view information regarding the user's field of view; an estimation unit that estimates a recognition position of a recognition target object recognized by the user in an out-of-field area of ​​the virtual space that is not included in the user's field of view; a generation unit that generates a saliency map representing saliency in the out-of-field area based on the estimated recognition position of the recognition target object in the out-of-field area; An information processing device comprising: (2) The information processing device according to (1), The estimation unit sets an object that has been a rendering target up to the current time as the recognition target object. Information processing device. (3) The information processing device according to (1) or (2), the two-dimensional video data is composed of a plurality of frame images that are successive in time series; The estimation unit estimates a recognition position of the recognition target object that is not included in the frame image at the current time based on a position in the virtual space that corresponds to a position of the recognition target object in the most recent past frame image that includes the recognition target object. Information processing device. (4) The information processing device according to (3), The estimation unit estimates, as the recognition position, a position in the virtual space corresponding to a position of the object to be recognized in the most recent frame image. Information processing device. (5) The information processing device according to (3) or (4), The estimation unit estimates, as the recognition position, a position shifted along a movement direction of the object to be recognized from a position in the virtual space corresponding to the position of the object to be recognized in the most recent frame image. Information processing device. (6) An information processing device according to any one of (3) to (5), When it is determined that the user has recognized a sound emitted by the recognition target object that is not included in the frame image at the current time, the estimation unit estimates the generation position of the sound in the virtual space as the recognition position. Information processing device. (7) An information processing device according to any one of (3) to (6), the three-dimensional space data includes three-dimensional space description data that defines a configuration of the virtual space and three-dimensional object data that defines three-dimensional objects in the virtual space; the three-dimensional space description data includes role information representing a role of the object to be recognized, and fixed position information representing a fixed position associated with the role; The estimation unit estimates, for the recognition target object to which predetermined role information that is not included in the frame image at the current time is set, the fixed position related to the role as the recognition position when the recognition target object to which the same role information is set has been rendered up to the current time. Information processing device. (8) The information processing device according to (7), The estimation unit estimates the fixed position related to the role as the recognition position when a state in which the recognition target object set with the same role information is at the fixed position related to the role has been rendered up to the current time. Information processing device. (9) An information processing device according to any one of (3) to (8), the three-dimensional space data includes three-dimensional space description data that defines a configuration of the virtual space and three-dimensional object data that defines three-dimensional objects in the virtual space; the three-dimensional space description data includes role information representing a role of the object to be recognized, and fixed position information representing a fixed position associated with the role; When it is determined that the user has recognized a sound emitted by the recognition target object to which predetermined role information that is not included in the frame image at the current time has been set, the estimation unit estimates the fixed position related to the role as the recognition position. Information processing device. (10) The information processing device according to (9), When it is determined that the user has recognized a sound emitted when the recognition target object, to which the same role information has been set up until the current time, is in the fixed position related to the role, the estimation unit estimates the fixed position related to the role as the recognition position. Information processing device. (11) An information processing device according to any one of (1) to (10), The estimation unit estimates the recognition position based on a position of the recognition target object in the two-dimensional video data in which the recognition target object is rendered. Information processing device. (12) An information processing device according to any one of (1) to (11), The generating unit generates the saliency map in which the saliency based on bottom-up attention in the out-of-field region is zero. Information processing device. (13) An information processing device according to any one of (1) to (12), The generation unit generates the saliency map representing saliency based on top-down attention in the out-of-field region based on the recognition position of the recognition target object in the out-of-field region. Information processing device. (14) The information processing device according to any one of (1) to (13), The generation unit generates the saliency map in the out-of-field area and a saliency map representing the saliency of the two-dimensional video data. Information processing device. (15) The information processing device according to any one of (1) to (14), further comprising: a prediction unit that generates future visual field information as predicted visual field information based on the saliency map; The rendering unit generates the two-dimensional video data based on the predicted field of view information. Information processing device. (16) The information processing device according to (15), The field of view information includes at least one of a position of a viewpoint, a direction of a line of sight, a rotation angle of a line of sight, a position of the user's head, or a rotation angle of the user's head. Information processing device. (17) The information processing device according to (16), the field of view information includes a rotation angle of the user's head; The prediction unit predicts a future head rotation angle of the user based on the saliency map. Information processing device. (18) The information processing device according to any one of (15) to (17), the two-dimensional video data is composed of a plurality of frame images that are successive in time series; The rendering unit generates a frame image based on the predicted view information and outputs it as a predicted frame image. Information processing device. (19) generating two-dimensional video data according to the user's field of view by performing a rendering process on three-dimensional space data constituting a virtual space based on field of view information relating to the user's field of view; Estimating a recognition position of a recognition target object recognized by the user in an out-of-field area of ​​the virtual space that is not included in the user's field of view; A saliency map representing saliency in the out-of-field area is generated based on the estimated recognition position of the object to be recognized in the out-of-field area. An information processing method implemented by a computer system. [Explanation of symbols]

[0149] S...Virtual space 1. Server-side rendering system 2...HMD 3...Client device 4. Server equipment 5...User 7...User's field of view 8...Rendered image 13...Prediction Department 14...Rendering section 17...Saliency map generation unit 18...Saliency map recording section 19...Recognition position estimation unit 20...Predicted frame image (frame image) 35...Full-dome saliency map 60...Computer

Claims

1. a rendering unit that generates two-dimensional video data according to the user's field of view by performing a rendering process on three-dimensional space data that constitutes a virtual space based on field of view information regarding the user's field of view; an estimation unit that estimates a recognition position of a recognition target object that is assumed to be recognized by the user as existing in the virtual space among objects existing in the virtual space, in an out-of-field area of ​​the virtual space that is not included in the user's field of view; and a generation unit that generates a saliency map representing saliency in the out-of-field area based on the estimated recognition position of the recognition target object in the out-of-field area; An information processing device comprising:

2. 2. The information processing device according to claim 1, The estimation unit sets an object that has been a rendering target up to the current time as the recognition target object. Information processing device.

3. 2. The information processing device according to claim 1, the two-dimensional video data is composed of a plurality of frame images that are successive in time series, The estimation unit estimates a recognition position of the recognition target object that is not included in the frame image at the current time based on a position in the virtual space that corresponds to a position of the recognition target object in the most recent past frame image that includes the recognition target object. Information processing device.

4. 4. The information processing device according to claim 3, The estimation unit estimates, as the recognition position, a position in the virtual space corresponding to a position of the object to be recognized in the most recent frame image. Information processing device.

5. 4. The information processing device according to claim 3, The estimation unit estimates, as the recognition position, a position shifted along a movement direction of the object to be recognized from a position in the virtual space corresponding to the position of the object to be recognized in the most recent frame image. Information processing device.

6. 4. The information processing device according to claim 3, When it is determined that the user has recognized a sound emitted by the recognition target object that is not included in the frame image at the current time, the estimation unit estimates the generation position of the sound in the virtual space as the recognition position. Information processing device.

7. 4. The information processing device according to claim 3, the three-dimensional space data includes three-dimensional space description data that defines a configuration of the virtual space and three-dimensional object data that defines three-dimensional objects in the virtual space; the three-dimensional space description data includes role information representing a role of the object to be recognized, and fixed position information representing a fixed position associated with the role; The estimation unit estimates, for the recognition target object to which predetermined role information that is not included in the frame image at the current time is set, the fixed position related to the role as the recognition position when the recognition target object to which the same role information is set has been rendered up to the current time. Information processing device.

8. 8. The information processing device according to claim 7, The estimation unit estimates the fixed position related to the role as the recognition position when a state in which the recognition target object set with the same role information is at the fixed position related to the role has been rendered up to the current time. Information processing device.

9. 4. The information processing device according to claim 3, the three-dimensional space data includes three-dimensional space description data that defines a configuration of the virtual space and three-dimensional object data that defines three-dimensional objects in the virtual space; the three-dimensional space description data includes role information representing a role of the object to be recognized, and fixed position information representing a fixed position associated with the role; When it is determined that the user has recognized a sound emitted by the recognition target object to which predetermined role information that is not included in the frame image at the current time has been set, the estimation unit estimates the fixed position related to the role as the recognition position. Information processing device.

10. 10. The information processing device according to claim 9, When it is determined that the user has recognized a sound emitted when the recognition target object, to which the same role information has been set up until the current time, is in the fixed position related to the role, the estimation unit estimates the fixed position related to the role as the recognition position. Information processing device.

11. 2. The information processing device according to claim 1, The estimation unit estimates the recognition position based on a position of the recognition target object in the two-dimensional video data in which the recognition target object is rendered. Information processing device.

12. 2. The information processing device according to claim 1, The generating unit generates the saliency map in which the saliency based on bottom-up attention in the out-of-field region is zero. Information processing device.

13. 2. The information processing device according to claim 1, The generation unit generates the saliency map representing saliency based on top-down attention in the out-of-field region based on the recognition position of the recognition target object in the out-of-field region. Information processing device.

14. 2. The information processing device according to claim 1, The generation unit generates the saliency map in the out-of-field area and a saliency map representing saliency of the two-dimensional video data. Information processing device.

15. The information processing device according to claim 1, further comprising: a prediction unit that generates future visual field information as predicted visual field information based on the saliency map; The rendering unit generates the two-dimensional video data based on the predicted view information. Information processing device.

16. 16. The information processing device according to claim 15, The field of view information includes at least one of a position of a viewpoint, a direction of a line of sight, a rotation angle of a line of sight, a position of the user's head, or a rotation angle of the user's head. Information processing device.

17. 17. The information processing device according to claim 16, the field of view information includes a rotation angle of the user's head; The prediction unit predicts a future head rotation angle of the user based on the saliency map. Information processing device.

18. 16. The information processing device according to claim 15, the two-dimensional video data is composed of a plurality of frame images that are successive in time series, The rendering unit generates a frame image based on the predicted view information and outputs it as a predicted frame image. Information processing device.

19. generating two-dimensional video data according to the user's field of view by performing a rendering process on three-dimensional space data constituting a virtual space based on field of view information relating to the user's field of view; Estimating a recognition position of a recognition target object that is assumed to be recognized by the user as existing in the virtual space among objects existing in the virtual space, in an out-of-field area of ​​the virtual space that is not included in the user's field of view, and that is assumed to be recognized by the user at the current time; A saliency map representing saliency in the out-of-field area is generated based on the estimated recognition position of the object to be recognized in the out-of-field area. An information processing method implemented by a computer system.

Citation Information

Patent Citations

  • Video data accuracy adjustment method and adjustment system linked to video signal processing capability of terminal means

    JP2007520925A

  • Method and apparatus for determining and varying the panning speed of an image based on saliency

    US20180189928A1