Information processing device and information processing system

The information processing apparatus addresses user-specific sensory differences by using reinforcement learning to optimize foveated rendering, improving immersion and reducing load in VR, AR, and MR systems.

JP2025103571APending Publication Date: 2025-07-09DENSO TEN LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023221038
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-07-09

AI Technical Summary

Technical Problem

Existing four-view rendering processes in VR, AR, and MR technologies fail to account for individual sensory differences among users, leading to a potential reduction in user immersion due to inappropriate rendering settings.

Method used

An information processing apparatus that includes a controller to perform foveated rendering, adjusting parameters based on the physical and mental state of the user through reinforcement learning to optimize the rendering process.

Benefits of technology

The apparatus ensures an appropriate four-view rendering process tailored to the user, enhancing immersion while reducing processing load by adapting to individual sensory differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025103571000001_ABST
    Figure 2025103571000001_ABST
Patent Text Reader

Abstract

To provide an information processing device and an information processing system that can perform foveated rendering processing appropriate for each user when playing back content.SOLUTION: An information processing device according to an embodiment includes a controller that executes a foveated rendering processing on a content. The controller changes characteristics of the foveated rendering processing in response to a mental and physical state of a user viewing the reproduced content.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus and an information processing system.

Background Art

[0002] Conventionally, technologies for providing digital content including virtual space experiences such as VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality) to users using an HMD (Head Mounted Display) or the like are known. Such digital content is also referred to as XR (Cross Reality) content. XR is an expression that summarizes all virtual space technologies including VR, AR, MR, as well as SR (Substitutional Reality), AV (Audio / Visual), and the like.

[0003] In addition, a technology for performing foveated rendering processing on content when the content is reproduced is also known. The foveated rendering processing, for example, renders only the area in front of the user's line of sight at a high resolution and reduces the resolution around such an area, thereby suppressing a decrease in the user's sense of immersion while reducing the processing load on the information processing apparatus.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, since there are individual sensory differences among users who view the content, depending on the content of the four-view rendering process, there is a risk that the user's sense of immersion may be excessively reduced. That is, the four-view rendering process in the prior art was not tailored to the user.

[0006] The present invention has been made in view of the above, and an object thereof is to provide an information processing apparatus and an information processing system capable of making the four-view rendering process appropriate according to the user in the reproduction of content.

Means for Solving the Problems

[0007] In order to solve the above problems and achieve the object, the information processing apparatus according to the present invention includes a controller that executes a four-view rendering process on content. The controller changes the characteristics in the four-view rendering process according to the physical and mental state of the user who views the reproduced content.

Effects of the Invention

[0008] According to the present invention, in the reproduction of content, the four-view rendering process can be made appropriate according to the user.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of the information processing apparatus and the information processing system disclosed in the present application will be described in detail with reference to the accompanying drawings. Note that the present invention is not limited by the embodiments described below.

[0011] First, with reference to FIGS. 1 and 2, an overview of the information processing system including the information processing apparatus according to the embodiment will be described. FIG. 1 is a diagram showing an overview of the information processing system. FIG. 2 is a diagram showing the data flow in the information processing system. Hereinafter, the case where the XR space (virtual space) is a VR space will be described.

[0012] As shown in FIG. 1, the information processing system 1 includes an information processing apparatus 10, a display device 3, a speaker 4, and a biosensor 5.

[0013] As shown in FIG. 2, the information processing apparatus 10 provides video data to the display device 3. Further, the information processing apparatus 10 provides audio data to the speaker 4. Further, the information processing apparatus 10 acquires the biological data of the user U from the biosensor 5.

[0014] As shown in FIG. 1, the display device 3 is, for example, a head-mounted display. The display device 3 displays an image according to the video signal of the content output from the information processing device 10. Specifically, the display device 3 is an information processing terminal that presents video data related to XR content provided from the information processing device 10 to the user U and allows the user to enjoy the XR experience. Note that the display device 3 may be a non-transmissive type that completely covers the field of view, or may be a video transmissive type or an optical transmissive type.

[0015] The speaker 4 outputs sound according to the audio signal of the content output from the information processing device 10. Specifically, the speaker 4 generates audio data related to XR content provided from the information processing device 10 as sound. The speaker 4 may be, for example, a headphone type worn on the ear of the user U, or a box type installed on the floor or the like. Also, the speaker 4 may be of a stereo audio or multi-channel audio type. Note that the speaker 4 is an example of an audio output device.

[0016] The biological sensor 5 detects biological data of the user U. Specifically, the biological sensor 5 includes a first biological sensor 5a, a second biological sensor 5b, and a third biological sensor 5c. In the following, when the biological sensors 5a to 5c are not particularly distinguished and described, they are referred to as "biological sensor 5".

[0017] The first biological sensor 5a is provided in the display device 3. The first biological sensor 5a detects, as biological data, for example, the direction of the user U's line of sight, the direction of the user U's head, etc. Also, the first biological sensor 5a may detect the user U's blink as biological data. Note that the first biological sensor 5a is realized by, for example, a camera, a motion sensor, etc., but is not limited thereto.

[0018] The second biological sensor 5b is a headgear-type electroencephalogram sensor. The second biological sensor 5b detects the electroencephalogram of the user U as biological data. The third biological sensor 5c is a wristband-type pulse sensor. The third biological sensor 5c detects the pulse of the user U as biological data. The biological sensor 5 outputs the detected biological data to the information processing device 10 via wired communication or wireless communication.

[0019] Note that the position where the biological sensor 5 is provided in FIG. 1 is merely an example and is not limited. Also, in the example of FIG. 1, an example where the first to third biological sensors 5a to 5c are separate bodies is shown, but it is not limited to this, and a part or all of the first to third biological sensors 5a to 5c may be integrated.

[0020] The information processing device 10 is composed of a computer. The information processing device 10 is connected to the display device 3 and the speaker 4 via wired communication or wireless communication. The information processing device 10 provides the video of the XR content to the display device 3. The information processing device 10 provides the audio of the XR content to the speaker 4. Thereby, the XR content is reproduced on the display device 3 and the speaker 4.

[0021] Also, the information processing device 10 acquires the changes in the biological data detected by, for example, the first biological sensor 5a at any time, and reflects such changes in the XR content. For example, the information processing device 10 can change the direction of the field of view in the virtual space of the XR content according to the changes in the orientation of the user U's head and the direction of the line of sight detected by the first biological sensor 5a.

[0022] When reproducing the XR content, the information processing device 10 performs foveated rendering processing on the XR content, in other words, central fovea rendering processing, in order to adjust the user U's immersion in the XR content and reduce the processing load of the information processing device 10. Note that hereinafter, the foveated rendering processing may be described as "FR processing".

[0023] Here, the FR process will be described with reference to FIG. 3. FIG. 3 is a diagram for explaining the FR process. Note that FIG. 3 shows the image A of the XR content displayed on the display device 3.

[0024] In the FR process, for example, only the region B1 at the tip of the line of sight of the user U is rendered at a high resolution. Also, in the FR process, in the region B2 around such a region B1, rendering is performed so as to reduce the resolution, in other words, rendering is performed at a low resolution. Hereinafter, the region B1 may be referred to as the "high-resolution region B1" and the region B2 may be referred to as the "low-resolution region B2". In FIG. 3, for the sake of convenience of understanding, the low-resolution region B2 is shaded to indicate that it is a region where the resolution has decreased. Also, in FIG. 3, an example in which the resolution in the FR process is in two stages of high resolution and low resolution is shown, but it is not limited to this, and the resolution may be three or more stages, or the resolution may be continuously changed.

[0025] In the information processing apparatus 10, by performing the FR process in which the tip of the line of sight of the user U is set as the high-resolution region B1 and the periphery thereof is set as the low-resolution region B2, it is possible to reduce the processing load of the information processing apparatus 10 while adjusting the sense of immersion of the user U.

[0026] However, there are individual sensory differences among users U who view XR content. Note that individual sensory differences include, for example, differences in the visual fields of each user U and the way of moving the line of sight. Due to such individual sensory differences of the user, depending on the content of the FR process, there was a risk that the sense of immersion of the user U would be excessively reduced. For example, there was a risk that the user U would be concerned about the presence of the low-resolution region B2 or the degree of low resolution (roughness) in the low-resolution region B2, and the sense of immersion and presence when viewing the XR content would be excessively reduced. That is, the FR process in the prior art was not adapted to the sensory differences of the user U.

[0027] Therefore, the information processing apparatus 10 according to the present embodiment is configured such that the FR processing can be made appropriate according to the user U in the reproduction of XR content. Specifically, while reproducing the XR content, the information processing apparatus 10 performs learning (reinforcement learning) according to the physical and mental state of the user U to optimize various parameters in the FR processing. In other words, the various parameters are made to correspond to the user.

[0028] The configuration of this information processing apparatus 10 will be specifically described with reference to FIG. 4 and the like. FIG. 4 is a block diagram showing a configuration example of the information processing apparatus 10 according to the embodiment. In the block diagram of FIG. 4, only the components necessary for explaining the features of the present embodiment are represented as functional blocks, and the description of general components is omitted.

[0029] In other words, each component shown in the block diagram of FIG. 4 is a functional concept, and it is not necessarily physically configured as shown in the figure. For example, the specific form of the distribution and integration of each functional block is not limited to that shown in the figure, and all or part of it can be functionally or physically distributed and integrated in any unit according to various loads and usage situations.

[0030] As shown in FIG. 4, the information processing apparatus 10 includes a controller (control unit) 20 and a storage unit 30.

[0031] The storage unit 30 is realized by, for example, a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk. In the example of FIG. 4, content information 31, parameter information 32, an emotion estimation model 33, and a reinforcement learning model 34 are stored in the storage unit 30. In addition, various data and various programs are stored in the storage unit 30.

[0032] Content information 31 is information regarding XR content to be displayed on the display device 3. Here, the content information 31 will be described with reference to FIG. 5. FIG. 5 is an explanatory diagram showing an example of the content information 31.

[0033] As shown in FIG. 5, in the content information 31, information on items such as "content name", "genre", "emotion label", and "biometric data target value" is associated with the "content ID".

[0034] The "content ID" is identification information for identifying XR content. The "content name" is information indicating the name of the XR content. As an example, when the XR content is content such as a movie or a game, information indicating the movie name or game name is stored in the "content name". In the example shown in FIG. 5, for convenience, the "content name" is described abstractly as "D1" or the like, but it is assumed that specific information is stored in "D1". Hereinafter, other information may also be described abstractly.

[0035] The "genre" is information indicating the genre of the XR content. Information such as horror, comedy, action, suspense, and healing is stored in the "genre".

[0036] The "emotion label" is information indicating the emotion of the user U that the XR content aims for. In other words, the "emotion label" is information indicating the emotion (assumed emotion) that the user U who has viewed the XR content is assumed to have (hold). To put it another way, the "emotion label" is information indicating the emotion expected of the user U who has viewed the XR content. Information such as horror, fun, anxiety, and relaxation is stored in the "emotion label". Thus, in the XR content according to the present embodiment, information indicating the emotion of the user U that the XR content aims for is given as an emotion label.

[0037] The "biometric data target value" is information indicating the target value of the biometric data of user U that the XR content aims for. In other words, the "biometric data target value" is information indicating the biometric data expected for user U who views the XR content. The "biometric data target value" stores information such as, for example, the heart rate increase rate (or decrease rate), and the decrease rate of the number of blinks (the number of blinks in an appropriate period set based on experiments, etc.). Thus, in the XR content according to this embodiment, information indicating the biometric data of user U that the XR content aims for is given as the biometric data target value.

[0038] Note that the number of pieces of information stored in the "genre", "emotion label", and "biometric data target value" may be one or plural. Also, the "genre", "emotion label", and "biometric data target value" may be set for each scene of the XR content. In this case, a scene ID is assigned to each scene, and the "genre", "emotion label", and "biometric data target value" are stored in association with the scene ID. Then, the parameters in the FR processing described later are adjusted for each scene.

[0039] In the example of FIG. 5, the content information identified by the content ID "C01" indicates that the content name is "D1", the genre is "horror", the emotion label is "fear", and the biometric data target value is "E1".

[0040] Returning to the description of FIG. 4, the parameter information 32 is information regarding the parameters in the FR processing. Such parameters are an example of the characteristics in the FR processing. Here, the parameter information 32 will be described with reference to FIG. 6. FIG. 6 is an explanatory diagram showing an example of the parameter information 32.

[0041] As shown in FIG. 6, the parameter information 32 is information on items such as "intensity", "processing area", and "parameter change period". Using each piece of data of this parameter information 32, the controller 20 performs FR processing. In other words, when each parameter value such as "intensity" is written as parameter information 32 in the storage unit 30 (for example, when each parameter value is stored in the parameter information storage area of the storage unit 30), the controller 20 uses these parameter values stored in the storage unit 30 to perform FR processing.

[0042] "Intensity" is information indicating the intensity (level) of FR processing. In FR processing, as the intensity increases, the resolution in the low-resolution area B2 (see FIG. 3) decreases, and accordingly, the processing load of the information processing apparatus 10 decreases. Conversely, in FR processing, as the intensity decreases, the resolution in the low-resolution area B2 increases, and accordingly, the processing load of the information processing apparatus 10 increases. Note that in the above, an example where the resolution of the low-resolution area B2 changes depending on the high or low intensity is shown, but it is not limited to this, and the resolution of the high-resolution area B1 (see FIG. 3) may also change.

[0043] "Processing area" is information indicating the area of the region in the video of the XR content where FR processing is executed. Specifically, the "processing area" includes the area of the high-resolution area B1 and the area of the low-resolution area B2. Since the area of the image A (see FIG. 3) displayed on the display device 3 during XR content playback is constant, when the area of the high-resolution area B1 increases, the area of the low-resolution area B2 decreases by the increased amount, and accordingly, the processing load of the information processing apparatus 10 increases. Also, when the area of the high-resolution area B1 decreases, the area of the low-resolution area B2 increases by the decreased amount, and accordingly, the processing load of the information processing apparatus 10 decreases. Note that the "processing area" may not be the value of the area itself, but a value indicating the area, for example, the radius (when the region is circular) or the side length (when the region is square or rectangular).

[0044] The "parameter change period" is information indicating the period for changing the parameters (characteristics) of the FR process. Specifically, the "parameter change period" is information indicating the timing for changing parameters by learning processes and the like described later.

[0045] In the example of FIG. 6, the parameters indicating the characteristics of the FR process show that the intensity is "F1", the processing area is "G1", and the parameter change period is "H1", and the controller 20 performs the FR process using these respective parameter values, "F1", "G1", and "H1".

[0046] Note that initial values used at the start of the FR process are stored in advance for "intensity", "processing area", and "parameter change period" respectively. Such initial values can be set to arbitrary values, and for example, general-purpose values are set based on experiments and the like.

[0047] Note that when changing the XR content, values corresponding to the content of the XR content may be set. As an example, when the XR content is "horror", "emotion label: fear", "biometric data target value: E1", the initial values of various parameters are set to values that can aim for an FR process suitable for the content of such XR content. Thereby, the controller 20 can execute an FR process corresponding to the content of the XR content by using the initial values of various parameters when changing the XR content. Therefore, in the present embodiment, it is possible to give the user U a sense of immersion, fear, etc. close to the target in the XR content even immediately after playing new XR content.

[0048] Note that the above-mentioned "intensity", "processing area", and "parameter change period" are each changed by learning (reinforcement learning) described later. Specifically, they are changed according to the physical and mental state of the user U who views the played XR content.

[0049] Returning to the description of FIG. 4, the emotion estimation model 33 is a model used for the process of estimating the emotion of the user U who views the XR content. Such an emotion estimation model 33 is stored in the storage unit 30 in advance. Here, the emotion estimation model 33 will be described with reference to FIG. 7. FIG. 7 is a diagram for explaining the emotion estimation model 33, and is a diagram showing an example of the coordinate plane of the model.

[0050] In the emotion estimation model 33 according to the present embodiment, the emotion of the user U is estimated based on the index values of a plurality of indexes based on the biological data of the user U.

[0051] Here, the index value is the value of an index related to the biological data. For example, the activity level of the autonomic nervous system (calculated by the standard deviation of the LF (Low Frequency) component of the heartbeat) and the arousal level (calculated by the β wave / α wave of the electroencephalogram) are indexes. Also, the specific value (for example, a numerical value) corresponding to each index is the index value. Note that the index value is the sensor value of the biological sensor 5 or a value calculated from the sensor value.

[0052] As shown in FIG. 7, the above-described indexes are used as the axes constituting the coordinate plane. Whether the index value is in the positive direction or the negative direction from the origin (the determination threshold of the activity level and the arousal level) of the coordinate plane has meaning. It is assumed that the index value is standardized so that the average is 0. That is, the above standardization is performed based on the estimation that no emotion associated with the index value has occurred when the index value is the average value. Note that the index value is not limited to standardization, and may be converted into a format that is easy to handle in the coordinate plane. Also, in the coordinate plane shown in FIG. 7, the vertical axis is set to "arousal - non - arousal (positive side - negative side)" and the horizontal axis is set to "strong activity level - weak activity level (positive side - negative side)".

[0053] The emotions of user U are known to appear in brain waves, heartbeats, etc. Specifically, if the measurement result (index value) of the biological sensor 5 (specifically, the second biological sensor 5b) shows the positive side that "the beta wave / alpha wave of the brain wave is large with respect to the threshold value (arousal threshold value)", it is presumed that user U may have emotions such as "happy, joyful, angry, sad, depressed" in the awakened state. Also, if the measurement result (index value) shows the negative side that "the beta wave / alpha wave of the brain wave is small with respect to the threshold value (arousal threshold value)", it is presumed that user U may have emotions such as "anxious, frightened, unpleasant, relaxed, calm" in the non-awakened state.

[0054] Also, the standard deviation of the LF component of the heartbeat is known to represent the activity level of the autonomic nervous system. On the other hand, it is considered that there is a correlation between the activity level of the autonomic nervous system and the emotion type. And as emotion types with strong activity levels, "happy, joyful, angry, sad, anxious, frightened, unpleasant" are known. Also, as emotion types with weak activity levels, "depressed, relaxed, calm" are known.

[0055] From the above, in the coordinate plane of the emotion estimation model 33, emotion types such as "happy, joyful, angry, sad" are arranged in the first quadrant 210, "depressed" in the second quadrant 220, "relaxed, calm" in the third quadrant 230, and "anxious, frightened, unpleasant" in the fourth quadrant 240. That is, by applying the arousal level and activity level calculated from the output value of the biological sensor 5 to the emotion estimation model 33, emotions can be estimated.

[0056] Returning to the explanation of FIG. 4, the reinforcement learning model 34 is a model used in the process of learning (reinforcement learning) described later. Such a reinforcement learning model 34 is stored in the memory unit 30 in advance. The reinforcement learning model 34 will be described later with reference to FIG. 10 and the like.

[0057] The controller 20 of the information processing apparatus 10 can estimate the emotion of the user U by using each index value calculated based on the biological data of the user U obtained from the biological sensor 5 and the above-described emotion estimation model 33. For example, when the arousal level (the β wave / α wave of the brain wave is in a non-aroused state lower than the threshold value) and the activity level of the autonomic nervous system (the standard deviation of the heart rate LF component) calculated based on the biological data of the user U are highly active (strong emotion) higher than the threshold value, the controller 20 estimates that emotions such as "anxiety, fear, and discomfort" are occurring in the user U.

[0058] Continuing with the description of FIG. 4, the controller 20 of the information processing apparatus 10 is realized, for example, by various programs (not shown) stored in the storage unit 30 being executed with the RAM as a work area by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or the like. Further, the controller 20 can also be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0059] While playing the XR content, the controller 20 performs learning (reinforcement learning) according to the physical and mental state of the user U to optimize various parameters in the FR process.

[0060] Reinforcement learning is performed by the controller 20, which is an "agent". Reinforcement learning has elements such as "state", "action", and "reward", and is a learning (machine learning) that is carried out through trial and error using these elements. For example, the controller 20 repeats a process of receiving, as a "reward", the change caused by a certain "action" in the current "state" of the controller 20. Note that the "state" indicates the current situation of the controller 20. The "reward" is a quantification of the change, but is not limited to this. The controller 20 determines an appropriate "action" by repeating trial and error, such as changing the "action" to change the "state" so that the received "reward" increases.

[0061] The controller 20 according to the present embodiment changes the parameters (characteristics) in the FR process as an "action". Specifically, the controller 20 performs learning to change the parameters (characteristics) in the FR process as an "action". Also, the controller 20 performs learning to receive, as a "reward", a change related to the mental and physical state of the user U (the reward increases as the difference between the target mental and physical state and the detected mental and physical state decreases). Specifically, the controller 20 performs learning to change the characteristics in the FR process according to the mental and physical state (the difference between the target mental and physical state and the actual mental and physical state of the user U) of the user U who is viewing the reproduced XR content. More specifically, the controller 20 performs learning (reinforcement learning) to change the parameters (characteristics) in the FR process as an "action" so that the "reward" indicating the change related to the mental and physical state (the difference between the target change state and the detected change state) increases. Thereby, in the present embodiment, in the reproduction of XR content, the FR process can be made appropriate according to the user U.

[0062] Here, the processing procedure of learning and the like executed by the controller 20 in the FR process will be described in detail with reference to FIG. 8. FIG. 8 is a flowchart showing the processing procedure executed by the controller 20 according to the embodiment. Note that the processing of the flowchart shown in FIG. 8 is executed when the information processing system 1 is powered on and is repeatedly executed by the controller 20.

[0063] First, the controller 20 executes XR content setting processing (step S10). Here, the XR content setting processing is setting processing for a content playback device 40 (see FIG. 10) that plays XR content. The XR content setting processing includes, for example, various processes related to each initial setting of the content playback device 40 for playing XR content, selection of XR content by the user U, and the like.

[0064] Next, the controller 20 acquires biometric data of the user U before playing the XR content from the biometric sensor 5 (step S11). Specifically, the controller 20 acquires biometric data indicating the direction of the user U's line of sight, the direction of the head, blinking, etc. from the first biometric sensor 5a. Further, the controller 20 acquires biometric data indicating the brain waves of the user U from the second biometric sensor 5b. Further, the controller 20 acquires biometric data indicating the pulse of the user U from the third biometric sensor 5c. Note that the controller 20 is not limited to acquiring all of the above-described biometric data, and may be configured to acquire a part of the biometric data.

[0065] Next, the controller 20 starts playing the XR content selected by the user U (step S12). Specifically, the controller 20 plays the XR content on the content playback device 40, outputs the video signal in the XR content to the display device 3 (see FIGS. 2 and 10), and outputs the audio signal to the speaker 4 (see FIG. 2).

[0066] When playing XR content, the controller 20 starts the FR process (step S13). Specifically, the controller 20 reads the initial values of various parameters in the parameter information 32, and uses the read parameters to perform the FR process on the video signal in the XR content. Note that the controller 20 may start the FR process at the same timing as the start of the playback of the XR content, or may start it at an arbitrary timing after the playback of the XR content. Thereby, the user U is provided with the XR content on which the FR process has been executed.

[0067] Here, the image of the XR content provided to the user U will be described with reference to FIG. 9. FIG. 9 is a diagram showing an example of the image of the XR content provided to the user U. In FIG. 9, the upper part shows the image A1 of the XR content before changing the parameters (for example, when the parameters are at the initial values), and the lower part shows the image A2 of the XR content after changing the parameters.

[0068] As shown in the upper part of FIG. 9, the controller 20 provides the user U with an image A1 including a high-resolution region B1 and a low-resolution region B2 by the FR process in step S13. Thereby, it is possible to reduce the processing load of the information processing apparatus 10 while suppressing a decrease in the immersion feeling of the user U.

[0069] Continuing the description of FIG. 8, next, the controller 20 acquires the biological data of the user U who is viewing the played XR content from the biological sensor 5 (step S14). Specifically, the controller 20 acquires biological data indicating the direction of the user U's line of sight, the direction of the head, blinking, etc. from the first biological sensor 5a. Also, the controller 20 acquires biological data indicating the brain waves of the user U from the second biological sensor 5b. Also, the controller 20 acquires biological data indicating the pulse of the user U from the third biological sensor 5c. Note that the controller 20 is not limited to acquiring all of the above-described biological data, and may be configured to acquire a part of the biological data.

[0070] Next, based on the acquired biological data, the controller 20 receives information regarding the mental and physical state of the user U who is viewing the reproduced XR content as "reward" related information (step S15). Next, the controller 20 performs learning to change the parameters (characteristics) in the FR process as an "action" (step S16). Specifically, the controller 20 performs learning (reinforcement learning) to change the "action" so that the received "reward" (the reward increases as the mental and physical state of the user U gets closer to the mental and physical state set as the goal for the content) increases. Then, the controller 20 executes the FR process using the changed parameters (step S17). Specifically, in step S17, the controller 20 stores each parameter calculated by the reinforcement learning model 34 in the storage unit 30 as parameter information 32. Then, the controller 20 executes the FR process using each parameter value of the parameter information 32 stored in this storage unit 30.

[0071] Here, the processing of steps S15 to S17 including the above-described reinforcement learning will be described with reference to FIG. 10. FIG. 10 is a diagram for explaining the reinforcement learning process and the like according to the embodiment.

[0072] In the example shown in FIG. 10, the controller 20 performs learning using the information regarding the emotion of the user U as the "reward". Specifically, the controller 20 acquires the information of the assumed emotion included in the "emotion label" (see FIG. 5) given to the XR content, and inputs the acquired information of the assumed emotion to the reinforcement learning model 34. Note that the information of the assumed emotion may be set for each scene of the XR content. In this case, the controller 20 acquires the information of the assumed emotion for each scene of the reproduced XR content, and inputs it to the reinforcement learning model 34.

[0073] Further, the controller 20 estimates the emotion of the user U who views the reproduced XR content, using the electroencephalogram and heart rate of the user U and the emotion estimation model 33 (see FIG. 7). Hereinafter, the estimated emotion may be referred to as the "estimated emotion". The controller 20 inputs the information on the estimated emotion of the estimated user U into the reinforcement learning model 34. In the reinforcement learning model 34, as a "reward", it calculates and receives the degree of coincidence between the estimated emotion of the estimated user U and the assumed emotion (the target emotion set by the content author or the like) that is assumed to be held by the user who views the XR content. The "reward" indicated by such a degree of coincidence increases as the degree of coincidence between the estimated emotion of the estimated user U and the assumed emotion becomes higher.

[0074] The controller 20 (specifically, the reinforcement learning model 34) performs learning (reinforcement learning) to change the parameters (characteristics) in the FR process as the "action" so that the "reward" increases as described above. Here, the controller 20 (reinforcement learning model 34) performs reinforcement learning to change the parameters so that the degree of coincidence between the estimated emotion and the assumed emotion increases. The controller 20 executes the FR process based on the changed parameters, and the video signal on which the FR process has been executed is output to the display device 3. The types of parameters to be changed will be described in detail later.

[0075] In the above, an example of performing learning by receiving the emotion of user U as a "reward" has been shown, but it is not limited to this. Specifically, the controller 20 performs learning using information regarding the heartbeat of user U as a "reward". For example, when the XR content is expected to increase the heartbeat of user U, such as "horror", the controller 20 calculates and receives, as a "reward", the rate of increase from the heartbeat (average value) of user U before the XR content is played. The "reward" indicated by such a rate of increase becomes larger as the degree of coincidence between the rate of increase from the heartbeat before XR content playback and the rate of increase of the heartbeat of the "biometric data target value" (see Fig. 5) increases. In this case, the controller 20 (reinforcement learning model 34) performs reinforcement learning to change the parameters so that the degree of coincidence between the rate of increase from the heartbeat before XR content playback and the rate of increase of the heartbeat of the "biometric data target value" increases.

[0076] Note that, for example, when the XR content is expected to decrease the heartbeat of user U, such as "healing", the controller 20 calculates and receives, as a "reward", the rate of decrease from the heartbeat of user U before the XR content is played. The "reward" indicated by such a rate of decrease becomes larger as the degree of coincidence between the rate of decrease from the heartbeat before XR content playback and the rate of decrease of the heartbeat of the "biometric data target value" (see Fig. 5) increases. In this case, the controller 20 (reinforcement learning model 34) performs reinforcement learning to change the parameters so that the degree of coincidence between the rate of decrease from the heartbeat before XR content playback and the rate of decrease of the heartbeat of the "biometric data target value" increases.

[0077] In addition, the controller 20 performs learning using the information on the number of blinks of the user U as "reward". The number of blinks is about 3 or 4 times per minute during normal times such as before the XR content is played. However, when the user U watches the XR content and the sense of immersion, horror, etc. increase and the concentration level rises, it is considered that the number of blinks decreases. Therefore, the controller 20 calculates and receives, as "reward", the reduction rate from the number of blinks (average value) of the user U before the XR content is played. The "reward" indicated by such a reduction rate increases as the degree of coincidence between the reduction rate from the number of blinks before the XR content is played and the reduction rate of the number of blinks of the "biometric data target value" (see FIG. 5) becomes higher. In this case, the controller 20 (reinforcement learning model 34) performs reinforcement learning to change the parameters so that the degree of coincidence between the reduction rate from the number of blinks before the XR content is played and the reduction rate of the number of blinks of the "biometric data target value" becomes higher.

[0078] Here, the types of parameters changed by the controller 20 will be specifically described. That is, the processes of steps S16 and S17 will be specifically described. The parameter (characteristic) to be changed is the area in the video of the XR content where the FR process is executed. Specifically, the controller 20 changes the area (parameter value of the processing area) of the area in the video of the XR content where the FR process is executed. For example, when the user U is concerned about the existence of the low-resolution area B2 (see FIG. 9), and the sense of immersion and presence during the viewing of the XR content are excessively reduced, and the "reward" received becomes relatively low. In such a case, as shown in the lower part of FIG. 9, the controller 20 increases the area of the high-resolution area B1 (see area B1a), decreases the area of the low-resolution area B2 (see area B2a), changes the "processing area" of the parameter information 32 (see FIG. 6), and executes the FR process.

[0079] Due to the decrease in the area of this low-resolution region B2, it becomes less likely for the user U to be concerned about the existence of the low-resolution region B2, thus improving the sense of immersion and presence when viewing XR content. Due to changes in the physical and mental state of the user U such as this improvement in the sense of immersion, the "reward" received by the controller 20 increases. In other words, the controller 20 can make the FR processing appropriate according to the user U by performing learning to change the parameters (characteristics, here the processing area) in the FR processing according to the physical and mental state of the user U during the reproduction of XR content.

[0080] Also, the parameter (characteristic) to be changed may be the intensity of the FR processing. Specifically, the controller 20 changes the intensity (parameter, characteristic) of the FR processing in the video of the XR content. For example, when the user U is concerned about the degree of low resolution (roughness) in the low-resolution region B2 (see FIG. 3), and the sense of immersion when viewing XR content excessively decreases, and the "reward" received becomes relatively low. In such a case, the controller 20 changes the "intensity" of the parameter information 32 (see FIG. 6) so that the intensity of the FR processing becomes low, in other words, increases the resolution of the low-resolution region B2, and executes the FR processing.

[0081] Due to this change in the intensity of the FR processing, it becomes less likely for the user U to be concerned about the degree of low resolution (roughness) in the low-resolution region B2, thus improving the sense of immersion and presence when viewing XR content. Due to changes in the physical and mental state of the user U such as this improvement in the sense of immersion, the "reward" received by the controller 20 increases. In other words, the controller 20 can make the FR processing appropriate according to the user U by performing learning to change the parameters (characteristics, here the intensity) in the FR processing according to the physical and mental state of the user U during the reproduction of XR content.

[0082] Also, the parameter (characteristic) to be changed may be the parameter (characteristic) of the FR process. Specifically, the controller 20 changes the period for changing the parameter (characteristic) of the FR process. For example, through the above-mentioned repeated learning, the "reward" received by the controller 20 may increase to a desired value. In such a case, it can be estimated that the FR process using the current parameters is appropriate for the user U. At this time, the controller 20 changes the "parameter change period" of the parameter information 32 (see FIG. 6) so that the period for changing the parameter (characteristic) of the FR process becomes longer, in other words, the learning frequency decreases, and then executes the FR process.

[0083] By changing the period for changing the parameter (characteristic) of this FR process, it becomes possible to further reduce the processing load of the information processing apparatus 10 while maintaining a relatively high sense of immersion of the user U.

[0084] Continuing the description of FIG. 8, the controller 20 determines whether the XR content has ended (step S18). If the controller 20 determines that the XR content has not ended (step S18, No), it returns to the process of step S14. In other words, the controller 20 performs learning (reinforcement learning) that repeats the processes of steps S14 to S18 until the XR content ends, the update process of the parameter value of the FR process, and the FR process using the updated parameter value.

[0085] If the controller 20 determines that the XR content has ended (step S18, Yes), it ends the process.

[0086] As described above, the information processing apparatus 10 according to the embodiment includes a controller 20 that executes a four-viewpoint rendering process (FR process) on content (XR content). The controller 20 changes the characteristics in the FR process according to the physical and mental state of the user U who views the reproduced XR content. Thereby, in the present embodiment, in the reproduction of XR content, the FR process can be made appropriate according to the user U.

[0087] Further, the controller 20 changes the characteristics in the FR process by the reinforcement learning model 34. The controller 20 performs reinforcement learning using, as a reward, the difference between the physical and mental state (for example, estimated emotion) of the user U who views the content (XR content) being reproduced and the physical and mental state (for example, assumed emotion) set for the XR content being reproduced. In this way, by the controller 20 changing the characteristics in the FR process by reinforcement learning using the reinforcement learning model 34, in the reproduction of XR content, the FR process can be made more appropriate according to the user U.

[0088] Further, the physical and mental state of the user U described above is an emotion. Thereby, in the present embodiment, in the reproduction of XR content, the FR process can be made appropriate according to the emotion of the user U.

[0089] Further, the controller 20 performs learning using, as a reward, information regarding the emotion of the user U. Thereby, in the present embodiment, it is possible to perform learning to change the characteristics in the FR process according to the physical and mental state of the user who views the reproduced XR content, specifically, the emotion of the user U, and as a result, the FR process can be made more appropriate according to the user U.

[0090] Further, the controller 20 estimates the emotion of the user U according to the brain waves and heart rate of the user U. Then, the controller 20 performs learning using, as a reward, information regarding the estimated emotion of the user U. In this way, the controller 20 can accurately estimate the emotion of the user U by using the brain waves and heart rate of the user U.

[0091] Also, the controller 20 performs learning using information regarding the user U's heartbeat as a reward. Thereby, in the present embodiment, it is possible to perform learning to change the characteristics in the FR process according to the physical and mental state of the user viewing the reproduced XR content, specifically, according to the heartbeat of the user U. As a result, the FR process can be made more appropriate according to the user U.

[0092] Also, the controller 20 performs learning using information regarding the number of blinks of the user U as a reward. Thereby, in the present embodiment, it is possible to perform learning to change the characteristics in the FR process according to the physical and mental state of the user viewing the reproduced XR content, specifically, according to the number of blinks of the user U. Therefore, the FR process can be made more appropriate according to the user U.

[0093] Note that in the above, an example was shown in which the parameters (characteristics) of the FR process to be changed as the "action" of the learning process are the intensity, processing area, and parameter change period of the FR process. However, it is not limited to changing all of these. That is, the "action" of the learning process may be to change a part of the intensity, processing area, and parameter change period of the FR process. Also, the parameters (characteristics) of the FR process to be changed are not limited to these.

[0094] Also, in the above, an example was shown in which the "reward" of the learning process is information regarding the user U's emotions, heartbeat, and number of blinks. However, the "reward" is not limited to receiving all of these. That is, the "reward" of the learning process may be a part of the information regarding the user U's emotions, heartbeat, and number of blinks. Also, the "reward" of the learning process is not limited to these, and may be other things such as information regarding the posture of the user U.

[0095] Next, the information processing apparatus 10 according to a modification of the embodiment will be described. In the embodiment, the parameters of the FR process are adjusted using the reinforcement learning model 34, but in the modification, a method of adjusting the parameters of the FR process by learning processing is also applicable.

[0096] Such a modification will be described with reference to FIGS. 11 and 12. FIG. 11 is a diagram for explaining the learning process and the like according to the modification, and FIG. 12 is an explanatory diagram showing an example of the table data 35. Note that the table data 35 is stored in the storage unit 30.

[0097] As shown in FIG. 11, the adjustment of the parameters of the FR process according to the modification is a method using the table data 35 instead of the reinforcement learning model 34. As shown in FIG. 12, the table data 35 is table data for determining a correction value of the parameter used for the FR process, with the estimated emotion and the assumed emotion (the emotion to be held by the target user in the content (or each scene)) as two parameters.

[0098] In the table data 35 shown in FIG. 12, the "estimated emotion" is the mental and physical state of the user viewing the content being played, and is an example of the first parameter. The "assumed emotion" is the assumed mental and physical state set for the content being played, and is an example of the second parameter. In the table data 35, the value selected by the combination of the estimated emotion (first parameter) and the assumed emotion (second parameter) is the correction value (parameter correction value) for the rendering parameter used in the FR process. In the example of FIG. 12, when the combination of the estimated emotion is "EF1" and the assumed emotion is "TF2", it shows that the parameter correction values are "intensity: +1", "processing area: +2", and "parameter change cycle: 0".

[0099] Continuing the description of FIG. 11, the controller 20 changes the parameters (characteristics) in the FR process by performing a selection process of correction values (parameter correction values) for the rendering parameters using this table data 35. Specifically, the controller 20 determines (selects) the correction value of the parameter using the table data 35 based on the assumed emotion in the content and the estimated emotion of the user at the time of content playback. Then, the controller 20 corrects the parameter stored as the parameter information 32 with the determined (selected) parameter correction value. The controller 20 performs the FR process using these corrected parameter information 32. As a result, the FR process is performed based on the parameters corrected by the parameter correction value.

[0100] As described above, in the modification example, by changing the parameters (characteristics) in the FR process using the table data 35, similarly to the embodiment, in the reproduction of the content, the FR process can be made appropriate according to the user.

[0101] Also, the mental and physical state of the user U in the modification example is emotion. Thus, in the modification example, similarly to the embodiment, in the reproduction of the XR content, the FR process can be made appropriate according to the emotion of the user U.

[0102] Further effects and modification examples can be easily derived by those skilled in the art. Therefore, the broader aspects of the present invention are not limited to the specific details and representative embodiments described and represented as above. Accordingly, various changes can be made without departing from the spirit or scope of the general inventive concept defined by the appended claims and their equivalents.

Explanation of Reference Numerals

[0103] 1 Information processing system 3 Display device 4 Speaker 10 Information processing device 20 Controller

Claims

1. An information processing apparatus comprising a controller that executes a four-dimensional rendering process on content, wherein the controller changes characteristics in the four-dimensional rendering process according to the physical and mental state of a user who views the reproduced content.

2. The physical and mental state is an emotion, The information processing apparatus according to claim 1.

3. The controller changes the characteristics in the four-dimensional rendering process by means of a reinforcement learning model, and performs reinforcement learning using, as a reward, the difference between the physical and mental state of the user who views the content being reproduced and the physical and mental state set for the content being reproduced. The information processing apparatus according to claim 1.

4. The physical and mental state is an emotion, The information processing apparatus according to claim 3.

5. The controller changes the characteristics in the four-dimensional rendering process by a selection process of a correction value for a rendering parameter using table data, wherein the table data has a first parameter that is the physical and mental state of the user who views the content being reproduced, a second parameter that is the assumed physical and mental state set for the content being reproduced, and a value selected by a combination of the first parameter and the second parameter is the correction value for the rendering parameter used in the four-dimensional rendering process. The information processing apparatus according to claim 1.

6. The physical and mental state is an emotion, The information processing apparatus according to claim 5.

7. The characteristic in the four-dimensional rendering process is the intensity of the four-dimensional rendering process, The information processing apparatus according to any one of claims 1 to 6.

8. The characteristic in the four-dimensional rendering process is the area in the video of the content where the four-dimensional rendering process is executed, The information processing apparatus according to any one of claims 1 to 6.

9. The characteristic in the four-dimensional rendering process is the period for changing the characteristic of the four-dimensional rendering process, The information processing apparatus according to any one of claims 1 to 6.

10. The controller estimates the emotion of the user according to the brain waves and heart rate of the user, The information processing apparatus according to claim 2, 4 or 6.

11. The controller ​ Performing the forbiated rendering process according to the content of the content The information processing apparatus according to any one of claims 1 to 6

12. An information processing apparatus that outputs a video signal and an audio signal in content, A display device that displays video according to the video signal of the content output from the information processing apparatus, An audio output device that outputs audio according to the audio signal of the content output from the information processing apparatus Comprising The information processing apparatus is Performing a forbiated rendering process on the video signal in the content, Changing the characteristics in the forbiated rendering process according to the physical and mental state of the user who views the reproduced content Information processing system

Citation Information

Patent Citations

  • Foveated rendering system and method

    JP2020042807A