Information processing device, information processing method, and program
By estimating user emotions in specific scenes and associating them with scene identification information, the method efficiently quantifies the sense of presence in XR experiences, addressing the limitations of conventional evaluation methods.
Patent Information
- Application Number
- JP2023189402
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-19
AI Technical Summary
Existing technologies face challenges in efficiently acquiring information about the sense of presence felt by users during virtual reality, augmented reality, and mixed reality experiences, due to the subjective and inaccurate nature of conventional evaluation methods.
An information processing method that estimates the emotion of a user in specific scenes of content and associates this emotion with scene identification information, allowing for the efficient quantification of the sense of presence.
This approach enables accurate and efficient acquisition of user sense of presence information, reducing the burden on users and improving the effectiveness of presence-enhancing parameters for sound and vibration.
Smart Images

Figure 2025077313000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.
Background Art
[0002] Conventionally, technologies for providing digital content including virtual space experiences such as VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality) to users using an HMD (Head Mounted Display) or the like are known. Such content is called XR (Cross Reality) content. XR is a representation that encompasses all virtual space technologies including VR, AR, MR, as well as SR (Substitutional Reality), AV (Audio / Visual), and the like.
[0003] Also, for example, a technology has been proposed in which a vibration device such as an exciter is controlled according to the video that the user watches, and the user is made to feel vibration, thereby improving the sense of presence with respect to the video (see, for example, Patent Document 1). In addition, improvement of the sense of presence is also realized by controlling the sound of the content.
[0004] Also, a technology is known in which parameters for controlling vibration and sound for improving the sense of presence are automatically set based on the scene of the content (see, for example, Patent Document 2).
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, in the prior art, there is a problem that information regarding the sense of presence felt by the user cannot be efficiently acquired.
[0007] In a system that outputs vibration and sound together with content, information (physical experience value) obtained by quantifying the sense of presence felt by the user is important information.
[0008] Conventionally, after a user experiences content, it has been common to obtain a physical experience value through subjective evaluation such as oral confirmation or a questionnaire. Also, in order to improve the sense of presence as a system, it is necessary to repeatedly perform the operation of obtaining the physical experience value and extract optimal parameters for sound and vibration.
[0009] On the other hand, it is considered that the physical experience value changes moment by moment according to various scenes included in the content. For this reason, it is necessary to obtain the physical experience value of each scene in the content. However, since the conventional subjective evaluation is based on the memory after experiencing the content, there are problems of lack of accuracy and a large burden on the user.
[0010] The present invention has been made in view of the above, and an object thereof is to provide a technique capable of efficiently acquiring information regarding the sense of presence felt by the user.
Means for Solving the Problems
[0011] The information processing method according to the present invention estimates the emotion of a content-using user in an arbitrary scene of the content being provided, and stores in a storage medium by associating scene identification information for identifying the scene with the emotion type estimated in the scene.
Effects of the Invention
[0012] According to the present invention, by estimating the emotion type, the sense of presence felt by the user in any scene of the content can be quantified. Therefore, compared with the conventional subjective evaluation, the sense of presence felt by the user can be accurately obtained while reducing the burden on the user. As a result, according to the present invention, information regarding the sense of presence felt by the user can be efficiently obtained.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Embodiments for Carrying Out the Invention
[0014] Hereinafter, with reference to the accompanying drawings, embodiments of an information processing apparatus, an information processing method, and a program disclosed in the present application will be described in detail. Note that the present invention is not limited by the embodiments shown below.
[0015] First, the outline of the information processing system according to the embodiment will be described with reference to FIG. 1. FIG. 1 is a diagram showing the outline of the information processing system.
[0016] Hereinafter, the case where the XR space (virtual space) is a VR space will be described. That is, the information processing system provides VR content to the user U1. However, the content provided by the information processing system is not limited to VR content, and may be, for example, video content displayed on a flat display.
[0017] As shown in FIG. 1, the information processing system 1 includes a first sensor 2, a second sensor 3, a display device 4, an acoustic output device 5, and a vibration device 6.
[0018] The information processing device 10 provides video data to the display device 4. Further, the information processing device 10 provides acoustic data to the acoustic output device 5. Further, the information processing device 10 provides vibration control data to the vibration device 6. Note that "acoustic" in the following description may be appropriately replaced with "voice" or "sound".
[0019] As shown in FIG. 1, the display device 4 is, for example, a head-mounted display. The display device 4 outputs an image based on the video data provided from the information processing device 10.
[0020] Note that the display device 4 may be a non-transmissive type that completely covers the field of view, or may be a video transmissive type or an optical transmissive type. Further, the display device 4 has a device that detects changes in the internal and external situations of the user U1 by a sensor unit, such as a camera or a motion sensor.
[0021] The acoustic output device 5 is, for example, headphones and is worn on the ears of the user U1. The acoustic output device 5 outputs an acoustic based on the acoustic data provided from the information processing device 10. Note that the acoustic output device 5 is not limited to the headphone type, and may be a box type (installed on the floor or the like). Further, the acoustic output device 5 may be a stereo audio or a multi-channel audio type.
[0022] The vibration device 6 vibrates by means of an electro-vibratory transducer including an electromagnetic circuit and a piezoelectric element. The vibration device 6 is provided, for example, on a seat on which the user U1 sits. The vibration device 6 vibrates in accordance with vibration control data provided from the information processing device 10.
[0023] The sound output from the acoustic output device 5 and the vibration of the vibration device 6 are waves. Also, the acoustic output device 5 and the vibration device 6 are examples of wave devices. The information processing system 1 can improve the sense of presence regarding the reproduction of video by adapting the waves by the wave device to the video and applying it to the user U1 of the content.
[0024] Also, the information processing device 10 estimates the emotion of the user U1 based on the sensor data acquired from the first sensor 2 and the second sensor 3. For example, the first sensor 2 is a headgear type electroencephalogram sensor. Also, for example, the second sensor 3 is a wristband type pulse sensor.
[0025] The configuration of the information processing device 10 will be described with reference to FIG. 2. FIG. 2 is a diagram showing a configuration example of the information processing device. As shown in FIG. 2, the information processing device 10 includes a controller 11, a memory 12, an input unit 13, and an output unit 14.
[0026] The controller 11 reads out and executes the program stored in the memory 12. The controller 11 is a microcomputer, a CPU (Central Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), a GPU (Graphics Processing Unit), a SoC (System on a Chip), or the like.
[0027] The controller 11 may be a single processor. The controller 11 may have a multiprocessor configuration. Also, the controller 11 may have a multi-core configuration having a plurality of cores in a single chip.
[0028] The memory 12 is a storage medium such as a flash memory. The memory 12 functions as a ROM (Read Only Memory) or a RAM (Random Access Memory). Note that for storing data that does not require high speed, such as storing the results of program processing, a storage medium such as a flash memory or a hard disk may be used, or an external storage medium (connectable and detachable to the information processing apparatus 10) such as a memory card may be used.
[0029] The input unit 13 and the output unit 14 are interfaces for inputting and outputting data between the information processing apparatus 10 and other apparatuses. For example, the input unit 13 receives the input of data from the first sensor 2 and the second sensor 3. Further, the output unit 14 outputs data to the display device 4, the acoustic output device 5, and the vibration device 6.
[0030] Using FIG. 3, the processing flow of the information processing apparatus 10 will be described. FIG. 3 is a functional block diagram for explaining the processing flow of the information processing apparatus. The main body of the processing shown in FIG. 3 is the controller 11.
[0031] As shown in FIG. 3, first, the controller 11 detects the type of scene based on the video data and acoustic data of the content provided to the user (step S1).
[0032] The scene type is predetermined by a designer or the like, and the information is registered in the scene information DB stored in the memory 12. FIG. 4 is a diagram showing an example of the scene information DB. As shown in FIG. 4, the scene information DB is a database that associates the scene type with the conditions for detecting (discriminating) each scene type. Note that a scene that does not correspond to the set scene type is determined as an other scene, and the acoustic and vibration parameters described later are default parameters (for example, no acoustic enhancement processing, no vibration). Note that the conditions for detecting each scene type are an example of scene specific information.
[0033] The scene example of the database shown in FIG. 4 is an example of a scene type that is often found in content that includes scenes where a moving object such as a racing game or a driving simulator travels. The conditions include the presence or absence of an image of an object or the like in the video of the content (utilization of image recognition technology), the presence or absence of the generated sound of an object or the like in the audio of the content, and the like. When the additional data of the content includes data indicating the scene type, the scene type may be determined based on the data.
[0034] The discrimination condition in the scene information DB is a condition for discriminating whether or not it is the corresponding scene type. For example, if the video and audio being played in the content satisfy the conditions, the scene type corresponding to those conditions becomes the scene type for the scene being played.
[0035] The scene example of the scene information DB shown in FIG. 4 is an example of a scene type that is often found in content that includes scenes where a moving object such as a racing game or a driving simulator travels. The conditions include the presence or absence of an image of an object or the like in the video of the content (utilization of image recognition technology), the presence or absence of the generated sound of an object or the like in the audio of the content, and the like. When the additional data of the content includes scene identification data indicating the scene type, the scene type may be determined based on the scene identification data (which is a condition in the scene information DB).
[0036] Next, the controller 11 extracts presence parameters related to the acoustic processing and vibration processing defined for the scene type (step S2). The presence parameters are parameters for processing the wave signals of the acoustic and vibration that are waves. For example, the presence parameters include at least any of the conversion parameters such as the intensity, frequency, and delay time of the wave, and the parameters of the residual wave characteristics (characteristics of the reverberant sound and residual vibration (for example, attenuation characteristics)). Here, it is assumed that the presence parameters include the above-described parameters belonging to the acoustic parameters and the vibration parameters. This makes it possible to arbitrarily process the acoustic and vibration that are waves.
[0037] Further, when the controller 11 detects a plurality of scene types at the same time, it may extract the presence parameter according to a predetermined priority. In this case, the data of the priority is added to the scene information DB shown in FIG. 4, and the scene type with the highest priority is selected based on the priority data.
[0038] Then, the controller 11 executes output control processing (step S3). The output control processing performs enhancement processing of the acoustic data based on the acoustic parameters for the acoustic data of the content, and conversion processing of the acoustic data into vibration control data based on the vibration parameters for the acoustic data of the content.
[0039] The controller 11 outputs the enhanced acoustic data to the acoustic output device 5. Further, the controller 11 outputs the vibration control data generated by the conversion processing to the vibration device 6.
[0040] Thereby, in the information processing method, it is possible to provide the user with enhanced sound according to the scene being viewed by the user and vibration according to the scene.
[0041] Note that the controller 11 can perform the processing of steps S1 to S3 by known acoustic processing methods and vibration generation methods described in Reference 1 (Japanese Patent Application Laid-Open No. 2023-51201) and the like. For example, FIG. 2 of Reference 1 shows the processing applicable to the scene detection, parameter extraction, and output control of the present embodiment.
[0042] Subsequently, the controller 11 estimates the user's emotion based on the first sensor data acquired from the first sensor 2 and the second sensor data acquired from the second sensor 3 (step S4).
[0043] The emotion estimation may identify emotion types such as "happy", "boring", and "scary", or may calculate the level for each emotion type.
[0044] Here, as an example of the emotion estimation method, a method using the user's brain waves and heartbeats and a method using the user's images and voices will be described.
[0045] (1) Emotion estimation method using the user's brain waves and heartbeats Figure 5 is a diagram for explaining the emotion estimation method. As shown in Figure 5, the controller 11 plots the coordinate points of the two index values corresponding to the acquired sensor data on the coordinate plane having an axis corresponding to the index value based on each of the two sensor data. Then, the emotion type assigned to the quadrant of the coordinate plane where the plotted coordinates exist is set as the estimated emotion (type). According to this method, it becomes possible to accurately capture the user's state and estimate emotions with high accuracy. It is also possible to estimate the intensity of the emotion by estimating the distance from the plotted coordinates.
[0046] The horizontal axis of the coordinate plane in Figure 5 corresponds to the standard deviation of the heart rate LF (Low Frequency) component. Also, the vertical axis of the coordinate plane in Figure 5 corresponds to the β wave / α wave of the brain wave (the relative magnitude of the β wave with respect to the α wave).
[0047] Both the standard deviation of the heart rate LF component and the β wave / α wave of the brain wave are examples of index values. Note that the standard deviation of the heart rate LF component is an example of an emotion intensity index value representing the intensity of an emotion. Also, the β wave / α wave of the brain wave is an example of an arousal degree index value representing the degree of arousal.
[0048] An emotion type is assigned in advance to each area of the coordinate plane. For example, an emotion type such as "happy" is assigned to the area R11 in the first quadrant. Also, for example, an emotion type "boring" is assigned to the area R12 straddling the second quadrant and the third quadrant. Also, for example, an emotion type such as "fear" is assigned to the area R13 in the fourth quadrant.
[0049] The controller 11 identifies the emotion type assigned to the quadrant in which the coordinates plotting the index values of emotion intensity and arousal exist as the user's emotion type. Further, according to the distance from the origin of the plotted coordinates, the level of the identified emotion type is identified. For example, the level increases as the plotted coordinates are farther from the origin (the greater the distance).
[0050] Note that the method for determining the emotion type using this coordinate plane can be determined by the relationship (greater than or less than relationship) of each index value with respect to the boundary value of each quadrant. It can also be determined using a matrix data table imitating the coordinate plane. Specifically, a data table is constructed with each index value as a parameter (the values of emotion intensity and arousal are the values of the vertical and horizontal axes), and the value for each combination of these parameters is the emotion type (it is also possible to include intensity). (Stored in the memory 12). Then, the data table is searched with the detected parameters, and the value for the corresponding combination of parameters is estimated as the emotion type.
[0051] Note that the controller 11 can also perform emotion estimation using various emotion estimation methods such as the methods described in Reference 2 (Japanese Patent Application Laid-Open No. 2019-63324) or Reference 3 (Japanese Patent Application Laid-Open No. 2022-134929).
[0052] (2) Emotion Estimation Method Using Appearance Information Such as the User's Image and Voice (Detection Information by a Non-Contact Sensor) An emotion estimation method using the user's image and voice will be described. Here, it is assumed that the first sensor 2 is a camera. Also, it is assumed that the second sensor 3 is a microphone. According to this method, it becomes possible to estimate emotions with a configuration using a non-contact sensor.
[0053] The controller 11 inputs the data of the captured image of the first sensor 2 or the gaze data (gaze data indicating an image of the eyeball part, gaze (direction), etc.) obtained by processing the captured image through image recognition processing or the like into an AI model (for example, a deep neural network). Also, the controller 11 inputs the voice data collected by the second sensor 3 into the AI model.
[0054] The AI model is a trained model that is trained using, as supervised learning data, a combination of image data and audio data (input data) and a label (correct data) representing the emotion type for the input data. Note that the supervised learning data can be generated, for example, from image data and audio data obtained from a subject for learning data generation and the emotion type obtained by the above-described method (emotion type estimated from brain waves and heartbeats) at the time of acquisition of the image data and audio data.
[0055] Based on the input data, the AI model outputs information for identifying emotions. The AI model may output an emotion type and its intensity. In this case, the AI model is a trained model that is trained using, as supervised learning data, a combination of image data and audio data (input data) and a label (correct data) representing the emotion type and its intensity for the input data.
[0056] Also, the AI model may output index values such as the intensity of emotion and the level of arousal. In this case, the AI model is a trained model that is trained using, as supervised learning data, a combination of image data and audio data (input data) and a label (correct data) representing each index value (for example, emotion intensity and level of arousal) for the input data. In this case, the controller 11 identifies the emotion type by plotting the output index values on the coordinate plane of FIG. 5.
[0057] Returning to FIG. 3, the controller 11 associates (correlates) the detected scene, the extracted sense of presence parameter, and the estimated emotion of the user (emotion estimation data) (step S5). Then, the controller 11 records the associated data obtained by the association by storing it in the memory 12 (step S6). The set of the recorded associated data is called time-series associated data. Note that the details of the association (correlation) process (associated data) will be described later.
[0058] Also, as shown in FIG. 3, the linking data is stored in the memory 12 so as to be identifiable and extractable for each user and for each content (type), and the data can be output based on the request of the data user (instruction operation to the user). Note that these data can be output (provided) to an external device via a memory card or communication.
[0059] Furthermore, the controller 11 analyzes the time-series linking data (step S7). Then, the controller 11 calibrates the presence parameters defined for the scene according to the analysis result (step S8).
[0060] Specifically, in step S8, the controller 11 updates the presence parameters defined for the scene based on the estimated user emotion. For example, the controller 11 can perform the update by correcting the presence parameters using the parameter correction value obtained based on the analysis result.
[0061] Here, a plurality of embodiments of the information processing apparatus 10 will be described. In each embodiment, the basic processing flow of the information processing apparatus 10 is as described with reference to FIG. 3. Between the embodiments, for example, the content of each data is different.
[0062] (Embodiment 1) FIG. 6 is a diagram for explaining the operation (transition of various data) in Embodiment 1. Note that each data will be stored in the memory 12 as a data table.
[0063] Step A1: Based on the content data from the content playback device CP, the controller 11 detects the playback location of the content (expressed in terms of the playback location time; the elapsed time from the start of the content). (Detect the playback location time data included in the content data, or measure the elapsed time since the start of content playback.) Note that the content playback device CP is a part of the information processing device 10 and includes a function of providing content data. The content data includes at least information for specifying the playback location of the content, video data, and audio data. Also, the playback location is an example of playback position information.
[0064] Step A2: For each playback location time (time at predetermined time intervals), determine the scene at that playback location time.
[0065] Step A3: The controller 11 searches the presence parameter table DA by the determined scene type and determines each presence parameter (acoustic / vibration) to be used for the output control (presence processing) S3.
[0066] Step A4: The controller 11 uses the playback location time data as a unique key (identification data) to generate a data group (one record data of the association data table DC) including the determined scene type, each presence parameter (acoustic / vibration) corresponding to the scene type (used for presence processing), and the emotion estimation result. In this example, the controller 11 generates the data of the association data table DC (temporary storage for creating the data records of the time-series association data table DD) only when a preset scene type (the scene type registered in the presence parameter table DA) is detected. Also, continuous scene types are grouped as one data, and the duration data indicating the scene continuation time is also stored as one data item of the association data table DC. The emotion estimation result may be the result of statistically processing the time lengths (number of determinations) of each emotion estimation result determined within the scene continuation time (for example, the emotion type with the most determination times).
[0067] Step A5: The controller 11 stores these data groups generated during content playback as data records in the time-series association data table DD with the playback location time data as the primary key (data for identification, and a data record is formed for each primary key).
[0068] Step A6: At a preset timing (when playback of one content ends, when a predetermined aggregation time has elapsed), the controller 11 performs analysis processing. In this example of analysis processing, the controller 11 aggregates the emotion estimation results in each same scene type stored in the time-series association data table DD, and determines the analysis results (correction content of the sense of presence parameter) corresponding to the aggregated emotion estimation results. Then, based on the results of these processes, data is stored (updated) in the analysis result data table DE composed of data records with the scene type as the primary key. As the process of aggregating the emotion estimation results in the analysis result data table DE, by statistically processing the total duration value of each emotion estimation result in each scene type, the emotion estimation analysis result of the corresponding scene type based on the analysis results is obtained. For example, the emotion estimation result corresponding to the longest total duration value is used as the emotion estimation analysis result for the corresponding scene type.
[0069] Step A7: The controller 11 searches the preset analysis data table DG (a matrix table of the sense of presence parameters with the scene type and the estimated emotion type as parameters) based on the emotion estimation analysis results (scene type and the corresponding emotion estimation analysis results), and determines the correction data for the sense of presence parameters. Note that each of these data is stored (updated) in the parameter correction value data table DF composed of data records with the scene type as the primary key. Note that each correction data in the analysis data table DG is a correction value for the sense of presence parameter for the emotion that the target user in each scene is expected to have, and is preset by experiments or the like.
[0070] Step A8: The controller 11 corrects the sense of presence parameters for each corresponding scene type in the sense of presence parameter table DA using the correction data of each scene type stored in the analysis result data table DE. Then, the controller 11 processes the content data (acoustic reproduction data) from the content playback device CP according to the scene using these corrected sense of presence parameters to generate acoustic data for output and vibration data, and outputs them to the acoustic output device 5 and the vibration device 6.
[0071] Note that in this example, the method of updating the sense of presence parameter data in the sense of presence parameter table DA is adopted. However, the sense of presence parameter data in the sense of presence parameter table DA may not be updated, but the correction value in the parameter correction value data table DF is updated, and the sense of presence parameter data in the sense of presence parameter table DA is corrected with the correction value. Then, using the corrected sense of presence parameter data, the content data (acoustic reproduction data) from the content playback device CP is processed to generate acoustic data for output and vibration data, and output to the acoustic output device 5 and the vibration device 6.
[0072] Note that in this embodiment, an example of the device until the acoustic data for output and the vibration data corresponding to the emotion are generated and output has been described. However, the processing until the generation of the association data table DC, the time-series association data table DD, or the analysis data table DE may be performed, and the generated data may be stored in a portable memory and provided, or provided online so that it can be utilized by another device. For example, an effective utilization method can be considered in which the time-series association data table DD is analyzed by a design developer and used for the design and development of the sense of presence device.
[0073] In Example 1, the scene detected by the controller 11 at the time "00:01:12" is "explosion" (Steps A1, A2). The presence parameters extracted by the controller 11 are "acoustic parameter +15dB" and "vibration parameter +20dB". Also, the emotion of the user estimated by the controller 11 is "bored". For example, the emotion estimation data is a label representing the emotion type.
[0074] The controller 11 adds the association data including the detected scene, the extracted presence parameters, and the estimated emotion to the time-series association data table DD stored in the memory 12 (Steps A3, A4, A5). In the example of FIG. 6, not only the association data at the time "00:01:12" but also the subsequent association data is shown.
[0075] Furthermore, the controller 11 analyzes the time-series association data (Step A6). Here, it is assumed that it is previously defined that it is an ideal state for the user's emotion to always be "happy". Also, it is known in advance that as the presence parameter increases, the user's emotion changes from "bored" to "happy" and "scary".
[0076] The controller 11 determines that it is necessary to enhance the presence parameter when the emotion is "bored". Also, the controller 11 determines that it is not necessary to change the presence parameter when the emotion is "happy". Also, the controller 11 determines that it is necessary to reduce the presence parameter when the emotion is "scary".
[0077] The controller 11 determines that it is necessary to reduce the presence parameter for the scene "shooting". Also, the controller 11 determines that it is necessary to enhance the presence parameter for the scene "explosion".
[0078] Furthermore, the controller 11 calculates a parameter correction value (step A7). Assume that for each of the presence parameters, an enhancement amplitude and a reduction amplitude are defined in advance in the analysis data table DG. The controller 11 acquires the predefined enhancement amplitude and reduction amplitude. Further, the controller 11 may use the acquired enhancement amplitude and reduction amplitude as the parameter correction value as they are, or may calculate the parameter correction value by adjusting the acquired enhancement amplitude and reduction amplitude with a coefficient.
[0079] For example, assume that the reduction amplitude of the vibration parameter is defined as "-6dB". Also, as described above, the controller 11 has determined that reduction of the presence parameter is necessary for the scene "shooting". From this, the controller 11 sets the parameter correction value of the vibration parameter for the scene "shooting" to "-6dB" (+12dB → +6dB).
[0080] For example, assume that the enhancement amplitude of the vibration parameter is defined as "+5dB". Also, as described above, the controller 11 has determined that enhancement of the presence parameter is necessary for the scene "explosion". From this, the controller 11 sets the parameter correction value of the vibration parameter for the scene "explosion" to "+5dB" (+20dB → +25dB).
[0081] In this way, the controller 11 corrects the presence parameter so that the emotion estimation result estimated in a certain scene of the content becomes the emotion type (for example, "fun") that the user targeted by the scene type has. Then, the controller 11 generates an acoustic signal and a vibration signal in each scene using the corrected presence parameter for each scene type and applies them to the user. Thereby, the user can feel the presence for each scene intended by the content producer.
[0082] (Example 2) FIG. 7 is a diagram for explaining Example 2. In Example 2, the sense of presence parameters are determined not only for each scene but also for each scene and user. Also, the user who receives the content is shown in the association data and the time-series association data. Note that descriptions of the same content as that in Example 1 described with reference to FIG. 6 are omitted as appropriate.
[0083] In Example 2, the controller 11 obtains the analysis results for each scene and user. Also, the controller 11 acquires the parameter correction values for each scene and user. Then, the controller 11 updates the pre-determined sense of presence parameters for each scene and user.
[0084] In Example 2, the configuration of some data tables is different from that in Example 1. The sense of presence parameter table DA_2 is a data table in which an item indicating the user is added to the sense of presence parameter table DA. The association data table DC_2 is a data table in which an item indicating the user is added to the association data table DC. The time-series association data table DD_2 is a data table in which an item indicating the user is added to the time-series association data table DD. The analysis result data table DE_2 is a data table in which an item indicating the user is added to the analysis result data table DE. The parameter correction value table DF_2 is a data table in which an item indicating the user is added to the parameter correction value table DF.
[0085] Also, in Example 2, step A1 in Example 1 changes to step A1_2. In step A1_2, in addition to the process of step A1, the controller 11 performs a process of determining the content-using user (before the start of content playback: determination by user input of user identification information or personal identification by face authentication, etc.).
[0086] Furthermore, in step A3_2, information indicating the user is added to the search key in step A3. Also, in steps A4_2, A5_2, A6_2, A7_2, and A8_2, information indicating the user is added to the information to be stored in steps A4, A5, A6, A7, and A8.
[0087] Even if the scene is the same, it is considered that the emotions of users are different. According to Example 2, it becomes possible to set presence parameters suitable for each user.
[0088] (Example 3) FIG. 8 is a diagram for explaining Example 3. In Example 3, the presence parameters are determined not only for each scene but also for each scene and user attribute. Also, in the association data and the time-series association data, the user who receives the content and the attributes of that user are shown.
[0089] In Example 3, the configuration of some data tables is different from that in Example 1. The presence parameter table DA_3 is a data table in which an item indicating the user attribute is added to the presence parameter table DA. The association data table DC_3 is a data table in which items indicating the user and attributes are added to the association data table DC. The time-series association data table DD_3 is a data table in which items indicating the user and attributes are added to the time-series association data table DD. The analysis result data table DE_3 is a data table in which items indicating the user and attributes are added to the analysis result data table DE. The parameter correction value table DF_3 is a data table in which items indicating the user and attributes are added to the parameter correction value table DF.
[0090] In addition, in Example 2, step A1 of Example 1 changes to step A1_3. In step A1_3, in addition to the processing of step A1, the controller 11 performs processing for determining the content-using user (before the start of content playback: determination by user input of user identification information or personal identification by face authentication, etc.). Further, in step A1_3, the controller 11 performs processing for determining the attributes of the content-using user (before the start of content playback: user input of various user information or extraction of various information of the user registered after user authentication).
[0091] Furthermore, in step A3_3, information indicating an attribute is added to the search key in step A3. Also, in steps A4_3, A5_3, A6_3, A7_3, and A8_3, information indicating the user and the attribute is added to the information stored in steps A4, A5, A6, and A7.
[0092] Step A8 of Example 1 is replaced by steps A8_3 and A9_3 in Example 3. The controller 11 acquires parameter correction values for each scene and attribute and stores them as learning data in the learning data table DH_3. The controller 11 corrects the presence parameter table DA_3 with the parameter correction values reflecting the learning data (step A9_3).
[0093] Even if the scenes are the same, the parameter correction values vary depending on the user. Also, depending on the attributes of the user, there is considered to be a certain tendency in the feelings remembered for the scene. The learning data can represent parameter correction values reflecting such tendencies for each attribute of the user.
[0094] For example, the learning data is the average value of the parameter correction values for each attribute of the user. The learning data may be used as the parameter correction value as it is. On the other hand, the learning data may be used for learning a model (for example, a regression model) with the user attributes as explanatory variables and the parameter correction values as target variables.
[0095] (Example 4) FIG. 9 is a diagram for explaining Example 4. In Example 4, the configuration of the data table and the processing content are common to those in Example 1. Example 4 can be realized by changing the target content and the content of the data table from those in Example 1.
[0096] When the content includes a scene where a moving object such as a racing game or a driving simulator travels, the scene may represent a predetermined place (course) where the moving object travels. For example, as shown in the on-site feeling parameter table DA in FIG. 9, "road driving", "dirt driving", "snow road driving", etc. can be scenes. Note that in FIG. 9, data such as analysis results is omitted.
[0097] According to such an example, for example, by appropriately setting the scene type to be processed according to the content type, evaluation suitable for the content type and acoustic / vibration can be performed.
[0098] (Example 5) FIG. 10 is a diagram for explaining Example 5. In Example 5, the configuration of the data table and the processing content are common to those in Example 1. Example 4 can be realized by changing the target content and the content of the data table from those in Example 1. However, as shown in the emotion estimation data table DB in FIG. 10, the emotion estimation result in Example 5 may represent not only the label indicating the emotion type but also the level (strength) of the emotion type. Note that in FIG. 10, data such as analysis results is omitted.
[0099] Furthermore, in Example 5, in the analysis data table DG, the correction value is further subdivided. This is realized, for example, by adding columns. As shown in FIG. 6, in Example 1, columns for each estimated emotion type such as "boring", "scary", and "fun" are registered in the analysis data table DG. On the other hand, in Example 5, for example, the column corresponding to "fun" is subdivided into three columns corresponding to three emotion types and levels such as "fun level 1", "fun level 2", and "fun level 3".
[0100] For example, as described with reference to FIG. 5, the controller 11 can determine the level of emotion according to the distance from the origin of the position where the point based on the index value is plotted on the coordinate plane.
[0101] In the first embodiment, the controller 11 performs calibration so that the emotion estimation result becomes the target emotion type (e.g., "happy"). On the other hand, in the fifth embodiment, the controller 11 can perform calibration so that the emotion estimation result becomes the target level of the target emotion type (e.g., "happy level 1").
[0102] In this case, the controller 11 can perform flexible calibration such that, for example, a sense of presence has a distinct contrast for each scene.
[0103] (Embodiment 6) FIG. 11 is a diagram for explaining the sixth embodiment. The configuration and processing content of the data table in the sixth embodiment are common to those in the fifth embodiment. That is, as in the fifth embodiment, in the sixth embodiment as well, the subdivided analysis data table DG is used.
[0104] In the sixth embodiment, the controller 11 detects the same scene from the content a plurality of times. The controller 11 extracts different presence parameters determined for each scene and detection count. The controller 11 estimates the emotion of the user in each of the scenes. The controller 11 associates and records each of the detected plurality of scenes, the presence parameters, and the estimated emotion of the user.
[0105] In the example of FIG. 11, the controller 11 detects the scene "dirt driving" a plurality of times. The presence parameters are determined for each detection count of the scene "dirt driving". That is, in this example, the change in the estimated emotion with respect to the occurrence count of the same scene, that is, the correlation of the occurrence count of the estimated emotion, is estimated and used for correction of sound and vibration, etc.
[0106] Example 6 is effective not only when the user receives content as a consumer but also when the user receives content as a developer. For example, it is conceivable that a user who is a developer (or a test player) may get bored while repeatedly experiencing the same scene. According to Example 6, it becomes possible to set a sense-of-presence parameter according to the degree of boredom due to repetition.
[0107] Using the flowchart of FIG. 12, the processing flow of the information processing apparatus 10 will be described. FIG. 12 is a flowchart showing the processing procedure executed by the information processing apparatus 10 (controller 11). This processing is executed by starting up the information processing apparatus 10 (a startup operation such as a content playback operation by the user).
[0108] The controller 11 performs startup processing of the information processing system 1 (power supply to each part in the system, various initial settings, etc.) and moves to step S101. In step S101, the controller 11 starts providing content and moves to step S102.
[0109] In step S102, the controller 11 detects the scene type of the content being provided (during playback) based on the information of the content being provided to the user and moves to step S103. That is, the controller 11 determines which scene type, such as "road driving", "dirt road driving", "snow road driving", "falling", "explosion", "shooting", "flying", the scene that the user is watching or playing belongs to.
[0110] In step S103, the controller 11 extracts the sense-of-presence parameter corresponding to the detected scene type from the sense-of-presence parameter table DA. Then, the controller 11 processes the audio data in the content using this extracted sense-of-presence parameter to generate an acoustic signal and a vibration signal. Then, these generated signals are output to the acoustic output device 5 and the vibration device 6, and the process moves to step S104.
[0111] In step S104, the controller 11 estimates the user's emotion based on the detection data of the first and second sensors 2 and 3, and then proceeds to step S105. That is, the controller 11 estimates the emotion type of the user's emotion by a method using brain waves and heartbeats, or a method using images and voices.
[0112] In step S105, the controller 11 associates the reproduction location data of the content, the scene type data, the extracted sense of presence parameters, and the estimated emotion type data to generate association data (temporarily stores it in the association data table DC). Then, this association data is stored in the time-series association data table DD as a data record with the reproduction location data as the primary key, and the process proceeds to step S106. That is, as described with reference to FIG. 6 and the like, the controller 11 combines the reproduction location data of the content, the scene type data, the extracted sense of presence parameters, and the estimated emotion type data into one record and stores it in the time-series association data table DD.
[0113] These data indicate at which position of the content the content user had what kind of emotion, what was the scene type at that time, and what were the sense of presence parameters used in the process at that time. Therefore, they are very useful data for evaluation and design and development in the sense of presence processing. Specifically, for example, these data are imported into a computer system for design and development (transmitted using a memory card or communication) and used. Note that the analysis results (data generated at each stage in the analysis) in the next step S106 are also very useful data in the same way.
[0114] In step S107, the controller 11 determines whether the provision (reproduction) of the content being provided has ended. If it has ended, the process proceeds to step S108. If it has not ended, the process returns to step S103 and repeats the processing from step S103 to step S107.
[0115] In step S108, the controller 11 performs the above-described analysis process on the association data stored in the time-series association data table DD in the reproduction of the content, and calculates a correction value for the presence parameter regarding each scene type. Then, these calculated correction values are stored (updated) in the parameter correction value data table DF, and the process proceeds to step S109.
[0116] In step S109, the controller 11 corrects the presence parameters of each scene type in the presence parameter table DA using the correction values stored in the parameter correction value data table DF, and ends the process. After that, the controller 11 processes the acoustic data of the content being provided using these corrected presence parameters for each scene type to generate acoustic data and vibration data. Then, the controller 11 outputs these generated signals to the acoustic output device 5 and the vibration device 6. Through the above-described process, based on the emotion estimated in the scene of each scene type in the content, the presence parameters for the corresponding scene type are corrected. Therefore, acoustic output and vibration output suitable for the scene can be expected.
[0117] [Other Embodiments] The configuration of the information processing system 1 is not limited to that shown in FIG. 1 and the like. For example, the functions of the information processing device 10 may be provided in a server. In that case, the display device 4, the acoustic output device 5, and the vibration device 6 receive video data, acoustic data, or vibration control data from the server via a terminal device capable of communicating with the server. Also, the first sensor 2 and the second sensor 3 transmit sensor data to the server via the terminal device.
[0118] Also, a configuration is possible in which a content providing device (VR playback device) is a separate device from the information processing system 1, and content information such as an acoustic signal and a video signal is input from the content providing device to the information processing system 1 (the information processing system 1 processes the content information to generate a vibration signal and the like).
[0119] In addition, the installation location of the information processing system 1 may be not only in buildings such as houses and entertainment facilities but also in vehicles. The information processing system 1 provided in a vehicle can provide content to a driver during a break.
[0120] Also, in a vehicle, notifications and warnings to the driver output by video, sound, vibration, etc. can also be regarded as content in a broad sense. For example, when the vehicle is about to come into contact with an object, the information processing system 1 outputs sound from a speaker and vibrates the seat to warn the driver.
[0121] At this time, if the intensity of the sound and vibration is too small, the driver's emotion will not change and the warning will have no meaning. On the other hand, if the intensity of the sound and vibration is too large, the driver may be overly surprised and may make a mistake in the driving operation. According to the information processing system 1, calibration can be performed so that the sound and vibration output as warnings in the vehicle have appropriate intensities based on the driver's emotion.
[0122] Further effects and modifications can be easily derived by those skilled in the art. Therefore, a broader aspect of the present invention is not limited to the specific details and representative embodiments described and represented as above. Accordingly, various changes can be made without departing from the spirit or scope of the general inventive concept defined by the appended claims and their equivalents.
Explanation of Reference Numerals
[0123] 1 Information processing system 2 First sensor 3 Second sensor 4 Display device 5 Sound output device 6 Vibration device 10 Information processing device 11 Controller 12 Memory 13 Input unit 14 Output unit
Claims
1. Estimating the emotion of a content user in any scene of the content being provided; scene identification information for identifying the scene and emotion type information estimated in the scene are stored in a storage medium in association with each other; Information processing device.
2. The information processing apparatus according to claim 1 , wherein the scene identification information is playback position information indicating a playback position of the content.
3. The information processing apparatus according to claim 1 , wherein the scene specification information is a scene type that indicates a type of the scene.
4. The scene identification information and the emotion type information related to each content user are stored in the storage medium so as to be extractable for each content user.
2. The information processing device according to claim 1.
5. The scene identification information and the emotion type information related to each content are stored in the storage medium so as to be extractable by each content.
2. The information processing device according to claim 1.
6. Get your heart rate information, Acquire brainwave information, Calculating an emotion intensity index value based on the heart rate information as a standard deviation of a heart rate LF (Low Frequency) component; Calculating a wakefulness index value based on the electroencephalogram information as a beta wave / alpha wave of the electroencephalogram; estimating the emotion based on a combination of a magnitude relationship of the emotion intensity index value relative to an emotion intensity threshold and a magnitude relationship of the arousal index value relative to an arousal threshold; The information processing device according to claim 1 .
7. The AI model has been trained using training data in which face images are input data and emotions are correct answer data, Acquire a facial image of the content user, Applying the acquired face image of the content user to the AI model to estimate the emotion. The information processing device according to claim 1 .
8. Estimating the emotion of a content user in any scene of the content being provided; scene identification information for identifying the scene and the emotion type estimated in the scene are stored in a storage medium in association with each other; Information processing methods.
9. Estimating the emotion of a content user in any scene of the content being provided; scene identification information for identifying the scene and the emotion type estimated in the scene are stored in a storage medium in association with each other; program.
Citation Information
Patent Citations
Bodily feeling experiencing video / sound system
JP2004081357A
Information processor, information processing system, and information processing method
JP2023051201A