Information processing device, information processing method, information processing program, and information processing system
The information processing method addresses the challenge of setting presence parameters by automatically adjusting them based on scene types and user emotions, enhancing the sense of presence for users in content systems.
Patent Information
- Application Number
- JP2023189403
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-19
AI Technical Summary
Existing systems face challenges in automatically setting appropriate presence parameters for users, as individual differences and environmental factors affect how users perceive presence, making it difficult for general users to adjust these parameters effectively.
An information processing method that detects scene types in content, generates vibration information based on acoustic and vibration parameters corresponding to the scene type, and estimates user emotions to adjust these parameters, thereby automatically setting presence parameters suitable for each user.
The method enables automatic correction of presence parameters based on user emotions and scene types, improving the sense of presence for users without requiring manual adjustments.
Smart Images

Figure 2025077314000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, an information processing program, and an information processing system.
Background Art
[0002] Conventionally, there has been known a technique for providing digital content including a virtual space experience such as VR (Virtual Reality), AR (Augmented Reality), or MR (Mixed Reality) to a user using an HMD (Head Mounted Display) or the like. Such content is called XR (Cross Reality) content. XR is an expression that encompasses all virtual space technologies including VR, AR, MR, as well as SR (Substitutional Reality), AV (Audio / Visual), and the like.
[0003] Also, for example, a technique has been proposed to improve the sense of immersion in a video by controlling a vibration device such as an exciter according to the video viewed by the user and making the user feel vibration (see, for example, Patent Document 1). The improvement of the sense of immersion can also be realized by controlling the sound of the content.
[0004] Also, a technique is known for automatically setting parameters for controlling vibration and sound to improve the sense of immersion based on the scene of the content (see, for example, Patent Document 2).
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, in the prior art, there is a problem that it is difficult to set presence parameters according to the user.
[0007] In order to set presence parameters according to the user, since there are individual differences and environmental differences in the way of feeling presence, in a system that outputs vibration and sound together with content, it is necessary for the user to set presence parameters by his own adjustment operation. However, the majority of general users have little knowledge about presence parameters and their adjustment methods, and it is very difficult to set appropriate presence parameters.
[0008] The present invention has been made in view of the above, and an object thereof is to provide a technology capable of automatically and appropriately setting presence parameters in a system that outputs vibration and sound together with content.
Means for Solving the Problems
[0009] In order to solve the above-described problems and achieve the object, an information processing method according to the present invention detects a scene type of each scene in content, and based on acoustic information in each scene of the content and vibration parameters corresponding to the detected scene type, generates vibration information in each scene of the content, and outputs the vibration information to a vibration device that generates vibration. The information processing method is such that in an arbitrary scene of the content, it estimates the emotion of the content user, generates association information associating the scene type of the scene and the emotion estimated in the scene, generates correction information for the vibration parameters for each scene type based on the association information, performs correction processing of the vibration parameters with the correction information to generate corrected vibration parameters, and generates the vibration information using the corrected vibration parameters.
Effects of the Invention
[0010] The present invention can automatically obtain information for correcting presence parameters according to scene types based on the estimated emotions of a user without requiring an operation by the user.
[0011] As a result, the parameters for giving a sense of presence can be adjusted based on the correction information, and the presence parameters can be set according to the user.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Modes for Carrying Out the Invention
[0013] Hereinafter, with reference to the accompanying drawings, embodiments of the information processing apparatus, information processing method, information processing program, and information processing system disclosed in the present application will be described in detail. Note that the present invention is not limited to the embodiments shown below.
[0014] First, the outline of the information processing system according to the embodiment will be described with reference to FIG. 1. FIG. 1 is a diagram showing the outline of the information processing system.
[0015] Hereinafter, the case where the XR space (virtual space) is the VR space will be described. That is, the information processing system provides VR content to the user U1. However, the content provided by the information processing system is not limited to VR content, and may be, for example, video content displayed on a flat display.
[0016] As shown in FIG. 1, the information processing system 1 includes a first sensor 2, a second sensor 3, a display device 4, an acoustic output device 5, and a vibration device 6.
[0017] The information processing apparatus 10 provides video data to the display device 4. The information processing apparatus 10 also provides acoustic data to the acoustic output device 5. The information processing apparatus 10 further provides vibration control data to the vibration device 6. Note that "acoustic" in the following description may be appropriately replaced with "voice" or "sound".
[0018] As shown in FIG. 1, the display device 4 is, for example, a head-mounted display. The display device 4 outputs an image based on the video data provided from the information processing apparatus 10.
[0019] Note that the display device 4 may be a non-transmissive type that completely covers the field of view, or may be a video transmissive type or an optical transmissive type. The display device 4 also has a device that detects changes in the internal and external situations of the user U1 by a sensor unit, such as a camera or a motion sensor.
[0020] The acoustic output device 5 is, for example, headphones and is worn on the ears of the user U1. The acoustic output device 5 outputs sound based on the acoustic data provided from the information processing device 10. Note that the acoustic output device 5 is not limited to the headphone type and may be a box type (installed on the floor or the like). Also, the acoustic output device 5 may be of the stereo audio or multi-channel audio type.
[0021] The vibration device 6 vibrates by means of an electro-vibration converter including an electromagnetic circuit and a piezoelectric element. The vibration device 6 is provided, for example, on the seat on which the user U1 is seated. The vibration device 6 vibrates in accordance with the vibration control data provided from the information processing device 10.
[0022] The sound output by the acoustic output device 5 and the vibration of the vibration device 6 are waves. Also, the acoustic output device 5 and the vibration device 6 are examples of wave devices. The information processing system 1 can improve the sense of presence regarding the reproduction of video by adapting the waves by the wave devices to the video and applying them to the user U1 of the content.
[0023] Also, the information processing device 10 estimates the emotion of the user U1 based on the sensor data acquired from the first sensor 2 and the second sensor 3. For example, the first sensor 2 is a headgear type electroencephalogram sensor. Also, for example, the second sensor 3 is a wristband type pulse sensor.
[0024] The configuration of the information processing device 10 will be described with reference to FIG. 2. FIG. 2 is a diagram showing a configuration example of the information processing device. As shown in FIG. 2, the information processing device 10 includes a controller 11, a memory 12, an input unit 13, and an output unit 14.
[0025] The controller 11 reads and executes the program stored in the memory 12. The controller 11 is a microcomputer, a CPU (Central Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), a GPU (Graphics Processing Unit), a SoC (System on a Chip), etc.
[0026] The controller 11 may be a single processor. The controller 11 may have a multiprocessor configuration. Also, the controller 11 may have a multi-core configuration with multiple cores in a single chip.
[0027] The memory 12 is a storage medium such as a flash memory. The memory 12 functions as a ROM (Read Only Memory) or a RAM (Random Access Memory). Note that for storing data such as the results of program processing where high speed is not highly required, a storage medium such as a flash memory or a hard disk may be used, or an external storage medium (connectable and detachable to the information processing apparatus 10) such as a memory card may also be used.
[0028] The input unit 13 and the output unit 14 are interfaces for inputting and outputting data between the information processing apparatus 10 and other devices. For example, the input unit 13 receives the input of data from the first sensor 2 and the second sensor 3. Also, the output unit 14 outputs data to the display device 4, the acoustic output device 5, and the vibration device 6.
[0029] Using FIG. 3, the processing flow of the information processing apparatus 10 will be described. FIG. 3 is a functional block diagram for explaining the processing flow of the information processing apparatus. The main body of the processing shown in FIG. 3 is the controller 11.
[0030] As shown in FIG. 3, first, the controller 11 detects the type of scene based on the video data and audio data of the content provided to the user (step S1).
[0031] The scene type is predetermined by a designer or the like, and the information is registered in the scene information DB stored in the memory 12. FIG. 4 is a diagram showing an example of the scene information DB. As shown in FIG. 4, the scene information DB is a database that associates a scene type with conditions for detecting (discriminating) each scene type. Note that a scene that does not correspond to the set scene type is determined as another scene, and the acoustic and vibration parameters described later are default parameters (for example, no acoustic enhancement process, no vibration). Note that the conditions for detecting each scene type are an example of scene specific information.
[0032] The scene examples in the database shown in FIG. 4 are examples of scene types that are often included in content that includes a scene in which a moving object such as a racing game or a driving simulator travels. The conditions include the presence or absence of an image of an object or the like that is the target in the video of the content (use of image recognition technology), the presence or absence of a generated sound of an object or the like that is the target in the audio of the content, and the like. Note that when the additional data of the content includes data indicating the scene type, the scene type may be discriminated based on the data.
[0033] The discrimination conditions in the scene information DB are conditions for discriminating whether or not it is the corresponding scene type. For example, if the video and audio being played in the content satisfy the conditions, the scene type corresponding to those conditions becomes the scene type for the scene being played.
[0034] The scene examples in the scene information DB shown in FIG. 4 are examples of scene types that are often included in content that includes a scene in which a moving object such as a racing game or a driving simulator travels. The conditions include the presence or absence of an image of an object or the like that is the target in the video of the content (use of image recognition technology), the presence or absence of a generated sound of an object or the like that is the target in the audio of the content, and the like. Note that when the additional data of the content includes scene identification data indicating the scene type, the scene type may be discriminated based on the scene identification data (which becomes the condition in the scene information DB).
[0035] Next, the controller 11 extracts presence parameters related to acoustic processing and vibration processing defined for the scene type (step S2). The presence parameters are parameters for processing the wave signals of the acoustic and vibration that are waves. For example, the presence parameters include at least any of the parameters for wave processing, such as conversion parameters such as the intensity, frequency, and delay time of the wave, and parameters of the residual wave characteristics (characteristics of the reverberant sound and residual vibration (for example, attenuation characteristics)). Here, it is assumed that the presence parameters include the above-described parameters belonging to the acoustic parameters and the vibration parameters. This makes it possible to arbitrarily process the acoustic and vibration that are waves.
[0036] Also, when the controller 11 detects a plurality of scene types at the same time, the controller 11 may extract the presence parameters according to a predetermined priority. In this case, the data of the priority is added to the scene information DB shown in FIG. 4, and the scene type with the higher priority is selected based on the priority data.
[0037] Then, the controller 11 executes output control processing (step S3). The output control processing performs an enhancement process on the acoustic data based on the acoustic parameters for the acoustic data of the content, and a conversion process to vibration control data based on the vibration parameters for the acoustic data of the content.
[0038] The controller 11 outputs the enhanced acoustic data to the acoustic output device 5. Also, the controller 11 outputs the vibration control data generated by the conversion process to the vibration device 6.
[0039] Thereby, in the information processing method, it is possible to provide the user with the enhanced sound according to the scene that the user views and the vibration according to the scene.
[0040] Note that the controller 11 can perform the processes of steps S1 to S3 by using known acoustic processing methods and vibration generation methods described in Reference 1 (Japanese Patent Application Laid-Open No. 2023-51201) and the like. For example, FIG. 2 of Reference 1 shows processes applicable to scene detection, parameter extraction, and output control of the present embodiment.
[0041] Subsequently, the controller 11 estimates the user's emotion based on the first sensor data acquired from the first sensor 2 and the second sensor data acquired from the second sensor 3 (step S4).
[0042] The emotion estimation may identify emotion types such as "happy", "boring", and "scary", or may calculate the level for each emotion type.
[0043] Here, as an example of the emotion estimation method, a method using the user's brain waves and heart rate and a method using the user's image and voice will be described.
[0044] (1) Emotion Estimation Method Using the User's Brain Waves and Heart Rate FIG. 5 is a diagram for explaining the emotion estimation method. As shown in FIG. 5, the controller 11 plots the coordinate points of two index values corresponding to the acquired sensor data on a coordinate plane having axes corresponding to the index values based on each of the two sensor data. Then, the emotion type assigned to the quadrant of the coordinate plane where the plotted coordinates exist is set as the estimated emotion (type). According to this method, it becomes possible to accurately capture the user's state and estimate the emotion with high accuracy. It is also possible to estimate the intensity of the emotion based on the distance from the plotted coordinates.
[0045] The horizontal axis of the coordinate plane in FIG. 5 corresponds to the standard deviation of the heart rate LF (Low Frequency) component. Also, the vertical axis of the coordinate plane in FIG. 5 corresponds to the β wave / α wave of the brain wave (relative magnitude of the β wave with respect to the α wave).
[0046] The standard deviation of the LF component of the heartbeat and the β / α wave of the electroencephalogram are both examples of index values. Note that the standard deviation of the LF component of the heartbeat is an example of an emotional intensity index value representing the strength of emotion. Also, the β / α wave of the electroencephalogram is an example of an arousal level index value representing the degree of arousal.
[0047] Each region of the coordinate plane has an emotion type assigned to it in advance. For example, the emotion type "happy" etc. is assigned to the region R11 in the first quadrant. Also, for example, the emotion type "boring" is assigned to the region R12 straddling the second and third quadrants. Also, for example, the emotion type "fear" etc. is assigned to the region R13 in the fourth quadrant.
[0048] The controller 11 identifies the emotion type assigned to the quadrant in which the coordinates plotting the index values of the emotional intensity and arousal level exist as the user's emotion type. Also, according to the distance from the origin of the plotted coordinates, the level of the identified emotion type is identified. For example, the level increases as the plotted coordinates are farther from the origin (the greater the distance).
[0049] Note that the method for determining the emotion type using this coordinate plane can be determined by the relationship (greater than or less than relationship) of each index value with respect to the boundary value of each quadrant. Also, it can be determined using a matrix data table imitating the coordinate plane. Specifically, a data table is constructed (stored in the memory 12) with each index value as a parameter (the values of the emotional intensity and arousal level are the values on the vertical and horizontal axes), and the value for the combination of each parameter is the emotion type (it is also possible to include the intensity). Then, the data table is searched with the detected parameters, and the value for the corresponding combination of parameters is estimated as the emotion type.
[0050] Note that the controller 11 can also perform emotion estimation using various emotion estimation methods such as the methods described in Reference 2 (Japanese Patent Application Laid-Open No. 2019-63324) or Reference 3 (Japanese Patent Application Laid-Open No. 2022-134929).
[0051] (2) Method for estimating emotion using appearance information such as user's image and voice (detection information by non-contact sensors) A method for estimating emotion using the user's image and voice will be described. Here, it is assumed that the first sensor 2 is a camera. Also, it is assumed that the second sensor 3 is a microphone. According to this method, it becomes possible to estimate emotion with a configuration using non-contact sensors.
[0052] The controller 11 inputs the data of the captured image of the first sensor 2 or the gaze data (gaze data indicating an image of the eye area, gaze (direction), etc.) obtained by processing the captured image by image recognition processing or the like into an AI model (for example, a deep neural network). Also, the controller 11 inputs the voice data collected by the second sensor 3 into the AI model.
[0053] Note that the AI model is a learned model that has been trained using, as teacher - supervised learning data, a combination of image data and voice data (input data) and a label (correct data) representing the emotion type for the input data. Note that the teacher - supervised learning data can be generated, for example, from image data and voice data obtained from subjects for learning data generation and the emotion type obtained by the above - mentioned method (emotion type estimated by electroencephalogram and heartbeat) at the time of acquisition of the image data and voice data.
[0054] Based on the input data, the AI model outputs information for identifying emotion. The AI model may output the emotion type and its intensity. In this case, the AI model is a learned model that has been trained using, as teacher - supervised learning data, a combination of image data and voice data (input data) and a label (correct data) representing the emotion type and its intensity for the input data.
[0055] In addition, the AI model may output index values such as the intensity of emotion and the level of arousal. In this case, the AI model is a trained model that has been trained using, as training data with labels (correct data), a combination of image data and audio data (input data) and each index value (for example, emotion intensity and arousal level) for the input data. In this case, the controller 11 identifies the emotion type by plotting the output index value on the coordinate plane of FIG. 5.
[0056] Returning to FIG. 3, the controller 11 associates (correlates) the detected scene, the extracted sense of presence parameter, and the estimated emotion of the user (emotion estimation data) (step S5). Then, the controller 11 records the associated data obtained by the association by storing it in the memory 12 (step S6). The set of the recorded associated data is called time-series associated data. The details of the association (correlation) process (associated data) will be described later.
[0057] In addition, as shown in FIG. 3, the associated data is stored in the memory 12 in an identifiable and extractable manner for each user and for each content (type), and the data can be output based on the request of the data user (user instruction operation). These data can be output (provided) to an external device via a memory card or communication.
[0058] Furthermore, the controller 11 analyzes the time-series associated data (step S7). Then, the controller 11 calibrates the sense of presence parameter defined for the scene according to the analysis result (step S8).
[0059] Specifically, in step S8, the controller 11 updates the sense of presence parameter defined for the scene based on the estimated emotion of the user. For example, the controller 11 can perform the update by correcting the sense of presence parameter using the parameter correction value obtained based on the analysis result.
[0060] Here, a plurality of embodiments of the information processing apparatus 10 will be described. In each embodiment, the basic processing flow of the information processing apparatus 10 is as described with reference to FIG. 3. Between the embodiments, for example, the content of each data is different.
[0061] (Embodiment 1) FIG. 6 is a diagram for explaining the operation (transition of various data) in Embodiment 1. Each data will be stored in the memory 12 as a data table.
[0062] Step A1: The controller 11 detects the playback position of the content (expressed by the playback position time; the elapsed time from the start of the content) based on the content data from the content playback device CP (detects the playback position time data included in the content data, or measures the elapsed time since the start of content playback). Note that the content playback device CP is a part of the information processing apparatus 10 and includes a function of providing content data. The content data includes at least information for specifying the playback position of the content, video data, and audio data. Also, the playback position is an example of playback position information.
[0063] Step A2: Determine the scene at the playback position time for each playback position time (time at predetermined time intervals).
[0064] Step A3: The controller 11 searches the presence parameter table DA by the determined scene type and determines each presence parameter (acoustic / vibration) to be used for the output control (presence processing) S3.
[0065] Step A4: The controller 11 uses the playback location time data as a unique key (identification data), and generates a data group (one record data of the association data table DC) including the determined scene type, each presence parameter (acoustic / vibration) corresponding to the scene type (used for presence processing), and the emotion estimation result. In this example, the controller 11 generates the data of the association data table DC (temporary storage for creating the data records of the time-series association data table DD) only when a preset scene type (the scene type registered in the presence parameter table DA) is detected. Also, continuous scene types are grouped as one data, and the duration data indicating the scene continuation time is also stored as one data item of the association data table DC. The emotion estimation result may be the result of statistically processing the time lengths (number of determinations) of each emotion estimation result determined within the scene continuation time (for example, the emotion type with the most determination times).
[0066] Step A5: The controller 11 stores these data groups generated during content playback as data records in the time-series association data table DD with the playback location time data as the primary key (identification data, and a data record is formed for each primary key).
[0067] Step A6: At a preset timing (when one content playback ends, when a preset summary time elapses), the controller 11 performs analysis processing. In this example of the analysis processing, the controller 11 summarizes the emotion estimation results for each same scene type stored in the time-series association data table DD, and determines the analysis result (correction content of the presence parameter) corresponding to the summarized emotion estimation result. Then, based on the results of these processes, data is stored (updated) in the analysis result data table DE composed of data records with the scene type as the primary key. As the process of summarizing the emotion estimation results in the analysis result data table DE, the statistical processing of the total duration values of each emotion estimation result for each scene type is used to obtain the emotion estimation analysis result of the scene type based on the analysis result. For example, the emotion estimation result corresponding to the longest total duration value is used as the emotion estimation analysis result for the scene type.
[0068] Step A7: The controller 11 searches a preset analysis data table DG (a matrix table of presence parameters with scene type and estimated emotion type as parameters) based on the emotion estimation analysis results (scene type and corresponding emotion estimation analysis results), and determines correction data for the presence parameters. Each of these data is stored (updated) in a parameter correction value data table DF composed of data records with the scene type as the primary key. Each correction data in the analysis data table DG is a correction value for the presence parameter to make the estimated emotion closer to the emotion that the target user has in each scene, and is preset by experiments or the like.
[0069] Step A8: The controller 11 corrects the presence parameters for the corresponding scene types in the presence parameter table DA using the correction data for each scene type stored in the analysis result data table DE. Then, the controller 11 processes the content data (acoustic reproduction data) from the content playback device CP according to the scene using these corrected presence parameters, generates output acoustic data and vibration data, and outputs them to the acoustic output device 5 and the vibration device 6.
[0070] Note that in this example, the presence parameter data in the presence parameter table DA is updated. However, the presence parameter data in the presence parameter table DA may not be updated, but the correction value in the parameter correction value data table DF is updated, and the presence parameter data in the presence parameter table DA is corrected with the correction value. Then, the corrected presence parameter data is used to process the content data (acoustic reproduction data) from the content playback device CP to generate output acoustic data and vibration data, and output them to the acoustic output device 5 and the vibration device 6.
[0071] In this embodiment, an example of the apparatus until generating and outputting acoustic data and vibration data for output according to emotions has been described. However, the processing until generating the association data table DC, the time-series association data table DD, or the analysis data table DE may be performed, and the generated data may be stored in a portable memory and provided, or provided online so that it can be utilized by another apparatus. For example, an effective utilization method such as an analysis developer analyzing the time-series association data table DD and using it for the design and development of the immersive device can be considered.
[0072] In the first embodiment, the scene detected at the time "00:01:12" detected by the controller 11 is "explosion" (steps A1, A2). The immersive parameters extracted by the controller 11 are "acoustic parameter +15 dB" and "vibration parameter +20 dB". Also, the emotion of the user estimated by the controller 11 is "bored". For example, the emotion estimation data is a label representing the emotion type.
[0073] The controller 11 adds the association data including the detected scene, the extracted immersive parameters, and the estimated emotion to the time-series association data table DD stored in the memory 12 (steps A3, A4, A5). In the example of FIG. 6, not only the association data at the time "00:01:12" but also the subsequent association data is shown.
[0074] Furthermore, the controller 11 analyzes the time-series association data (step A6). Here, it is assumed that it is predetermined that the ideal state is that the user's emotion is always "happy". Also, it is assumed that it is known in advance that as the immersive parameter increases, the user's emotion changes from "bored" to "happy" and "scary".
[0075] When the emotion is "boring", the controller 11 determines that it is necessary to enhance the sense of presence parameter. Also, when the emotion is "fun", the controller 11 determines that it is not necessary to change the sense of presence parameter. Further, when the emotion is "scary", the controller 11 determines that it is necessary to reduce the sense of presence parameter.
[0076] The controller 11 determined that it is necessary to reduce the sense of presence parameter for the scene "shooting". Also, the controller 11 determined that it is necessary to enhance the sense of presence parameter for the scene "explosion".
[0077] Furthermore, the controller 11 calculates a parameter correction value (step A7). Assume that for each of the sense of presence parameters, the enhancement amplitude and the reduction amplitude are defined in the analysis data table DG in advance. The controller 11 acquires the predefined enhancement amplitude and reduction amplitude. Further, the controller 11 may use the acquired enhancement amplitude and reduction amplitude as the parameter correction value as they are, or may calculate the parameter correction value by adjusting the acquired enhancement amplitude and reduction amplitude with a coefficient.
[0078] For example, assume that the reduction amplitude of the vibration parameter is defined as "-6dB". Also, as described above, the controller 11 determined that it is necessary to reduce the sense of presence parameter for the scene "shooting". From this, the controller 11 sets the parameter correction value of the vibration parameter for the scene "shooting" to "-6dB" (+12dB → +6dB).
[0079] For example, assume that the enhancement amplitude of the vibration parameter is defined as "+5dB". Also, as described above, the controller 11 determined that it is necessary to enhance the sense of presence parameter for the scene "explosion". From this, the controller 11 sets the parameter correction value of the vibration parameter for the scene "explosion" to "+5dB" (+20dB → +25dB).
[0080] In this way, the controller 11 corrects the presence parameter so that the emotion estimation result estimated in a scene with content becomes the emotion type (for example, "happy") that the user targeted by the scene type has. Then, the controller 11 generates an acoustic signal and a vibration signal in each scene using the corrected presence parameter for each scene type and applies them to the user. As a result, the user can feel the presence for each scene intended by the content producer.
[0081] (Example 2) FIG. 7 is a diagram for explaining Example 2. In Example 2, the presence parameter is determined not only for each scene but also for each scene and user. Also, the user who receives the content is shown in the association data and the time-series association data. Regarding the same content as that described in Example 1 using FIG. 6, the explanation will be omitted as appropriate.
[0082] In Example 2, the controller 11 obtains an analysis result for each scene and user. Also, the controller 11 acquires a parameter correction value for each scene and user. Then, the controller 11 updates the predefined presence parameter for each scene and user.
[0083] In Example 2, the configuration of some data tables is different from that in Example 1. The presence parameter table DA_2 is a data table in which an item indicating the user is added to the presence parameter table DA. The association data table DC_2 is a data table in which an item indicating the user is added to the association data table DC. The time-series association data table DD_2 is a data table in which an item indicating the user is added to the time-series association data table DD. The analysis result data table DE_2 is a data table in which an item indicating the user is added to the analysis result data table DE. The parameter correction value table DF_2 is a data table in which an item indicating the user is added to the parameter correction value table DF.
[0084] In addition, in Example 2, step A1 of Example 1 changes to step A1_2. In step A1_2, in addition to the processing of step A1, the controller 11 performs processing for determining the content-using user (before starting content playback: determination by user input of user identification information or personal identification by face authentication or the like).
[0085] Furthermore, in step A3_2, information indicating the user is added to the search key in step A3. Also, in steps A4_2, A5_2, A6_2, A7_2, and A8_2, information indicating the user is added to the information stored in steps A4, A5, A6, A7, and A8.
[0086] Even if the scene is the same, it is considered that the emotions of users are different. According to Example 2, it becomes possible to set the presence parameters suitable for each user.
[0087] (Example 3) FIG. 8 is a diagram for explaining Example 3. In Example 3, the presence parameters are determined not only for each scene but also for each scene and user attribute. Also, the user who receives the content and the attributes of that user are shown in the association data and the time-series association data.
[0088] In Example 3, the configuration of some data tables is different from that in Example 1. The presence parameter table DA_3 is a data table in which items indicating user attributes are added to the presence parameter table DA. The association data table DC_3 is a data table in which items indicating the user and attributes are added to the association data table DC. The time-series association data table DD_3 is a data table in which items indicating the user and attributes are added to the time-series association data table DD. The analysis result data table DE_3 is a data table in which items indicating the user and attributes are added to the analysis result data table DE. The parameter correction value table DF_3 is a data table in which items indicating the user and attributes are added to the parameter correction value table DF.
[0089] Also, in Example 3, step A1 of Example 1 changes to step A1_3. In step A1_3, in addition to the processing of step A1, the controller 11 performs processing for determining the content-using user (before starting content playback: determination by user input of user identification information or personal identification by face authentication, etc.). Further, in step A1_3, the controller 11 performs processing for determining the attributes of the content-using user (before starting content playback: user input of various user information or extraction of various information of the registered user after user authentication).
[0090] Furthermore, in step A3_3, information indicating an attribute is added to the search key of step A3. Also, in steps A4_3, A5_3, A6_3, A7_3, and A8_3, information indicating the user and the attribute is added to the information stored in steps A4, A5, A6, and A7.
[0091] Step A8 of Example 1 is replaced by steps A8_3 and A9_3 in Example 3. The controller 11 acquires parameter correction values for each scene and attribute, and stores them as learning data in the learning data table DH_3. The controller 11 corrects the presence parameter table DA_3 with the parameter correction values reflecting the learning data (step A9_3).
[0092] Even if the scenes are the same, the parameter correction values change depending on the user. Also, it is considered that there is a certain tendency in the emotions remembered for the scene depending on the attributes of the user. The learning data can represent parameter correction values reflecting such tendencies for each attribute of the user.
[0093] For example, the learning data is the average value of the parameter correction values for each attribute of the user. The learning data may be used as the parameter correction value as it is. On the other hand, the learning data may be used for learning a model (for example, a regression model) with the user attribute as the explanatory variable and the parameter correction value as the objective variable.
[0094] (Example 4) FIG. 9 is a diagram for explaining Example 4. The configuration and processing content of the data table in Example 4 are common to those in Example 1. Example 4 can be realized by changing the target content and the content of the data table from those in Example 1.
[0095] When the content includes a racing game or a scene where a moving object runs such as a driving simulator, the scene may represent a predetermined place (course) where the moving object runs. For example, as shown in the immersive parameter table DA in FIG. 9, "road driving", "dirt driving", "snow road driving", etc. can be scenes. Note that in FIG. 9, data such as analysis results is omitted.
[0096] According to such an example, for example, by appropriately setting the scene type to be processed according to the content type, evaluation suitable for the content type and acoustic / vibration can be performed.
[0097] (Example 5) FIG. 10 is a diagram for explaining Example 5. The configuration and processing content of the data table in Example 5 are common to those in Example 1. Example 4 can be realized by changing the target content and the content of the data table from those in Example 1. However, as shown in the emotion estimation data table DB in FIG. 10, the emotion estimation result in Example 5 may represent not only the label representing the emotion type but also the level (strength) of the emotion type. Note that in FIG. 10, data such as analysis results is omitted.
[0098] Furthermore, in Example 5, in the analysis data table DG, the correction values are further subdivided. This is achieved, for example, by adding columns. As shown in FIG. 6, in Example 1, columns for each estimated emotion type such as "boring", "scary", and "fun" are registered in the analysis data table DG. On the other hand, in Example 5, for example, the column corresponding to "fun" is subdivided into three columns corresponding to three emotion types and levels such as "fun level 1", "fun level 2", and "fun level 3".
[0099] For example, as described with reference to FIG. 5, the controller 11 can determine the emotion level according to the distance from the origin of the position where the point based on the index value is plotted on the coordinate plane.
[0100] In Example 1, the controller 11 performs calibration so that the emotion estimation result becomes the target emotion type (for example, "fun"). On the other hand, in Example 5, the controller 11 can perform calibration so that the emotion estimation result becomes the target level of the target emotion type (for example, "fun level 1").
[0101] In this case, the controller 11 can perform flexible calibration such that, for example, a sense of presence with a sense of tension occurs for each scene.
[0102] (Example 6) FIG. 11 is a diagram for explaining Example 6. Example 6 has the same configuration and processing content of the data table as Example 5. That is, also in Example 6, similar to Example 5, the subdivided analysis data table DG is used.
[0103] In Example 6, the controller 11 detects the same scene from the content a plurality of times. The controller 11 extracts different sense-of-presence parameters determined for each scene and detection count. The controller 11 estimates the emotion of the user in each scene. The controller 11 associates and records each of the detected plurality of scenes, the sense-of-presence parameters, and the estimated emotion of the user.
[0104] In the example of FIG. 11, the controller 11 detects the scene "dirt driving" a plurality of times. And the sense-of-presence parameter is determined for each number of detections of the scene "dirt driving". That is, in this example, the change in the estimated emotion with respect to the number of occurrences of the same scene, that is, the correlation of the number of occurrences of the estimated emotion, is estimated and used for correction of sound and vibration, etc.
[0105] Example 6 is effective not only when the user receives content as a consumer but also when the user receives content as a developer. For example, it is conceivable that a user who is a developer (or a test player) gets bored while repeatedly experiencing the same scene. According to Example 6, it becomes possible to set a sense-of-presence parameter according to the degree of boredom due to repetition.
[0106] The flow of processing of the information processing apparatus 10 will be described using the flowchart of FIG. 12. FIG. 12 is a flowchart showing the processing procedure executed by the information processing apparatus 10 (controller 11). This processing is executed by the activation of the information processing apparatus 10 (an activation operation such as a content reproduction operation by the user).
[0107] The controller 11 performs activation processing of the information processing system 1 (power supply to each part in the system, various initial settings, etc.) and moves to step S101. The controller 11 starts providing content in step S101 and moves to step S102.
[0108] In step S102, the controller 11 detects the scene type of the content being provided (being reproduced) based on the information of the content being provided to the user and moves to step S103. That is, the controller 11 determines which of the scene types such as "road driving", "dirt driving", "snow road driving", "fall", "explosion", "shooting", "flight" the scene type of the scene that the user is watching or playing corresponds to.
[0109] In step S103, the controller 11 extracts the presence parameter corresponding to the detected scene type from the presence parameter table DA. Then, the controller 11 processes the audio data in the content using this extracted presence parameter to generate an acoustic signal and a vibration signal. Then, these generated signals are output to the acoustic output device 5 and the vibration device 6, and the process proceeds to step S104.
[0110] In step S104, the controller 11 estimates the user's emotion based on the detection data of the first and second sensors 2 and 3, and proceeds to step S105. That is, the controller 11 estimates the emotion type of the user's emotion by a method using brain waves and heartbeats, or a method using images and voices.
[0111] In step S105, the controller 11 associates the reproduction location data of the content, the scene type data, the extracted presence parameter, and the estimated emotion type data to generate association data (temporarily stores it in the association data table DC). Then, this association data is stored in the time-series association data table DD as a data record with the reproduction location data as the primary key, and the process proceeds to step S106. That is, as described with reference to FIG. 6 and the like, the controller 11 combines the reproduction location data of the content, the scene type data, the extracted presence parameter, and the estimated emotion type data into one record and stores it in the time-series association data table DD.
[0112] These data indicate at which position of the content, what kind of emotion the content user had, what was the scene type at that time, and what was the presence parameter used in the processing at that time, so they are very useful data for evaluation and design and development in presence processing. Specifically, for example, these data will be taken into a computer system for design and development (transmitted using a memory card or communication) and used. Note that the analysis results (data generated at each stage in the analysis) in the next step S106 are also very useful data in the same way.
[0113] In step S107, the controller 11 determines whether the provision (reproduction) of the content being provided has ended. If it has ended, the process proceeds to step S108. If it has not ended, the process returns to step S103 and the processing from step S103 to step S107 is repeated.
[0114] In step S108, the controller 11 performs the above-described analysis process on the association data stored in the time-series association data table DD in the reproduction of the content, and calculates correction values for the presence parameters regarding each scene type. Then, these calculated correction values are stored (updated) in the parameter correction value data table DF, and the process proceeds to step S109.
[0115] In step S109, the controller 11 corrects the presence parameters of each scene type in the presence parameter table DA using the correction values stored in the parameter correction value data table DF, and ends the process. After that, the controller 11 processes the acoustic data of the content being provided using these corrected presence parameters for each scene type to generate acoustic data and vibration data. Then, the controller 11 outputs these generated signals to the acoustic output device 5 and the vibration device 6. Through the above-described processing, based on the emotions estimated in the scene of each scene type in the content, the presence parameters for the corresponding scene type are corrected. Therefore, acoustic output and vibration output suitable for the scene can be expected.
[0116] [Other Embodiments] The configuration of the information processing system 1 is not limited to that shown in FIG. 1 and the like. For example, the functions of the information processing device 10 may be provided in a server. In that case, the display device 4, the acoustic output device 5, and the vibration device 6 receive video data, acoustic data, or vibration control data from the server via a terminal device capable of communicating with the server. Also, the first sensor 2 and the second sensor 3 transmit sensor data to the server via the terminal device.
[0117] Alternatively, the content providing device (VR playback device) may be a separate device from the information processing system 1, and content information such as an acoustic signal and a video signal may be input from the content providing device to the information processing system 1 (the information processing system 1 processes the content information to generate a vibration signal or the like).
[0118] In addition, the installation location of the information processing system 1 may be not only a building such as a house or an entertainment facility but also a vehicle. The information processing system 1 provided in the vehicle can provide content to a driver during a break.
[0119] Also, in a vehicle, notifications and warnings to the driver output by video, sound, vibration, etc. can also be regarded as content in a broad sense. For example, when the vehicle is about to contact an object, the information processing system 1 outputs sound from a speaker and vibrates the seat to warn the driver.
[0120] At this time, if the intensity of the sound and vibration is too small, the driver's emotion will not change and the warning will be meaningless. On the other hand, if the intensity of the sound and vibration is too large, the driver may be overly surprised and may make a mistake in the driving operation. According to the information processing system 1, calibration can be performed so that the sound and vibration output as a warning in the vehicle have an appropriate intensity based on the driver's emotion.
[0121] Further effects and modifications can be easily derived by those skilled in the art. Therefore, the broader aspects of the present invention are not limited to the specific details and representative embodiments represented and described as above. Accordingly, various changes can be made without departing from the spirit or scope of the general inventive concept defined by the appended claims and their equivalents.
Explanation of Reference Numerals
[0122] 1 Information processing system 2 First sensor 3 Second sensor 4 Display device 5 Audio output device 6 Vibration device 10 Information processing device 11 Controller 12 Memory 13 Input unit 14 Output unit
Claims
1. Detecting a scene type for each scene in the content; generating vibration information for each scene of the content based on the acoustic information for each scene of the content and vibration parameters corresponding to the detected scene type; An information processing device that outputs the vibration information to a vibration device that generates vibrations, Estimating the emotion of a content user in any scene of the content; generating association information that associates a scene type of the scene with an emotion estimated in the scene; generating correction information for the vibration parameters for each scene type based on the association information; performing a correction process on the vibration parameters using the correction information to generate corrected vibration parameters; The vibration information is generated using the corrected vibration parameters. Information processing device.
2. The vibration parameter correction information is generated based on a data table that defines the correction vibration parameters using the scene type and emotion type as parameters.
2. The information processing device according to claim 1.
3. Get your heart rate information, Acquire brainwave information, Calculating an emotion intensity index value based on the heart rate information as a standard deviation of a heart rate LF (Low Frequency) component; Calculating a wakefulness index value based on the electroencephalogram information as a beta wave / alpha wave of the electroencephalogram; estimating the emotion based on a combination of a magnitude relationship of the emotion intensity index value relative to an emotion intensity threshold and a magnitude relationship of the arousal index value relative to an arousal threshold; 2. The information processing device according to claim 1.
4. The AI model has been trained using training data in which face images are input data and emotions are correct answer data, Acquire a facial image of the content user, Applying the acquired face image of the content user to the AI model to estimate the emotion.
2. The information processing device according to claim 1.
5. Detecting a scene type for each scene in the content; generating vibration information for each scene of the content based on the acoustic information for each scene of the content and vibration parameters corresponding to the detected scene type; An information processing method for outputting vibration information to a vibration device that generates vibrations, comprising: Estimating the emotion of a content user in any scene of the content; generating association information that associates a scene type of the scene with an emotion estimated in the scene; generating correction information for the vibration parameters for each scene type based on the association information; performing a correction process on the vibration parameters using the correction information to generate corrected vibration parameters; The vibration information is generated using the corrected vibration parameters. Information processing methods.
6. Detecting a scene type for each scene in the content; generating vibration information for each scene of the content based on the acoustic information for each scene of the content and vibration parameters corresponding to the detected scene type; An information processing program for outputting the vibration information to a vibration device that generates vibration, Estimating the emotion of a content user in any scene of the content; generating association information that associates a scene type of the scene with an emotion estimated in the scene; generating correction information for the vibration parameters for each scene type based on the association information; performing a correction process on the vibration parameters using the correction information to generate corrected vibration parameters; The vibration information is generated using the corrected vibration parameters. Information processing program.
7. a display device that displays an image based on the image signal; an audio output device that outputs audio based on an audio signal; a vibration device that outputs vibration based on a vibration signal; An information processing system including an information processing device that receives content information from a content providing device, generates an image signal, an audio signal, and a vibration signal, and outputs the generated image signal, audio signal, and vibration signal to the display device, the audio output device, and a vibration device, The information processing device includes: Detecting a scene type for each scene in the content; generating vibration information for each scene of the content based on the acoustic information for each scene of the content and vibration parameters corresponding to the detected scene type; A realistic sensation providing device that outputs the vibration information to a vibration device that generates vibration, Estimating the emotion of a content user in any scene of the content; generating association information that associates a scene type of the scene with an emotion estimated in the scene; generating correction information for the vibration parameters for each scene type based on the association information; performing a correction process on the vibration parameters using the correction information to generate corrected vibration parameters; generating the vibration information using the corrected vibration parameters; Information processing system.
Citation Information
Patent Citations
Bodily feeling experiencing video / sound system
JP2004081357A
Information processor, information processing system, and information processing method
JP2023051201A