Information processing device, information processing system, information processing method, and program
By controlling the amplitude and delay of vibrators based on the directional component of sound sources, the information processing device enhances the realism of vibration experiences in content playback.
Patent Information
- Application Number
- JP2024504315
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-04
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-03-04
AI Technical Summary
Conventional techniques struggle to provide users with a natural sense of vibration positioning, lacking elements that give the user a realistic sensation of vibration propagation.
An information processing device identifies the directional component of a sound source relative to a vibration device in input content and controls the amplitude and delay of each vibrator based on this directional component.
This approach enables the generation of vibrations that simulate propagation, providing users with a more realistic and localized vibration experience.
Smart Images

Figure 0007689622000001 
Figure 0007689622000002 
Figure 0007689622000003
Abstract
Description
[Technical field]
[0001] The present invention relates to an information processing device, an information processing system, and an information processing method. [Background technology]
[0002] Conventionally, there is known a technology that provides users with digital content that includes virtual space experiences such as VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality), so-called XR (Cross Reality) content, using a head mounted display (HMD) etc. XR is a collective term that includes all virtual space technologies including VR, AR, MR, SR (Substitutional Reality), AV (Audio / Visual), etc.
[0003] Furthermore, for example, a technique has been proposed for improving the sense of realism of an image by providing the user with vibrations corresponding to the image the user is viewing (see, for example, Patent Document 1).
[0004] Furthermore, a technology has been proposed in which the vibration of each of a plurality of cells arranged on the seat surface is controlled and a signal is presented to the user (see, for example, Patent Document 2). [Prior art documents] [Patent documents]
[0005] [Patent Document 1] JP 2004-081357 A [Patent Document 2] JP 2021-158392 A Summary of the Invention [Problem to be solved by the invention]
[0006] However, with conventional techniques, it has been difficult to provide the user with a natural sense of vibration positioning.
[0007] Fig. 3 is a diagram showing a conventional method of providing vibration, and is a diagram showing the seat surface of a seat on which a user sits, viewed vertically downward from above the user's head.
[0008] 3, the seat is provided with a vibrator 51a_FL (left front), a vibrator 51a_RL (left rear), a vibrator 51a_FR (right front), and a vibrator 51a_RR (right rear). The output vibration of each vibrator is controlled according to the control content, that is, to the position of the vibration source in the content (position based on the user position) so that the user feels a sense of localization of the vibration.
[0009] For example, if the vibration source (the object emitting vibration) in the content is located to the front right of the user, in order to create a sense of the position of the vibration source, the vibration strength of the vibrator 51 at each position is controlled to, for example, 1 for FL (front left), 0 for RL (rear left), 8 for FR (front right), and 1 for RR (rear right), as shown in FIG. 23.
[0010] In this case, the user feels a strong vibration at the front right of the seat, and is therefore able to recognize that the vibration source is located at the front right.
[0011] However, with this type of vibration delivery method, although the user can sense to some extent the position of the vibration source due to differences in vibration intensity at different seat positions, there are not many elements that give the user the sensation of the vibration propagating, so there is a demand for a vibration delivery method that provides a more realistic sensation and gives the user a more localized sense of the vibration.
[0012] The present invention has been made in view of the above, and has an object to provide a user with a highly realistic vibration when playing back content or the like. [Means for solving the problem]
[0013] In order to solve the above-mentioned problems and achieve the objective, an information processing device according to the present invention identifies a directional component of a sound source relative to a vibration device in input content, and controls the amplitude and delay of each vibrator based on the directional component. Effect of the Invention
[0014] According to the present invention, it is possible to generate vibrations that include components that give the sensation of vibrations propagating, and to provide a user with vibrations that feel more realistic. [Brief description of the drawings]
[0015] [Figure 1] FIG. 1 is a diagram illustrating an overview of an information processing system. [Diagram 2] FIG. 2 is a diagram showing a data flow in the information processing system. [Diagram 3] FIG. 3 is a diagram illustrating an example of the configuration of the vibration device. [Figure 4] FIG. 4 is a diagram showing an outline of the information processing method. [Diagram 5] FIG. 5 is a block diagram of an information processing device. [Figure 6] FIG. 6 is a diagram showing an example of the scene information DB. [Figure 7] FIG. 7 is a diagram showing an example of the scene information DB. [Figure 8] FIG. 8 is a diagram showing an example of the scene information DB. [Figure 9] FIG. 9 is a diagram showing an example of the priority order information DB. [Figure 10] FIG. 10 is a diagram illustrating an example of the parameter information DB. [Figure 11] FIG. 11 is a diagram showing an example of the transducer information DB. [Figure 12] FIG. 12 is a block diagram of the scene detection unit. [Figure 13] FIG. 13 is a block diagram of the priority setting unit. [Figure 14] FIG. 14 is a diagram showing an example of a method for determining a priority object. [Figure 15] FIG. 15 is a block diagram of the parameter extraction unit. [Figure 16] FIG. 16 is a block diagram of the output section. [Figure 17]FIG. 17 is a diagram showing an example of a vibration localization processing method. [Figure 18] FIG. 18 is a diagram illustrating an example of a signal processing method. [Figure 19] FIG. 19 is a flowchart showing a processing procedure executed by the information processing device. [Figure 20] FIG. 20 is a flowchart showing the procedure of the vibration localization process. [Figure 21] FIG. 21 is a diagram showing an example of the transducer information DB. [Figure 22] FIG. 22 is a diagram illustrating an example of a vibration control method. [Figure 23] FIG. 23 is a diagram showing a conventional method for providing vibration. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] Hereinafter, an information processing device, an information processing system, and an information processing method disclosed in the present application will be described in detail with reference to the accompanying drawings. Note that the present invention is not limited to the following embodiments.
[0017] [First embodiment] First, an overview of an information processing system and an information processing method according to an embodiment will be described with reference to Figs. 1, 2, 3, and 4. Fig. 1 is a diagram showing an overview of the information processing system. Fig. 2 is a diagram showing a data flow in the information processing system. Fig. 3 is a diagram showing an example of the configuration of a vibration device. Fig. 4 is a diagram showing an overview of an information processing method. Note that the following describes a case where the XR space (virtual space) is a VR space.
[0018] As shown in FIG. 1, the information processing system 1 includes a display device 3, a speaker 4, and a vibration device 5.
[0019] 2, the information processing device 10 provides video data to the display device 3. The information processing device 10 also provides audio data to the speaker 4. The information processing device 10 also provides vibration data to the vibration device 5.
[0020] 1, the display device 3 is, for example, a head mounted display. The display device 3 is an information processing terminal that presents video data related to XR content provided by the information processing device 10 to the user and allows the user to enjoy a VR experience.
[0021] The display device 3 may be a non-transmissive type that completely covers the field of view, or may be a video-transmissive type or an optical-transmissive type. The display device 3 also has a device that detects changes in the user's internal and external conditions using a sensor unit, such as a camera or a motion sensor.
[0022] The speaker 4 is an audio output device that outputs audio, and is provided, for example, as a headphone type, and is worn on the user's ear. The speaker 4 generates audio data provided from the information processing device 10 as audio. The speaker 4 is not limited to a headphone type, and may be a box type (installed on the floor, etc.). The speaker 4 may also be a stereo audio or multi-channel audio type.
[0023] The vibration device 5 includes a plurality of vibrators. Each vibrator is made up of an electric vibration converter including an electric magnetic circuit and a piezoelectric element, and is provided, for example, on a seat on which a user sits, and vibrates in accordance with vibration data provided by the information processing device 10. The information processing device 10 controls each vibrator of the vibration device 5 individually.
[0024] 3 is a view of the seat surface of the seat on which the user sits, viewed vertically downward from above the user's head. As shown in FIG. 3, the vibrators of the vibration device 5, vibrators 51_FL, 51_RL, 51_FR, and 51_RR, are provided at the left front, left rear, right front, and right rear positions on the seat surface.
[0025] When the user sits on the seat, each transducer comes into contact with a different body part and applies vibrations to the user. For example, transducer 51_FL, transducer 51_RL, transducer 51_FR, and transducer 51_RR apply vibrations to the left thigh, left buttock, right thigh, and right buttock of the user, respectively, seated on the seat.
[0026] By applying the sound from the speaker 4 and the vibration from the vibration device 5, that is, the waves from the wave device, to the content user in a manner that matches the reproduced video, it is possible to enhance the sense of realism regarding the video reproduction.
[0027] The information processing device 10 is configured with a computer, is connected to the display device 3 by wire or wirelessly, and provides images of XR content to the display device 3. In addition, the information processing device 10, for example, acquires changes in the situation detected by a sensor unit provided in the display device 3 at any time, and reflects such changes in the situation in the XR content.
[0028] For example, the information processing device 10 can change the direction of the field of view in the virtual space of the XR content in response to changes in the user's head or line of sight detected by the sensor unit.
[0029] Incidentally, when providing XR content, the sense of realism of the XR content can be improved by emphasizing the sound generated from the speaker 4 in accordance with the scene, or by vibrating the vibration device 5 in accordance with the scene.
[0030] However, the parameters used to control the sense of realism to improve the sense of realism (hereinafter referred to as "realism parameters") had to be set manually after the XR content was produced, requiring a huge amount of work to set the realism parameters.
[0031] Therefore, in the information processing method, it is intended to automate the setting of these realism parameters. For example, as shown in Fig. 4, in the information processing method according to the embodiment, first, a scene that satisfies a predetermined condition is detected from video data and audio data related to the XR content (step S1).
[0032] The predetermined condition here is, for example, a condition regarding whether the corresponding video data or audio data is a scene that requires the setting of a realism parameter, and is defined, for example, by a conditional expression regarding the situation inside the XR content.
[0033] That is, in the information processing method, when the situation inside the XR content satisfies the condition defined by the conditional expression, it detects it as a scene that satisfies a predetermined condition. As a result, the information processing method does not require processing such as detailed analysis of the video data, and therefore it is possible to reduce the processing load of scene detection.
[0034] Next, in the information processing method, a priority order is set for the scenes detected by the scene detection (step S2). Here, the priority order indicates the order of the realism parameters of which scenes should be prioritized. That is, in the information processing method, when multiple scenes overlap in time, the realism parameters of which scenes should be prioritized are defined in advance for each scene.
[0035] As a result, even when a plurality of scenes overlap, it is possible to provide the user with an appropriate sense of realism. As will be described later, in the information processing method, a priority order regarding sound and a priority order regarding vibration are set separately.
[0036] Next, in the information processing method, a realism parameter is extracted for each scene (step S3). For example, in the information processing method, a realism parameter is extracted for each scene using parameter information in which a relationship between the scene and the realism parameter is predefined.
[0037] In this case, the information processing method extracts a corresponding realism parameter according to the priority. Specifically, for example, in the case where a scene with a low priority overlaps with a scene with a high priority, the information processing method extracts the realism parameter of the scene with the high priority.
[0038] In the information processing method, a voice enhancement process is performed to enhance the voice data using a voice enhancement parameter from among the extracted realism parameters (step S4), and the voice data is output to the speaker 4. In addition, in the information processing method, a vibration conversion process is performed to convert the voice data into vibration data, and the vibration data is enhanced using a vibration parameter from among the extracted realism parameters (step S5), and the vibration data is output to the vibration device 5.
[0039] As a result, the information processing method can provide the user with sound that is emphasized in accordance with the scene that the user is viewing, or vibration that corresponds to the scene.
[0040] In this way, in the information processing method according to the embodiment, scenes are detected from the XR content, a priority is set, and then a sense of realism parameter related to wave control including sound processing and vibration processing is extracted for the scene. Therefore, according to the information processing method according to the embodiment, it is possible to automate the setting of the sense of realism parameter related to the improvement of the sense of realism of the content.
[0041] Furthermore, in step S5, the information processing device 10 identifies a directional component of a sound source in the input content with respect to the vibration device 5. Then, the information processing device 10 controls the output vibration of the multiple vibrators based on the identified directional component. This allows the information processing device 10 to provide the user with a sense of vibration localization.
[0042] Next, a configuration example of the information processing device 10 according to the embodiment will be described with reference to Fig. 5. Fig. 5 is a block diagram of the information processing device 10. As shown in Fig. 5, the information processing device 10 includes a control unit 120 and a storage unit 130.
[0043] The storage unit 130 is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. In the example of FIG. 5, the storage unit 130 has an XR content DB (database) 131, a scene information DB 132, a priority order information DB 133, a parameter information DB 134, and a transducer information DB 135.
[0044] The XR content DB 131 is a database that stores a group of XR contents to be displayed on the display device 3. The scene information DB 132 is a database that stores various information related to the scene to be detected.
[0045] 6 to 8 are diagrams showing an example of the scene information DB 132. As shown in Fig. 6, for example, the scene information DB 132 stores information on items such as "detected scene", "condition category", "object", "condition parameter", "threshold value", and "condition formula" in association with each other.
[0046] "Detected scene" indicates the name of the scene to be detected. Note that "detected scene" acts as an identification symbol, and although a code such as a numerical value is normally used, in this example, a name (duplication prohibited) is used to make the explanation easier to understand. "Condition category" indicates the category based on what information the scene is detected based on. In the example shown in the figure, categories are broadly divided into the positional relationship between the user and the object, the user's actions, spatial information about the user's presence, time information about the user's presence, and sound emitted from the object. Note that the user here refers to the operator himself in the XR space.
[0047] "Object" indicates an object for scene detection. In the example shown in the figure, information such as object 1, object 2, user, space 1, space 1+object 3, content 1, object 4, object 5, object 6, etc. corresponds to the object. Here, object 1, object 2, object 3, object 4, object 5, and object 6 each indicate a different object in the XR space. Also, space 1 indicates, for example, a space in the XR space where the user exists, and content 1 indicates, for example, a predetermined event in the XR space.
[0048] "Condition parameters" indicate conditions related to parameters, such as which parameters are used for scene detection. As shown in the figure, information such as distance, angle, speed, acceleration, rotation speed, presence of objects in space, quantity, start time-end time, and sound pattern is associated with each parameter.
[0049] The "threshold" indicates a threshold value corresponding to the condition parameter. Furthermore, the "condition formula" indicates a condition formula for detecting a detection scene, and for example, the relationship between the condition parameter and the threshold value is defined as the condition formula.
[0050] In FIG. 6, for the purpose of explanation, each item value is represented using symbols such as "W," "4," and "w," such as "Scene W," "Object 4," and "Pattern w." In reality, however, data for each item value will be stored in a form that allows the specific meaning to be understood.
[0051] For example, "Scene W," "Scene X," "Scene Y," and "Scene Z" would actually be data such as an "elephant walking scene," a "horse walking scene," a "car driving scene," and a "car making a sharp turn scene," respectively.
[0052] In this case, "Object 4," "Object 5," and "Object 6" would actually be data such as "Horse," "Elephant," and "Car," respectively.
[0053] Furthermore, "Pattern w," "Pattern x," "Pattern y," and "Pattern z" are actually data such as "pattern of the sound of a horse walking," "pattern of the sound of an elephant walking," "pattern of the sound of a car running," and "pattern of the sound of tires squealing."
[0054] The voice pattern is represented by, for example, a feature vector having voice features as elements. For example, the feature may be obtained by performing spectral decomposition on the voice signal (for example, a Mel filter bank or a cepstrum).
[0055] If the similarity (for example, cosine similarity or Euclidean distance) between the feature vectors corresponding to two voice patterns is equal to or greater than a threshold, the two voice patterns can be said to be similar.
[0056] For example, "the audio pattern is similar to pattern w" means that the similarity between the feature vector calculated from the audio occurring in the scene and the feature vector of the audio corresponding to pattern w is equal to or greater than a threshold value.
[0057] The threshold value for the similarity of the sound pattern may also be included in the “threshold value” of the scene information DB 132 .
[0058] In addition, the information processing device 10 may detect a scene by combining, for example, the condition categories or condition parameters shown in Fig. 6. For example, as shown in Fig. 7, a detection scene may be set by combining condition categories of a plurality of scenes, or as shown in Fig. 8, a detection scene may be set by combining condition parameters of a plurality of scenes.
[0059] For example, by combining condition categories and condition parameters in this way, it is possible to simplify the setting of a new detection scene.
[0060] Returning to the description of Fig. 5, the priority order information DB 133 will be described. For example, in the information processing device 10 according to the embodiment, a priority order is set for each scene on a rule basis. The priority order information DB 133 stores various information related to the priority order of the realism parameters. Fig. 9 is a diagram showing an example of the priority order information DB 133.
[0061] 9, for example, the priority information DB 133 stores information items such as "rule number" and "priority rule" in association with each other. The "rule number" indicates a number for identifying a priority rule, and the "priority rule" indicates a rule regarding the priority.
[0062] In the figure, "Give priority to the scene detected earlier" and "Give priority to the scene detected later (switch to the later scene)" indicate that the sense of presence parameter of the scene that comes earlier or later in time is given priority, respectively. This makes it possible to simplify the rules for setting scene priorities, for example.
[0063] Moreover, "give priority to a specific parameter with a larger weight" indicates that the realism parameter of a scene in which either the voice emphasis parameter or the vibration parameter is larger is given priority.
[0064] That is, in this case, the realism parameter extracted for the scene in which the voice emphasis parameter or the vibration parameter is large is set, so that it is possible to provide a realism parameter linked to the voice data or large vibration data that should be highly emphasized.
[0065] In addition, "prioritize the larger weight of each parameter" indicates that the realism parameters of the scene in which the voice emphasis parameters or the vibration parameters are larger are prioritized. In the case of this rule, the voice emphasis parameters and the vibration parameters may each be used for different scenes.
[0066] In other words, in this case, the vibration data and the sound data can be emphasized by the realism parameter having a large value, so that the realism of each of the vibration data and the sound data can be improved. Note that the magnitude of the weight here indicates, for example, the magnitude of the parameter value.
[0067] "Priority is given to parameters of a scene with a shorter duration" indicates that priority is given to the realism parameters of a scene with a shorter duration. When a scene with a longer duration is being played back and a scene with a shorter duration is interrupted, the realism parameters of the scene with the shorter duration are set with priority.
[0068] This allows, for example, a scene having a short duration to be appropriately emphasized. Note that a rule may be set such that a parameter for a scene having a long duration is given priority.
[0069] Additionally, "give priority to scenes with larger amplitude in the low frequencies" indicates that when multiple scenes in which objects are emitting sound occur simultaneously, priority is given to the scene corresponding to an object emitting sound with a larger amplitude in the low frequencies (e.g., below 500 Hz).
[0070] In general, it is considered that the larger the creature, the greater the amplitude of the low-frequency range of the walking sound of the creature. For this reason, for example, if a walking scene of an elephant and a walking scene of a horse are detected, the walking scene of the elephant will be prioritized according to the rule that "the one with the larger amplitude of the low-frequency range is prioritized."
[0071] Additionally, "give priority to scenes with large temporal variations in sound and video" indicates that priority is given to scenes with large variations per unit time in the volume of sound produced by an object or the position of an object in a video.
[0072] Also, "Give priority to scenes of objects close to the center of the field of view" indicates that scenes of objects located close to the center of the screen are given priority in the video of the content. This rule will be explained later with reference to FIG. 14.
[0073] Also, "Give priority to scene X over scene W" indicates that when scene W and scene X are detected, scene X is given priority. In this manner, a person (designer, developer) may manually define a priority rule in advance for two or more specific scenes.
[0074] Returning to the explanation of Fig. 5, the following will explain the parameter information DB 134. The parameter information DB 134 is a database that stores information related to the realism parameters for each scene. Fig. 10 is a diagram showing an example of the parameter information DB 134.
[0075] As shown in FIG. 10, the parameter information DB 134 stores information items such as "scene name," "audio emphasis parameters," and "vibration parameters" in association with one another.
[0076] The "scene name" indicates the name of the detected scene described above, and corresponds to, for example, the "detected scene" shown in Fig. 6. Note that, in order to make the explanation easier to understand, the "scene names" are shown here as an explosion scene, a concert hall scene, an elephant walking scene, a horse walking scene, a car driving scene, and a car turning sharply scene.
[0077] The "voice emphasis parameters" indicate the voice emphasis parameters to be set in the corresponding scene. For example, as shown in Fig. 10, the voice emphasis parameters are stored for each speaker 4 according to the number of speakers 4, such as "for speaker 1", "for speaker 2", etc.
[0078] Also, for each speaker 4, parameter values related to audio processing, such as "delay" and "band emphasis / attenuation" are stored. For example, "delay" indicates a parameter related to the delay time, and "band emphasis / attenuation" indicates a parameter regarding the extent to which sound in which band is emphasized or attenuated.
[0079] The "vibration parameters" indicate parameters related to vibrations set in the corresponding scene. For example, parameters such as "LPF (Low Pass Filter)", "amplitude emphasis coefficient (ω)", and "delay emphasis coefficient (γ)" are stored as the "vibration parameters".
[0080] "LPF" indicates a parameter related to a low-pass filter used for vibration generation (cutoff frequency in the example shown in FIG. 10). "Amplitude emphasis coefficient (ω)" indicates a parameter related to amplification and attenuation of the amplitude of vibration used for vibration generation. "Delay emphasis coefficient (ω)" indicates a parameter related to the delay in the generation time of vibration used for vibration generation.
[0081] Returning to the description of Fig. 5, a description will be given of the transducer information DB 135. The transducer information DB 135 is a database that stores information on the transducer included in the vibration device 5. Fig. 11 is a diagram showing an example of the transducer information DB.
[0082] As shown in FIG. 11, the transducer information DB 135 stores information on items such as "transducer" and "position coordinates" in association with each other.
[0083] "Vibrator" indicates information for identifying a vibrator included in the vibration device 5. Furthermore, "position coordinates" indicates the position of the vibrator by coordinates.
[0084] Here, "FL", "RL", "FR", and "RR" shown in "Vibrator" correspond to the vibrator 51_FL, the vibrator 51_RL, the vibrator 51_FR, and the vibrator 51_RR in Fig. 3, respectively. Furthermore, the "position coordinates" may be set by an installer when each vibrator is installed in the vibration device 5.
[0085] For example, by referring to the transducer information DB 135, the positional relationship between the transducers can be grasped.
[0086] Returning to the explanation of Fig. 5, the control unit 120 will be explained. The control unit 120 is a controller, and is realized, for example, by a CPU (Central Processing Unit) or an MPU (Micro Processing Unit) executing various programs (not shown) stored in the storage unit 11 using a RAM as a working area. The control unit 120 can also be realized, for example, by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0087] The control unit 120 has a content generation unit 121, a rendering processing unit 122, a scene detection unit 123, a priority setting unit 124, a parameter extraction unit 125, and an output unit 126, and realizes or executes the functions and actions of the information processing described below.
[0088] The content generation unit 121 generates a 3D model of the space in the XR content. For example, the content generation unit 121 refers to the XR content DB 131 and generates a 3D model of the space in the XR content according to the user's current field of view in the XR content. The content generation unit 121 passes the generated 3D model to the rendering processing unit 122.
[0089] The rendering processing unit 122 performs rendering processing to convert the 3D model received from the content generation unit 121 into video data and audio data. For example, the rendering processing unit 122 outputs the converted video data to the display device 3 (see FIG. 2) and passes it to the scene detection unit 123. The rendering processing unit 122 also passes the converted audio data to the output unit 126 and the scene detection unit 123. The content generation unit 121 and the rendering processing unit 122 function as a calculation unit that calculates condition data for items of a conditional formula from the content.
[0090] The scene detection unit 123 detects a scene that satisfies a predetermined condition from the input content. For example, the scene detection unit 123 detects a scene for which a realism parameter should be set, using the video data input from the rendering processing unit 122 and a conditional expression stored in the scene information DB 132.
[0091] At this time, for example, the scene detection unit 123 receives coordinate information of an object in the XR space and information on the object type from the rendering processing unit 122, and detects a scene for which a realism parameter should be set using a conditional expression.
[0092] In addition, for example, when the XR content is MR content, the scene detection unit 123 may recognize objects in the MR space or calculate the coordinates of the objects by performing image analysis on an image captured in the MR space.
[0093] Fig. 12 is a block diagram of the scene detection unit 123. As shown in Fig. 12, for example, the scene detection unit 123 includes a scene determination unit 123a and a condition setting unit 123b. The scene determination unit 123a uses each condition data (conditional expression) for scene determination stored in the scene information DB 132 to determine whether or not the situation in the video data satisfies the detection condition for each scene.
[0094] More specifically, for example, as shown in FIG. 6, the scene determination unit 123a determines whether or not the current situation in the XR space corresponds to each predefined detection scene based on data (calculated from the content by the content generation unit 121 or the rendering processing unit 122) for items of conditional expressions such as the positional relationship between the user and the target object (object in the XR space), the user's actions, and spatial information in which the user exists.
[0095] Here, the scene determination unit 123a performs scene detection processing using text information-like data already calculated by the content generation unit 121 or the rendering processing unit 122, such as the user's movement in the XR space, object coordinate information, information about the object type, and spatial information.
[0096] This makes it possible to perform processes such as scene detection and extraction of realism parameters in parallel with processes that have a relatively heavy processing load, such as rendering by the rendering processor 122, even if the CPU performance is relatively low.
[0097] Also, at this time, for example, the scene determination unit 123a may determine whether or not the current situation in the XR space corresponds to each detected scene based on scene determination information including a combination of condition categories as shown in FIG. 7, or a combination of condition parameters as shown in FIG. 8.
[0098] If the scene determination unit 123a determines that the scene corresponds to a detected scene, it passes the detected scene information for the video data to the priority order setting unit 124 (see FIG. 5). If the scene determination unit 123a determines that the scene does not correspond to any detected scene, the scene is not the detected scene, and the realism parameters are returned to the initial state (the realism parameters when the scene is not the detected scene). If the scene determination unit 123a determines that the current situation in the XR space corresponds to multiple detected scenes, it passes the multiple detected scenes determined to be the detected scenes to the priority order setting unit 124.
[0099] Further, here, the case has been described where the scene determining unit 123a determines whether or not a scene is detected based on video data, but the scene determining unit 123a may determine whether or not a scene is detected based on audio data.
[0100] The scene determination unit 123a detects a scene in which a sound is generated from an object from the input content. The detected scenes in this case correspond to the scenes W, X, Y, and Z (the elephant walking scene, the horse walking scene, the car running scene, and the car turning sharply scene) in FIG. 6.
[0101] For example, the scene determination unit 123a calculates the similarity between a feature vector obtained from an audio signal of a content and a predetermined feature vector (eg, pattern w), and determines whether the similarity is equal to or greater than a threshold value.
[0102] The condition setting unit 123b sets various conditional expressions for scene detection. The condition setting unit 123b sets the conditional expressions based on, for example, information input by the creator of the XR content or the user.
[0103] For example, the condition setting unit 123b receives input information from a producer or a user, such as what kind of realism parameters are to be set for what kind of scenes, and converts the situation of the scene into a conditional formula. Then, for each setting of a conditional formula, the condition setting unit 123b writes information about the conditional formula into the scene information DB 132 and writes the corresponding realism parameters into the parameter information DB 134.
[0104] Furthermore, the condition setting unit 123b may set the scene information DB 132 and the parameter information DB 134 in advance based on the content that the user views.
[0105] The condition setting unit 123b can set conditions for detecting a scene in which an object generates a specified sound in a low-frequency region. For example, the condition setting unit 123b adds a record in which a scene including an elephant's walking sound including a sound in a low-frequency region is detected as a detected scene to the scene information DB 132 (corresponding to the record of scene W in FIG. 6).
[0106] The condition setting unit 123b can identify that an object (for example, an elephant) is included in the scene and that sound in the low frequency range is occurring by recognizing the images and sounds included in the content.
[0107] Moreover, the condition setting unit 123b determines the value of the "vibration parameter" in the parameter information DB 134 according to the size of the object and the amplitude for each frequency band in the low frequency region.
[0108] This makes it possible to automate the setting of the scene information DB 132 and the parameter information DB 134.
[0109] Furthermore, the condition setting unit 123b sets a condition based on a scene in which the amplitude of a sound in a low frequency region generated from an object exceeds a threshold value, among scenes in the content.
[0110] For example, the threshold value here may be the same as the threshold value used when cutting off low-frequency regions in noise cancellation.
[0111] The threshold value may also be determined according to the type of content, such as a music video that allows users to mainly listen to music, or an animal documentary that explains the biology of an animal.
[0112] In a music video, it may be better not to generate excessive vibrations in a scene where an elephant is walking, so as not to interfere with the music. On the other hand, in an animal documentary, it may be better to generate vibrations in a scene where an elephant is walking, to create a sense of realism.
[0113] By setting the threshold for music videos lower than the threshold for animal documentaries, the condition setting unit 123b is less likely to regard a scene in a music video in which an elephant is walking as a scene in which vibrations should be generated.
[0114] This makes it possible to generate vibrations that are suited to the content.
[0115] The above-mentioned setting process of the scene information DB 132 and the parameter information DB 134 may be performed by a person, instead of the condition setting unit 123b, actually viewing the content and operating an input device.
[0116] This makes it possible for the information processing device 10 to detect a scene desired by the producer or user, and to set the sense of presence parameters desired by the producer or user for the detected scene.
[0117] 5, the priority order setting unit 124 will be described. The priority order setting unit 124 sets priorities for the scenes detected by the scene detection unit 123.
[0118] For example, the priority setting unit 124 refers to the priority information DB 133 and selects which scene should be given priority in processing when multiple types of scenes are detected simultaneously by the scene detection unit 123. Note that when the scene detection unit 123 detects only one scene, that scene is given the highest priority.
[0119] Fig. 13 is a block diagram of the priority order setting unit 124. For example, as shown in Fig. 13, the priority order setting unit 124 has a timing detection unit 124a and a rule setting unit 124b.
[0120] The timing detection unit 124a detects the occurrence and end timing of a scene detected by the scene detection unit 123. For example, the timing detection unit 124a detects each scene that exists at each time point (including grasping overlapping states), the occurrence timing of an existing scene, the deletion timing of an existing scene, etc., based on the scene information at each time point from the scene detection unit 123. In other words, the timing detection unit 124a grasps the state of all scenes that exist at each time point, including the order of occurrence.
[0121] The rule setting unit 124b sets a priority order of the scenes to be used for determining the realism parameters for the scenes detected by the scene detection unit 123. That is, based on the state of all the existing scenes grasped by the timing detection unit 124a, the rule setting unit 124b sets a priority order for the detected scenes in order to determine which scene-linked parameter is to be used preferentially for the realism parameters to be used at that time point. This allows the information processing device 10 to set the realism parameters according to the priority order.
[0122] That is, in the information processing device 10, by setting a priority condition for each scene in advance, when scene A and scene B overlap in time, it is possible to appropriately determine which scene's realism parameter should be used preferentially.
[0123] For example, the rule setting unit 124b sets the priority of the scene for determining the parameter to be used for each of the voice emphasis parameter and the vibration parameter with reference to the priority information DB 133. At this time, the rule setting unit 124b may set the scene to be used for parameter selection based on the priority rule independent for each speaker 4, for example.
[0124] As a result, the realism parameters are set for each speaker 4 according to its own rules, so that the realism can be further improved compared to a case where the realism parameters are set uniformly.
[0125] Furthermore, rule setting unit 124b passes information on the set rules to parameter extraction unit 125 (see FIG. 5) in association with the video data and audio data.
[0126] 5, the parameter extraction unit 125 will be described. The parameter extraction unit 125 extracts a sense of presence parameter for the scene detected by the scene detection unit 123.
[0127] Fig. 15 is a block diagram of the parameter extraction unit 125. As shown in Fig. 15, the parameter extraction unit 125 has a vibration parameter extraction unit 125a, a voice emphasis parameter extraction unit 125b, and a learning unit 125c.
[0128] The vibration parameter extraction unit 125a refers to the parameter information DB 134 and extracts vibration parameters corresponding to the scene that has been set to the highest priority by the priority setting unit 124. For example, the vibration parameter extraction unit 125a extracts vibration parameters corresponding to the "detected scene" with the highest priority received from the priority setting unit 124 from the parameter information DB 134, thereby extracting vibration parameters corresponding to the scene.
[0129] In other words, when the scene detection unit 123 detects multiple scenes in which objects generating sounds are different from each other and overlap in time, the parameter extraction unit 125 can select a scene with a high priority, i.e., a scene that is estimated to make the user feel more realistic due to vibration, and extract parameters for vibration generation corresponding to that scene. As a result, realistic vibrations can be generated with appropriate parameters even during a content playback period in which multiple scenes overlap.
[0130] Specifically, the scene detection unit 123 can perform such scene selection processing based on the priority rules of the priority information DB shown in Figure 9 and the priority conditions for each scene (which are set and stored in the scene information DB shown in Figure 4).
[0131] For example, when the scene detection unit 123 detects a scene in which an elephant makes a walking sound (an elephant walking scene) and a scene in which a horse makes a walking sound (a horse walking scene), the parameter extraction unit 125 prioritizes the elephant walking scene according to the rule of "prioritizing the scene with a larger amplitude in the low frequency range." As a result, vibrations that reproduce the vibrations caused by an elephant's walking, which are primarily felt in the real world, are applied to the user in the content playback (for example, in a virtual space), and the user can obtain a highly realistic vibration sensation, that is, a sense of vibration close to reality.
[0132] In addition, when the scene detection unit 123 detects multiple scenes that overlap in time and in which an object that generates sound is present, the parameter extraction unit 125 can also apply a method of extracting parameters corresponding to a scene selected from the multiple scenes based on the type and position of the object corresponding to each of the multiple scenes in the images included in the content.
[0133] Specifically, the scene detection unit 123 can perform such scene selection processing by setting the priority rules of the priority information DB shown in FIG. 9 and the priority conditions for each scene (which are set and stored in the scene information DB shown in FIG. 4) (in this example, the function value F(M, d) of the type of object (m) and the distance to the object (d) is added to the priority conditions, and a condition based on the function value F(M, d) (for example, the larger the function value “F(M, d)” is given priority)) in the priority rules).
[0134] A method for determining a prioritized scene based on the position of an object will be described using a specific example shown in Fig. 14. Fig. 14 is a diagram showing an example of a method for determining a prioritized object.
[0135] 14, it is assumed that an image 31 of a content being played back is displayed on the display device 3. An object 311 (a horse) and an object 312 (an elephant) are shown in the image 31. At this time, it is assumed that the scene detection unit 123 has detected both a horse walking scene and an elephant walking scene that satisfy the conditions as target scenes for vibration control.
[0136] Also, assume that the distance from the reference position (the user's position relative to the content image, for example, the position of an avatar corresponding to the user in XR content) to object 311 is L1. Meanwhile, assume that the distance from the reference position to object 312 is L2. Also, assume that the reference vibration strengths (the strength of the low frequency components of the audio signal of the object in the content) of objects 311 and 312 are V1 and V2, respectively. Furthermore, assume that the priority condition is set as "prioritize the one with a larger value of the function F(Ln, Vn) = Vn / (Ln Ln)".
[0137] The distance from the reference position to the object is calculated based on information added to the content (for example, calculated based on the position information of each object used to generate an image in the XR content). The reference vibration strength of the object can be determined by reading the reference vibration strength preset for each object type from a data table in which it is stored according to the type of the target object, or by adding it to the content as content information. In many cases, audio data is added to the content for audio playback, so it is possible to calculate the reference vibration strength based on the low-frequency characteristics (audio intensity level, low-frequency signal level, etc.) of the audio data (the vibration mode is highly correlated with the low-frequency components of the audio, and vibration is often generated based on the low-frequency components of the audio).
[0138] In this way, the information processing device 10 can estimate the low-frequency characteristics of the sound generated by the vibration generating object in the content. In this case, the information processing device 10 selects the vibration generating object based on the estimated low-frequency characteristics. This makes it possible to select a more appropriate vibration generating object.
[0139] For example, the low-frequency characteristic of the sound is a low-frequency signal level. In this case, the information processing device 10 selects a vibration generating object whose estimated low-frequency signal level exceeds a threshold. The information processing device 10 can extract the low-frequency signal level from the sound data. This makes it possible to easily select a vibration generating object by using the low-frequency signal level included in the sound data.
[0140] In addition, the threshold value of the low-frequency signal level is set according to the type of content. As mentioned above, it is often better to generate vibrations in a music video compared to an animal documentary, even if the object is the same. In this way, it is possible to select a vibration object suitable for the content type (music video, animal documentary, etc.).
[0141] In this case, if the relationship between the function values of the object 311 (horse) and the object 312 (elephant) is function F(L1, V1)>function F(L2, V2), a scene in which the object 311 generates sound (vibration), that is, a horse walking scene, is preferentially selected, and the parameter extraction unit 125 extracts vibration parameters corresponding to the horse walking scene. Then, the vibration corresponding to the horse walking scene is applied to the user. After that, for example, if the object 312 (elephant) approaches the reference position and the relationship changes to function F(L1, V1)<function F(L2, V2), a scene in which the object 311 generates sound (vibration), that is, an elephant walking scene, is preferentially selected, and the parameter extraction unit 125 extracts vibration parameters corresponding to the elephant walking scene. Then, the vibration corresponding to the elephant walking scene is applied to the user.
[0142] In addition, when the function F(Ln, Vn) is smaller than a predetermined threshold value, that is, when the vibration caused by the object at the user's position in the content (such as the virtual space of the game) is small (the user does not feel it much, i.e. there is little need to apply vibration), it is also effective to use a method of not selecting it as an object that generates vibration. In other words, it is also effective to use a method of selecting only objects of content that cause a certain degree of vibration caused by the object at the user's position in the content (such as the virtual space of the game) as an object that generates vibration. In other words, it is effective to select an object that has a large effect on the vibration signal generated from the object candidate that is a candidate for the vibration generating object (a vibration object whose vibration is strongly felt by the user).
[0143] In this way, the information processing device 10 can estimate the candidate object that has a large influence on the vibration signal generated from the candidate object that is the vibration generating object, and select it as the vibration generating object. As a result, vibration that matches the user's sensation in real space can be applied to the user, making it possible to play content with a rich sense of realism.
[0144] In this case, it is preferable to change the threshold value for selecting an object that generates vibration based on the type of content. In other words, depending on the content, it may be preferable to suppress or emphasize the reproduction of vibrations caused by objects that appear in the content, and it is preferable to adjust the determination content (judgment level) of the object that generates vibration.
[0145] In other words, the principle of vibration generation is as follows: The object that will generate vibration in the content (in each scene) is determined based on the content. Then, a vibration signal (vibration data) is generated (the low-frequency components of the object's audio signal are extracted and appropriately amplified, etc.) based on the acoustic signal corresponding to the determined object (audio data of the object included in the content, or audio data of the object generated from audio data in the scene) (e.g., extracted by filtering the low-frequency range).
[0146] In addition, as a method for determining the object that generates vibration, the low-frequency characteristics (e.g., volume level) of the voice produced by the sound-generating object in the content are estimated (in the above example, estimation is based on a reference vibration strength based on the type of object and the distance between a reference position (such as the user's location in the content's virtual space) and the object), and the object is determined (the sound-generating object with the greater low-frequency volume level of the voice produced is determined as the object that generates vibration).
[0147] In this way, by determining the prioritized scene based on the position of the object, vibrations that are more suited to the user's visual intuition, that is, vibrations that match the user's sensations in real space, can be applied to the user, making it possible to play content with a rich sense of realism.
[0148] At this time, the vibration parameter extraction unit 125a extracts vibration parameters corresponding to each of the vibration devices 5. This makes it possible to further improve the sense of realism compared to the case where vibration parameters are uniformly extracted.
[0149] The voice emphasis parameter extraction unit 125b refers to the parameter information DB 134 and extracts a voice emphasis parameter corresponding to a scene that has been set to the highest priority by the priority setting unit 124. The voice emphasis parameter extraction unit 125b extracts a voice emphasis parameter individually for each speaker 4, and determines a voice emphasis parameter to be extracted based on the priority set by the priority setting unit 124 (based on the scene with the highest priority) in the same manner as the vibration parameter extraction unit 125a.
[0150] The learning unit 125c learns the relationship between the scenes and the realism parameters stored in the parameter information DB 134. For example, the learning unit 125c learns the relationship between the scenes and the realism parameters by performing machine learning on each scene stored in the parameter information DB 134 and each corresponding realism parameter using the user's reaction to the realism control by the parameter as learning data.
[0151] In this case, for example, the learning unit 125c may use user evaluations of the realism parameters (user adjustment operations after realism control, user input such as questionnaires) as learning data. That is, the learning unit 125c may learn the relationship between the scene and the realism parameters from the viewpoint of what kind of realism parameters should be set for what kind of scene to obtain a high user evaluation (i.e., whether a high realism was obtained).
[0152] Furthermore, the learning unit 125c can also determine what kind of realism parameters should be set when a new scene is input, based on the learning results. As a specific example, the realism parameters for a fireworks scene can be determined using the learning results of realism control of a similar situation such as an explosion scene. Also, it is possible to learn rules regarding priority order based on the presence or absence and the degree of factors that change the priority order in the user's adjustment operation after realism control or in user input such as a questionnaire (when the user's adjustment operation approaches parameters corresponding to other scenes that exist simultaneously, or when there is a response in the questionnaire that other scenes should be prioritized, etc.).
[0153] This enables the information processing device 10 to automatically optimize, for example, rules regarding priority order and realistic parameters.
[0154] 5, the following describes the output unit 126. The output unit 126 outputs the realism parameters extracted by the parameter extraction unit 125 to the speaker 4 and the vibration device 5.
[0155] Fig. 16 is a block diagram of the output unit 126. As shown in Fig. 16, the output unit 126 has a voice emphasis processing unit 126a, a voice vibration conversion processing unit 126b, and a vibration localization processing unit 126c.
[0156] The voice enhancement processing unit 126a performs enhancement processing on the voice data received from the rendering processing unit 122 using the voice enhancement parameters extracted by the parameter extraction unit 125. For example, the voice enhancement processing unit 126a performs enhancement processing on the voice data by performing delay or band enhancement / attenuation processing based on the voice enhancement parameters.
[0157] At this time, the voice emphasis processing unit 126a performs voice emphasis processing for each speaker 4, and outputs the voice data that has been subjected to the voice emphasis processing to each corresponding speaker 4.
[0158] The sound vibration conversion processing unit 126b converts the sound data received from the rendering processing unit 122 into vibration data by performing band limiting processing suitable for vibration such as an LPF.
[0159] The vibration localization processing unit 126c performs processing related to the localization of vibration on the vibration data obtained by the conversion by the sound vibration conversion processing unit 126b. The vibration localization processing unit 126c then outputs the vibration data for each transducer that has been subjected to amplitude and delay processing by this processing. The vibration device 5 vibrates each transducer according to the vibration data output by the vibration localization processing unit 126c.
[0160] A vibration localization processing method by the vibration localization processing unit 126c will be described with reference to Fig. 17. Fig. 17 is a diagram showing an example of the vibration localization processing method.
[0161] 17, first, the vibration localization processing unit 126c identifies the directional component of the vibration to be provided to the user (content viewer) (step S11). Specifically, since the sense of localization of the vibration is based on the location of an object that is the vibration source, the location (direction) of the object (vibration source) is estimated from the directional component of the sound based on the same object, and the directional component of the vibration is estimated (identified) from the estimated location and the user position (user position in the content space).
[0162] In addition, there may be multiple vibration sources (objects) of vibrations provided to the user in the content, but for ease of understanding, the process will be described in which one main vibration source (object expected to have the greatest effect of improving the sense of realism) is selected by the above-mentioned method. Also, by performing similar processes in parallel for multiple vibration sources, it is possible to effectively provide vibrations based on multiple vibration sources to the user and play back content with a rich sense of realism.
[0163] Therefore, the direction 52 of the virtual vibration source (which reproduces the state in which vibration is generated from this virtual vibration source) based on the user is the direction from the user to the object that is the sound source in the XR space (virtual space), that is, the directional component of the sound.
[0164] The vibration localization processing unit 126c can specify the directional component of the sound (vibration) based on the position data of the sound source (position of the target object) received from the rendering processing unit 122, for example, in the same manner as in the case of the sound localization processing.
[0165] Also, for example, the vibration localization processing unit 126c can identify the position of an object (sound source) based on the spectrum of each of the audio signals of multiple channels included in the audio data, and identify the directional component of the sound (vibration) based on the identified position.
[0166] Furthermore, the vibration localization processing unit 126c can identify the directional component of the sound (vibration) based on metadata of the content (metadata including data indicating the position of the target object).
[0167] In other words, content developed using a 3D engine contains information indicating the position of the object in virtual space as well as data on the sound generated by the object at each time.
[0168] For example, in the case of a horse walking scene, the data of the content for that scene includes data on the horse's hoofbeats and data on the horse's position (as metadata), so the vibration localization processing unit 126c uses this horse's position data to identify the sound source position of the horse's hoofbeats (the position of the horse, which is the sound source).
[0169] Then, the vibration localization processing unit 126c identifies the direction connecting the user's position in the virtual space to the sound source position of the horse's hoofbeats (the horse's position) as the directional component of the sound, and determines this as the directional component of the vibration (localization sense direction).
[0170] In addition, by performing image recognition processing on the image of the content, it is possible to recognize the sound source object and its position, identify the directional component of the sound, and determine this as the directional component of the vibration (direction of perceived positioning).
[0171] Next, the vibration localization processing unit 126c determines various processing data such as coefficient values, correction values, etc., used for vibration control of each vibrator 51 of the vibrating device 5 (vibration data (signal) generation processing).
[0172] For example, since each vibrator 51 has individual differences in characteristics (relationship between input signal and vibration output, for example, ratio between input signal level and vibration output level), correction data for correcting the characteristic difference is determined. Specifically, since the output vibration level has a large effect in this embodiment, vibrator characteristic data is determined based on the ratio between the input signal level and vibration output level (amplitude) (hereinafter referred to as vibrator sensitivity).
[0173] In addition, the transducer sensitivity data can be calculated by measuring the vibration amplitude when a test vibration signal is applied to the transducer and using the test vibration signal amplitude and the vibration amplitude, and the calculated data is stored in the memory unit 130 (transducer information DB 135) for use.
[0174] In addition, the vibration localization processing unit 126c determines sensitivity characteristic correction data for correcting differences in sensitivity characteristics, which are the characteristics of the sensation received by the user in response to vibration, and sensitivity characteristic correction data for correcting differences in vibration transmission characteristics to the user due to the contact state between the user and each vibrator.
[0175] One type of sensitivity characteristic correction data is data for correcting differences in how users feel vibration due to individual differences or differences in body parts, and the vibration localization processing unit 126c determines the sensitivity characteristic, which is the intensity characteristic of the vibration sensation, as the sensitivity characteristic correction data.
[0176] The sensitivity characteristics can be determined by a user's input operation before viewing the content. Specifically, the sensitivity characteristics can be measured by a method in which vibrations of a predetermined intensity are provided to the user from each vibrator, and the user inputs the sensation of the vibration.
[0177] In addition, the other sensitivity characteristic correction data is data for correcting differences in how vibrations are felt depending on the contact state between the user and each transducer 51, and in this embodiment, it is pressure that each transducer receives when the user is seated, which has a large effect on the intensity characteristics of the vibration sensation, that is, pressure distribution data on the seating surface when the user is seated, and the vibration localization processing unit 126c determines the pressure value of the portion of the seating surface where each transducer 51 is installed as the sensitivity characteristic correction data.
[0178] The pressure value can be determined by a method of measuring the pressure by placing a pressure sensor on the seat surface on which the user sits when viewing content.
[0179] It is also possible to determine the sensitivity characteristic correction data by combining the sensitivity characteristic and the pressure value, for example by providing the user with vibrations of a predetermined strength from each vibrator and having the user input their sensation of the vibrations.
[0180] In addition, since the sensitivity characteristic correction data is correction data according to the characteristics and state of the user (the user's seated state), the ratio of the vibration level (the level of seismic intensity when a vibration signal is input to a transducer with a standard characteristic) to the user's sensation (vibration level) is hereinafter referred to as the user sensitivity. This user sensitivity is stored in the storage unit 130 (transducer information DB 135) and is used when playing back content.
[0181] Then, the vibration localization processing unit 126c calculates the output level correction value of each vibrator using the above-mentioned vibrator sensitivity and user sensitivity, and stores it in the storage unit 130 (vibrator information DB 135). Specifically, the vibration localization processing unit 126c stores the inverse value of the product of the vibrator sensitivity and the user sensitivity as the output level correction value in the storage unit 130 (vibrator information DB 135). In other words, each vibrator vibrates based on vibration data, but the vibration sensitivity characteristic (the relationship between the vibration data and the user's vibration intensity feeling, in this case the characteristic element (characteristic) of the vibrator is also taken into consideration) of how the user feels the vibration (what vibration level) is stored in the storage unit 130 as the output level correction value. Since the transducer sensitivity is the ratio of the vibration signal level to the vibration level (amplitude), and the user sensitivity is the ratio of the vibration level (amplitude) to the user's sensation, if a vibration signal of the same vibration signal level is corrected (divided) by the output level correction value corresponding to each transducer and input to each transducer, the user will feel the same level of vibration from each transducer. In the example shown in S12 of Fig. 17, the output level correction value 61A for each transducer is calculated as 2 for transducer 51FL, 4 for transducer 51FR, 1 for transducer 51RL, and 3 for transducer 51RR.
[0182] Then, the vibration localization processing unit 126c performs signal processing (step S13). The signal processing will be described with reference to FIG. 18. FIG. 18 is a diagram showing the principle of the signal processing method in this embodiment, and the processing method is realized by the control unit 120 (CPU) executing an arithmetic expression and a processing program based on this principle. In order to make the explanation easier to understand, the processing is performed in a horizontal two-dimensional space in the content reproduction space (processing while ignoring the height direction). Note that most contents are widely distributed on the plane (ground) of the vibration object, and the movement direction is often on the plane (ground), so processing in a horizontal two-dimensional space is sufficient for approximation processing.
[0183] 18, first, the vibration localization processing unit 126c plots the position of each transducer in a coordinate space for calculation processing (step S21). That is, the vibration localization processing unit 126c plots points according to the position coordinate data (points 53_FL, 53_RL, 53_FR, and 53_RR) of each transducer acquired from the transducer information DB 135.
[0184] Next, the vibration localization processing unit 126c calculates the center of gravity of the position coordinate points of each of the plotted multiple transducers (average coordinates of X and Y coordinate values of each point). In addition, the vibration localization processing unit 126c draws straight lines 535a, 535b, 535c, and 535d (lines other than diagonals) that connect the position coordinate points of each of the multiple transducers and form the periphery of a polygon (rectangle). Furthermore, the vibration localization processing unit 126c draws a straight line 525 that passes through the center of gravity and extends in the direction 52 determined in step S11 of FIG. 17.
[0185] Then, the vibration localization processing unit 126c plots the intersections (points 531 and 532) between the straight line 525 and the straight lines 535a and 535c (step S22).
[0186] Of the line segments whose intersections are plotted in step S22, the line segment (straight line 535a) that passes through point 531 on the direction 52 side is referred to as the no-delay side line segment (535a). In addition, the end points (points 53_FL and 53_FR) of the no-delay side line segment 535a are referred to as the no-delay side points (53_FL, point 53_FR).
[0187] Among the line segments whose intersections are plotted in step S22, the line segment (straight line 535c) that passes through point 532 on the opposite side of direction 52 is called the delayed side line segment (535c). The end points (points 53_RL and 53_RR) of the delayed side line segment 535c are called the delayed side points (53_RL, point 53_RR).
[0188] In addition, the figure formed by the straight lines (straight lines forming the perimeter) connecting the position coordinate points of each transducer does not have to be a rectangle, and may be a polygon or polyhedron other than a rectangle. In other words, it will be a polygon according to the number of transducers to be controlled (for example, a pentagon when there are five transducers to be controlled). Among the sides of the polygon that pass through the center of gravity of the polygon and intersect with a straight line 525 extending in the direction 52 obtained in step S11 of FIG. 17, the side on the direction 52 side will be a line segment on the non-delay side, and the side opposite to the direction 52 will be a line segment on the delayed side. In addition, the end points of the line segment on the non-delay side will be points on the non-delay side, and the end points of the line segment on the delayed side will be points on the delayed side.
[0189] In addition, when selecting a line segment on the non-delay side and a line segment on the delayed side from the sides of the polygon that pass through the center of gravity of the polygon and intersect with the straight line 525 extending in the direction 52 determined in step S11 of FIG. 17, the selected line segments are not adjacent to each other. Therefore, as described later, one transducer will not be subjected to both vibration control for the non-delay side transducer and vibration control for the delayed side transducer, so that the calculation processing for control is simplified. In addition, since one transducer will not share the operations of two transducers, the control accuracy is also improved. In addition, since the center of gravity is in an equivalent relationship with each side of the polygon, the processing content is the same for vibration sources in any direction, so that the effect of making it easier to create a processing program can be expected.
[0190] Next, the vibration localization processing unit 126c performs processing related to the control of the vibration perception position based on the technical concept of phantom sensation (Phs). Phantom sensation is that "when the same stimulation (e.g., vibration) is given to two points at the same time, the user feels as if the stimulation is being received at the center of the two points. Also, when the magnitude of each stimulation (e.g., amplitude in the case of vibration) differs, the point where the stimulation is felt (hereinafter referred to as the stimulation sensitive point) moves toward the larger stimulation." The position of the stimulation sensitive point is approximately estimated to be inversely proportional to the distance ratio from each stimulation point to the stimulation intensity ratio (amplitude in the case of vibration). In this embodiment, the vibration stimulation given to the user is controlled based on this concept.
[0191] Furthermore, vibration localization processor 126c performs processing related to control of vibration directional sense based on the technical concept of haptic apparent movement. Haptic apparent movement is a movement sensation that is generated by providing a time difference between vibrations at two points.
[0192] That is, the vibration localization processing unit 126c performs processing to generate vibrations that give the user a highly realistic sense of localization based on the concepts of phantom sensation and haptic apparent movement.
[0193] 18 are set as vibration perception positions, and a time difference is provided between the vibration times at points 531 and 532, so that the user feels a sense of movement between points 531 and 532. In other words, the user feels a localized vibration moving in the direction of the straight line connecting points 531 and 532, that is, to the position of the vibration source (the vibration generating object in the content).
[0194] First, processing based on the technical concept of phantom sensation will be described in detail using a concrete example.
[0195] The vibration localization processing unit 126c performs processing to set the stimulation sensitive points to points 531 and 532. Based on the technical concept of phantom sensation, when the ratio of the distance (L1) between the point 531 and the transducer position 53_FL to the distance (L2) between the point 531 and the transducer position 53_FR, and the ratio of the amplitude generated by the transducer 51_FL to the amplitude generated by the transducer 51_FR are inversely related, the stimulation sensitive point becomes the point 531. Therefore, the vibration localization processing unit 126c calculates L2 / (L1+L2) as the correction value 60AFL (correction value to be integrated into the vibration signal) for the transducer 51_FL. Also, the vibration localization processing unit 126c calculates L1 / (L1+L2) as the correction value 60AFR for the transducer 51_FR.
[0196] The vibration localization processing unit 126c performs the same processing for the processing in which the stimulation sensitive point is the point 532, and calculates L4 / (L3+L4) as the correction value 60ARL for the transducer 51_RL. The vibration localization processing unit 126c also performs the same processing for the stimulation sensitive point 532, and calculates L3 / (L3+L4) as the correction value 60ARR for the transducer 51_RR. Note that L3 is the distance between the point 532 and the transducer position 53_RL, and L4 is the distance between the point 532 and the transducer position 53_RR.
[0197] Therefore, by outputting a vibration signal generated by accumulating the vibration data VD generated by the above-mentioned method by the correction value 60AFL to the transducer 51_FL and outputting a vibration signal generated by accumulating the vibration data BD by the correction value 60AFR to the transducer 51_FR, the user's stimulation sensitive point becomes the position of point 531. Similarly, by outputting a vibration signal generated by accumulating the vibration data VD by the correction value 60ARL to the transducer 51_RL and outputting a vibration signal generated by accumulating the vibration data BD by the correction value 60ARR to the transducer 51_RR, the user's stimulation sensitive point becomes the position of point 532.
[0198] For example, if the above-mentioned distances L1, L2, L3, and L4 are 3k, 2k, 2k, and 3k, respectively, the correction value 60AFL for the transducer 51_FL is 3k / (3k+2k)=0.6, and the correction value 60AFR for the transducer 51_FR is 2k / (3k+2k)=0.4.
[0199] Further, the correction value 60AFL for the transducer 51_RL is 2k / (3k+2k)=0.4, and the correction value 60AFR for the transducer 51_RR is 3k / (3k+2k)=0.6.
[0200] However, in reality, due to individual differences in vibrators, the sensitivity of the user, and the state of sitting on the vibrating seat, errors occur in the intensity of the vibration felt by the user, and the user's stimulation sensitive points end up being points 541 and 542.
[0201] Therefore, the vibration localization processing unit 126c corrects the vibration signal to each transducer using the output level correction values (61AFL, 61AFR, 61ARL, 61ARR) of each transducer (51_FL, 51_FR, 51_RL, 51_RR) calculated in advance by the processing described in Figure 17.
[0202] Specifically, the vibration localization processing unit 126c sets the vibration data for the transducer 51_FL as an integrated value of the vibration data VD, the correction value 60ARL, and the correction value 61AFL, and outputs a vibration signal to the transducer 51_FL. Similarly, the vibration localization processing unit 126c sets the vibration data for the transducer 51_FR as an integrated value of the vibration data VD, the correction value 60AFR, and the correction value 61AFR, and outputs a vibration signal to the transducer 51_FR. The vibration localization processing unit 126c sets the vibration data for the transducer 51_RL as an integrated value of the vibration data VD, the correction value 60ARL, and the correction value 61ARL, and outputs a vibration signal to the transducer 51_RL. The vibration localization processing unit 126c sets the vibration data for the transducer 51_RR as an integrated value of the vibration data VD, the correction value 60ARR, and the correction value 61ARR, and outputs a vibration signal to the transducer 51_RR.
[0203] For example, if the above-mentioned correction values 61AFR, 61AFL, 61ARL, 61ARR are 2, 4, 1, 3 as shown in step S12 of Figure 17, and the above-mentioned distances L1, L2, L3, L4 are 3k, 2k, 2k, 3k, respectively, the vibration data 53DFL, 53DFR, 53DRL, 53DRR of the vibration signals output to each transducer 51_FL, 51_FR, 51_RL, 51_RR will be as follows, assuming that the original vibration data is VD, and vibration signals based on these vibration data 53D are output to each transducer 51. 53DFL=VD·3k / (3k+2k) / 2=0.3·VD 53DFR=VD·2k / (3k+2k) / 4=0.1·VD 53DRL=VD·2k / (2k+3k) / 1=0.4·VD 53DRR=VD·3k / (2k+3k) / 3=0.2·VD
[0204] This corrects errors due to individual differences in the transducers and the user's sensitivity, and as shown in step S24 of FIG. 18, the user's stimulation sensitive points move from points 541 and 542 to the desired positions of points 531 and 532.
[0205] In this way, the vibration localization processing unit 126c controls the amplitude and delay of the output vibration of each transducer based on the arrangement of each transducer.
[0206] This makes it possible to control the amplitude to match the actual placement of the transducers, allowing the user to feel a more natural vibration positioning.
[0207] The vibration localization processing unit 126c controls the amplitude and delay of the output vibration of each transducer based on the user's vibration sensitivity characteristics to the output vibration of each transducer.
[0208] For example, the information processing device 10 stores in advance the sensitivity characteristics for each part of the user's body. Then, the vibration localization processing unit 126c uses different sensitivity characteristics depending on whether the part to which each vibrator is in close contact is the left or right side of the user's body, or the thigh or buttocks.
[0209] In this way, it is possible to control the amplitude to match the actual vibration sensitivity characteristics of the user, that is, to control the amplitude taking into consideration the relationship between the vibration signal and the vibration sensation felt by the user, making it possible to bring the way the user feels the vibration closer to what was intended in the design.
[0210] Furthermore, the vibration sensitivity characteristics are determined by taking into consideration the individual differences between transducers and the individual differences between users.
[0211] For example, vibration sensitivity characteristics are estimated based on the relationship between the input signal and output vibration level of each vibrator, the user's weight, physical condition, posture, etc., and the vibration localization processing unit 126c uses the vibration sensitivity characteristics to control the amplitude of each vibrator.
[0212] This makes it possible to provide vibrations that take into account the individual differences between transducers and the personal differences between users.
[0213] It should be noted that it is difficult to set the posture state of the user when viewing content through user input, etc. Therefore, the vibration localization processing unit 126c vibrates the vibrator for calibration in the posture state when the user is viewing content, and measures the vibration sensitivity characteristic of the user.
[0214] For example, the information processing device 10 instructs the user to assume a viewing posture, and then has the user actually view sample content (for calibrating vibration sensitivity characteristics for the user's posture) and causes the vibrator to generate vibrations for calibration. Then, based on the user's impressions or bio-information, the information processing device 10 estimates a vibration sensitivity characteristic correction value related to the user's viewing posture. The information processing device 10 then stores the obtained vibration sensitivity characteristic correction value and uses it in subsequent calculation processing of the vibration sensitivity characteristic.
[0215] This makes it possible to provide vibrations that are precisely tailored to the posture of each user when viewing content.
[0216] Furthermore, the vibration localization processing unit 126c corrects the amplitudes of all the transducers in accordance with the scene, based on the amplitude emphasis coefficient acquired from the parameter information DB 134.
[0217] Next, processing based on the technical concept of tactile apparent movement will be described in detail using a concrete example.
[0218] The vibration localization processing unit 126c calculates the delay time (Δt: delay time from the vibration timing of the transducers 51_FL and 51_FR corresponding to the points on the non-delay side) of the transducers corresponding to the points on the delayed side, in this example, the transducers 51_RL and 51_RR (step S25). Note that the delay time from the vibration generation timing (sound generation of the same target object) in the content of the transducers 51_FL and 51_FR corresponding to the points on the non-delay side is set to 0, but it is effective to delay or advance the vibration generation timing of the transducers 51_FL and 51_FR corresponding to the points on the non-delay side depending on the scene of the content.
[0219] Although a predetermined fixed delay time is effective in improving the sense of realism, in order to achieve greater effectiveness, the delay time may be calculated using, for example, the following formula: Delay time Δti = ai · yi · Y (i indicates each timing)
[0220] Here, ai is a value indicating whether or not a delay is required, and is 1 when a delay process is required, and is 0 when it is not required. Also, yi is a vibration emphasis coefficient, which is a value for appropriately emphasizing the vibration generated by the vibration generating object in the target scene of the content. For example, in a scene where it is desired to strongly emphasize the vibration, the emphasis coefficient ai becomes a large value, and the delay time is also lengthened according to the degree of emphasis, so that the difference is easily felt. Also, Y is a theoretical value of the time required for the vibration to travel the distance between the points 531 and 532 that are the stimulation sensitive points, but it is advantageous to use an appropriate constant to reduce the processing load. Note that, as the constant for Y, the intermediate value of the distance between the points 531 and 532 (the average of the distance between the transducer position 53_FL and the transducer position 53_RL and the distance between the transducer position 53_FR and the transducer position 53_RL) or a value determined to be appropriate by a sensitivity test or the like may be used.
[0221] These delay necessity value ai and emphasis coefficient yi are values that are determined according to the state of the vibration generating object in the target scene of the content, and are determined based on, for example, the results of image analysis of the content image, the results of analysis of the content audio (especially the audio generated by the vibration generating object), or additional information of the content (which is added to the content in advance as control data), etc.
[0222] The vibration localization processing unit 126c calculates the delay time Δti for each scene using this formula, and outputs a corresponding vibration signal to each transducer 51 at each timing based on the calculated delay time Δti, thereby vibrating each transducer 51 (step S25).
[0223] For example, in the above vibration data example, if the calculated delay time Δti is 1 second, each transducer will vibrate as follows. Transducer 51_FL: Vibration data 0.3 VD, vibration timing 0 seconds (delay time from the playback timing of the target scene in the content, same for the following transducers) Transducer 51_FR: Vibration data 0.1 VD, vibration timing 0 seconds Transducer 51_RL: Vibration data 0.4 VD, vibration timing 1 second Transducer 51_RR: Vibration data 0.2 VD, vibration timing 1 second
[0224] As a result, the user feels vibration at point 531 when a target scene in the content is played back, and one second later at point 532, so that the user feels vibration in the direction from the vibration source. Therefore, the user can properly sense the position of the vibration source (vibration generating object), and can enjoy content playback with a rich sense of realism.
[0225] Next, a modified example of the above-mentioned delayed drive of the transducer 51 will be described. In the above-mentioned processing example, the vibration localization processing unit 126c (output unit 126) performs processing to move the vibration discontinuously from the transducer corresponding to the point on the non-delay side to the transducer corresponding to the point on the delayed side, but in this example, the user is made to feel the sensation that the vibration position is gradually moving. This processing is also based on the technical idea of phantom sensation.
[0226] Specifically, the vibration localization processing unit 126c attenuates the amplitudes of the transducers 51_FL and 51_FR corresponding to the points on the non-delay side over the delay time Δt. Also, the vibration localization processing unit 126c increases the amplitudes of the transducers 51_RL and 51_RR corresponding to the points on the delayed side over the delay time Δt.
[0227] For example, in the vibration data example described above, if the delay time Δti is 1 second, each vibrator will vibrate based on the following vibration data T seconds (T is equal to or less than the delay time Δti (1 second)) after the playback timing of the target scene in the content. Oscillator 51_FL: Vibration data 0.3·VD·((1-T) / 1) Transducer 51_FR: Vibration data 0.1·VD·((1-T) / 1) Transducer 51_RL: Vibration data 0.4 VD (T / 1) Transducer 51_RR: Vibration data 0.2 VD (T / 1)
[0228] It is also effective to set the final attenuation value for the transducers 51_FL and 51_FR corresponding to the points on the non-delay side to a moderately weak sound level rather than a silent level, or to set the final attenuation value before the delay time Δt has elapsed from the playback timing of the target scene in the content. It is also effective to start amplitude boosting for the transducers 51_RL and 51_RR corresponding to the points on the delayed side from a moderately weak sound level rather than a silent level, or to set the attenuation start value after a predetermined time (delay time Δt or less) has elapsed from the playback timing of the target scene in the content.
[0229] Thus, according to this embodiment, based on the technical concepts of phantom sensation and tactile apparent movement, the content viewing user can be made to feel the position of the vibration source and the movement of the vibration appropriately according to the content, so that the user can enjoy realistic playback of the content.
[0230] In addition, these operations can be conceptualized as follows: "controlling the amplitude of output vibration of each transducer so that the positional relationship (line connecting points 551 and 552) between a first resultant vibration position (point 551) determined based on the vibration level of each transducer in a first transducer group consisting of a plurality of transducers (transducers 51_FL and transducer 51_FR) and a first resultant vibration position (point 552) determined based on the vibration level of each transducer in a second transducer group consisting of a plurality of transducers (transducers 51_RL and transducer 51_RR) coincides with the specified directional component (direction 52) of the vibration source; The delay in the output vibration of each transducer (the vibration timing of transducer 51_FL and transducer 51_FR (e.g., delay 0 from vibration occurrence in the content) and the vibration timing of transducer 51_RL and transducer 51_RR (e.g., delay Δt from vibration occurrence in the content)) is controlled according to the directional component of the vibration source.
[0231] Next, a process procedure executed by the information processing device 10 according to the embodiment will be described with reference to Fig. 19. Fig. 19 is a flowchart showing the process procedure executed by the information processing device 10. Note that the process procedure shown below is repeatedly executed by the control unit 120.
[0232] 19 is repeatedly executed in a power-on state of the information processing system 1. When the process starts, it is determined whether or not there is an operation to start playing XR content. If the start operation is detected (step S101, Yes), the process proceeds to step S102, and if not, the process ends (step S101, No).
[0233] Then, first, an XR content setting process is executed (step S102). Note that the XR content setting process here includes, for example, various processes related to the initial settings of the device for playing XR content, the selection of XR content by the user, and the like.
[0234] Next, the information processing device 10 starts playing the XR content (step S103), and performs a scene detection process on the XR content being played (step S104). Next, the information processing device 10 performs a priority setting process on the result of the scene detection process (step S105), and executes a realism parameter extraction process (step S106).
[0235] Then, the information processing device 10 executes output processing of various vibration data or audio data reflecting the processing result of the realism parameter extraction processing (step S107). Then, the information processing device 10 determines whether the XR content has ended (step S108), and ends the processing if it is determined that the XR content has ended (step S108; Yes).
[0236] Furthermore, when the information processing device 10 determines in step S108 that the XR content has not ended (step S108; No), the information processing device 10 again transitions to the processing of step S104.
[0237] The procedure of the vibration localization process will be described with reference to Fig. 20. Fig. 20 is a flowchart showing the procedure of the vibration localization process. The vibration localization process corresponds to the process executed by the vibration localization processing unit 126c (control unit 120). This process is also performed as a part of the process of steps S106 and S107 in the process shown in Fig. 19. The specific detailed process contents of each step are the process contents described above.
[0238] First, as shown in FIG. 20, the vibration localization processing unit 126c identifies a directional component of a sound (vibration) (step S201).
[0239] Next, the vibration localization processing unit 126c determines a correction value CI for correcting the difference in vibration level felt by the user due to the individual difference of each vibrator, the individual difference of the user, the user's content viewing state, etc. (step S202). Note that this correction value CI is calculated and stored before the content is played (when the user sits down, etc.), and the stored correction value CI is read out in this step S202.
[0240] Next, the vibration localization processing unit 126c calculates a correction value FS for correcting the vibration level of each transducer, based on the phantom sensation technical concept, using the directional component of the sound determined in step S201 and the installation position information of each transducer (step S203).
[0241] Then, the vibration localization processing unit 126c determines a correction value CV for correcting vibration data for each transducer from (accumulating) the correction value CI for correcting the influence of individual differences of the transducers determined in step S202 and the correction value FS based on the phantom sensation technical concept calculated in step S203. Then, the vibration data determined separately based on the content is corrected (accumulated) with the correction value CV for the vibration of each transducer, and output data for each transducer is determined (step S204).
[0242] Next, the vibration localization processing unit 126c calculates the vibration timing of each vibrator (the timing of outputting a vibration signal to each vibrator) based on the technical concept of haptic apparent movement. That is, it calculates the delay time from the vibration generation timing of the vibration generating object in the content scene (the vibration signal is generated based on the audio signal in this embodiment, so it becomes the audio generation timing) (step S205).
[0243] Then, the vibration localization processing unit 126c provides the vibration data and the vibration timing data for each transducer as output data, and the output unit 126 outputs an output signal to each transducer (step S107 in FIG. 19).
[0244] As described above, the vibration localization processing unit 126c of the information processing device 10 according to the embodiment includes a plurality of vibrators, identifies the directional components of the vibration source in the input content, and controls the amplitude and delay of the output vibration of each vibrator based on the directional components.
[0245] Through such control, the information processing device 10 can provide the user with a sense of localization (sense of position) of the sound source and a sense of vibration transmission (sense of vibration movement) by adjusting the amplitude and delay of the output vibration of the multiple vibrators. In other words, the information processing device 10 provides the user with a sense of localization of the sound source based on the relationship between the amplitudes of the output vibrations of the respective vibrators, and provides the user with a sense of vibration movement based on the difference in the timing of the output vibrations of the respective vibrators. As a result, the information processing device 10 can provide the user with a sense of realism of the vibration in the content.
[0246] [Second embodiment] In the second embodiment, the calculation process of the correction value based on the phantom sensation technical concept is simplified so that it can be handled even by a relatively slow arithmetic processing device (CPU, etc.), for example.
[0247] In general, the directional components of sound (vibration) are identified in units of an appropriate number of areas (in this embodiment, eight angle areas, i.e., eight stages), and the correction values are then calculated by model processing of each angle area, for example by using a data table in which control values for each angle area are stored, or by using a calculation processing routine designed for each angle area, thereby simplifying the processing and reducing the processing load on the calculation processing device.
[0248] A method for setting the amplitude and delay time in advance will be described with reference to Fig. 22. Fig. 22 is a diagram showing a method for determining a directional component of vibration.
[0249] The vibration localization processing unit 126c (control unit 120) determines which of the angular regions r1 to r8 obtained by dividing the user's surroundings into eight regions the directional component of the vibration in the content (estimated from the sound in the content in this embodiment) belongs to (step S31). Note that data defining the angular regions is stored in advance (at the time of design, etc.) in the storage unit 130, and the angular region of the directional component of the vibration is determined using the stored data. Also, in this embodiment, as shown in step S31, angular regions r1 to r8 are set at 45 degree intervals based on an angular region r1 of 45 degrees in front.
[0250] Then, the vibration localization processing unit 126c determines the direction (d1 to d8: referred to as the representative direction) at the center of the angle region to which the vibration directional component is determined to belong as the vibration directional data used for calculating the correction value based on the phantom sensation technology concept. Note that this process can be realized by a method such as storing a data table showing the relationship between the angle regions r1 to r8 and the representative directions d1 to d8 in the storage unit 130 in advance and collating the data in the data table.
[0251] For example, if the front of the content viewing user is 0° and clockwise is expressed as a positive angle, the region r1 is a region from -30° to 30°, and its representative direction d1 is the direction of 0°. The region r2 is a region from 30° to 60°, and its direction d1 is the direction of 45°. For example, if the directional component of the vibration is 45°, the vibration localization processing unit 126c determines that the representative direction is d2, and d2 is used as the representative direction in subsequent processing.
[0252] Then, the information processing device 10 uses the representative direction data of the directional components of the vibration determined by the above method to perform signal processing equivalent to the method shown in Fig. 17. At this time, since there are only eight representative directions d, in this embodiment, correction values based on the phantom sensation technology concept are calculated in advance for each of the eight representative directions d1 to d8, and are stored in the storage unit 130 as a data table.
[0253] Furthermore, a correction value based on the technical concept of haptic apparent movement is also calculated in advance, and an integrated correction value calculated based on the calculated correction value and the correction value based on the technical concept of phantom sensation obtained as described above (for example, accumulating the correction value for the vibration level) is stored as a data table in storage unit 130. In this case, the data table of correction values based on the technical concept of phantom sensation can be omitted.
[0254] It is also possible to implement a method in which, instead of calculating an integrated correction value, a correction value based on the phantom sensation technical concept is calculated using a data table, and a correction value based on the haptic apparent movement technical concept is calculated without using a data table, thereby correcting the vibration data with each correction value.
[0255] FIG. 21 is a diagram showing an example of a data table of correction values based on the phantom sensation technical concept and the haptic apparent movement technical concept.
[0256] The data table stores correction values for vibration amplitude and vibration timing (delay) calculated in advance (at the time of design, etc.) for each transducer (51FL, 51FR, 51RL, 51RR) and for each representative direction (d1 to d8).
[0257] Then, the vibration localization processing unit 126c extracts the amplitude and delay correction values corresponding to the determined representative direction d from the data table for each transducer 51, and corrects the vibration data.
[0258] For example, when the representative direction is direction d2, the amplitude correction value for transducer 51FL is -2db, and the delay time is 0ms, the amplitude correction value for transducer 51FR is +4db, and the delay time is 0ms, the amplitude correction value for transducer 51RL is +4db, and the delay time is 50ms, and the amplitude correction value for transducer 51RR is -4db, and the delay time is 50ms. The vibration data is corrected by these correction values, and the corresponding vibration signals are output to each transducer 51.
[0259] Note that the data table shown in FIG. 21 excludes factors that vary depending on the content viewing situation, such as the user's sensitivity and the seating condition (pressure distribution on the seat surface) as correction factors, and also excludes the seat type (the vibrator itself or its arrangement, etc.) as correction factors. However, by adding these variable factors as parameters to the data table, it is possible to perform control that corresponds to these variable factors.
[0260] Thus, in the second embodiment, the information processing device 10 determines which of a plurality of predetermined angle regions the directional component of the vibration belongs to, selects a model (a data group of the corresponding direction in the data table of FIG. 21) corresponding to the angular region of the determined directional component from models (data table of FIG. 21) preset for each of the angular regions, and controls the amplitude and delay of the output vibration of each of the transducers based on the selected model.
[0261] Specifically, the model has a data table (FIG. 21 data table) in which an amplitude correction value for the amplitude and a delay correction value for the delay are stored for each of a plurality of angle regions. The information processing device 10 controls the amplitude and delay of the output vibration of each transducer based on the amplitude correction value and delay correction value for each transducer stored in the data table corresponding to the angle region of the directional component.
[0262] Therefore, in the second embodiment, control can be performed by processing using a model (data table) that has been generated in advance corresponding to the angle range to which the directional components of the vibration belong, without performing complex processing using the directional components of the vibration, thereby reducing the processing load, such as by reducing the amount of calculations.
[0263] In the above embodiment, the content is XR content, but the present invention is not limited to this. That is, the content may be 2D video and audio, or only video, or only audio.
[0264] Further advantages and modifications may readily occur to those skilled in the art. Thus, the invention in its broader aspects is not limited to the specific details and representative embodiments shown and described above. Accordingly, various modifications may be made without departing from the spirit or scope of the general inventive concept as defined by the appended claims and equivalents thereof. [Explanation of symbols]
[0265] 1. Information Processing Systems 3 Display device 4 Speakers 5. Vibration Devices 10. Information processing device 31 images 121 Content Generation Department 122 Rendering Processor 123 Scene detection section 123a Scene determination section 123b Condition setting section 124 Priority Setting Section 124a Timing detection unit 124b Rule setting section 125 Parameter Extraction Unit 125a Vibration parameter extraction unit 125b Speech enhancement parameter extraction unit 125c Learning Department 126 Output section 126a Voice enhancement processing unit 126b Voice vibration conversion processing unit 126c Vibration localization processing unit 131 XR Content DB 132 Scene Information DB 133 Priority Information DB 134 Parameter Information DB 311, 312 Objects
Claims
1. An information processing device that controls a plurality of vibrators provided in a vibration device that applies vibrations to a user according to content, A controller is provided. The controller: Identifying a directional component of a vibration source of the content relative to a position of the user in a content space of the content; determining which of a plurality of predetermined angular regions the directional component belongs to; selecting a model corresponding to the angular region of the determined directional component from among models preset for each angular region; Controlling the amplitude and delay of output oscillation of each of said oscillators based on said selected model. Information processing device.
2. the model has a data table in which an amplitude correction value for an amplitude and a delay correction value for a delay are stored for each of the plurality of angular regions; The controller: The amplitude and delay of the output vibration of each of the transducers are controlled based on the amplitude correction value and the delay correction value stored in the data table corresponding to the angle region of the directional component. The information processing device according to claim 1 .
3. The controller: The amplitude and delay of the output vibration of each of the vibrators are controlled based on the user's vibration sensitivity characteristics to the output vibration of each of the vibrators.
3. The information processing device according to claim 1 or 2.
4. The controller: Calculating the vibration sensitivity characteristic of the user based on the characteristics of the vibrator and the state of the user. The information processing device according to claim 3 .
5. The controller: The vibrator is vibrated for calibration while the user is in a posture in which the user is viewing content, and the vibration sensitivity characteristics of the user are measured. The information processing device according to claim 4.
6. The controller: controlling an amplitude of output vibration of each of the plurality of transducers so that a relationship between a first resultant vibration position determined based on a vibration level of each transducer in a first transducer group included in the plurality of transducers and a second resultant vibration position determined based on a vibration level of each transducer in a second transducer group included in the plurality of transducers coincides with a directional component of the identified vibration source; A delay in output vibration of each of the plurality of transducers is controlled according to a directional component of the vibration source. The information processing device according to claim 1 .
7. An information processing device that reproduces XR content; a vibration device that includes a plurality of vibrators and applies vibrations to a user in response to a vibration signal output from the information processing device; Equipped with The controller of the information processing device Detecting a scene in which a sound is generated from an object from the XR content; extracting vibration parameters corresponding to the scene, the vibration parameters being used to control the vibration device; a signal obtained by processing a sound signal generated from the object is emphasized using the vibration parameters; Identifying a directional component of the object in the XR content with respect to a position of the user in a content space of the XR content; Identifying an angle region to which the specified directional component belongs from a plurality of predetermined angle regions; selecting a model process corresponding to the specified angular region from among model processes preset for each angular region; A signal obtained by processing the amplitude and delay of the signal obtained by the enhancement processing using the selected model processing is output to the vibration device as the vibration signal. Information processing system.
8. further comprising an audio output device that generates audio in response to an audio signal output from the information processing device; The controller of the information processing device Extracting audio parameters related to audio processing corresponding to the scene; The audio signal that has been enhanced using the audio parameters is output to the audio output device. The information processing system according to claim 7.
9. An information processing method for controlling a plurality of vibrators provided in a vibration device that applies vibrations to a user according to content, comprising: Identifying a directional component of a vibration source of the content relative to a position of the user in a content space of the content; determining which of a plurality of predetermined angular regions the directional component belongs to; selecting a model corresponding to the angular region of the determined directional component from among models preset for each angular region; Controlling the amplitude and delay of output oscillation of each of said oscillators based on said selected model. An information processing method in which processing is performed by a computer.
10. the model has a data table in which an amplitude correction value for an amplitude and a delay correction value for a delay are stored for each of the plurality of angular regions; The computer, The amplitude and delay of the output vibration of each of the transducers are controlled based on the amplitude correction value and the delay correction value stored in the data table corresponding to the angle region of the directional component. The information processing method according to claim 9.
11. The computer, The amplitude and delay of the output vibration of each of the vibrators are controlled based on the user's vibration sensitivity characteristics to the output vibration of each of the vibrators.
11. The information processing method according to claim 9 or 10.
12. The computer, Calculating the vibration sensitivity characteristic of the user based on the characteristics of the vibrator and the state of the user. The information processing method according to claim 11.
13. The computer, The vibrator is vibrated for calibration while the user is in a posture in which the user is viewing content, and the vibration sensitivity characteristics of the user are measured. The information processing method according to claim 12.
14. The computer, controlling an amplitude of output vibration of each of the plurality of transducers so that a relationship between a first resultant vibration position determined based on a vibration level of each transducer in a first transducer group included in the plurality of transducers and a second resultant vibration position determined based on a vibration level of each transducer in a second transducer group included in the plurality of transducers coincides with a directional component of the identified vibration source; A delay in output vibration of each of the plurality of transducers is controlled according to a directional component of the vibration source. The information processing method according to claim 9.
15. A program for controlling a plurality of vibrators provided in a vibration device that applies vibrations to a user according to content, comprising: Identifying a directional component of a vibration source of the content relative to a position of the user in a content space of the content; determining which of a plurality of predetermined angular regions the directional component belongs to; selecting a model corresponding to the angular region of the determined directional component from among models preset for each angular region; Controlling the amplitude and delay of output oscillation of each of said oscillators based on said selected model. A program that causes a computer to carry out processing.
16. the model has a data table in which an amplitude correction value for an amplitude and a delay correction value for a delay are stored for each of the plurality of angular regions; The amplitude and delay of the output vibration of each of the transducers are controlled based on the amplitude correction value and the delay correction value stored in the data table corresponding to the angle region of the directional component. The program according to claim 15, which causes the computer to execute a process.
17. The amplitude and delay of the output vibration of each of the vibrators are controlled based on a user's vibration sensitivity characteristic to the output vibration of each of the vibrators.
17. The program according to claim 15 or 16, which causes the computer to execute a process.
18. The vibration sensitivity characteristic of the user is calculated based on the characteristics of the vibrator and the state of the user. The program according to claim 17, which causes the computer to execute a process.
19. The vibrator is vibrated for calibration purposes while the user is in a posture when viewing content, thereby measuring the vibration sensitivity characteristics of the user. The program according to claim 18, which causes the computer to execute a process.
20. Controlling the amplitude of output vibration of each of the plurality of oscillators so that a relationship between a first composite vibration position determined based on the vibration level of each oscillator in a first oscillator group included in the plurality of oscillators and a second composite vibration position determined based on the vibration level of each oscillator in a second oscillator group included in the plurality of oscillators coincides with a directional component of the identified vibration source; A delay in output vibration of each of the plurality of transducers is controlled according to a directional component of the vibration source. The program according to claim 15, which causes the computer to execute a process.
Citation Information
Patent Citations
Stereoscopic display device and enjoying chair
JP1996098957A
Sound reproducing device
JP2000013900A
Direction contact feedback for tactile feedback interface device
JP2003199974A
Bodily feeling experiencing video / sound system
JP2004081357A
Rocking device and method, and audiovisual system
JP2007068881A