Content reproduction device, vibration control signal generation device, server device, vibration control signal generation method, content reproduction system, and design assistance device
Patent Information
- Application Number
- JP2024551047
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-05
AI Technical Summary
Existing content playback devices require significant manual effort and man-hours to design and adjust vibration output mechanisms, as different configurations necessitate tailored vibration control signals, leading to inefficiencies in generating optimal vibration experiences for users.
A content playback system that includes a controller capable of detecting vibration output devices and adjusting vibration control parameters automatically, based on scene recognition and user feedback, to generate tailored vibration signals for improved realism.
This approach enables efficient selection and adjustment of vibration output devices and parameters, enhancing the sense of presence by providing customized vibrations for various content scenarios, reducing manual labor and improving user experience.
Abstract
Description
Content playback device, vibration control signal generating device, server device, vibration control signal generating method, content playback system, and design support device
[0001] The present invention relates to a content playback device, a vibration control signal generating device, a server device, a vibration control signal generating method, a content playback system, and a design support device.
[0002] Conventionally, there have been proposed techniques for improving the sense of realism of content by applying vibrations to a user in accordance with the content being viewed by the user. For example, there is known a technique for improving the sense of realism by vibrating the air around the user, allowing the user to experience the air vibrations with their entire body (see, for example, Patent Document 1).
[0003] Japanese Patent Application Publication No. 11-46391
[0004] In content playback devices that apply vibrations to a user to enhance the sense of realism of content, the components that apply (output) the vibrations to the user (hereinafter referred to as vibration output mechanisms) are not necessarily the same. For example, different models of chair-type vibration output mechanisms differ in the type of vibration output device used, the material and shape of the chair on which the user sits, the mounting position of the vibration output device, and so on. Therefore, the optimal vibration generated varies depending on the type of vibration output mechanism, and the vibration output value for vibration control used to generate a vibration signal also differs. For this reason, it has been necessary to design and adjust a dedicated content playback device depending on the vibration output mechanism used. Alternatively, it has been necessary to design and adjust a dedicated vibration output mechanism compatible with a vibration control signal generating device that generates a vibration control signal. Conventionally, these adjustments have been performed manually, resulting in the problem of requiring a huge amount of work.
[0005] In view of the above-mentioned problems, an object of the present invention is to provide a technique that enables efficient design and adjustment of a vibration output mechanism.
[0006] An exemplary embodiment of the present invention is a content playback device that provides a user with vibrations corresponding to content being played back, the content playback device comprising: a vibration output mechanism that generates vibrations; and a controller. The controller detects a vibration output device of the vibration output mechanism and controls the vibration generated by the vibration output mechanism in accordance with the detected vibration output device.
[0007] According to the present invention, it is possible to efficiently select a vibration output device and adjust vibration control parameter values related to improving the sense of realism of content.
[0008] 6. FIG. 6 is a diagram showing an example of a content reproduction system according to an embodiment; FIG. 6 is a diagram showing an overview of vibration control signal generation processing performed by the content reproduction device of FIG. 1; FIG. 6 is a structural diagram showing an example of the content reproduction device of FIG. 1; FIG. 6 is a diagram showing an example of a scene information DB; FIG. 6 is a diagram showing an example of a parameter information DB; FIG. 6 is a diagram showing an example of an arrangement of vibration output devices in a vibrating sheet of a content reproduction system;
[0009] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the drawings. However, the present invention is not limited to the contents of the embodiments shown below.
[0010] 1. Overview of vibration control signal generation method> Fig. 1 is an explanatory diagram showing an example of a content playback system PS according to an embodiment. As shown in Fig. 1, the content playback system PS includes a content playback device 10, a display device (video display) P1, a speaker (audio output device) P2, and a vibrating sheet (vibration output mechanism) P3.
[0011] The content playback device 10 is a device that generates control signals for actuators for video playback, audio playback, and vibration playback according to the content to be played. The display device P1 is a device that provides user U1 with video according to the content to be played (video based on the video (control) signal of the content playback device 10). The speaker P2 is a device that provides user U1 with audio according to the content to be played (audio based on the audio (control) signal of the content playback device 10).
[0012] The display device P1 is, for example, a head-mounted display. The display device P1 outputs an image corresponding to the content being played, allowing the user U1 to enjoy, for example, an XR (Cross Reality) experience. The display device P1 includes a device, such as a camera, a microphone, a motion sensor, etc., that detects changes in the internal and external circumstances of the user U1 using a sensor unit.
[0013] The content provided to the user U1 is not limited to XR content, but may be content displayed on a normal display, such as a movie, concert video, or game. In this case, the display device P1 may be a display such as a television set installed on the floor or a desk, or hung on a wall.
[0014] The speaker P2 is, for example, a headphone type speaker that is worn on the ears of the user U1. The speaker P2 outputs sound corresponding to the content being played back, thereby providing the sound of the content to the user U1. The speaker P2 is not limited to a headphone type speaker, and may be, for example, a box-shaped speaker that is installed on the floor or a desk, or hung on a wall, a so-called box speaker.
[0015] The vibrating seat P3 is, for example, a chair-type vibration output mechanism, and includes a seat (seat (chair) body) 20 on which the user U1 sits, and a plurality of vibration output devices 30. The vibration output devices 30 are installed inside or outside the seat 20. The vibration output devices 30 are configured by electric vibration converters including, for example, an electric magnetic circuit, a piezoelectric element, and an electric cylinder. The vibration output devices 30 generate vibrations according to the content being played (vibrations based on vibration (control) signals from the content playback device 10), and impart (output) the vibrations to the user U1.
[0016] Fig. 2 is an explanatory diagram showing an outline of the vibration control signal generation process performed by the content playback device 10 of Fig. 1. The content playback device 10 first detects scenes that satisfy predetermined conditions from the video data and audio data related to the content (step S1).
[0017] The predetermined conditions of a scene are conditions under which a scene is determined to be one in which vibration should be generated (vibration should be applied to the user). Specifically, the predetermined conditions of a scene are conditions related to the characteristics of an object in the content (weight, moving speed, etc.), the distance between the avatar corresponding to the user and the object in the content space (hereinafter referred to as the object distance), etc. For example, if the conditions "the square of the object weight (kg) / object distance (m) is 10 or more" and "the object is moving" are satisfied in the content space, the content playback device 10 determines that a scene is one in which vibration should be generated.
[0018] Next, the content playback device 10 sets priorities for events to be subjected to vibration for the scenes detected by scene detection (step S2). That is, when a scene is determined to be one in which vibration should be generated, multiple events may satisfy the vibration generation conditions. For example, if an elephant approaches while driving on a rough road in the content space, vibrations corresponding to the driving on the rough road and vibrations corresponding to the approaching elephant will be generated. However, if vibrations corresponding to both events are generated, the cause of the vibrations (which object the vibrations are directed at) will not be perceived based on human sensitivity to vibrations. For this reason, the content playback device 10 assigns a higher priority to the event in which the user should feel the vibration, and controls the generation of vibrations corresponding to that event.
[0019] Next, the content playback device 10 extracts vibration control parameter values corresponding to the highest-priority scene (the event for which vibration is to be generated) (step S3). The vibration control parameters are data used when generating vibration, and include parameters such as low-pass filter characteristics (cutoff frequency, etc.), delay characteristics, and amplification characteristics. These parameter values are associated with the scene and stored in a data table, etc. Note that, in the case of processing to generate vibrations corresponding to multiple scenes in descending order of priority, a similar vibration generation process is performed for each of the multiple highest-priority scenes.
[0020] Next, the content playback device 10 generates a vibration control signal for the vibration output device 30 of the vibrating sheet P3 based on the extracted vibration control parameter value (step S4), and outputs the vibration control signal to the vibration output device 30 of the vibrating sheet P3. Specifically, among the information contained in much of the content, there is a relatively high correlation between sound and vibration. For this reason, the content playback device 10 processes the sound data contained in the content based on the content data (object distance, object weight, object type, etc.) and vibration control parameter values corresponding to the scene, to generate a vibration control signal. Then, the generated vibration control signal is output to the vibration output device 30 of the vibrating sheet P3. In this way, the content playback device 10 can impart vibrations to the user according to the content being played back.
[0021] 2. Content Reproducing Device Fig. 3 is a configuration diagram showing an example of the content reproducing device 10 of Fig. 1. Fig. 3 shows components necessary for explaining the features of this embodiment, and omits the description of general components.
[0022] As shown in FIG. 3 , the content playback device 10 includes a storage unit 12 and a controller 13. In this description, the content input to the content playback device 10 is described as "XR content." That is, the "user" refers to the operator himself / herself (a virtual character or avatar for the content viewing user) in the XR space. The content viewing user (the operator himself / herself) hears the sounds that the virtual character (avatar) hears (surrounding sounds (such as the vocalizations of other characters), the vocalizations of the virtual character). For this reason, the user in the XR content has a function (microphone function) to convert the user's own voice and the sounds of the surroundings into audio (electrical) signals.
[0023] The storage unit 12 includes a volatile memory and a non-volatile memory. The volatile memory is, for example, a random access memory (RAM). The non-volatile memory is, for example, a read-only memory (ROM), a flash memory, or a hard disk drive. The non-volatile memory stores programs and data that can be read by the controller 13. At least some of the programs and data stored in the non-volatile memory may be obtained from another computer device (server device) connected via a wired or wireless connection, or from a portable recording medium.
[0024] The storage unit 12 is provided with a plurality of databases (hereinafter sometimes referred to as "DBs (Databases)") for various processes. The databases provided include a content DB 121, a scene information DB 122, a priority information DB 123, a parameter information DB 124, and a scene recognition DB 125.
[0025] The content DB 121 is a database that stores data of a group of content to be played back by the content playback device 10. Based on this data of the content group, the video, audio, vibration, etc. of each content is played back. Note that this content data may be obtained from an external server (the external server is treated as the content DB 121), or the content DB 121 and an external server may be used together.
[0026] The scene information DB 122 is a database that stores various information related to a scene for which vibration is to be generated. FIG. 4 is a diagram showing an example of the scene information DB 122.
[0027] As shown in Figure 4, the scene information DB 122 contains data for the items "detected scene," "condition category," "object," "condition parameter," "threshold," and "condition expression," and each piece of information is stored in association with the "detected scene" information.
[0028] The "detected scene" item in the scene information DB 122 is the name of a scene, which serves as identification information for the scene. The "detected scene" also serves as identification information for identifying a data record in the scene information DB 122. In other words, a data record in the scene information DB 122 is generated for each "detected scene" data, and data for the items "condition category," "object," "condition parameter," "threshold," and "condition formula" corresponding to the data record is stored. While a scene identification code such as a numerical value is usually used for the "detected scene," in this embodiment, a distinctive name is used to make the explanation easier to understand.
[0029] The "condition category" item in the scene information DB 122 indicates a category of scene detection information, such as the type of information used to detect a scene. In the example shown in Fig. 4, the data in the "condition category" is broadly categorized into categories such as the positional relationship between the user and an object in the XR space, the user's actions, spatial information about the user's presence, time information about the user's presence, and whether a sound is being emitted from an object.
[0030] The "object" item in the scene information DB 122 indicates the type of object used for scene detection. In the example shown in FIG. 4 , "object" corresponds to information such as object 1, object 2, user, space 1, space 1 + object 3, content 1, object 4, object 5, and object 6. Here, object 1, object 2, object 3, object 4, object 5, and object 6 each represent a different object in the XR space. Furthermore, space 1 represents, for example, the space in the XR space where the user exists, and content 1 represents the content itself.
[0031] The item "condition parameter" in the scene information DB 122 indicates a condition related to a parameter, such as which parameter is to be used for an object (data of "object") when detecting a scene. In the example shown in Fig. 4, the "condition parameter" corresponds to parameter type information such as distance, angle, speed, acceleration, rotation speed, in space, presence of object, quantity, start time to end time, and sound pattern.
[0032] The "threshold" item in the scene information DB 122 indicates a threshold value corresponding to a condition parameter for determining a detected scene. The "conditional formula" item in the scene information DB 122 indicates a conditional formula for detecting a detected scene, and for example, the relationship between the conditional parameter and the threshold value is defined and stored as a conditional formula.
[0033] For ease of explanation, in Figure 4, each item value is represented using symbols such as "W", "4", and "w", such as "Scene W", "Object 4", and "Pattern w", but in reality, each item value will be stored as data in a form that allows the specific meaning to be understood.
[0034] Specifically, for example, the detected scenes "Scene W," "Scene X," "Scene Y," and "Scene Z" are actually data such as an "elephant walking scene," a "horse walking scene," a "car driving scene," and a "car making a sharp turn scene," respectively. In this case, the targets "Object 4," "Object 5," and "Object 6" are actually data such as an "elephant," a "horse," and a "car," respectively. Furthermore, the conditional expressions "Pattern w," "Pattern x," "Pattern y," and "Pattern z" are actually data such as a "horse walking sound pattern," a "elephant walking sound pattern," a "car driving sound pattern," and a "tire squealing sound pattern," respectively.
[0035] A voice pattern is represented by, for example, a feature vector having voice features as elements. If the similarity (e.g., cosine similarity or Euclidean distance) between the feature vectors corresponding to two voice patterns is equal to or greater than a threshold, the two voice patterns can be determined to be similar. For example, the conditional expression "voice pattern is similar to pattern w" means that the similarity between the feature vector calculated from the voice occurring in the scene and the feature vector of the voice corresponding to pattern w is equal to or greater than a threshold.
[0036] Furthermore, the content playback device 10 may detect a scene by combining the condition categories or condition parameters shown in Fig. 4. For example, a scene α may be detected in which the condition category is the positional relationship between the user and an object and the condition parameters are position and angle, i.e., the scene α satisfies the conditions of scene A and scene B.
[0037] The priority information DB 123 is a database that stores various information related to the priority of vibration generation. The content playback device 10 sets a priority of vibration generation for each scene in which vibration should be generated based on a predetermined rule. The rules related to the priority of vibration generation are stored in the priority information DB 123. Although a detailed description will be omitted here, the priority rules may include, for example, "give priority to the scene detected earlier (or later)," "give priority to a scene with a shorter duration," "give priority to a scene with a larger low-frequency amplitude," and "give priority to a scene that ends first."
[0038] The parameter information DB 124 is a database that stores information about vibration control parameters for each scene. Fig. 5 is a diagram showing an example of the parameter information DB 124. As shown in Fig. 5, the parameter information DB 124 includes items such as "scene type" and "vibration control parameter," and stores each piece of information in association with information about the "scene type."
[0039] The "Scene Type" field in the parameter information DB 124 indicates the type of scene. The "Detected Scene" data shown in FIG. 4 is associated with the "Scene Type" data in a predetermined manner (for example, a data table showing the correspondence). In other words, the "Scene Type" field in the parameter information DB 124 and the "Detected Scene" field in the scene information DB 122 are associated with each other in a predetermined manner, and as a result, the data records in the scene information DB 122 and the parameter information DB 124 are associated (linked).
[0040] The item "vibration control parameters" in the parameter information DB 124 indicates vibration control parameters to be set for the corresponding scene, and data (values) of each parameter are stored individually for each vibration output device 30 of the vibrating sheet P3. As the "vibration control parameters," data on items such as "LPF (Low Pass Filter, low frequency characteristics)," "delay (delay characteristics)," and "amplification (amplification factor)" are stored. Note that while FIG. 5 shows "vibration control parameters" for two types of vibration output devices, "vibration control parameters" are stored for the vibration output devices to be individually controlled.
[0041] The data example shown in Figure 5 is parameter values for vibration generation processing based on content audio. "LPF" indicates the cutoff frequency of a low-pass filter that extracts low-frequency components from audio. "Delay" indicates the time by which vibration is delayed relative to audio. "Amplification" indicates the amplification factor, i.e., the degree to which the original vibration generated from audio is amplified or attenuated to control the vibration.
[0042] Returning to Fig. 3, the explanation will be continued. The controller 13 realizes various functions of the content playback device 10 and includes a processor that performs arithmetic processing and the like. The processor includes, for example, a CPU (Central Processing Unit). The controller 13 may be configured with one processor or multiple processors. When the controller 13 is configured with multiple processors, these processors are connected to each other so that they can communicate with each other and cooperate to execute processing.
[0043] The controller 13 includes, as its functions, a scene detection unit 131, a priority setting unit 132, a parameter extraction unit 133, and an output unit 134. In this embodiment, the functions of the controller 13 are realized by a processor executing arithmetic processing in accordance with a program stored in the storage unit 12.
[0044] The scene detection unit 131 includes a scene determination unit 131a that determines whether a scene in the content being played is a scene in which vibration control (generation) should be performed, and a parameter setting unit 131b that sets parameter values to be used in processing for vibration generation processing.
[0045] The scene determination unit 131a determines whether a scene in the currently played content satisfies predetermined conditions. The scene determination unit 131a determines whether a scene in the currently played content should be subjected to vibration control (i.e., should generate vibrations) (detects a scene in which vibration control should be performed) by using, for example, video data and audio data related to the content and a conditional expression stored in the scene information DB 122. Specifically, the scene determination unit 131a determines whether a scene in the currently played content should be subjected to vibration control by using a conditional expression stored in the scene information DB 122, based on, for example, coordinate information of an object (object subject to vibration generation) in the XR space and information related to the object type. Furthermore, if a scene in the currently played content is a scene in which vibration control should be performed, the scene determination unit 131a determines which of the "detected scenes" stored in the scene information DB 122 the scene in question is, based on this information.
[0046] As a specific example, the scene determination unit 131a detects a scene in the currently played content in which sound is generated from an object. For the sound generation scene, scene W, scene X, scene Y, and scene Z, which are in the condition category "sound generated from an object" shown in FIG. 4, are selected as candidate scenes. The scene determination unit 131a calculates the similarity between a feature vector obtained from the audio signal of the content and a predetermined feature vector of sound in the candidate scene (sound pattern of the condition parameter), and determines whether the candidate scene satisfies the sound pattern condition based on the determination result of whether the similarity is equal to or greater than a predetermined similarity threshold. Furthermore, the scene determination unit 131a determines whether the candidate scene satisfies the object distance condition based on the determination result of whether the object distance in the currently played content is equal to or less than a predetermined threshold. Then, the scene determination unit 131a determines a candidate scene that satisfies the sound pattern condition and the object distance condition (satisfies the conditional expression) as a detection scene for generating vibration.
[0047] If the scene determination unit 131a determines that the detected scene does not correspond to any of the detected scenes, it is assumed that there is no corresponding detected scene and vibration is not generated (the vibration control parameter is set to a no-vibration value).
[0048] The parameter setting unit 131b sets (initializes or changes) the value of the "vibration control parameter" in the parameter information DB 124. The main methods for setting the parameter value include a setting method based on input information by an XR content developer or a user, and an automatic setting method based on the content type or the like.
[0049] Specifically, in the method for setting parameter values based on information input by the user, the user selects a scene for which parameter values are to be set (adjusted) and a parameter type for which they are to be set (adjusted), and sets the parameters to be set in that scene by operating the up / down operation button, etc. When setting, it is preferable to display a test image of the scene for which the parameters are to be set, and generate vibrations based on the parameters being set, so that the user can set the parameters while feeling the vibrations.
[0050] In addition, in a method for automatically setting parameter values based on the scene type of content, the scene type of the content to be played is first detected. The content type is detected based on scene type information added to the content information, or is inferred by analyzing part of the content video and audio. Then, the automatic setting method sets each parameter value according to the detected content type.
[0051] The parameter values can be obtained from a server (which collects parameter information from each device and stores appropriate parameter values according to the scene type of the content by performing statistical processing, etc.) based on the scene type information of the content (by transmitting a parameter value request signal including the scene type information), etc. This allows the parameter information DB 124 to be configured more appropriately.
[0052] Furthermore, it is efficient for the parameter setting unit 131b to set vibration control parameter values for scenes in the content in which the amplitude of low-frequency sound generated from an object exceeds a predetermined threshold as scenes in which vibration is generated. Objects that generate large vibrations are highly correlated with objects that generate large low-frequency sound, and the magnitude of the vibrations is also correlated with the magnitude of the low-frequency sound. Therefore, it is estimated that scenes in which the amplitude of low-frequency sound exceeds a predetermined threshold will also have large amplitude vibrations that should be generated to improve the sense of realism, and it is efficient to set vibration control parameter values for scenes in which vibration is generated.
[0053] Such scenes can be set by the user or content developer, or obtained from a server (which collects scene information and parameter information for various content from each device and stores appropriate scene information and parameter information by performing statistical processing, etc.).
[0054] The audio amplitude threshold may be determined based on the type (details) of content. Specifically, a data table of content types (details) and intensity thresholds is created in advance, and when selecting a scene for which conditions are to be set, the intensity threshold corresponding to the target content is searched for in the data table, and the scene for which conditions are to be set is selected using the searched intensity threshold.
[0055] For example, types of content include music videos that mainly allow users to listen to music, and animal documentaries that explain the biology of animals. If a music video includes a scene of an elephant walking, it is often better not to generate excessive vibrations so as not to interfere with the music. On the other hand, if an elephant is walking in an animal documentary, it is often better to generate vibrations to create a sense of realism.
[0056] Therefore, the parameter setting unit 131b sets the threshold for the music video lower than the threshold for the animal documentary. As a result, a scene in which an elephant walks in a music video is less likely to be set as a scene in which vibrations should be generated than a scene in which an elephant walks in an animal documentary, and the generation of unnecessary vibrations in the scene in which an elephant walks in a music video is suppressed. This makes it possible to generate vibrations that are appropriate for the content.
[0057] In addition, each parameter value in the scene information DB 122 and the parameter information DB 124 may be updated by calculating (correcting) a new parameter value (for example, the adjustment value itself, or a value with an offset, etc. added) based on various vibration adjustments (vibration level adjustment, delay adjustment, etc.) actually performed by the user while watching content.
[0058] Returning to Fig. 3, the explanation will be continued. The priority setting unit 132 sets priorities for the scenes detected by the scene detection unit 131. When the scene detection unit 131 simultaneously detects multiple types of scenes, the priority setting unit 132 refers to the priority information DB 123, for example, and selects which scene should be given priority in processing. Note that when the scene detection unit 131 detects only one scene, that scene has the highest priority.
[0059] The parameter extraction unit 133 extracts vibration control parameter values for scenes for which priorities have been set by the priority setting unit 132. More specifically, the parameter extraction unit 133 refers to the parameter information DB 124 and extracts, from the parameter information DB 124, vibration control parameter values corresponding to the "detected scene" that has been set to the highest priority by the priority setting unit 132. In this case, the parameter extraction unit 133 extracts vibration control parameter values individually corresponding to each of the multiple vibration output devices 30, so that each vibration output device 30 can be controlled with its own dedicated vibration control parameter value. This makes it possible to further improve the sense of realism compared to when each vibration output device 30 is controlled with a uniform vibration control parameter value.
[0060] The content playback device 10 can estimate, from among the candidate objects that are vibration-generating objects, object candidates that will have a large impact on the user when they generate vibrations, and select the candidate object as the vibration-generating object, based on the priorities set by the priority setting unit 132. In this case, it is preferable to change the threshold for selecting an object that generates vibration based on the type of content. That is, depending on the content, it may be preferable to reduce or emphasize the reproduction of vibrations caused by objects that appear in the content, and it is therefore preferable to adjust the determination content (determination level) of the object that generates vibration.
[0061] In other words, the vibration generation principle is as follows: The object that will generate vibration in the content (in each scene) is determined based on the content of the content. Then, a vibration signal (vibration data) is generated based on the audio signal corresponding to the determined object. In this case, the audio signal corresponding to the object is audio data of the object included in the content, or audio data of the object generated from audio data in the scene (for example, extracted by filtering the low-frequency range). In addition, the vibration signal (vibration data) is generated by extracting the low-frequency components of the audio signal of the object and amplifying it appropriately.
[0062] In addition, as a method for determining the object that generates vibration, the low-frequency characteristics (e.g., volume level) of the vocalizations of the sound-generating objects of the content are estimated to determine the object. In this case, the low-frequency characteristics of the vocalizations of the sound-generating objects are estimated based on, for example, a reference vibration intensity based on the type of object (object) in the virtual space and the distance between a reference position (e.g., a user position in the virtual space) and the object (object). In determining the object, the sound-generating object with the highest low-frequency volume level of the vocalizations is determined to be the object that generates vibration.
[0063] At this time, the parameter extraction unit 133 extracts a dedicated vibration control parameter value for each vibration output device 30. This makes it possible to further improve the sense of realism compared to when each vibration output device 30 is controlled by a uniform vibration control parameter value.
[0064] The parameter extraction unit 133 also includes a learning unit 133a. The learning unit 133a learns the relationship between the scenes stored in the parameter information DB 124 and the vibration control parameter values.
[0065] The learning unit 133a updates the vibration control parameter values stored in the parameter information DB 124 by performing machine learning using, for example, scenes stored in the parameter information DB 124, the corresponding vibration control parameter values, and the reaction of the user (the user watching the content to which vibration is applied) to the vibration control of the vibration output device 30 based on the parameter values as learning data.
[0066] In this case, the learning unit 133a may use, for example, user evaluations of vibration control parameter values (vibrations given to the user) (vibration adjustment operations by the user after vibration control, survey results by the user, etc.) as learning data.
[0067] In this way, the learning unit 133a learns (updates) vibration control parameter values according to the scene from the perspective of what vibration control parameter values should be set for each scene to obtain a high user rating, i.e., a high sense of realism.
[0068] Furthermore, the learning unit 133a determines, from the learning results, what vibration control parameter values should be set when a new scene is played. Specifically, for example, when a viewer user adjusts vibrations during playback of a fireworks scene that is not registered as a vibration-generating scene, the learning unit 133a calculates vibration control parameter values using the new scene and the adjustment details as learning data, and stores data such as vibration control parameter values based on the learning results in the parameter information DB 124, etc. Note that it is also possible to learn (generate) vibration control parameter values for a new scene using vibration control parameter values for a similar scene. For example, when a new fireworks scene is played, vibration control can be performed using vibration control parameter values for a similar situation, such as an explosion scene, and the learning results, i.e., the user's reaction, can be used to learn (generate) vibration control parameter values for the fireworks scene.
[0069] The output unit 134 generates a vibration control signal for each vibration output device 30 using the vibration control parameter values extracted by the parameter extraction unit 133, and outputs the signal to each vibration output device 30. Specifically, the output unit 134 performs band limiting processing suitable for vibration using an LPF on audio data in a vibration generation scene of the content being played, thereby converting the audio data into original vibration data. Furthermore, the output unit 134 performs vibration adjustment processing on the original vibration data based on the vibration control parameter values extracted by the parameter extraction unit 133, and generates a vibration control signal.
[0070] Specifically, the output unit 134 performs vibration adjustment processing on the original vibration data, such as adding frequency characteristics such as low-frequency emphasis, delaying, and amplifying the original vibration data in accordance with the vibration control parameter value. In this way, the output unit 134 outputs vibration control signals (vibration control data) obtained by adjusting signals suitable for vibration obtained by processing audio signals generated from objects in the XR space of the content in accordance with the vibration control parameter value to each of the multiple vibration output devices 30. At this time, the output unit 134 outputs the vibration control data that has been individually adjusted for each vibration output device 30 to the corresponding vibration output device 30.
[0071] 5, the vibration control parameter values are set for each scene, but it is also effective to make corrections according to the detailed situation of the scene (which can also be called the detailed scene type). For example, in a scene in which a vibration target (e.g., an elephant) exists in an XR space, the vibration characteristics may be adjusted by increasing or decreasing the values of the vibration control parameters "LPF," "delay," and "amplification" according to the distance between the user and the target (detailed scene by distance).
[0072] As a result, the content playback device 10 generates an original vibration signal based on the audio data of a scene in which vibration should be generated during content playback, and further processes the original vibration signal according to the scene type to generate a vibration control signal. As a result, even for general content that does not have dedicated data for vibration control in the content data, it is possible to generate vibrations that are carefully adapted to each scene in the content.
[0073] 3. Setting of Vibration Control Parameters Next, setting of vibration control parameter values for the plurality of vibration output devices 30 that constitute the vibration output mechanism (vibrating sheet P3) will be described.
[0074] In order to improve the sense of realism of the content, the content playback device 10 breaks down the vibrations imparted to the user (person) into aerial vibrations and structural vibrations. Aerial vibrations refer to vibrations that are transmitted from a vibration source to the body via the air. Structural vibrations refer to vibrations that are transmitted to the body through direct contact with the vibration source or through contact via structural components or the ground.
[0075] When playing content such as XR content, in scenes involving strong vibrations accompanied by the user's body movements, structural vibration is appropriate for the vibration to be applied to the user. In scenes involving vibrations other than structural vibrations, aerial vibration is appropriate. Scenes involving strong vibrations accompanied by the user's body movements include, for example, scenes in which a heavy object in the content shakes, the entire video (camera) shakes, the user's movement in the XR content exceeds a predetermined threshold, or an object including the user in the content or the camera comes into contact with another object. Furthermore, with regard to the user in the XR content, the user's grounding state (standing on the ground, flying in the air, etc.) affects the strength of the structural vibration and aerial vibration.
[0076] The content playback device 10 separates the vibrations to be applied to the user for each content scene into aerial vibrations and structural vibrations, and assigns each of the multiple vibration output devices 30 to handle the aerial vibrations and structural vibrations. That is, the content playback device 10 generates a vibration control signal for each of the multiple vibration output devices 30 using a vibration control parameter value corresponding to the aerial vibrations, and drives the vibration output device 30, or generates a vibration control signal using a vibration control parameter value corresponding to the structural vibrations, and drives the vibration output device 30.
[0077] This allows the content viewing user to be given an appropriate combination of aerial vibrations, which are transmitted to the body through the air, and structural vibrations, which are transmitted directly to the body through contact with a vibrating object, etc. Furthermore, by appropriately setting the vibration forms, such as the positions at which the aerial vibrations and structural vibrations are generated (applied), the vibration frequency, and the vibration waveform, it becomes possible to apply vibrations to the user in various forms. This can improve the sense of realism of the content.
[0078] 6 is an explanatory diagram showing an example of the arrangement of vibration output devices 30 in a vibrating seat P3 of a content reproduction system PS. In the content reproduction system PS, vibration output devices 31, 32, each consisting of an exciter (diaphragm), are installed inside the seat surface 21 and back surface 22 of the seat 20. In addition, a vibration output device 33, each consisting of a six-axis electric cylinder, is installed in the lower portion 23 of the seat 20.
[0079] The vibration output devices 31 on the seat surface 21 of the seat 20 are arranged at five locations: the four corners (vibration output devices 31s) and the center (vibration output device 31c) of the seat surface 21. The vibration output device 31c at the center of the seat surface 21 has a larger diaphragm than the other four surrounding vibration output devices 31s, making it easier to generate low-frequency, large-amplitude vibrations. The vibration output devices 32 on the back surface 22 of the seat 20 are arranged at four locations: the four corners of the back surface 22. In addition, vibration output devices 33 are arranged in the lower part 23 of the seat 20, generating vibrations (low frequency, large amplitude) that rock the entire seat 20. If the entire seat 20 is structured to be mounted on the vibration output devices 33, the vibrations generated by the vibration output devices 33 can rock the entire seat 20.
[0080] Structural vibrations are handled by vibration output devices that are advantageous for reproducing low frequencies, and in this case, vibration output device 31c and vibration output device 33 are responsible for them. In particular, strong structural vibrations that accompany the user's body movements are handled by vibration output device 33, which applies vibrations to the user that shake the user's entire body. Air vibrations are handled by vibration output devices that do not press strongly against the user, i.e., have low contact with the user, and in this case, vibration output device 32, which is less likely to be subjected to pressure from the user's weight, is responsible for them. Considering that vibration output devices 31s at the four corners of seat surface 21 do not apply vibrations that shake the user's entire body and have relatively high contact with the user, here, they are responsible for either structural vibrations or air vibrations as appropriate depending on the content being played and the scene.
[0081] 6 are merely an example, and the type, installation location, number, etc. of the vibration output devices 30, the state of contact between the vibration output devices 30 and the viewing user, and the content to be played (particularly when only specific content is played, such as in a dedicated game machine) determine whether structural vibration or air vibration is to be used, and the vibrating sheet P3 is designed and assembled accordingly. The state of contact between the vibration output devices 30 and the viewing user can be detected based on operation input by the user or images of the user's viewing state captured by a camera, and if the contact state is close, the purpose is structural vibration, and if the contact state is not close, the purpose is air vibration.
[0082] 7 is an explanatory diagram showing an overview of operations such as setting vibration control parameter values performed by the content playback device 10 of FIG. 3 and operations of the content playback device 10. These operations are roughly divided into a preparation stage before content playback (step S11) and operations during content playback (step S12).
[0083] In step S11, a preparation stage before content playback, the content viewing user or the like sets hardware conditions (such as the type, installation location, and number of vibration output devices 30, as well as the contact state between the vibration output devices 30 and the viewing user). These hardware conditions (such as the installation conditions of the vibration output devices 30) are input by the content viewing user or the like operating an input device such as a keyboard, or by searching a database based on the model of the vibrating sheet P3. It is also possible to automatically input the hardware conditions by inputting information provided by the connected vibrating sheet P3 (such as configuration information for the vibrating sheet P3 stored in a storage device and read and provided when the content playback device 10 is connected). The contact state between the vibration output devices 30 and the viewing user is detected based on the result of the content viewing user's operation of an input device such as a keyboard, or based on the analysis of a camera image capturing the state of the content viewing user.
[0084] Based on the information on the installation conditions and contact states, the controller 13 determines the role of each vibration output device 30, i.e., whether it is to generate structural vibration, aerial vibration, or both, and stores the determined information in the storage unit 12. The information on the role of each vibration output device 30 includes information such as a correction coefficient for each vibration generation (e.g., a coefficient of 1 when structural vibration occurs, and a coefficient of 0.5 when aerial vibration occurs), and this information is used when generating a vibration signal (e.g., multiplying the basic structural vibration signal by the coefficient to generate the vibration signal of the corresponding vibration output device 30). These determinations are made based on, for example, a data table (generated through experiments by the designers and developers of the content playback device 10) that stores data indicating the relationship between the installation conditions and contact states and the role of each installed vibration output device 30.
[0085] Next, the controller 13 analyzes the content to extract each scene in which vibration should be generated, and detects conditions related to vibration in each scene, such as the user's ground contact state in the scene. Then, for each scene in which vibration should be generated, the controller 13 determines vibration control parameter values for each vibration output device 30 based on the conditions related to vibration in the scene and the above-mentioned assigned information for each vibration output device 30. Furthermore, the controller 13 stores the determined parameter information for each scene and for each vibration output device 30 in the parameter information DB 124.
[0086] 9 , various acoustic features of the scene, vibration classification, the ground contact state of the user in the content, and vibration control parameter values are stored in the scene recognition DB 125 of the content playback device 10 in association with the scene type. Note that the data in the scene recognition DB 125 is generated by, for example, a designer or developer of the content playback device 10 through experiments or the like, and is stored in the scene recognition DB 125.
[0087] If necessary (for example, depending on the physique of the content viewing user and the ground environment of the content playback device 10), the content viewing user manually adjusts the assignment of each vibration output device 30 and the vibration control parameter value (step S13).
[0088] During content playback in step S12, the controller 13 (scene recognition unit 135) of the content playback device 10 determines, based on the acoustic features and video features of the scene in the content being played, which scene of the content being played corresponds to which of the scenes stored in the scene recognition DB 125. In other words, the controller 13 (scene recognition unit 135) compares the acoustic features and video features of the scene in the content being played with the feature data in the scene recognition DB 125, and detects a scene in a data record in the scene recognition DB 125 with matching features.
[0089] The controller 13 (output unit 134) then reads out vibration control parameter values corresponding to the determined (detected) scene (vibration control parameter values of the same data record in the scene recognition DB 125). If the content being played back contains multiple scenes for which vibration should be generated, the controller 13 (priority setting unit 132) selects the scene with the highest priority as the target scene for generating vibration. The controller 13 (output unit 134) then generates vibration control signals for each vibration output device 30 based on the read vibration control parameter values and audio information of the content being played back, and outputs the signals to each vibration output device 30.
[0090] That is, for each scene recognized from the content, the controller 13 generates a vibration control signal for each vibration output device 30 based on the vibration control parameter value for each vibration output device 30, and controls the vibration generated by each vibration output device 30. The vibration control parameter value is set based on conditions such as the installation state of each vibration output device 30 on the vibrating sheet P3 (vibration output mechanism). Therefore, the controller 13 generates a vibration control signal for each vibration output device 30 based on the scene of the content being played back, as well as the type, installation location, and number of vibration output devices 30, and controls the vibration generated by the vibration output device 30.
[0091] This allows each vibration output device 30 to vibrate appropriately for various vibrating sheets P3 (various vibration output mechanisms) that differ in type, installation location, and number of vibration output devices 30. Also, each vibration output device 30 can be vibrated appropriately for a variety of content and its scenes that differ in type and content. Furthermore, by automating the setting of vibration control parameter values for each vibration output device 30 by the controller 13, it becomes possible to efficiently set and adjust vibration control parameter values in accordance with changes and modifications to the vibrating sheet P3 (various vibration output mechanisms).
[0092] Furthermore, the controller 13 sets a vibration control parameter value for each scene of the content being played, and controls the vibration generated by the vibration output device 30 based on the vibration control parameter value. For example, if the content is a music video, appropriate vibrations can be applied to the user for slow songs such as ballads and other up-tempo songs. Also, for example, if the content is an animal documentary, appropriate vibrations can be applied to the user for scenes of elephants walking and scenes of horses running. In other words, by applying vibrations appropriate for the scenes of the content being played to the user, the sense of realism of the content can be improved.
[0093] The controller 13 may set a vibration control parameter value for each piece of content being played back according to the content type (the overall content content) rather than for each scene of the content being played back, and vibrate each vibration output device 30 based on the vibration control parameter value. In this case, for example, if the content is a music video, a vibration suitable for allowing the user to listen to music can be imparted to the user. Also, for example, if the content is an animal documentary, a vibration suitable for allowing the user to observe or explain the living animals can be imparted to the user. In other words, imparting vibrations suitable for the content content to the user can improve the sense of realism of the content.
[0094] 4. Scene Recognition Processing During Content Playback> FIG. 8 is a flowchart showing scene recognition processing for content currently being played back, executed by the controller 13 of the content playback device 10 of FIG. 3. More specifically, FIG. 8 is a flowchart showing scene recognition processing for determining an appropriate vibration classification for vibrations occurring for each scene of content currently being played back by the controller 13 of the content playback device 10. This flowchart illustrates the technical content of a computer program that causes a computer device to perform the scene recognition processing. The computer program may be provided (sold, distributed, etc.) in the form of, for example, various readable non-volatile recording media on which the computer program is stored, or downloaded via a communication line from a server on which the computer program is stored. The computer program may consist of only one program, or may consist of multiple programs that work together.
[0095] This scene recognition is started in the pre-content playback preparation stage S11, for example, by a user's operation to start pre-content playback preparation.
[0096] FIG. 9 is a diagram showing an example of the scene recognition DB 125. More specifically, FIG. 9 is a database of information for determining the vibration classification (air vibration, structural vibration) and vibration control parameter value for each scene based on the values of each item resulting from the scene recognition processing of FIG. 8. The vibration classification data of a data record whose value of each item in FIG. 9 matches the scene recognition processing result becomes the vibration type of generated vibration appropriate for that scene, and the vibration control parameter value becomes the vibration control parameter value of generated vibration appropriate for that scene. It is also possible to set a scene so that it corresponds to both vibration classifications (air vibration, structural vibration). In that case, the vibration output devices 30 responsible for air vibration and structural vibration will generate vibrations according to the corresponding vibration control parameter values.
[0097] In this example, the scene recognition DB 125 is a database (data table) of recognized scenes corresponding to the audio in the content and the sitting / standing status of the virtual character relative to the content viewing user. To make it easier to understand the relationship between the value of each item and the vibration classification, data records with the same data for each item are displayed in a format where the frames for each item are appropriately combined. Furthermore, since the data for the standing status of the virtual character is in the same data record format as the sitting status, the entire data is not displayed and details are omitted.
[0098] As shown in FIG. 9 , the data items in the scene recognition DB 125 include characteristics of the sound (audio data) in a content scene, such as "frequency," "amplitude," "steadiness of the sound source," "steadiness of the heard sound," "sense of pitch," "number of simultaneous directions," and "volume of low-frequency sounds (of the heard sound)," as well as a "grounding" item indicating the seated or standing position of the virtual character relative to the viewing user, which can be determined from image data in the content, and vibration classifications of vibrations generated appropriate for the content situation indicated by the other items, which are associated with the recognized scene (stored in data records generated for each recognized scene). Another data item in the scene recognition DB 125 is "vibration control parameters," which store vibration control parameter values appropriate for the scene in the corresponding data record. The various data in the scene recognition DB 125 are generated and stored by designers of the content playback device 10, who select appropriate scenes based on experiments and other factors, and generate and store audio characteristic data corresponding to the scene, vibration control parameter values appropriate for the scene, and so on.
[0099] It should be noted that the "vibration control parameters" may be stored in a separate database as parameter information DB 124 as shown in FIG. 5, and the data records of both databases may be associated by scene type (data for the items "recognition scene" and "scene type").
[0100] The scene recognition DB 125 also has an item for "recognized scene" that indicates the scene type, and scene type identification data (scene name, etc.) that matches the data for each item related to the characteristics of the audio (audio data) in the scene of the above-mentioned content is stored as "recognized scene" data.
[0101] The process shown in Figure 8 is a flowchart showing vibration control processing during content playback, and is executed repeatedly by the controller 13 during content playback (at a period that does not cause discomfort to the content viewing user due to the impact of delays in changing vibration control parameter values associated with scene changes).
[0102] In step S101, the controller 13 (scene recognition unit 135) inputs and analyzes the audio signal of a scene of the content being played back, extracts data related to each item of audio characteristics, and then proceeds to step S102. Note that this audio signal analysis can be realized by digitizing the audio signal and performing various processes such as frequency resolution through arithmetic processing. Alternatively, a video signal of a scene of the content being played back may be input and analyzed, and used to determine the audio characteristics described below.
[0103] In step S102, the controller 13 (scene recognition unit 135) determines whether the frequency band (main band) of the sound in the scene is high or low (determined based on its hierarchical relationship with a threshold), and then proceeds to step S103. That is, for example, the threshold is set to a frequency of 20 Hz, and the scene recognition unit 135 determines that the sound has a low frequency when the sound intensity distribution below 20 Hz is high, and determines that the sound has a high frequency when the sound intensity distribution above 20 Hz is low.
[0104] In step S103, the controller 13 (scene recognition unit 135) determines whether the amplitude (average value or maximum value) of the sound is large or small (determined based on the relationship between the amplitude and the threshold), and then proceeds to step S104.
[0105] In step S104, the controller 13 (scene recognition unit 135) determines whether the sound source is stationary or non-stationary (whether the sound source emits a continuous sound or a sudden sound), and proceeds to step S105. A sound source is an object within the content that is emitting a sound. A sound source that emits a similar sound that continues continuously, such as the sound of a car running or an animal's footsteps, is considered stationary, while a sound source that emits a sudden sound, such as a car horn or an animal's cry, is considered non-stationary. As an example of this determination method, the scene recognition unit 135 determines that a sound source is stationary when the dynamic range of the sound of the sound source is below a predetermined threshold, and determines that the sound is non-stationary when the dynamic range exceeds the threshold.
[0106] In step S105, the controller 13 (scene recognition unit 135) determines whether the audible sound is stationary or non-stationary, and proceeds to step S106. The audible sound is the sound heard by a user (an avatar corresponding to the user) in content (e.g., an XR content space) and the sound recorded by a microphone worn by the user in the content. The scene recognition unit 135 determines that the audible sound is stationary when the dynamic range of the audible sound is below a predetermined threshold, and determines that the audible sound is non-stationary when the dynamic range exceeds the threshold.
[0107] In step S106, the controller 13 (scene recognition unit 135) determines whether the sense of pitch of the audio is strong or weak, and then proceeds to step S107. For example, the scene recognition unit 135 determines that the sense of pitch is strong when the audio has large pitch fluctuations (frequency fluctuations) (such as animal cries with large pitch fluctuations over a wide frequency range), and determines that the sense of pitch is low when the pitch fluctuations are small (such as steady mechanical sounds with a narrow frequency range).
[0108] In step S107, the controller 13 (scene recognition unit 135) determines whether the sound is coming from a single direction or from multiple directions simultaneously, and then proceeds to step S108. The scene recognition unit 135 determines that the sound is coming from multiple directions simultaneously when, for example, multiple sound sources are present at a distance closer to the user than a predetermined threshold, and otherwise determines that the sound is coming from a single direction.
[0109] In step S108, the controller 13 (scene recognition unit 135) determines whether the signal level of the low-frequency components of the audio is high or low, and proceeds to step S109. The scene recognition unit 135 determines that the low-frequency components of the audio are high when the sound pressure of audio components below a predetermined frequency exceeds a predetermined threshold, for example, and determines that the low-frequency components of the audio are low when the sound pressure is below the threshold.
[0110] In step S109, the controller 13 (scene recognition unit 135) performs data matching processing on the scene recognition DB 125 using the determination results from steps S101 to S108, determines the corresponding scene, and proceeds to step S110.
[0111] In step S110, the controller 13 (output unit 134) extracts vibration control parameter values corresponding to each vibration output device 30 according to the determined scene from the scene recognition DB 125, processes the audio signal of the content using the vibration control parameter values, and generates a vibration signal.The generated vibration signal is then output to each corresponding vibration output device 30 to generate a desired vibration, and the process ends.Note that the controller 13 (output unit 134) continues to generate vibration according to the content (audio) of the content being played using the vibration control parameter values before the update, until the vibration control parameter values are updated.
[0112] 8, appropriate vibration control parameter values are set for each vibration output device 30 in the content reproduction system PS according to the hardware configuration of the content reproduction system PS and the characteristics of the audio during content reproduction, and each vibration output device 30 vibrates with vibrations generated based on the set vibration control parameter values. Therefore, appropriate vibrations according to the hardware configuration of the content reproduction system PS and the content details can be imparted to the content viewing user, allowing the content viewing user to enjoy content reproduction with a rich sense of realism.
[0113] In this processing example, the vibration control parameter values are determined (calculated) when the content is played back, but the vibration control parameter values of the content to be played back may be calculated in advance using a similar method and stored in association with the playback scene (playback time, scene, etc.) of the content. Then, when the content is played back, the vibration control parameter values associated with the corresponding scene may be extracted to perform vibration control. In this case, the vibration control parameter values may be recorded as part of the content information on a content recording medium, for example, a content optical disc, together with the content body information (video and audio information).
[0114] 10A, 10B, and 10C are explanatory diagrams showing first, second, and third examples of the role (structural vibration, air vibration) of each vibration output device 30 in the content reproduction device 10 of Fig. 6. The role of each vibration output device 30 differs depending on the hardware configuration of the content reproduction system PS and the content (scene) details.
[0115] In this embodiment, in Fig. 10A, the vibration output devices 31s at the four corners of the seat 21 and the four vibration output devices 32 on the back surface 22 are responsible for air vibrations and structural vibrations, respectively. In Fig. 10B, the vibration output devices 31s at the four corners of the seat 21 and the four vibration output devices 32 on the back surface 22 are responsible for air vibrations, and the vibration output device 31c at the center of the seat 21 is responsible for structural vibrations. In Fig. 10C, the vibration output devices 31s at the four corners of the seat 21 and the four vibration output devices 32 on the back surface 22 are responsible for air vibrations, the vibration output device 31c at the center of the seat 21 is responsible for structural vibrations (frequencies from 20 Hz to 40 Hz), and the vibration output device 33 at the lower part 23 of the seat 20 is responsible for structural vibrations (frequencies less than 20 Hz). In other words, the content playback device 10 generates and outputs vibration signals for each vibration output device 30, 31, 32 according to vibration control parameter values (set in the scene recognition DB 125) corresponding to the role of each vibration output device 30, 31, 32, thereby generating structural vibration or air vibration.
[0116] 11 is an explanatory diagram showing an example of a content reproduction system PS according to a modified example. Note that components in the modified example that are common to the embodiment described above may be given the same reference numerals or names, and descriptions thereof may be omitted.
[0117] As shown in FIG. 11, the content reproduction system PS of the modified example is connected to a server device 40 via a communication line.
[0118] The content playback device 10 includes components common to the previously described embodiments. The content playback device 10 outputs a vibration signal corresponding to the content to be played back to the vibration output device 30, and the vibration output device 30 applies vibration to the user U1.
[0119] The server device 40 is connected to the content playback device 10 via a network N so as to be able to perform two-way communication. The server device 40 may be a physical server or a virtual server. The network N may be, for example, a local area network or the Internet.
[0120] Fig. 12 is a configuration diagram showing an example of the server device 40 of Fig. 11. Fig. 12 shows components necessary for explaining the features of this embodiment, and omits the description of general components.
[0121] 12, the server device 40 includes a communication unit 41, a storage unit 42, and a controller 43. The communication unit 41 is an interface for performing data communication with other devices via a network N. The communication unit 41 is, for example, a network interface card (NIC).
[0122] The server device 40 includes components equivalent to those of the content playback device 10 of the embodiment described above. Equivalent components (same structure, operation, etc.) are designated by the same names as those in Fig. 3, and the reference numerals are prefixed with SV, and their explanations are omitted.
[0123] In the content reproduction system PS of this modified example, hardware configuration information of the content reproduction device 10, video information showing the seating state of the content viewing user U1 on the vibrating seat P3, and information on the content to be reproduced are transmitted from the content reproduction device 10 to the server device 40. In addition, vibration control signals for each vibration output device 30 of the vibrating seat P3 are transmitted from the server device 40 to the content reproduction device 10. Then, the content reproduction device 10 synchronizes the image signal and audio signal of the content and the vibration control signal from the server device 40 and outputs them to the display device P1, the speaker P2, and each vibration output device 30.
[0124] The division of roles between the content playback device 10 and the server device 40 is not limited to this modification and can be set as appropriate. For example, it is possible to give the server a content playback function as well, and for the server device 40 to transmit video signals and audio signals of the content, as well as vibration control signals to each vibration output device 30, to the content playback device 10.
[0125] By using the server device 40, the server device 40 can have information and programs that can be compatible with various types of content reproduction systems PS with different configurations, and can perform processing that corresponds to the hardware configuration of the content reproduction system PS and the content to be reproduced in response to a request from the content reproduction system PS. Therefore, according to this modification, there is an advantage that each content reproduction system PS does not need to have its own dedicated configuration, and updates of various information and programs can also be managed on the server device 40 side.
[0126] 6. Design Process of Vibrating Sheet (Vibration Output Mechanism) of Content Reproduction System Next, the design process of the vibration output mechanism will be described using the design process of the vibrating sheet P3 of the content reproduction system PS as an example.
[0127] In this design processing example, an example will be described in which server device 40 is used as a design support device for a vibration output mechanism, but it is also possible to use content playback device 10. Furthermore, this design processing can also be realized by a design system using a computer system that has components equivalent to those of server device 40 used in the processing in the following description.
[0128] 13 is a flowchart showing the design process of the vibrating sheet P3 executed by the controller 43 of the server device 40 in FIG. 12. This flowchart shows the technical content of a computer program that causes the server device 40 to realize the design process of the content playback device 10. The computer program is provided (sold, distributed, etc.) in the form of, for example, various readable non-volatile recording media on which the computer program is stored, or in the form of being downloaded via a communication line from a server on which the computer program is stored. The computer program may be composed of only one program, or may be composed of multiple programs that work together.
[0129] The process shown in FIG. 13 is executed in the server device 40 when the designer of the content playback device 10 executes the design process and performs a process start operation using an operation unit such as a keyboard.
[0130] In step S201, the controller 43 inputs audio data of the content, and then proceeds to step S202. At this time, video data of the content may also be input and used for subsequent determination processing, etc. Note that if the content to be used is content that is frequently used in the content playback system PS or content of a similar type, it is possible to design the system to be suitable for the frequently used content. For example, when designing a content playback system PS dedicated to a certain game, the game content will be used.
[0131] In step S202, the controller 43 (scene detection unit SV131) determines whether the main component of the vibration to be generated is air vibration or structural vibration based on the audio data of the content and, if necessary, with reference to the video data, and stores the result in the memory unit 12, and proceeds to step S203.
[0132] In step S203, the controller 43 (scene detection unit SV131) determines whether playback of the content has been completed (a predetermined amount required for the design), and if not completed, the process returns to step S202, and if completed, the process proceeds to step S204. In other words, by the processes of steps S202 and S203, it is possible to grasp the number of situations in which the main component of the vibration to be generated in the target content is air vibration, and the number of situations in which it is structural vibration.
[0133] In step S204, the controller 43 (scene detection unit SV131) calculates the ratio between vibration situations in which air vibrations are the main component and vibration situations in which structural vibrations are the main component in the target content, and then proceeds to step S205.
[0134] In step S205, the controller 43 inputs data on the state of the vibrating sheet P3 when all the vibration output devices 30 that can be installed thereon are installed (the position and vibration effect level of each vibration output device 30), and also inputs data such as the component price and installation cost of each vibration output device 30 and the target price of the vibrating sheet P3 to be designed, and then proceeds to step S206. Note that this information on the vibrating sheet P3 is input, for example, by the designer of the vibrating sheet P3 operating a keyboard or the like.
[0135] In step S206, the controller 43 determines the order of elimination of each vibration output device 30 based on the ratio of air vibration to structural vibration in the target content calculated in step S204 and the vibration effect level (degree of contribution to improving the sense of realism) of each vibration output device 30 input in step S205, and proceeds to step S207. In other words, the lower the vibration generation ratio of the vibration type (air vibration, structural vibration) in charge of the target content and the lower the vibration effect on the content viewing user, the earlier the device will be deleted (the higher the deletion priority).
[0136] For example, the vibration component ratio, which is the ratio of the main components of air vibration and structural vibration in the target content, is 8:3. The vibration output devices responsible for air vibration are designated A1, A2, and A3 in order of decreasing vibration effect level, and the vibration output devices responsible for structural vibration are designated B1, B2, and B3 in order of decreasing vibration effect level. Since the proportion of air vibration is high in the target content, the first deletion priority is assigned to vibration output device B3, which is the vibration output device responsible for structural vibration and has the lowest vibration effect level. Then, once one deletion priority for the vibration output device responsible for structural vibration has been determined, the proportion value of air vibration is reduced (e.g., halved) to lower the dominance of air vibration (the vibration component ratio becomes 4:3). Continuing this process, the next content also has a high proportion of air vibration, so the second deletion priority is assigned to vibration output device B2, which is the vibration output device responsible for structural vibration and has the lowest vibration effect level. The new vibration component ratio then becomes 2:3. Next, the proportion of structural vibration is high, so the third deletion priority is assigned to vibration output device A3, which is the vibration output device responsible for air vibration and has the lowest vibration effect level. This process is continued until all the deletion priorities are determined. In this case, the deletion priorities are as follows, from highest to lowest: B3, B2, A3, B1, A2, A1.
[0137] The vibration effect level is determined, for example, by subjective evaluation by members of a design and development group, for example, by statistical processing of the results of each subject's evaluation of the vibration effect level when multiple vibration output devices 30 are set to a vibration generating state in turn in response to the playback of a certain content (such as the results of a questionnaire given to each subject).
[0138] In step S207, the controller 43 eliminates the vibration output device 30 highest in the reduction order determined in step S206 among the vibration output devices 30 mounted on the vibrating sheet P3, calculates the manufacturing price of the vibrating sheet P3 in this case based on the input prices of each vibration output device 30 of the vibrating sheet P3, and proceeds to step S208. Note that in this example, the vibration output devices 30 are eliminated simply according to the reduction priority order of the vibration output devices 30, but the reduction may also take into account factors such as the cost of the vibration output devices 30 and the remaining price until the target described below is achieved.
[0139] Furthermore, this reduction in the number of components of the vibrating sheet P3 is a hypothetical process for calculating the manufacturing price, and does not actually involve a reduction in the number of components of the vibrating sheet P3 (entity). In reality, for example, a designer or the like will decide on the final specifications and design based on the results of the above-mentioned evaluation process.
[0140] In step S208, the controller 43 determines whether a predetermined target cost for the vibrating sheet P3 or a predetermined target number of vibration output devices 30 has been reached (or whether it has fallen below the target), and if it has been reached, it notifies (displays) the result (the configuration of the vibrating sheet P3 in which each vibration output device 30 has been appropriately reduced) and terminates the processing, and if it has not been reached, it returns to step S207 and continues the reduction processing and its evaluation processing.
[0141] This makes it possible to perform a simulation to reduce the vibration output devices 30 mounted on the vibrating sheet P3 in an appropriate order, while confirming the configuration of the vibrating sheet P3 that meets the target, thereby improving the efficiency of the design of the content playback system PS (vibrating sheet P3).
[0142] 7. Notes, etc. The various technical features disclosed as embodiments in this specification can be modified in various ways without departing from the spirit of the technical creation. In other words, the above-described embodiments are illustrative in all respects and are not limiting. The technical scope of the present invention is defined by the claims, not by the description of the above-described embodiments, and includes all modifications that fall within the meaning and scope of the claims. Furthermore, the multiple embodiments described in this specification may be combined as appropriate to the extent possible.
[0143] In the above embodiment, various functions are realized by software through the arithmetic processing of the CPU in accordance with a program, but at least some of these functions may be realized by electrical hardware resources. Examples of hardware resources include an ASIC (Application Specific Integrated Circuit) and an FPGA (Field Programmable Gate Array). Conversely, at least some of the functions realized by hardware resources may be realized by software.
[0144] The scope of the present embodiment may also include a computer program that causes a processor (computer) to realize at least some of the functions of the content playback device 10. The scope of the present embodiment may also include a computer-readable nonvolatile recording medium that records such a computer program. The nonvolatile recording medium may be, for example, the nonvolatile memory described above, an optical recording medium (e.g., an optical disk), a magneto-optical recording medium (e.g., a magneto-optical disk), a USB memory, an SD card, or the like.
[0145] REFERENCE SIGNS LIST 10 Content playback device 12 Storage unit 13 Controller 20 Seat 21 Seat surface 22 Back surface 23 Lower portion 30, 31, 31c, 31s, 32, 33 Vibration output device 40 Server device (design support device) P1 Display device P2 Speaker P3 Vibrating seat (vibration output mechanism) PS Content playback system U1 User
Claims
1. A vibration signal generating device for generating a vibration signal according to content, A controller is provided. The controller: Acquire an audio signal of the content; Detecting a plurality of vibration output devices installed in the vibration output mechanism; Acquire a parameter value for each of the detected vibration output devices; generating the vibration signal for each of the plurality of vibration output devices based on the audio signal and the parameter value for each of the vibration output devices; Vibration signal generator.
2. A memory for storing the parameter values for each of the plurality of vibration output devices, The controller: obtaining the parameter value corresponding to the vibration output device from the memory; The vibration signal generating device according to claim 1 .
3. The memory stores the parameter values according to the scenes of the content, The controller: Detecting the scene of the content being played; obtaining the parameter value corresponding to the scene and the vibration output device from the memory; The vibration signal generating device according to claim 2 .
4. The controller: Separating the audio signal into an air vibration audio signal and a structural vibration audio signal; generating an air vibration signal based on the sound signal for air vibration and the parameter value corresponding to the air vibration; generating a structural vibration signal based on the structural vibration sound signal and the parameter value corresponding to the structural vibration; The vibration signal generating device according to claim 1 .
5. The parameter value is a cutoff frequency of a low-pass filter that extracts low-frequency components from the audio signal, a delay time that delays the audio signal, or an amplification factor for the audio signal. The vibration signal generating device according to any one of claims 1 to 4.
6. A content playback device that outputs a video signal, an audio signal, and a vibration signal according to a content, A controller is provided. The controller: Acquire an audio signal of the content; Detecting a plurality of vibration output devices installed in the vibration output mechanism; Acquire a parameter value for each of the detected vibration output devices; generating the vibration signal for each of the plurality of vibration output devices based on the audio signal and the parameter value for each of the vibration output devices; Content playback device.
7. The parameter value is a cutoff frequency of a low-pass filter that extracts low-frequency components from the audio signal, a delay time that delays the audio signal, or an amplification factor for the audio signal. The content reproducing device according to claim 6.
8. a video display that displays a video according to the video signal; an audio output device that outputs audio in response to an audio signal; a vibration output mechanism that outputs a vibration in response to a vibration signal; a content reproducing device that outputs the video signal, the audio signal, and the vibration signal according to a content to the video display device, the audio output device, and the vibration output mechanism; A content playback system including: The content playback device includes a controller, The controller: obtaining the audio signal of the content; Detecting a plurality of vibration output devices installed in the vibration output mechanism; Acquire a parameter value for each of the detected vibration output devices; generating the vibration signal for each of the plurality of vibration output devices based on the audio signal and the parameter value for each of the vibration output devices; Content playback system.
9. The parameter value is a cutoff frequency of a low-pass filter that extracts low-frequency components from the audio signal, a delay time that delays the audio signal, or an amplification factor for the audio signal. The content reproduction system according to claim 8.
10. A vibration signal generating method executed by a controller to generate a vibration signal according to content, comprising: Acquire an audio signal of the content; Detecting a plurality of vibration output devices installed in the vibration output mechanism; Acquire a parameter value for each of the detected vibration output devices; generating the vibration signal for each of the plurality of vibration output devices based on the audio signal and the parameter value for each of the vibration output devices; Vibration signal generation method.
11. The parameter value is a cutoff frequency of a low-pass filter that extracts low-frequency components from the audio signal, a delay time that delays the audio signal, or an amplification factor for the audio signal. The vibration signal generating method according to claim 10.
12. A vibration signal generating program executed by a controller to generate a vibration signal according to content, Acquire an audio signal of the content; Detecting a plurality of vibration output devices installed in the vibration output mechanism; Acquire a parameter value for each of the detected vibration output devices; generating the vibration signal for each of the plurality of vibration output devices based on the audio signal and the parameter value for each of the vibration output devices; Vibration signal generator.
13. The parameter value is a cutoff frequency of a low-pass filter that extracts low-frequency components from the audio signal, a delay time that delays the audio signal, or an amplification factor for the audio signal. The vibration signal generating program according to claim 12.
14. A design support device for a vibration output mechanism that has a plurality of vibration output devices and provides a user with vibration corresponding to content being played, comprising: A controller is provided. The controller: Detecting an installation state of the plurality of vibration output devices relative to the vibration output mechanism; determining a reduction priority order for each of the vibration output devices based on vibration effect levels of the vibration output devices; setting the vibration output device to be used according to the reduction priority order; outputting a vibration control signal generated according to a matching target content for matching a vibration operation of the vibration output mechanism for each of the set plurality of vibration output devices to each of the plurality of vibration output devices; Calculating cost information when the vibration output mechanism is configured with the plurality of vibration output devices that have been set. Design support equipment.
15. The controller: The vibration output devices are divided into those for generating air vibrations and those for generating structural vibrations, and the reduction priority order is determined; setting the vibration output device to be used according to a ratio of a scene suitable for generating air vibration to a scene suitable for generating structural vibration in the matching target content suitable for generating air vibration and structural vibration; The design support device according to claim 14.