Acoustic signal generation device, acoustic signal generation method, acoustic signal generation program, and content playback system
The acoustic signal generation device addresses the issue of unaccounted sounds from vibrations by combining content and vibration-acoustic signals, enhancing the sense of realism during content playback.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- DENSO TEN LTD
- Filing Date
- 2024-11-21
- Publication Date
- 2026-06-02
AI Technical Summary
Existing techniques fail to consider the sounds generated by vibrations transmitted to users, leading to a diminished sense of presence during content playback.
An acoustic signal generation device that combines content acoustic signals with vibration-acoustic signals generated by objects in the content, accounting for vibrations to enhance the sense of realism.
Enhances the sense of realism by outputting sound based on vibrations occurring in the content, improving the overall user experience.
Smart Images

Figure 2026089845000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an acoustic signal generation device, an acoustic signal generation method, an acoustic signal generation program, and a content playback system.
Background Art
[0002] Conventionally, there has been proposed a technique for improving the sense of presence of content by transmitting vibrations and sounds corresponding to the content viewed by the user to the user. For example, there is known a technique for improving the sense of presence by extracting vibration parameters and voice parameters corresponding to a scene detected from the content and outputting vibration data and voice data subjected to enhancement processing using the parameters (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In reality, when vibrations are transmitted to the user, objects around the user (for example, walls, etc.) also vibrate, and sounds are generated by the vibrations. However, in the prior art, there has been a problem that the sounds generated by the vibrations are not taken into consideration. As a result, there has been a concern that the user cannot feel a sufficient sense of presence.
[0005] In view of the above problems, an object of the present invention is to provide a technique capable of improving the sense of presence during content playback.
Means for Solving the Problems
[0006] An exemplary acoustic signal generation apparatus of the present invention is an acoustic signal generation apparatus that generates an acoustic signal according to content, which acquires a content vibration signal according to the content, generates a vibration acoustic signal for sound generated by vibration in the content based on the content vibration signal, and generates the acoustic signal by combining the content acoustic signal included in the content and the vibration acoustic signal. [Effects of the Invention]
[0007] According to the present invention, a vibration-acoustic signal is generated based on vibrations occurring in the content being played, which corresponds to the sound generated as a result of those vibrations. Then, an acoustic signal can be generated by combining this vibration-acoustic signal with the content acoustic signal included in the content and output during content playback. Therefore, by outputting sound based on the acoustic signal affected by vibrations during content playback, it becomes possible to enhance the sense of realism. [Brief explanation of the drawing]
[0008] [Figure 1] Overall configuration diagram of the content playback system of this embodiment [Figure 2] Block diagram showing the configuration of the acoustic signal generation device in Figure 1. [Figure 3] Figure 1 is an explanatory diagram showing an overview of the acoustic signal generation method in the acoustic signal generation device. [Figure 4] A schematic diagram showing an example of content video displayed on the display device in Figure 1. [Figure 5] Block diagram showing the configuration of the scene detection unit of the acoustic signal generation device of Example 1. [Figure 6] Block diagram showing the configuration of the vibration acoustic signal generation unit of the acoustic signal generation device of Example 1. [Figure 7] An explanatory diagram showing the acoustic signal generation method in the acoustic signal generation device of Example 1. [Figure 8A] A diagram showing an example of a vibroacoustic data table. [Figure 8B] Figure 8A shows a modified example of the vibration-acoustic data table. [Figure 9] A flowchart showing the acoustic signal generation process performed by the acoustic signal generation device of Example 1. [Figure 10] Block diagram showing the configuration of the acoustic signal generation device in Example 2 [Figure 11] An explanatory diagram showing the acoustic signal generation method in the acoustic signal generation device of Example 2. [Figure 12] A diagram showing an example of a weighted data table. [Figure 13] A flowchart illustrating the acoustic signal generation process performed by the acoustic signal generation device of Example 2. [Modes for carrying out the invention]
[0009] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the drawings. However, the present invention is not limited to the embodiments described below.
[0010] <1. Content Playback System> Figure 1 is an overall configuration diagram of the content playback system 1 of this embodiment. In this embodiment, the content playback system 1 is a system that outputs video, audio, and vibration according to the content. The content playback system 1 comprises a content playback device 10, a display device 20, a speaker device 30, a vibration device 40, and an acoustic signal generation device 50.
[0011] The content playback device 10 acquires video signals, audio signals, and vibration signals from the data of the content to be played back. The content playback device 10 outputs the video signal to the display device 20, the audio signal to the speaker device 30, and the vibration signal to the vibration device 40. For example, the content playback device 10 reads video data, audio data, and vibration data from a recording medium such as an optical disc, or acquires video data, audio data, and vibration data distributed via the internet, etc., and generates and outputs video signals, audio signals, and vibration signals based on this data.
[0012] The display device 20 is a device that provides the user U1 with video corresponding to the content to be played, which is based on the video signal of the content playback device 10. The speaker device 30 is a device that provides the user U1 with sound corresponding to the content to be played, which is based on the sound signal of the content playback device 10.
[0013] The vibration device 40 is provided on the seating surface (seat part) of the seat Se on which the user U1 sits. The vibration device 40 is constituted by, for example, an electric vibration converter including an electromagnetic circuit, a piezoelectric element, and an electric cylinder. The vibration device 40 generates vibration according to the vibrator drive signal and outputs the vibration. Note that the vibrator drive signal is a signal obtained by performing appropriate amplification processing, frequency characteristic adjustment processing, etc. on the vibration signal of the content playback device 10 according to the characteristics of the vibration device 40 (such as the conversion efficiency from signal to vibration). In other words, the vibration device 40 generates vibration corresponding to the content to be played, which is based on the vibrator drive signal output from the content playback device 10, and outputs the vibration toward the user U1.
[0014] The acoustic signal generation device 50 is a device that generates a control signal (acoustic signal) for driving the speaker device 30 for acoustic reproduction according to the content to be played. The acoustic signal generation device 50 and the speaker device 30 are connected by wire via the amplifier 21S (see FIG. 2). Note that the amplifier 21S amplifies the acoustic signal in power and outputs it to the speaker device 30, causing the speaker device 30 to output sound (voice) to the outside.
[0015] <2. Acoustic Signal Generation Device> <2-1. Outline of Acoustic Signal Generation Device> FIG. 2 is a block diagram showing the configuration of the acoustic signal generation device 50 of FIG. 1. In FIG. 2, the components necessary for explaining the features of the present embodiment are shown, and the description of general components is omitted.
[0016] The acoustic signal generation device 50 includes a communication unit 51, a storage unit 52, and a controller 53.
[0017] The communication unit 51 is an interface for communicating data with other devices (content playback device 10, speaker device 30) via a communication network. The communication unit 51 includes communication equipment for wired and wireless communication with other devices. The wireless communication equipment consists of, for example, a transceiver for a 5G (fifth-generation mobile communication system) mobile telephone network. For example, when using an external content provision service, content (data) will be received from the content provision service server via 5G communication or the like through the communication unit 51.
[0018] The storage unit 52 is configured to include volatile memory and non-volatile memory, and stores various information necessary for content playback processing. The volatile memory is composed of, for example, RAM (Random Access Memory). The non-volatile memory is composed of, for example, ROM (Read Only Memory), flash memory, or a hard disk drive. Programs and data that can be read by the controller 53 are stored in the non-volatile memory. At least a portion of the programs and data stored in the non-volatile memory may be obtained from other computer devices connected by wired or wireless connections, or from portable recording media.
[0019] The memory unit 52 stores an acoustic signal generation program 521 and a vibration acoustic data table 522. The contents of these programs, data tables, etc., stored in the memory unit 52 are described separately. Furthermore, the memory unit 52 also stores data tables, etc. (not shown) for various processing.
[0020] The controller 53 consists of a processor that performs calculations and other processing, and controls various operations in the sound signal generation device 50. The processor includes, for example, a CPU (Central Processing Unit). The controller 53 executes the sound signal generation program 521 stored in the memory unit 52 and performs sound signal generation processing. When generating sound signals, the controller 53 communicates necessary data with the content playback device 10, the display device 20, and the vibration device 40 as needed to adjust the timing and level of sound playback, video playback, and vibration playback. The sound signal generation program 521 includes various programs that realize various functions of the sound signal generation device 50.
[0021] The controller 53 includes, as its functions, an acquisition unit 531, a scene detection unit 532, a vibration acoustic signal generation unit 533, a synthesis unit 534, and an output unit 535. In this embodiment, the functions of the controller 53 are realized by the processor executing calculation processing according to the data of the acoustic signal generation program 521 and the vibration acoustic data table 522 stored in the storage unit 52.
[0022] The acquisition unit 531 acquires (receives) content data corresponding to the content to be played back from the content playback device 10 via the communication network and the communication unit 51. This content data includes a content video signal PS, a content audio signal SS, and a content vibration signal VS, and may also include content supplementary data CD containing various data related to the content, such as data about objects existing in the content space (position, shape, etc.). These various content data will be created as appropriate by the content creator.
[0023] The scene detection unit 532 performs scene detection of the content being played based on the content data acquired by the acquisition unit 531. For example, the scene detection unit 532 performs image analysis processing, audio analysis processing, 3D model analysis processing, etc. on the content data, or extracts corresponding data from the content supplement data CD, to acquire scene information such as the timing of scene changes, the duration of the scene, the content of the scene, and the vibrating objects included in the scene. In particular, by performing scene detection, the scene detection unit 532 acquires information related to the vibrating objects included in the scene as scene information. A vibrating object is an object that vibrates in response to applied vibrations, and that vibration causes the surrounding air to vibrate, generating sound at a level that the user can perceive.
[0024] The vibration acoustic signal generation unit 533 generates a vibration acoustic signal VSS based on the information relating to the vibrating object acquired by the scene detection unit 532 and the content vibration signal VS. Unlike the content acoustic signal SS which is originally included in the content itself, the vibration acoustic signal VSS is an acoustic signal generated by an object that vibrates based on the vibrations that occur in the content being played.
[0025] The synthesis unit 534 synthesizes the content sound signal SS and the vibration sound signal VSS (signal addition processing) to generate a sound signal (synthesized sound signal CSS) corresponding to the content. The output unit 535 outputs the synthesized sound signal CSS synthesized by the synthesis unit 534 to the speaker device 30 via the amplifier 21S.
[0026] <2-2. Overview of signal transitions in the acoustic signal generation device 50> Next, the signal transitions in the acoustic signal generation device 50 will be explained. When vibrations are transmitted to the user, objects around the user also vibrate, generating sound. Therefore, the acoustic signal generation device 50 performs acoustic signal generation processing that takes into account the sound generated based on the vibrations. This acoustic signal generation method is executed continuously in real time when content is played.
[0027] In this explanation, the content played back by the content playback device 10 is content that includes video signals, audio signals, and vibration signals, and includes, for example, XR (Cross Reality) content, movies, concert videos, games, and other similar content. In this embodiment, the content played back by the content playback device 10 is described as a "racing game." The audio signals originally included in the content are referred to as "content audio signals," and the vibration signals originally included in the content are referred to as "content vibration signals."
[0028] Furthermore, the operator (user: content viewer) within the content space is referred to as the "user avatar," while the user in the real world is simply referred to as the "user" unless otherwise specified. While viewing content on Content Playback System 1, the user will hear sounds that simulate the surrounding sounds heard by the user avatar in the content space.
[0029] Figure 3 is an explanatory diagram showing an overview of the acoustic signal generation method in the acoustic signal generation device 50 shown in Figure 1.
[0030] The content playback device 10 outputs a content video signal PS, a content audio signal SS, a content vibration signal VS, and content supplementary data CD, which contains various information about the content, as included in the content (content data stored on an optical disc, etc.). The display device 20 displays an image based on the content video signal PS and provides the image to the user viewing the display device 20. The vibration device 40 vibrates a vibrating member based on the content vibration signal VS and applies vibration to the user in contact with the vibrating member. The sound signal generation device 50 acquires (inputs) the content audio signal SS, the content vibration signal VS, and the content supplementary data CD and performs the following processing.
[0031] The acoustic signal generator 50 (scene detection unit 532) detects the scene being played from the content-added data CD acquired by the acquisition unit 531 from the content playback device 10, and acquires scene information SD. The scene information SD includes information about the vibration-sound generating object included in the scene. The vibration-sound generating object is an object that generates sound when affected by vibrations in the content space.
[0032] For example, in the case of a racing game, as shown in Figure 4, the objects that generate vibration and sound include other vehicles Ve2 around the player's own vehicle Ve1, buildings St1, and interior components Iv of the vehicle's cabin. These objects can be detected based on content-added data CD (which may include object data such as the location of objects in the content space), or they can be detected using data such as the location of objects in the content space calculated by image analysis of the content video signal PS. Figure 4 is a schematic diagram showing an example of content video (racing game video) displayed on the display device 20 in Figure 1.
[0033] Next, the acoustic signal generation device 50 (vibration acoustic signal generation unit 533) generates a vibration acoustic signal VSS for vibrations generated by an object that generates vibration sounds, based on the content vibration signal VS from the content playback device 10, the scene information SD detected by the scene detection unit 532, and the vibrations corresponding to the content vibration signal VS. Specifically, the vibration acoustic signal VSS is an acoustic signal that reproduces the sound generated by the object that generates vibration sounds due to vibrations occurring in the scene being played in the content.
[0034] The sound generated by a vibratory acoustics generating object is primarily determined by the vibrations applied to the object, the vibration characteristics of the object (determined by magnitude, physical properties, etc.), and the position of the object (the positional relationship between the user avatar and the vibratory acoustics generating object in the content space). Therefore, by calculating this data based on scene information SD and content vibration signal VS, and then applying this calculated data to appropriately defined calculation formulas or artificial intelligence, a vibratory acoustics signal VSS can be generated.
[0035] For example, in a simplified manner, the type of object generating vibration and sound, and the positional relationship between the user avatar and the object generating vibration and sound in the content space are used as two parameters. The parameter values of the vibration and sound signal calculation formula (for example, determined by the design and development team based on experiments, etc.) determined by these two parameter values are stored in a data table (the parameter values for vibration and sound signal calculation are set by the design and development team based on experiments, etc.). Then, by processing the content vibration signal VS with the vibration and sound signal calculation formula to which these determined parameter values are applied, a vibration and sound signal VSS can be generated.
[0036] For example, in the racing game shown in Figure 4, when the vehicle Ve1 drives on an unpaved road such as a gravel road, vibrations occur, causing the dashboard Iv1 in the cabin, which is the object that generates vibration-acoustic sounds, to vibrate and generate a vibration-acoustic signal VSS related to the creaking sound of the dashboard Iv1. This creaking sound is reproduced by the vibration-acoustic signal VSS, which is generated by processing the content vibration signal VS (corresponding to vibrations that occur when driving on an unpaved road) with a vibration-acoustic signal calculation formula that applies parameter values determined based on the type of object that generates vibration-acoustic sounds in the data table, which is dashboard Iv1, and the positional relationship (e.g., distance) between dashboard Iv1 and the driver (user avatar).
[0037] Furthermore, there may be multiple objects being vibrated, resulting in the generation of multiple VSS (Vibration Acoustic Signals) corresponding to each object. Also, since the objects being vibrated differ from scene to scene, the VSS will vary from scene to scene.
[0038] Next, the sound signal generation device 50 (synthesis unit 534) synthesizes the content sound signal SS and the vibration sound signal VSS to generate a sound signal (synthesized sound signal CSS) corresponding to the content. For example, in the case of a racing game (a scene where there are no other vehicles or other objects that generate vibration sounds in the surroundings, and the only object that generates vibration sounds inside the car is the dashboard Iv1 in the car), the sounds heard by the user avatar are the sound of the vehicle Ve1 driving, which is based on the content sound signal SS, and the creaking sound generated by the vibration of the dashboard Iv1 in the car due to driving vibrations. In this scene, the sound signal generation device 50 synthesizes the content sound signal SS corresponding to the sound of the vehicle Ve1 driving and the vibration sound signal VSS corresponding to the creaking sound of the dashboard Iv1 generated by the vehicle vibrations, and as a result, a synthesized sound signal CSS corresponding to the sounds heard by the user avatar is generated.
[0039] The acoustic signal generator 50 (output unit 535) then outputs the generated synthesized acoustic signal CSS to the speaker device 30. The synthesized acoustic signal CSS generated by the acoustic signal generator 50 is subjected to necessary processing such as power amplification before being input to the speaker device 30.
[0040] According to the above configuration, a vibration acoustic signal VSS is generated based on vibrations occurring in the content being played, which corresponds to the sound generated by those vibrations. Then, by combining the vibration acoustic signal VSS with the content acoustic signal SS, a synthesized acoustic signal CSS can be generated that corresponds to the sound that the user avatar will hear. Therefore, by outputting sound based on the synthesized acoustic signal CSS, which is affected by vibrations, during content playback, it becomes possible to enhance the sense of realism.
[0041] <3. Example 1> <3-1. Details of the acoustic signal generation device in Example 1> Regarding the acoustic signal generation device 50, we will now describe Embodiment 1, which differs from the embodiments shown in Figures 2 and 3. Figures 5 and 6 are block diagrams showing the configuration of the scene detection unit 532 and the vibration acoustic signal generation unit 533 of the acoustic signal generation device 50 (Figure 2) of Embodiment 1. Figure 7 is an explanatory diagram showing the acoustic signal generation method in the acoustic signal generation device 50 of Embodiment 1.
[0042] In Example 1, the scene detection unit 532 includes a peripheral object extraction unit 532a and an acoustic object extraction unit 532b, as shown in Figure 5.
[0043] The surrounding object extraction unit 532a extracts vibrating objects that exist in the vicinity of the user avatar (within a threshold distance from the user avatar) in the content space. The surrounding object extraction unit 532a extracts vibrating objects by performing image analysis processing, 3D model analysis processing, etc., on the content space.
[0044] The vibrating objects vibrate in response to the applied vibration, causing the surrounding air to vibrate and generate sound. The vibrating objects are selected based on the user avatar's threshold distance, size, type, etc., so that objects that generate sound at a level perceptible to the user avatar are extracted. For example, in a racing game, vibrating objects include other vehicles Ve2, buildings St1, and interior components Iv within the vicinity of the user's vehicle Ve1 (e.g., within 10m) (see Figure 4). The extracted vibrating objects are stored in the data table of the storage unit 52 as needed.
[0045] The silent object extraction unit 532b extracts objects (silent objects) from among the vibrating objects extracted by the surrounding object extraction unit 532a, in which the content sound signal SS does not contain an acoustic signal generated by the vibrating object. In other words, in Embodiment 1, vibrating objects that generate sound are excluded from the vibrating objects because the sound generated by vibrations transmitted from the surroundings is considered negligible due to their own vibrations (sound source) and is masked by the sound they generate themselves.
[0046] For example, in the racing game shown in Figure 4 (although this may vary depending on the content creation), the acoustic signals of sounds generated by other vehicles Ve2 around the player's own vehicle Ve1 are included in the content acoustic signal SS and are therefore excluded from the objects that vibrate. However, the acoustic signals of sounds generated by buildings St1 and interior components Iv in the vehicle cabin are not included in the content acoustic signal SS and are therefore included in the objects that vibrate. The extracted non-sounding objects are stored in the data table of the storage unit 52 as needed. The vibration acoustic signal VSS is generated by the processing described below for non-sounding objects, but it is also possible to generate vibration acoustic signals VSS for objects that vibrate other than non-sounding objects.
[0047] The vibration acoustic signal generation unit 533 generates a vibration acoustic signal VSS based on information about silent objects acquired by the scene detection unit 532. The vibration acoustic signal VSS is distinct from the content acoustic signal SS originally included in the content itself, and is an acoustic signal for the sound generated by silent objects when vibrations generated in the content being played cause silent objects to vibrate. As shown in Figure 6, the vibration acoustic signal generation unit 533 includes a vibration acoustic acquisition unit 533a and a vibration acoustic adjustment unit 533b.
[0048] The vibration-acoustic acquisition unit 533a acquires basic vibration-acoustic data from the vibration-acoustic data table 522 according to the type of silent object extracted from the content by the silent object extraction unit 532b. The vibration-acoustic data table 522 is a data table in which basic vibration-acoustic data is stored for each type of silent object.
[0049] Figure 8 shows an example of a vibration-acoustic data table 522. As shown in Figure 8, the items in the vibration-acoustic data table 522 include "Data ID," "Category Type," "Object Type," and "Acoustic Signal File."
[0050] The "Data ID" is identification information used to identify each data record in the vibration-acoustic data table 522. The data for the Data ID is also the primary key of the data record in the vibration-acoustic data table 522. In other words, in the vibration-acoustic data table 522, a data record is created for each Data ID, and the data for each item associated with the Data ID is stored in that data record.
[0051] "Category Type" and "Object Type" are data items related to the category type and the type of object for an object that can be an acoustic object. In the racing game of this embodiment, for example, category types include surrounding vehicles, surrounding buildings, and interior components of a car. As for object types, for example, the surrounding vehicle category includes passenger cars, trucks, buses, etc., the surrounding building category includes buildings with windows, buildings with corrugated iron roofs, signs, etc., and the interior component category includes dashboards, center consoles, doors, etc.
[0052] An "acoustic signal file" is a data file of the acoustic signal corresponding to the sound generated by an object of a given object type when vibration is applied to that object. Strictly speaking, the sound generated by an object will differ depending on the type of vibration applied to it, but since the intensity of the sound component will be high around the object's resonant frequency, the same acoustic signal is used for each object type in this example for the sake of simplicity in processing. It is also possible to apply more detailed (for example, by vehicle type for passenger cars) or simpler object types.
[0053] The vibration-acoustic acquisition unit 533a acquires an acoustic signal data file corresponding to the type of silent object from the vibration-acoustic data table 522 as the basic vibration-acoustic data for each silent object in the content space. The vibration-acoustic data table 522 will be set with appropriate data based on experiments and analyses of objects in various content by the design and development team.
[0054] The vibration-acoustic adjustment unit 533b generates a vibration-acoustic signal VSS by performing adjustment processing on the basic vibration-acoustic signal of an acoustically neutral object acquired by the vibration-acoustic acquisition unit 533a. In other words, the basic vibration-acoustic signal is a vibration-acoustic signal VSS corresponding to the object type, and does not take into account the size of the object, the distance from the user avatar, etc.
[0055] Therefore, the vibration acoustic adjustment unit 533b adjusts the basic vibration acoustic signal according to the content vibration signal intensity in the content space, the size of the object, the distance from the user avatar, etc., to bring it closer to the sound heard by the user avatar. More specifically, the vibration acoustic adjustment unit 533b adjusts various acoustic parameter values such as the strength and length (duration) of the basic vibration acoustic signal according to the type of object, according to the content vibration signal intensity, the size of the silent object obtained from the content data, etc., and the distance from the user avatar.
[0056] For example, the intensity of the basic vibration acoustic signal is changed by multiplying the decoded acoustic digital signal value of the acoustic signal file according to the content vibration signal intensity, the size of the silent object, and the distance from the user avatar. Additionally, the length of the basic vibration acoustic signal is changed by changing the number of repetitions (decimal repetitions are allowed) of the acoustic signal file according to the size of the silent object and the distance from the user avatar. The above processing of the basic vibration acoustic signal can be performed by using a data table of intensity change values and length change values with the content vibration signal intensity, the size of the silent object, and the distance from the user avatar as parameters, or by using calculation formulas for intensity change values and length change values with the content vibration signal intensity, the size of the silent object, and the distance from the user avatar as parameters.
[0057] Furthermore, the vibration acoustic signal data of vibrating objects in the content space may be generated using AI (Artificial Intelligence). In the case of a racing game, for example, such an AI model would assume the vibration generation conditions of each object in a wide variety of driving situations, apply these vibration generation conditions to a simulation device to estimate the vibration sound generation state of each object in those conditions, and further estimate the vibration sound heard by the user avatar in this vibration sound generation state. Then, by generating a large amount of training data consisting of such vibration generation conditions and vibration sound heard by the user avatar, and training the AI model with this large amount of training data, it is possible to create an AI model that generates vibration acoustic signal data by training with training data containing many vibration acoustic signals (VSS) from a wide variety of vibrating objects.
[0058] Furthermore, in the above example, the vibration acoustic signal VSS was generated using a basic vibration acoustic signal corresponding to the object type of the unachoic object. However, a method of generating the vibration acoustic signal VSS using the content vibration signal VS is also applicable. For example, as shown in Figure 8B, the acoustic signal file of the vibration acoustic data table 522 shown in Figure 8A is changed to vibration acoustic conversion filter data. Figure 8B is a diagram showing a modified version of the vibration acoustic data table 522 in Figure 8A. The vibration acoustic conversion filter data is a filter (various parameter values that determine the filter characteristics) for generating an acoustic signal from the vibration signal when vibration is applied to an object of the target object type.
[0059] The vibration acoustic signal generation unit 533 extracts a filter from the vibration acoustic data table 522 corresponding to the object type of the silent object, and generates a basic vibration acoustic signal by filtering the content vibration signal VS with the said filter. Furthermore, the vibration acoustic signal generation unit 533 (vibration acoustic adjustment unit 533b) adjusts the generated basic vibration acoustic signal according to the size of the object, the distance from the user avatar, etc., to make it closer to the sound heard by the user avatar. In this example, since the content vibration signal VS is included in the basic vibration acoustic signal, the vibration acoustic adjustment unit 533b does not need to perform adjustment processing based on the content vibration signal VS.
[0060] The synthesis unit 534 synthesizes the content sound signal SS and the vibration sound signal VSS (signal addition process) to generate a sound signal (synthesized sound signal CSS) corresponding to the content.
[0061] The output unit 535 outputs the synthesized sound signal CSS, synthesized in the synthesis unit 534, to the speaker device 30 via the amplifier 21S.
[0062] As described above, in Example 1, since the source of the vibration acoustic signal VSS is narrowed down to anechoic objects and the vibration acoustic signal VSS is generated, it is possible to focus the processing on objects that have a significant influence as sources of the vibration acoustic signal VSS, thereby reducing the processing load. Furthermore, since the vibration acoustic signal VSS is generated using an acoustic signal file based on the type of anechoic object or a vibration acoustic conversion filter, it is possible to suppress the use of complex waveform analysis and calculation processing for acoustic and vibration signals, thereby reducing the processing load.
[0063] <3-2. Example of operation of the acoustic signal generator in Example 1> Figure 9 is a flowchart showing the acoustic signal generation process performed by the acoustic signal generation device 50 (controller 53) of Example 1. The operation shown in this flowchart is realized by a computer program (acoustic signal generation program 521) executed by the controller 53 (the computer constituting the controller 53).
[0064] The computer program that implements the acoustic signal generation method according to this embodiment in a computer device is installed in a computer device such as the acoustic signal generation device 50 to realize the various functions described above. Furthermore, such a computer program is provided to the computer device via a computer-readable non-volatile recording medium. For example, optical discs on which the computer program is recorded are distributed and sold, or computer programs stored on the hard disk of a server device are distributed and sold via a network environment. Also, the computer program that implements the calibration method according to this embodiment in a computer device may consist of only one program, or it may consist of multiple programs.
[0065] The process shown in Figure 9 is repeatedly executed when the content playback system 1 (content playback device 10, display device 20, speaker device 30, vibration device 40, acoustic signal generation device 50) is running and content is being played by the content playback device 10.
[0066] In step S101, the controller 53 (acquisition unit 531) acquires (receives) content data, content audio signal SS, content vibration signal VS, and content video signal PS from the content playback device 10, corresponding to the content to be played, and proceeds to step S102. The acquired content data, content audio signal SS, content vibration signal VS, and content video signal PS are stored in the data table of the storage unit 52 as needed.
[0067] In step S102, the controller 53 (scene detection unit 532) performs scene detection based on the content data acquired in step S101, extracts vibration-sensitive objects that exist around the user in the content space (where the vibration-acoustic signal VSS has a significant impact (the user avatar can hear the sound produced by the vibration-acoustic signal VSS)), and proceeds to step S103. The detected scene information SD and the extracted vibration-sensitive objects are stored in the data table of the storage unit 52 as needed.
[0068] In step S103, the controller 53 (scene detection unit 532) extracts objects from the vibration targets extracted in step S102 that do not contain an acoustic signal component of that object in the content acoustic signal SS (silent objects), and then proceeds to step S104. The extracted vibration targets (silent objects) are stored in the data table of the storage unit 52 as needed.
[0069] In step S104, the controller 53 (vibration acoustic signal generation unit 533) determines whether the silent object extracted in step S103 is registered in the vibration acoustic data table 522. If it is registered, the process proceeds to step S105; otherwise, the process shown in Figure 9 is terminated. The process shown in Figure 9 is also terminated if no silent object is extracted.
[0070] In other words, if a silent object is not registered in the vibration acoustic data table 522, the vibration acoustic signal VSS is not generated. The content acoustic signal SS is output to the speaker device 30 without the vibration acoustic signal VSS being synthesized. If a silent object extracted by the scene detection unit 532 is not registered in the vibration acoustic data table 522 (but a silent object has been extracted), the process may proceed to step S107 by selecting a predetermined vibration acoustic signal VSS for unregistered objects.
[0071] In step S105, the controller 53 (vibration acoustic signal generation unit 533) obtains vibration acoustic data (basic vibration acoustic signal) of the acoustically neutral object from the vibration acoustic data table 522, and then proceeds to step S106.
[0072] In step S106, the controller 53 (vibration acoustic signal generation unit 533) performs adjustment processing on the basic vibration acoustic signal acquired in step S105 to generate the vibration acoustic signal VSS, and then proceeds to step S107. More specifically, the controller 53 (vibration acoustic signal generation unit 533) performs processing such as adjusting various parameter values such as strength and length for the basic vibration acoustic signal based on the content data acquired in step S101.
[0073] In step S107, the controller 53 (synthesizer 534) synthesizes (synthesizes synchronized data) the content sound signal SS acquired in step S101 and the vibration sound signal VSS generated by the vibration sound signal generation unit 533 in step S106 to generate a synthesized sound signal CSS, and then proceeds to step S108.
[0074] In step S108, the controller 53 (output unit 535) outputs each content-related signal in a synchronized state to the subsequent output device, and the process shown in Figure 9 is completed. Specifically, the controller 53 (output unit 535) outputs the synthesized sound signal CSS synthesized in step S107 to the speaker device 30 via the amplifier 21S that performs power amplification, outputs the content vibration signal VS to the vibration device 40, and outputs the content video signal PS to the display device 20.
[0075] The audio signal generation for the speaker device 30 by the audio signal generation device 50 needs to be generated in real time during content playback. Therefore, the process shown in Figure 9 is executed repeatedly at high speed during content playback to prevent the sound from becoming unnatural.
[0076] <4. Example 2> Next, the acoustic signal generation device 50 of Example 2 will be described. The configuration of the acoustic signal generation device 50 of Example 2 (see Figure 10) is basically similar to the configuration of the acoustic signal generation device 50 of Example 1 described earlier (see Figures 2, 5, and 6). Also, the acoustic signal generation method in Example 2 (see Figure 11) is basically similar to the acoustic signal generation method in Example 1 described earlier (see Figure 7). Therefore, in the following, common components and processing steps are denoted by the same reference numerals, and their descriptions are omitted.
[0077] <4-1. Acoustic signal generation apparatus and acoustic signal generation method of Example 2> Figure 10 is a block diagram showing the configuration of the acoustic signal generation device 50 of Embodiment 2. Figure 11 is an explanatory diagram showing the acoustic signal generation method in the acoustic signal generation device 50 of Embodiment 2. As shown in Figure 8, the controller 53 of the acoustic signal generation device 50 of Embodiment 2 includes a vibration signal generation unit 536 as one of its functions. In addition, a weighted data table 523 is stored in the storage unit 52 of the acoustic signal generation device 50 of Embodiment 2.
[0078] The vibration signal generation unit 536 generates a content vibration signal VS based on the content sound signal SS acquired by the acquisition unit 531. In other words, vibrations that occur in real space are caused by force applied to an object, and these vibrations cause the surrounding air to vibrate, resulting in sound. For this reason, there is a relatively high correlation between sound and vibration in real space. Therefore, by generating a content vibration signal based on the audio signal in the content, it is possible to generate a vibration signal that is relatively consistent with the content.
[0079] This allows the content vibration signal VS to be generated based on the content acoustic signal SS, even if the content played by the content playback device 10 does not originally contain a content vibration signal VS. The content vibration signal VS generated by the vibration signal generation unit 536 is output to the vibration device 40 via the output unit 535.
[0080] More specifically, the vibration signal generation unit 536 extracts a vibration-corresponding frequency band signal from the content audio signal SS. The vibration-corresponding frequency band is a frequency range that is appropriate for vibrations experienced by the user, and is a frequency range in which human tactile sensation is highly sensitive to vibrations. Specifically, it is a frequency range of 30 to 130 Hz, which is the frequency range of so-called deep bass in an audio signal. The vibration signal generation unit 536 generates a content vibration signal VS based on the vibration-corresponding frequency band signal extracted from the content audio signal SS.
[0081] Furthermore, the vibration signal generation unit 536 acquires the scene information SD of the content being played, which is detected by the scene detection unit 532. Then, the vibration signal generation unit 536 performs adjustment processing on the content vibration signal VS based on each scene of the content. The vibration signal generation unit 536 obtains the parameter values used for this adjustment processing from the weighted data table 523.
[0082] The weighted data table 523 is a data table related to the adjustment process of acoustic signals and vibration signals corresponding to each scene of the content to be played, which were acquired as content data by the acquisition unit 531.
[0083] Figure 12 shows an example of a weighted data table 523. As shown in Figure 12, the items in the weighted data table 523 include "Data ID", "Scene Type", and "Parameter Value" ("For Acoustic Signal", "For Vibration Signal").
[0084] The "Data ID" is identification information used to identify data records in weighted data. The Data ID data is also the primary key of the data record in the weighted data table 523. In other words, in the weighted data table 523, a data record is created for each Data ID, and the data for each item associated with the Data ID is stored in that data record.
[0085] "Scene type" is a data item that indicates the type of each scene included in the content to be played. In the case of the racing game of this embodiment, for example, scene types include road driving, dirt driving, overtaking other vehicles, etc.
[0086] "Parameter values" are predefined parameter values applied to each signal during the adjustment process of acoustic and vibration signals. In this example, "parameter values" are signal level adjustment values, and include "for acoustic signals" and "for vibration signals" as needed. In the racing game of this embodiment, for example, the parameter value for the acoustic signal in a road driving scene is +3dB, and the parameter value for the vibration signal is +5dB. Also, the parameter value for the acoustic signal in an overtaking scene is +10dB, and the parameter value for the vibration signal is +5dB. "Parameter values" can be set in advance based on experiments, etc. In addition to signal levels, frequency characteristics and other parameters can also be used.
[0087] The vibration signal generation unit 536 uses the parameter values for vibration signals in the parameter weighting data table 523 to perform adjustment processing on the content vibration signal VS based on each scene (type) of content. For example, the vibration signal generation unit 536 performs vibration level adjustment processing (a process that multiplies the acoustic signal, or the vibration signal generated based on the acoustic signal, by a parameter value (level value)). As a result, the content vibration signal VS is generated based on parameter values that are appropriately defined in advance according to each scene of the content. Therefore, in Embodiment 2, even if the content itself does not contain a content vibration signal VS, a content vibration signal VS is generated according to each scene type of content and the content acoustic signal SS, and the vibration device 40 can be appropriately vibrated by the content vibration signal VS.
[0088] Furthermore, the weighted data table 523 is also used in the vibration acoustic signal generation process by the vibration acoustic signal generation unit 533 (vibration acoustic adjustment unit 533b).
[0089] The vibration-acoustic adjustment unit 533b performs adjustment processing on the basic vibration-acoustic signal (acoustic signal from the acoustic signal data file) obtained from the vibration-acoustic data table 522, similar to the vibration-acoustic adjustment unit 533b of Embodiment 1 shown in Figures 6 and 7. In the vibration-acoustic adjustment unit 533b of Embodiment 2, further adjustment processing is performed on the basic vibration-acoustic signal according to the content scene.
[0090] In other words, the basic vibration acoustic signal is a vibration acoustic signal VSS corresponding to the object type, and does not take into account the size of the object, the distance from the user avatar, etc. Therefore, the vibration acoustic adjustment unit 533b adjusts the basic vibration acoustic signal according to the content vibration signal intensity, the size of the object, the distance from the user avatar, etc., to make it closer to the sound heard by the user avatar. In addition, the basic vibration acoustic signal changes depending on the environment such as weather (changes in sound due to changes in the characteristics of acoustic objects and changes in the acoustic transmission path characteristics) and is transmitted to the user avatar. Therefore, in Example 2, the basic vibration acoustic signal is adjusted according to the content scene to make it closer to the sound heard by the user avatar. More specifically, the vibration acoustic adjustment unit 533b acquires the content vibration signal VS generated by the vibration signal generation unit 536 from the vibration signal generation unit 536. Then, the vibration acoustic adjustment unit 533b applies the processing shown in Example 1 (adjustment of the basic vibration acoustic signal according to the signal level of the content vibration signal VS) to the acquired content vibration signal VS.
[0091] Furthermore, the vibration-acoustic adjustment unit 533b acquires the scene information SD of the content being played, which is detected by the scene detection unit 532. Based on the acquired scene information SD, the vibration-acoustic adjustment unit 533b searches the weighted data table 523 and extracts the vibration signal parameter values corresponding to the scene. The vibration-acoustic adjustment unit 533b then adjusts the basic vibration-acoustic signal using these vibration signal parameter values to bring it closer to the sound heard by the user avatar. In other words, the vibration-acoustic adjustment unit 533b adjusts the basic vibration-acoustic signal according to the content vibration signal intensity in the content space, the size of the object, the distance from the user avatar, the surrounding environment, etc., to bring it closer to the sound heard by the user avatar.
[0092] <4-2. Example of operation of the acoustic signal generator in Example 2> Figure 13 is a flowchart showing the acoustic signal generation process performed by the acoustic signal generation device 50 (controller 53) of Example 2. The operation shown in this flowchart is realized by a computer program (acoustic signal generation program 521) executed by the controller 53 (the computer constituting the controller 53). In the following, the same reference numerals are used for processing steps that are common to the operation flow described in Example 1 (see Figure 9), and detailed explanations are omitted.
[0093] Furthermore, the content used in the description of Example 2 is assumed to be configured not to include the content vibration signal VS.
[0094] In step S201, the controller 53 (acquisition unit 531) acquires (receives) content data, content audio signal SS, and content video signal PS corresponding to the content to be played back from the content playback device 10, and then proceeds to step S202. The acquired content data, content audio signal SS, and content video signal PS are stored in the data table of the storage unit 52 as needed.
[0095] In step S202, the controller 53 (vibration signal generation unit 536) generates a content vibration signal VS based on the content sound signal SS acquired in step S101, and proceeds to step S102. The generated content vibration signal VS is stored in the data table of the storage unit 52 as needed.
[0096] In step S102, the controller 53 (scene detection unit 532) performs scene detection based on the content data acquired in step S101, extracts vibrating objects present around the user in the content space, and proceeds to step S203.
[0097] In step S203, the controller 53 (vibration signal generation unit 536) performs adjustment processing on the content vibration signal VS generated in step S202 based on the scene information SD detected in step S102 (parameter values for vibration signals corresponding to the detected scene information (type) in the weighted data table 523), and then proceeds to step S103.
[0098] In step S103, the controller 53 (scene detection unit 532) extracts silent objects from the vibrating objects extracted in step S102, and proceeds to step S104.
[0099] In steps S104 and S105, the controller 53 (vibration acoustic signal generation unit 533) obtains vibration acoustic data (basic vibration acoustic signal) of an acoustic object from the vibration acoustic data table 522 if the acoustic object is pre-registered in the vibration acoustic data table 522, and then proceeds to step S204. In step S104, if the controller 53 determines that the acoustic object is not registered in the vibration acoustic data table 522, it terminates the process shown in Figure 13.
[0100] In step S204, the controller 53 (vibration acoustic signal generation unit 533) performs adjustment processing on the basic vibration acoustic signal acquired in step S105 to generate a vibration acoustic signal VSS, and then proceeds to step S107. More specifically, the controller 53 (vibration acoustic signal generation unit 533) performs adjustment processing on the basic vibration acoustic signal based on the scene information SD (parameter value for vibration signal corresponding to the detected scene information (type) in the weighting data table 523) detected in step S102, and also based on the content vibration signal VS (signal level) generated in step S202 (for example, multiplying the basic vibration acoustic signal by the weighting value and a value obtained by multiplying the signal level value of the content vibration signal VS by an appropriate coefficient (determined based on experiments, etc.)).
[0101] In step S107, the controller 53 (synthesizer 534) synthesizes the content sound signal SS acquired in step S201 and the vibration sound signal VSS generated in step S204 (synthesizes synchronized data) to generate a synthesized sound signal CSS, and then proceeds to step S108.
[0102] In step S108, the controller 53 (output unit 535) outputs each content-related signal in a synchronized state to the subsequent output device, and the process shown in Figure 13 is completed. Specifically, the controller 53 (output unit 535) outputs the synthesized sound signal CSS synthesized in step S107 to the speaker device 30 via the amplifier 21S that performs power amplification, outputs the content vibration signal VS to the vibration device 40, and outputs the content video signal PS to the display device 20.
[0103] <5. Things to keep in mind> The various technical features disclosed as embodiments herein can be modified in various ways without departing from the spirit of the technical creation. That is, the above embodiments are illustrative in all respects and not restrictive. The technical scope of the present invention is indicated by the claims rather than by the above descriptions of embodiments, and includes all modifications that fall within the meaning and scope equivalent to the claims. Furthermore, the multiple embodiments shown herein may be combined as appropriate to the extent possible.
[0104] Furthermore, although the above embodiment explains that various functions are implemented in software through CPU arithmetic processing according to a program, at least some of these functions may be implemented by electrical hardware resources. These hardware resources may be implemented entirely or partially by, for example, ASICs (Application Specific Integrated Circuits) or FPGAs (Field Programmable Gate Arrays). Conversely, at least some of the functions implemented by hardware resources may be implemented in software.
[0105] Furthermore, the system may include a computer program that enables a processor (computer) to implement at least some of the functions of each component of the content playback system 1 (content playback device 10, sound signal generation device 50). Such a computer program can be stored on a computer-readable non-volatile recording medium (for example, in addition to the non-volatile memory mentioned above, it can be provided (sold, etc.) on an optical recording medium (e.g., optical disc), a magneto-optical recording medium (e.g., magneto-optical disc), a USB memory stick, or an SD card, etc.), and it can also be provided from a server device via a communication line such as the Internet, a method known as download. [Explanation of Symbols]
[0106] 1. Content Playback System 10. Content playback device 20 Display device 30 Speaker System 40 Vibration device 50 Acoustic signal generation device 51 Communications Department 52 Storage section 53 Controllers 521 Acoustic Signal Generation Program 522 Vibration Acoustics Data Table 523 Weighted Data Table 531 Acquisition Department 532 Scene detection unit 532a Peripheral object extraction unit 532b No-acoustic object extraction part 533 Vibration Acoustic Signal Generation Unit 533a Vibroacoustic acquisition unit 533b Vibroacoustic adjustment section 534 Synthesis section 535 Output section 536 Vibration signal generation unit
Claims
1. An acoustic signal generating device that generates an acoustic signal corresponding to the content, A content vibration signal corresponding to the aforementioned content is determined, Based on the content vibration signal, a vibration acoustic signal is generated for the sound generated by the vibration in the content. The audio signal is generated by combining the content audio signal included in the content with the vibration audio signal. Acoustic signal generation device.
2. The content includes the content vibration signal, The acoustic signal generating apparatus according to claim 1.
3. The content vibration signal is generated based on the content sound signal. The acoustic signal generating apparatus according to claim 1.
4. The content vibration signal is generated based on the scene of the content. The acoustic signal generating apparatus according to claim 1.
5. A method for generating an audio signal that generates an audio signal corresponding to the content, A content vibration signal corresponding to the aforementioned content is determined, Based on the content vibration signal, a vibration acoustic signal is generated for the sound generated by the vibration in the content. The audio signal is generated by combining the content audio signal included in the content with the vibration audio signal. The method by which the controller generates acoustic signals.
6. A sound signal generation program executed by a controller that generates sound signals according to the content, A content vibration signal corresponding to the aforementioned content is determined, Based on the content vibration signal, a vibration acoustic signal is generated for the sound generated by the vibration in the content. The computer is instructed to perform a process of generating the sound signal by combining the content sound signal included in the content with the vibration sound signal. A program for generating acoustic signals.
7. A content playback system that outputs video, audio, and vibration corresponding to the content, It comprises a content playback device, an acoustic signal generation device, a display device, a speaker device, and a vibration device. The aforementioned content playback device is The content audio signal and content video signal included in the aforementioned content are output. The aforementioned acoustic signal generating device is A content vibration signal is generated based on the content audio signal output from the content playback device. Based on the content vibration signal, a vibration acoustic signal is generated for the sound generated by the vibration in the content. The content audio signal and the vibration audio signal are combined to generate a synthesized audio signal. The aforementioned display device is Output video based on the aforementioned content video signal, The speaker device is Outputting sound based on the aforementioned synthesized sound signal, The aforementioned vibration device is Outputting vibrations based on the aforementioned content vibration signal, Content playback system.
8. A content playback system that outputs video, audio, and vibration corresponding to the content, It comprises a content playback device, an acoustic signal generation device, a display device, a speaker device, and a vibration device. The aforementioned content playback device is The content audio signal, content vibration signal, and content video signal included in the aforementioned content are output. The aforementioned acoustic signal generating device is Based on the content vibration signal output from the content playback device, a vibration acoustic signal is generated for the sound generated by the vibration in the content. The content audio signal and the vibration audio signal are combined to generate a synthesized audio signal. The aforementioned display device is Output video based on the aforementioned content video signal, The speaker device is Outputting sound based on the aforementioned synthesized sound signal, The aforementioned vibration device is Outputting vibrations based on the aforementioned content vibration signal, Content playback system.