Signal processing device, signal processing method, and program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-03-16
- Publication Date
- 2026-03-10
AI Technical Summary
Existing augmented reality technologies cannot accurately simulate sound changes caused by virtual objects interacting with real objects, as they only account for sound diffraction when a virtual sound source is hidden behind a real object, failing to replicate the sound as if the virtual object were present in the real world.
A signal processing device that acquires sound generation characteristics by vibrating real objects in the augmented reality environment, applies these characteristics to virtual sound sources, and reproduces the sound through headphones or speakers, mimicking the interaction of virtual objects with real-world objects.
The device effectively reproduces sound as if virtual objects were present in the real world by applying sound generation characteristics due to object vibrations, enhancing the realism of augmented reality experiences.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a signal processing device, a control method for a signal processing device, and a program. [Background technology]
[0002] In a system that displays a virtual object superimposed on the real world, known as an augmented reality (AR) display, there is a technology that changes the sound emitted by the virtual object to reflect the appearance of the real world. Patent Document 1 discloses a technology that adjusts a synthetic voice in an augmented reality display system so that when the position of a virtual sound source is hidden behind a real object as seen by the user, the synthetic voice is heard going around the real object. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2021-175043 A Summary of the Invention [Problem to be solved by the invention]
[0004] When a virtual object is superimposed on the real world and displayed in augmented reality, if a sound appropriate to the real object that the virtual object touches can be produced, it is possible to produce an effect as if the virtual object is actually present there. The technology of Patent Document 1 can only calculate and reflect the diffraction of sound by the real object only when there is a real object between the virtual sound source and the viewing position, and cannot express the change in sound caused by the virtual sound source touching the real object. Therefore, the object of the present invention is to be able to reproduce a sound that makes it seem as if the virtual sound source exists in the real world. [Means for solving the problem]
[0005] The signal processing device according to the present invention is characterized by having a characteristic acquisition means for acquiring sound characteristics due to vibration corresponding to an object in the real world that a virtual object touches in a virtual reality display in which an image of the virtual object is superimposed on the real world, a characteristic application means for applying the acquired sound characteristics due to vibration to a sound source signal of the virtual object, and a first playback means for playing back a sound corresponding to the sound source signal to which the sound characteristics due to vibration have been applied. Effect of the Invention
[0006] According to the present invention, it is possible to reproduce sounds that make it seem as if a virtual sound source exists in the real world. [Brief description of the drawings]
[0007] [Figure 1] FIG. 2 is a diagram illustrating an example of a functional configuration of a signal processing device. [Diagram 2] 1A and 1B are diagrams for explaining acquisition of sound generation characteristics by vibration. [Diagram 3] FIG. 4 is a diagram showing an example of frequency characteristics of sound generated by vibration. [Figure 4] FIG. 2 is a diagram illustrating an example of a hardware configuration of a signal processing device. [Diagram 5] 1 is a flowchart illustrating an example of a sound reproduction process in a signal processing device. [Figure 6] 13 is a flowchart illustrating an example of a sound generation characteristic acquisition process. [Figure 7] FIG. 2 is a diagram illustrating an example of a functional configuration of a signal processing device. [Figure 8] FIG. 11 is a diagram showing an example of a data configuration of sound generation characteristic information by vibration. [Figure 9] 1 is a flowchart showing an example of video and audio reproduction processing in a signal processing device. [Figure 10] 13 is a flowchart showing an example of a process for acquiring a pronunciation characteristic from a pronunciation characteristic DB. [Figure 11] 13 is a flowchart illustrating an example of a sound generation characteristic acquisition process. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0008] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Note that the embodiment described below does not limit the present invention, and not all of the combinations of features described in the present embodiment are necessarily essential configurations. Note that the same components will be described with the same reference numerals.
[0009] (Embodiment 1) In this embodiment, an example will be described in which sound characteristics of a real object (real world object) that a virtual sound source (virtual object) touches in a virtual reality display in which an image of the virtual sound source (virtual object) is superimposed on the real world and applied to a sound source signal is obtained. In this embodiment, for simplicity of explanation, it will be described as being possible to obtain a sound source signal that is a sound associated with a virtual sound source (virtual object).
[0010] FIG. 1 is a diagram illustrating an example of a functional configuration of a signal processing device according to the first embodiment. The transducer 101 is installed on a real object and vibrates the installed real object according to an input measurement signal (for example, a measurement acoustic signal). This vibration generates a sound that takes into account vibration-based sound characteristics determined by various attributes of the object, such as its size, area, material, structure, etc. The microphone 102 acquires (collects) the sound generated by the real object vibrated by the transducer 101.
[0011] The sound characteristic acquisition unit 103 acquires characteristics of a sound generated by vibrating an object. Hereinafter, the characteristics of a sound generated by vibrating an object are also referred to as sound characteristics due to vibration. The sound characteristic acquisition unit 103 vibrates a real object from which sound characteristics are to be acquired by outputting a measurement signal to the transducer 101. The sound characteristic acquisition unit 103 also acquires the generation characteristics due to vibration by analyzing the sound acquired by the microphone 102 based on the measurement signal. The sound characteristic acquisition unit 103 outputs the acquired generation characteristics due to vibration to the sound characteristic application unit 106.
[0012] An image of acquiring sound characteristics due to vibration is shown in Fig. 2. In Fig. 2, the same components as those shown in Fig. 1 are given the same reference numerals. In the example shown in Fig. 2, 201 is a desk from which the sound characteristics due to vibration are acquired. The sound characteristics acquisition unit 103 outputs a measurement signal to a vibrator 101 placed on the desk 201 to vibrate the desk 201, and acquires the sound characteristics due to the vibration of the desk 201 by analyzing the sound acquired by the microphone 102 at that time. In this manner, the sound characteristics due to the vibration of a real object can be acquired.
[0013] Here, the sound characteristics due to vibration will be explained. Figures 3(a) and 3(b) are diagrams showing an example of the frequency characteristics of a sound generated by actually vibrating an object in the real world. In Figures 3(a) and 3(b), the vertical axis is sound pressure, and the horizontal axis is frequency. Figure 3(a) shows the frequency characteristics of a sound generated by vibrating a plywood desk that is commonly used in offices. It can be seen that the sound characteristics generated by vibrating this desk have a relatively flat frequency characteristic up to about 3 kHz. Figure 3(b) shows the frequency characteristics of a sound generated by vibrating a small cardboard box with a bottom of about B4 size and a height of about 10 cm. The sound characteristics generated by vibrating this cardboard box show a characteristic with a large peak in the vicinity of 3 kHz to 4 kHz, and there are almost no components in other frequency bands. In other words, the sound generated from the vibrated cardboard box becomes a harsh and rustling sound. In this way, the sound quality changes depending on the attributes of the object to be vibrated. It can also be seen that the sound quality becomes as if it expresses the attributes of the object to be vibrated.
[0014] 1, the sound source signal acquisition unit 105 acquires a sound source signal 104 such as a piece of music, and outputs it to the sound characteristic application unit 106. The sound source signal acquisition unit 105 appropriately reads out the sound source signal 104 from a storage unit (not shown) or the like, and outputs it to the sound characteristic application unit 106. Note that the sound source signal 104 may include music related to the virtual sound source (virtual object), such as background music, in addition to the sound emitted by the virtual sound source (virtual object).
[0015] The sound characteristic application unit 106 applies the sound characteristic due to vibration output from the sound characteristic acquisition unit 103 to the sound source signal input from the sound source signal acquisition unit 105. The sound source signal to which the sound characteristic due to vibration has been applied is output to the sound reproduction unit 107. The sound reproduction unit 107 appropriately amplifies the sound source signal to which the sound characteristic due to vibration has been applied, input from the sound characteristic application unit 106, and outputs it to an audio output device such as headphones 108 or a speaker. The headphones 108 are worn on the head of a listener, and convert the sound source signal output from the sound reproduction unit 107 into sound and output it to both ears of the listener.
[0016] 4 is a diagram showing an example of the hardware configuration of the signal processing device in this embodiment. The signal processing device in this embodiment has an input / output unit 401, a CPU 402, a RAM 403, an external storage unit 404, an operation unit 405, a display unit 406, a ROM 407, a communication IF unit 408, and a bus 409. The input / output unit 401, the CPU 402, the RAM 403, the external storage unit 404, the operation unit 405, the display unit 406, the ROM 407, and the communication IF unit 408 are connected via the bus 409 so as to be able to communicate with each other.
[0017] The input / output unit 401 receives input of a microphone signal and a sound source signal from the outside, and sends them to other components via a bus 409 according to appropriate instructions from a CPU 402. The input / output unit 401 also processes sound source signals stored in a RAM 403 or an external storage unit 404 according to appropriate instructions from the CPU 402, and sends them via the bus 409 to external audio output devices such as headphones and speakers.
[0018] A CPU (Central Processing Unit) 402 performs overall control of each component of the signal processing device. The CPU 402 sends control signals to other components via a bus 409 to control them according to a program, and performs various calculations. In this embodiment, the CPU 402 executes each function described in FIG. 1 by executing processing according to a program stored in a ROM 407 or an external storage unit 404.
[0019] A RAM (Random Access Memory) 403 temporarily stores a part of a program being executed, associated data, and calculation results of the CPU 402. The CPU 402 loads necessary programs and data into the RAM 403 and executes the programs by reading and writing as necessary.
[0020] The external storage unit 404 stores the program itself and data to be accumulated for a long period of time. The external storage unit 404 is, for example, an HDD (hard disk drive) or an SSD (solid state drive). The operation unit 405 accepts various instructions and operations from the user, converts them into control signals, and transmits them to the CPU 402 via the bus 409. The CPU 402 controls the running program and issues control instructions for other components in accordance with the control signals.
[0021] A display unit 406 displays to the user the status of a running program and the output of the program. A ROM (Read Only Memory) 407 stores fixed programs and fixed parameters, such as programs for starting and shutting down the hardware device and programs for controlling basic input and output. A communication IF unit 408 can input and output data to and from a communication network such as the Internet.
[0022] The sound reproduction process by the signal processing device in this embodiment will be described below. Fig. 5 is a flowchart showing an example of the sound reproduction process by the signal processing device in this embodiment. In step S501, a vibrator 101 is placed on a real object that is to generate sound, and a microphone 102 is placed in a position near the real object that can pick up the sound generated by the real object. Here, the real object that generates sound is an object in the real world that is touched by a virtual sound source (virtual object) in virtual reality display. The microphone 102 is placed, for example, at a position about 30 cm to 50 cm away from the surface on which the vibrator 101 is placed.
[0023] In step S502, the sound characteristic acquisition unit 103 outputs a measurement signal to the vibrator 101 installed on the real object, and the microphone 102 picks up the sound generated by the vibration of the vibrator 101. The sound characteristic acquisition unit 103 then analyzes the sound collected by the microphone 102 to acquire the sound characteristic generated by the vibration of the real object. The acquired sound characteristic generated by the vibration is output to the sound characteristic application unit 106. The sound characteristic acquisition unit 103 outputs, for example, a TSP (Time Stretched Pulse) signal or pink noise as a measurement signal to the vibrator 101, and analyzes a signal generated by the sound generated from the real object and collected by the microphone 102. As a result, an impulse response is obtained. Furthermore, by removing aliasing distortion and adjusting the time length to an appropriate length, a FIR (Finite Impulse Response) filter coefficient is obtained as the sound characteristic generated by the vibration. Details of the sound characteristic acquisition process performed in step S502 will be described later with reference to FIG. 6.
[0024] In step S503, the sound source signal acquisition unit 105 acquires the sound source signal for the next processing unit time from the sound source signal 104. The sound source signal acquisition unit 105 outputs the sound source signal for the acquired processing unit time to the sound characteristic application unit 106. The processing unit time is a time corresponding to the time for performing a series of processes from step S503 to S505, for example.
[0025] In step S504, the sound characteristic application unit 106 applies the sound characteristic due to vibration acquired in step S502 to the sound source signal acquired in step S503. The sound characteristic application unit 106 performs a filter process on the sound source signal acquired in step S503 using an FIR filter having the filter coefficient obtained in step S502. The sound source signal after the application of the sound characteristic due to vibration is output to the sound reproduction unit 107.
[0026] In step S505, the sound reproducing unit 107 appropriately adjusts and amplifies the sound source signal to which the sound generation characteristics due to vibration have been applied in step S504, and outputs the sound to the headphones 108. The headphones 108 convert the sound source signal input from the sound reproducing unit 107 into sound and delivers the sound to both ears of the listener. This allows the listener to hear a sound to which the sound generation characteristics due to vibration of a real object that the virtual object touches in the virtual reality display have been applied.
[0027] In step S506, the CPU 402 of the signal processing device determines whether or not an instruction to end the sound reproduction process has been issued by a user operation of the operation unit 405. If the signal processing device determines that there has been no instruction to end the sound reproduction process (NO in step S506), the sound reproduction process continues, and the process returns to step S503, where the sound reproduction process is executed on the sound source signal for the next processing unit time. If the signal processing device determines that there has been an instruction to end the sound reproduction process (YES in step S506), the sound reproduction process is terminated.
[0028] Fig. 6 is a flowchart showing an example of the pronunciation characteristic acquisition process in step S502 in Fig. 5. All the processes in the pronunciation characteristic acquisition process shown in Fig. 6 are executed in the pronunciation characteristic acquisition unit 103.
[0029] In step S601, the sound characteristic acquisition unit 103 outputs a measurement signal (measurement acoustic signal) such as a TSP signal to the transducer 101 installed on a real object. For example, a signal in an audible frequency band of about 5 Hz to 20 kHz is used as the measurement signal. Note that this is just one example, and a signal in another frequency band, for example, a signal including a band other than the audible frequency band, may also be used as the measurement signal.
[0030] In step S602, the sound generation characteristic acquisition unit 103 acquires a signal obtained by collecting a sound generated by vibration of the vibrator 101 with the microphone 102. In step S603, the sound characteristic acquisition unit 103 performs an analysis process on the picked-up signal acquired by the microphone 102 in step S602 based on the measurement signal output to the transducer 101 in step S601, and calculates an impulse response. This process is a process that is usually performed in the field of acoustic analysis and is well known, so a description thereof will be omitted.
[0031] In step S604, the sound characteristic acquisition unit 103 performs processing to remove high frequencies outside the audible frequency band using a low-pass filter to remove aliasing distortion from the impulse response obtained in step S603. Furthermore, the sound characteristic acquisition unit 103 creates FIR filter coefficients with a default coefficient length (number of taps) by performing processing such as arranging the impulse response processed by the low-pass filter to a predetermined time length. Note that the sound characteristic acquisition unit 103 usually adjusts the length of the coefficients to the Nth power of 2, taking into account processing when applying the FIR filter, etc.
[0032] In step S605, the sound characteristic acquisition unit 103 outputs the FIR filter coefficients created in step S604 to the sound characteristic application unit 106 as sound characteristics due to vibration of the target real object. When the processing of step S605 is completed, the sound characteristic acquisition process is terminated and the process returns to the sound reproduction process shown in Fig. 5. The FIR filter obtained in this manner is applied to the sound source signal in step S504 in Fig. 5, whereby the sound characteristics due to vibration are applied and a change in sound due to the sound source touching the real object is expressed.
[0033] According to this embodiment, by acquiring the sound characteristics due to the vibration of a real object touched by a virtual sound source (virtual object) in a virtual reality display and applying them to a sound source signal, it is possible to express a change in sound caused by the sound source touching a real object. This makes it possible to generate a sound as if the virtual sound source (virtual object) displayed in virtual reality were touching a real object, and to reproduce a sound as if the virtual sound source (virtual object) exists in the real world.
[0034] In the above example, an example of acquiring FIR filter coefficients in the sound generation characteristic acquisition process has been described. However, the present invention is not limited to this example, and other methods may be used, such as designing an IIR (Infinite Impulse Response) filter that approximates the frequency characteristics of the sound generated by vibration and acquiring the IIR filter coefficients.
[0035] (Embodiment 2) In the first embodiment, an example is described in which characteristics of a sound emitted when a real object is vibrated are acquired and the acquired characteristics are applied to a sound source signal. In the second embodiment, an example is described in which, when a virtual object having a three-dimensional shape serving as a virtual sound source is displayed in augmented reality together with a real-world image, sound characteristics suitable for an object at the display position of the virtual object are acquired from a database and applied to a sound source signal. Note that a description of the same configuration and processing as in the first embodiment is omitted. In the following, a virtual object serving as a virtual sound source is also referred to as a virtual sound source object.
[0036] Fig. 7 is a diagram showing an example of the functional configuration of a signal processing device in embodiment 2. In Fig. 7, components having the same functions as those shown in Fig. 1 are given the same reference numerals, and duplicated explanations will be omitted.
[0037] The camera 701 is installed on the surface of the VR goggles 709. The camera 701 captures an image of the real world in the front direction of the VR goggles, and outputs an electrical signal related to the captured image captured by a built-in sensor to an image acquisition unit 702. The image acquisition unit 702 acquires an image signal of the real world by performing development processing on the electrical signal received from the camera 701, and outputs the image signal to a display position determination unit 705.
[0038] Distance sensor 703 is installed on the surface of VR goggles 709. Distance sensor 703 scans the state of the real world in the front direction of the VR goggles within a range equivalent to the angle of view of camera 701, acquires distance information at each point in the scan range, and outputs the acquired distance information to distance map generation unit 704. Distance map generation unit 704 generates a distance map of a range equivalent to the angle of view of camera 701 based on the distance information received from distance sensor 703, and outputs the distance map to display position determination unit 705.
[0039] The display position determination unit 705 determines a display position at which the virtual sound source object is to be synthesized and displayed in the image of the real world captured by the camera 701. The display position determination unit 705 analyzes the state of the real world based on the image signal output from the image acquisition unit 702 and the distance map output from the distance map generation unit 704. Then, the display position determination unit 705 searches for a position in the real world where the virtual sound source object is likely to be naturally synthesized and displayed, and determines it as the display position. For example, the display position determination unit 705 searches for a contact surface (e.g., a horizontal surface) in the real world where the virtual sound source object is likely to be synthesized and displayed, and determines it as the display position of the virtual sound source object. Such processing is a processing that is usually performed in the field of augmented reality image generation and is well known, so a description will be omitted. The display position determination unit 705 outputs the image of the real world and the determined display position of the virtual sound source object to the image synthesis unit 707 and the real object analysis unit 710.
[0040] The video synthesis unit 707 renders the virtual sound source object as a 3D object based on the virtual sound source 3D model 706, and synthesizes it at the display position of the real world video output from the display position determination unit 705. The virtual sound source 3D model 706 stores information such as the color and shape of the virtual sound source object. The video synthesis unit 707 outputs the video generated by synthesizing the virtual sound source object to the video playback unit 708. The video playback unit 708 converts the video output from the video synthesis unit 707 into a format to be displayed by the VR goggles 709, and outputs it to the VR goggles 709. The VR goggles 709 are worn by the viewer, and display the video received from the video playback unit 708. This allows the viewer to see the video in which the virtual object serving as the virtual sound source is synthesized into the real world. In this embodiment, the VR goggles 709 that perform display using a video see-through method will be described as an example, but they may also perform display using an optical see-through method.
[0041] The real object analysis unit 710 analyzes what real object is at the display position of the virtual sound source object based on the image of the real world output from the display position determination unit 705 and the display position of the virtual sound source object. The real object analysis unit 710 can obtain the model number and specifications of the object by, for example, extracting a part of the object from the image of the real world and performing an image search on a cloud server on a network. Such image analysis processing is a technology that is generally performed using a smartphone or the like and is well known, so a description thereof will be omitted. The real object analysis unit 710 outputs the analysis result of the real object, for example, the model number and specifications of the real object, to the sound characteristic acquisition unit 712.
[0042] The pronunciation characteristic database (pronunciation characteristic DB) 711 stores the pronunciation characteristics due to vibration of a plurality of real objects used for displaying a virtual sound source object in contact. The pronunciation characteristic DB 711 is an example of a characteristic storage means. The pronunciation characteristics due to vibration of various real objects are stored in the pronunciation characteristic DB 711 together with attribute information of the objects. Each piece of data stored in the pronunciation characteristic DB 711 is created by measuring the pronunciation characteristics due to vibration in advance for various real objects and acquiring various attributes from a specification table, actual measurement, etc. Hereinafter, the data stored in this pronunciation characteristic DB 711 is also referred to as pronunciation characteristic information due to vibration. Details of the pronunciation characteristic information due to vibration will be described later.
[0043] The sound characteristic acquisition unit 712 searches the sound characteristic DB 711 based on the analysis result of the real object output from the real object analysis unit 710, and acquires sound characteristics due to vibration corresponding to the real object at the display position of the virtual sound source object. The sound characteristic acquisition unit 712 searches the sound characteristic DB 711 by attribute information (various attributes) of the real object based on the analysis result of the real object output from the real object analysis unit 710. As a result, the sound characteristic acquisition unit 712 acquires sound characteristics due to vibration of an object having the highest correlation (the closest attribute) with the attribute information of the real object at the display position of the virtual sound source object from the data in the sound characteristic DB 711. The sound characteristic acquisition unit 712 outputs the acquired generation characteristics due to vibration to the sound characteristic application unit 106.
[0044] In this embodiment, an example of the hardware configuration of the signal processing device is omitted since it is the same as that of the first embodiment.
[0045] The data structure of the sound characteristic information by vibration stored in the sound characteristic DB 711 will be described with reference to Fig. 8. Fig. 8 is a diagram showing an example of the data structure of the sound characteristic information by vibration. As shown in Fig. 8, the sound characteristic information by vibration includes a sound characteristic ID 801, and attribute information such as object type 802, material 803, size 804, and internal space volume 805, and sound characteristic by vibration 806. The sound characteristic DB 711 stores a plurality of pieces of sound characteristic information by vibration as shown in Fig. 8.
[0046] The sound characteristic ID 801 is an ID number for identifying the sound characteristic information by this vibration, and is uniquely assigned to each piece of sound characteristic information by vibration. The type of object 802 indicates the type of the real object from which this sound characteristic information by vibration was obtained. The type of object 802 is information indicating, for example, a desk, a shelf, a box, etc. The material 803 indicates the material that forms the real object from which this sound characteristic information by vibration was obtained. The material 803 is information indicating, for example, steel, plywood, cedar board, plastic, corrugated cardboard, etc. The size 804 indicates the size (for example, width, depth, height) of the real object from which this sound characteristic information by vibration was obtained. The internal space volume 805 indicates the volume of the space inside the real object from which this sound characteristic information by vibration was obtained. For an object without an internal space, the value of the internal space volume 805 is 0. If an object has an internal space, the internal reflection of the sound due to this space changes when it is vibrated, so this attribute is given to the object.
[0047] The components of the sound characteristic information by vibration (type 802 of object, material 803, size 804, and internal space volume 805) described above each have a large effect on the sound characteristic by vibration. Therefore, it can be considered that the higher the similarity in these elements between two objects, the closer the sound characteristic by vibration becomes. Therefore, the sound characteristic acquisition unit 712 can acquire sound characteristics suitable for a real object by searching for data with high similarity of each component among the sound characteristic information by vibration stored in the sound characteristic DB 711 using these components as keys. Here, an example is shown in which the type 802, material 803, size 804, and internal space volume 805 of an object are included as attribute information, but the data included as attribute information is not limited to these and may include other attributes.
[0048] The sound characteristics due to vibration 806 are characteristics of a sound generated by vibrating an object (sound characteristics due to vibration) acquired by a method similar to the sound characteristics acquisition process shown in Fig. 6 in the first embodiment. The sound characteristics due to vibration 806 are stored in a state that can be applied to the sound source signal as it is, such as FIR filter coefficients or IIR filter coefficients.
[0049] The video / audio reproduction process by the signal processing device in this embodiment will be described below. Fig. 9 is a flowchart showing an example of the video / audio reproduction process by the signal processing device in this embodiment. In step S901, the image acquisition unit 702 acquires an image of the real world captured by the camera 701. The image acquisition unit 702 outputs the acquired image of the real world to the display position determination unit 705.
[0050] In step S902, distance map generation unit 704 generates a distance map of a range corresponding to the angle of view of camera 701. First, distance sensor 703 scans the real world in a range corresponding to the angle of view of camera 701, acquires distances at each point in the scan range, and outputs the acquired distances to distance map generation unit 704. Based on the distance information received from distance sensor 703, distance map generation unit 704 generates a distance map of a range corresponding to the angle of view of camera 701. Distance map generation unit 704 outputs the generated distance map to display position determination unit 705.
[0051] In step S903, the display position determination unit 705 analyzes the state of the real world based on the image of the real world acquired in step S901 and the distance map generated in step S902, and determines the display position of the virtual sound source object. For example, the display position determination unit 705 performs analysis using the image of the real world and the distance map, searches for a horizontal plane having an area that is likely to be able to display a 3D model of the virtual sound source, and determines the display position of the virtual sound source object. This processing is a processing that is generally performed in the field of augmented reality display, and is well known, so a description thereof will be omitted. The display position determination unit 705 outputs the determined display position of the virtual sound source object and the image of the real world to the image synthesis unit 707 and the real object analysis unit 710.
[0052] In step S904, the image synthesis unit 707 synthesizes an image related to the virtual sound source object and an image of the real world. The image synthesis unit 707 renders the virtual sound source object into a three-dimensional shape based on the virtual sound source 3D model 706, and scales the virtual sound source object so as to fit the surface size of the display position determined in step S903. Then, the image synthesis unit 707 synthesizes the virtual sound source object scaled according to the surface size of the display position at the display position of the real world image determined in step S903. Such processing is a processing generally performed in the field of virtual reality display, and is well known, so a description thereof will be omitted. The image synthesis unit 707 outputs the synthesized image to the image playback unit 708.
[0053] In step S905, the real object analysis unit 710 analyzes, as described above, what real object is located at the display position of the virtual sound source object determined in step S903 in the image of the real world, and acquires the product name, model number, etc. as the analysis result of the real object. The real object analysis unit 710 outputs this information of the acquired real object to the sound characteristic acquisition unit 712.
[0054] In step S906, the sound characteristic acquisition unit 712 searches for sound characteristic information stored in the sound characteristic DB 711 based on information such as the product name and model number of the real object obtained in step S905. As a result, the sound characteristic acquisition unit 712 acquires the sound characteristic due to vibration that is most suitable for the real object at the display position of the virtual sound source object from among the sound characteristics due to vibration stored in the sound characteristic DB 711. The sound characteristic acquisition unit 712 outputs the acquired sound characteristic due to vibration to the sound characteristic application unit 106. Details of the sound characteristic acquisition process from the sound characteristic DB 711 performed in step S906 will be described later with reference to FIG. 10.
[0055] The processes in steps S907 and S908 are similar to the processes in steps S503 and S504 in FIG. 5 in the first embodiment, respectively, and therefore will not be described.
[0056] In step S909, the video playback unit 708 and the audio playback unit 107 perform playback processing of video and audio related to the virtual reality display. The video playback unit 708 converts the video synthesized in step S904 (video synthesized from a video related to the virtual sound source object and a video of the real world) into an video for the VR goggles 709, plays it, and causes the VR goggles 709 to display the video. The audio playback unit 107 also appropriately amplifies and plays the sound source signal to which the sound characteristics due to vibration have been applied in step S908, and causes the headphones 108 to output the sound. This allows the viewer to view a virtual reality video in which the virtual sound source object is displayed, as well as a sound with a sound quality appropriate for the real object that the virtual sound source object is touching.
[0057] The process of step S910 is similar to the process of step S506 in FIG. 5 in the first embodiment, and therefore a description thereof will be omitted.
[0058] Fig. 10 is a flowchart showing an example of the process of acquiring pronunciation characteristics from the pronunciation characteristics DB in step S906 in Fig. 9. All the processes in the process of acquiring pronunciation characteristics from the pronunciation characteristics DB shown in Fig. 10 are executed by the pronunciation characteristics acquisition unit 712.
[0059] In step S1001, the pronunciation characteristic acquisition unit 712 initializes the pronunciation characteristic ID and the maximum score stored in the pronunciation characteristic acquisition unit 712. These values are stored in later processing as the pronunciation characteristic ID and score that are the search results of the pronunciation characteristic DB 711. In the initialization processing in step S1001, the pronunciation characteristic acquisition unit 712 stores, for example, an invalid value in the pronunciation characteristic ID and stores 0 in the maximum score.
[0060] In step S1002, the sound characteristic acquisition unit 712 acquires data of the sound characteristic information for the real object that can be acquired from outside, based on the information such as the product name and model number of the real object acquired in step S905. For example, the sound characteristic acquisition unit 712 performs an Internet search using the information such as the product name and model number to acquire information such as a specification sheet of the real object, and acquires information on the type, material, and size of the object in the sound characteristic information.
[0061] In step S1003, the sound characteristic acquisition unit 712 calculates the internal space volume of the physical object based on the information acquired in step S1002. For example, if the type of object is a desk, the sound characteristic acquisition unit 712 sets the internal space volume to 0, and if the type of object is a box, it calculates the volume of a rectangular parallelepiped of the same size as the internal space volume.
[0062] In step S1004, the pronunciation characteristic acquisition unit 712 selects the first pronunciation characteristic information from among the pronunciation characteristic information by vibration stored in the pronunciation characteristic DB 711, and sets it as a search processing target.
[0063] In step S1005, the pronunciation characteristic acquisition unit 712 calculates the similarity between each data of the pronunciation characteristic information of the real object acquired in steps S1002 and S1003 and each data of the pronunciation characteristic information in the pronunciation characteristic DB 711 selected in step S1004. The similarity calculated here is the similarity between the pronunciation characteristics influenced by each data. The similarity can be calculated, for example, by the following calculation. In the case of the type and material of the object, the similarity is determined by calculating a characteristic obtained by synthesizing the pronunciation characteristics for data having the same values stored in the pronunciation characteristic DB 711, and calculating the correlation between the synthesized pronunciation characteristics. The size and the internal space are also categorized into S, M, and L according to the size, and the similarity is determined by calculating a characteristic obtained by synthesizing the pronunciation characteristics in each category, and calculating the correlation between the synthesized pronunciation characteristics. The similarity between the values of these data can be calculated in advance, and can be stored as a table in the pronunciation characteristic DB 711. This allows the pronunciation characteristic acquisition unit 712 to check to which category each piece of data of the two pieces of pronunciation characteristic information to be compared belongs, and to acquire the degree of similarity by searching the comparison table.
[0064] In step S1006, the pronunciation characteristic acquisition unit 712 calculates a score for the pronunciation characteristic information to be searched based on the weight of each data determined in advance and the similarity of each data calculated in step S1005. The pronunciation characteristic acquisition unit 712 calculates the product of the weight of each data of the pronunciation characteristic information and the similarity, and calculates the score by adding the products for each data. Here, the weight of each data is determined in advance depending on the influence that the data value has on the pronunciation characteristic. The weight can be calculated by setting a reference pronunciation characteristic and calculating the degree of deformation of the pronunciation characteristic obtained by combining each data value or the pronunciation characteristic obtained by combining the categories of each data value. The weight can also be calculated in advance and stored in the pronunciation characteristic DB 711 as a table. The pronunciation characteristic acquisition unit 712 can acquire the weight of each data by referring to this table.
[0065] In step S1007, the pronunciation characteristic acquisition unit 712 compares the score calculated in step S1006 with the maximum score stored in the pronunciation characteristic acquisition unit 712. If the pronunciation characteristic acquisition unit 712 determines that the score calculated in step S1006 is greater than the maximum score stored as a result of the comparison (YES in step S1007), the process of step S1008 is executed. If the pronunciation characteristic acquisition unit 712 determines that the score calculated in step S1006 is greater than the maximum score stored (NO in step S1007), the process of step S1010 is executed.
[0066] In step S1008, the pronunciation characteristic acquisition unit 712 saves the pronunciation characteristic ID of the pronunciation characteristic information to be searched as a search result in a predetermined area on the RAM 403. If there is already a saved pronunciation characteristic ID, the pronunciation characteristic acquisition unit 712 overwrites and saves the pronunciation characteristic ID of the pronunciation characteristic information to be searched.
[0067] In step S1009, the pronunciation characteristic acquisition unit 712 saves the score related to the pronunciation characteristic information of the search process target calculated in step S1006 as the maximum score. If a maximum score has already been saved, the pronunciation characteristic acquisition unit 712 overwrites the score related to the pronunciation characteristic information of the search process target and saves it.
[0068] In step S1010, the pronunciation characteristic acquisition unit 712 judges whether or not the search process has been completed for all the pronunciation characteristic information stored in the pronunciation characteristic DB 711. If the pronunciation characteristic acquisition unit 712 judges that there is unprocessed pronunciation characteristic information for which the search process has not been executed among the pronunciation characteristic information stored in the pronunciation characteristic DB 711 (NO in step S1010), the process of step S1011 is executed. On the other hand, if the pronunciation characteristic acquisition unit 712 judges that the search process has been completed for all the pronunciation characteristic information stored in the pronunciation characteristic DB 711 (YES in step S1010), the process of step S1012 is executed.
[0069] In step S1011, the pronunciation characteristic acquisition unit 712 selects the next pronunciation characteristic information from among the pronunciation characteristic information due to vibration stored in the pronunciation characteristic DB 711, and sets it as a search processing target. After performing the processing of step S1011, the processing of step S1005 is executed. That is, the pronunciation characteristic acquisition unit 712 selects unprocessed pronunciation characteristic information from among the pronunciation characteristic information due to vibration stored in the pronunciation characteristic DB 711 in step S1011, and then executes the search processing for the pronunciation characteristic information selected in step S1011.
[0070] In step S1012, the pronunciation characteristic acquisition unit 712 judges whether or not the maximum value of the score obtained by the processing up to this point is equal to or greater than a predetermined threshold value. This threshold judgment judges whether or not the pronunciation characteristic due to vibration obtained as the search result is suitable for a real object. If the pronunciation characteristic acquisition unit 712 judges that the maximum value of the score is equal to or greater than the threshold value (YES in step S1012), the processing of step S1013 is executed. If the pronunciation characteristic acquisition unit 712 judges that the maximum value of the score is not equal to or greater than the threshold value, that is, that the maximum value of the score is less than the threshold value (NO in step S1012), the processing of step S1014 is executed.
[0071] In step S1013, the pronunciation characteristic acquisition unit 712 outputs the pronunciation characteristic due to vibration stored in the pronunciation characteristic information of the pronunciation characteristic ID saved in the pronunciation characteristic acquisition unit 712 as the search result to the pronunciation characteristic application unit 106. When the processing of step S1013 is completed, the pronunciation characteristic acquisition processing from the pronunciation characteristic DB is terminated, and the process returns to the video / audio reproduction processing shown in FIG.
[0072] In step S1014, the pronunciation characteristic acquisition unit 712 outputs to the pronunciation characteristic application unit 106 a predetermined pronunciation characteristic due to general vibration, rather than the pronunciation characteristic due to vibration of the pronunciation characteristic ID obtained as the search result. Here, the pronunciation characteristic due to general vibration is defined as a pronunciation characteristic having a frequency characteristic that is flat up to around 4 kHz and gradually decreases in a band above that. This makes it possible to produce a change in sound by applying the pronunciation characteristic while avoiding a sound that is far removed from a real object. When the process of step S1014 is completed, the process of acquiring the pronunciation characteristic from the pronunciation characteristic DB is terminated, and the process returns to the video / audio reproduction process shown in FIG. 9.
[0073] 10, the sound characteristic acquisition unit 712 can acquire sound characteristic information in the sound characteristic DB 711 that is most suitable for the real object at the display position of the virtual sound source object, and output the sound characteristic due to the vibration. Even if the sound characteristic DB 711 does not contain a sound characteristic suitable for the real object, it is possible to express a change in sound by applying the sound characteristic due to vibration while avoiding a sound that is unsuitable for the real object.
[0074] 10, the pronunciation characteristic information by a plurality of vibrations stored in the pronunciation characteristic DB 711 is selected in order from the first pronunciation characteristic information and search processing is performed, but this is not limited to this. It is sufficient to perform search processing on all the pronunciation characteristic information by vibrations stored in the pronunciation characteristic DB 711, and the pronunciation characteristic information for search processing may be selected in any order.
[0075] According to this embodiment, by acquiring a sound characteristic by excitation suitable for a real object present at the display position of a virtual sound source (virtual object) in virtual reality display and applying it to a sound source signal, it is possible to express how the sound changes to a quality suitable for that object. This makes it possible to reproduce a sound as if the virtual sound source (virtual object) displayed in virtual reality exists in the real world.
[0076] (Other embodiments) In the above-mentioned embodiment, only the sound characteristics due to vibration are applied to the sound source signal, but if the sound characteristics due to vibration have a narrow bandwidth and poor sound quality as shown in Fig. 3(b), applying them as they are will not produce a very good sound. Therefore, the original sound source signal may be added appropriately. For example, in the sound characteristics acquisition process in step S502 of Fig. 5, the original sound source signal may be added as shown in the flowchart of Fig. 11. Fig. 11 is a flowchart showing another example of the sound characteristics acquisition process in step S502 of Fig. 5.
[0077] The processing in steps S1101 to S1103 is similar to the processing in steps S601 to S603 in FIG. 6, respectively, and therefore description thereof will be omitted. In step S1104, the sound characteristic acquisition unit 103 adds a pulse based on the original sound source signal to the impulse response of the sound produced by the vibration obtained in step S1103. This makes it possible to add the original sound source signal, and also to adjust the addition ratio by adjusting the amplitude of the pulse to be added. The processes in steps S1105 and S1106 are similar to the processes in steps S604 and S605 in FIG. 6, respectively, and therefore will not be described.
[0078] Further, in the above-described embodiment, the sound generation characteristics due to vibration are described as impulse responses, but other characteristics, such as frequency characteristics obtained by exciting a frequency sweep signal, may be used. In addition, in the above-described embodiment, the sound source signal is repeatedly processed for each processing unit time, but a memory unit such as a large-capacity buffer may be provided and processing may be performed in units of songs, etc. Furthermore, in the above-described embodiment, headphones are used as the sound output destination, but sound may be output to other output devices (output equipment) such as speakers. Furthermore, in the above-described embodiment, when virtual reality is displayed, in addition to the virtual object that serves as a virtual sound source (virtual sound source object), other virtual objects may be displayed together with an image of the real world.
[0079] In the above-mentioned second embodiment, the pronunciation characteristic acquisition unit 712 calculates the score for all the pronunciation characteristic information stored in the pronunciation characteristic DB 711. However, this is not limited to this. For example, the pronunciation characteristic acquisition unit 712 may perform clustering in advance for the pronunciation characteristic information stored in the pronunciation characteristic DB 711, first check the approximate score with each cluster, and perform search processing only for the pronunciation characteristic information included in the most approximate cluster. In this case, the amount of processing required for the search processing can be reduced.
[0080] Also, in the above-mentioned second embodiment, the signal processing device may predetermine the range in which the viewer moves around, and may pre-extract physical objects within the camera's predicted shooting range. Furthermore, the signal processing device may perform a search process for the extracted physical objects in advance, and store a correspondence table between the physical objects in the shooting range and the search results in the sound characteristic DB. In this way, when playing back sound, the sound characteristic acquisition unit 712 refers to this table and quickly acquires the sound characteristic due to vibration, making it possible to perform processing in real time without delay.
[0081] In the above-mentioned second embodiment, the display position of the virtual sound source object is not limited to a horizontal plane, but may be a slope by calculating the inclination angle of the slope and appropriately correcting the display position of the virtual sound source object accordingly. Alternatively, the display position may be a vertical plane or the lower side of a ceiling by rotating the virtual sound source object so that it appears to be standing on the vertical plane or the lower side of a ceiling, or various planes with respect to a real object may be used as the display position according to the presentation in the video.
[0082] In the above-mentioned second embodiment, when the pronunciation characteristic DB711 does not contain a pronunciation characteristic suitable for a real object, a pronunciation characteristic due to vibration having a flat characteristic is applied, but the pronunciation characteristic due to vibration may not be applied to the sound source signal. In this way, when the pronunciation characteristic DB711 does not contain a pronunciation characteristic suitable for a real object, it is not possible to produce a change in sound according to the real object, but it is possible to avoid a sound that is far removed from reality.
[0083] In the above-mentioned second embodiment, the virtual sound source object is rendered based on a 3D model, but it is also possible to use a two-dimensional image extracted from a green or blue screen, a two-dimensional CG character, etc. Furthermore, it is also possible to generate a model of an object using CG or the like, or to generate a similar model, and play back sound using sound characteristic information that is adapted to or similar to the model of the object.
[0084] The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiment is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) for implementing one or more of the functions.
[0085] It should be noted that the above-mentioned embodiments are merely examples of the implementation of the present invention, and the technical scope of the present invention should not be interpreted as being limited by these. In other words, the present invention can be implemented in various forms without departing from its technical concept or main features.
[0086] The disclosure of the present embodiment includes the following configurations, methods, etc. (Configuration 1) a characteristic acquisition means for acquiring sound characteristics generated by excitation corresponding to an object in the real world that is in contact with the virtual object in a virtual reality display in which an image of the virtual object is superimposed on the real world; A characteristic application means for applying the obtained sound generation characteristic due to the vibration to a sound source signal of the virtual object; and a first reproduction means for reproducing a sound corresponding to the sound source signal to which the sound generation characteristic due to the vibration has been applied. (Configuration 2) a determination means for determining a display position for displaying the virtual object based on the image of the real world; 2. The signal processing device according to configuration 1, wherein the characteristic acquisition means acquires characteristics of sound produced by the vibration based on an object present at a display position of the virtual object. (Configuration 3) an analysis means for analyzing an object present at a display position of the virtual object based on the image of the real world; 3. The signal processing device according to configuration 2, wherein the characteristic acquisition means acquires characteristics of sound produced by the vibration based on a result of the analysis by the analysis means. (Configuration 4) A characteristic storage means for storing the sound generation characteristics of a plurality of objects generated by the vibration, The signal processing device according to any one of configurations 1 to 3, wherein the characteristic acquisition means acquires a sound characteristic caused by vibration corresponding to an object touched by the virtual object from among a plurality of sound characteristics caused by the vibration stored in the characteristic storage means. (Configuration 5) The signal processing device according to configuration 4, wherein the characteristic acquisition means searches the characteristic storage means based on an object that is present at the display position of the virtual object in the image of the real world, and acquires sound generation characteristics due to vibration corresponding to the object that the virtual object touches. (Configuration 6) The characteristic storage means stores the sound generation characteristics of the object caused by the vibration and attribute information of the object, The signal processing device according to configuration 4, wherein the characteristic acquisition means searches the characteristic storage means based on attribute information of an object present at the display position of the virtual object in the image of the real world, and acquires the sound characteristic caused by the vibration having the highest correlation with the attribute information. (Configuration 7) The signal processing device according to any one of configurations 1 to 6, wherein the characteristic acquisition means acquires the sound generation characteristic due to the vibration based on a sound generated by installing a vibrator on an object and vibrating the object. (Configuration 8) The signal processing device according to any one of configurations 1 to 7, wherein the characteristic application means performs filtering on the sound source signal of the virtual object using an FIR filter or an IIR filter having a filter coefficient corresponding to the acquired sound generation characteristic due to the vibration. (Configuration 9) The signal processing device according to any one of configurations 1 to 8, wherein the sound generation characteristic due to vibration is an impulse response obtained by installing a vibrator that vibrates an object on the object and analyzing a signal collected by a microphone. (Configuration 10) The signal processing device according to any one of configurations 1 to 8, characterized in that the sound generation characteristics due to vibration are created based on a response obtained by adding a pulse to an impulse response obtained by attaching a vibrator that vibrates an object to an object and analyzing a signal collected by a microphone. (Configuration 11) A synthesis means for synthesizing an image of a virtual object into a real world; 11. The signal processing device according to any one of configurations 1 to 10, further comprising a second reproduction means for reproducing the video generated by the synthesis means. (Method 1) a characteristic acquisition step of acquiring sound characteristics caused by vibration according to an object in the real world that the virtual object touches in a virtual reality display in which an image of the virtual object is superimposed on the real world; a characteristic application step of applying the acquired sound generation characteristics due to the vibration to a sound source signal of the virtual object; and a first reproduction step of reproducing a sound corresponding to the sound source signal to which the sound generation characteristic due to the vibration has been applied. (Program 1) a characteristic acquisition step of acquiring sound characteristics caused by vibration according to an object in the real world that the virtual object touches in a virtual reality display in which an image of the virtual object is superimposed on the real world; a characteristic applying step of applying the acquired sound generation characteristic due to the vibration to a sound source signal of the virtual object; A first reproduction step of reproducing a sound corresponding to the sound source signal to which the sound generation characteristic due to the vibration is applied. [Explanation of symbols]
[0087] 101: Transducer 102: Microphone 103: Pronunciation characteristic acquisition unit 104: Sound source signal 105: Sound source signal acquisition unit 106: Pronunciation characteristic application unit 107: Sound reproduction unit 108: Headphones 701: Camera 702: Video acquisition unit 703: Distance sensor 704: Distance map generation unit 705: Display position determination unit 706: Virtual sound source 3D model 707: Video synthesis unit 708: Video reproduction unit 709: VR goggles 710: Real object analysis unit 711: Pronunciation characteristic DB 712: Pronunciation characteristic acquisition unit
Claims
1. a first acquisition means for acquiring a virtual object; a second acquisition means for acquiring a captured image using a camera; a determination means for determining a material of an object in the real world that is to come into contact with the virtual object and that is displayed in the captured image; a display means for displaying a superimposed image in which the virtual object is superimposed on the captured image; a reproduction means for reproducing a contact sound between the virtual object and the physical object based on the determined material; A signal processing device comprising:
2. a determination unit that determines whether the virtual object is in contact with an object in the real world; The reproduction means reproduces the contact sound in response to the determination means determining that the virtual object is in contact with the object in the real world.
2. The signal processing device according to claim 1.
3. 2. The signal processing device according to claim 1, wherein the determining means determines the material by classifying the real-world object into one of a plurality of material classes including at least wood, metal, plastic, paper, glass, stone, and fabric.
4. The signal processing device according to claim 1 , wherein the determining unit determines the material by performing image recognition on a region of interest in the captured image that corresponds to the display position of the virtual object.
5. A distance sensor; a distance map generating means for generating a distance map corresponding to the angle of view of the camera; The determining means determines the material based on the captured image and the distance map.
2. The signal processing device according to claim 1.
6. a display position determining means for determining a display position of the virtual object in the captured image; The determining means determines the material of the real-world object present at the display position.
2. The signal processing device according to claim 1.
7. further comprising a storage means for storing a plurality of acoustic characteristics each associated with a material; The reproduction means selects an acoustic characteristic corresponding to the determined material from a plurality of acoustic characteristics stored in the storage means, and reproduces the contact sound based on the selected acoustic characteristic.
2. The signal processing device according to claim 1.
8. 8. The signal processing apparatus according to claim 7, wherein the selected acoustic characteristic is an impulse response corresponding to the determined material.
9. the storage means stores attribute information of real-world objects together with the plurality of acoustic characteristics; The reproduction means selects the acoustic characteristics based on the determined material and at least one additional attribute including object type, size, or interior volume.
8. The signal processing device according to claim 7,
10. The signal processing device according to claim 7, wherein the reproduction means, when the reliability of the determined material is less than a threshold value, selects predetermined default acoustic characteristics and reproduces the contact sound based on the default acoustic characteristics.
11. 2. The signal processing device according to claim 1, further comprising a mixing unit that mixes the sound generated based on the determined material with an original sound source signal associated with the virtual object at a mixing ratio.
12. a feedback obtaining means for obtaining user feedback indicating whether the played contact sound matches the real-world object; and updating means for updating the correspondence between materials and acoustic characteristics based on the user feedback.
2. The signal processing device according to claim 1.
13. 2. The signal processing device according to claim 1, wherein, when the determining means determines that the real-world object includes a plurality of materials, the reproducing means reproduces the contact sound by switching or blending between a plurality of acoustic characteristics corresponding to the plurality of materials.
14. 2. The signal processing device according to claim 1, wherein the display means combines a virtual object image corresponding to the virtual object with the captured image to form a single composite image, and outputs the composite image to a display device.
15. 2. The signal processing device according to claim 1, wherein the reproduction means controls at least one of a volume, an equalization characteristic, and a decay time of the contact sound based on the determined material.
16. The method further includes reading means for reading an identifier attached to the real-world object, 2. The signal processing device according to claim 1, wherein the determining means acquires the identifier and determines the material based on the identifier.
17. 2. The signal processing device according to claim 1, further comprising a log storage means for storing a log including the result of determining the material and an identifier of the acoustic characteristic used to reproduce the contact sound in the storage means.
18. obtaining a virtual object; acquiring a captured image generated by a camera; determining the material of a real-world object shown in the captured image captured by the camera; displaying a superimposed image in which the virtual object is superimposed on the captured image; reproducing a contact sound between the virtual object and the physical object based on the determined material; A signal processing method comprising:
19. obtaining a virtual object; acquiring a captured image generated by a camera; determining the material of a real-world object shown in the captured image captured by the camera; displaying a superimposed image in which the virtual object is superimposed on the captured image; reproducing a contact sound between the virtual object and the physical object based on the determined material; A program for causing a computer of a signal processing device to execute the above.