Sound processing apparatus, sound processing method, and program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-11-18
- Publication Date
- 2026-08-14
AI Technical Summary
【0010】 本開示によれば、より適切に立体的な音をユーザに知覚させることが可能となる。
Smart Images

Figure 0007905365000002 
Figure 0007905365000003 
Figure 0007905365000004
Abstract
Description
Technical Field
[0001] The present disclosure relates to an acoustic processing device, as well as an acoustic processing method and a program related to the acoustic device. process
Background Art
[0002] Conventionally, there has been known a technique related to acoustic reproduction for causing a user to perceive stereophonic sound by controlling the position of a sound image, which is a virtual sound source object, in a virtual three-dimensional space (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] On the other hand, when causing a user to perceive sound as stereophonic sound in a three-dimensional sound field, there may be a case where sound that is difficult for the user to perceive is generated. In the information processing method in a conventional acoustic reproduction device or the like, there may be a case where appropriate processing is not performed on such difficult-to-perceive sound.
[0005] In view of the above, an object of the present disclosure is to provide an acoustic processing device or the like that allows a user to more appropriately perceive stereophonic sound.
Means for Solving the Problems
[0006] An acoustic processing device according to one aspect of the present disclosure is an acoustic processing device that causes a user to perceive a reproduced sound as a sound arriving from a predetermined direction in a three-dimensional sound field, and comprises: a first processing unit that generates a first output sound signal by convolving a first head-direction transfer function into sound information including the reproduced sound, to localize the sound contained in the information as a sound arriving from the predetermined direction; a second processing unit that generates a second output sound signal by convolving a second head-direction transfer function into the sound information, to localize the sound contained in the information as a sound arriving from a first direction having a first angle with respect to the predetermined direction greater than 0 degrees and less than 360 degrees, and having a first delay time greater than 0 and a first volume attenuation greater than 0 with respect to the reproduced sound perceived by the first output sound signal; and a combiner that outputs an output sound signal obtained by combining the generated first output sound signal and the second output sound signal.
[0007] Furthermore, an acoustic processing method according to one aspect of the present disclosure is an acoustic processing method that causes a user to perceive a reproduced sound as a sound arriving from a predetermined direction in a three-dimensional sound field, wherein a first head-direction transfer function for localizing the sound contained in the information as a sound arriving from the predetermined direction is convolved onto sound information including the reproduced sound to generate a first output sound signal; a second output sound signal for localizing the sound contained in the information as a sound arriving from a first direction having a first angle with respect to the predetermined direction that is greater than 0 degrees and less than 360 degrees, and having a first delay time greater than 0 and a first volume attenuation greater than 0 with respect to the reproduced sound perceived by the first output sound signal; and an output sound signal obtained by combining the generated first output sound signal and the second output sound signal.
[0008] Furthermore, one aspect of this disclosure can also be implemented as a program for causing a computer to execute the sound processing method described above.
[0009] These comprehensive or specific embodiments may be implemented as systems, devices, methods, integrated circuits, computer programs, or non-temporary recording media such as computer-readable CD-ROMs, or as any combination of systems, devices, methods, integrated circuits, computer programs, and recording media. [Effects of the Invention]
[0010] This disclosure makes it possible to enable users to perceive three-dimensional sound more appropriately. [Brief explanation of the drawing]
[0011] [Figure 1] Figure 1 is a schematic diagram showing an example of use of the acoustic processing device according to the embodiment. [Figure 2] Figure 2 is a block diagram showing the functional configuration of an audio playback device according to an embodiment. [Figure 3] Figure 3 is a block diagram showing a more detailed functional configuration of the acoustic processing device according to the embodiment. [Figure 4] Figure 4 is a diagram illustrating the sound attenuation according to the embodiment. [Figure 5] Figure 5 is a diagram illustrating the direction of sound output by the sound processing device according to the embodiment. [Figure 6] Figure 6 is a flowchart showing the operation of the sound processing apparatus according to the embodiment. [Figure 7] Figure 7 illustrates an appropriate first angle according to the embodiment. [Figure 8] Figure 8 illustrates an appropriate first delay time according to the embodiment. [Figure 9] Figure 9 illustrates the appropriate first sound attenuation according to the embodiment. [Figure 10] Figure 10 is a block diagram showing the functional configuration of an audio playback device according to a modified embodiment. [Figure 11] Figure 11 is a block diagram showing the detailed functional configuration of an acoustic processing device according to a modified embodiment. [Figure 12] Figure 12 is a diagram illustrating the direction of arrival of sound output by an acoustic processing device according to a modified embodiment. [Figure 13] Figure 13 is a flowchart showing the operation of an acoustic apparatus according to a modified embodiment. [Modes for carrying out the invention]
[0012] (Knowledge that formed the basis of the disclosure) Conventionally, there are known technologies for acoustic reproduction that allow users to perceive three-dimensional sound by controlling the position of a sound image, which is a sound source object perceived by the user, within a virtual three-dimensional space (hereinafter sometimes referred to as a three-dimensional sound field) (see, for example, Patent Document 1). By localizing a sound image to a predetermined position in a virtual three-dimensional space, the user can perceive this sound as if it were arriving from a direction parallel to the line connecting that predetermined position and the user (i.e., a predetermined direction). In order to localize a sound image to a predetermined position in a virtual three-dimensional space in this way, for example, computational processing is required to generate a difference in the arrival time of the sound between the two ears, and a difference in the sound level (or sound pressure difference) between the two ears, so that the recorded sound is perceived as three-dimensional.
[0013] Here, in recent years, so-called online conferencing systems that transmit and receive video and audio bidirectionally using a communication line to communicate with a communication partner have been widely used. In such an online conferencing system, a head-mounted acoustic playback device such as headphones is often used. As represented by the above online conferencing system, when listening to sound with headphones, it is difficult to expand the sound into a three-dimensional sound field and make the user perceive it. For example, simply using the direction of the display device on which the communication partner is displayed as the sound arrival direction and only convolution of the head transfer function to make the sound be perceived as arriving from this direction is known not to obtain a sufficient sense of presence outside the head. That is, since the sound image is localized inside the user's head, there is a sense of incongruity between the video of the communication partner in front of the display device and the sound localized inside the head. And when continuing to listen while having such a sense of incongruity, there may be a case where one becomes overly fatigued. Similar problems can also occur when listening to the sound of content using a three-dimensional video space such as VR or AR with an acoustic playback device such as headphones.
[0014] Conventionally, a technique that enables expanding sound into a three-dimensional sound field even when using headphones has been known. For example, there is a method of simulating how reflected sound is generated assuming a virtual room, artificially creating and synthesizing these reflected sounds, and having the user listen to them. Then, the user can perceive the original sound as if it is arriving from a predetermined direction inside the virtual room by the sound including the synthesized reflected sound. However, in this method, it is necessary to calculate the reflected sound generated inside the virtual room by complex calculations, and also to perform a large number of convolutions of the head transfer function to create such reflected sounds. The process of convolving the head transfer function for making the sound be perceived as the reflected sound arriving from a certain direction with the signal of the target sound usually requires a huge amount of calculation, so a large-scale computing device is required.
[0015] On the other hand, it is also possible to create a sound similar to a reflected sound by causing a time delay in the sound signal and performing filter processing to attenuate the volume. However, this filter processing has a low effect of expanding the sound into the three-dimensional sound field and lacks practicality.
[0016] In view of the above, in the present disclosure, when making a user perceive a sound as a sound from a predetermined direction within a three-dimensional sound field using an acoustic reproduction device such as headphones, by creating and synthesizing about one to several reflected sounds, an acoustic processing device that can sufficiently obtain the effect of expanding into the three-dimensional sound field without requiring a large-scale computing device will be described.
[0017] More specifically, the acoustic processing device according to the first aspect of the present disclosure is an acoustic processing device that makes a user perceive a reproduced sound as a sound arriving from a predetermined direction on a three-dimensional sound field. A first processing unit that generates a first output sound signal by convolution of a first head-related transfer function for localizing the sound included in the information as a sound arriving from a predetermined direction with respect to the sound information including the reproduced sound; a second processing unit that generates a second output sound signal by convolution of a second head-related transfer function for localizing the sound included in the information as a sound having a first delay time greater than 0 and a first volume attenuation greater than 0 with respect to the reproduced sound perceived by the first output sound signal and arriving from a first direction having a first angle greater than 0 degrees and less than 360 degrees with respect to the predetermined direction; and a combiner that outputs an output sound signal obtained by combining the generated first output sound signal and the second output sound signal.
[0018] In such an acoustic processing device, the second output sound signal is localized as a sound arriving from the first direction, having a first delay time and a first volume attenuation, and is therefore perceived by the user as a reflected sound, where the reproduced sound has been reflected by a pseudo-reflective wall. As a result, the reflected sound is perceived along with the reproduced sound as the direct sound, with the first delay time and first volume attenuation, improving the sense of the sound image of the direct sound being outside the head. In particular, in this process, it is sufficient that at least the second output sound signal is synthesized and output together with the first output sound signal, and the effect of improving the sense of the direct sound being outside the head can be obtained if the computational processing to generate the second output sound signal is possible. Thus, it is possible to make the user perceive a more appropriate three-dimensional sound while keeping the computational cost required for processing low.
[0019] Furthermore, for example, the sound processing apparatus according to the second aspect of this disclosure is the sound processing apparatus according to the first aspect, wherein the output sound signal is reproduced using headphones or earphones worn on the user's head.
[0020] According to this, it becomes possible to allow users to perceive three-dimensional sound more accurately using headphones or earphones worn on their heads.
[0021] Furthermore, for example, an acoustic apparatus according to a third aspect of this disclosure is an acoustic apparatus according to the first or second aspect, wherein the first angle is an angle within an angular range greater than 90 degrees and less than 270 degrees with respect to a predetermined direction.
[0022] According to this, when the predetermined direction from which the reproduced sound arrives coincides with the user's frontal direction, the angle range between 90 degrees and 270 degrees corresponds to the user's rear. Therefore, when the user is facing the direction of the reproduced sound, the reflected sound will arrive from the user's rear. When localizing reflected sound, it is effective to localize it to the user's rear in order to make the presence of the reflected sound itself less noticeable, and by doing so as described above, it becomes possible to make the user perceive a more appropriate three-dimensional sound.
[0023] Furthermore, for example, the acoustic processing apparatus according to the fourth aspect of the present disclosure further comprises a third processing unit that generates a third output sound signal by convolving a third head transfer function onto sound information to localize the sound contained in the information as a sound that arrives from a second direction having a second angle greater than 0 degrees and less than 360 degrees with respect to a predetermined direction, and having a second delay time greater than 0 and a second volume attenuation greater than 0 with respect to the reproduced sound perceived by the first output sound signal, and a combiner outputs an output sound signal which is a combination of the first output sound signal, the second output sound signal and the third output sound signal.
[0024] According to this, the third output sound signal arrives from the second direction and is further localized as a sound with a second delay time and a second volume attenuation, so that the user perceives the reproduced sound as a reflected sound that has been further reflected by a pseudo-reflective wall. Therefore, along with the reproduced sound as the direct sound and the reflected sound from the second output sound signal, a further reflected sound with a second delay time and a second volume attenuation is perceived, further improving the sense of being outside the head at the position where the sound image of the direct sound is localized. By generating and perceiving two or more small reflected sounds in this way, a high effect of improving the sense of being outside the head can be obtained with relatively low computational cost, making it possible to perceive three-dimensional sound more appropriately to the user.
[0025] Furthermore, for example, the acoustic apparatus according to the fifth aspect of this disclosure is the acoustic apparatus according to the fourth aspect, wherein the second angle is an angle within an angular range where the angle with respect to a predetermined direction is greater than 90 degrees and less than 270 degrees, and the difference angle obtained by subtracting the second angle from 360 degrees does not coincide with the first angle.
[0026] According to this, when the predetermined direction from which the reproduced sound arrives coincides with the user's frontal direction, the angle range between 90 degrees and 270 degrees corresponds to the user's rear. Therefore, when the user is facing the direction of the reproduced sound, the reflected sound will arrive from the user's rear. When localizing reflected sound, it is effective to localize it to the user's rear in order to make the presence of the reflected sound itself less noticeable, and by doing so as described above, it becomes possible to make the user perceive a more appropriate three-dimensional sound.
[0027] Furthermore, for example, the sound processing apparatus according to the sixth aspect of this disclosure is the sound processing apparatus according to the fourth or fifth aspect, wherein the first delay time and the second delay time are different delay times.
[0028] According to this, the possibility that the reflected sound from the second output sound signal and the reflected sound from the third output sound signal will be perceived as the same single reflected sound can be reduced, making it possible to allow the user to perceive a more appropriately three-dimensional sound using the two reflected sounds.
[0029] Furthermore, for example, the sound processing apparatus according to the seventh aspect of this disclosure is the sound processing apparatus according to any one of the fourth to sixth aspects, wherein the first volume attenuation and the second volume attenuation are different volume attenuation amounts.
[0030] According to this, the possibility that the reflected sound from the second output sound signal and the reflected sound from the third output sound signal will be perceived as the same single reflected sound can be reduced, making it possible to allow the user to perceive a more appropriately three-dimensional sound using the two reflected sounds.
[0031] Furthermore, for example, the sound processing apparatus according to the eighth aspect of this disclosure further includes a reverberation suppression processing unit that performs reverberation suppression processing on sounds contained in information to reduce the reverberation components contained in the information, and the sound information is generated by performing reverberation suppression processing on original sound information containing reverberation components, and includes sounds other than the reduced reverberation components from the sounds contained in the original sound information as reproduced sound, as described in any one of the first to seventh aspects.
[0032] According to this method, if the original sound information contains reverberation components, it is possible to reduce these reverberation components and generate sound information. Then, by generating both the reproduced sound and the reflected sound from the sound information, it becomes possible to allow the user to perceive a more appropriate three-dimensional sound.
[0033] Furthermore, for example, the acoustic processing apparatus according to the ninth aspect of the present disclosure further includes an acquisition unit that acquires sensing results from a sensor that detects the movement of the user's head, and a second processing unit convolves a second head-related transfer function obtained by changing the volume attenuation amount of the first volume attenuation with respect to sound information based on the acquired sensing results, as described in any one of the first to eighth aspects.
[0034] According to this, the amount of volume attenuation in the second output sound signal can be changed based on the user's head movements. For example, when the user moves their head and the direction from which the reflected sound arrives becomes closer to the direction in front of the user, the user may become aware of the presence of the reflected sound itself, and the effect of improving the sense of externality of the reproduced sound may not be properly achieved. According to this embodiment, in the above case, by increasing the amount of volume attenuation of the reflected sound (reducing the volume), the possibility of the user's attention being drawn to the reflected sound can be reduced. Therefore, it becomes possible to make the user perceive a more appropriate three-dimensional sound.
[0035] Furthermore, for example, the acoustic processing apparatus according to the 10th aspect of this disclosure is an acoustic processing apparatus according to the 9th aspect, wherein a first head-level transfer function is convolved to localize the sound contained in the information as a sound arriving from a predetermined direction and having a third volume attenuation of 0 or more, and the first processing unit convolves a first head-level transfer function with a reduced volume attenuation of the third volume attenuation onto the sound information when the volume attenuation amount of the first volume attenuation in the second processing unit increases, and convolves a first head-level transfer function with an increased volume attenuation of the third volume attenuation onto the sound information when the volume attenuation amount of the first volume attenuation in the second processing unit decreases.
[0036] According to this, the volume attenuation of the reproduced sound can be changed in sync with the volume attenuation of the reflected sound. Specifically, when the volume attenuation of the reflected sound increases (the volume decreases), the volume attenuation of the reproduced sound decreases (the volume is amplified). Also, when the volume attenuation of the reflected sound decreases (the volume increases), the volume attenuation of the reproduced sound decreases (the volume decreases). In this way, the volume of the reproduced sound and the reflected sound can complement each other so that the total volume of the entire sound in the three-dimensional sound field does not change drastically.
[0037] Furthermore, for example, the acoustic processing apparatus according to the 11th aspect of this disclosure further includes an acquisition unit that acquires sensing results from a sensor that detects the movement of the user's head, and a third processing unit convolves a third head-related transfer function obtained by changing the volume attenuation amount of the second volume attenuation based on the acquired sensing results onto the sound information, the acoustic processing apparatus according to any one of the 5th to 10th aspects referencing the 4th aspect.
[0038] According to this, the amount of volume attenuation in the third output sound signal can be changed based on the user's head movements. For example, when the user moves their head and the direction from which the reflected sound arrives becomes closer to the direction in front of the user, the user may become aware of the presence of the reflected sound itself, and the effect of improving the sense of externality of the reproduced sound may not be properly achieved. According to this embodiment, in the above case, by increasing the amount of volume attenuation of the reflected sound (reducing the volume), the possibility of the user's awareness being directed towards the reflected sound can be reduced. Therefore, it becomes possible to make the user perceive a more appropriate three-dimensional sound.
[0039] Furthermore, for example, the acoustic apparatus according to the twelfth aspect of this disclosure is an acoustic apparatus according to any one of the first to eleventh aspects, wherein at least one of the first angle, first delay time, and first volume attenuation is adjusted by the user.
[0040] According to this, the user can adjust at least one of the following to suit their own preferences: the first angle, the first delay time, and the first volume attenuation.
[0041] Furthermore, for example, the acoustic apparatus according to the 13th aspect of this disclosure is an acoustic apparatus according to any one of the 5th to 12th aspects, referencing the 4th aspect, wherein at least one of the second angle, second delay time, and second volume attenuation is adjusted by the user.
[0042] According to this, the user can adjust at least one of the following to suit their own preferences: the second angle, the second delay time, and the second volume attenuation.
[0043] Furthermore, for example, the sound processing apparatus according to the 14th aspect of this disclosure is the sound processing apparatus according to any one of the 1st to 13th aspects, wherein sound information is generated based on original sound information including a reproduced sound and reverberation components, and the first delay time is a delay time smaller than the delay time of the reverberation component relative to the reproduced sound.
[0044] According to this, the sound perceived by the user by the first output sound signal has a shorter delay time than the reverberation that occurs in the original sound information's recording environment. Therefore, the sound perceived by the first output sound signal is less likely to be perceived as noise like reverberation. In other words, the sound perceived by the first output sound signal can be appropriately perceived by the user as reflected sound.
[0045] Furthermore, for example, an acoustic processing apparatus according to the 15th aspect of this disclosure is an acoustic processing apparatus according to any one of the 5th to 14th aspects, referencing the 4th aspect, wherein sound information is generated based on original sound information including a reproduced sound and reverberation components, and the second delay time is a delay time smaller than the delay time of the reverberation component relative to the reproduced sound.
[0046] According to this, the sound perceived by the user by the second output sound signal has a shorter delay time than the reverberation that occurs in the original sound information's recording environment. Therefore, the sound perceived by the first output sound signal is less likely to be perceived as noise like reverberation. In other words, the sound perceived by the first output sound signal can be appropriately perceived by the user as reflected sound.
[0047] Furthermore, the sixteenth aspect of the present disclosure is an acoustic processing method that causes a user to perceive a reproduced sound as a sound arriving from a predetermined direction in a three-dimensional sound field, and generates a first output sound signal by convolving a first head-direction transfer function into sound information including a reproduced sound to localize the sound contained in the information as a sound arriving from a predetermined direction, generates a second output sound signal by convolving a second head-direction transfer function into sound information to localize the sound contained in the information as a sound arriving from a first direction having a first angle with respect to the predetermined direction that is greater than 0 degrees and less than 360 degrees, and having a first delay time greater than 0 and a first volume attenuation greater than 0 with respect to the reproduced sound perceived by the first output sound signal, and outputs an output sound signal obtained by combining the generated first output sound signal and the second output sound signal.
[0048] Such an acoustic processing method can achieve the same effects as the acoustic processing device described above.
[0049] Furthermore, the program relating to the 17th aspect of this disclosure is a program for causing a computer to execute the sound processing method described above.
[0050] Such a program can use a computer to produce the same effects as the sound processing device described above.
[0051] (Embodiment) [overview] First, an overview of the sound reproduction device according to the embodiment will be described. Figure 1 is a schematic diagram showing an example of use of the sound reproduction device according to the embodiment. In Figure 1, (a) shows a user 99 using one of the two examples of sound reproduction device 100, and (b) shows a user 99 using the other example of sound reproduction device 100.
[0052] As described above, the sound reproduction device 100 shown in Figure 1 is used simultaneously with a display device for displaying images and a device for reproducing stereoscopic images (neither of which are shown).
[0053] The sound reproduction device 100 is a sound presentation device worn on the head of the user 99. Therefore, the sound reproduction device 100 moves integrally with the head of the user 99. For example, the sound reproduction device 100 in this embodiment may be a so-called over-ear headphone type device, as shown in Figure 1(a), or it may be two earplug type devices (in-ear headphone type devices) that are independently worn on the left and right ears of the user 99, as shown in Figure 1(b). These two devices communicate with each other to present the sound for the right ear and the sound for the left ear in synchronous manner.
[0054] Furthermore, the sound reproduction device described herein is not limited to head-mounted sound reproduction devices such as over-ear headphones and in-ear headphones. For example, it can also be applied to sound reproduction devices that are installed close to both ears of the user 99, such as headrest speakers, when the speakers are not attached to the user 99.
[0055] The sound reproduction device 100 changes the sound it presents in response to the user 99's head movements, thereby making the user 99 perceive that they are moving their head within a three-dimensional sound field. To this end, as described above, the sound reproduction device 100 moves the three-dimensional sound field in the opposite direction to the user's movements in response to the user 99's movements.
[0056] [composition] Next, the configuration of the sound reproduction device 100 according to this embodiment will be described with reference to Figures 2 and 3. Figure 2 is a block diagram showing the functional configuration of the sound reproduction device according to this embodiment. Figure 3 is a block diagram showing a more detailed functional configuration of the sound processing device according to this embodiment.
[0057] As shown in Figure 2, the sound reproduction device 100 according to this embodiment comprises a sound processing device 101, a communication module 102, a sensor 103, and a driver 104.
[0058] The sound processing device 101 is a computing device for performing various signal processing in the sound reproduction device 100. The sound processing device 101 includes, for example, a processor and memory, and performs various functions when a program stored in memory is executed by the processor.
[0059] The sound processing device 101 includes an acquisition unit 111, a first processing unit 121, a second processing unit 131, and a combiner 150. The acquisition unit 111 will be described later in conjunction with the description of the communication module 102, and the combiner 150 will be described later in conjunction with the description of the driver 104.
[0060] The first processing unit 121 generates the output sound signal for playback. The first processing unit 121 is a functional unit that generates the first output sound signal by convolving a first head-related transfer function that localizes the sound contained in the information as sound arriving from a predetermined direction. The first processing unit 121 convolves a head-related transfer function that localizes the sound in a predetermined direction onto the input sound information and outputs a attenuated first output sound signal via volume attenuation α (third volume attenuation). This processing by the first processing unit 121 is collectively understood as convolving the first head-related transfer function. The first output sound signal is input to the first EQ 122, where the low and high frequencies are adjusted before being sent to the combiner 150.
[0061] The second processing unit 131 generates the output sound signal of the first reflected sound. The second processing unit 131 is a functional unit that generates the second output sound signal by convolving a second head-related transfer function into the input sound information, which is used to localize the sound contained in the information as a sound that arrives from a first direction having a first angle greater than 0 degrees and less than 360 degrees with respect to the playback sound perceived by the first output sound signal, and has a first delay time greater than 0 and a first volume attenuation greater than 0. The second processing unit 131 convolves a head-related transfer function into the input sound information to localize the sound in the first direction, and outputs a attenuated second output sound signal via volume attenuation β (first volume attenuation). This processing by the second processing unit 131 is collectively understood as the convolution of the second head-related transfer function. The second output sound signal is input to the second EQ 132, where the low and high frequencies are adjusted, and then it is sent to the combiner 150. Furthermore, before the sound information is input to the second processing unit 131, the first angle determination unit 130 adds information specifying the head-related transfer function to be convolved thereafter.
[0062] The communication module 102 is an interface device for receiving sound information input to the sound reproduction device 100. The communication module 102 includes, for example, an antenna and a signal converter, and receives sound information from an external device via wireless communication. More specifically, the communication module 102 receives a wireless signal indicating sound information converted into a format for wireless communication using the antenna, and the signal converter converts the wireless signal back into sound information. As a result, the sound reproduction device 100 acquires sound information from an external device via wireless communication. The sound information acquired by the communication module 102 is acquired by the acquisition unit 111. In this way, the sound information is input to the sound processing device 101. Note that communication between the sound reproduction device 100 and the external device may be performed by wired communication.
[0063] Furthermore, the sound processing device 101 includes a reverberation suppression processing device 120 as shown in Figure 3. When generating and synthesizing reflected sound, if the original sound contains reverberation components, that is, components of sound that are input to the sound receiver with a delay due to reflection or other factors in the sound pickup environment, the effect of improving the sense of sound being outside the head by synthesizing reflected sound is reduced. For this reason, the sound processing device 101 uses the reverberation suppression processing device 120 to perform reverberation suppression processing on the sound contained in the information to reduce the reverberation components contained in the information. By performing reverberation suppression processing on the original sound information, which includes the playback sound to be reproduced and the reverberation components, sound information can be generated that includes sounds other than the reduced reverberation components from the sound information contained in the original sound information as the playback sound, and this can be input to the first processing device 121 and the second processing device 131. The reverberation suppression processing device 120 may be inserted before the acquisition device 111 or after the acquisition device 111.
[0064] The sound information acquired by the sound reproduction device 100 is encoded in a predetermined format, such as MPEG-H 3D Audio (ISO / IEC 23008-3). As an example, the encoded sound information includes information about the reproduced sound played by the sound reproduction device 100, and information about the localization position when the sound image of the sound is localized to a predetermined position in a three-dimensional sound field (i.e., perceived as a sound arriving from a predetermined direction), i.e., information about the predetermined direction. For example, the sound information includes information about multiple sounds, including a first reproduced sound and a second reproduced sound, and the sound image is localized so that the sound image when each sound is played is perceived as a sound arriving from a different direction in a three-dimensional sound field.
[0065] This three-dimensional sound can enhance the sense of presence of content being viewed, for example, in conjunction with images viewed using a display device. The sound information may include only information about the reproduced sound. In this case, information regarding a predetermined direction may be acquired separately. As mentioned above, the sound information includes first sound information related to the first reproduced sound and second sound information related to the second reproduced sound, but multiple sound information sets containing these separately may be acquired and played back simultaneously to localize sound images at different positions within the three-dimensional sound field. Thus, there are no particular limitations on the form of the input sound information, and it is sufficient that the sound reproduction device 100 (especially the sound processing device 101) is equipped with an acquisition unit 111 corresponding to various forms of sound information.
[0066] In this embodiment, the acquisition unit 111 includes, for example, an encoded sound information input unit, a decoding processing unit, and a sensing information input unit.
[0067] The encoded sound information input unit is a processing unit that receives encoded sound information acquired by the acquisition unit 111. The encoded sound information input unit outputs the input sound information to the decoding processing unit. The decoding processing unit is a processing unit that decodes the sound information output from the encoded sound information input unit to generate information about a predetermined sound and information about a predetermined direction contained in the sound information in a format to be used in subsequent processing. The sensing information input unit will be described below along with the function of the sensor 103.
[0068] Sensor 103 is a device for detecting the speed of the user 99's head movement. Sensor 103 is composed of a combination of various sensors used for motion detection, such as a gyro sensor and an accelerometer. In this embodiment, sensor 103 is built into the sound playback device 100, but it may also be built into an external device, such as a stereoscopic image playback device that operates in response to the user 99's head movement, similar to the sound playback device 100. In this case, sensor 103 does not have to be included in the sound playback device 100. Alternatively, as sensor 103, an external imaging device or the like may be used to capture images of the user 99's head movement, and the user 99's head movement may be detected by processing the captured images.
[0069] The sensor 103 is, for example, integrally fixed to the housing of the sound reproduction device 100 and detects the speed of movement of the housing. Since the sound reproduction device 100, including the housing, moves integrally with the user 99's head after the user 99 puts it on, the sensor 103 can consequently detect the speed of movement of the user 99's head.
[0070] The sensor 103 may, for example, detect the amount of rotation of the user 99's head, with at least one of the three mutually orthogonal axes in three-dimensional space as the axis of rotation, or it may detect the amount of displacement with at least one of the three axes as the direction of displacement. Alternatively, the sensor 103 may detect both the amount of rotation and the amount of displacement as the amount of movement of the user 99's head.
[0071] The sensing information input unit of the acquisition unit 111 acquires the movement velocity of the user 99's head from the sensor 103. More specifically, the sensing information input unit acquires the amount of movement of the user 99's head detected by the sensor 103 per unit time as the movement velocity. In this way, the sensing information input unit acquires at least one of the rotational velocity and displacement velocity as sensing results from the sensor 103. The amount of movement of the user 99's head acquired here is used to determine the coordinates and orientation of the user 99 in the three-dimensional sound field. In the sound reproduction device 100, the relative position of the sound image is determined based on the determined coordinates and orientation of the user 99, and sound is reproduced.
[0072] Furthermore, in this embodiment, the sensing results acquired from the sensor 103 by the sensing information input unit of the acquisition unit 111 are used to control the volume attenuation amounts of volume attenuation α and volume attenuation β. In other words, the volume attenuation amounts of volume attenuation α and volume attenuation β change automatically according to the sensing results. This is because if the reflected sound is clearly audible from the direction in which the user 99 is facing, the user 99 may feel uncomfortable. Therefore, when the user 99 rotates their head, the volume of the reflected sound is controlled to be attenuated as the direction in which the user 99 is facing approaches the direction of the reflected sound. At the same time, the volume of the reproduced sound is amplified (volume attenuation is reduced) so that the overall volume does not change. In other words, the first processing unit 121 decreases the volume attenuation amount of volume attenuation α when the volume attenuation amount of volume attenuation β in the second processing unit 131 increases, and increases the volume attenuation amount of volume attenuation α when the volume attenuation amount of volume attenuation β in the second processing unit 131 decreases.
[0073] Figure 4 illustrates the sound attenuation according to the embodiment. In the figure, the sound attenuation amount of sound attenuation α (dashed line) and the sound attenuation amount of sound attenuation β (solid line) are shown with respect to the rotation angle (yaw angle) when the user 99's head rotates around an axis parallel to the vertical direction of the user 99's head. The first angle here is set to 120 degrees. Here, the sound attenuation amount of sound attenuation α and the sound attenuation amount of sound attenuation β are calculated based on the following equation (1).
[0074]
number
[0075] In the above equation, α represents the volume attenuation amount (gain) of volume attenuation α, and β represents the volume attenuation amount (gain) of volume attenuation β. In this example, it can be seen that when the user 99 rotates their head to half that angle, 60 degrees, for a reflected sound set at an angle of 120 degrees with respect to a predetermined direction, the reflected sound disappears. In this way, the sound processing device 101 appropriately changes the volume attenuation amounts of the reproduced sound and the reflected sound so that the reflected sound itself does not become a source of discomfort. The relationship between volume attenuation α and volume attenuation β explained using equation (1) above is just one example, and any relationship can be used as long as the volume attenuation amount of the reflected sound increases as the user 99 rotates their head toward the direction of the reflected sound. Furthermore, the above relationship may also hold not only for the relationship between volume attenuation α and volume attenuation β, but also for the relationship between volume attenuation α and volume attenuation γ when generating other reflected sounds with volume attenuation γ (described later in the modified examples).
[0076] The combiner 150 is a functional unit that synthesizes the generated output sound signals and outputs them to the driver 104. The combiner 150 outputs a synthesized output sound signal by adding the first output sound signal and the second output sound signal. The combiner 150 further generates a waveform signal by performing signal conversion from a digital signal to an analog signal based on the output sound signal, and generates sound waves in the driver 104 based on the waveform signal, presenting sound to the user 99. The driver 104 has, for example, a diaphragm and a drive mechanism such as a magnet and a voice coil. The driver 104 operates the drive mechanism according to the waveform signal, and the drive mechanism vibrates the diaphragm. In this way, the driver 104 generates sound waves by the vibration of the diaphragm according to the output sound signal, the sound waves propagate through the air and are transmitted to the user 99's ears, and the user 99 perceives the sound.
[0077] As described above, when the output sound signal from the combiner 150 is reproduced by the driver 104, a sound field like that shown in Figure 5 is formed. Figure 5 is a diagram illustrating the direction of arrival of sound output by the sound processing device according to the embodiment. Figure 5 shows a plan view of a virtual three-dimensional sound field from a direction along the vertical direction of the user's 99 head. Figure 5 shows the user 99 in a posture with the top of the paper facing forward, and this user 99 is in an upright posture perpendicular to the plane of the paper. A predetermined direction in which the reproduced sound is localized is set in the direction in front of the user 99. The position P1 where the reproduced sound is localized is shown as a black circle, and a virtual speaker is also shown.
[0078] As shown in the figure, the first reflected sound is localized in a direction having a first angle clockwise from a predetermined direction (position P2).
[0079] Furthermore, the dashed lines extending to the left and right of User 99 in the figure represent a hypothetical boundary that divides User 99's head into front and back sections. This boundary may be a surface that follows User 99's external auditory canal, a surface that passes through the point at the rearmost end of User 99's auricle, or simply a surface that passes through the center of gravity of User 99's head. It is known that there is a difference in the ease of hearing sound on and off such a boundary, that is, on the front and back of User 99. When localizing reflected sound, it is effective to localize it on the rear side of User 99 in order to make the presence of the reflected sound itself less noticeable. Therefore, the first angle should be set to an angle within the range of angles greater than 90 degrees and less than 270 degrees with respect to a predetermined direction.
[0080] Although the first angle, first delay time, and first volume attenuation described above were explained as values preset by the sound processing device 101 or values that change according to the sensing results from the sensor 103, at least one of these may be configured to be adjustable by a value arbitrarily input by the user 99. In other words, the sound processing device 101 may accept input from the user 99 to adjust at least one of the first angle, first delay time, and first volume attenuation.
[0081] [Operation] Next, with reference to Figure 6, the operation of the sound reproduction device 100 described above will be explained. Figure 6 is a flowchart showing the operation of the sound processing device according to the embodiment. First, when the operation of the sound reproduction device 100 is started, the acquisition unit 111 acquires original sound information via the communication module 102. Since the original sound information includes reverberation components in addition to the reproduced sound, sound information including the reproduced sound with the reverberation components reduced is generated by the reverberation suppression processing unit 120.
[0082] The first processing unit 121 generates a first output sound signal by convolving a first head-direction transfer function onto the sound information to localize the sound contained in the information as a sound arriving from a predetermined direction (S101). Next, the second processing unit 131 generates a second output sound signal by convolving a second head-direction transfer function onto the sound information to localize the sound contained in the information as a sound arriving from a first direction and having a first delay time greater than 0 and a first volume attenuation greater than 0 with respect to the reproduced sound perceived by the first output sound signal (S102).
[0083] Steps S101 and S102 described above may be performed in any order, or in parallel. The combiner 150 then synthesizes the generated first output sound signal and the second output sound signal and outputs the synthesized output sound signal (step S103). When the output sound signal thus output is played back by the driver 104, the reflected sound is superimposed on the played sound, and the user 99 perceives it as a three-dimensional sound. In particular, since only one reflected sound is generated, a large-scale computing device is not required, and effective three-dimensional sound can be perceived by the user 99.
[0084] [Examples] Figure 7 shows the implementation example Figure 8 illustrates an appropriate first angle related to the above. Figure 9 illustrates an appropriate first delay time related to the above.
[0085] Figure 7 shows the perceived distance to the sound image position (perceived distance) as perceived by the subjects when the first angle is rotated from 0 to 180 degrees, i.e., how far away the sound was perceived in a given direction. A larger perceived distance indicates a stronger sense of externality and a more effective perception of three-dimensional sound. Here, the conditions were set to a first volume attenuation of -3dB and a first delay time of 2.2ms. As shown in Figure 7, a high sense of externality was obtained by setting the first direction to 105 degrees or 120 degrees.
[0086] Figure 8 shows the perceived distance as perceived by subjects when the first delay time was varied from 0ms to 3.4ms. Here, the conditions were set to a first volume attenuation of -3dB and a first angle of 105 degrees. As shown in Figure 8, a high sense of external presence was obtained by setting the first delay time from 2.4ms to 2.8ms, and a sufficient sense of external presence was obtained by setting the first delay time from 1.8ms to 3.0ms. However, increasing the delay time leads to a deterioration of sound quality, so a relatively short first delay time is appropriate. Therefore, it is best to set the first delay time to 1.8ms to 2.4ms, for example, 2.2ms.
[0087] Figure 9 shows the perceived distance as perceived by subjects when the volume attenuation of the first volume attenuation was varied from -30 dB to 0 dB. Here, the conditions were set to a first delay time of 2.2 ms and a first angle of 105 degrees. As shown in Figure 9, a high sense of external presence was obtained by setting the volume attenuation of the first volume attenuation to -5 dB to -3 dB, and no further improvement in the sense of external presence was observed even when the volume attenuation was set to -3 dB or higher. Furthermore, since loud reflected sound is a factor in the degradation of sound quality, it is considered best to keep the volume attenuation as small as possible.
[0088] [Differentiation] Next, an acoustic processing apparatus according to a modified version of the embodiment described above will be described. In the modified version described below, the configuration is substantially the same as that of the embodiment described above, so the explanation here will be omitted by referring to the above description. Figure 10 is a block diagram showing the functional configuration of the acoustic playback apparatus according to a modified version of the embodiment. Figure 11 is a block diagram showing the detailed functional configuration of the acoustic processing apparatus according to a modified version of the embodiment. As shown in Figures 10 and 11, the acoustic playback apparatus 100a according to the modified version includes an acoustic processing apparatus 101a. Furthermore, the acoustic processing apparatus 101a differs from the configuration of the acoustic processing apparatus 101 according to the above embodiment in that it has a third processing unit 141.
[0089] The third processing unit 141 generates the output sound signal of the second reflected sound. The third processing unit 141 is a functional unit that generates the third output sound signal by convolving a third head-related transfer function to localize the sound contained in the information as a sound that arrives from a second direction having a second angle greater than 0 degrees and less than 360 degrees with respect to a predetermined direction, and having a second delay time greater than 0 with respect to the reproduced sound perceived by the first output sound signal, and a second volume attenuation greater than 0, and a second volume attenuation different from the first volume attenuation. The third processing unit 141 convolves a head-related transfer function to localize the sound in the second direction with respect to the input sound information, and calculates the volume attenuation γ (third 2 The attenuated third output sound signal is output via volume attenuation. This processing by the third processing unit 141 is collectively understood as the convolution of the third head-of-function. The third output sound signal is input to the third EQ 142, where the low and high frequencies are adjusted before being sent to the combiner 150. Before being input to the third processing unit 141, the sound information is supplemented by the second angle determination unit 140 with information specifying the head-of-function to be convolved thereafter.
[0090] The combiner 150 is a functional unit that synthesizes the generated output sound signals and outputs them to the driver 104. The combiner 150 outputs a synthesized output sound signal by adding the first output sound signal, the second output sound signal, and the third output sound signal. In other words, the sound processing device 101a generates two different reflected sounds in the second processing unit 131 and the third processing unit 141, and the combiner 150 superimposes these onto the reproduced sound. As in this modified example, when two reflected sounds are generated and superimposed onto the reproduced sound, it is possible to further improve the effect of unfolding them into a three-dimensional sound field depending on the conditions.
[0091] As described above, when the output sound signal from the combiner 150 is reproduced by the driver 104, a sound field like that shown in Figure 12 is formed. Figure 12 is a diagram illustrating the direction of arrival of sound output by the sound processing device according to this embodiment. Figure 12 shows a plan view of a virtual three-dimensional sound field from the same viewpoint as in Figure 5.
[0092] As shown in the figure, the first reflected sound is localized in a direction having a first angle clockwise from a predetermined direction (position P2). The second reflected sound is localized in a direction having a second angle clockwise from the predetermined direction (position P3). As shown in the figure, the first and second angles do not coincide, and they are not symmetrical with respect to the dashed line parallel to the front rear of the user 99 (and also parallel to the predetermined direction). If the first and second directions were symmetrical, under certain conditions, the two reflected sounds might be superimposed and localized as a single reflected sound behind the user 99. Therefore, the second angle is such that the difference angle obtained by subtracting the second angle from 360 degrees does not coincide with the first angle.
[0093] Furthermore, as shown in the figure, the first and second angles are positioned on the rear side of the user 99, relative to the virtual boundary that divides the user 99's head into front and rear halves. Therefore, both the first and second angles are set to angles within a range greater than 90 degrees and less than 270 degrees relative to a predetermined direction.
[0094] The second angle, second delay time, and second volume attenuation described above are either values preset by the sound processing device 101a or values that change according to the sensing results from the sensor 103, similar to the first angle, first delay time, and first volume attenuation. However, at least one of these may be configured to be adjustable by a value arbitrarily input by the user 99. In other words, the sound processing device 101a may accept input from the user 99 to adjust at least one of the second angle, second delay time, and second volume attenuation.
[0095] [Operation] Next, the operation of the sound reproduction device 100a described above will be explained with reference to Figure 13. Figure 13 is a flowchart showing the operation of the sound processing device according to the embodiment. First, steps S101 and S102 are performed, similar to the operation of the sound processing device 101 described with reference to Figure 6. Next, the third processing unit 141 generates a third output sound signal (S201) by convolving the sound information with a third head transfer function that localizes the sound contained in the information as a sound that arrives from a second direction and has a second delay time greater than 0 and a second volume attenuation greater than 0 with respect to the reproduced sound perceived by the first output sound signal.
[0096] Steps S101, S102, and S201 described above may be performed in any order, or in parallel. The combiner 150 then synthesizes the generated first output sound signal, second output sound signal, and third output sound signal, and outputs the synthesized output sound signal (step S202). When the output sound signal thus output is played back by the driver 104, the reflected sound is superimposed on the played sound, and the user 99 perceives it as a three-dimensional sound. In particular, since only two reflected sounds are generated, a large-scale computing device is not required in this case, and effective three-dimensional sound can be perceived by the user 99.
[0097] Furthermore, the processing unit may be increased to superimpose three or more reflected sounds onto the reproduced sound.
[0098] (Other embodiments) Although embodiments have been described above, this disclosure is not limited to the embodiments described above.
[0099] For example, the above embodiment describes an example where the sound does not follow the user's head movements, but the contents of this disclosure are also valid when the sound follows the user's head movements. That is, in an operation in which a predetermined sound is perceived by the user as a sound arriving from a first position that moves relatively with the user's head movements, a spatial sound filter may be selected to emphasize the fluctuation when the amount of variation in the arrival direction of the predetermined sound is smaller than a threshold.
[0100] Furthermore, for example, the sound reproduction device described in the above embodiment may be realized as a single device comprising all the components, or it may be realized by assigning each function to multiple devices and having these multiple devices cooperate. In the latter case, an information processing device such as a smartphone, tablet terminal, or PC may be used as the device corresponding to the processing module.
[0101] As a configuration different from the above-described embodiment, for example, the decode processing unit can correct the original sound information to select a modified stereophonic filter. Specifically, the decode processing unit in this example generates information about a predetermined direction included in the sound information and corrects the original sound information. The decode processing unit calculates the amount of angle of variation in the predetermined direction on the time axis, and corrects the information about the predetermined direction so that, when the calculated amount of angle of variation in the predetermined direction is smaller than a threshold, the predetermined sound is perceived by the user as being more emphasized compared to when the amount of angle of variation in the predetermined direction is greater than or equal to the threshold. As a result, the modified stereophonic filter in the above-described embodiment is applied simply by selecting a stereophonic filter that defines the direction to which the predetermined sound arrives based on the corrected information about the predetermined direction output from the decode processing unit.
[0102] Thus, the information processing method disclosed in this application may be realized by correcting the information regarding a predetermined direction in the original sound information. The above-described decoding processing unit can be used to realize an audio playback device that can achieve the same effects as disclosed in this application simply by replacing, for example, the decoding processing unit of a conventional stereophonic sound playback device with the above-described decoding processing unit.
[0103] Furthermore, the sound playback device of this disclosure can also be implemented as a sound processing device that is connected to a playback device equipped only with a driver and outputs an output sound signal to the playback device using a stereophonic sound filter selected based on the acquired sound information. In this case, the sound processing device may be implemented as hardware equipped with a dedicated circuit, or as software that causes a general-purpose processor to perform specific processing.
[0104] Furthermore, in the above embodiment, the processing performed by a specific processing unit may be performed by another processing unit. Also, the order of multiple processing units may be changed, or multiple processing units may be executed in parallel.
[0105] Furthermore, in the above embodiment, each component may be realized by executing a software program suitable for each component. Each component may also be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0106] Furthermore, each component may be implemented by hardware. For example, each component may be a circuit (or integrated circuit). These circuits may form a single circuit as a whole, or they may be separate circuits. Also, each of these circuits may be a general-purpose circuit or a dedicated circuit.
[0107] Furthermore, the general or specific aspects of this disclosure are: systemThese may be implemented as devices, methods, integrated circuits, computer programs, or recording media such as computer-readable CD-ROMs. Furthermore, the general or specific embodiments of this disclosure may be implemented as any combination of devices, methods, integrated circuits, computer programs, and recording media.
[0108] For example, the present disclosure may be implemented as a method for reproducing an audio signal executed by a computer, or as a program for causing a computer to execute an audio signal reproduction method. The present disclosure may also be implemented as a computer-readable non-temporary recording medium on which such a program is recorded.
[0109] Furthermore, this disclosure also includes forms obtained by applying various modifications to each embodiment that a person skilled in the art could conceive, or forms realized by arbitrarily combining the components and functions of each embodiment without departing from the spirit of this disclosure. [Industrial applicability]
[0110] This disclosure is useful for sound reproduction, such as making users perceive three-dimensional sound. [Explanation of Symbols]
[0111] 99 users 100, 100a Sound reproduction device 101, 101a Acoustic processing device 102 Communication Module 103 Sensor 104 Driver 111 Acquisition Department 120 Reverberation suppression processing unit 121 First Processing Unit 122 1st EQ 130 First Angle Determination Unit 131 Second Processing Unit 132 2nd EQ 140 Second Angle Determination Section 141 Third Processing Unit 142 3rd EQ 150 Combine Na
Claims
1. An acoustic processing device that causes the user to perceive the reproduced sound as sound arriving from a predetermined direction in a three-dimensional sound field, A first processing unit generates a first output sound signal by convolving a first head-direction transfer function with respect to sound information including the reproduced sound, which localizes the sound contained in the information as sound arriving from a predetermined direction. A second processing unit generates a second output sound signal by convolving a second head-related transfer function onto the aforementioned sound information, which causes the sound contained in the information to arrive from a first direction having a first angle greater than 0 degrees and less than 360 degrees with respect to the predetermined direction, and to localize the reproduced sound perceived by the first output sound signal as a sound having a first delay time greater than 0 and a first volume attenuation greater than 0. The system includes a combiner that outputs an output sound signal obtained by combining the generated first output sound signal and the second output sound signal, Acoustic processing device.
2. The output sound signal is played back using headphones or earphones worn on the user's head. The acoustic apparatus according to claim 1.
3. The first angle is an angle within an angular range greater than 90 degrees and less than 270 degrees with respect to the predetermined direction. The acoustic apparatus according to claim 1.
4. Furthermore, the system includes a third processing unit that generates a third output sound signal by convolving the sound information with a third head-related transfer function that causes the sound contained in the information to arrive from a second direction having a second angle greater than 0 degrees and less than 360 degrees with respect to the predetermined direction, and which is different from the first angle, and to localize the reproduced sound perceived by the first output sound signal as a sound having a second delay time greater than 0 and a second volume attenuation greater than 0. The combiner outputs an output sound signal obtained by combining the first output sound signal, the second output sound signal, and the third output sound signal. The acoustic apparatus according to claim 1.
5. The second angle is an angle within an angle range greater than 90 degrees and less than 270 degrees with respect to the predetermined direction, and is an angle in which the difference angle obtained by subtracting the second angle from 360 degrees does not coincide with the first angle. The acoustic apparatus according to claim 4.
6. The first delay time and the second delay time are different delay times. The acoustic apparatus according to claim 4.
7. The first volume attenuation and the second volume attenuation are each different amounts of volume attenuation. The acoustic apparatus according to claim 4.
8. Furthermore, it includes a reverberation suppression processing unit that performs reverberation suppression processing on the sound contained in the information to reduce the reverberation components contained in the information. The aforementioned information is, The reverberation suppression process is performed on the original sound information including the reverberation component to generate the following: The reproduced sound includes, among the sounds contained in the original sound information, sounds other than the reduced reverberation component. The acoustic apparatus according to claim 1.
9. Furthermore, it includes an acquisition unit that acquires sensing results from a sensor that detects the user's head movements, The second processing unit convolves the second head-level transfer function, obtained by changing the volume attenuation amount of the first volume attenuation, onto the sound information based on the acquired sensing results. The acoustic apparatus according to claim 1.
10. The first head-related transfer function, by being convolved, localizes the sound contained in the information as a sound that arrives from a predetermined direction and has a third volume attenuation of 0 or more. The first processing unit is, If the volume attenuation amount of the first volume attenuation in the second processing unit increases, the first head-level transfer function obtained by reducing the volume attenuation amount of the third volume attenuation is convolved with the sound information. If the volume attenuation amount of the first volume attenuation in the second processing unit decreases, the first head-level transfer function with an increased volume attenuation amount of the third volume attenuation is convolved with the sound information. The acoustic apparatus according to claim 9.
11. Furthermore, it includes an acquisition unit that acquires sensing results from a sensor that detects the user's head movements, The third processing unit convolves the third head-level transfer function, obtained by changing the volume attenuation amount of the second volume attenuation based on the acquired sensing results, onto the sound information. The acoustic apparatus according to claim 5.
12. At least one of the first angle, the first delay time, and the first volume attenuation is adjusted by the user. The acoustic apparatus according to claim 1.
13. At least one of the second angle, the second delay time, and the second volume attenuation is adjusted by the user. The acoustic apparatus according to claim 5.
14. The aforementioned sound information is generated based on the original sound information, which includes the reproduced sound and reverberation components. The first delay time is a delay time smaller than the delay time of the reverberation component relative to the reproduced sound. The acoustic apparatus according to claim 1.
15. The aforementioned sound information is generated based on the original sound information, which includes the reproduced sound and reverberation components. The second delay time is a delay time smaller than the delay time of the reverberation component relative to the reproduced sound. The acoustic apparatus according to any one of claims 5 to 7, 11, or 13.
16. An acoustic processing method that causes the user to perceive the reproduced sound as sound arriving from a predetermined direction in a three-dimensional sound field, A first output sound signal is generated by convolving a first head-direction transfer function, which localizes the sound contained in the information as sound arriving from a predetermined direction, onto the sound information including the reproduced sound. A second output sound signal is generated by convolving a second head-direction transfer function with respect to the aforementioned sound information, which causes the sound contained in the information to arrive from a first direction having a first angle greater than 0 degrees and less than 360 degrees with respect to the predetermined direction, and to localize the reproduced sound perceived by the first output sound signal as a sound having a first delay time greater than 0 and a first volume attenuation greater than 0. The system outputs an output sound signal obtained by combining the generated first output sound signal and the second output sound signal. Sound processing methods.
17. To cause a computer to execute the sound processing method described in claim 16. program.
Citation Information
Patent Citations
Voice generation program in virtual space, generation method of quadtree, and voice generation device
JP2020018620A
Virtual sound source device and acoustic device comprising the same
WO2000045619A1
Acoustic reproduction method, program, and acoustic reproduction system
WO2021187147A1