Acoustic reproduction method, program, and acoustic reproduction system
By generating output sound signals that localize sounds from a third position when the user's head movement exceeds a threshold, the acoustic reproduction method and system enhance computational efficiency and maintain a high sense of presence for three-dimensional sound perception.
Patent Information
- Application Number
- JP2022508208
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-17
- Filing Date
- 2021-03-04
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-03-04
AI Technical Summary
Conventional acoustic reproduction methods require significant computational processing to generate three-dimensional sound, often leading to inadequate processing and a compromised sense of presence for the user.
An acoustic reproduction method and system that acquire the movement speed of the user's head and generate an output sound signal to perceive sounds as originating from a third position between two original positions, reducing computational load by unifying sound localization processes when the head movement speed exceeds a threshold.
This approach allows users to perceive three-dimensional sound with improved computational efficiency, reducing processing load and maintaining a high sense of presence, even during rapid head movements.
Smart Images

Figure 0007692402000001 
Figure 0007692402000002 
Figure 0007692402000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an acoustic reproduction system and an acoustic reproduction method.
Background Art
[0002] Conventionally, there has been known a technique related to acoustic reproduction for causing a user to perceive three-dimensional sound by controlling the position of a sound image, which is a virtual sound source object, in a virtual three-dimensional space (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] On the other hand, when generating sound for causing a user to perceive three-dimensional sound, enormous computational processing is required. Here, in conventional acoustic reproduction methods and the like, appropriate computational processing may not be performed.
[0005] In view of the above, an object of the present disclosure is to provide an acoustic reproduction method and the like that cause a user to perceive three-dimensional sound by more appropriate computational processing.
Means for Solving the Problems
[0006] An acoustic reproduction method according to one aspect of the present disclosure is an acoustic reproduction method for causing a user to perceive a first sound as a sound reaching from a first position on a three-dimensional sound field and causing the user to perceive a second sound as a sound reaching from a second position different from the first position, the method including an acquisition step of acquiring a movement speed of the user's head and a generation step of generating an output sound signal for causing the user to perceive a sound reaching from a predetermined position on the three-dimensional sound field, wherein in the generation step, when the acquired movement speed is greater than a first threshold value, the output sound signal is generated for causing the user to perceive the first sound and the second sound as a sound reaching from a third position between the first position and the second position.
[0007] An acoustic reproduction system according to one aspect of the present disclosure is an acoustic reproduction system for causing a user to perceive a first sound as a sound reaching from a first position on a three-dimensional sound field and causing the user to perceive a second sound as a sound reaching from a second position different from the first position, the system including an acquisition unit that acquires a movement speed of the user's head and a generation unit that generates an output sound signal for causing the user to perceive a sound reaching from a predetermined position on the three-dimensional sound field, wherein the generation unit generates the output sound signal for causing the user to perceive the first sound and the second sound as a sound reaching from a third position between the first position and the second position when the acquired movement speed is greater than a first threshold value.
[0008] One aspect of the present disclosure can also be realized as a program for causing a computer to execute the acoustic reproduction method described above.
[0009] These general or specific aspects may be implemented by a system, apparatus, method, integrated circuit, computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, apparatus, method, integrated circuit, computer program, and recording medium.
Advantages of the Invention
[0010] According to the present disclosure, it is possible to make a user perceive three-dimensional sound through more appropriate calculation processing.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 6C
Modes for Carrying Out the Invention
[0012] (Findings on which the disclosure is based) Conventionally, there has been known a technique related to audio reproduction for causing a user to perceive three-dimensional sound by controlling the position of a sound image, which is a sound source object in the user's perception, in a virtual three-dimensional space (hereinafter sometimes referred to as a three-dimensional sound field) (see, for example, Patent Document 1). By localizing the sound image at a predetermined position in the virtual three-dimensional space, the user can perceive this sound as if it were a sound emitted from the predetermined position. In order to localize the sound image at a predetermined position in such a virtual three-dimensional space, for example, calculation processing for generating a time difference of arrival of sound between both ears and a level difference of sound between both ears, which are perceived as three-dimensional sound, for the recorded sound is required.
[0013] As an example of such calculation processing, a process of convolving a head-related transfer function for causing a sound arriving from a predetermined position to be perceived with a signal of the target sound is known. By performing the convolution process of this head-related transfer function with higher resolution, the sense of presence experienced by the user is improved. On the other hand, the convolution of the head-related transfer function is relatively computationally intensive as a calculation process, and resources for calculation are required. That is, in order to perform the process of convolving the head-related transfer function with high resolution, a high-performance computing device and power associated with the use of the computing device are required.
[0014] In recent years, the development of technologies related to virtual reality (VR) has been actively carried out. In virtual reality, the main focus is on the ability to feel as if the user is moving within the virtual space without the position of the virtual three-dimensional space following the user's movement. In particular, attempts have been made to enhance the sense of presence by incorporating auditory elements into visual elements in this virtual reality technology. For example, when a sound image is localized in front of the user, when the user turns right, the sound image moves to the left direction of the user, and when the user turns left, the sound image moves to the right direction of the user. Thus, it becomes necessary to move the localization position of the sound image in the virtual space in the direction opposite to the user's movement with respect to the user's movement.
[0015] In order to improve the sense of presence in a virtual space, it is required to increase the spatial resolution and perform the convolution process of the head-related transfer function. Therefore, in order to perform acoustic reproduction that allows a user to perceive three-dimensional sound with a high sense of presence, such as in the above virtual reality, constraints such as those on the computing device and power consumption become more prominent.
[0016] Therefore, in view of the above, the present disclosure performs more appropriate computational processing by suppressing a decrease in the sense of presence and reducing the load of computational processing. An object of the present disclosure is to provide an acoustic reproduction method or the like that allows a user to perceive three-dimensional sound by this appropriate computational processing.
[0017] More specifically, an acoustic reproduction method according to one aspect of the present disclosure is an acoustic reproduction method that causes a user to perceive a first sound as a sound arriving from a first position on a three-dimensional sound field and causes the user to perceive a second sound as a sound arriving from a second position different from the first position, the method including: an acquisition step of acquiring a movement speed of the user's head; and a generation step of generating an output sound signal for causing the user to perceive a sound arriving from a predetermined position on the three-dimensional sound field, wherein in the generation step, when the acquired movement speed is greater than a first threshold value, the output sound signal for causing the user to perceive the first sound and the second sound as sounds arriving from a third position between the first position and the second position is generated.
[0018] According to such an acoustic playback method, when the movement speed of the user's head is greater than the first threshold value, the first sound perceived as the sound reaching from the first position and the second sound perceived as the sound reaching from the second position can be made to be perceived as the sound reaching from the third position. At this time, the process for localizing the sound image of the first sound to the first position and the process for localizing the sound image of the second sound to the second position can both be unified into the process for localizing to the third position, so that the processing amount can be reduced. Also, here, if the first threshold value is set to a value such that when the movement speed of the user's head exceeds this value, the perception of the sound image position of the user becomes ambiguous, even if the above processing is performed, the influence on the sense of presence due to the change in the sound image position is suppressed. Thereby, it is also possible to reduce the discomfort of the user that may occur by reducing the processing amount. Therefore, it becomes possible to make the user perceive three-dimensional sound by more appropriate calculation processing.
[0019] Further, for example, in the generation step, when the acquired movement speed is less than or equal to the first threshold value, the first head-related transfer function for localizing the sound to the first position is convolved with the first sound signal related to the first sound, and the second head-related transfer function for localizing the sound to the second position is convolved with the second sound signal related to the second sound to generate the output sound signal. When the acquired movement speed is greater than the first threshold value, the output sound signal may be generated by convolving the third head-related transfer function for localizing the sound to the third position with the added sound signal obtained by adding the first sound signal and the second sound signal.
[0020] When localizing the sound image of the first sound to the first position, convolve the first head-related transfer function with the first sound signal related to the first sound. When localizing the sound image of the second sound to the second position, convolve the second head-related transfer function with the second sound signal related to the second sound. According to the above, when localizing the sound images of the first sound and the second sound to the third position, it is only necessary to perform a process of convolving the third head-related transfer function for localizing the sound to the third position with the added sound signal obtained by adding the first sound signal and the second sound signal. That is, the convolution process of the first head-related transfer function with respect to the first sound signal and the convolution process of the second head-related transfer function with respect to the second sound signal can be made common to the convolution process of the third head-related transfer function with respect to the added sound signal. Therefore, the processing amount can be reduced, and it becomes possible to make the user perceive three-dimensional sound by more appropriate calculation processing.
[0021] Also, for example, the movement speed is the rotational speed of the user's head around the first axis passing through the user's head, and the third position may be a position on the bisector that bisects the angle formed by the straight lines connecting the first position and the second position to the user, respectively, in a virtual plane when viewing the three-dimensional sound field from the direction of the first axis.
[0022] According to this, the set third position can be used in correspondence with the rotational movement of the user's head. At this time, the third position is set at a position on the bisector that bisects the angle formed by the straight lines connecting the first position and the second position to the user, respectively, in a virtual plane when viewing the three-dimensional sound field from the direction of the first axis, which is the rotation axis. Therefore, the third position can be set in the direction between the direction of the first position and the direction of the second position as seen from the user, in accordance with the arrival direction of the sound that becomes ambiguous due to the rotational movement of the user. Thus, it becomes possible to make the user perceive three-dimensional sound while suppressing the discomfort in the arrival direction of the sound while reducing the processing amount.
[0023] Also, for example, the rotational speed may be obtained as the rotational amount per unit time detected by a detector that moves integrally with the user's head and detects the rotational amount around at least one of three mutually orthogonal axes as the rotation axis.
[0024] According to this, as the movement speed, the rotation speed of the user's head can be obtained using a detector. Therefore, based on the rotation speed obtained as described above, it is possible to suppress the sense of incongruity in the sound arrival direction and let the user perceive three-dimensional sound.
[0025] Also, for example, the movement speed is the displacement speed of the user's head along the second axis direction passing through the user's head, and the displacement speed is obtained as the displacement amount per unit time detected by a detector that moves integrally with the user's head and detects a displacement amount having at least one of three mutually orthogonal axes as the displacement direction.
[0026] The set third position can be used corresponding to the movement of the displacement of the user's head. At this time, the displacement speed of the user's head can be obtained using a detector. Therefore, based on the displacement speed obtained as described above, it is possible to suppress the sense of incongruity in the sound arrival direction and let the user perceive three-dimensional sound.
[0027] Also, for example, in the acoustic reproduction method, a plurality of sounds reaching from each position within a predetermined region on the three-dimensional sound field including the first position and the second position, and including at least the first sound and the second sound, are made to be perceived by the user, and in the generation step, when the movement speed is greater than the first threshold value, an output sound signal for making all of the plurality of sounds be perceived by the user as sounds reaching from the third position may be generated.
[0028] According to this, all of the plurality of sounds within a predetermined range can be made to be perceived by the user as sounds reaching from the third position. For this reason, the head-related transfer functions convolved with each of the sounds within the predetermined range can be made common by the head-related transfer function for localizing the sound image at the third position. Therefore, the processing amount of the convolution of the head-related transfer functions is reduced, and it becomes possible to let the user perceive three-dimensional sound by more appropriate computational processing.
[0029] Further, for example, in the acoustic playback method, the user is made to perceive a first intermediate sound as a sound reaching from a first intermediate position between the first position and the third position, and the user is made to perceive a second intermediate sound as a sound from a second intermediate position between the second position and the third position. In the generation step, further, when the movement speed is equal to or less than the first threshold value and greater than a second threshold value smaller than the first threshold value, the output sound signal for making the user perceive the first intermediate sound and the second intermediate sound as sounds reaching from the third position may be generated.
[0030] According to this, the same processing as described above can be applied within a narrow range including a first intermediate position and a second intermediate position closer to the third position than the first position and the second position respectively. Here, since the movement speed of the user's head is smaller than the first threshold value, aggregating the sounds at the first position and the second position etc. to the third position may cause a sense of discomfort because the change in the sound image position can be perceived. On the other hand, since the movement speed of the user's head is greater than the second threshold value, even if the sounds within a narrow range narrower than a predetermined range including the first position and the second position etc. are aggregated to the third position, the change in the sound image position is not perceived. Therefore, when the movement speed is equal to or less than the first threshold value and greater than the second threshold value smaller than the first threshold value, the sounds at the first intermediate position and the second intermediate position included within such a narrow range can be aggregated to the third position to reduce the processing amount of the calculation process. Thus, it becomes possible to make the user perceive three-dimensional sound by more appropriate calculation processing.
[0031] Also, an acoustic reproduction system according to one aspect of the present disclosure is an acoustic reproduction system that causes a user to perceive a first sound as a sound reaching from a first position on a three-dimensional sound field and causes the user to perceive a second sound as a sound reaching from a second position different from the first position, the acoustic reproduction system including: an acquisition unit that acquires a movement speed of the user's head; and a generation unit that generates an output sound signal for causing the user to perceive a sound reaching from a predetermined position on the three-dimensional sound field, wherein when the acquired movement speed is greater than a first threshold value, the generation unit generates the output sound signal for causing the user to perceive the first sound and the second sound as sounds reaching from a third position between the first position and the second position.
[0032] According to this, an acoustic reproduction system that exhibits the same effect as the acoustic reproduction method described above can be realized.
[0033] Also, one aspect of the present disclosure can be realized as a program for causing a computer to execute the acoustic reproduction method described above.
[0034] According to this, the same effect as the acoustic reproduction method described above can be achieved using a computer.
[0035] Furthermore, these general or specific aspects may be realized by a system, apparatus, method, integrated circuit, computer program, or non-transitory recording medium such as a computer-readable CD-ROM, or may be realized by any combination of a system, apparatus, method, integrated circuit, computer program, and recording medium.
[0036] Hereinafter, embodiments will be specifically described with reference to the drawings. Note that all the embodiments described below show comprehensive or specific examples. Numerical values, shapes, materials, components, arrangement positions and connection forms of components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. In addition, among the components in the following embodiments, components not described in the independent claims are described as optional components. Note that each drawing is a schematic diagram and is not necessarily drawn precisely. Also, in each drawing, substantially the same configuration is denoted by the same reference numeral, and duplicate descriptions may be omitted or simplified.
[0037] In addition, in the following description, ordinal numbers such as first, second, and third may be attached to elements. These ordinal numbers are attached to the elements to identify them and do not necessarily correspond to a meaningful order. These ordinal numbers may be appropriately interchanged, newly assigned, or removed.
[0038] (Embodiment) [Overview] First, an overview of the audio reproduction system according to the embodiment will be described. FIG. 1 is a schematic diagram showing a usage example of the audio reproduction system according to the embodiment. In FIG. 1, a user 99 who uses the audio reproduction system 100 is shown.
[0039] The audio reproduction system 100 shown in FIG. 1 is used simultaneously with the stereoscopic video reproduction system 200. As described above, in this embodiment, by simultaneously viewing a stereoscopic image and stereoscopic sound, the image enhances the auditory sense of presence, and the sound enhances the visual sense of presence, so that the user can feel as if they are at the scene where the image and sound were captured. For example, when an image (moving image) of a person talking is displayed, it is known that the user 99 can perceive the conversation sound emitted from the mouth of the person even when the sound localization of the conversation sound is deviated from the mouth of the person. In this way, the sense of presence may be enhanced by combining the image and the sound, such as correcting the position of the sound image by visual information.
[0040] The stereoscopic video playback system 200 is an image display device worn on the head of the user 99. Therefore, the stereoscopic video playback system 200 moves integrally with the head of the user 99. For example, as shown in the figure, the stereoscopic video playback system 200 is a glasses-type device supported by the ears and nose of the user 99.
[0041] The stereoscopic video playback system 200 changes the image to be displayed according to the movement of the head of the user 99, so that the user 99 is perceived as moving the head in the three-dimensional image space. That is, when an object in the three-dimensional image space is located in front of the user 99, when the user 99 turns to the right, the object moves to the left direction of the user 99, and when the user 99 turns to the left, the object moves to the right direction of the user. In this way, the stereoscopic video playback system 200 moves the three-dimensional image space in the direction opposite to the movement of the user 99 with respect to the movement of the user 99.
[0042] The stereoscopic video playback system 200 displays two images with a disparity shift in the left and right eyes of the user 99, respectively. The user 99 can perceive the three-dimensional position of the object on the image based on the disparity shift of the displayed images. Note that when the user 99 uses it with eyes closed, such as using the audio playback system 100 to play soothing sounds for sleep induction, etc., it is not necessary to use the stereoscopic video playback system 200 at the same time. That is, the stereoscopic video playback system 200 is not an essential component of the present disclosure.
[0043] The audio playback system 100 is a sound presentation device worn on the head of the user 99. Therefore, the audio playback system 100 moves integrally with the head of the user 99. For example, the audio playback system 100 is two earplug-type devices independently worn on the left and right ears of the user 99. These two devices communicate with each other to present the sound for the right ear and the sound for the left ear synchronously.
[0044] The audio reproduction system 100 makes the user 99 perceive as if the user 99 is moving their head within the three-dimensional sound field by changing the sound presented according to the movement of the user 99's head. For this reason, as described above, the audio reproduction system 100 moves the three-dimensional sound field in the direction opposite to the movement of the user with respect to the movement of the user 99.
[0045] Here, when the movement of the user 99's head exceeds a certain level, it is known that the discrimination of the position of the sound image within the three-dimensional sound field becomes ambiguous for the user 99. The audio reproduction system 100 according to the present embodiment reduces the load of the calculation process by utilizing this phenomenon. That is, the audio reproduction system 100 acquires the movement speed of the user 99's head, and when the acquired movement speed is greater than the first threshold value, causes a plurality of sounds perceived as sounds arriving from within a predetermined region on the three-dimensional sound field to be perceived as sounds arriving from one location within the predetermined region.
[0046] This predetermined region corresponds to a range in which the perception of the sound image position by the user 99 becomes ambiguous due to the high movement speed of the head. Therefore, since it needs to be set for each user 99, for example, it may be set by conducting an experiment or the like in advance. Further, since the predetermined region is also affected by the amount of movement of the user 99's head, a predetermined region corresponding to the amount of movement may be set by detecting the amount of movement of the user 99's head.
[0047] Similarly, for the first threshold value with respect to the movement speed, a user-specific numerical setting is required as to from what movement speed the perception of the sound image position by the user 99 becomes ambiguous. Therefore, a value set by conducting an experiment or the like in advance may be adopted. Note that a generalized predetermined region and first threshold value may be set by averaging from the experimental results of a plurality of users 99.
[0048] [Configuration] Next, with reference to FIG. 2, the configuration of the audio reproduction system 100 according to the present embodiment will be described. FIG. 2 is a block diagram showing the functional configuration of the audio reproduction system according to the embodiment.
[0049] As shown in FIG. 2, the acoustic reproduction system 100 according to the present embodiment includes a processing module 101, a communication module 102, a detector 103, and a driver 104.
[0050] The processing module 101 is an arithmetic device for performing various signal processes in the acoustic reproduction system 100. The processing module 101 includes, for example, a processor and a memory, and various functions are exhibited by the program stored in the memory being executed by the processor.
[0051] The processing module 101 has an input unit 111, an acquisition unit 121, a generation unit 131, and an output unit 141. Details of each functional unit included in the processing module 101 will be described below together with details of other configurations of the processing module 101.
[0052] The communication module 102 is an interface device for receiving an input of an audio signal to the acoustic reproduction system 100. The communication module 102 includes, for example, an antenna and a signal converter, and receives an audio signal from an external device by wireless communication. More specifically, the communication module 102 receives a radio signal indicating an audio signal converted into a format for wireless communication using an antenna, and reconverts the radio signal into an audio signal by the signal converter. Thereby, the acoustic reproduction system 100 acquires an audio signal from an external device by wireless communication. The audio signal acquired by the communication module 102 is input to the input unit 111. In this way, the audio signal is input to the processing module 101. Note that communication between the acoustic reproduction system 100 and an external device may be performed by wired communication.
[0053] The audio signal acquired by the audio reproduction system 100 is encoded in a predetermined format such as MPEG-H Audio, for example. As an example, the encoded audio signal includes information about the sound reproduced by the audio reproduction system 100 and information about the localization position when localizing the sound image of the sound at a predetermined position in the three-dimensional sound field. For example, the audio signal includes information about a plurality of sounds including a first sound and a second sound, and the sound images when each sound is reproduced are localized at different positions in the three-dimensional sound field.
[0054] With this three-dimensional sound, for example, the sense of presence of content to be viewed, such as in combination with an image viewed using the stereoscopic video reproduction system 200, can be improved. Note that the audio signal may include only information about the sound. In this case, the information about the localization position may be acquired separately. Also, as described above, the audio signal includes a first audio signal related to the first sound and a second audio signal related to the second sound, but a plurality of audio signals including these separately may be acquired and reproduced simultaneously to localize the sound image at different positions in the three-dimensional sound field. Thus, the form of the input audio signal is not particularly limited, and the audio reproduction system 100 may be provided with an input unit 111 corresponding to various forms of audio signals.
[0055] The detector 103 is a device for detecting the movement speed of the head of the user 99. The detector 103 is configured by combining various sensors used for detecting movement, such as a gyro sensor and an acceleration sensor. In the present embodiment, the detector 103 is built into the audio reproduction system 100, but may be built into an external device such as the stereoscopic video reproduction system 200 that operates according to the movement of the head of the user 99 in the same manner as the audio reproduction system 100. In this case, the detector 103 may not be included in the audio reproduction system 100. Also, as the detector 103, an external imaging device or the like may be used to image the movement of the head of the user 99 and detect the movement of the user 99 by processing the captured image.
[0056] The detector 103 is, for example, fixedly integrated with the housing of the audio playback system 100 and detects the speed of movement of the housing. Since the audio playback system 100 moves integrally with the head of the user 99 after the user 99 wears it, as a result, the speed of movement of the head of the user 99 can be detected.
[0057] The detector 103 may, for example, detect, as the amount of movement of the head of the user 99, the amount of rotation having at least one of three axes orthogonal to each other in a three-dimensional space as the rotation axis, or may detect the amount of displacement having at least one of the three axes as the displacement direction. Further, the detector 103 may detect both the amount of rotation and the amount of displacement as the amount of movement of the head of the user 99.
[0058] The acquisition unit 121 acquires the movement speed of the head of the user 99 from the detector 103. More specifically, the acquisition unit 121 acquires, as the movement speed, the amount of movement of the head of the user 99 detected by the detector 103 per unit time. In this way, the acquisition unit 121 acquires at least one of the rotation speed and the displacement speed from the detector 103.
[0059] Here, the generation unit 131 determines whether or not the acquired movement speed of the head of the user 99 is greater than the above-described first threshold value. The generation unit 131 determines whether or not to reduce the load amount of the calculation process based on the result of this determination. More detailed operations of the generation unit 131 will be described later. The generation unit 131 performs a calculation process on the input audio signal according to the above-described determination content and generates an output audio signal for presenting sound.
[0060] The output unit 141 is a functional unit that outputs the generated output sound signal to the driver 104. The driver 104 performs signal conversion from a digital signal to an analog signal based on the output sound signal, generates a waveform signal, generates a sound wave based on the waveform signal, and presents sound to the user 99. The driver 104 has, for example, a driving mechanism such as a diaphragm, a magnet, and a voice coil. The driver 104 operates the driving mechanism according to the waveform signal, and vibrates the diaphragm by the driving mechanism. In this way, the driver 104 generates a sound wave due to the vibration of the diaphragm according to the output sound signal, the sound wave propagates through the air and is transmitted to the ear of the user 99, and the user 99 perceives the sound.
[0061] [Operation] Next, with reference to FIG. 3, the operation of the acoustic reproduction system 100 described above will be described. FIG. 3 is a flowchart showing the operation of the acoustic reproduction system according to the embodiment. As shown in FIG. 3, first, when the operation of the acoustic reproduction system 100 is started, a first sound signal related to the first sound and a second sound signal related to the second sound are acquired (step S101). Here, the sound signal acquired by the communication module 102 from an external device is input to the input unit 111, and the processing module 101 acquires the sound signal including the first sound signal and the second sound signal.
[0062] Subsequently, acquisition unit 121 acquires, as a detection result from detector 103, the movement speed of user 99's head (acquisition step S102). Generation unit 131 compares the acquired movement speed with a first threshold value to determine whether the movement speed is greater than the first threshold value (step S103). When the movement speed is less than or equal to the first threshold value (No in step S103), acoustic reproduction system 100 causes user 99 to perceive the first sound and the second sound as sounds arriving from the first position and the second position, which are their original sound image positions. For this purpose, generation unit 131 convolves the first sound signal with a first head-related transfer function for localizing the sound image at the first position. Also, generation unit 131 convolves the second sound signal with a second head-related transfer function for localizing the sound image at the second position (step S104). Generation unit 131 generates an output sound signal including the first sound signal and the second sound signal that have been subjected to the convolution process in this way (step S105).
[0063] On the other hand, when the movement speed is greater than the first threshold value (Yes in step S103), acoustic reproduction system 100 causes user 99 to perceive the first sound and the second sound as sounds arriving from a third position between these positions with respect to the first position and the second position, which are the original sound image positions of the first sound and the second sound. For this purpose, generation unit 131 generates an addition sound signal related to the sound in which the first sound and the second sound are superimposed by adding the first sound signal and the second sound signal. Here, the area between the first position and the second position means, for example, the area sandwiched between a virtual straight line passing through the first position and another virtual straight line parallel to the virtual straight line and passing through the second position. At this time, it may be assumed that the virtual straight line and the other virtual straight line are included in the area.
[0064] Generation unit 131 further convolves this addition sound signal with a third head-related transfer function for localizing the sound image at the third position (step S107). Generation unit 131 generates an output sound signal including the addition sound signal that has been subjected to the convolution process in this way (step S108). Note that steps S103 to S108 are also collectively referred to as a generation step.
[0065] The output unit 141 drives the driver 104 by outputting the output sound signal generated by the generation unit 131 to the driver 104, so as to present a sound based on the output sound signal (step S106). In this way, since the first sound and the second sound can be perceived together as a sound reaching from the third position, compared with the case where the first sound is perceived as a sound reaching from the first position and the second sound is perceived as a sound reaching from the second position, the calculation process for localizing the sound image can be simplified. As a result, the required processing capacity can be temporarily reduced, and heat generation due to the driving of the processor, power consumption associated with the calculation process, etc. can be reduced. Also, as described above, due to the simplification of the calculation process, the perception of the sound image position of the user 99 is ambiguous, so the impact on the sense of presence is small. In the acoustic reproduction system 100, in this way, the calculation process can be simplified as needed, so that a three-dimensional sound can be perceived by the user through a more appropriate calculation process.
[0066] Here, the third position described above will be described in more detail with reference to FIG. 4. FIG. 4 is a diagram for explaining the third position where the sound image is localized by the third head transfer function according to the embodiment. In FIG. 4, the sound image position in the three-dimensional sound field is indicated by black dots, and the arrow extending from the black dots toward the user 99 indicates the direction of arrival of the sound to the user 99. Note that a virtual speaker is also shown at the black dots indicating the sound image position.
[0067] In the example shown in FIG. 4, it is assumed that the user 99 is rotating the head, and the rotation speed of this rotation is greater than the first threshold value. Note that when the user 99 displaces the head and the displacement speed of this displacement is greater than the first threshold value, the following operations may be performed. In this example, as shown by the white double-headed arrow, the head of the user 99 is rotating around the first axis perpendicular to the paper surface. At this time, as shown in the figure, the third position P3 or P3a in this example is the position on the bisector line indicated by the hatched arrow in the figure that bisects the angle formed by the straight line connecting the first position P1 or P1a and the user 99 and the straight line connecting the second position P2 or P2a and the user 99.
[0068] In this way, by simplifying the calculation process of the convolution of the head-related transfer function, it becomes possible to make the user 99 perceive three-dimensional sound through more appropriate calculation processing. When the head-related transfer function includes information regarding the distance at which the sound image is localized, among a plurality of head-related transfer functions prepared for localizing the sound image at positions of a plurality of distances in the same sound arrival direction, one head-related transfer function selected therefrom may be convolved. In this case, since the arrival directions of the first sound and the second sound and the distances to the sound image positions are averaged, the user 99 is likely to feel discomfort, and thus a configuration for reducing discomfort such as setting a narrower predetermined region may be further included.
[0069] When the user 99 displaces the head, the displacement speed of this displacement is described as being greater than the first threshold value. In this example, for instance, the head of the user 99 displaces along the second axis in the vertical direction along the plane of the paper. At this time, the third position P3 in this example is a position on the equidistant line that is orthogonal to the second axis direction and has equal distances from the first position P1 and the second position P2. By localizing the sound image at such a position, an average third position P3 can be set in a region of distance where discrimination becomes ambiguous in accordance with the displacement of the head of the user 99. Note that the displacement direction of the head of the user 99 may be in one direction.
[0070] Also, when setting the third position, a position corresponding to either one of the first position and the second position itself may be set. For example, when the first sound is the dialogue of a person in the content and the second sound is the environmental sound in the content, etc., the first sound is prioritized, and the sound image position set for the first sound is set as the third position. According to this, the first sound and the second sound are perceived as sounds arriving from the first position set as the third position. At this time, the first head-related transfer function for making the user 99 perceive the sound as a sound arriving from the first position is used as it is.
[0071] That is, in this example, since the already used head-related transfer function is employed, for example, as shown in the above example, there is no need to set a third position at a position that does not correspond to any sound image position such as the first position and the second position that are originally set by the sound signal. In other words, the sound image position originally set by the sound signal can be set as the third position. Therefore, since the head-related transfer function for localizing the sound image at the originally set sound image position can be diverted, there is no need to use mapping information obtained by mapping the head-related transfer function for making the user 99 perceive sound as sound arriving from an arbitrary point within the three-dimensional sound field. Thus, the process of determining the head-related transfer function for the set third position is simplified, and it becomes possible to make the user 99 perceive stereophonic sound through more appropriate calculation processing. As described above, the range between the first position and the second position means a range including the first position and the second position themselves.
[0072] Also, as the third position, an intermediate point on the line segment connecting the first position and the second position spatially may be set, or simply a random position between the first position and the second position may be set.
[0073] [Modification Example] Hereinafter, the operation of the acoustic reproduction system according to the modification example of the present embodiment will be described with reference to FIGS. 5 and 6 A~Figure 6C In the description of the following modification example of the embodiment, the description will focus on the differences compared with the above embodiment, and the substantially equivalent points will be omitted or simplified.
[0074] FIG. 5 is a flowchart showing the operation of the acoustic reproduction system according to a modification of the embodiment. FIG. 6A is a first diagram for explaining a third position where a sound image is localized by a third head-related transfer function according to a modification of the embodiment. FIG. 6B is a second diagram for explaining a third position where a sound image is localized by a third head-related transfer function according to a modification of the embodiment. FIG. 6C is a third diagram for explaining a third position where a sound image is localized by a third head-related transfer function according to a modification of the embodiment. The acoustic reproduction system according to this modification is different from the acoustic reproduction system 100 according to the above-described embodiment in that the sound to which the head-related transfer function is convolved with respect to the sound signal changes at the boundaries of the first threshold and the second threshold.
[0075] More specifically, in the acoustic reproduction system according to this modification, a second threshold smaller than the first threshold is set. The first threshold is used, as in the above-described embodiment, to determine whether to apply a third head-related transfer function for making the user 99 perceive the first sound and the second sound as sounds arriving from the third position. In this modification, further, by the determination using the second threshold, a third head-related transfer function for making the user 99 perceive the first intermediate sound and the second intermediate sound, which are localized at a first intermediate position and a second intermediate position closer to the third position than the first sound and the second sound, as sounds arriving from the third position is convolved, thereby realizing a reduction in the amount of calculation processing.
[0076] Here, a determination is made based on the movement speed of the head of the user 99. When the movement speed is equal to or lower than the second threshold, the first sound is localized at the first position P1, the second sound is localized at the second position P2, the first intermediate sound is localized at the first intermediate position P1m (see FIG. 6A, etc.), and the second intermediate sound is localized at the second intermediate position P2m (see FIG. 6A, etc.). On the other hand, when the movement speed of the head of the user 99 is greater than the first threshold, as described above, a process of convolving the third head-related transfer function with the sound signals related to the first sound and the second sound (that is, the first sound signal and the second sound signal) is applied. At this time, the third head-related transfer function is also convolved with the sound signals related to the first intermediate sound and the second intermediate sound (that is, the first intermediate sound signal and the second intermediate sound signal), and the first sound, the second sound, the first intermediate sound, and the second intermediate sound are all localized at the third position P3.
[0077] In addition, in this modified example, when the movement speed of the head of the user 99 is greater than the second threshold value and equal to or less than the first threshold value, the first sound is localized at the first position P1, the second sound is localized at the second position P2, and the first intermediate sound and the second intermediate sound are localized at the third position P3. That is, in this modified example, when the movement speed of the head of the user 99 is not so fast as to be equal to or less than the second threshold value, the calculation process of the convolution of the head transfer function is simplified for a narrower predetermined region (that is, a narrow region) that does not include the first position P1 and the second position P2 and includes the first intermediate position P1m and the second intermediate position P2m.
[0078] As an operation in the acoustic reproduction system according to this modified example, as shown in FIG. 5, after the acquisition unit 121 acquires the movement speed (step S102), the generation unit 131 determines whether the movement speed is greater than the second threshold value (step S201). When the movement speed is equal to or less than the second threshold value (No in step S201), the process proceeds to step S202, and an operation (step S202) of convolving the head transfer function for localizing the sound image at the position where each sound signal should be originally localized is performed in the same manner as in the above-described embodiment. That is, the first head transfer function for localizing the sound image at the first position P1 is convolved with the first sound signal related to the first sound, the second head transfer function for localizing the sound image at the second position P2 is convolved with the second sound signal related to the second sound, the first intermediate head transfer function for localizing the sound image at the first intermediate position P1m is convolved with the first intermediate sound signal related to the first intermediate sound, and the second intermediate head transfer function for localizing the sound image at the second intermediate position P2m is convolved with the second intermediate sound signal related to the second intermediate sound.
[0079] On the other hand, when the movement speed is greater than the second threshold (Yes in step S201), the generation unit 131 further determines whether the movement speed is greater than the first threshold (step S204). When the movement speed is equal to or less than the first threshold (No in step S204), the acoustic reproduction system 100 causes the user 99 to perceive the first intermediate sound and the second intermediate sound as sounds arriving from the third position. For this purpose, the generation unit 131 convolves the addition sound signal obtained by adding the first intermediate sound signal related to the first intermediate sound and the second intermediate sound signal related to the second intermediate sound with the third head transfer function (step S205). The generation unit 131 generates an output sound signal including the first sound signal, the second sound signal, and the addition sound signal obtained by adding the first intermediate sound signal and the second intermediate sound signal, which have been subjected to the convolution process in this way (step S206). Thereafter, the process proceeds to step S106, and the same operation as in the above-described embodiment is performed.
[0080] On the other hand, when the movement speed is greater than the first threshold (Yes in step S204), the process proceeds to step S207, and the same operation as in the above-described embodiment is performed to convolve the addition sound signal obtained by adding the first sound signal and the second sound signal with the third head transfer function. In this modification, further, the first intermediate sound signal and the second intermediate sound signal are also added to this addition sound signal, and the first sound, the second sound, the first intermediate sound, and the second intermediate sound are perceived by the user 99 as sounds arriving from the third position P3.
[0081] As a result of the above operation, in the acoustic reproduction system according to the modification of the present embodiment, when the movement speed of the user 99 is equal to or less than the second threshold, an acoustic image as shown in FIG. 6A is formed in the three-dimensional sound field. Note that in FIG. 6A, a view of the three-dimensional sound field as seen from the first axis direction is shown in the same manner as in FIG. 4. As shown in FIG. 6A, when the movement speed of the user 99 is equal to or less than the second threshold, each of the first sound, the second sound, the first intermediate sound, and the second intermediate sound is perceived by the user 99 as a sound arriving from the original acoustic image position.
[0082] In the acoustic reproduction system according to this modified example, when the movement speed of the user 99 is equal to or lower than the first threshold value and higher than the second threshold value, an acoustic image shown in FIG. 6B is formed in the three-dimensional sound field. Note that FIG. 6B shows a view of the three-dimensional sound field as seen from the first axis direction, similar to FIG. 4.
[0083] As shown in FIG. 6B, when the movement speed of the user 99 is equal to or lower than the first threshold value and higher than the second threshold value, a first intermediate sound that is originally perceived by the user 99 as a sound reaching from a first intermediate position P1m closer to a third position P3 than a first position P1 is perceived by the user 99 as a sound reaching from the third position P3. Similarly, when the movement speed is equal to or lower than the first threshold value and higher than the second threshold value, a second intermediate sound that is originally perceived by the user 99 as a sound reaching from a second intermediate position P2m closer to the third position P3 than a second position P2 is perceived by the user 99 as a sound reaching from the third position P3.
[0084] Furthermore, in the acoustic reproduction system according to this modified example, when the movement speed of the user 99 is higher than the first threshold value, an acoustic image shown in FIG. 6C is formed in the three-dimensional sound field. Note that FIG. 6C shows a view of the three-dimensional sound field as seen from the first axis direction, similar to FIG. 4.
[0085] As shown in FIG. 6C, when the movement speed of the user 99 is higher than the first threshold value, all the sounds that are originally localized at acoustic image positions included in a predetermined region that includes the first position P1 and the second position P2, including the first intermediate position P1m and the second intermediate position P2m, are perceived by the user 99 as sounds reaching from the third position P3.
[0086] By doing so, when the movement speed exceeds the second threshold value, the sounds within a predetermined region having a width that gradually corresponds to the movement speed of the user 99 are perceived by the user 99 as sounds reaching from the third position P3. For example, in the figure, in the case of a movement speed exceeding the first threshold value, the sounds within the predetermined region indicated by the long dashed line are perceived by the user 99 as sounds reaching from the third position P3. Also, in the case of a movement speed exceeding the second threshold value and equal to or lower than the first threshold value, the sounds within the narrow predetermined region (i.e., the narrow region) indicated by the dashed line are perceived by the user 99 as sounds reaching from the third position P3.
[0087] At this time, as the third position P3, the first intermediate position P1m and the second intermediate position P2m are considered. That is, the third position P3 is set based on the four positions of the first position P1, the second position P2, the first intermediate position P1m, and the second intermediate position P2m. Here, for example, as the third position P3, on the straight line connecting the center between the first position P1, the second position P2, the first intermediate position P1m, and the second intermediate position P2m and the user 99, and from each of the first position P1, the second position P2, the first intermediate position P1m, and the second intermediate position P2m, a position having the same distance as the shortest distance among the distances to the position of the user 99 is set. Further, the third position P3 may be set to the average coordinates of the coordinates corresponding to the four positions in the plane coordinates viewed from the first axis direction, and the like.
[0088] Furthermore, three or more levels such as a third threshold value with respect to the movement speed of the user 99 may be provided, and the sound within a further narrow predetermined region may be configured to be perceived by the user 99 as the sound reaching from the third position P3. There is no particular limitation on the number of levels in the relationship between the movement speed and the width of the predetermined region.
[0089] Regarding the second threshold value, similar to the first threshold value in the description of the above embodiment, it may be set based on a user-specific numerical setting or the like regarding at what movement speed the perception of the sound image position by the user 99 becomes ambiguous, or a generalized numerical value may be set.
[0090] (Other Embodiments) As described above, the embodiments have been described, but the present disclosure is not limited to the above embodiments.
[0091] For example, in the above embodiment, an example where the sound does not follow the movement of the user's head was described. However, the content of the present disclosure is also effective when the sound follows the movement of the user's head. That is, the user is made to perceive a first sound as the sound reaching from a first position that relatively moves along with the movement of the user's head, and a second sound as the sound reaching from a second position that relatively moves along with the movement of the user's head. Among these operations, when the movement speed of the head is greater than a first threshold value, the first sound and the second sound are made to be perceived as the sound reaching from a third position that relatively moves along with the movement of the user's head.
[0092] Even in this case, a process of convolving the head transfer functions for localizing the first sound and the second sound to the first position and the second position, respectively, with each sound signal is performed. Since the head transfer functions convolved with the sound signals are made common with the first threshold as a boundary, the calculation process is simplified. That is, similar to the above embodiment, the required processing ability can be temporarily reduced, and heat generation due to driving the processor, power consumption associated with the calculation process, etc. can be reduced. On the other hand, even if such simplification of the calculation process is performed, if the movement speed of the user's head is high, it becomes difficult to accurately perceive the position of the sound image, so the user's discomfort with respect to the sound image position is less likely to increase. Therefore, it becomes possible to make the user perceive three-dimensional sound by more appropriate calculation processing.
[0093] Also, for example, the acoustic reproduction system described in the above embodiment may be realized as one device including all the components, or may be realized by allocating each function to a plurality of devices and having these plurality of devices cooperate. In the latter case, an information processing device such as a smartphone, a tablet terminal, or a PC may be used as the device corresponding to the processing module.
[0094] In addition, the acoustic reproduction system of the present disclosure can also be realized as an acoustic processing device that is connected to a playback device having only a driver and outputs only an output audio signal obtained by performing convolution processing of the head-related transfer function on the acquired audio signal with respect to the playback device. In this case, the acoustic processing device may be realized as hardware including a dedicated circuit, or may be realized as software for causing a general-purpose processor to execute specific processing.
[0095] In addition, in the above-described embodiments, the processing executed by a specific processing unit may be executed by another processing unit. Also, the order of a plurality of processes may be changed, or a plurality of processes may be executed in parallel.
[0096] In addition, in the above-described embodiments, each component may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.
[0097] In addition, each component may be realized by hardware. For example, each component may be a circuit (or an integrated circuit). These circuits may constitute one circuit as a whole, or may be separate circuits respectively. Also, these circuits may be general-purpose circuits or dedicated circuits respectively.
[0098] In addition, the general or specific aspects of the present disclosure may be realized by a system, a device, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM. Also, the general or specific aspects of the present disclosure may be realized by any combination of a system, a device, a method, an integrated circuit, a computer program, and a recording medium.
[0099] For example, the present disclosure may be implemented as a method for reproducing an audio signal executed by a computer, or may be implemented as a program for causing a computer to execute the audio signal reproduction method. The present disclosure may be implemented as a computer-readable non-transitory recording medium on which such a program is recorded.
[0100] In addition, forms obtained by applying various modifications that can be conceived by those skilled in the art to each embodiment, or forms realized by arbitrarily combining the components and functions in each embodiment without departing from the spirit of the present disclosure are also included in the present disclosure.
Industrial Applicability
[0101] The present disclosure is useful in acoustic reproduction for allowing a user to perceive three-dimensional sound accompanied by movement of the user's head.
Explanation of Signs
[0102] 99 User 100 Acoustic reproduction system 101 Processing module 102 Communication module 103 Detector 104 Driver 111 Input unit 121 Acquisition unit 131 Generation unit 141 Output unit 200 Stereoscopic video reproduction system P1, P1a First position P2, P2a Second position P3, P3a Third position P1m First intermediate position P2m Second intermediate position
Claims
**Claim 1**: An acoustic playback method executed by an acoustic playback system, which causes a user to perceive a first sound as a sound reaching from a first position on a three-dimensional sound field and causes the user to perceive a second sound as a sound reaching from a second position different from the first position, the method comprising: an acquisition step of acquiring a movement speed of the user's head; a generation step of generating an output sound signal for causing the user to perceive a sound reaching from a predetermined position on the three-dimensional sound field; and in the generation step, when the acquired movement speed is greater than a first threshold, generating the output sound signal for causing the user to perceive the first sound and the second sound as sounds reaching from a third position between the first position and the second position An acoustic playback method. **Claim 2**: In the generation step, when the acquired movement speed is less than or equal to the first threshold, generating the output sound signal by convolving a first head transfer function for localizing sound at the first position with a first sound signal related to the first sound and convolving a second head transfer function for localizing sound at the second position with a second sound signal related to the second sound; when the acquired movement speed is greater than the first threshold, generating the output sound signal by convolving a third head transfer function for localizing sound at the third position with an added sound signal obtained by adding the first sound signal and the second sound signal The acoustic playback method according to claim 1. **Claim 3**: The movement speed is a rotation speed of the user's head around a first axis passing through the user's head, and the third position is a position on a bisector that bisects an angle formed by straight lines connecting the user to the first position and the second position, respectively, in a virtual plane when the three-dimensional sound field is viewed from the direction of the first axis The acoustic playback method according to claim 1 or 2. **Claim 4**: The rotation speed is acquired as a rotation amount per unit time detected by a detector that moves integrally with the user's head and detects a rotation amount around at least one of three mutually orthogonal axes as a rotation axis The acoustic playback method according to claim 3. **Claim 5**: The movement speed is a displacement speed of the user's head along a second axis direction passing through the user's head, and the displacement speed is acquired as a displacement amount per unit time detected by a detector that moves integrally with the user's head and detects a displacement amount along at least one of three mutually orthogonal axes as a displacement direction The acoustic reproduction method according to claim 1 or 2.
6. In the acoustic reproduction method, a plurality of sounds reaching from each position within a predetermined region on the three-dimensional sound field, including the first position and the second position, and including at least the first sound and the second sound, are made perceptible to the user. In the generation step, when the movement speed is greater than the first threshold value, an output sound signal for making all of the plurality of sounds perceptible to the user as sounds reaching from the third position is generated. The acoustic reproduction method according to any one of claims 1 to 5.
7. In the acoustic reproduction method, a first intermediate sound is made perceptible to the user as a sound reaching from a first intermediate position between the first position and the third position, and a second intermediate sound is made perceptible to the user as a sound reaching from a second intermediate position between the second position and the third position. In the generation step, further, when the movement speed is less than or equal to the first threshold value and greater than a second threshold value smaller than the first threshold value, an output sound signal for making the first intermediate sound and the second intermediate sound perceptible to the user as sounds reaching from the third position is generated. The acoustic reproduction method according to any one of claims 1 to 6.
8. For causing a computer to execute the acoustic reproduction method according to any one of claims 1 to 7 Program.
9. An acoustic reproduction system that makes a first sound perceptible to the user as a sound reaching from a first position on a three-dimensional sound field and makes a second sound perceptible to the user as a sound reaching from a second position different from the first position, An acquisition unit that acquires the movement speed of the user's head, A generation unit that generates an output sound signal for making a sound reaching from a predetermined position on the three-dimensional sound field perceptible to the user, and includes: When the acquired movement speed is greater than the first threshold value, the generation unit generates the output sound signal for making the first sound and the second sound perceptible to the user as sounds reaching from a third position between the first position and the second position. Acoustic reproduction system.
Citation Information
Patent Citations
Methods, apparatuses and computer programs relating to spatial audio
EP3503592A1
Video game authoring system
JP1994233395A
Information processor, information processing method, and program
JP2013005021A
Simulation system and program
JP2017184174A
Voice generation program in virtual space, generation method of quadtree, and voice generation device
JP2020018620A
Cited By
Acoustic reproduction method, program, and acoustic reproduction system
JP2025128231A