Information processing method, information processing system, and program
By employing a distributed processing method with panning and head-related transfer functions, the method addresses high computational demands in virtual three-dimensional sound reproduction, ensuring realistic and efficient sound adaptation to user movements.
Patent Information
- Application Number
- PCT/JP2025/025424
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-14
- Filing Date
- 2025-07-16
- Publication Date
- 2026-01-22
AI Technical Summary
Existing sound reproduction technologies for virtual three-dimensional spaces require extensive processing to maintain realistic sound environments, especially when user movements change the sound transmission path, leading to high computational demands and potential degradation of sound quality.
A method involving multiple information processing terminals, where a first terminal performs panning processing on sound information and transmits it to a second terminal worn by the user, which detects head position and generates output sound signals using head-related transfer functions, reducing processing load and maintaining sound quality.
This approach reduces processing requirements by consolidating sounds from fewer representative directions, allowing instant updates to sound direction based on head movement while maintaining high sound quality.
Smart Images

Figure JP2025025424_22012026_PF_FP_ABST
Abstract
Description
Information processing method, information processing system, and program
[0001] The present disclosure relates to an information processing method, an information processing system, and a program.
[0002] Conventionally, there are known techniques for reproducing sound in a virtual three-dimensional space to allow a user to perceive stereoscopic sound (see, for example, Patent Document 1). Furthermore, in order to perceive sound as if it is coming from a sound source object to the user in such a three-dimensional space, processing is required to generate output sound information from original sound information. In particular, reproducing stereoscopic sound in response to the user's body movements in a virtual space requires extensive processing. Advances in computer graphics (CG) have made it relatively easy to create visually complex virtual environments, making technology for realizing corresponding auditory information important. Additionally, when processing from sound information to output sound information is performed in advance, a large memory area is required to store the pre-calculated processing results. Furthermore, transmitting such large amounts of processed data may require a wide communication bandwidth.
[0003] To realize a more realistic sound environment, the number of objects that emit sound in the virtual three-dimensional space increases, secondary sounds based on acoustic effects such as reflected sound, diffracted sound, and reverberation increase, and these secondary sounds must be appropriately changed in response to the user's movements, requiring a large amount of processing.To reduce this large amount of processing, a conversion technique known as panning processing (or simply panning) is known, which represents sounds in a three-dimensional space using sounds from several representative points that are set in advance in the three-dimensional space.
[0004] Japanese Patent Application Laid-Open No. 2020-18620
[0005] However, in a conversion process such as panning, a reduction in the amount of processing may result in a degradation of sound quality. Therefore, an object of the present disclosure is to provide an information processing method for appropriately performing the conversion process.
[0006] An information processing method according to one aspect of the present disclosure is an information processing method executed by a plurality of information processing terminals, the information processing method including a step of acquiring, in a first terminal that is one of the plurality of information processing terminals, first sound information including an acoustic signal and information about a position of a sound source object in a three-dimensional sound field; the first sound information being information for causing the sound source object in the three-dimensional sound field to emit a reproduced sound using the acoustic signal, and converting, in the first terminal, the first sound information into second sound information for generating a representative sound arriving at a reference position from a representative point set in the three-dimensional sound field using the acoustic signal. the first terminal transmitting the second sound information to a second terminal which is another information processing terminal among the plurality of information processing terminals; the second terminal detecting a user's position or head direction in the three-dimensional sound field; the second terminal calculating a position of a reproduction representative point corresponding to the position of the representative point based on the detected user's position or head direction and the reference position; and the second terminal generating an output sound signal using a head related transfer function corresponding to the calculated position of the reproduction representative point and the received second sound information.
[0007] Moreover, an information processing system according to one aspect of the present disclosure is an information processing system including a first terminal and a second terminal, wherein the first terminal comprises: an acquisition unit that acquires sound information including an acoustic signal and information about the position of a sound source object within a three-dimensional sound field; the first sound information is information for causing the sound source object within the three-dimensional sound field to emit a playback sound using the acoustic signal; a conversion unit that converts the first sound information into second sound information for generating a representative sound that arrives at a reference position from a representative point set within the three-dimensional sound field using the acoustic signal; and a transmission unit that transmits the second sound information to the second terminal, wherein the second terminal comprises: a detector that detects a user's position or head direction within the three-dimensional sound field; a calculation unit that calculates a position of a playback representative point corresponding to the position of the representative point based on the detected user's position or head direction and the reference position; and an output unit in the second terminal that outputs an output sound signal using a head-related transfer function corresponding to the calculated position of the playback representative point and the received second sound information.
[0008] Furthermore, one aspect of the present disclosure can also be realized as a program for causing a computer to execute the information processing method described above.
[0009] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.
[0010] According to the present disclosure, it is possible to appropriately perform conversion processing.
[0011] FIG. 1 is a schematic diagram illustrating a use example of a sound reproduction system according to an embodiment. FIG. 2 is a block diagram illustrating a functional configuration of the sound reproduction system according to an embodiment. FIG. 3 is a diagram illustrating an example of an audio signal according to an embodiment. FIG. 4 is a block diagram illustrating a functional configuration of an acquisition unit according to an embodiment. FIG. 5 is a block diagram illustrating a functional configuration of an output sound generation unit according to an embodiment. FIG. 6 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 7 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 8 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 9 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 10 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 11 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 12 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 13 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 14 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 15 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 16 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 17A is a flowchart showing an example of operation of an information processing device according to an embodiment. FIG. 17B is a flowchart showing an example of operation of an information processing device according to an embodiment. FIG. 18 is a flowchart of audio reproduction processing according to an embodiment. FIG. 19 is a diagram for explaining synthesis of head-related transfer functions in audio reproduction processing according to an embodiment. FIG. 20 is a diagram for explaining arrangement of representative directions according to an embodiment. FIG. 21 is a diagram for explaining an example of time shift value calculation according to an embodiment. FIG. 22 is a diagram for explaining an example of time shift value calculation according to an embodiment. FIG. 23 is a diagram for explaining an example of gain value calculation according to an embodiment. FIG. 24 is a diagram showing a result of gain value calculation according to an embodiment. FIG. 25 is a diagram showing a result of gain value calculation according to an embodiment. FIG. 26 is a diagram for verifying a result of time shift value calculation according to an embodiment.FIG. 27 is a diagram for verifying the results of gain value calculation according to an embodiment. FIG. 28 is a diagram showing the results of a localization experiment for verifying the results of gain value calculation according to an embodiment. FIG. 29 is a diagram showing the results of a localization experiment for verifying the results of gain value calculation according to an embodiment. FIG. 30 is a diagram showing the results of a localization experiment for verifying the results of gain value calculation according to an embodiment. FIG. 31 is a diagram showing the results of a localization experiment for verifying the results of gain value calculation according to an embodiment. FIG. 32 is a diagram showing set fixed gain values according to an embodiment. FIG. 33 is a diagram showing the results of a localization experiment for verifying the results of gain value calculation according to an embodiment. FIG. 34 is a diagram showing the results of a localization experiment for verifying the results of gain value calculation according to an embodiment.
[0012] (Knowledge that forms the basis of the disclosure) Conventionally, a technology related to sound reproduction that allows a user to perceive stereoscopic sound in a virtual three-dimensional space (hereinafter sometimes referred to as a three-dimensional sound field) has been known (see, for example, Patent Document 1). Using this technology, a user can perceive sound as if a sound source object exists at a predetermined position in the virtual space and the sound is coming from that direction. In order to localize a sound image at a predetermined position in the virtual three-dimensional space in this way, for example, calculation processing is required for a sound signal emitted by a sound source object (also referred to as a sound emitted from the sound source object or a reproduced sound) to generate a sound arrival time difference between the two ears and a sound level difference (or sound pressure difference) between the two ears that causes the sound to be perceived as stereoscopic sound. Such calculation processing is performed by applying a stereophonic filter. A stereophonic filter is an information processing filter that, when an output sound signal obtained by applying the filter to original sound information is reproduced, causes the position (such as the direction and distance) of the sound, the size of the sound source, the width of the space, and the like to be perceived with a three-dimensional effect.
[0013] As an example of the computational process for applying such a stereophonic filter, a process is known in which a head-related transfer function (HRTF) is convolved with a target sound signal to make the sound perceived as coming from a predetermined direction. By performing this HRTF convolution process at a sufficiently fine angle with respect to the direction of arrival of the reproduced sound from the position of the sound source object to the user's position, the sense of realism experienced by the user is improved.
[0014] In recent years, there has been active development of technologies related to virtual reality (VR). Virtual reality focuses on appropriately changing the position of a sound source object in a virtual three-dimensional space in response to a user's movements, allowing the user to experience the sensation of moving within the virtual space. To achieve this, it is necessary to move the localization position of a sound image in the virtual space relative to the user's movements. This processing has been performed by applying a stereophonic filter, such as the head-related transfer function described above, to the original sound information. However, when a user moves within a three-dimensional space, the sound transmission path changes from moment to moment due to changes in the positional relationship between the sound source object and the user, such as due to sound reverberation and interference. This requires determining the sound transmission path from the sound source object based on the positional relationship between the sound source object and the user, and convolving the transfer function to take into account sound reverberation and interference. However, such information processing requires a huge amount of processing, and an improvement in the sense of realism may not be achieved without a large-scale processing device.
[0015] Therefore, in order to reduce such an enormous amount of processing, attempts have been made to apply a panning process to the reproduced sound to reduce the amount of convolution of the head-related transfer function. Specifically, instead of convolving the reproduced sound with a head-related transfer function for each of a number of sound source objects in a three-dimensional space, the reproduced sound from the sound source object is re-expressed using sounds (representative sounds) from several representative points pre-set in the three-dimensional space. Then, by simply convolving the representative sounds with the head-related transfer functions from the representative points to the user's position, it becomes possible to allow the user to perceive a three-dimensional sound that is comparable to that of the original sound source objects. If the number of representative points is smaller than the number of original sound source objects, the number of targets for convolution of the head-related transfer function will naturally be reduced, which is advantageous in terms of processing amount.
[0016] A more specific outline of the present disclosure is as follows.
[0017] an information processing method according to a first aspect of the present disclosure, which is an information processing method executed by a plurality of information processing terminals, including: a step of acquiring, at a first terminal that is one of the plurality of information processing terminals, first sound information including an acoustic signal and information about the position of a sound source object within a three-dimensional sound field; the first sound information being information for causing the sound source object within the three-dimensional sound field to emit a playback sound using the acoustic signal; a step of converting, at the first terminal, the first sound information into second sound information for generating a representative sound that arrives at a reference position from a representative point set within the three-dimensional sound field using the acoustic signal; a step of transmitting, at the first terminal, the second sound information to a second terminal that is another of the plurality of information processing terminals; a step of detecting, at the second terminal, a user's position or head direction within the three-dimensional sound field; a step of calculating, at the second terminal, a position of a playback representative point corresponding to the position of the representative point based on the detected user's position or head direction and the reference position; and a step of generating, at the second terminal, an output sound signal using a head-related transfer function corresponding to the calculated position of the playback representative point and the received second sound information.
[0018] According to this information processing method, when the second terminal is worn by a user, the first terminal can be provided separately from the second terminal worn by the user. Therefore, the first terminal, which is not restricted by the need for a user to wear the first terminal and can have relatively high processing performance compared to the second terminal, can perform panning processing on an audio signal, including the movement of a sound source object defined in content, which requires relatively high processing resources. Furthermore, panning processing can reduce the amount of information by consolidating sounds from fewer representative directions, thereby reducing the impact of communication bandwidth restrictions when transmitting and receiving data between the first terminal and the second terminal. Furthermore, when the second terminal is worn by a user, the second terminal worn by the user can easily detect the position and orientation of the user's head and use this information to output an output sound signal from the panned audio signal. Therefore, the sound arrival direction can be instantly updated in response to head movement. Furthermore, when outputting an output sound signal, it is only necessary to convolve head-related transfer functions for a relatively small number of representative directions, i.e., a few directions after panning processing, thereby reducing the processing resources required for the second terminal. In this way, according to the present disclosure, the conversion process can be performed appropriately so as to obtain multiple advantages.
[0019] An information processing method according to a second aspect is the information processing method according to the first aspect, in which the converting step applies time shift adjustment and gain adjustment to the reproduced sound to convert it into a representative sound.
[0020] According to this, by applying time shift adjustment and gain adjustment to the reproduced sound, it is possible to convert it into a representative sound that is less likely to cause discomfort.
[0021] An information processing method according to a third aspect is the information processing method according to the second aspect, wherein the adjustment amount of the time shift adjustment is set to 0 at the representative point.
[0022] This allows the adjustment amount of the time shift adjustment to be set so that it becomes 0 at the representative point, which is an adjustment amount that is substantially the correct value.
[0023] An information processing method according to a fourth aspect is the information processing method according to the second or third aspect, wherein the adjustment amount of the time shift adjustment at the predetermined position is set using the adjustment amount at a position closer to the representative point than the predetermined position.
[0024] According to this, when setting the adjustment amount of the time shift adjustment at a predetermined position, by using the adjustment amount at a position close to the representative point as viewed from the predetermined position, it is possible to set an adjustment amount closer to the adjustment amount at a position close to the representative point, for example, and as a result, it is possible to make it difficult for jumps to occur, where the adjustment amount of the time shift adjustment changes suddenly.
[0025] Furthermore, the information processing method according to the fifth aspect is an information processing method according to any one of the second to fourth aspects, in which the adjustment amount for gain adjustment at a predetermined position is calculated by using two of three adjacent representative points surrounding the predetermined position to calculate the adjustment amount for each of the two representative points that minimizes the error, and then fixing the ratio of the calculated adjustment amounts to calculate the adjustment amount for the remaining representative point.
[0026] According to this method, after calculating the adjustment amount for each of the two representative points that minimizes the error using two representative points, the ratio of the calculated adjustment amounts can be fixed and the adjustment amount for the remaining representative point can be calculated sequentially. This makes it possible to adjust the influence of the two representative points initially used to calculate the adjustment amount and the remaining representative point on the calculation of the adjustment amount for gain adjustment at a predetermined position.
[0027] Furthermore, an information processing method according to a sixth aspect is an information processing method according to any one of the second to fifth aspects, in which the adjustment amount for gain adjustment at a predetermined position is calculated for a horizontal position obtained by removing the vertical component from the predetermined position, using two of three adjacent representative points surrounding the predetermined position that are located horizontally to minimize the error, and then the ratio of the calculated adjustment amounts is fixed and the adjustment amount for the remaining representative point is calculated.
[0028] According to this method, after calculating the amount of gain adjustment for a horizontal position at two representative points located in the horizontal direction, it is possible to fix the ratio of the calculated adjustment amounts and calculate the adjustment amount for the remaining representative point. The sense of sound localization in the horizontal direction is relatively accurate. Therefore, by first calculating the adjustment amount for the horizontal position and fixing the ratio of the adjustment amounts, the influence of the two representative points in the horizontal direction on the calculation of the adjustment amount for gain adjustment at a predetermined position is made greater than the influence of the remaining representative point on the calculation of the adjustment amount for gain adjustment at a predetermined position, thereby increasing the accuracy of the sense of localization in the horizontal direction and making it possible to output an output sound signal in which the sound image is more appropriately localized.
[0029] In addition, the information processing method according to the seventh aspect is an information processing method according to any one of the second to sixth aspects, in which the amount of time shift adjustment is corrected based on the expected value of the error so as to reduce the error from the calculated correct value.
[0030] This allows the amount of time shift adjustment to be set so as to reduce the error from the calculated correct value.
[0031] In addition, the information processing method according to the eighth aspect is an information processing method according to any one of the second to seventh aspects, in which the amount of gain adjustment at a predetermined position is corrected based on the expected value of the error so as to reduce the error from the calculated correct value.
[0032] This allows the amount of gain adjustment at a predetermined position to be set so as to reduce the error from the calculated correct value.
[0033] An information processing method according to a ninth aspect is an information processing method according to any one of the second to sixth aspects, in which the adjustment amount for gain adjustment at a predetermined position is set using two or more of adjustment amounts at a plurality of virtual representative points corresponding to a plurality of directions within a predetermined angle range in the horizontal direction centered on the direction of the representative point of the starting point, with the representative point for the predetermined position as the starting point.
[0034] According to this, as the adjustment amount of gain adjustment at a predetermined position, two or more adjustment amounts at a plurality of virtual representative points corresponding to a plurality of directions within a predetermined angle range in the horizontal direction centered on the direction of the representative point of origin can be used, thereby making it possible to set the adjustment amount of gain adjustment at a predetermined position so that all of the two or more adjustment amounts have an effect.
[0035] An information processing method according to a tenth aspect is the information processing method according to the ninth aspect, wherein the adjustment amount for gain adjustment at a predetermined position is set using all of the adjustment amounts at a plurality of virtual representative points corresponding to a plurality of directions within a predetermined angle range in the horizontal direction centered on the direction of the representative point of the starting point, with the representative point for the predetermined position as the starting point.
[0036] This allows the adjustment amounts for gain adjustment at a predetermined position to be all used, as adjustment amounts for gain adjustment at a plurality of virtual representative points corresponding to a plurality of directions within a predetermined angle range in the horizontal direction centered on the direction of the representative point of origin, thereby enabling the adjustment amount for gain adjustment at a predetermined position to be set so that all adjustment amounts have an effect.
[0037] An information processing method according to an eleventh aspect is the information processing method according to the ninth aspect, wherein the adjustment amount for gain adjustment at a predetermined position is calculated by averaging the adjustment amounts at a plurality of virtual representative points.
[0038] According to this, the adjustment amount of the gain adjustment at a predetermined position can be calculated by averaging two or more adjustment amounts.
[0039] An information processing method according to a twelfth aspect is the information processing method according to the ninth aspect, in which the adjustment amount for gain adjustment at a predetermined position is calculated by taking a weighted average of the adjustment amounts at a plurality of virtual representative points based on the angle difference between the direction of each virtual representative point and the original direction corresponding to that virtual representative point.
[0040] According to this, the adjustment amount of the gain adjustment at a predetermined position can be calculated by taking a weighted average of two or more adjustment amounts.
[0041] An information processing method according to a thirteenth aspect is an information processing method according to any one of the second to twelfth aspects, in which the amount of gain adjustment at a predetermined position is set to a predetermined value that does not depend on the direction of the predetermined position in the horizontal plane, by adjusting the amount of adjustment of an elevation angle representative point that includes a vertical component among three adjacent representative points that surround the predetermined position.
[0042] According to this, by setting the adjustment amount of the elevation angle representative point to a predetermined value that does not depend on the direction of the predetermined position in the horizontal plane, it is possible to set the adjustment amount of the gain adjustment at the predetermined position.
[0043] An information processing method according to a fourteenth aspect is the information processing method according to the thirteenth aspect, wherein the predetermined value is a value that changes depending on the direction of the predetermined position in a vertical plane.
[0044] This allows the amount of adjustment of the elevation angle representative point to be set to a predetermined value that does not depend on the direction of the predetermined position in the horizontal plane, but changes depending on the direction of the predetermined position in the vertical plane, thereby making it possible to set the amount of adjustment for gain adjustment at a predetermined position.
[0045] In addition, an information processing method according to a fifteenth aspect is the information processing method according to the fourteenth aspect, in which the predetermined value is a value that gradually increases between 0 and 1 when the direction of the predetermined position in the vertical plane is from 0° to 90°.
[0046] According to this, the adjustment amount of the elevation angle representative point does not depend on the direction of the specified position in the horizontal plane, but is set to a predetermined value that gradually increases between 0 and 1 when the direction of the specified position in the vertical plane is from 0° to 90°, thereby making it possible to set the adjustment amount of the gain adjustment at the specified position.
[0047] Further, an information processing system according to a sixteenth aspect is an information processing system including a first terminal and a second terminal, wherein the first terminal comprises an acquisition unit that acquires sound information including an acoustic signal and information about the position of a sound source object within a three-dimensional sound field, a conversion unit that converts the first sound information into second sound information using the acoustic signal to generate a representative sound that arrives at a reference position from a representative point set within the three-dimensional sound field, and a transmission unit that transmits the second sound information to the second terminal, and the second terminal comprises a detector that detects the position of a user or the direction of the head within the three-dimensional sound field, a calculation unit that calculates the position of a reproduction representative point corresponding to the position of the representative point based on the detected position or direction of the user's head and the reference position, and an output unit in the second terminal that outputs an output sound signal using a head related transfer function corresponding to the calculated position of the reproduction representative point and the received second sound information.
[0048] This can achieve the same effects as the information processing method described above.
[0049] A program according to a seventeenth aspect is a program for causing a computer to execute the information processing method described above.
[0050] This makes it possible to achieve the same effects as the above-described information processing method using a computer.
[0051] Furthermore, these comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.
[0052] Hereinafter, embodiments will be described in detail with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not recited in independent claims will be described as optional components. Note that each figure is a schematic diagram and is not necessarily an exact illustration. Furthermore, in each figure, substantially identical components are assigned the same reference numerals, and duplicated descriptions may be omitted or simplified.
[0053] In the following description, elements may be assigned ordinal numbers such as first, second, and third. These ordinal numbers are assigned to elements in order to identify them and do not necessarily correspond to a meaningful order. These ordinal numbers may be rearranged, newly added, or removed as appropriate.
[0054] In the following description, an acoustic signal included in sound information may be described, but the acoustic signal may also be referred to as a voice signal or a sound signal. In other words, in the present disclosure, the acoustic signal has the same meaning as the voice signal or the sound signal.
[0055] (Embodiment) [Overview] First, an overview of an audio reproduction system according to an embodiment will be described. Fig. 1 is a schematic diagram showing an example of use of an audio reproduction system according to an embodiment. Fig. 1 shows a user 99 using an audio reproduction system 100.
[0056] The sound reproduction system 100 shown in FIG. 1 is used simultaneously with, for example, a three-dimensional video reproduction device 300. By simultaneously viewing three-dimensional images and three-dimensional sound, the image enhances the auditory sense of realism, and the sound enhances the visual sense of realism, allowing the user to experience the image and sound as if they were actually at the scene where they were captured. For example, when an image (moving image) of people having a conversation is displayed, it is known that even if the localization of the sound image (sound source object) of the conversation sound is not aligned with the person's mouth, the user 99 will perceive the conversation sound as coming from the person's mouth. In this way, the visual information can correct the position of the sound image, and the image and sound can be combined to enhance the sense of realism.
[0057] The three-dimensional video reproduction device 300 is an image display device worn on the head of the user 99. Therefore, the three-dimensional video reproduction device 300 moves integrally with the head of the user 99. For example, as shown in the figure, the three-dimensional video reproduction device 300 is a glasses-type device that is supported by the ears and nose of the user 99.
[0058] The three-dimensional video reproduction device 300 changes the displayed image in accordance with the movement of the user 99's head, thereby making the user 99 perceive the movement of his or her head in the three-dimensional image space. In other words, when an object in the three-dimensional image space is located in front of the user 99, if the user 99 turns to the right, the object moves to the left of the user 99, and if the user 99 turns to the left, the object moves to the right of the user 99. In this way, the three-dimensional video reproduction device 300 moves the three-dimensional image space in the opposite direction to the movement of the user 99.
[0059] The 3D video playback device 300 displays two images with a parallax difference to each of the user's 99's left and right eyes. The user 99 can perceive the three-dimensional position of an object on the image based on the parallax difference between the displayed images. Note that when the user 99 uses the audio playback system 100 with their eyes closed, for example, when using it to play healing sounds for sleep induction, the 3D video playback device 300 does not need to be used at the same time. In other words, the 3D video playback device 300 is not an essential component of the present disclosure. In addition to dedicated video display devices, the 3D video playback device 300 may also be a general-purpose mobile terminal owned by the user 99, such as a smartphone or tablet device.
[0060] Such general-purpose mobile terminals are equipped with not only a display for displaying images but also various sensors for detecting the terminal's posture and movement. Furthermore, they are also equipped with a processor for information processing, and are capable of connecting to a network to transmit and receive information to and from a server device such as a cloud server. In other words, the 3D video playback device 300 and the audio playback system 100 can be realized by combining a smartphone with general-purpose headphones or the like that do not have an information processing function.
[0061] As in this example, the head movement detection function, the video presentation function, the video information processing function for presentation, the sound presentation function, and the sound information processing function for presentation may be appropriately arranged in one or more devices to realize the 3D video reproduction device 300 and the sound reproduction system 100. If the 3D video reproduction device 300 is not required, it is sufficient to appropriately arrange the head movement detection function, the sound presentation function, and the sound information processing function for presentation in one or more devices. For example, the sound reproduction system 100 can be realized by a processing device such as a computer or smartphone having a sound information processing function for presentation, and headphones or the like having a head movement detection function and a sound presentation function.
[0062] The sound reproduction system 100 is a sound presentation device that is worn on the head of the user 99. Therefore, the sound reproduction system 100 moves integrally with the head of the user 99. For example, the sound reproduction system 100 in this embodiment is a so-called over-ear headphone type device. Note that there are no particular limitations on the form of the sound reproduction system 100, and it may be, for example, two earplug-type devices that are worn independently on the left and right ears of the user 99.
[0063] The sound reproduction system 100 changes the sound presented in accordance with the movement of the head of the user 99, thereby making the user 99 perceive as if he or she is moving his or her head within the three-dimensional sound field. For this reason, as described above, the sound reproduction system 100 moves the three-dimensional sound field in the opposite direction to the movement of the user 99.
[0064] Here, when the user 99 moves within the three-dimensional sound field, the position of the sound source object relative to the position of the user 99 within the three-dimensional sound field changes. As a result, each time the user 99 moves, it is necessary to perform calculation processing based on the positions of the sound source object and the user 99 to generate an output sound signal for playback. Since such processing typically requires a huge amount of processing, in the present disclosure, a panning process is applied as one type of conversion processing to represent the reproduced sound as a representative sound from a representative point, in order to reduce the amount of processing. As a result, it is possible to allow the user 99 to perceive the reproduced sound from the sound source object simply by convolving a head-related transfer function with the representative sound. Hereinafter, in the present embodiment, a case will be described in which panning processing is used as an example of conversion processing. However, the conversion processing is not limited to panning processing, and any conversion processing can be applied as long as the conversion processing is expected to reduce the amount of processing, depending on the conditions. Furthermore, the panning process will be described using specific examples, but the panning process is not limited to the specific example described below, and existing panning process techniques such as VBAP (Vector Based Amplitude Panning), DBAP (Distance Based Amplitude Panning), and Ambisonics can also be applied.
[0065] [Configuration] Next, the configuration of the sound reproduction system 100 according to this embodiment will be described with reference to Fig. 2. Fig. 2 is a block diagram showing the functional configuration of the sound reproduction system according to this embodiment.
[0066] As shown in FIG. 2, the sound reproduction system 100 according to this embodiment includes an information processing device 101, a communication module 102, a detector 103, a driver 104, and a database 105.
[0067] The information processing device 101 is an arithmetic device for performing various signal processing in the sound reproduction system 100. The information processing device 101 includes a processor and a memory, such as a computer, and is realized by the processor executing a program stored in the memory. Execution of this program provides functions related to each functional unit described below.
[0068] The information processing device 101 includes an acquisition unit 111, a path calculation unit 121, an output sound generation unit 131, and a signal output unit 141. Details of each functional unit included in the information processing device 101 will be described below together with details of the configuration other than the information processing device 101.
[0069] The communication module 102 is an interface device for accepting input of sound information to the sound reproduction system 100. The communication module 102 includes, for example, an antenna and a signal converter, and receives sound information from an external device via wireless communication. More specifically, the communication module 102 receives, using the antenna, a wireless signal representing sound information converted into a format for wireless communication, and reconverts the wireless signal into sound information using the signal converter. In this way, the sound reproduction system 100 acquires sound information from the external device via wireless communication. The sound information acquired by the communication module 102 is acquired by the acquisition unit 111. In this way, the acquisition unit 111 is an example of a sound acquisition unit. The sound information is input to the information processing device 101 in the above manner. Note that communication between the sound reproduction system 100 and the external device may be performed via wired communication.
[0070] The sound information acquired by the sound reproduction system 100 is encoded in a predetermined format, such as MPEG-H 3D Audio (ISO / IEC 23008-3). As an example, the encoded sound information includes information about the sound reproduced by the sound reproduction system 100 and information about the localization position when the sound image of the sound is localized at a predetermined position in a three-dimensional sound field (i.e., perceived as sound coming from a predetermined direction). The sound information can also be interpreted as information about a sound source object. In other words, the sound information includes the position of the sound source object in the three-dimensional sound field and the sound produced by the sound source object. The sound information may also include a flag for determining whether or not to apply panning processing. This flag will be described later.
[0071] As described above, sound information is obtained as input data, and includes an audio signal (acoustic signal), which is information about the reproduced sound, and other information, such as information about the position of a sound source object in a three-dimensional sound field. The other information may also include information for defining a three-dimensional sound field. Therefore, the other information may be collectively referred to as information about space (spatial information), including information about the position of a sound source object and information for defining a three-dimensional sound field. When the audio signal is viewed as the main focus, the input data can be said to be sound information in which other information (metadata) is attached to the audio signal. When the spatial information is viewed as the main focus, the input data can be said to be information in which the audio signal is attached to the spatial information. Alternatively, since the input data has both aspects, the input data can also be considered as sound spatial information.
[0072] As a specific example, the sound information includes information about a plurality of sounds including a first reproduced sound and a second reproduced sound, and the sound images of the respective sounds when reproduced are localized so that they are perceived as coming from different positions in a three-dimensional sound field. Therefore, the sound source object of the first reproduced sound is localized at a first position in the three-dimensional sound field, and the sound source object of the second reproduced sound is localized at a second position in the three-dimensional sound field. Thus, the sound information may include a plurality of sounds. In other words, the sound information may include a plurality of audio signals corresponding to the first reproduced sound and the second reproduced sound, respectively, and the positions of a plurality of sound source objects at first and second positions that correspond one-to-one to the plurality of audio signals.
[0073] FIG. 3 is a diagram illustrating an example of an audio signal according to an embodiment. For example, as shown in (a) of FIG. 3, the audio information may include an audio signal of a first direct sound arriving from a first position (from a first direction) to the position of the user 99 and an audio signal of a second direct sound arriving from a second position (from a second direction) to the position of the user 99. The immediately acquired audio information may include only information about the reproduced sound. In this case, information about a predetermined position may be acquired separately, and subsequent processing may be performed once the information is collected. Alternatively, the audio information may include multiple audio signals and the position of a sound source object that corresponds to the multiple audio signals in a many-to-one relationship. For example, such audio information is used in a situation where multiple reproduced sounds are generated from a certain sound source object. For example, each of the multiple audio signals corresponds to a direct sound that arrives directly from the position of the sound source object to the position of the user 99 and a secondary sound (sound generated by indirect propagation) that accompanies the direct sound and arrives via a path different from the direct sound.
[0074] For example, as shown in (b) of FIG. 3 , the sound information immediately after acquisition includes an audio signal related to the direct sound, and is converted into sound information including audio signals such as reverberation, primary reflection, and diffraction through a conversion process that calculates secondary sounds. This conversion process that calculates the secondary sounds uses information on the spatial environment conditions of the three-dimensional sound field (e.g., the position, reflection, and diffraction characteristics of objects in the three-dimensional sound field). Thus, secondary sounds are computationally generated from sound information related to a single reproduced sound based on the spatial environment conditions of the three-dimensional sound field. Therefore, the secondary sounds are not included in the sound information immediately after acquisition, and sound information including these secondary sounds is generated through the conversion process that calculates the secondary sounds. Another secondary sound may be generated from one secondary sound through its propagation. Note that the information on the spatial environment conditions is part of the spatial information and may be acquired together with the audio signal from the input sound information. Alternatively, the audio signal and the spatial information may be acquired separately. That is, the sound information may be acquired from a single file or bitstream, or may be acquired separately by dividing it into multiple files or bitstreams. For example, the audio signal and the spatial information may be acquired from separate files or bitstreams, or the audio signal and the spatial information may each be acquired from a plurality of files or bitstreams.
[0075] As described above, there is no particular limitation on the form of the input sound information, and the sound reproduction system 100 may be provided with an acquisition unit 111 that can accommodate various forms of sound information.
[0076] An example of the acquisition unit 111 will now be described with reference to Fig. 4. Fig. 4 is a block diagram showing the functional configuration of the acquisition unit according to the embodiment. As shown in Fig. 4, the acquisition unit 111 in this embodiment includes, for example, an encoded sound information input unit 112, a decoding processing unit 113, and a sensing information input unit 114.
[0077] The encoded sound information input unit 112 is a processing unit to which the coded (in other words, encoded) sound information acquired by the acquisition unit 111 is input. The encoded sound information input unit 112 outputs the input sound information to the decoding processing unit 113. The decoding processing unit 113 is a processing unit that decodes (in other words, decodes) the sound information output from the encoded sound information input unit 112 to generate a playback sound and the position of a sound source object included in the sound information in a format that is used for subsequent processing. The sensing information input unit 114 will be described below together with the function of the detector 103.
[0078] The detector 103 is a device for detecting the speed of movement of the head of the user 99. The detector 103 is configured by combining various sensors used for detecting movement, such as a gyro sensor and an acceleration sensor. In this embodiment, the detector 103 is built into the sound reproduction system 100. However, the detector 103 may be built into an external device, such as a three-dimensional image reproduction device 300 that operates in response to the movement of the head of the user 99 in the same way as the sound reproduction system 100. In this case, the detector 103 does not need to be included in the sound reproduction system 100. Alternatively, the detector 103 may be an external imaging device or the like that captures the movement of the head of the user 99 and detects the movement of the user 99 by processing the captured image.
[0079] The detector 103 is, for example, fixed integrally to the housing of the sound reproduction system 100 and detects the speed of movement of the housing. After the sound reproduction system 100 including the housing is worn by the user 99, it moves integrally with the head of the user 99, and as a result, the detector 103 can detect the speed of movement of the head of the user 99.
[0080] The detector 103 may detect, for example, the amount of head movement of the user 99 as the amount of rotation about at least one of three axes that are orthogonal to each other in three-dimensional space as the rotation axis, or may detect the amount of displacement about at least one of the three axes as the displacement direction. Furthermore, the detector 103 may detect both the amount of rotation and the amount of displacement as the amount of head movement of the user 99.
[0081] The sensing information input unit 114 acquires the movement speed of the user 99's head from the detector 103. More specifically, the sensing information input unit 114 acquires the amount of head movement of the user 99 detected by the detector 103 per unit time as the movement speed. In this way, the sensing information input unit 114 acquires at least one of the rotation speed and the displacement speed from the detector 103. The amount of head movement of the user 99 acquired here is used to determine the position and posture (in other words, coordinates and orientation) of the user 99 in the three-dimensional sound field. Therefore, the acquisition unit 111 also functions as a position acquisition unit via the sensing information input unit 114. In the sound reproduction system 100, the relative position of the sound image object with respect to the user 99 is determined based on the determined coordinates and orientation of the user 99, and sound is reproduced. Specifically, the above functions are realized by the path calculation unit 121 and the output sound generation unit 131.
[0082] The path calculation unit 121 includes an arrival direction calculation function that calculates the relative arrival direction of the reproduced sound from the position of the sound source object to the position of the user 99 based on the determined coordinates and orientation of the user 99, and the conversion process that calculates the secondary sound described above. Therefore, the path calculation unit 121 includes a function that calculates a propagation path from the sound source object and calculates the secondary sound and the arrival direction of the secondary sound that arrives at the position of the user 99 through indirect propagation of the reproduced sound according to the calculated propagation path of the reproduced sound. The arrival direction of the secondary sound includes additional information, such as what object the sound is reflected from in the case of a reflected sound and the attenuation rate at the time of reflection. The additional information is included in the arrival direction of the secondary sound calculated from the input sound information. In other words, the additional information is computationally generated and acquired from the sound information.
[0083] To summarize the spatial information, the spatial information includes further information such as the spatial position of a sound source object in a space (three-dimensional sound field) (information on the position of the sound source object), sound reflection and diffraction characteristics at the sound source object (together with information on the conditions of the spatial environment), and the size of the three-dimensional sound field. Based on the spatial information, the path calculation unit 121 generates secondary sounds depending on which sound source object the reproduced sound is reflected or diffracted by, and calculates, as additional information, the direction from which the secondary sounds arrive and the volume of the secondary sounds after attenuation by reflection or diffraction. The sound information (input data) includes spatial information in the form of audio signals and accompanying metadata. As described above, the spatial information includes, as information other than the audio signal, information necessary to turn the sound into stereophonic sound and position the sound source object in a three-dimensional sound field, and / or information used to calculate information necessary to turn the sound into stereophonic sound and position the sound source object in a three-dimensional sound field.
[0084] The path calculation unit 121 may be realized by any processing as long as it can calculate the arrival direction of the reproduced sound when the reproduced sound reaches the user as a direct sound and can calculate the arrival direction of a secondary sound that arrives at the position of the user 99 due to secondary propagation of the reproduced sound. The path calculation unit 121 determines from which direction in the three-dimensional sound field the reproduced sound and the secondary sound are to be perceived by the user 99 as coming from, based on the coordinates and orientation of the user 99, and processes the sound information so that the output sound signal is perceived as such a sound when it is reproduced.
[0085] The output sound generating unit 131 is a processing unit that processes information about the reproduced sound included in the sound information to generate an output sound signal.
[0086] An example of the output sound generation unit 131 will now be described with reference to Fig. 5. Fig. 5 is a block diagram showing the functional configuration of the output sound generation unit according to the embodiment. As shown in Fig. 5, the output sound generation unit 131 in the present embodiment includes, for example, a generation unit 134 and a synthesis unit 135.
[0087] The generation unit 134 is a processing unit used when applying a panning process to perform a conversion process to convert a playback sound into a representative sound, and then convolving a head-related transfer function with the converted representative sound. The generation unit 134 acquires the playback sound and the position of the representative point, and performs a conversion process to convert the playback sound into a representative sound for reproducing the playback sound using sound from the representative point. Note that the generation unit 134 has the same function as the panning unit described later in FIG. 15 .
[0088] For example, if a sound source object is located midway between two representative points, a sound is generated so that the same sound as the playback sound is emitted from each of the two representative points. In other words, the playback sound is distributed to the two representative points. Then, a representative sound can be generated by adjusting the gain of the generated sound to match the position of the sound source object. Conversion of the playback sound to a representative sound is not limited to this example. For example, as described below, conversion of the playback sound to a representative sound may be performed by time shift adjustment and gain adjustment, or any other existing conversion may be used as long as the playback sound can be converted to a representative sound that is reproduced as a sound from a representative point. Furthermore, in this specification, the conversion process of the playback sound to a representative sound may be interpreted as a process of distributing the playback sound to representative points (representative directions). Specifically, sound signals of the playback sound associated with the positions of each sound source object are distributed to the positions of the representative points, and a representative sound arriving from the representative points (representative directions) to the listener is generated. Here, the representative direction refers to the direction of the representative point as seen from the listener, or the direction of the listener as seen from the representative point. An example of the conversion that performs the time shift adjustment and the gain adjustment will be described later. The generation unit 134 acquires the same number of representative sounds as the number of representative points obtained by the conversion and head-related transfer functions corresponding to the representative directions from each representative point to the position of the user 99, and performs a convolution process of the acquired head-related transfer functions on the representative sounds to generate a sound signal.
[0089] That is, the panning unit performs panning to represent the sound source by panning sounds from a specific representative direction based on the sound source directions of multiple sound sources (target signals) acquired by the path calculation unit 121, by time shifting the sound sources and adjusting their gains. Specifically, the panning unit synthesizes the sound source (target signal) by panning in a representative direction that approximates the sound source direction of the sound source. As a result, the panning unit generates an HRIR for the sound source direction equivalently. Here, in this embodiment, "equivalent" and "equivalently" refer to signals with an error below a specific level and substantially similar signals, as shown in the examples described below. Specifically, the panning unit generates an HRIR for the sound source direction equivalently by panning the sound source by synthesizing HRIRs for several directions that are closest to the sound source direction of the sound source or that are most similar to the HRIR for the sound source direction. In this embodiment, this direction will be referred to as a "specific representative direction" (hereinafter simply referred to as a "representative direction") described below. This reduces the amount of calculation required to generate the ear signal.
[0090] That is, the panning unit synthesizes a sound image from multiple sound sources using sounds from multiple representative directions. For example, two or three representative directions can be used, but the number of representative directions is not limited to this. Specifically, the panning unit can group together the sound sources into a number of representative points that is fewer than the number of sound sources, and synthesize a sound image using only the HRIRs of the representative directions for these representative points.
[0091] At this time, the panning unit calculates a time shift (delay) that maximizes the cross-correlation between the HRIR in the sound source direction and the HRIR in the representative direction. The time shift obtained here, or a time shift obtained by adding a negative sign to this time shift, is applied to the sound source, and the subsequent processing is performed assuming that the signal after the time shift is in the representative direction.
[0092] This time shift may be a time shift shorter than the sampling frequency (a shift in which the sample position is expressed by a decimal number; hereinafter, referred to as a "decimal shift"). This decimal shift can be performed by oversampling.
[0093] Here, the panning unit applies a gain to the signal of the representative direction obtained by time-shifting the sound source, and calculates the sum of the values calculated for each representative point convolved with the HRIR at each representative point, thereby synthesizing a signal equivalent to the sound source convolved with the HRIR of the sound source direction.
[0094] On the other hand, when synthesizing the HRIR (vector) of the sound source direction by the sum of the HRIR (vector) of the representative direction, the panning unit may calculate the gain by orthogonalizing the error signal vector between the synthesized HRIR (vector) and the HRIR (vector) of the sound source direction to the HRIR (vector) of the representative direction. Note that the HRIR (vector) is a time waveform of the HRIR that is considered to be a vector. Hereinafter, this HRIR (vector) will also be referred to as an "HRIR vector."
[0095] The panning unit corrects this gain so that the energy balance of the HRIRs for the left and right ears from the sound source position is maintained in the HRIR substantially synthesized by panning using HRIRs from multiple representative points. In other words, the panning unit may correct the gain so that the energy balance of the HRIRs for the left and right ears of the listener from the sound source is maintained in the HRIR substantially synthesized by panning.
[0096] In this embodiment, the panning unit can calculate the gain value of the HRIR gain in the representative direction and the time shift value corresponding to the time shift of the HRIR for each sound source direction of the sound source, and store them in the HRIR table 200 described later.
[0097] The panning unit then time-shifts each sound source using a time shift value and gain value corresponding to the sound source direction of each sound source, multiplies the gain, and sums them to generate a sum signal. The panning unit treats this sum signal as being present at the position of the representative point. The panning unit can convolve the HRIR at the position of the representative point with this sum signal to generate a signal at the listener's ear.
[0098] The synthesis unit 135 generates an output sound signal. The synthesis unit 135 may perform EQ adjustment on the sound signal. Specifically, in the panning process, EQ adjustment may be performed to increase the gain of a high-frequency domain that is likely to be attenuated, thereby emphasizing this high-frequency domain. Therefore, the synthesis unit 135 functions as an EQ adjustment unit. Note that, when there are multiple sound signals, the EQ adjustment performed by the synthesis unit 135 may be performed on only some or all of the multiple sound signals.
[0099] Referring again to FIG. 2 , the output sound generation unit 131 acquires a head-related transfer function used to generate an output sound signal from the database 105. The database 105 is an information storage device that functions both as a storage device for storing information and as a storage controller that reads out the stored information and outputs it to an external component. The database 105 stores a head-related transfer function for each direction of arrival of the sound from the user 99. The head-related transfer functions included in the database 105 are a set of general-purpose head-related transfer functions that can be used by everyone, a set of head-related transfer functions optimized for each individual user 99, or a set of head-related transfer functions that are publicly available. The database 105 receives an inquiry from the output sound generation unit 131 using the direction of arrival as a query, and outputs a head-related transfer function corresponding to the direction of arrival to the output sound generation unit 131. In addition, the output sound generation unit 131 may output the entire set of head-related transfer functions, or may output the characteristics of the set of head-related transfer functions itself.
[0100] The signal output unit 141 is a functional unit that outputs the generated output sound signal to the driver 104. The signal output unit 141 generates a waveform signal by performing signal conversion from a digital signal to an analog signal based on the output sound signal, and then causes the driver 104 to generate sound waves based on the waveform signal, thereby presenting sound to the user 99. The driver 104 includes, for example, a diaphragm and a drive mechanism such as a magnet and a voice coil. The driver 104 operates the drive mechanism in response to the waveform signal, causing the drive mechanism to vibrate the diaphragm. In this way, the driver 104 generates sound waves by vibrating the diaphragm in response to the output sound signal (this means "reproducing" the output sound signal; in other words, "reproducing" does not include perception by the user 99). The sound waves propagate through the air and reach the ears of the user 99, and the user 99 perceives the sound.
[0101] [Another Configuration Example] In the above example, the sound reproduction system 100 according to the present embodiment is a sound presentation device, and has been described as including an information processing device 101, a communication module 102, a detector 103, a database 105, and a driver 104. However, the functions of the sound reproduction system 100 may be realized by a plurality of devices or by a single device. Specific examples will be described using Figures 6 to 15. Figures 6 to 15 are diagrams for explaining other examples of the sound reproduction system according to the embodiment.
[0102] For example, the information processing device 601 may be included in the audio presentation device 602, and the audio presentation device 602 may perform both acoustic processing and sound presentation. Alternatively, the information processing device 601 and the audio presentation device 602 may share the acoustic processing described in the present disclosure, or a server connected to the information processing device 601 or the audio presentation device 602 via a network may perform part or all of the acoustic processing described in the present disclosure.
[0103] In the above explanation, the information processing device 601 is referred to as such, but when the information processing device 601 performs acoustic processing by decoding a bitstream generated by encoding at least a portion of the data of an audio signal or spatial information used for acoustic processing, the information processing device 601 may be referred to as a decoding device, and the acoustic reproduction system 100 (i.e., the stereophonic sound reproduction system 600 in the figure) may be referred to as a decoding processing system.
[0104] Here, an example will be described in which the sound reproduction system 100 functions as a decoding processing system.
[0105] <Example of Encoding Device> FIG. 7 is a functional block diagram showing the configuration of an encoding device 700 that is an example of an encoding device according to the present disclosure.
[0106] Input data 701 is data to be coded, including spatial information and / or an audio signal, that is input to an encoder 702. Details of the spatial information will be described later.
[0107] The encoder 702 encodes the input data 701 to generate encoded data 703. The encoded data 703 is, for example, a bit stream generated by the encoding process.
[0108] The memory 704 stores the encoded data 703. The memory 704 may be, for example, a hard disk or a solid-state drive (SSD), or may be any other storage device.
[0109] In the above description, a bitstream generated by an encoding process is given as an example of the encoded data 703 stored in memory 704, but data other than a bitstream may also be used. For example, the encoding device 700 may convert a bitstream into a predetermined data format and store the converted data in memory 704. The converted data may be, for example, a file or multiplexed stream storing one or more bitstreams. Here, the file may have a file format such as ISOBMFF (ISO Base Media File Format). The encoded data 703 may also be in the form of multiple packets generated by dividing the bitstream or file. When converting the bitstream generated by the encoder 702 into data other than the bitstream, the encoding device 700 may be provided with a conversion unit (not shown), or the conversion process may be performed by a CPU (Central Processing Unit).
[0110] <Example of Decoding Device> FIG. 8 is a functional block diagram showing the configuration of a decoding device 800 that is an example of a decoding device according to the present disclosure.
[0111] The memory 804 stores, for example, the same data as the coded data 703 generated by the coding device 700. The memory 804 reads out the stored data and inputs it as input data 803 to the decoder 802. The input data 803 is, for example, a bitstream to be decoded. The memory 804 may be, for example, a hard disk or an SSD, or may be another storage device.
[0112] Note that the decoding device 800 may not use the data stored in the memory 804 as input data 803 as is, but may convert the read data and generate converted data as input data 803. The data before conversion may be, for example, multiplexed data storing one or more bitstreams. Here, the multiplexed data may be, for example, a file having a file format such as ISOBMFF. The data before conversion may also be in the form of multiple packets generated by dividing the bitstream or file. When converting data different from the bitstream read from the memory 804 into a bitstream, the decoding device 800 may be provided with a conversion unit (not shown), or the conversion process may be performed by a CPU.
[0113] Decoder 802 decodes input data 803 to generate audio signal 801 that is presented to a listener.
[0114] <Another Example of Encoding Device> Fig. 9 is a functional block diagram showing the configuration of an encoding device 900, which is another example of an encoding device of the present disclosure. In Fig. 9, components having the same functions as those in Fig. 7 are assigned the same reference numerals as those in Fig. 7, and descriptions of these components will be omitted.
[0115] The coding device 900 differs from the coding device 700 in that the coding device 700 includes a memory 704 for storing coded data 703, whereas the coding device 900 includes a transmitting unit 901 for transmitting coded data 703 to the outside.
[0116] The transmitter 901 transmits a transmission signal 902 to another device or a server based on the encoded data 703 or data in another data format generated by converting the encoded data 703. The data used to generate the transmission signal 902 is, for example, the bit stream, multiplexed data, file, or packet described in the encoding device 700.
[0117] 10 is a functional block diagram showing the configuration of a decoding device 1000, which is another example of a decoding device according to the present disclosure. In Fig. 10, components having the same functions as those in Fig. 8 are assigned the same reference numerals, and description of these components will be omitted.
[0118] The decoding device 1000 differs from the decoding device 800 in that the decoding device 800 includes a memory 804 for reading out input data 803, whereas the decoding device 1000 includes a receiving unit 1001 for receiving input data 803 from the outside.
[0119] The receiving unit 1001 receives a received signal 1002, acquires received data, and outputs input data 803 to be input to the decoder 802. The received data may be the same as the input data 803 to be input to the decoder 802, or may be data in a data format different from that of the input data 803. If the received data is data in a data format different from that of the input data 803, the receiving unit 1001 may convert the received data into the input data 803, or a conversion unit or CPU (not shown) included in the decoding device 1000 may convert the received data into the input data 803. The received data is, for example, a bit stream, multiplexed data, a file, or a packet, as described in the encoding device 900.
[0120] <Functional Description of Decoder> FIG. 11 is a functional block diagram showing the configuration of a decoder 1100, which is an example of the decoder 802 in FIG. 8 or 10.
[0121] The input data 803 is an encoded bitstream, and includes encoded audio data, which is an encoded audio signal, and metadata used for acoustic processing.
[0122] The spatial information management unit 1101 acquires metadata included in the input data 803 and analyzes the metadata. The metadata includes information describing elements that act on sounds arranged in a sound space. The spatial information management unit 1101 manages spatial information necessary for acoustic processing obtained by analyzing the metadata and provides the spatial information to the rendering unit 1103. Note that, although the information used for acoustic processing is referred to as spatial information in this disclosure, it may be called by other names. The information used for acoustic processing may be called, for example, sound space information or scene information. Furthermore, when the information used for acoustic processing changes over time, the spatial information input to the rendering unit 1103 may be called a space state, a sound space state, a scene state, or the like.
[0123] Furthermore, spatial information may be managed for each sound space or for each scene. For example, when different rooms are represented as virtual spaces, each room may be managed as a scene of a different sound space, or even if the same space is represented, spatial information may be managed as different scenes depending on the situation being represented. In managing spatial information, an identifier for identifying each piece of spatial information may be assigned. The spatial information data may be included in a bitstream, which is one form of input data 803, or the bitstream may include an identifier for the spatial information, and the spatial information data may be acquired from a source other than the bitstream. When the bitstream includes only the identifier for the spatial information, the identifier for the spatial information may be used during rendering to acquire the spatial information data stored in the memory of the acoustic signal processing device or an external server as input data.
[0124] Note that the information managed by the spatial information management unit 1101 is not limited to information included in the bitstream. For example, the input data 803 may include data indicating the characteristics or structure of a space acquired from a software application or server providing VR or AR, as data not included in the bitstream. Furthermore, for example, the input data 803 may include data indicating the characteristics or position of a listener or object, as data not included in the bitstream. Furthermore, the input data 803 may include, as information indicating the position of the listener, information acquired by a sensor provided in a terminal including a decoding device, or information indicating the position of the terminal estimated based on information acquired by the sensor. In other words, the spatial information management unit 1101 may communicate with an external system or server to acquire spatial information and the position of the listener. Furthermore, the spatial information management unit 1101 may acquire clock synchronization information from an external system and execute a process of synchronizing with the clock of the rendering unit 1103. Note that the space in the above description may be a virtually formed space, i.e., a VR space, or may be a real space or a virtual space corresponding to a real space, i.e., an AR space or an MR (Mixed Reality) space. The virtual space may also be called a sound field or a sound space. Furthermore, the information indicating a position in the above description may be information such as coordinate values indicating a position within a space, information indicating a relative position with respect to a predetermined reference position, or information indicating the movement or acceleration of a position within a space.
[0125] The audio data decoder 1102 decodes the encoded audio data included in the input data 803 to obtain an audio signal.
[0126] The encoded audio data acquired by the stereophonic sound reproduction system 600 is a bitstream encoded in a predetermined format such as MPEG-H 3D Audio (ISO / IEC 23008-3). Note that MPEG-H 3D Audio is merely one example of an encoding method that can be used to generate the encoded audio data included in the bitstream, and the encoded audio data may be included in a bitstream encoded in another encoding method. For example, the encoding method used may be a lossy codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), or Vorbis, or a lossless codec such as ALAC (Apple Lossless Audio Codec) or FLAC (Free Lossless Audio Codec), or any other encoding method may be used. For example, PCM (Pulse Code Modulation) data may be a type of encoded audio data. In this case, the decoding process may be, for example, a process of converting an N-bit binary number into a number format (e.g., floating-point format) that can be processed by the rendering unit 1103, where the number of quantization bits of the PCM data is N.
[0127] The rendering unit 1103 receives an audio signal and spatial information as input, performs acoustic processing on the audio signal using the spatial information, and outputs an audio signal 801 after acoustic processing.
[0128] Before starting rendering, the spatial information management unit 1101 reads metadata of the input signal, detects rendering items such as objects or sounds defined in the spatial information, and transmits the detected items to the rendering unit 1103. After starting rendering, the spatial information management unit 1101 grasps changes over time in the spatial information and the position of the listener, and updates and manages the spatial information. The spatial information management unit 1101 then transmits the updated spatial information to the rendering unit 1103. The rendering unit 1103 generates and outputs an audio signal to which acoustic processing has been applied based on the audio signal included in the input data and the spatial information received from the spatial information management unit 1101.
[0129] The spatial information update process and the audio signal output process with added acoustic processing may be executed in the same thread, or the spatial information management unit 1101 and the rendering unit 1103 may be allocated to independent threads. When the spatial information update process and the audio signal output process with added acoustic processing are executed in different threads, the thread startup frequency may be set individually, or the processes may be executed in parallel.
[0130] By having the spatial information management unit 1101 and the rendering unit 1103 execute their processes in different, independent threads, computational resources can be preferentially allocated to the rendering unit 1103. Therefore, in the case of sound output processing in which even a slight delay cannot be tolerated, for example, sound output processing in which a delay of even one sample (0.02 msec) would cause a popping noise, can be safely performed. In this case, the allocation of computational resources to the spatial information management unit 1101 is limited. However, compared to audio signal output processing, updating spatial information is a less frequent process (e.g., a process such as updating the listener's facial orientation). Therefore, unlike audio signal output processing, updating spatial information does not necessarily require an instantaneous response, and therefore limiting the allocation of computational resources does not significantly affect the acoustic quality provided to the listener.
[0131] The space information may be updated periodically at preset times or intervals, or when preset conditions are met. The space information may be updated manually by a listener or a sound space manager, or may be triggered by a change in an external system. For example, when a listener operates a controller to instantly warp the position of their avatar, instantly advance or rewind the time, or when a virtual space manager suddenly changes the environment of the space, the thread in which the space information management unit 1101 is located may be started as a one-off interrupt process in addition to being started periodically.
[0132] The role of the information update thread that executes the spatial information update process is, for example, to update the position or orientation of the listener's avatar placed in the virtual space based on the position or orientation of the VR goggles worn by the listener, and to update the position of objects moving in the virtual space. These tasks are handled within a processing thread that runs relatively infrequently, on the order of several tens of Hz. Processing to reflect the properties of direct sound may be performed in such an infrequently occurring processing thread. This is because the properties of direct sound change less frequently than the frequency of audio processing frames for audio output. Doing so can relatively reduce the computational load of the process and also avoid the risk of pulsive noise occurring when information is updated at an unnecessarily fast frequency.
[0133] FIG. 12 is a functional block diagram showing the configuration of a decoder 1200, which is another example of the decoder 802 in FIG. 8 or 10.
[0134] Figure 12 differs from Figure 11 in that the input data 803 includes an unencoded audio signal rather than encoded audio data. The input data 803 includes a bitstream including metadata and an audio signal.
[0135] The spatial information management unit 1201 is the same as the spatial information management unit 1101 in FIG. 11, and therefore a description thereof will be omitted.
[0136] The rendering unit 1202 is the same as the rendering unit 1103 in FIG. 11, and therefore a description thereof will be omitted.
[0137] In the above description, the configuration in Fig. 12 is called a decoder, but it may also be called an audio processing unit that performs audio processing. Furthermore, a device including an audio processing unit may also be called an audio processing device rather than a decoding device. Furthermore, the audio signal processing device (information processing device 601) may also be called an audio processing device.
[0138] <Physical Configuration of Encoding Apparatus> Fig. 13 is a diagram showing an example of the physical configuration of an encoding apparatus. The encoding apparatus shown in Fig. 13 is an example of the encoding apparatuses 700 and 900 described above.
[0139] The encoding device of FIG. 13 includes a processor, a memory, and a communication IF.
[0140] The processor may be, for example, a CPU (Central Processing Unit), a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit), and the CPU, DSP, or GPU may execute a program stored in a memory to perform the encoding process of the present disclosure. Alternatively, the processor may be a dedicated circuit that performs signal processing on an audio signal, including the encoding process of the present disclosure.
[0141] The memory may be configured, for example, by RAM (Random Access Memory) or ROM (Read Only Memory). The memory may also include a magnetic storage medium such as a hard disk or a semiconductor memory such as an SSD (Solid State Drive). The term "memory" may also refer to an internal memory built into a CPU or GPU.
[0142] The communication IF (Interface) is a communication module compatible with a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The encoding device has a function of communicating with other communication devices via the communication IF and transmits an encoded bitstream.
[0143] The communication module is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. In the above example, Bluetooth (registered trademark) or WIGIG (registered trademark) is used as the communication method, but other communication methods such as LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark) may also be supported. Furthermore, the communication IF may be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface) instead of the wireless communication method described above.
[0144] <Physical configuration of audio signal processing device> Fig. 14 is a diagram showing an example of the physical configuration of an audio signal processing device. Note that the audio signal processing device in Fig. 14 may be a decoding device. Furthermore, part of the configuration described here may be provided in the audio presentation device 602. Furthermore, the audio signal processing device shown in Fig. 14 is an example of the audio signal processing device 601 described above.
[0145] The acoustic signal processing device of FIG. 14 includes a processor, a memory, a communication IF, a sensor, and a speaker.
[0146] The processor may be, for example, a CPU (Central Processing Unit), a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit), and the CPU, DSP, or GPU may execute a program stored in a memory to perform the audio processing or decoding processing of the present disclosure. Alternatively, the processor may be a dedicated circuit that performs signal processing on audio signals, including the audio processing of the present disclosure.
[0147] The memory may be configured, for example, by RAM (Random Access Memory) or ROM (Read Only Memory). The memory may also include a magnetic storage medium such as a hard disk or a semiconductor memory such as an SSD (Solid State Drive). The term "memory" may also refer to an internal memory built into a CPU or GPU.
[0148] The communication IF (Interface) is a communication module compatible with a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The acoustic signal processing device shown in Fig. 14 has a function of communicating with other communication devices via the communication IF and acquires a bitstream to be decoded. The acquired bitstream is stored in a memory, for example.
[0149] The communication module is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. In the above example, Bluetooth (registered trademark) or WIGIG (registered trademark) is used as the communication method, but other communication methods such as LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark) may also be supported. Furthermore, the communication IF may be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface) instead of the wireless communication method described above.
[0150] The sensor performs sensing to estimate the position or orientation of the listener. Specifically, the sensor estimates the position and / or orientation of the listener based on one or more detection results of the position, orientation, movement, velocity, angular velocity, or acceleration of a part or the entire body of the listener, such as the head, and generates position information indicating the position and / or orientation of the listener. Note that the position information may be information indicating the position and / or orientation of the listener in real space, or information indicating a displacement of the position and / or orientation of the listener based on the position and / or orientation of the listener at a predetermined time. Furthermore, the position information may be information indicating the position and / or orientation relative to the stereophonic sound reproduction system or an external device equipped with the sensor.
[0151] The sensor may be, for example, an imaging device such as a camera or a ranging device such as LiDAR (Light Detection and Ranging), and may capture an image of the listener's head movement and detect the movement of the listener's head by processing the captured image. Alternatively, the sensor may be a device that performs position estimation using wireless signals of any frequency band, such as millimeter waves.
[0152] The audio signal processing device shown in Fig. 14 may acquire position information from an external device equipped with a sensor via a communication IF. In this case, the audio signal processing device does not need to include a sensor. Here, the external device is, for example, the audio presentation device 602 described in Fig. 6 or a 3D video playback device worn on the listener's head. In this case, the sensor is configured by combining various sensors such as a gyro sensor and an acceleration sensor.
[0153] The sensor may, for example, detect the angular velocity of rotation around at least one of three mutually perpendicular axes in the sound space as the axis of rotation as the speed of movement of the listener's head, or may detect the acceleration of displacement with at least one of the three axes as the direction of displacement.
[0154] For example, the sensor may detect the amount of rotation about at least one of three mutually orthogonal axes in the sound space as the rotation axis, or the amount of displacement about at least one of the three axes as the displacement direction, as the amount of movement of the listener's head. Specifically, the sensor detects 6 DoF (position (x, y, z) and angle (yaw, pitch, roll)) as the position of the listener. The sensor is configured by combining various sensors used for detecting movement, such as a gyro sensor and an acceleration sensor.
[0155] The sensor may be any device capable of detecting the position of the listener, such as a camera or a GPS (Global Positioning System) receiver. Alternatively, the sensor may use location information obtained by performing self-position estimation using LiDAR (Laser Imaging Detection and Ranging). For example, when the audio signal reproduction system is implemented by a smartphone, the sensor is built into the smartphone.
[0156] The sensor may also include a temperature sensor such as a thermocouple that detects the temperature of the acoustic signal processing device shown in Figure 14, and a sensor that detects the remaining charge of a battery provided in or connected to the acoustic signal processing device.
[0157] A speaker has, for example, a diaphragm, a drive mechanism such as a magnet or a voice coil, and an amplifier, and presents an audio signal after acoustic processing to a listener as sound. The speaker operates the drive mechanism in response to the audio signal (more specifically, a waveform signal indicating the waveform of the sound) amplified by the amplifier, and the drive mechanism vibrates the diaphragm. In this way, the diaphragm vibrating in response to the audio signal generates sound waves, which propagate through the air and reach the listener's ears, causing the listener to perceive the sound.
[0158] Note that, although the description has been given here of an example in which the acoustic signal processing device shown in FIG. 14 includes a speaker and presents an audio signal after acoustic processing via the speaker, the means for presenting the audio signal is not limited to the above configuration. For example, the audio signal after acoustic processing may be output to an external audio presentation device 602 connected via a communication module. Communication via the communication module may be wired or wireless. As another example, the acoustic signal processing device shown in FIG. 14 may include a terminal for outputting an analog audio signal, and a cable such as earphones may be connected to the terminal to present the audio signal from the earphones. In the above case, the audio signal may be reproduced by headphones, earphones, a head-mounted display, a neck speaker, a wearable speaker, a surround speaker composed of multiple fixed speakers, or the like, which are worn on the head or part of the body of the listener, which is the audio presentation device 602.
[0159] <Functional Description of Rendering Unit> FIG. 15 is a functional block diagram showing an example of the detailed configuration of the rendering units 1103 and 1202 in FIGS.
[0160] The rendering section is composed of an analysis section, a panning section, and a synthesis section (different from the synthesis section 135), and applies acoustic processing to the sound data contained in the input signal and outputs the result.
[0161] The information contained in the input signal will now be described.
[0162] The input signal may be composed of, for example, spatial information, sensor information, and sound data. The input signal may also include a bitstream composed of sound data and metadata (control information), in which case the metadata may include spatial information.
[0163] Spatial information is information about the sound space (three-dimensional sound field) created by the stereophonic playback system, and is composed of information about the objects included in the sound space and information about the listener. Objects include sound source objects that emit sound and act as sound sources, and non-sound-emitting objects that do not emit sound. Non-sound-emitting objects function as obstacle objects that reflect sounds emitted by sound source objects, but sound source objects may also function as obstacle objects that reflect sounds emitted by other sound source objects.
[0164] Information commonly assigned to sound source objects and non-sound generating objects includes position information, shape information, and the rate of attenuation of the volume when the object reflects sound.
[0165] The position information is expressed as coordinate values on three axes, for example, the X-axis, Y-axis, and Z-axis, in Euclidean space, but it does not necessarily have to be three-dimensional information. For example, it may be two-dimensional information expressed as coordinate values on two axes, the X-axis and the Y-axis. The position information of an object is determined by a representative position of a shape expressed by a mesh or voxels.
[0166] The shape information may include information about the surface material.
[0167] The information may also include information indicating whether the object belongs to a living thing, information indicating whether the object is a moving object, etc. If the object is a moving object, the position information may change over time, and the changed position information or the amount of change is transmitted to the rendering unit.
[0168] The information about the sound source object includes the information commonly given to the sound source object and the non-sound generating object, as well as sound data and information required to radiate the sound data into the sound space.
[0169] The sound data is data that represents the sound perceived by a listener, including information about the frequency and intensity of the sound. The sound data is typically a PCM signal, but may also be data compressed using an encoding method such as MP3. In this case, the signal must be decoded at least before it reaches the synthesis unit, so the rendering unit may include a decoding unit (not shown). Alternatively, the signal may be decoded by the audio data decoder 1102.
[0170] At least one piece of sound data may be set for one sound source object, but multiple pieces of sound data may also be set. Furthermore, identification information for identifying each piece of sound data may be assigned, and the identification information for the sound data may be stored as information about the sound source object.
[0171] The information necessary for radiating sound data into a sound space may include, for example, information on a reference volume that serves as a reference when playing back sound data, information indicating the properties (also called characteristics) of the sound data, information on the position of the sound source object, information on the orientation of the sound source object, information on the directivity of the sound emitted by the sound source object, etc. The information on the reference volume is, for example, the effective value of the amplitude value of the sound data at the sound source position when radiating the sound data into a sound space, and may be expressed as a floating-point decibel (dB) value.
[0172] For example, if the reference volume is 0 dB, it may indicate that sound is to be emitted into the sound space from the position indicated by the information about the position at the same volume as the signal level indicated by the sound data, without increasing or decreasing the volume, or if it is -6 dB, it may indicate that sound is to be emitted into the sound space from the position indicated by the information about the position with the volume of the signal level indicated by the sound data reduced to about half. These pieces of information are assigned to one piece of sound data or to multiple pieces of sound data collectively.
[0173] The information indicating the properties of the sound data may be, for example, information regarding the volume of the sound source, and may be information indicating time-series fluctuations. For example, if the sound space is a virtual conference room and the sound source is a speaker, the volume will transition intermittently over a short period of time. To put it more simply, this can be said to be alternating between sound and silence.
[0174] If the sound space is a concert hall and the sound source is a performer, the volume is maintained for a certain period of time. If the sound space is a battlefield and the sound source is an explosive, the volume of the explosion increases for a moment and then remains silent. In this way, the volume information of the sound source includes not only information about the volume of the sound but also information about the transition of the volume of the sound, and such information may be used as information indicating the properties of the sound data.
[0175] Here, the information on the transition in loudness of a sound may be data showing frequency characteristics in a time series. It may be data showing the duration of a sound section. It may be data showing a time series of the duration of a sound section and the duration of a silent section. It may be data listing multiple sets of data on the duration for which the amplitude of a sound signal can be considered steady (considered to be roughly constant) and the amplitude values of the signal during that time in a time series. It may be data on the duration for which the frequency characteristics of a sound signal can be considered steady. It may be data listing multiple sets of data on the duration for which the frequency characteristics of a sound signal can be considered steady and the frequency characteristics during that time in a time series.
[0176] The data format may be, for example, data indicating the outline of a spectrogram. Furthermore, the volume that serves as a reference for the frequency characteristics may be used as the reference volume. Information on the reference volume and information indicating the properties of the sound data may be used to calculate the volume of the direct sound or reflected sound to be perceived by the listener, as well as in a selection process for selecting whether or not to perceive the direct sound or reflected sound. Other examples of information indicating the properties of the sound data and specific uses for the selection process will be described later.
[0177] Orientation information is typically expressed using yaw, pitch, and roll. Alternatively, the roll rotation may be omitted and the information may be expressed using azimuth (yaw) and elevation (pitch). Orientation information may change over time, and if it does change, it is transmitted to the rendering unit.
[0178] The information about the listener is information about the position and orientation of the listener in sound space. The position information is expressed as a position on the XYZ axes in Euclidean space, but it does not necessarily have to be three-dimensional information and may be two-dimensional information. The information about orientation is typically expressed using yaw, pitch, and roll. Alternatively, the roll rotation may be omitted and the information may be expressed using azimuth (yaw) and elevation (pitch). The position information and orientation information may change over time, and if they change, they are transmitted to the rendering unit.
[0179] The sensor information includes the amount of rotation or displacement detected by a sensor worn by the listener and the position and orientation of the listener. The sensor information is transmitted to a rendering unit, which updates the information on the position and orientation of the listener based on the sensor information. The sensor information may be, for example, position information obtained by a mobile terminal performing self-position estimation using a GPS, a camera, or LiDAR (Laser Imaging Detection and Ranging). Furthermore, information obtained from an external source other than a sensor via a communication module may be detected as sensor information. Information indicating the temperature of the audio signal processing device and information indicating the remaining battery level may be obtained from the sensor. Computing resources (CPU capacity, memory resources, PC performance) of the audio signal processing device and the audio signal presentation device may be obtained in real time.
[0180] The analysis unit performs the same function as the acquisition unit 111 in the above example. That is, it analyzes the input signal and acquires information required by the path calculation unit 121 and the output sound generation unit 131.
[0181] The synthesis unit performs the same functions as the output sound generation unit 131 and the signal output unit 141 in the above example. Based on the audio signal of the direct sound and information on the direct sound arrival time and volume at the time of direct sound arrival calculated by the analysis unit, the synthesis unit processes the input audio signal to generate direct sound. Also, based on information on the reflected sound arrival time and volume at the time of reflected sound arrival calculated by the analysis unit, the synthesis unit processes the input audio signal to generate reflected sound. The synthesis unit synthesizes the generated direct sound and reflected sound and outputs the synthesized sound.
[0182] The panning unit performs the same function as the generation unit 134 in the above example. That is, based on the sound source directions of the multiple sound sources (target signals) acquired by the analysis unit, panning by sound from a specific representative direction is performed by time shifting the sound source and adjusting the gain, thereby performing panning to represent the sound source. The processing performed by the above-mentioned panning unit may be executed as part of pipeline processing such as that described in International Publication No. 2021 / 180938, for example.
[0183] FIG. 16 is a block diagram showing an example of the configuration for the rendering unit 1300 to perform pipeline processing.
[0184] The rendering unit 1300 in Fig. 16 includes a reverberation processing unit 1311, an early reflection processing unit 1312, a distance attenuation processing unit 1313, a selection unit 1314, a generation unit 1315, and a binaural processing unit 1316. These multiple components may be configured from the multiple components of the rendering unit shown in Fig. 15, or may be configured from at least some of the multiple components of the acoustic signal processing device shown in Fig. 14.
[0185] Pipeline processing refers to dividing the process for applying sound effects into multiple processes and executing the multiple processes one by one in sequence. Each of the multiple processes performs, for example, signal processing on an audio signal or generation of parameters used in the signal processing.
[0186] The rendering unit 1300 may perform reverberation processing, early reflection processing, distance attenuation processing, binaural processing, and the like as pipeline processing. However, these processes are merely examples, and the pipeline processing may include other processes or may not include some of the processes. For example, the pipeline processing may include diffraction processing and occlusion processing. Furthermore, for example, reverberation processing may be omitted if it is not necessary. Furthermore, not all sounds may be processed in the binaural processing stage.
[0187] Each process may be expressed as a stage. An audio signal such as a reflected sound generated as a result of each process may be expressed as a rendering item. The multiple stages in the pipeline process and their order are not limited to the example shown in FIG. 16. For example, the processing of the panning unit may be executed in a binaural processing stage, which is one of the multiple stages included in the pipeline process. The binaural processing unit performs a function equivalent to that of the synthesis unit described above.
[0188] In the stereophonic sound reproduction system 600 described above, in order to make the user 99 perceive as if they are moving their head within a three-dimensional sound field by changing the sound presented in accordance with the movement of the user 99's head, as explained in the example of the sound reproduction system 100 above, it is necessary to detect the position and orientation of the user 99's head (orientation relative to the position of the sound source object).
[0189] At this time, it is necessary to obtain the detection results of the position and orientation of the head of the user 99 and perform information processing accordingly, and therefore, ideally, all of the components would be built into a device related to the final output portion, such as a headphone, i.e., the audio presentation device 602 in Fig. 6, as shown in Figures 1 and 2. However, due to constraints such as power supply availability, information processing performance, and housing size and weight, it is necessary to divide the processing so that the information processing portion is performed in the information processing device 601 and the audio output is performed in the audio presentation device 602. Furthermore, in recent years, it has become desirable to connect the audio presentation device 602 to the information processing device 601 via wireless communication, and therefore, when transmitting and receiving information between the audio presentation device 602 and the information processing device 601 via such wireless communication, limitations on the amount of information that can be transmitted and received simultaneously (i.e., limitations on the communication bandwidth) become a bottleneck.
[0190] For example, if the information processing device 601 outputs an output sound signal, the output sound signal must be transmitted to the audio presentation device 602 via wireless communication. This creates delays due to encoding and decoding processes conforming to wireless communication standards, as well as transmission delays due to the large amount of information in the output sound signal itself, which contains a three-dimensional audio signal. This potentially detracts from the user's experience. Furthermore, the information processing device 601 must first acquire the detection results of the user's head position and orientation before generating the output sound signal, making it difficult to instantaneously track the user's head movement. Therefore, the movement of a sound source object specified in the content is processed in the information processing device 601, generating a three-dimensional audio signal, which is then panned to compress the amount of information. The compressed, panned audio signal is then transmitted to the audio presentation device 602, thereby avoiding limitations on the transmission path bandwidth. Because the processing up to this point is a predetermined part of the content, the information can be transmitted to the audio presentation device 602 in advance and buffered.
[0191] Then, the movement of the user's 99's head is detected by the audio presentation device 602, and the detection result is used to convolve the head-related transfer function corresponding to the representative direction (after the head has moved) according to the detection result into the audio signal that has been panned on the audio presentation device 602, thereby making it possible to instantaneously track the direction of arrival of the sound to the movement of the user's head. At this time, by performing the panning process in advance in the information processing device 601, only the process of convolving the head-related transfer functions for a not-so-large number of representative directions is executed in the audio presentation device 602, which makes it difficult to increase the processing resources required on the audio presentation device 602 side, and significantly reduces the delay due to information processing on the audio presentation device 602. Because the representative direction associated with head movement is updated on the audio presentation device 602 side, no communication is required between the information processing device 601 and the audio presentation device 602 after head movement is detected and before the head-related transfer function is updated, making it possible to minimize the time required to update the direction of arrival of the sound.
[0192] In this way, in a stereophonic sound reproduction system 600 divided into an information processing device 601 and an audio presentation device 602, by making the audio signals transmitted and received between the information processing device 601 and the audio presentation device 602 audio signals that have been panned, it is possible to split the information processing between the two devices while still allowing the direction of sound arrival to instantly follow the movement of the user's 99 head.
[0193] When the information processing device 601 is the first terminal and the audio presentation device 602 is the second terminal, the above configuration allows (1) the first terminal, which is not required to be worn by the user 99 and therefore has relatively high processing performance compared to the second terminal, to perform panning processing on the audio signal, including the movement of a sound source object specified in the content, which requires relatively high processing resources. Furthermore, (2) the panning processing compresses the amount of information by consolidating sounds from, for example, 10 or fewer representative directions, reducing the communication bandwidth constraints when transmitting and receiving between the first terminal and the second terminal. Furthermore, (3) the second terminal worn by the user 99 (on the head where the ears are located) can easily detect the position and orientation of the user 99's head and use this information to output an output sound signal from the panned audio signal, thereby instantly updating the sound arrival direction in response to head movement. Furthermore, (4) the output sound signal can be output by simply convolving head-related transfer functions for a small number of representative directions when panning is performed, thereby reducing the processing resources required for the second terminal. The above benefits can be obtained.
[0194] In the configurations shown in FIGS. 2 to 5, the first terminal includes the communication module 102, some functions of the acquisition unit 111, and functions other than the convolution processing of the path calculation unit 121 and the output sound generation unit 131 shown in FIG. 2. The path calculation unit 121 calculates the relative arrival direction of the sound source object from the position of the reference position using a reference position, i.e., the coordinate position or orientation of the user 99 already acquired by the first terminal before transmitting sound information to the second terminal, the coordinate position or orientation of the user 99 at the time of initialization of the system (the first terminal and the second terminal), or a predetermined coordinate position or direction determined in advance by the system. The reference position may be determined in any manner as long as it is the same coordinate position or orientation shared by the first terminal and the second terminal. Therefore, the orientation as the reference position may be determined as a specific direction such as "north" in absolute coordinates, or as an average direction calculated from the direction in which the user 99 faced the user 99 during a predetermined period of time (e.g., one minute) in the past. In the latter case, the orientation as the reference position is updated every predetermined period. In this way, the first terminal and the second terminal share the same coordinate position or orientation using the reference position, and then perform a series of rendering processes, so long as the shared reference position does not change during the series of rendering processes. When updating the reference position, the update may be performed by including it in the information update thread in the spatial information update process mentioned in <Description of Decoder Function>, or by including it in a thread for updating other information, or by using a dedicated thread for updating the reference position.
[0195] 2 , the functions of other parts of the acquisition unit 111, the function of convolution processing of head-related transfer functions in the output sound generation unit 131, the signal output unit 141, the database 105, and the driver 104. The signal output unit 141 converts the direction to set the reference position to the determined coordinates and orientation of the user 99, and outputs an output sound signal according to the detection result by the detector 103.
[0196] In the configurations of Figures 6 to 15, the first terminal is equipped with the analysis unit and panning unit shown in Figure 15. The analysis unit calculates the relative arrival direction from the position of the sound source object to the reference position using a reference position based on first sound information, which is an input signal. The first sound information includes sound source object position information and an audio signal. The analysis unit has functions equivalent to the path calculation unit 121 described above. The panning unit determines a representative direction from a representative point to the reference position based on the arrival direction calculated by the analysis unit, and performs panning processing to distribute the audio signal to the representative point (representative direction). In other words, by performing panning processing, the first sound information is converted into second sound information including position information of the representative point and a panned audio signal. The second terminal is equipped with a synthesis unit. The synthesis unit converts the direction to match the reference position with the coordinates and orientation of the user 99 detected by the second terminal, and performs convolution processing of the head-related transfer function based on the coordinates and orientation of the user 99.
[0197] [Operation] Here, an example of operation of the stereophonic sound reproduction system 600 when processing is divided between the information processing device 601 and the audio presentation device 602 will be described.
[0198] The operation of the information processing device 601 will be described with reference to FIG. 17A. FIG. 17A is a flowchart showing an example of the operation of the information processing device 601 (first terminal) according to an embodiment. In the example of operation shown in the figure, an acquisition unit (not shown in FIG. 15) acquires first sound information including information about the playback sound and information about the position of a sound source object via a communication module (step S101). Based on the first sound information, which is an input signal, the analysis unit uses a reference position to calculate the relative arrival direction from the position of the sound source object to the reference position. The analysis unit may, for example, identify a direct sound and one or more secondary sounds and calculate the propagation path from the sound source object position to the reference position. Based on the arrival direction calculated by the analysis unit, the panning unit determines a representative direction from a representative point to the reference position and performs panning processing to distribute the audio signal to the representative point (representative direction) (step S102). In other words, by performing panning processing, the first sound information is converted into second sound information including position information of the representative point and a panned audio signal. A transmitting unit (not shown in FIG. 15) transmits the second sound information to the sound presentation device 602 (second terminal) (step S103).
[0199] Next, the operation of the audio presentation device 602 (second terminal) will be described with reference to FIG. 17B. A receiving unit (not shown in FIG. 15) receives second sound information from the first terminal (step S104). Next, a sensing information input unit (not shown in FIG. 15) acquires information about the user's position (at least one of position information and facial orientation information) (step S105). Next, a synthesis unit converts the direction of the reference position to the determined user's coordinates and orientation, performs convolution processing of the head-related transfer function based on the user's coordinates and orientation (step S106), and synthesizes and outputs a sound signal (step S107).
[0200] [Specific Example of Panning Processing] To reiterate, in panning processing, reproduced sounds from multiple sound source objects are represented by representative sounds from multiple representative directions. For example, two or three directions can be used as these representative directions. Specifically, in panning processing, the number of representative points is reduced to a number less than the number of sound source objects, and the reproduced sounds can be perceived as sounds coming from the direction of arrival using only the head-related transfer functions of the representative directions for these representative points.
[0201] In this case, the panning process calculates a time shift (delay) that maximizes the cross-correlation between the head-related transfer function of the sound source object in the arrival direction and the head-related transfer function of the representative direction. The time shift obtained here, or a time shift obtained by adding a negative sign to this time shift, is applied to the reproduced sound from the sound source object, and the subsequent processing is performed assuming that the signal after the time shift is in the representative direction.
[0202] This time shift may be a time shift shorter than the sampling frequency (a shift in which the sample position is expressed by a decimal number; hereinafter, referred to as a "decimal shift"). This decimal shift can be performed by oversampling.
[0203] Here, in the panning process, a gain is applied to the signal of a representative direction obtained by time-shifting the reproduced sound of the sound source object, and the sum of these values calculated for each representative point is calculated, and the sum is convolved with the head-related transfer function at each representative point, thereby synthesizing a signal equivalent to the reproduced sound of the sound source object convolved with the head-related transfer function of the arrival direction.
[0204] On the other hand, in the panning process, when synthesizing the head-related transfer function (vector) of the arrival direction by the sum of the head-related transfer functions (vector) of the representative direction, the gain may be calculated so that the error signal vector between the synthesized head-related transfer function (vector) and the head-related transfer function (vector) of the arrival direction is orthogonal to the head-related transfer function (vector) of the representative direction. Note that the head-related transfer function (vector) is a time waveform of a head impulse response, which is an expression of the head-related transfer function in the time domain, regarded as a vector. Hereinafter, this head-related transfer function (vector) will also be simply referred to as a "head-related transfer function vector."
[0205] In the panning process, this gain is corrected so that the energy balance of the head-related transfer functions from the position of the sound source object to the left and right ears of the user 99 is maintained even in the head-related transfer functions substantially synthesized by the panning process using head-related transfer functions from multiple representative points. In other words, in the panning process, the gain may be corrected so that the energy balance of the head-related transfer functions of the left and right ears of the user 99 due to the sound source object is maintained even in the head-related transfer functions substantially synthesized by the panning process.
[0206] In this embodiment, in the panning process, for each direction of arrival of the sound source object, a gain value to be multiplied by the head transfer function of the representative direction and a time shift value to be applied to the head transfer function of the representative direction can be calculated and stored in a head transfer function table described later.
[0207] Then, in the panning process, each sound source object is time-shifted by a time shift value and a gain value corresponding to the direction of arrival of each sound source object, and the time shifts and gains are multiplied and summed to generate a sum signal. In the panning process, this sum signal is treated as being present at the position of the representative point. In the panning process, the head-related transfer function at the position of the representative point is convolved with this sum signal to generate a signal at the ear of the user 99.
[0208] The panning process and the associated audio playback process will be described in detail below, step by step, with reference to the flowchart in Fig. 18. First, a sound source and direction acquisition process is performed (step S201). For example, the path calculation unit 121 acquires the direction of a sound source object as seen by the user 99.
[0209] Specifically, the acquisition unit acquires an audio signal (target signal) of the sound source object. This audio signal may have any sampling frequency and any quantization bit rate. In this embodiment, an example will be described in which an audio signal with a sampling frequency of 48 kHz and a quantization bit rate of 16 bits is used. Furthermore, the path calculation unit 121 acquires directional information of the sound source object added to the audio signal of the content or the audio signals of the participants in the remote call.
[0210] Then, the path calculation unit 121 grasps the spatial arrangement of the sound source object and the user 99. As described above, this arrangement may be an arrangement within a space including a virtual space set in content or the like. Then, the path calculation unit 121 calculates the direction of the sound source object as seen by the user 99, i.e., the arrival direction, according to the grasped arrangement within the space. Similarly, the path calculation unit 121 can also calculate the arrival direction for the audio signal of the content based on the arrangement of the user 99 by referring to directional information of the audio signal of the sound source object.
[0211] The path calculation unit 121 may also calculate the direction of the user 99 from the sound source object.
[0212] Next, the panning unit (the generation unit 134 in the functional block diagram of FIG. 5 ), which is a processing unit that executes the panning process, performs the panning process (step S202). Here, the panning unit performs the panning process on the sound source object using the direction information. In this embodiment, the panning unit performs the panning process from the viewpoint of how closely the sound synthesized at the ear by the panning process can be made to resemble the sound that should be heard at the ear.
[0213] The calculation performed by the panning unit when panning a sound source object (sound source S-1) using representative points R-1 and R-2 will be described with reference to Fig. 19. In Fig. 19, the signal to be panned is sound source S-1, but in order to calculate the optimum shift amount and optimum gain for this, calculations are performed using head-related transfer functions from sound source S-1, representative point R-1, and representative point R-2 to the ears.
[0214] 19, the head-related transfer function with P sampling points (number of taps) from the sound source S-1 to the ear is a P-dimensional vector, which is denoted as v{x} (in the following embodiments, a vector will be represented as "v{}").
[0215] Here, the panning unit calculates the head-related transfer function from the representative point R-1 to the ear of the user 99 as v{x 01}, and the head-related transfer function from the representative point R-2 to the ear is v{x 02}. v{x} and v{x01}, and v{x 01} is time-shifted to v{x1}. Similarly, v{x} and v{x 02}, and v{x 02} is calculated as v{x2} by shifting the time.
[0216] This v{x1} is multiplied by gain A, and v{x2} is multiplied by gain B, and v{x} is approximated by the sum of these. In other words, v{x} is approximated as the approximate value of v{x} = A × v{x1} + B × v{x2}. This makes it possible to achieve panning processing with reduced error.
[0217] The calculation of the gain and the time shift will be described in detail below. First, the calculation of the gain will be described. The error vector obtained by approximating v{x} is expressed by the following equation (1).
[0218]
[0219] In the above formula (1), the arrows on the variables indicate vectors. Here, when A and B are optimally sized, that is, when the magnitude of the error vector is minimized, the error vector v{e} is orthogonal to the plane spanned by the original vectors v{x1} and v{x2}. Therefore, the relationship in the following formula (2) holds.
[0220]
[0221] As a result, the following equation (3) is calculated.
[0222]
[0223] By modifying this equation (3), the following equation (4) is obtained.
[0224]
[0225] For the above equation (4), |v{x2}| 2 When v{x1}·v{x2} is calculated for the equation below, the following equation (5) is obtained.
[0226]
[0227] A can be calculated by subtracting the lower equation from the upper equation of equation (5) and eliminating B. This is shown in equation (6).
[0228]
[0229] Therefore, the gain A is expressed by the following equation (7).
[0230]
[0231] Similarly, by eliminating the gain A, the gain B can be calculated as shown in the following equation (8).
[0232]
[0233] In this way, the gains A and B are determined so that the error vector between the composite signal and the target signal is orthogonal to the representative direction vector used.
[0234] The gains A and B obtained by this calculation are multiplied by the waveform of the head-related transfer function of v{x1} after the time shift due to cross-correlation and the waveform of the head-related transfer function of v{x2}, and it becomes possible to synthesize the head-related transfer function to be output. In other words, these time shift amounts (time shift values) and gains A and B are applied to the sound source S-1 to perform panning processing.
[0235] Next, a specific calculation process for the time shift that maximizes the cross-correlation will be described. 01} handles the head-related transfer function with the number of samples at P points as a vector. Therefore, the subscript of the time (position of the sample point) of the head-related transfer function can be explicitly written as in the following equation (9).
[0236]
[0237] Then, the cross-correlation between the two vectors in equation (9) is defined as a function of "k" as shown in equation (10) below.
[0238]
[0239] where φ xx01 The k that gives the maximum value of (k) is k max01The panning unit, for example, substitutes each value into k, and max01 Similarly, φ xx02 The k that gives the maximum value of (k) is k max02 The panning section writes this k max02 k max01 This k is calculated in the same way. max01 and k max02 Hereinafter, either of the above will be referred to simply as "k max " should be written.
[0240] The panning unit may, for example, calculate gains A, B, and k for the arrival direction of each sound source object that differs every 2 degrees around 360 degrees. max01 , k max02 are stored in the head-related transfer function table as gain values and time shift values, respectively, and are used in the output process described below. max01 , k max02 It is also possible to perform only the following audio output process using a head-related transfer function table in which the values of (a) and (b) have already been calculated and stored.
[0241] Next, the panning unit and the output unit perform audio output processing (step S203). First, the panning unit obtains a gain value and a time shift value corresponding to the obtained arrival direction from the head-related transfer function table for each sound source object. Then, the panning unit multiplies each sampling point (sample) of the waveform of the sound source object by this gain value.
[0242] At this time, the panning unit may correct the gain so that the energy balance of the left and right ear objects due to the sound source object is maintained in the head-related transfer functions synthesized by the panning process. That is, each gain value may be multiplied by an adjustment coefficient that matches the energy balance between the left and right head-related transfer functions with the original head-related transfer functions.
[0243] Next, the panning unit performs a time shift on the signal multiplied by this gain value.
[0244] The details of this time shift are as follows: 01} element k maxA vector v{x1} shifted by a sample is generated by the following procedure.
[0245] First, when the phase is advanced, that is, k max If ≧0, then k is added to the end of the vector. max Only the sample is set to zero, and the length of the vector is maintained. On the other hand, if the phase is delayed, that is, k max If <0, the vector begins with k max Only the samples are set to zero and the length of the vector is maintained, that is, set as in the following equation (11).
[0246]
[0247] In this way, a time-shifted vector v{x1} is generated. The positive or negative polarity of the time shift amount value is reversed depending on which is used as the reference for calculating the cross-correlation. Also, when convolving the head-related transfer function with the sound source signal, attention must be paid to the polarity of the time shift amount.
[0248] The panning unit may perform this time shift by a decimal multiple of the number of taps, rather than by an integer multiple of the number of taps. Alternatively, the time shift may be multiplied by a gain value after the time shift.
[0249] The panning unit treats the signal that has been calculated in this way and that has undergone a gain and time shift as a representative point signal that exists at the position of representative point R. The panning unit then takes the sum of the representative point signals of the sound source objects that are grouped together at representative point R to generate a sum signal. The panning unit then convolves this sum signal with the head-related transfer function at the position of representative point R (head-related transfer function in the representative point direction) to generate a signal at the ear of the user 99.
[0250] The signals generated by the panning unit are reproduced by outputting them to the ears of the user 99. The output may be, for example, a two-channel analog audio signal corresponding to the left and right ears of the user 99.
[0251] This makes it possible to reproduce audio signals corresponding to a virtual sound field as two-channel audio signals through headphones. This completes the audio reproduction process.
[0252] The above configuration can provide the following effects.
[0253] In recent years, when content such as movies, AR, VR, MR, and games is played using VR headphones or HMDs, a rendering technology (binauralization technology) that appropriately describes and plays back the entire 3D sound field has been required. Conventional 3D stereophonic sound (binaural signals) has been generated by individually convolving a head-related transfer function (HRTF) of the corresponding direction of arrival with a plurality of sound source signals. In this way, convolving a HRTF with each individual sound source object has posed a problem in that a huge amount of calculation is required to follow human movement (6 DoF: 6 Degrees of Freedom) with a high sense of realism.
[0254] On the other hand, in the conventional speaker panning process, a sound image is created between speakers by controlling the volume balance of the speakers using the sine law, tangent law, etc. (The sound source object is localized.) However, simply controlling the volume balance is not enough to properly reproduce a stereophonic sound image through headphones.
[0255] In contrast, the above-described audio playback process is characterized by using a path calculation unit 121 that acquires the direction of arrival of a sound source object, and a panning unit that expresses the sound source object by performing panning processing using sound from a specific representative direction based on the direction of arrival acquired by the path calculation unit 121 through time shifting and gain adjustment of the sound source object.
[0256] This configuration enables more efficient and effective rendering by synthesizing sound source objects through panning of representative directions and reducing the number of arrival directions. This reduces the amount of computation compared to conventional methods that individually convolve head-related transfer functions into the signals of each sound source object. That is, the panning unit equivalently synthesizes head-related transfer functions of representative directions that approximate the arrival directions acquired by the path calculation unit 121 through panning processing, thereby generating head-related transfer functions for the arrival directions. By reducing the amount of computation in this way, the system can be applied to VR / AR applications such as games and movies as a 3D sound field playback system. Furthermore, by applying it to smartphones and home appliances, the amount of computation required to generate stereophonic sound can be reduced, thereby reducing costs. Furthermore, as a method with even lower computational complexity, the system can be applied to international standardization, etc.
[0257] In the above-described embodiment, an example has been described in which the panning unit expresses a sound source signal by panning processing using representative points in two directions, left and right, that is, an example in which a vector of a head-related transfer function in the direction of arrival is equivalently synthesized using a vector of a head-related transfer function in the left and right directions. That is, in the above-described embodiment, an example has been described in which the angular directions of the left and right of the user 99 are taken into consideration as direction information.
[0258] However, the vertical direction can also be considered as the direction of arrival. Specifically, it is also possible to equivalently synthesize the vector of the head-related transfer function of the direction of arrival by interpolation using the vectors of the head-related transfer functions in three directions. In other words, the panning unit can also perform panning processing using representative points in three directions including the elevation angle direction (or depression angle direction).
[0259] In this case, similar to the interpolation from two directions, the head related transfer functions in the representative directions are time-shifted so that the cross-correlation with v{x} is maximized, and are expressed as vectors v{x1}, v{x2}, and v{x3}. In this case, the error vector v{e} is expressed by the following equation (12).
[0260]
[0261] This is applied to the following equation (13) and solved.
[0262]
[0263] Specifically, the optimal gains A, B, and C can be calculated by the following equation (14).
[0264]
[0265] In the above equation (14), the "-1" on the right shoulder of the matrix means the inverse matrix. The time shift amount k of the HRIR in the representative direction determined so as to maximize the cross-correlation is max01 , k max02 , k max03 Similarly to the values in the two directions, the values are calculated before the gain values are calculated.
[0266] In the above embodiment, an example in which two to four representative points R are used has been described.
[0267] However, it is of course possible to use two or more representative points R. For example, as shown in the examples described later, it is also possible to use four to six representative points R corresponding to range angles of 90° and 60°, etc. Furthermore, even in the case of four representative points, it is also possible to set the representative points R at different positions, such as diagonally (45°, 135°, 225°, and 315°) or vertically and horizontally (0°, 90°, 180°, and 270°) relative to the user 99. It is also possible to select two or three points from the four to six representative points R that are closest to the direction of arrival and use them as representative points R for synthesizing the sound source.
[0268] That is, in the audio reproduction process, the panning process may use a gain calculated so as to minimize the energy or L2 norm of the error signal vector between the synthesized HRIR vector and the HRIR vector in the sound source direction.
[0269] [Weighting filter when calculating time shift and gain] In the above example, when calculating the time shift and gain that maximize the cross-correlation, the head-related transfer function itself is used. On the other hand, the time shift and / or gain may be subjected to a weighting filter on the frequency axis and then the cross-correlation may be calculated.
[0270] That is, when calculating the time shift and gain that maximize the cross-correlation, it is possible to use a filter that has been subjected to a weighting filter on the frequency axis (hereinafter also referred to as a "frequency weighting filter").
[0271] It is preferable to use a frequency weighting filter that has a cutoff frequency near or slightly higher than the frequency band where human hearing sensitivity is high, and attenuates the higher frequency band, i.e., the frequency band where human hearing sensitivity decreases. For example, it is preferable to use a low-pass filter (LPF) with a cutoff frequency of 3000 Hz to 6000 Hz and approximately 6 dB / oct (octave) to 12 dB / oct.
[0272] Specifically, v{x} and v{x 01} handles the head-related transfer function of point P as a vector, so it is possible to explicitly write the time subscript of the head-related transfer function and write it as in the above equation (9). Here, if the impulse response w of the frequency weighting filter is added to the two vectors in the above equation (9), c (n) is convoluted and truncated to a length P, as shown in the following equation (15).
[0273]
[0274] Here, the operation "*" indicates convolution. Then, the cross-correlation of the two vectors in equation (15) is defined as a function of "k" as shown in equation (16) below.
[0275]
[0276] Here, φ according to equation (16) xx01 The k that gives the maximum value of (k) is k max The panning unit is, for example, a vector v{x 01} element k max A vector v{x1} shifted by one sample is generated in the following procedure, similar to the above equation (11).
[0277] Specifically, when the phase is advanced, that is, k max If ≧0, k maxTo ensure that the vector is a sample, zeros are added to the end of the vector to maintain its length. max If ≧0, the vector v{x1} is v{x1}=(x 01 (0+k max ), x 01 (1+k max ), x 01 (2+k max ), …… x 01 (P-1), ... 0,0,0).
[0278] Also, if the phase is delayed, that is, k max If <0, pad the beginning of the vector with zeros and max Keep the length of the vector to be k samples. max < 0, the vector v{x1} is v{x1} = (0,0,0, ..., x 01 (0), x 01 (1), x 01 (2), ……, x 01 (P-1+k max )) becomes.
[0279] In the above, the vector v{x 01w} is a vector v{x 01}. In this way, it is possible to generate a vector v{x1}, i.e., a cross-correlation can be calculated and used to calculate the time shift, similar to what was described above.
[0280] In the above-described system, when calculating the error (similarity) between the synthesized head-related transfer function and the original head-related transfer function, |v{e}| of the error signal vector (error vector) v{e} is calculated as in the above-described equation (12). 2 A, B, and C that minimize the above were calculated.
[0281] In this case, v{e} may be filtered by a frequency weighting filter. Specifically, when v{e} is waveform data on the time axis, v{e} convolved with the impulse response w(n) of the weighting filter is used as v{e}. w}, then v{e w} is expressed by the following formula (17).
[0282]
[0283] The operator "*" indicates convolution. Here, the operator "*" is used for vectors, but this is a vector representation of the sequence obtained by convolving the sequence representations of the vectors on the left and right of the operator. In other words, v{x} * v{y} is the vector representation of the result of x(n) * y(n). Hereinafter, unless otherwise specified, the operator "*" for vectors will be treated in the same way.
[0284] On this basis, v{e w} into the following equation (18) and solving it, the gains A, B, and C can be calculated.
[0285]
[0286] Or equivalently, v{e} w It is also possible to calculate
[0287]
[0288] Using the time shift and gain thus determined, it becomes possible to distribute (pan) the target signal to a representative direction.
[0289] The target signal to be panned and the head-related transfer function to be convolved may be the same as those described above. That is, the target signal and the head-related transfer function to be convolved do not need to be convolved with a weighting filter.
[0290] By introducing such frequency weighting, it is possible to set the frequency band for approximation with smaller errors (higher accuracy). In particular, since the main energy of music and voice signals is concentrated in the low frequency range, good performance can be obtained by using a weighting filter that weights the low frequency range.
[0291] Furthermore, if the convolution of a weighting filter whose impulse response is w(n) and a vector is expressed as a convolution matrix W in which each row has the impulse response w(n) of the weighting filter time-shifted by one sample, then equation (17) can be transformed into equation (20) below.
[0292]
[0293] Then, in the following formula (21), |v{e}| 2 can be calculated.
[0294]
[0295] Here, W T represents the transpose matrix of W.
[0296] Furthermore, the weighting filter used when calculating the cross-correlation and when calculating the gain may have the same characteristics, or may have different characteristics. If the same weighting filter is used, the weighting filter w may be convoluted with the entire set of original head-related transfer functions, and then the time shift amount and the gain may be calculated by the same process as described above.
[0297] In addition, when the cross-correlation and the optimum gain are calculated by weighting the low frequency band with an LPF as the weighting filter as described above, if the effective band is limited to about 3000 Hz, the decimal shift described above does not need to be performed. In this case, oversampling is also not required.
[0298] In the above-described embodiment, the audio signal is panned and distributed in a plurality of representative directions, and the head-related transfer functions of each representative direction are convolved and expressed. Specifically, the head-related transfer function of the target direction is simulated by the sum of the head-related transfer functions of the representative directions, with the approximate value of v{x} in three directions = A × v{x1} + B × v{x2} + C × v{x3}.
[0299] In such cases, the amplitude characteristics of the high frequencies of the HRTF tend to be lower in level than the original HRTF compared to the low frequencies. This is because even a slight time error caused by a slight shift in the listening point can cause a large phase rotation of the high frequency components of the HRTF, which tends to be canceled out by the addition caused by the panning process.
[0300] In contrast to this, in the audio reproduction process according to this embodiment, the tendency for high frequencies to attenuate may be compensated for by a reproduction high-frequency emphasis filter.
[0301] Specifically, it is possible to compensate for the tendency of high frequencies to attenuate by applying a high-frequency emphasis filter to a signal obtained by performing panning processing and convolving a head-related transfer function in a representative direction. Alternatively, equivalently, the head-related transfer function in the representative direction itself may be subjected to a high-frequency emphasis filter process in advance to emphasize the high frequencies. This high-frequency emphasis filter may be, for example, an impulse response weighting filter that emphasizes the high frequencies by about +1 to +1.5 dB with a turnover frequency of 5000 to 15000 Hz or more.
[0302] In this way, by performing a filter process that emphasizes the high frequencies of the synthesized sound using the panning process, it is possible to further enhance the stereoscopic effect perceived by the listener.
[0303] Even when a decimal shift similar to that described above is performed, mismatches in the high frequency components of head-transmitted signals remain with normal 8 to 16 times oversampling, so a high frequency emphasis filter may be applied.
[0304] In the panning process, the adjustment amounts in the time shift adjustment and the gain adjustment may be determined according to the head-related transfer functions included in the database 105, and the time shift adjustment and the gain adjustment may be applied to the reproduced sound using the determined adjustment amounts to convert it into a representative sound. Since the optimal values of the adjustment amounts in the time shift adjustment and the gain adjustment used in the panning process change according to the head-related transfer functions, first, when the head-related transfer functions included in the database 105 are read out, the adjustment amounts in the time shift adjustment and the gain adjustment corresponding to the read out head-related transfer functions can be determined, and thereafter, the same adjustment amounts can be reused as long as these head-related transfer functions are used, which is advantageous in terms of the amount of processing.
[0305] The head-related transfer function table is an example of table data including head-related transfer functions stored in the database 105. The head-related transfer function table stores the head-related transfer functions together with the adjustment amounts in the time shift adjustment and the gain adjustment determined according to the head-related transfer functions, which are linked to each other. That is, the head-related transfer function table may be constructed by calculating the adjustment amounts in the time shift adjustment and the gain adjustment in advance for each head-related transfer function included in the database 105. In this way, table data of the head-related transfer function table linking each head-related transfer function with the adjustment amount may be stored in the database 105. In this way, the database 105 is an example of a storage unit. The calculation of the adjustment amount for each head-related transfer function may be performed by the generation unit 134 or the decoding processing unit 113. Alternatively, the calculation of the adjustment amount may be performed by an external device and stored in the memory of the external device. In this case, the memory of the external device corresponds to an example of a storage unit.
[0306] Furthermore, adjustment amounts in the time shift adjustment and the gain adjustment may be calculated in advance, and an adjustment amount table linked to each of a plurality of representative directions may be constructed and stored in the database 105. The adjustment amount table may include table data linking the head related transfer functions of each of the plurality of representative directions with the adjustment amounts in the time shift adjustment and the gain adjustment, or the head related transfer functions of each of the plurality of representative directions may be extracted from head related transfer functions of the entire celestial sphere (multiple directions) that are acquired in advance and stored in the database 105 at the time of rendering or at the time of system initialization.
[0307] Furthermore, the adjustment amount table may be a table including, for example, information as to which of a plurality of representative directions the signal is to be distributed to, for a sound signal arriving at the position of the listener from the direction of each head-related transfer function in the spherical head-related transfer function database, and information on a time shift adjustment amount and a gain adjustment amount to be multiplied by the sound signal for each representative direction when distributing.
[0308] When performing the convolution process of the head-related transfer function, the adjustment amount table stored in the database 105 is referenced, and the adjustment amounts for the time shift adjustment and gain adjustment linked to the head-related transfer function of the direction to be applied are used, which eliminates the need to calculate the adjustment amount for each convolution process and contributes to reducing the amount of processing.
[0309] Note that the embodiment of the present invention can also be applied to new head-related transfer functions that are not included in the database 105. When decoding a sound signal, when powering on the sound reproduction system 100, or when initializing the sound reproduction system 100, the head-related transfer functions of the entire three-dimensional sound field may be newly read, and the adjustment amount for each head-related transfer function may be calculated using the method disclosed in this embodiment or another method. In this case, table data linking the head-related transfer functions with the adjustment amounts may be stored in the database 105. Alternatively, the adjustment amounts may be calculated by an external device and stored in the memory of the external device. When performing convolution processing of head-related transfer functions, by referencing the adjustment amounts for the time shift adjustment and gain adjustment linked to the applied head-related transfer function, it is not necessary to calculate the adjustment amount for each convolution processing, which can contribute to reducing the amount of processing.
[0310] In this way, when a new head-related transfer function that is not stored in the database 105 is read, adjustment amounts in the time shift adjustment and gain adjustment used in the panning process may be determined for the new head-related transfer function before storing it in the database 105, and a head-related transfer function table may be constructed by linking the new head-related transfer function with the determined adjustment amounts, and the head-related transfer function table may be stored in the database 105. Then, when performing the panning process, these adjustment amounts are read from the database 105, and the shift adjustment and gain adjustment are applied based on these adjustment amounts. Note that the new head-related transfer function may be one that was previously stored in the database 105, but was temporarily removed from the database 105 when decoding a sound signal, when powering on the sound reproduction system 100, when initializing the sound reproduction system 100, or the like, and then re-stored in the database 105. Instead of the generation unit 134, a second generation unit may be provided which applies time shift adjustment and gain adjustment to the reproduced sound using adjustment amounts linked to a new head-related transfer function stored in the database to convert it into a representative sound, and generates an output sound signal by convolving a head-related transfer function corresponding to a representative direction from each position of the representative point toward the user's position into the representative sound.
[0311] Below, when using a stereophonic reproduction system 600 divided into an information processing device 601 and an audio presentation device 602, the above-described panning process may cause degradation in sound quality. Specifically, the panning process described above involves performing time shift adjustment and gain adjustment. The time shift adjustment and gain adjustment are performed when the user 99 is facing a fixed direction, by convolving a head-related transfer function with sounds from multiple preset representative directions, allowing the user 99 facing the fixed direction to perceive a sound source object located at a position other than the representative direction. More specifically, when the user 99 is facing a 0° direction, a stationary sound source object located at a 60° direction is generated by outputting an output sound signal in which a head-related transfer function is convolved with sounds from representative directions set at 30° and 90°. By multiplying the sound from the 30° representative direction and the sound from the 90° representative direction by the time shift value and gain value calculated for 30° and 90°, an output sound signal in which a sound image of the sound source object is formed at 60° can be generated.
[0312] However, when the user 99 rotates his / her head by -30 degrees, the positional relationship is calculated by convolving a 60° head-related transfer function with the time shift value and gain value calculated for 30 degrees, and convolving a 120° head-related transfer function with the time shift value and gain value calculated for 90 degrees. In this case, the user 99 may not perceive the sound source object as being exactly at a 90° position, with 0° being the front of the user 99 after the head rotation. In other words, the time shift value and gain value calculated for 30 degrees are calculated assuming the convolution of a 30° head-related transfer function, and therefore are not appropriate for the convolution of a 60° head-related transfer function, and the time shift value and gain value calculated for 90 degrees are calculated assuming the convolution of a 90° head-related transfer function, and therefore are not appropriate for the convolution of a 120° head-related transfer function. As a result, the user 99 may misperceive the localized position of the sound source object.
[0313] This is particularly noticeable when the position of a sound source object is perceived when an elevation component is included.
[0314] For example, a case will be described in which six representative points (representative directions) arranged horizontally as shown in FIG. 20 are used. FIG. 20 is a diagram for explaining the arrangement of representative directions in this embodiment. FIG. 20 shows representative directions R1, R2, R3, R4, R5, and R6 arranged at positions of 30°, 90°, 150°, 210°, 270°, and 330°, respectively, horizontally surrounding a user 99. When a sound source object is located in area A1, an output sound signal is output using representative directions R1 and R2. That is, the reproduced sound emitted by a sound source object located in area A1 is distributed to representative directions R1 and R2, and a 30° head-related transfer function is convolved with the sound distributed to representative direction R1, and a 90° head-related transfer function is convolved with the sound distributed to representative direction R2. Conversely, the reproduced sound of a sound source object within a 120° range between 330° and 90° is distributed to representative direction R1.
[0315] Here, (a) of each of Figures 21 and 22 shows a case where, as in the above example, a time shift value is calculated such that the cross-correlation between head-related transfer functions is maximized when calculating the time shift value. Figures 21 and 22 are diagrams for explaining an example of time shift value calculation according to an embodiment. Here, the time shift value (vertical axis) is shown for each direction (horizontal axis) of the sound source object after head rotation when the user 99 rotates their head and the angle shown at the right shoulder in each figure is used as the representative direction. Note that Figure 21 shows an example of the left ear, and Figure 22 shows an example of the right ear. As shown in (a) of each figure, the time shift value has a discontinuity at some point within 360°, which is presumed to be the cause of the loss of the sense of localization. Therefore, (b) of each of Figures 21 and 22 shows the results of examining the calculation direction of the time shift value to reduce the occurrence of such discontinuities.
[0316] Here, since the head-related transfer function of the representative direction itself is convoluted in the representative direction, based on the premise that the time shift value can be set to 0, when calculating the time shift value of a predetermined position so that outliers are less likely to occur, the time shift value of an adjacent position closer to the representative direction than the predetermined position is used. More specifically, the time shift in the representative direction is set to 0, and the time shift value of the next position closest (adjacent) to the representative direction is calculated. The calculation is performed so that the calculated value is closer to 0, which is the time shift value in the representative direction, than the time shift value when the above-mentioned cross-correlation is maximized.
[0317] More specifically, when calculating the time shift value at a predetermined position, the time shift value of the peak closest to the time shift value at a position adjacent to the predetermined position among several peaks indicated by the cross-correlation is calculated as the time shift value at the predetermined position. By doing so, as shown in FIGS. 21 and 22 , the time shift value is less likely to form a discontinuous value. When calculating the time shift value at a predetermined position, the time shift value of the peak closest to the time shift value at a position adjacent to the predetermined position among several peaks indicated by the cross-correlation is calculated as the time shift value at the predetermined position within a threshold range based on the maximum peak. The threshold range may be, for example, within a coefficient multiple of the correlation value at the maximum peak (the coefficient is less than 1). If the coefficient is set to, for example, 0.8, a time shift value that is relatively close to the correlation value at the maximum cross-correlation, that is, within 0.8 times the correlation value at the maximum cross-correlation, and close to the time shift value at the adjacent position, is obtained.
[0318] Next, the calculation of gain values when the time shift value is determined as described above will be described. FIG. 23 is a diagram illustrating an example of gain value calculation according to an embodiment. As shown in FIG. 23 , in reality, the arrival direction of a sound source object is surrounded by three representative directions in three-dimensional space. In other words, by multiplying each of the three representative directions by a certain gain and distributing the sound, the sound source object can be localized in the arrival direction using sounds from the three representative directions. Here, the three gains may be calculated as shown in equations (12) to (15). However, for example, gain values for localizing the sound source object in the arrival direction in two of the three representative directions may be calculated as follows, and then the gain value for the remaining representative direction may be calculated.
[0319] Here, the representative direction closest to the zenith is excluded from the two representative directions selected first. In this way, the gain values of the two representative directions close to the horizontal direction are determined first, and the gain values of the two representative directions close to the horizontal direction are emphasized, and the gain value of the representative direction closest to the zenith can be determined accordingly. As already explained, the sense of localization of the user 99 in the horizontal direction is relatively more stable than in the vertical direction. Therefore, if the gain values of the two representative directions close to the horizontal direction are emphasized and the gain value of the representative direction closest to the zenith is determined accordingly, it is easy to suppress the influence of the gain value of the representative direction closest to the zenith on the deterioration of the sense of localization.
[0320] Specifically, the calculation is performed as follows: First, to determine the gain values of the two representative directions, the following equation (22) is solved.
[0321]
[0322] G 1 is the gain value for one of the two representative directions, and G 2 is the gain value for the other of the two representative directions.
[0323] And G 1 and G 2 While maintaining the relationship (i.e., ratio), the gain value of the representative direction closest to the zenith is determined by solving the following equation (23).
[0324]
[0325] α is the gain value for the resultant vector obtained by multiplying the two representative directions by their respective gain values and taking the sum, and G t is the gain value for the representative direction closest to the zenith.
[0326] By doing so, for example, effects such as those shown in FIGS. 24 and 25 were confirmed. FIGS. 24 and 25 are diagrams showing the results of gain value calculation according to the embodiment. FIG. 24 shows the results of plotting the SNR at an elevation angle of 0°, and FIG. 25 shows the results of plotting the SNR at an elevation angle of 46°. In FIGS. 24 and 25, Comparative Example 1 shows the SNR calculated using the calculation method already described, Comparative Example 2 shows the SNR using a different calculation method, and Example 1 shows the SNR calculated using a method in which gain values for localizing a sound source object in the arrival direction in the above two representative directions are calculated, and then a gain value for the remaining one representative direction is calculated. In each case, the larger the value, the better the result.
[0327] As shown in the figure, at an elevation angle of 0°, there was no significant difference between Comparative Example 1 and Example 1, but both Comparative Example 1 and Example 1 showed better results than Comparative Example 2. At an elevation angle of 46°, Example 1 showed better results than Comparative Examples 1 and 2, especially as the rotation angle approached 90°.
[0328] Next, another example of a method for calculating the time shift value and the gain value will be described. In the following, a head-related transfer function vector at a horizontal position, which is a virtual position in the horizontal direction, is set, and similarly to the above, the cross-correlation with the head-related transfer function vectors of two representative directions close to the horizontal direction, excluding the representative direction closest to the zenith, is calculated, and the time shift value S 1 and S 2 Then, the gain values G in the two representative directions are calculated by solving the following equation (24) using the head-related transfer function vector to which this time shift has been applied. 1 and G 2 Calculate.
[0329]
[0330] Furthermore, a new vector is defined as shown in the following equation (25), and the cross-correlation between the arrival direction, the representative direction closest to the zenith, and the defined new vector is calculated, and the time shift value S that maximizes the correlation value is obtained. t and S X Ask for.
[0331]
[0332] Then, using the head-related transfer function vector to which this time shift has been applied, the representative direction closest to the zenith and the gain value of the newly defined vector are calculated by solving the following equation (26).
[0333]
[0334] As a result, the time shift values and gain values in the three representative directions are obtained by the following equation (27).
[0335]
[0336] By doing this, it is possible to obtain time shift values and gain values for the horizontal representative direction in which the user 99's sense of positioning is relatively stable, so that the sense of positioning is less likely to be impaired even if a head transfer function different from the head transfer function assumed by the information processing device 601 is convolved into the representative direction signal by the audio presentation device 602 as the head rotates.
[0337] It can be seen that the time shift value and gain value calculated as described above have a slight numerical difference when compared to when the angle of the user's 99 head after rotation is replaced with the angle before rotation.
[0338] The case where the arrival direction is 15°, the representative directions are 30° and 330°, and the time shift values and gain values calculated using equations (24) to (27) are used to generate sounds distributed in these representative directions, assuming that the head has rotated 60°, by convolving the head-related transfer functions of the representative directions of 30° and 90° (Example 2) is compared with the case where the arrival direction is 75°, the representative directions are 30° and 90°, and the head-related transfer functions of the representative directions of 30° and 90° are simply convolved to generate sounds distributed in these representative directions (calculated correct value: Comparative Example 3). The results are shown in FIGS. 26 and 27. FIG. 26 is a diagram for verifying the results of the time shift value calculation according to the embodiment. FIG. 27 is a diagram for verifying the results of the gain value calculation according to the embodiment.
[0339] In FIGS. 26 and 27 , the dashed lines indicate the calculation results for Comparative Example 3, and the solid lines indicate the calculation results for Example 2. As shown in FIGS. 26 and 27 , slight errors occur in both the time shift values and gain values between Example 2 and Comparative Example 3. In other words, it can be seen that errors occur between the actual correct values. Therefore, for example, in order to reduce such errors, the time shift values and gain values calculated in Example 2 may be corrected so as to reduce the expected error values. Specifically, for example, the time shift values and gain values may be corrected by subtracting the average value of the errors at each representative point. Alternatively, the expected values may be calculated so as to weight the direction with a larger gain value, and the above correction may be performed. In other words, the direction in which signals are concentrated within a 120° range distributed in a certain representative direction may be estimated from the distribution of gain values, and the correction amounts for the time shift values and gain values may be determined assuming that the signals are in that direction.
[0340] Incidentally, the method described below can also be considered as a method for suppressing the deterioration in sound quality caused by the above-mentioned panning process when using a stereophonic reproduction system 600 divided into an information processing device 601 and an audio presentation device 602.
[0341] As described above, the deterioration of sound quality occurs due to the use of gain values and shift values that are different from the gain values and shift values that should be used. Therefore, if the gain values and shift values that are used are corrected so that they are closer to the gain values and shift values that should be used, it can be said that such deterioration of sound quality can be suppressed. The following description will focus specifically on the gain values, but the same can be applied to the shift values.
[0342] For example, consider a sound source object whose direction in the horizontal plane (i.e., azimuth direction or bearing) is 15°. In the horizontal plane, the representative directions are set to −150°, −90°, −30°, 30°, 90°, and 150°.
[0343] When this sound source object is expressed as a representative sound from a representative direction, for example, a representative sound panned in representative directions of -30° and 30° is generated. The gain values at this time are set to a gain value that takes into account the representative direction of -30° and a gain value that takes into account the representative direction of 30°. However, if the user 99 rotates their head by 60°, the sound source object changes to a position of -45°, and the representative directions used are -90° and -30°. However, at this time, the gain values that take into account the representative direction of -30° and the gain value that takes into account the representative direction of 30° are set, resulting in a deterioration in sound quality.
[0344] Therefore, as the gain value to be set for the sound source object, it is advisable to select multiple directions (virtual representative points) within a range of, for example, ±60°, starting from −30° and 30°, i.e., −90° to 30° and −30° to 90°, respectively, and calculate gain values for each of these directions. A single gain value (hereinafter referred to as a global gain value) can be set using all or a plurality of these gain values. For example, if multiple directions are selected in 5° increments within the range of −90° to 30°, gain values for 25 directions, i.e., −90°, −85°, −80°, ..., 20°, 25°, and 30°, can be calculated. If a single gain value using all 25 gain values is set as the gain value for the sound source object, even if the head is rotated 60° as described above, the global gain value includes the gain value for −90° in the calculation, thereby suppressing deterioration in sound quality compared to simply using a gain value for −30°. Even in the range of -30° to 90° starting from 30°, the gain value for -30° is used in the calculation of the comprehensive gain value, so deterioration in sound quality can be suppressed compared to simply using the gain value for 30°.
[0345] Also, for example, consider a case where the head rotation amount is 120°. The sound source object changes to a position of -105°, and -150° and -90° are used as the representative directions. The global gain value takes into account gain values outside the ranges of -150° and -90°, from -90° to 30° and -30° to 90°. Even in this case, since gain values up to -90° and -30° are taken into account, deterioration in sound quality can be suppressed compared to when gain values for -30° and 30° are simply used. Alternatively, the range to be considered from the direction of the starting point may be set to ±120°, so that even when the head rotation amount is 120°, the gain value for the representative direction used at that time is included in the global gain value. In other words, the range to be considered from the direction of the starting point needs to be less than ±180°, and can be set arbitrarily, such as ±170°, ±160°, ±150°, ±140°, ±130°, ±120°, ±110°, ±100°, ±90°, ±80°, ±70°, ±60°, ±50°, ±40°, ±30°, ±20°, ±10°, and ±5°.
[0346] Furthermore, within these ranges, the density of the selected direction is not limited to 5° increments. For example, the density of the selected direction within the range may be arbitrarily set at 10°, 9°, 8°, 7°, 6°, 5°, 4°, 3°, 2°, 1°, or increments of less than 1°. As an example, the density of the selected direction within the range may be consistent with the density of the selectable directions in the set of head-related transfer functions that has been loaded. However, since the denser the density of the selected direction within the range, the greater the amount of calculation processing. Therefore, the density of the selected direction within the range may be set to a density of directions that has been thinned out at a rate according to processing capacity from the density of the selectable directions in the set of head-related transfer functions that has been loaded.
[0347] The following describes a method for calculating a global gain value and its effects based on examples. In the following examples, an example will be described in which the density of directions selected within a range of ±60° from the starting point is in 5° increments. The global gain value is calculated, for example, as an average value obtained by averaging the gain values for each of the multiple directions selected within the range. This average may be a simple average, or a weighted average in which a smaller rotation angle, which is considered to occur more frequently, approaches 1 and a larger rotation angle, which is considered to occur less frequently, approaches 0. Therefore, this weighted average can be considered to be a weighted average based on the angular difference between each direction of the virtual representative points and their corresponding original directions. More specifically, the global gain value may be a weighted average calculated so that the weight assigned to the adjustment amount at a virtual representative point with a small angular difference is greater than the weight assigned to the adjustment amount at a virtual representative point with a large angular difference. Alternatively, instead of averaging, a single gain value obtained by performing a calculation to minimize the error vector as described above for all selected virtual representative points may be used as the global gain value.
[0348] In the following example, the global gain value is calculated by averaging the two representative points that horizontally sandwich the position of the sound source object. However, the global gain value may be calculated not only for the two points that sandwich the sound source object but also for three or more representative points that surround the sound source object.
[0349] Two global gain values corresponding to two representative directions that sandwich a certain sound source object in the horizontal direction can be calculated by the following equation (28).
[0350]
[0351] Note that φ indicates one of the two representative directions, and α is the difference between the representative direction and the direction of the sound source object, where the direction of the sound source object (target direction) is φ-α. Note that the other of the two representative directions is expressed as φ-60°. Equation (28) can be converted to the average value of gain values in 5° increments by replacing M with M=5k (-12≦k≦12). The following equation (29) is obtained by replacing equation (28) with M=5k (-12≦k≦12).
[0352]
[0353] The results of a localization experiment using the global gain values calculated based on equation (29) are shown in Figures 28 to 31. Figures 28 to 31 are diagrams showing the results of a localization experiment for verifying the results of the gain value calculation according to the embodiment. Note that, below, an example using the global gain values is shown as an example, and a comparative example simply using the gain values before head rotation is shown.
[0354] 28 shows the relationship between the front-to-back error rate between the direction (azimuth) perceived by the subject and the actual direction (azimuth) for a sound source object at an elevation angle of 0°, i.e., in a horizontal plane, and the rotation angle of the subject's head. As shown in FIG. 28, the example showed a smaller error rate than the comparative example, especially when the head was rotated.
[0355] 29 shows the relationship between the average localization error between the direction (azimuth) perceived by the subject and the actual direction (azimuth) for a sound source object at an elevation angle of 0°, and the rotation angle of the subject's head. As shown in Fig. 29, in the example, the error was smaller than in the comparative example, especially when the head was rotated.
[0356] 30 shows the relationship between the front-to-back error rate between the direction (azimuth) perceived by the subject and the actual direction (azimuth) for a sound source object at an elevation angle of 40°, and the rotation angle of the subject's head. As shown in FIG. 30, the example showed a smaller improvement in the error rate than the comparative example.
[0357] 31 shows the relationship between the average localization error between the direction (azimuth) perceived by the subject and the actual direction (azimuth) for a sound source object at an elevation angle of 40°, and the rotation angle of the subject's head. As shown in Fig. 31, the example showed a smaller effect of reducing the error than the comparative example.
[0358] In particular, based on the results of the localization experiment on the sound source object at an elevation angle of 40° shown in Figures 30 and 31, it was predicted that the gain value of the representative direction in the elevation angle direction (for example, the representative direction of the zenith), in other words, the representative point regarding the elevation angle (elevation angle representative point), reduces the effect of using the global gain value, and the following improvement was devised.
[0359] In other words, if the gain value of the elevation representative point (a representative point that includes a vertical component among three or more representative points surrounding the sound source object) is set to a constant value (fixed gain value) regardless of the direction of the sound source object, it is thought that the influence of the elevation representative point on the sound localization in the horizontal plane relative to the sound source object can be suppressed. The same is thought to be true for the depression representative point on the nadir side opposite the zenith, so here we will focus on the elevation representative point only.
[0360] The gain values and shift values of the elevation representative points include components that affect the horizontal sound localization during the calculation process. Therefore, even if the gain values to be used in the horizontal direction are adjusted using the global gain values as described above, the gain values and shift values for the elevation representative points maintain the gain values and shift values corresponding to the direction before head rotation, which is thought to reduce the effectiveness of using the global gain values. Therefore, it is thought to be effective to divide the calculation into stages: in a first stage, the gain values for the elevation representative points are set to maintain a predetermined value regardless of the direction (target direction) of the sound source object, and then in a second stage, the global gain values are set corresponding to that predetermined gain value. In other words, by calculating the gain values for the elevation representative points separately from the gain values for the representative points in the horizontal plane, the influence of the gain values for the elevation representative points on the horizontal sound localization is reduced. This is effective not only when global gain values are used, but also in applications involving head rotation, particularly when the elevation component is included. In other words, it is also effective to apply the use of a fixed gain value as the gain value for the elevation angle representative point to the above-mentioned separate gain value adjustment.
[0361] A specific method for setting the fixed gain value and the results of a localization experiment using the set fixed gain value will be described below.
[0362] If the fixed gain value is set to be independent of the sound source object's direction in the horizontal plane but also independent of the sound source object's direction in the vertical plane, it would be difficult to localize the sound source object in the elevation direction. In other words, the fixed gain value needs to be set to be dependent on the sound source object's direction in the vertical plane. Therefore, appropriate fixed gain values are set for each of the sound source objects in several directions in the vertical plane. If the sound source object is in the horizontal plane (elevation angle 0°), the fixed gain value needs to be 0, and if the sound source object is in the zenith direction (elevation angle 90°), the fix gain value needs to be 1. In other words, the fix gain value needs to be a value that gradually increases between 0 and 1 as the sound source object's direction in the vertical plane changes from 0° to 90°.
[0363] Hereinafter, a range in which the elevation angle direction of the sound source object is greater than 0° and less than 90° will be considered. When the sound source object has an elevation angle of 20°, 0.05, 0.10, 0.15, and 0.30 are set as the fixed gain values. For each head rotation angle, the phase shift is calculated by taking the azimuth shift value that maximizes the cross-correlation between the ILD (Interaural Level Difference) in the head related transfer function at the original (i.e., appropriate gain value according to the angle after rotation) and the ILD in the head related transfer function with the fixed gain value. As a result, the gain value that minimizes the phase shift is adopted as the fixed gain value when the sound source object has an elevation angle of 20°. As a result, the fixed gain value when the sound source object has an elevation angle of 20° was set to 0.10.
[0364] Experiments were also conducted by setting several fixed gain values for the sound source object at an elevation angle of 40°, 60°, and 80°, respectively, and the fixed gain values were set to 0.15, 0.30, and 0.60. FIG. 32 is a diagram showing the set fixed gain values according to the embodiment. As shown in the figure, it is appropriate for the fixed gain value to be set to a relatively small value when the elevation angle of the sound source object is small (e.g., an elevation angle of 40° or less), while it is appropriate for the value to increase rapidly when the elevation angle is large (e.g., an elevation angle of 60° or more). In other words, it can be said that it is more appropriate for the fixed gain value to curve exponentially rather than simply being proportional to the elevation angle of the sound source object.
[0365] 33 and 34 are diagrams showing the results of localization experiments for verifying the results of gain value calculations according to the embodiment. Note that, in the following, an example in which an appropriate fixed gain value is used in combination with a global gain value is shown as an example, and a comparative example in which only the global gain value is used is shown.
[0366] 33 shows the relationship between the front-to-back error rate between the direction (azimuth) perceived by the subject and the actual direction (azimuth) for a sound source object at an elevation angle of 40°, and the rotation angle of the subject's head. As shown in FIG. 33, the example showed a smaller error rate than the comparative example, especially when the head was rotated.
[0367] 34 shows the relationship between the average localization error between the direction (azimuth) perceived by the subject and the actual direction (azimuth) of a sound source object at an elevation angle of 40° and the rotation angle of the subject's head. As shown in FIG. 34, the example shows that the error is smaller than that of the comparative example, especially when the head is rotated. In this way, by using a combination of fixed gain values, it is possible to localize the sound source object more appropriately than when simply using a global gain value.
[0368] When using the head-related transfer functions of the elevation angle representative point and the depression angle representative point, instead of using the fixed gain value as described above, an initial delay may be corrected according to the head rotation angle. The initial delay may be corrected by adjusting the initial delay of the head-related transfer functions of the elevation angle representative point and the depression angle representative point between the ears so that a change in ITD (Interaural Time Difference) equal to the change in ITD when a sound directly in front of the listener (azimuth 0°) moves due to head rotation is applied to the ITD of the head-related transfer functions of the elevation angle representative point and the depression angle representative point. The amount of adjustment here varies depending on the direction of the sound source object. However, the directions of representative sound source objects may be calculated in advance and stored as a table on the receiving side (i.e., the audio presentation device 602 side), so that the direction may be referenced, selected, and used according to the estimated sound source direction.
[0369] Other Embodiments Although the embodiments have been described above, the present disclosure is not limited to the above-described embodiments.
[0370] For example, the information processing system or sound reproduction system described in the above embodiments may be realized as a single device including all of the components, or may be realized by allocating each function to multiple devices and coordinating these multiple devices. In the latter case, the information processing device may be an information processing device such as a smartphone, tablet terminal, PC, or base station. For example, in the sound reproduction system 100 having a function as a renderer that generates a sound signal with added sound effects, all or part of the renderer's functions may be performed by a server. That is, all or part of the acquisition unit 111, path calculation unit 121, output sound generation unit 131, and signal output unit 141 may reside in a server (not shown). In this case, the sound reproduction system 100 is realized by combining, for example, an information processing device such as a computer or smartphone, a sound presentation device such as a head-mounted display (HMD) or earphones worn by the user 99, and a server (not shown). Note that the computer, sound presentation device, and server may be connected to each other so as to be able to communicate with each other via the same network, or may be connected via different networks. When the sound reproduction system 100 is connected via different networks, the possibility of communication delays increases, so processing on the server may be permitted only when the computer, sound presentation device, and server are connected so as to be able to communicate with each other via the same network. Also, depending on the amount of bitstream data received by the sound reproduction system 100, it may be determined whether the server will take on all or part of the functions of the renderer.
[0371] The information processing system or sound reproduction system of the present disclosure can also be realized as an information processing device that is connected to a reproduction device having only a driver and that simply reproduces an output sound signal generated based on acquired sound information to the reproduction device. In this case, the information processing device may be realized as hardware having a dedicated circuit, or as software that causes a general-purpose processor to execute specific processing.
[0372] In the above-described embodiment, the processing performed by a specific processing unit may be performed by another processing unit. The order of multiple processing operations may be changed, or multiple processing operations may be performed in parallel.
[0373] In the above-described embodiments, each component may be realized by executing a software program suitable for that component, or by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0374] Furthermore, each component may be realized by hardware. For example, each component may be a circuit (or integrated circuit). These circuits may form a single circuit as a whole, or each may be a separate circuit. Furthermore, each of these circuits may be a general-purpose circuit or a dedicated circuit.
[0375] Furthermore, the general or specific aspects of the present disclosure may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, etc. Furthermore, the general or specific aspects of the present disclosure may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.
[0376] For example, the present disclosure may be realized as an information processing method or an audio signal reproducing method executed by a computer, or as a program for causing a computer to execute the information processing method or the audio signal reproducing method. The present disclosure may also be realized as a computer-readable non-transitory recording medium on which such a program is recorded.
[0377] In addition, this disclosure also includes forms obtained by applying various modifications to each embodiment that a person skilled in the art would think of, or forms realized by arbitrarily combining the components and functions of each embodiment within the scope that does not deviate from the intent of this disclosure.
[0378] In the present disclosure, the encoded sound information can be rephrased as a bitstream containing a sound signal, which is information about a predetermined sound to be reproduced by the sound reproduction system 100, and metadata, which is information about the localization position when the sound image of the predetermined sound is localized at a predetermined position within a three-dimensional sound field. For example, the sound information may be acquired by the sound reproduction system 100 as a bitstream encoded in a predetermined format such as MPEG-H 3D Audio (ISO / IEC 23008-3). As an example, the encoded sound signal includes information about the predetermined sound to be reproduced by the sound reproduction system 100. The predetermined sound here refers to a sound emitted by a sound source object present in the three-dimensional sound field or a natural environmental sound, and may include, for example, a mechanical sound or the sounds of animals, including humans. When multiple sound source objects are present in the three-dimensional sound field, the sound reproduction system 100 acquires multiple sound signals corresponding to the multiple sound source objects.
[0379] On the other hand, metadata is, for example, information used to control acoustic processing of a sound signal in the sound reproduction system 100. Metadata may be information used to describe a scene expressed in a virtual space (three-dimensional sound field). Here, a scene is a term that refers to a collection of all elements representing three-dimensional images and acoustic events in a virtual space, modeled by the sound reproduction system 100 using metadata. In other words, the metadata here may include not only information for controlling acoustic processing, but also information for controlling video processing. Of course, the metadata may include information for controlling only either audio processing or video processing, or may include information used to control both. In the present disclosure, the bitstream acquired by the sound reproduction system 100 may include such metadata. Alternatively, the sound reproduction system 100 may acquire the metadata separately from the bitstream, as described below.
[0380] The sound reproduction system 100 generates virtual sound effects by performing sound processing on the sound signal using metadata included in the bitstream and additionally acquired position information of the interactive user 99. For example, sound effects such as early reflection sound generation, late reverberation sound generation, diffraction sound generation, distance attenuation effect, localization, sound image localization processing, and Doppler effect may be added. Information for switching on and off all or part of the sound effects may also be added as metadata.
[0381] All or part of the metadata may be obtained from sources other than the bitstream of audio information. For example, either the metadata controlling audio or the metadata controlling video may be obtained from sources other than the bitstream, or both may be obtained from sources other than the bitstream.
[0382] Furthermore, if metadata for controlling the video is included in the bitstream acquired by the audio reproduction system 100, the audio reproduction system 100 may have a function for outputting the metadata that can be used for controlling the video to a display device that displays images or a 3D video reproduction device that reproduces 3D video.
[0383] As an example, the encoded metadata includes information about a three-dimensional sound field including a sound source object emitting a sound and an obstacle object, and information about a localization position when a sound image of the sound is localized at a predetermined position in the three-dimensional sound field (i.e., the sound is perceived as arriving from a predetermined direction), i.e., information about the predetermined direction. Here, an obstacle object is an object that can affect the sound perceived by the user 99 by, for example, blocking or reflecting the sound emitted by the sound source object before it reaches the user 99. Obstacle objects may include not only stationary objects but also animals such as people or moving objects such as machines. Furthermore, when multiple sound source objects exist in a three-dimensional sound field, the other sound source objects may be obstacle objects for any one sound source object. Furthermore, both non-sound-source objects such as building materials or inanimate objects and sound-emitting sound objects may be obstacle objects.
[0384] The spatial information constituting the metadata may include not only the shape of the three-dimensional sound field, but also information representing the shape and position of obstacle objects present in the three-dimensional sound field and the shape and position of sound source objects present in the three-dimensional sound field. The three-dimensional sound field may be either a closed space or an open space, and the metadata may include information representing the reflectance of structures that can reflect sound in the three-dimensional sound field, such as floors, walls, or ceilings, and the reflectance of obstacle objects present in the three-dimensional sound field. Here, the reflectance is the ratio of the energy of reflected sound to incident sound, and is set for each frequency band of sound. Of course, the reflectance may be set uniformly regardless of the frequency band of the sound. Furthermore, when the three-dimensional sound field is an open space, parameters such as a uniform attenuation rate, diffracted sound, or early reflection sound may be used.
[0385] In the above description, reflectance is cited as a parameter related to an obstacle object or a sound source object included in the metadata, but the metadata may include information other than reflectance. For example, information about the material of the object may be included as metadata related to both the sound source object and the non-sound source object. Specifically, the metadata may include parameters such as diffusion rate, transmittance, or sound absorption rate.
[0386] Information about the sound source object may include information such as volume, radiation characteristics (directivity), playback conditions, the number and type of sound sources emitted from a single object, or information specifying a sound source area within the object. The playback conditions may, for example, determine whether the sound is a continuous sound or an event-triggering sound. The sound source area within the object may be determined based on the relative relationship between the position of the user 99 and the position of the object, or may be determined based on the object. When the sound source area is determined based on the relative relationship between the position of the user 99 and the position of the object, the surface from which the user 99 is viewing the object is used as the reference, and the user 99 can be made to perceive sound X as emanating from the right side of the object and sound Y as emanating from the left side of the object as viewed from the user 99. When the sound source area is determined based on the object, the surface from which the user 99 is viewing the object is used as the reference, and the sound emitted from which area of the object can be fixed regardless of the direction the user 99 is viewing. For example, the user 99 can be made to perceive a high-pitched sound coming from the right side and a low-pitched sound coming from the left side when viewing the object from the front. In this case, when the user 99 goes around to the back of the object, the user 99 can be made to perceive a low-pitched sound coming from the right side and a high-pitched sound coming from the left side as viewed from the back.
[0387] The spatial metadata may include the time to early reflections, the reverberation time, or the ratio of direct sound to diffuse sound. If the ratio of direct sound to diffuse sound is zero, the user 99 will perceive only direct sound.
[0388] Furthermore, information indicating the position and orientation of the user 99 in the three-dimensional sound field may be included in the bitstream as metadata as an initial setting, or may not be included in the bitstream. If the information indicating the position and orientation of the user 99 is not included in the bitstream, the information indicating the position and orientation of the user 99 is acquired from information other than the bitstream. For example, the position information of the user 99 in the VR space may be acquired from an app that provides VR content. The position information of the user 99 for presenting sound as AR may be position information obtained by a mobile terminal performing self-position estimation using GPS, a camera, LiDAR (Laser Imaging Detection and Ranging), or the like. Note that the sound signal and metadata may be stored in a single bitstream or may be stored separately in multiple bitstreams. Similarly, the sound signal and metadata may be stored in a single file or may be stored separately in multiple files.
[0389] When an audio signal and metadata are stored separately in multiple bitstreams, information indicating other related bitstreams may be included in one or some of the multiple bitstreams in which the audio signal and metadata are stored. Also, information indicating other related bitstreams may be included in metadata or control information of each of the multiple bitstreams in which the audio signal and metadata are stored. When an audio signal and metadata are stored separately in multiple files, information indicating other related bitstreams or files may be included in one or some of the multiple files in which the audio signal and metadata are stored. Also, information indicating other related bitstreams or files may be included in metadata or control information of each of the multiple bitstreams in which the audio signal and metadata are stored.
[0390] Here, the related bitstreams or files are, for example, bitstreams or files that may be used simultaneously during audio processing. Furthermore, information indicating other related bitstreams may be collectively described in the metadata or control information of one bitstream among multiple bitstreams storing audio signals and metadata, or may be separately described in the metadata or control information of two or more bitstreams among the multiple bitstreams storing audio signals and metadata. Similarly, information indicating other related bitstreams or files may be collectively described in the metadata or control information of one file among multiple files storing audio signals and metadata, or may be separately described in the metadata or control information of two or more files among the multiple files storing audio signals and metadata. Furthermore, a control file collectively describing information indicating other related bitstreams or files may be generated separately from the multiple files storing audio signals and metadata. In this case, the control file does not need to store the audio signal and metadata.
[0391] Here, the information indicating the other related bitstream or file may be, for example, an identifier indicating the other bitstream, a file name indicating the other file, a URL (Uniform Resource Locator), or a URI (Uniform Resource Identifier). In this case, the acquisition unit identifies or acquires the bitstream or file based on the information indicating the other related bitstream or file. Furthermore, the information indicating the other related bitstream may be included in metadata or control information of at least some of the bitstreams among a plurality of bitstreams storing audio signals and metadata, and the information indicating the other related file may be included in metadata or control information of at least some of the files among a plurality of files storing audio signals and metadata. Here, the file including information indicating the related bitstream or file may be, for example, a control file such as a manifest file used for content distribution.
[0392] The present disclosure is useful in reproducing sound, for example, by allowing a user to perceive stereoscopic sound.
[0393] 99 User 100 Sound reproduction system 101 Information processing device 102 Communication module 103 Detector 104 Driver 105 Database 111 Acquisition unit 112 Encoded sound information input unit 113 Decode processing unit 114 Sensing information input unit 121 Path calculation unit 131 Output sound generation unit 134 Generation unit 135 Synthesis unit 141 Signal output unit 300 3D video reproduction device
Claims
1. An information processing method executed by a plurality of information processing terminals, comprising: a step of acquiring, at a first terminal which is one of the plurality of information processing terminals, first sound information including an acoustic signal and information on the position of a sound source object in a three-dimensional sound field; the first sound information is information for causing the sound source object in the three-dimensional sound field to emit a reproduced sound using the acoustic signal; a step of converting, at the first terminal, the first sound information into second sound information for generating a representative sound arriving at a reference position from a representative point set in the three-dimensional sound field; a step of transmitting, at the first terminal, the second sound information to a second terminal which is another information processing terminal among the plurality of information processing terminals; a step of detecting, at the second terminal, the position or head direction of a user in the three-dimensional sound field; and a step of calculating, at the second terminal, the position of a reproduction representative point corresponding to the position of the representative point based on the detected position or head direction of the user and the reference position. generating, in the second terminal, an output sound signal using a head-related transfer function according to the calculated position of the reproduction representative point and the received second sound information.
2. The information processing method according to claim 1, wherein in the converting step, time shift adjustment and gain adjustment are applied to the reproduced sound to convert it into the representative sound.
3. The information processing method according to claim 2, wherein the amount of time shift adjustment is set to 0 at the representative point.
4. The information processing method according to claim 2, wherein the amount of time shift adjustment at a predetermined position is set using the amount of adjustment at a position closer to the representative point than the predetermined position.
5. The information processing method of claim 2, wherein the amount of gain adjustment at a predetermined position is calculated by using two of the three adjacent representative points surrounding the predetermined position to calculate the amount of adjustment for each of the two representative points that minimizes the error, and then fixing the ratio of the calculated adjustment amounts to calculate the amount of adjustment for the remaining representative point.
6. The information processing method of claim 2, wherein the amount of gain adjustment at a predetermined position is calculated by calculating the amount of adjustment for each of two representative points located horizontally that minimize the error, out of three representative points adjacent to each other that surround the predetermined position, for a horizontal position obtained by removing the vertical component from the predetermined position, and then fixing the ratio of the calculated adjustment amounts and calculating the amount of adjustment for the remaining representative point.
7. The information processing method according to claim 2, wherein the amount of time shift adjustment is corrected based on an expected value of the error so as to reduce the error from the calculated correct value.
8. The information processing method according to claim 2, wherein the amount of gain adjustment at a predetermined position is corrected based on an expected value of the error so as to reduce the error from the calculated correct value.
9. The information processing method according to claim 2, wherein the amount of gain adjustment at a predetermined position is set using two or more of adjustment amounts at a plurality of virtual representative points corresponding to a plurality of directions within a predetermined angle range in the horizontal direction centered on the direction of the representative point of the starting point, with the representative point for the predetermined position as the starting point.
10. An information processing method as described in claim 9, wherein the adjustment amount for the gain adjustment at a predetermined position is set using all of the adjustment amounts at multiple virtual representative points corresponding to multiple directions within a predetermined angle range in the horizontal direction centered on the direction of the representative point of the starting point, with the representative point for the predetermined position as the starting point.
11. The information processing method according to claim 9, wherein the gain adjustment amount at a predetermined position is calculated by averaging the adjustment amounts at a plurality of the virtual representative points.
12. The information processing method according to claim 9, wherein the amount of gain adjustment at a predetermined position is calculated by taking a weighted average of the amounts of adjustment at a plurality of the virtual representative points based on the angle difference between the direction of each of the virtual representative points and the original direction corresponding to that virtual representative point.
13. An information processing method according to any one of claims 2 to 12, wherein the amount of gain adjustment at a predetermined position is set to a predetermined value that is independent of the direction of the predetermined position in a horizontal plane, by adjusting the amount of adjustment at an elevation angle representative point that includes a vertical component among three adjacent representative points that surround the predetermined position.
14. The information processing method according to claim 13, wherein the predetermined value is a value that varies depending on the direction of the predetermined position in a vertical plane.
15. The information processing method according to claim 14, wherein the predetermined value is a value that gradually increases between 0 and 1 when the direction of the predetermined position in the vertical plane is from 0° to 90°.
16. An information processing system including a first terminal and a second terminal, wherein the first terminal comprises: an acquisition unit that acquires first sound information including an acoustic signal and information about the position of a sound source object within a three-dimensional sound field; the first sound information is information for causing the sound source object within the three-dimensional sound field to emit a playback sound using the acoustic signal; a conversion unit that converts the first sound information into second sound information for generating a representative sound arriving at a reference position from a representative point set within the three-dimensional sound field using the acoustic signal; and a transmission unit that transmits the second sound information to the second terminal, wherein the second terminal comprises: a detector that detects a user's position or head direction within the three-dimensional sound field; a calculation unit that calculates a position of a playback representative point corresponding to the position of the representative point based on the detected user's position or head direction and the reference position; and an output unit in the second terminal that outputs an output sound signal using a head related transfer function corresponding to the calculated position of the playback representative point and the received second sound information.
17. A program for causing a computer to execute the information processing method set forth in claim 1.
Citation Information
Patent Citations
Sound source management device, sound source management method, and sound source management system
JP2014236259A
Spatial audio reproduction by positioning at least part of a sound field
JP2023070650A
Sound generation apparatus, sound reproducing apparatus, sound generation method, and sound signal processing program
JP2023164284A
Audio processing device and method, and program
WO2017119318A1