Acoustic information processing method, information processing device, and program

The acoustic information processing method addresses the challenges of high processing demands in virtual environments by determining representative directions for panning processing, resulting in reduced complexity and improved sound localization.

WO2025135070A1PCT designated stage expired Publication Date: 2025-06-26PANASONIC HOLDINGS CORP +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/044753
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-12-18
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing acoustic reproduction technologies face challenges in efficiently processing and transmitting three-dimensional sound information in virtual environments, particularly due to the high processing demands and large storage requirements for realistic sound environments with multiple sound sources and complex acoustic effects.

Method used

An acoustic information processing method that determines representative directions based on the position of sound sources and users in a three-dimensional sound field, using panning processing to distribute sound signals to these directions, thereby reducing processing complexity and improving sound localization.

Benefits of technology

The method effectively reduces processing complexity and improves sound localization by selecting appropriate representative directions, allowing for more efficient conversion processing and enhanced auditory presence in virtual environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024044753_26062025_PF_FP_ABST
    Figure JP2024044753_26062025_PF_FP_ABST
Patent Text Reader

Abstract

This acoustic information processing method includes a step (S201) for acquiring position information of a sound source object within a three-dimensional sound field based on acoustic signals, a step (S201) for acquiring position information of a user within the three-dimensional sound field, a step (S202) for determining a plurality of representative directions, and a step (S203) for performing panning processing for distributing signals from the sound source object to the plurality of representative directions on the basis of the position information of the user and the position information of the sound source object, wherein, in the determining step (S202), if one candidate among selectable representative direction candidates is selected, the one candidate is determined as one of the plurality of representative directions, and if two or more candidates are selected, one candidate close to the forward direction of the user, among the two or more candidates, is determined as one of the plurality of representative directions.
Need to check novelty before this filing date? Find Prior Art

Description

Acoustic information processing method, information processing device, and program

[0001] The present disclosure relates to an acoustic information processing method, an information processing device, and a program.

[0002] Conventionally, there has been known a technology relating to sound reproduction that allows a user (or listener, as the case may be) to perceive three-dimensional sound in a virtual three-dimensional space (see, for example, Patent Document 1). Furthermore, in order to perceive sound as if it is coming from a sound source object to the user in such a three-dimensional space, processing is required to generate output sound information from original sound information. In particular, reproducing three-dimensional sound in response to the user's body movements in a virtual space requires extensive processing. Advances in computer graphics (CG) have made it relatively easy to create visually complex virtual environments, making technology that realizes corresponding auditory information important. In addition, when processing from sound information to output sound information is performed in advance, a large memory area is required to store the pre-calculated processing results. Furthermore, transmitting such large amounts of processing result data may require a wide communication bandwidth.

[0003] To realize a more realistic sound environment, the number of objects that emit sound in the virtual three-dimensional space increases, secondary sounds based on acoustic effects such as reflected sound, diffracted sound, and reverberation increase, and these secondary sounds must be appropriately changed in response to the user's movements, requiring a large amount of processing.To reduce this large amount of processing, a conversion technique known as panning processing (or simply panning) is known, which represents sounds in a three-dimensional space using sounds from several representative points (representative directions) set in advance in the three-dimensional space.

[0004] Japanese Patent Application Laid-Open No. 2020-18620

[0005] However, there is room for improvement in conversion processes such as panning. Therefore, an object of the present disclosure is to provide an acoustic information processing method and the like for appropriately performing conversion processes.

[0006] An acoustic information processing method according to one aspect of the present disclosure is an acoustic information processing method executed by an information processing terminal, and includes the steps of: acquiring position information of a sound source object within a three-dimensional sound field based on an acoustic signal; acquiring position information of a user within the three-dimensional sound field; determining a plurality of representative directions; and performing a panning process that distributes the signal of the sound source object to the plurality of representative directions based on the position information of the user and the position information of the sound source object. In the determining step, one or more candidates from among selectable representative direction candidates are selected based on the smallest deviation from an ideal direction in a horizontal plane, and if only one candidate is selected, the selected candidate is determined to be one of the plurality of representative directions; and if two or more candidates are selected, one of the two or more candidates that is closest to a frontal direction of the user is determined to be one of the plurality of representative directions.

[0007] Moreover, an information processing device according to one aspect of the present disclosure is an information processing system including a first terminal and a second terminal, wherein the first terminal comprises an acquisition unit that acquires sound information including an acoustic signal and information about the position of a sound source object within a three-dimensional sound field, the first sound information being information for causing the sound source object within the three-dimensional sound field to emit a playback sound using the acoustic signal, a conversion unit that converts the first sound information into second sound information for generating a representative sound that arrives at a reference position from a representative point set within the three-dimensional sound field using the acoustic signal, and a transmission unit that transmits the second sound information to the second terminal, wherein the second terminal comprises a detector that detects a user's position or head direction within the three-dimensional sound field, a calculation unit that calculates a position of a playback representative point corresponding to the position of the representative point based on the detected user's position or head direction and the reference position, and an output unit in the second terminal that outputs an output sound signal using a head related transfer function corresponding to the calculated position of the playback representative point and the received second sound information.

[0008] Furthermore, one aspect of the present disclosure can also be realized as a program for causing a computer to execute the above-described acoustic information processing method.

[0009] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0010] According to the present disclosure, it is possible to appropriately perform conversion processing.

[0011] FIG. 1 is a schematic diagram illustrating a use example of a sound reproduction system according to an embodiment. FIG. 2 is a block diagram illustrating a functional configuration of the sound reproduction system according to an embodiment. FIG. 3 is a diagram illustrating an example of an audio signal according to an embodiment. FIG. 4 is a block diagram illustrating a functional configuration of an acquisition unit according to an embodiment. FIG. 5 is a block diagram illustrating a functional configuration of an output sound generation unit according to an embodiment. FIG. 6 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 7 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 8 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 9 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 10 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 11 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 12 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 13 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 14 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 15 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 16 is a diagram illustrating another example of a sound reproduction system according to an embodiment. Fig. 17 is a flowchart of audio reproduction processing according to an embodiment. Fig. 18 is a diagram for explaining synthesis of head-related transfer functions in audio reproduction processing according to an embodiment. Fig. 19 is a diagram for explaining arrangement of a representative direction according to an embodiment. Fig. 20 is a diagram for explaining arrangement of a representative direction according to an embodiment. Fig. 21 is a diagram for explaining determination of a representative direction according to an embodiment. Fig. 22 is a diagram for explaining determination of a representative direction according to an embodiment. Fig. 23 is a diagram for explaining cross-fade processing according to an embodiment. Fig. 24A is a flowchart of operations related to cross-fade processing according to an embodiment. Fig. 24B is another example of a flowchart of operations related to cross-fade processing according to an embodiment.

[0012] (Knowledge that forms the basis of the disclosure) Conventionally, a technology related to sound reproduction that allows a user to perceive stereoscopic sound in a virtual three-dimensional space (hereinafter sometimes referred to as a three-dimensional sound field) has been known (see, for example, Patent Document 1). Using this technology, a user can perceive sound as if a sound source object exists at a predetermined position in the virtual space and the sound is coming from that direction. In order to localize a sound image at a predetermined position in the virtual three-dimensional space in this way, for example, calculation processing is required for a sound signal emitted by a sound source object (also referred to as a sound emitted from the sound source object or a reproduced sound) to generate a sound arrival time difference between the two ears and a sound level difference (or sound pressure difference) between the two ears that causes the sound to be perceived as stereoscopic sound. Such calculation processing is performed by applying a stereophonic filter. A stereophonic filter is an information processing filter that, when an output sound signal obtained by applying the filter to original sound information is reproduced, causes the position (such as the direction and distance) of the sound, the size of the sound source, the width of the space, and the like to be perceived with a three-dimensional effect.

[0013] As an example of the computational process for applying such a stereophonic filter, a process is known in which a head-related transfer function (HRTF) is convolved with a target sound signal to make the sound perceived as coming from a predetermined direction. By performing this HRTF convolution process at a sufficiently fine angle with respect to the direction of arrival of the reproduced sound from the position of the sound source object to the user's position, the sense of realism experienced by the user is improved.

[0014] In recent years, there has been active development of technologies related to virtual reality (VR). Virtual reality focuses on appropriately changing the position of a sound source object in a virtual three-dimensional space in response to a user's movements, allowing the user to experience the sensation of moving within the virtual space. To achieve this, it is necessary to move the localization position of a sound image in the virtual space relative to the user's movements. This processing has been performed by applying a stereophonic filter, such as the head-related transfer function described above, to the original sound information. However, when a user moves within a three-dimensional space, the sound transmission path changes from moment to moment due to changes in the positional relationship between the sound source object and the user, such as due to sound reverberation and interference. This requires determining the sound transmission path from the sound source object based on the positional relationship between the sound source object and the user, and convolving the transfer function to take into account sound reverberation and interference. However, such information processing requires a huge amount of processing, and an improvement in the sense of realism may not be achieved without a large-scale processing device.

[0015] Therefore, in order to reduce such an enormous amount of processing, attempts have been made to apply a panning process to the reproduced sound to reduce the amount of convolution of the head-related transfer function. Specifically, instead of convolving the reproduced sound with a head-related transfer function for each of a number of sound source objects in a three-dimensional space, the reproduced sound from the sound source object is re-expressed using sounds (representative sounds) from several representative points (hereinafter, treated as synonymous with the representative direction from the representative points toward the user) previously set in the three-dimensional space. Then, simply by convolving the representative sound with the head-related transfer function from the representative points to the user's position, it becomes possible to allow the user to perceive a three-dimensional sound that is comparable to that of the original sound source object. If the number of representative points is smaller than the number of original sound source objects, the number of targets for convolution of the head-related transfer function will naturally be reduced, which is advantageous in terms of processing amount.

[0016] Here, the representative direction is naturally determined on the assumption that a head-related transfer function for convolution into the representative direction exists. Furthermore, an ideal representative direction can be determined from the viewpoints of processing resources available for information processing and sound quality improvement. However, if a set of head-related transfer functions that can actually be used does not include a head-related transfer function for the ideal representative direction, the representative direction cannot be used. In this case, it is necessary to select and determine an alternative representative direction from among representative directions that are available (in other words, selectable) in the set of head-related transfer functions used for processing. This disclosure describes an acoustic information processing method and the like that can appropriately perform conversion processing by more appropriately determining a representative direction from such selectable representative directions.

[0017] A more specific outline of the present disclosure is as follows.

[0018] An acoustic information processing method according to a first aspect of the present disclosure is an acoustic information processing method executed by an information processing terminal, and includes the steps of: acquiring position information of a sound source object within a three-dimensional sound field based on an acoustic signal; acquiring position information of a user within the three-dimensional sound field; determining a plurality of representative directions; and performing a panning process that distributes the signal of the sound source object to the plurality of representative directions based on the position information of the user and the position information of the sound source object. In the determining step, one or more candidates from among selectable representative direction candidates are selected based on the smallest deviation from an ideal direction in a horizontal plane, and if only one candidate is selected, the selected candidate is determined to be one of the plurality of representative directions; and if two or more candidates are selected, one of the two or more candidates that is closest to a frontal direction of the user is determined to be one of the plurality of representative directions.

[0019] According to this acoustic information processing method, when determining a representative direction to be used in panning, a candidate representative direction that is closest to the ideal direction (with a small deviation) can be selected and used. Even when two or more candidate representative directions closer to the ideal direction are selected, one candidate closest to the front direction of the user can be used. The closer a candidate is to the front direction of the user (closer to the front side of the user), the easier it is for the user to identify the direction, making it more appropriate as a representative direction for panning. Thus, according to this aspect, a representative direction can be more appropriately determined from selectable representative directions, thereby enabling more appropriate panning (conversion) processing.

[0020] Furthermore, an acoustic information processing method according to a second aspect is the acoustic information processing method according to the first aspect, and in the determining step, one candidate is selected from among the selectable representative direction candidates that has the smallest angle with respect to the user's median plane and that is closest to an elevation angle of 90° from among the direction candidates with an elevation angle of 0 to 90° when the user's horizontal plane direction is an elevation angle of 0°, and the selected candidate is determined as one of the plurality of representative directions.

[0021] According to this, one candidate that has the smallest angle with the user's median plane and is closest to an elevation angle of 90° from among the candidate directions within the range of elevation angles of 0 to 90° when the user's horizontal plane direction is set to an elevation angle of 0° can also be used for the panning process.

[0022] Furthermore, an acoustic information processing method according to a third aspect is the acoustic information processing method according to the first or second aspect, and in the determining step, one candidate of the selectable representative directions is selected from among the candidates of the direction that has the smallest angle with the user's median plane and is closest to an elevation angle of -90° from among candidates of the direction that is within a range of -90° to 0° when the horizontal plane direction of the user has an elevation angle of 0°, and the selected candidate is determined as one of the plurality of representative directions.

[0023] According to this, it is possible to use one candidate that is closest to an elevation angle of -90° from among the candidates for directions that form the smallest angle with the user's median plane and are within the range of elevation angles of -90° to 0° when the user's horizontal plane direction is set to an elevation angle of 0° for the panning process.

[0024] An acoustic information processing method according to a fourth aspect is an acoustic information processing method according to any one of the first to third aspects, in which the ideal directions in a horizontal plane are six ideal directions that divide 360° of the horizontal plane into six equal parts, and in the determining step, one of a plurality of representative directions is determined for each of the six ideal directions.

[0025] This allows one of a plurality of representative directions to be determined for each of six ideal directions that divide 360° of the horizontal plane into six equal parts, and as a result, six representative directions that are close to each of the six ideal directions (with small deviations) can be determined and used for panning processing.

[0026] An acoustic information processing method according to a fifth aspect is the acoustic information processing method according to the fourth aspect, wherein the ideal directions in the horizontal plane are six ideal directions corresponding to 30, 90, 150, 210, 270, and 330° on the horizontal plane when the front direction of the user is set to 0°, and in the determining step, one of a plurality of representative directions is determined for each of the six ideal directions.

[0027] This allows one of multiple representative directions to be determined for each of six ideal directions corresponding to 30, 90, 150, 210, 270, and 330 degrees of the horizontal plane when the user's front direction is 0 degrees.

[0028] In addition, an information processing device according to a sixth aspect includes a first acquisition unit that acquires position information of a sound source object within a three-dimensional sound field based on an acoustic signal, a second acquisition unit that acquires position information of a user within the three-dimensional sound field, a representative direction determination unit that determines a plurality of representative directions, and a panning unit that performs a panning process to distribute the signal of the sound source object to the plurality of representative directions based on the position information of the user and the position information of the sound source object, wherein the representative direction determination unit selects one or more candidates from among the selectable representative direction candidates based on the smallest deviation from an ideal direction in a horizontal plane, and if only one candidate is selected, determines that one candidate as one of the plurality of representative directions, and if two or more candidates are selected, determines that one candidate from the two or more candidates that is closest to a frontal direction of the user as one of the plurality of representative directions.

[0029] This can achieve the same effects as the acoustic information processing method described above.

[0030] A program according to a seventh aspect is a program for causing a computer, such as an information processing terminal, to execute the acoustic information processing method according to any one of the above aspects.

[0031] This makes it possible to achieve the same effects as the above-described acoustic information processing method using a computer.

[0032] Furthermore, an acoustic information processing method according to another first aspect of the present disclosure is an acoustic information processing method executed by an information processing terminal, and includes the steps of: acquiring sound information including a first acoustic signal and positional information of a sound source object within a three-dimensional sound field, the sound information being information for causing the sound source object within the three-dimensional sound field to emit a reproduced sound using the first acoustic signal; acquiring positional information of a user within the three-dimensional sound field; and performing a panning process that distributes the first acoustic signal related to the sound source object in a plurality of representative directions using parameters based on the user's positional information and the positional information of the sound source object, wherein in the panning process, a time fade process is performed on the first acoustic signal and a second acoustic signal that is continuous with the first acoustic signal in the time domain using the parameters.

[0033] According to this, by using the parameters used in the panning process, it is possible to perform time fade processing on the first acoustic signal and the second acoustic signal. For example, the time fade processing includes fading out the first acoustic signal and fading in the second acoustic signal. The time fade processing involves a determination of whether or not to perform the time fade processing, and this determination may be made using the parameters used in the panning process. In this way, it is possible to more appropriately perform the conversion process by performing the time fade processing using the parameters used in the panning process.

[0034] Furthermore, an acoustic information processing method according to yet another first aspect of the present disclosure is an acoustic information processing method executed by an information processing terminal, and includes the steps of: acquiring sound information including an acoustic signal and positional information of a sound source object within a three-dimensional sound field; acquiring positional information of a user within the three-dimensional sound field; performing a panning process to distribute the acoustic signal related to the sound source object to the plurality of representative directions based on the positional information of the user and the positional information of the sound source object; and acquiring a data set of head-related transfer functions to be used for the distributed acoustic signal, wherein the step of performing the panning process is executed when the number of head-related transfer functions included in the acquired data set of head-related transfer functions satisfies a predetermined condition.

[0035] According to this, when the number of head-related transfer functions included in the data set of the acquired head-related transfer functions satisfies a predetermined condition, it is possible to execute the step of performing the panning process. In other words, it is possible to switch whether or not to execute the step of performing the panning process depending on a predetermined condition related to the number of head-related transfer functions. For example, by setting a condition related to the number of head-related transfer functions that reflects a situation suitable for executing the step of performing the panning process, it is possible to execute the step of performing the panning process in such a situation suitable for execution, and not execute the step of performing the panning process in a situation unsuitable for execution.

[0036] Furthermore, an acoustic information processing method according to yet another second aspect of the present disclosure is the acoustic information processing method according to yet another first aspect, wherein, when the number of head-related transfer functions included in the data set of the head-related transfer functions does not satisfy the predetermined condition, instead of performing the panning processing, the acoustic signal is processed using a head-related transfer function corresponding to the direction of arrival based on position information of the listener and position information of the sound source.

[0037] According to this, it is possible to process acoustic signals using head-related transfer functions according to the direction of arrival instead of executing a step of performing panning processing, depending on a predetermined condition related to the number of head-related transfer functions. For example, by setting a condition related to the number of head-related transfer functions that reflects a situation suitable for executing a step of performing panning processing, it is possible to execute the step of performing panning processing in such a situation suitable for execution, while processing acoustic signals using head-related transfer functions according to the direction of arrival instead of executing a step of performing panning processing in a situation unsuitable for execution.

[0038] Furthermore, these comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0039] Hereinafter, embodiments will be described in detail with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not recited in independent claims will be described as optional components. Note that each figure is a schematic diagram and is not necessarily an exact illustration. Furthermore, in each figure, substantially identical components are assigned the same reference numerals, and duplicated descriptions may be omitted or simplified.

[0040] In the following description, elements may be assigned ordinal numbers such as first, second, and third. These ordinal numbers are assigned to elements in order to identify them and do not necessarily correspond to a meaningful order. These ordinal numbers may be rearranged, newly added, or removed as appropriate.

[0041] In the following description, an acoustic signal included in sound information may be described, but the acoustic signal may also be referred to as a voice signal or a sound signal. In other words, in the present disclosure, the acoustic signal has the same meaning as the voice signal or the sound signal.

[0042] (Embodiment) [Overview] First, an overview of an audio reproduction system according to an embodiment will be described. Fig. 1 is a schematic diagram showing an example of use of an audio reproduction system according to an embodiment. Fig. 1 shows a user 99 using an audio reproduction system 100.

[0043] The sound reproduction system 100 shown in FIG. 1 is used simultaneously with, for example, a three-dimensional video reproduction device 300. By simultaneously viewing three-dimensional images and three-dimensional sound, the image enhances the auditory sense of realism, and the sound enhances the visual sense of realism, allowing the user to experience the image and sound as if they were actually at the scene where they were captured. For example, when an image (moving image) of people having a conversation is displayed, it is known that even if the localization of the sound image (sound source object) of the conversation sound is not aligned with the person's mouth, the user 99 will perceive the conversation sound as coming from the person's mouth. In this way, the visual information can correct the position of the sound image, and the image and sound can be combined to enhance the sense of realism.

[0044] The three-dimensional video reproduction device 300 is an image display device worn on the head of the user 99. Therefore, the three-dimensional video reproduction device 300 moves integrally with the head of the user 99. For example, as shown in the figure, the three-dimensional video reproduction device 300 is a glasses-type device that is supported by the ears and nose of the user 99.

[0045] The three-dimensional video reproduction device 300 changes the displayed image in accordance with the movement of the user 99's head, thereby making the user 99 perceive the movement of his or her head in the three-dimensional image space. In other words, when an object in the three-dimensional image space is located in front of the user 99, if the user 99 turns to the right, the object moves to the left of the user 99, and if the user 99 turns to the left, the object moves to the right of the user 99. In this way, the three-dimensional video reproduction device 300 moves the three-dimensional image space in the opposite direction to the movement of the user 99.

[0046] The 3D video playback device 300 displays two images with a parallax difference to each of the user's 99's left and right eyes. The user 99 can perceive the three-dimensional position of an object on the image based on the parallax difference between the displayed images. Note that when the user 99 uses the audio playback system 100 with their eyes closed, for example, when using it to play healing sounds for sleep induction, the 3D video playback device 300 does not need to be used at the same time. In other words, the 3D video playback device 300 is not an essential component of the present disclosure. In addition to dedicated video display devices, the 3D video playback device 300 may also be a general-purpose mobile terminal owned by the user 99, such as a smartphone or tablet device.

[0047] Such general-purpose mobile terminals are equipped with not only a display for displaying images but also various sensors for detecting the terminal's posture and movement. Furthermore, they are also equipped with a processor for information processing, and are capable of connecting to a network to transmit and receive information to and from a server device such as a cloud server. In other words, the 3D video playback device 300 and the audio playback system 100 can be realized by combining a smartphone with general-purpose headphones or the like that do not have an information processing function.

[0048] As in this example, the head movement detection function, the video presentation function, the video information processing function for presentation, the sound presentation function, and the sound information processing function for presentation may be appropriately arranged in one or more devices to realize the 3D video reproduction device 300 and the sound reproduction system 100. If the 3D video reproduction device 300 is not required, it is sufficient to appropriately arrange the head movement detection function, the sound presentation function, and the sound information processing function for presentation in one or more devices. For example, the sound reproduction system 100 can be realized by a processing device such as a computer or smartphone having a sound information processing function for presentation, and headphones or the like having a head movement detection function and a sound presentation function.

[0049] The sound reproduction system 100 is a sound presentation device that is worn on the head of the user 99. Therefore, the sound reproduction system 100 moves integrally with the head of the user 99. For example, the sound reproduction system 100 in this embodiment is a so-called over-ear headphone type device. Note that there are no particular limitations on the form of the sound reproduction system 100, and it may be, for example, two earplug-type devices that are worn independently on the left and right ears of the user 99.

[0050] The sound reproduction system 100 changes the sound presented in accordance with the movement of the head of the user 99, thereby making the user 99 perceive as if he or she is moving his or her head within the three-dimensional sound field. For this reason, as described above, the sound reproduction system 100 moves the three-dimensional sound field in the opposite direction to the movement of the user 99.

[0051] Here, when the user 99 moves within the three-dimensional sound field, the position of the sound source object relative to the position of the user 99 within the three-dimensional sound field changes. As a result, each time the user 99 moves, it is necessary to perform calculation processing based on the positions of the sound source object and the user 99 to generate an output sound signal for playback. Since such processing typically requires a huge amount of processing, in the present disclosure, a panning process is applied as one type of conversion processing to represent the reproduced sound as a representative sound from a representative point, in order to reduce the amount of processing. As a result, it is possible to allow the user 99 to perceive the reproduced sound from the sound source object simply by convolving a head-related transfer function with the representative sound. Hereinafter, in the present embodiment, a case will be described in which panning processing is used as an example of conversion processing. However, the conversion processing is not limited to panning processing, and any conversion processing can be applied as long as the conversion processing is expected to reduce the amount of processing, depending on the conditions. Furthermore, the panning process will be described using specific examples, but the panning process is not limited to the specific example described below, and existing panning process techniques such as VBAP (Vector Based Amplitude Panning), DBAP (Distance Based Amplitude Panning), and Ambisonics can also be applied.

[0052] [Configuration] Next, the configuration of the sound reproduction system 100 according to this embodiment will be described with reference to Fig. 2. Fig. 2 is a block diagram showing the functional configuration of the sound reproduction system according to this embodiment.

[0053] As shown in FIG. 2, the sound reproduction system 100 according to this embodiment includes an information processing device 101, a communication module 102, a detector 103, a driver 104, and a database 105.

[0054] The information processing device 101 is an arithmetic device for performing various signal processing in the sound reproduction system 100. The information processing device 101 includes a processor and a memory, such as a computer, and is realized by the processor executing a program stored in the memory. Execution of this program provides functions related to each functional unit described below.

[0055] The information processing device 101 includes an acquisition unit 111, a path calculation unit 121, an output sound generation unit 131, and a signal output unit 141. Details of each functional unit included in the information processing device 101 will be described below together with details of the configuration other than the information processing device 101.

[0056] The communication module 102 is an interface device for accepting input of sound information to the sound reproduction system 100. The communication module 102 includes, for example, an antenna and a signal converter, and receives sound information from an external device via wireless communication. More specifically, the communication module 102 receives, using the antenna, a wireless signal indicating sound information converted into a format for wireless communication, and reconverts the wireless signal into sound information using the signal converter. In this way, the sound reproduction system 100 acquires sound information from the external device via wireless communication. The sound information acquired by the communication module 102 is acquired by the acquisition unit 111. In this way, the acquisition unit 111 is an example of a sound acquisition unit. The sound information is input to the information processing device 101 in the above manner. Note that communication between the sound reproduction system 100 and the external device may be performed via wired communication.

[0057] The sound information acquired by the sound reproduction system 100 is encoded in a predetermined format, such as MPEG-H 3D Audio (ISO / IEC 23008-3). As an example, the encoded sound information includes information about the sound reproduced by the sound reproduction system 100 and information about the localization position when the sound image of the sound is localized at a predetermined position in a three-dimensional sound field (i.e., perceived as sound coming from a predetermined direction). The sound information can also be interpreted as information about a sound source object. In other words, the sound information includes the position of the sound source object in the three-dimensional sound field and the sound produced by the sound source object. The sound information may also include a flag for determining whether or not to apply panning processing. This flag will be described later.

[0058] As described above, sound information is obtained as input data, and includes an audio signal (acoustic signal), which is information about the reproduced sound, and other information, such as information about the position of a sound source object in a three-dimensional sound field. The other information may also include information for defining a three-dimensional sound field. Therefore, the other information may be collectively referred to as information about space (spatial information), including information about the position of a sound source object and information for defining a three-dimensional sound field. When the audio signal is viewed as the main focus, the input data can be said to be sound information in which other information (metadata) is attached to the audio signal. When the spatial information is viewed as the main focus, the input data can be said to be information in which the audio signal is attached to the spatial information. Alternatively, since the input data has both aspects, the input data can also be considered as sound spatial information.

[0059] As a specific example, the sound information includes information about a plurality of sounds including a first reproduced sound and a second reproduced sound, and the sound images of the respective sounds when reproduced are localized so that they are perceived as coming from different positions in a three-dimensional sound field. Therefore, the sound source object of the first reproduced sound is localized at a first position in the three-dimensional sound field, and the sound source object of the second reproduced sound is localized at a second position in the three-dimensional sound field. Thus, the sound information may include a plurality of sounds. In other words, the sound information may include a plurality of audio signals corresponding to the first reproduced sound and the second reproduced sound, respectively, and the positions of a plurality of sound source objects at first and second positions that correspond one-to-one to the plurality of audio signals.

[0060] 3 is a diagram illustrating an example of an audio signal according to an embodiment. For example, as shown in (a) of FIG. 3, the sound information may include an audio signal of a first direct sound arriving from a first position (from a first direction) to the position of the user 99 and an audio signal of a second direct sound arriving from a second position (from a second direction) to the position of the user 99. Note that the sound information immediately after acquisition may include only information about the reproduced sound. In this case, information about the predetermined position may be acquired separately, and the subsequent processing may be performed once the information is collected.

[0061] Alternatively, the sound information may include multiple audio signals and the position of a single sound source object that corresponds to the multiple audio signals in a many-to-one relationship. For example, such sound information is used in a situation where multiple sounds are reproduced from a sound source object. For example, each of the multiple audio signals corresponds to a direct sound that arrives directly from the position of the sound source object to the position of the user 99, and a secondary sound (sound generated by indirect propagation) that accompanies the direct sound and arrives via a path different from that of the direct sound.

[0062] For example, as shown in (b) of FIG. 3 , the sound information immediately after acquisition includes an audio signal related to the direct sound, and is converted into sound information including audio signals such as reverberation, primary reflection, and diffraction through a conversion process that calculates secondary sounds. This conversion process that calculates the secondary sounds uses information on the spatial environment conditions of the three-dimensional sound field (e.g., the position, reflection, and diffraction characteristics of objects in the three-dimensional sound field). Thus, secondary sounds are computationally generated from sound information related to a single reproduced sound based on the spatial environment conditions of the three-dimensional sound field. Therefore, the secondary sounds are not included in the sound information immediately after acquisition, and sound information including these secondary sounds is generated through the conversion process that calculates the secondary sounds. Another secondary sound may be generated from one secondary sound through its propagation. Note that the information on the spatial environment conditions is part of the spatial information and may be acquired together with the audio signal from the input sound information. Alternatively, the audio signal and the spatial information may be acquired separately. That is, the sound information may be acquired from a single file or bitstream, or may be acquired separately by dividing it into multiple files or bitstreams. For example, the audio signal and the spatial information may be acquired from separate files or bitstreams, or the audio signal and the spatial information may each be acquired from a plurality of files or bitstreams.

[0063] As described above, there is no particular limitation on the form of the input sound information, and the sound reproduction system 100 may be provided with an acquisition unit 111 that can accommodate various forms of sound information.

[0064] An example of the acquisition unit 111 will now be described with reference to Fig. 4. Fig. 4 is a block diagram showing the functional configuration of the acquisition unit according to the embodiment. As shown in Fig. 4, the acquisition unit 111 in this embodiment includes, for example, an encoded sound information input unit 112, a decoding processing unit 113, and a sensing information input unit 114.

[0065] The encoded sound information input unit 112 is a processing unit to which the coded (in other words, encoded) sound information acquired by the acquisition unit 111 is input. The encoded sound information input unit 112 outputs the input sound information to the decoding processing unit 113. The decoding processing unit 113 is a processing unit that decodes (in other words, decodes) the sound information output from the encoded sound information input unit 112 to generate a playback sound included in the sound information and the position of a sound source object in a format used for subsequent processing. In this way, the acquisition unit 111 functions as a first acquisition unit that acquires the position of a sound source object using the encoded sound information input unit 112 and the decoding processing unit 113. The sensing information input unit 114 will be described below together with the function of the detector 103.

[0066] The detector 103 is a device for detecting the speed of movement of the head of the user 99. The detector 103 is configured by combining various sensors used for detecting movement, such as a gyro sensor and an acceleration sensor. In this embodiment, the detector 103 is built into the sound reproduction system 100. However, the detector 103 may be built into an external device, such as a three-dimensional image reproduction device 300 that operates in response to the movement of the head of the user 99 in the same way as the sound reproduction system 100. In this case, the detector 103 does not need to be included in the sound reproduction system 100. Alternatively, the detector 103 may be an external imaging device or the like that captures the movement of the head of the user 99 and detects the movement of the user 99 by processing the captured image.

[0067] The detector 103 is, for example, fixed integrally to the housing of the sound reproduction system 100 and detects the speed of movement of the housing. After the sound reproduction system 100 including the housing is worn by the user 99, it moves integrally with the head of the user 99, and as a result, the detector 103 can detect the speed of movement of the head of the user 99.

[0068] The detector 103 may detect, for example, the amount of head movement of the user 99 as the amount of rotation about at least one of three axes that are orthogonal to each other in three-dimensional space as the rotation axis, or may detect the amount of displacement about at least one of the three axes as the displacement direction. Furthermore, the detector 103 may detect both the amount of rotation and the amount of displacement as the amount of head movement of the user 99.

[0069] The sensing information input unit 114 acquires the movement speed of the user 99's head from the detector 103. More specifically, the sensing information input unit 114 acquires the amount of head movement of the user 99 detected by the detector 103 per unit time as the movement speed. In this way, the sensing information input unit 114 acquires at least one of the rotation speed and the displacement speed from the detector 103. The amount of head movement of the user 99 acquired here is used to determine the position and posture (in other words, coordinates and orientation) of the user 99 in the three-dimensional sound field. Therefore, the acquisition unit 111 also functions as a position acquisition unit (second acquisition unit) by the sensing information input unit 114. In the sound reproduction system 100, the relative position of the sound image object with respect to the user 99 is determined based on the determined coordinates and orientation of the user 99, and sound is reproduced. Specifically, the above functions are realized by the path calculation unit 121 and the output sound generation unit 131.

[0070] The path calculation unit 121 includes an arrival direction calculation function that calculates the relative arrival direction of the reproduced sound from the position of the sound source object to the position of the user 99 based on the determined coordinates and orientation of the user 99, and the conversion process that calculates the secondary sound described above. Therefore, the path calculation unit 121 includes a function that calculates a propagation path from the sound source object and calculates the secondary sound and the arrival direction of the secondary sound that arrives at the position of the user 99 through indirect propagation of the reproduced sound according to the calculated propagation path of the reproduced sound. The arrival direction of the secondary sound includes additional information, such as what object the sound is reflected from in the case of a reflected sound and the attenuation rate at the time of reflection. The additional information is included in the arrival direction of the secondary sound calculated from the input sound information. In other words, the additional information is computationally generated and acquired from the sound information.

[0071] To summarize the spatial information, the spatial information includes further information such as the spatial position of a sound source object in a space (three-dimensional sound field) (information on the position of the sound source object), sound reflection and diffraction characteristics at the sound source object (together with information on the conditions of the spatial environment), and the size of the three-dimensional sound field. Based on the spatial information, the path calculation unit 121 generates secondary sounds depending on which sound source object the reproduced sound is reflected or diffracted by, and calculates, as additional information, the direction from which the secondary sounds arrive and the volume of the secondary sounds after attenuation by reflection or diffraction. The sound information (input data) includes spatial information in the form of audio signals and accompanying metadata. As described above, the spatial information includes, as information other than the audio signal, information necessary to turn the sound into stereophonic sound and position the sound source object in a three-dimensional sound field, and / or information used to calculate information necessary to turn the sound into stereophonic sound and position the sound source object in a three-dimensional sound field.

[0072] The path calculation unit 121 may be realized by any processing as long as it can calculate the arrival direction of the reproduced sound when the reproduced sound reaches the user as a direct sound and can calculate the arrival direction of a secondary sound that arrives at the position of the user 99 due to secondary propagation of the reproduced sound. The path calculation unit 121 determines from which direction in the three-dimensional sound field the reproduced sound and the secondary sound are to be perceived by the user 99 as coming from, based on the coordinates and orientation of the user 99, and processes the sound information so that the output sound signal is perceived as such a sound when it is reproduced.

[0073] The output sound generating unit 131 is a processing unit that processes information about the reproduced sound included in the sound information to generate an output sound signal.

[0074] An example of the output sound generation unit 131 will now be described with reference to Fig. 5. Fig. 5 is a block diagram showing the functional configuration of the output sound generation unit according to the embodiment. As shown in Fig. 5, the output sound generation unit 131 in the present embodiment includes, for example, a representative direction determination unit 133, a generation unit 134, and a synthesis unit 135.

[0075] The representative direction determination unit 133 is a processing unit that acquires a set of head-related transfer functions as information corresponding to selectable representative direction candidates, and determines a representative direction (hereinafter also referred to as a representative direction to be used) corresponding to a head-related transfer function to be convolved with a representative sound after panning processing from among the selectable representative direction candidates (included in the acquired set of head-related transfer functions). The representative direction determination unit 133 performs processing in four stages. First, in the first stage, if an ideal representative direction is included in the selectable representative direction candidates under the condition that the head-related transfer function is included, the representative direction determination unit 133 determines the representative direction candidate as the representative direction to be used.

[0076] As will be described in detail later, the representative direction determination unit 133 determines representative directions to be used in eight directions corresponding to a total of eight ideal representative directions, including six directions in a horizontal plane and two directions in the median plane of the user 99. The number of directions, eight, here is just an example, and the number may be two or more but less than eight, or nine or more. Furthermore, the breakdown of the eight ideal representative directions is not limited to the two mentioned above, and may be, for example, three directions in total, including two directions in a horizontal plane and one direction in the median plane of the user 99.

[0077] Next, in the second stage, if the ideal representative direction is not included in the selectable representative direction candidates, the representative direction determination unit 133 first selects one or more representative direction candidates for each ideal representative direction in the horizontal plane based on the smallness of deviation from the ideal representative direction. For example, the representative direction determination unit 133 selects a representative direction candidate with the smallest deviation from the ideal representative direction. Alternatively, the representative direction determination unit 133 may select a representative direction candidate with a deviation from the ideal representative direction within a predetermined range. Here, "deviation" may refer only to deviations in components in the horizontal plane direction, or may refer to deviations including deviations in components outside the horizontal plane direction. Then, if one representative direction candidate is selected in the above selection, that one representative direction candidate is determined to be one of the representative directions to be used.

[0078] In the third step, if there are two or more candidates for the selected representative direction (if the candidates for the selected representative direction are competing), the candidate representative direction that is closest to the front direction of the user 99 among the two or more selected representative directions is determined as one of the representative directions to be used.

[0079] The representative direction determiner 133 sequentially or simultaneously selects and determines candidates for the second and third stages for each of the six ideal representative directions in the horizontal plane. That is, at least one of the second and third stages is performed six times. In this way, the representative direction to be used for the six ideal representative directions in the horizontal plane is determined.

[0080] Next, in a fourth step, if the ideal representative direction is not included in the selectable representative direction candidates, the representative direction determination unit 133 determines a representative direction to be used for each of the two ideal representative directions in the median plane of the user 99. Specifically, the ideal representative directions in the two directions in the median plane of the user 99 are directions with an elevation angle of 90° and an elevation angle of −90° (depression angle of 90°) when the horizontal plane direction of the user 99 is set to an elevation angle of 0°.

[0081] Here, the representative direction determination unit 133 selects only candidate representative directions on the front side of the user 99 (the side in front of the so-called coronal plane), and excludes candidate representative directions on the back side of the user 99 (the side behind the so-called coronal plane). This is because the back side of the user 99 is an area where the user 99 has a relatively low directional resolution (where it is difficult to perceive differences in the direction from which sounds are coming). In other words, if a candidate representative direction on the back side of the user 99 is used as the representative direction to be used, the sense of direction of the sound is likely to be lost, which in turn leads to a deterioration in sound quality. Therefore, by using a candidate representative direction on the front side of the user 99 as the representative direction to be used, the sense of direction of the sound is less likely to be lost, which in turn leads to an effect of easily suppressing a deterioration in sound quality.

[0082] To achieve the above, the representative direction determination unit 133 selects one representative direction candidate closest to an elevation angle of 90° from among representative direction candidates within an elevation angle range of 0 to 90° when the horizontal plane direction of the user 99 is set to an elevation angle of 0°, and selects one representative direction candidate closest to an elevation angle of -90° from among representative direction candidates within an elevation angle range of -90 to 0°.

[0083] Note that there may be cases where there are no selectable representative direction candidates within the median plane. In such cases, the representative direction determination unit 133 selects the representative direction candidate that forms the smallest angle with the median plane. In other words, the representative direction determination unit 133 selects one candidate that forms the smallest angle with the median plane of the user 99 and is closest to an elevation angle of 90° from among the candidate directions that form the smallest angle with the median plane of the user 99 and are within a range of elevation angles of 0 to 90° when the horizontal plane direction of the user 99 is set to an elevation angle of 0°, and determines this candidate as one of the representative directions to be used. Furthermore, the representative direction determination unit 133 selects one candidate that forms the smallest angle with the median plane of the user 99 and is closest to an elevation angle of -90° from among the candidate directions that form the smallest angle with the median plane of the user 99 and are within a range of elevation angles of -90 to 0° when the horizontal plane direction of the user 99 is set to an elevation angle of 0°, and determines this candidate as one of the representative directions to be used.

[0084] The representative direction determination unit 133 determines a representative direction to be used for the ideal representative directions in two directions within the median plane of the user 99 by selecting and determining candidates in the fourth stage described above. The generation unit 134 is a processing unit used when applying a panning process to perform a conversion process to convert the reproduced sound into a representative sound, and then convolving a head-related transfer function with the converted representative sound. The generation unit 134 acquires the reproduced sound and the position of the representative point, and performs a conversion process to convert the reproduced sound into a representative sound for reproducing the sound from the representative point. Note that the generation unit 134 has the same function as the panning unit described later in FIG. 15.

[0085] For example, if a sound source object is located midway between two representative points, a sound is generated so that the same sound as the playback sound is emitted from each of the two representative points. In other words, the playback sound is distributed to the two representative points. Then, a representative sound can be generated by adjusting the gain of the generated sound to match the position of the sound source object. Conversion of the playback sound to a representative sound is not limited to this example. For example, as described below, conversion of the playback sound to a representative sound may be performed by time shift adjustment and gain adjustment, or any other existing conversion may be used as long as the playback sound can be converted to a representative sound that is reproduced as a sound from a representative point. Furthermore, in this specification, the conversion process of the playback sound to a representative sound may be interpreted as a process of distributing the playback sound to representative points (representative directions). Specifically, sound signals of the playback sound associated with the positions of each sound source object are distributed to the positions of the representative points, and a representative sound arriving from the representative points (representative directions) to the listener is generated. Here, the representative direction refers to the direction of the representative point as seen from the listener, or the direction of the listener as seen from the representative point. An example of the conversion that performs the time shift adjustment and the gain adjustment will be described later. The generation unit 134 acquires the same number of representative sounds as the number of representative points obtained by the conversion and head-related transfer functions corresponding to the representative directions from each representative point to the position of the user 99, and performs a convolution process of the acquired head-related transfer functions on the representative sounds to generate a sound signal.

[0086] That is, the panning unit performs panning to represent the sound source by panning sounds from a specific representative direction based on the sound source directions of multiple sound sources (target signals) acquired by the path calculation unit 121, by time shifting the sound sources and adjusting their gains. Specifically, the panning unit synthesizes the sound source (target signal) by panning in a representative direction that approximates the sound source direction of the sound source. As a result, the panning unit generates an HRIR for the sound source direction equivalently. Here, in this embodiment, "equivalent" and "equivalently" refer to signals with an error below a specific level and substantially similar signals, as shown in the examples described below. Specifically, the panning unit generates an HRIR for the sound source direction equivalently by panning the sound source by synthesizing HRIRs for several directions that are closest to the sound source direction of the sound source or that are most similar to the HRIR for the sound source direction. In this embodiment, this direction will be referred to as a "specific representative direction" (hereinafter simply referred to as a "representative direction") described below. This reduces the amount of calculation required to generate the ear signal.

[0087] That is, the panning unit synthesizes a sound image from multiple sound sources using sounds from multiple representative directions. For example, two or three representative directions can be used, but the number of representative directions is not limited to this. Specifically, the panning unit can group together the sound sources into a number of representative points that is fewer than the number of sound sources, and synthesize a sound image using only the HRIRs of the representative directions for these representative points.

[0088] At this time, the panning unit calculates a time shift (delay) that maximizes the cross-correlation between the HRIR in the sound source direction and the HRIR in the representative direction. The time shift obtained here, or a time shift obtained by adding a negative sign to this time shift, is applied to the sound source, and the subsequent processing is performed assuming that the signal after the time shift is in the representative direction.

[0089] This time shift may be a time shift shorter than the sampling frequency (a shift in which the sample position is expressed by a decimal number; hereinafter, referred to as a "decimal shift"). This decimal shift can be performed by oversampling.

[0090] Here, the panning unit applies a gain to the signal of the representative direction obtained by time-shifting the sound source, and calculates the sum of the values ​​calculated for each representative point convolved with the HRIR at each representative point, thereby synthesizing a signal equivalent to the sound source convolved with the HRIR of the sound source direction.

[0091] On the other hand, when synthesizing the HRIR (vector) of the sound source direction by the sum of the HRIR (vector) of the representative direction, the panning unit may calculate the gain by orthogonalizing the error signal vector between the synthesized HRIR (vector) and the HRIR (vector) of the sound source direction to the HRIR (vector) of the representative direction. Note that the HRIR (vector) is a time waveform of the HRIR that is considered to be a vector. Hereinafter, this HRIR (vector) will also be referred to as an "HRIR vector."

[0092] The panning unit corrects this gain so that the energy balance of the HRIRs for the left and right ears from the sound source position is maintained in the HRIR substantially synthesized by panning using HRIRs from multiple representative points. In other words, the panning unit may correct the gain so that the energy balance of the HRIRs for the left and right ears of the listener from the sound source is maintained in the HRIR substantially synthesized by panning.

[0093] In this embodiment, the panning unit can calculate the gain value of the HRIR gain in the representative direction and the time shift value corresponding to the time shift of the HRIR for each sound source direction of the sound source, and store them in the HRIR table 200 described later.

[0094] The panning unit then time-shifts each sound source using a time shift value and gain value corresponding to the sound source direction of each sound source, multiplies the gain, and sums them to generate a sum signal. The panning unit treats this sum signal as being present at the position of the representative point. The panning unit can convolve the HRIR at the position of the representative point with this sum signal to generate a signal at the listener's ear.

[0095] The synthesis unit 135 generates an output sound signal. The synthesis unit 135 may perform EQ adjustment on the sound signal. Specifically, in the panning process, EQ adjustment may be performed to increase the gain of a high-frequency domain that is likely to be attenuated, thereby emphasizing this high-frequency domain. Therefore, the synthesis unit 135 functions as an EQ adjustment unit. Note that, when there are multiple sound signals, the EQ adjustment performed by the synthesis unit 135 may be performed on only some or all of the multiple sound signals.

[0096] Referring again to FIG. 2 , the output sound generation unit 131 acquires a head-related transfer function used to generate an output sound signal from the database 105. The database 105 is an information storage device that functions both as a storage device for storing information and as a storage controller that reads out the stored information and outputs it to an external component. The database 105 stores a head-related transfer function for each direction of arrival of the sound from the user 99. The head-related transfer functions included in the database 105 are a set of general-purpose head-related transfer functions that can be used by everyone, a set of head-related transfer functions optimized for each individual user 99, or a set of head-related transfer functions that are publicly available. The database 105 receives an inquiry from the output sound generation unit 131 using the direction of arrival as a query, and outputs a head-related transfer function corresponding to the direction of arrival to the output sound generation unit 131. In addition, the output sound generation unit 131 may output the entire set of head-related transfer functions, or may output the characteristics of the set of head-related transfer functions itself.

[0097] The signal output unit 141 is a functional unit that outputs the generated output sound signal to the driver 104. The signal output unit 141 generates a waveform signal by performing signal conversion from a digital signal to an analog signal based on the output sound signal, and then causes the driver 104 to generate sound waves based on the waveform signal, thereby presenting sound to the user 99. The driver 104 includes, for example, a diaphragm and a drive mechanism such as a magnet and a voice coil. The driver 104 operates the drive mechanism in response to the waveform signal, causing the drive mechanism to vibrate the diaphragm. In this way, the driver 104 generates sound waves by vibrating the diaphragm in response to the output sound signal (this means "reproducing" the output sound signal; in other words, "reproducing" does not include perception by the user 99). The sound waves propagate through the air and reach the ears of the user 99, and the user 99 perceives the sound.

[0098] [Another Configuration Example] In the above example, the sound reproduction system 100 according to the present embodiment is a sound presentation device, and has been described as including an information processing device 101, a communication module 102, a detector 103, a database 105, and a driver 104. However, the functions of the sound reproduction system 100 may be realized by a plurality of devices or by a single device. Specific examples will be described using Figures 6 to 15. Figures 6 to 15 are diagrams for explaining other examples of the sound reproduction system according to the embodiment.

[0099] For example, the information processing device 601 may be included in the audio presentation device 602, and the audio presentation device 602 may perform both acoustic processing and sound presentation. Alternatively, the information processing device 601 and the audio presentation device 602 may share the acoustic processing described in the present disclosure, or a server connected to the information processing device 601 or the audio presentation device 602 via a network may perform part or all of the acoustic processing described in the present disclosure.

[0100] In the above explanation, the information processing device 601 is referred to as such, but when the information processing device 601 performs acoustic processing by decoding a bitstream generated by encoding at least a portion of the data of an audio signal or spatial information used for acoustic processing, the information processing device 601 may be referred to as a decoding device, and the acoustic reproduction system 100 (i.e., the stereophonic sound reproduction system 600 in the figure) may be referred to as a decoding processing system.

[0101] Here, an example will be described in which the sound reproduction system 100 functions as a decoding processing system.

[0102] <Example of Encoding Device> FIG. 7 is a functional block diagram showing the configuration of an encoding device 700 that is an example of an encoding device according to the present disclosure.

[0103] Input data 701 is data to be coded, including spatial information and / or an audio signal, that is input to an encoder 702. Details of the spatial information will be described later.

[0104] The encoder 702 encodes the input data 701 to generate encoded data 703. The encoded data 703 is, for example, a bit stream generated by the encoding process.

[0105] The memory 704 stores the encoded data 703. The memory 704 may be, for example, a hard disk or a solid-state drive (SSD), or may be any other storage device.

[0106] In the above description, a bitstream generated by an encoding process is given as an example of the encoded data 703 stored in memory 704, but data other than a bitstream may also be used. For example, the encoding device 700 may convert a bitstream into a predetermined data format and store the converted data in memory 704. The converted data may be, for example, a file or multiplexed stream storing one or more bitstreams. Here, the file may have a file format such as ISOBMFF (ISO Base Media File Format). The encoded data 703 may also be in the form of multiple packets generated by dividing the bitstream or file. When converting the bitstream generated by the encoder 702 into data other than the bitstream, the encoding device 700 may be provided with a conversion unit (not shown), or the conversion process may be performed by a CPU (Central Processing Unit).

[0107] <Example of Decoding Device> FIG. 8 is a functional block diagram showing the configuration of a decoding device 800 that is an example of a decoding device according to the present disclosure.

[0108] The memory 804 stores, for example, the same data as the coded data 703 generated by the coding device 700. The memory 804 reads out the stored data and inputs it as input data 803 to the decoder 802. The input data 803 is, for example, a bitstream to be decoded. The memory 804 may be, for example, a hard disk or an SSD, or may be another storage device.

[0109] Note that the decoding device 800 may not use the data stored in the memory 804 as input data 803 as is, but may convert the read data and generate converted data as input data 803. The data before conversion may be, for example, multiplexed data storing one or more bitstreams. Here, the multiplexed data may be, for example, a file having a file format such as ISOBMFF. The data before conversion may also be in the form of multiple packets generated by dividing the bitstream or file. When converting data different from the bitstream read from the memory 804 into a bitstream, the decoding device 800 may be provided with a conversion unit (not shown), or the conversion process may be performed by a CPU.

[0110] Decoder 802 decodes input data 803 to generate audio signal 801 that is presented to a listener.

[0111] <Another Example of Encoding Device> Fig. 9 is a functional block diagram showing the configuration of an encoding device 900, which is another example of an encoding device of the present disclosure. In Fig. 9, components having the same functions as those in Fig. 7 are assigned the same reference numerals as those in Fig. 7, and descriptions of these components will be omitted.

[0112] The coding device 900 differs from the coding device 700 in that the coding device 700 includes a memory 704 for storing coded data 703, whereas the coding device 900 includes a transmitting unit 901 for transmitting coded data 703 to the outside.

[0113] The transmitter 901 transmits a transmission signal 902 to another device or a server based on the encoded data 703 or data in another data format generated by converting the encoded data 703. The data used to generate the transmission signal 902 is, for example, the bit stream, multiplexed data, file, or packet described in the encoding device 700.

[0114] <Another Example of a Multifunction Apparatus> Fig. 10 is a functional block diagram showing the configuration of a decoding apparatus 1000, which is another example of a decoding apparatus according to the present disclosure. In Fig. 10, components having the same functions as those in Fig. 8 are assigned the same reference numerals as those in Fig. 8, and descriptions of these components will be omitted.

[0115] The decoding device 1000 differs from the decoding device 800 in that the decoding device 800 includes a memory 804 for reading out input data 803, whereas the decoding device 1000 includes a receiving unit 1001 for receiving input data 803 from the outside.

[0116] The receiving unit 1001 receives a received signal 1002, acquires received data, and outputs input data 803 to be input to the decoder 802. The received data may be the same as the input data 803 to be input to the decoder 802, or may be data in a data format different from that of the input data 803. If the received data is data in a data format different from that of the input data 803, the receiving unit 1001 may convert the received data into the input data 803, or a conversion unit or CPU (not shown) included in the decoding device 1000 may convert the received data into the input data 803. The received data is, for example, a bit stream, multiplexed data, a file, or a packet, as described in the encoding device 900.

[0117] <Functional Description of Decoder> FIG. 11 is a functional block diagram showing the configuration of a decoder 1100, which is an example of the decoder 802 in FIG. 8 or 10.

[0118] The input data 803 is an encoded bitstream, and includes encoded audio data, which is an encoded audio signal, and metadata used for acoustic processing.

[0119] The spatial information management unit 1101 acquires metadata included in the input data 803 and analyzes the metadata. The metadata includes information describing elements that act on sounds arranged in a sound space. The spatial information management unit 1101 manages spatial information necessary for acoustic processing obtained by analyzing the metadata and provides the spatial information to the rendering unit 1103. Note that, although the information used for acoustic processing is referred to as spatial information in this disclosure, it may be called by other names. The information used for acoustic processing may be called, for example, sound space information or scene information. Furthermore, when the information used for acoustic processing changes over time, the spatial information input to the rendering unit 1103 may be called a space state, a sound space state, a scene state, or the like.

[0120] Furthermore, spatial information may be managed for each sound space or for each scene. For example, when different rooms are represented as virtual spaces, each room may be managed as a scene of a different sound space, or even if the same space is represented, spatial information may be managed as different scenes depending on the situation being represented. In managing spatial information, an identifier for identifying each piece of spatial information may be assigned. The spatial information data may be included in a bitstream, which is one form of input data 803, or the bitstream may include an identifier for the spatial information, and the spatial information data may be acquired from a source other than the bitstream. When the bitstream includes only the identifier for the spatial information, the identifier for the spatial information may be used during rendering to acquire the spatial information data stored in the memory of the acoustic signal processing device or an external server as input data.

[0121] Note that the information managed by the spatial information management unit 1101 is not limited to information included in the bitstream. For example, the input data 803 may include data indicating the characteristics or structure of a space acquired from a software application or server providing VR or AR, as data not included in the bitstream. Furthermore, for example, the input data 803 may include data indicating the characteristics or position of a listener or object, as data not included in the bitstream. Furthermore, the input data 803 may include, as information indicating the position of the listener, information acquired by a sensor provided in a terminal including a decoding device, or information indicating the position of the terminal estimated based on information acquired by the sensor. In other words, the spatial information management unit 1101 may communicate with an external system or server to acquire spatial information and the position of the listener. Furthermore, the spatial information management unit 1101 may acquire clock synchronization information from an external system and execute a process of synchronizing with the clock of the rendering unit 1103. Note that the space in the above description may be a virtually formed space, i.e., a VR space, or may be a real space or a virtual space corresponding to a real space, i.e., an AR space or an MR (Mixed Reality) space. The virtual space may also be called a sound field or a sound space. Furthermore, the information indicating a position in the above description may be information such as coordinate values ​​indicating a position within a space, information indicating a relative position with respect to a predetermined reference position, or information indicating the movement or acceleration of a position within a space.

[0122] The audio data decoder 1102 decodes the encoded audio data included in the input data 803 to obtain an audio signal.

[0123] The encoded audio data acquired by the stereophonic sound reproduction system 600 is a bitstream encoded in a predetermined format such as MPEG-H 3D Audio (ISO / IEC 23008-3). Note that MPEG-H 3D Audio is merely one example of an encoding method that can be used to generate the encoded audio data included in the bitstream, and the encoded audio data may be included in a bitstream encoded in another encoding method. For example, the encoding method used may be a lossy codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), or Vorbis, or a lossless codec such as ALAC (Apple Lossless Audio Codec) or FLAC (Free Lossless Audio Codec), or any other encoding method may be used. For example, PCM (Pulse Code Modulation) data may be a type of encoded audio data. In this case, the decoding process may be, for example, a process of converting an N-bit binary number into a number format (e.g., floating-point format) that can be processed by the rendering unit 1103, where the number of quantization bits of the PCM data is N.

[0124] The rendering unit 1103 receives an audio signal and spatial information as input, performs acoustic processing on the audio signal using the spatial information, and outputs an audio signal 801 after acoustic processing.

[0125] Before starting rendering, the spatial information management unit 1101 reads metadata of the input signal, detects rendering items such as objects or sounds defined in the spatial information, and transmits the detected items to the rendering unit 1103. After starting rendering, the spatial information management unit 1101 grasps changes over time in the spatial information and the position of the listener, and updates and manages the spatial information. The spatial information management unit 1101 then transmits the updated spatial information to the rendering unit 1103. The rendering unit 1103 generates and outputs an audio signal to which acoustic processing has been applied based on the audio signal included in the input data and the spatial information received from the spatial information management unit 1101.

[0126] The spatial information update process and the audio signal output process with added acoustic processing may be executed in the same thread, or the spatial information management unit 1101 and the rendering unit 1103 may be allocated to independent threads. When the spatial information update process and the audio signal output process with added acoustic processing are executed in different threads, the thread startup frequency may be set individually, or the processes may be executed in parallel.

[0127] By having the spatial information management unit 1101 and the rendering unit 1103 execute their processes in different, independent threads, computational resources can be preferentially allocated to the rendering unit 1103. Therefore, in the case of sound output processing in which even a slight delay cannot be tolerated, for example, sound output processing in which a delay of even one sample (0.02 msec) would cause a popping noise, can be safely performed. In this case, the allocation of computational resources to the spatial information management unit 1101 is limited. However, compared to audio signal output processing, updating spatial information is a less frequent process (e.g., a process such as updating the listener's facial orientation). Therefore, unlike audio signal output processing, updating spatial information does not necessarily require an instantaneous response, and therefore limiting the allocation of computational resources does not significantly affect the acoustic quality provided to the listener.

[0128] The space information may be updated periodically at preset times or intervals, or when preset conditions are met. The space information may be updated manually by a listener or a sound space manager, or may be triggered by a change in an external system. For example, when a listener operates a controller to instantly warp the position of their avatar, instantly advance or rewind the time, or when a virtual space manager suddenly changes the environment of the space, the thread in which the space information management unit 1101 is located may be started as a one-off interrupt process in addition to being started periodically.

[0129] The role of the information update thread that executes the spatial information update process is, for example, to update the position or orientation of the listener's avatar placed in the virtual space based on the position or orientation of the VR goggles worn by the listener, and to update the position of objects moving in the virtual space. These tasks are handled within a processing thread that runs relatively infrequently, on the order of several tens of Hz. Processing to reflect the properties of direct sound may be performed in such an infrequently occurring processing thread. This is because the properties of direct sound change less frequently than the frequency of audio processing frames for audio output. Doing so can relatively reduce the computational load of the process and also avoid the risk of pulsive noise occurring when information is updated at an unnecessarily fast frequency.

[0130] FIG. 12 is a functional block diagram showing the configuration of a decoder 1200, which is another example of the decoder 802 in FIG. 8 or 10.

[0131] Figure 12 differs from Figure 11 in that the input data 803 includes an unencoded audio signal rather than encoded audio data. The input data 803 includes a bitstream including metadata and an audio signal.

[0132] The spatial information management unit 1201 is the same as the spatial information management unit 1101 in FIG. 11, and therefore a description thereof will be omitted.

[0133] The rendering unit 1202 is the same as the rendering unit 1103 in FIG. 11, and therefore a description thereof will be omitted.

[0134] In the above description, the configuration in Fig. 12 is called a decoder, but it may also be called an audio processing unit that performs audio processing. Furthermore, a device including an audio processing unit may also be called an audio processing device rather than a decoding device. Furthermore, the audio signal processing device (information processing device 601) may also be called an audio processing device.

[0135] <Physical Configuration of Encoding Apparatus> Fig. 13 is a diagram showing an example of the physical configuration of an encoding apparatus. The encoding apparatus shown in Fig. 13 is an example of the encoding apparatuses 700 and 900 described above.

[0136] The encoding device of FIG. 13 includes a processor, a memory, and a communication IF.

[0137] The processor may be, for example, a CPU (Central Processing Unit), a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit), and the CPU, DSP, or GPU may execute a program stored in a memory to perform the encoding process of the present disclosure. Alternatively, the processor may be a dedicated circuit that performs signal processing on an audio signal, including the encoding process of the present disclosure.

[0138] The memory may be configured, for example, by RAM (Random Access Memory) or ROM (Read Only Memory). The memory may also include a magnetic storage medium such as a hard disk or a semiconductor memory such as an SSD (Solid State Drive). The term "memory" may also refer to an internal memory built into a CPU or GPU.

[0139] The communication IF (Interface) is a communication module compatible with a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The encoding device has a function of communicating with other communication devices via the communication IF and transmits an encoded bitstream.

[0140] The communication module is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. In the above example, Bluetooth (registered trademark) or WIGIG (registered trademark) is used as the communication method, but other communication methods such as LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark) may also be supported. Furthermore, the communication IF may be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface) instead of the wireless communication method described above.

[0141] <Physical configuration of audio signal processing device> Fig. 14 is a diagram showing an example of the physical configuration of an audio signal processing device. Note that the audio signal processing device in Fig. 14 may be a decoding device. Furthermore, part of the configuration described here may be provided in the audio presentation device 602. Furthermore, the audio signal processing device shown in Fig. 14 is an example of the audio signal processing device 601 described above.

[0142] The acoustic signal processing device of FIG. 14 includes a processor, a memory, a communication IF, a sensor, and a speaker.

[0143] The processor may be, for example, a CPU (Central Processing Unit), a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit), and the CPU, DSP, or GPU may execute a program stored in a memory to perform the audio processing or decoding processing of the present disclosure. Alternatively, the processor may be a dedicated circuit that performs signal processing on audio signals, including the audio processing of the present disclosure.

[0144] The memory may be configured, for example, by RAM (Random Access Memory) or ROM (Read Only Memory). The memory may also include a magnetic storage medium such as a hard disk or a semiconductor memory such as an SSD (Solid State Drive). The term "memory" may also refer to an internal memory built into a CPU or GPU.

[0145] The communication IF (Interface) is a communication module compatible with a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The acoustic signal processing device shown in FIG. 2I has a function of communicating with other communication devices via the communication IF and acquires a bitstream to be decoded. The acquired bitstream is stored in a memory, for example.

[0146] The communication module is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. In the above example, Bluetooth (registered trademark) or WIGIG (registered trademark) is used as the communication method, but other communication methods such as LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark) may also be supported. Furthermore, the communication IF may be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface) instead of the wireless communication method described above.

[0147] The sensor performs sensing to estimate the position or orientation of the listener. Specifically, the sensor estimates the position and / or orientation of the listener based on one or more detection results of the position, orientation, movement, velocity, angular velocity, or acceleration of a part or the entire body of the listener, such as the head, and generates position information indicating the position and / or orientation of the listener. Note that the position information may be information indicating the position and / or orientation of the listener in real space, or information indicating a displacement of the position and / or orientation of the listener based on the position and / or orientation of the listener at a predetermined time. Furthermore, the position information may be information indicating the position and / or orientation relative to the stereophonic sound reproduction system or an external device equipped with the sensor.

[0148] The sensor may be, for example, an imaging device such as a camera or a ranging device such as LiDAR (Light Detection and Ranging), and may capture an image of the listener's head movement and detect the movement of the listener's head by processing the captured image. Alternatively, the sensor may be a device that performs position estimation using wireless signals of any frequency band, such as millimeter waves.

[0149] The audio signal processing device shown in Fig. 14 may acquire position information from an external device equipped with a sensor via a communication IF. In this case, the audio signal processing device does not need to include a sensor. Here, the external device is, for example, the audio presentation device 602 described in Fig. 6 or a 3D video playback device worn on the listener's head. In this case, the sensor is configured by combining various sensors such as a gyro sensor and an acceleration sensor.

[0150] The sensor may, for example, detect the angular velocity of rotation around at least one of three mutually perpendicular axes in the sound space as the axis of rotation as the speed of movement of the listener's head, or may detect the acceleration of displacement with at least one of the three axes as the direction of displacement.

[0151] For example, the sensor may detect the amount of rotation about at least one of three mutually orthogonal axes in the sound space as the rotation axis, or the amount of displacement about at least one of the three axes as the displacement direction, as the amount of movement of the listener's head. Specifically, the sensor detects 6 DoF (position (x, y, z) and angle (yaw, pitch, roll)) as the position of the listener. The sensor is configured by combining various sensors used for detecting movement, such as a gyro sensor and an acceleration sensor.

[0152] The sensor may be any device capable of detecting the position of the listener, such as a camera or a GPS (Global Positioning System) receiver. Alternatively, the sensor may use location information obtained by performing self-position estimation using LiDAR (Laser Imaging Detection and Ranging). For example, when the audio signal reproduction system is implemented by a smartphone, the sensor is built into the smartphone.

[0153] The sensor may also include a temperature sensor such as a thermocouple that detects the temperature of the acoustic signal processing device shown in Figure 14, and a sensor that detects the remaining charge of a battery provided in or connected to the acoustic signal processing device.

[0154] A speaker has, for example, a diaphragm, a drive mechanism such as a magnet or a voice coil, and an amplifier, and presents an audio signal after acoustic processing to a listener as sound. The speaker operates the drive mechanism in response to the audio signal (more specifically, a waveform signal indicating the waveform of the sound) amplified by the amplifier, and the drive mechanism vibrates the diaphragm. In this way, the diaphragm vibrating in response to the audio signal generates sound waves, which propagate through the air and reach the listener's ears, causing the listener to perceive the sound.

[0155] Note that, although the description has been given here of an example in which the acoustic signal processing device shown in FIG. 14 includes a speaker and presents an audio signal after acoustic processing via the speaker, the means for presenting the audio signal is not limited to the above configuration. For example, the audio signal after acoustic processing may be output to an external audio presentation device 602 connected via a communication module. Communication via the communication module may be wired or wireless. As another example, the acoustic signal processing device shown in FIG. 14 may include a terminal for outputting an analog audio signal, and a cable such as earphones may be connected to the terminal to present the audio signal from the earphones. In the above case, the audio signal may be reproduced by headphones, earphones, a head-mounted display, a neck speaker, a wearable speaker, a surround speaker composed of multiple fixed speakers, or the like, which are worn on the head or part of the body of the listener, which is the audio presentation device 602.

[0156] <Functional Description of Rendering Unit> FIG. 15 is a functional block diagram showing an example of the detailed configuration of the rendering units 1103 and 1202 in FIGS.

[0157] The rendering section is composed of an analysis section, a panning section, and a synthesis section (different from the synthesis section 135), and applies acoustic processing to the sound data contained in the input signal and outputs the result.

[0158] The information contained in the input signal will now be described.

[0159] The input signal may be composed of, for example, spatial information, sensor information, and sound data. The input signal may also include a bitstream composed of sound data and metadata (control information), in which case the metadata may include spatial information.

[0160] Spatial information is information about the sound space (three-dimensional sound field) created by the stereophonic playback system, and is composed of information about the objects included in the sound space and information about the listener. Objects include sound source objects that emit sound and act as sound sources, and non-sound-emitting objects that do not emit sound. Non-sound-emitting objects function as obstacle objects that reflect sounds emitted by sound source objects, but sound source objects may also function as obstacle objects that reflect sounds emitted by other sound source objects.

[0161] Information commonly assigned to sound source objects and non-sound generating objects includes position information, shape information, and the rate of attenuation of the volume when the object reflects sound.

[0162] The position information is expressed as coordinate values ​​on three axes, for example, the X-axis, Y-axis, and Z-axis, in Euclidean space, but it does not necessarily have to be three-dimensional information. For example, it may be two-dimensional information expressed as coordinate values ​​on two axes, the X-axis and the Y-axis. The position information of an object is determined by a representative position of a shape expressed by a mesh or voxels.

[0163] The shape information may include information about the surface material.

[0164] The information may also include information indicating whether the object belongs to a living thing, information indicating whether the object is a moving object, etc. If the object is a moving object, the position information may change over time, and the changed position information or the amount of change is transmitted to the rendering unit.

[0165] The information about the sound source object includes the information commonly given to the sound source object and the non-sound generating object, as well as sound data and information required to radiate the sound data into the sound space.

[0166] The sound data is data that represents the sound perceived by a listener, including information about the frequency and intensity of the sound. The sound data is typically a PCM signal, but may also be data compressed using an encoding method such as MP3. In this case, the signal must be decoded at least before it reaches the synthesis unit, so the rendering unit may include a decoding unit (not shown). Alternatively, the signal may be decoded by the audio data decoder 1102.

[0167] At least one piece of sound data may be set for one sound source object, but multiple pieces of sound data may also be set. Furthermore, identification information for identifying each piece of sound data may be assigned, and the identification information for the sound data may be stored as information about the sound source object.

[0168] The information necessary for radiating sound data into a sound space may include, for example, information on a reference volume that serves as a reference when playing back sound data, information indicating the properties (also called characteristics) of the sound data, information on the position of the sound source object, information on the orientation of the sound source object, information on the directivity of the sound emitted by the sound source object, etc. The information on the reference volume is, for example, the effective value of the amplitude value of the sound data at the sound source position when radiating the sound data into a sound space, and may be expressed as a floating-point decibel (dB) value.

[0169] For example, if the reference volume is 0 dB, it may indicate that sound is to be emitted into the sound space from the position indicated by the information about the position at the same volume as the signal level indicated by the sound data, without increasing or decreasing the volume, or if it is -6 dB, it may indicate that sound is to be emitted into the sound space from the position indicated by the information about the position with the volume of the signal level indicated by the sound data reduced to about half. These pieces of information are assigned to one piece of sound data or to multiple pieces of sound data collectively.

[0170] The information indicating the properties of the sound data may be, for example, information regarding the volume of the sound source, and may be information indicating time-series fluctuations. For example, if the sound space is a virtual conference room and the sound source is a speaker, the volume will transition intermittently over a short period of time. To put it more simply, this can be said to be alternating between sound and silence.

[0171] If the sound space is a concert hall and the sound source is a performer, the volume is maintained for a certain period of time. If the sound space is a battlefield and the sound source is an explosive, the volume of the explosion increases for a moment and then remains silent. In this way, the volume information of the sound source includes not only information about the volume of the sound but also information about the transition of the volume of the sound, and such information may be used as information indicating the properties of the sound data.

[0172] Here, the information on the transition in loudness of a sound may be data showing frequency characteristics in a time series. It may be data showing the duration of a sound section. It may be data showing a time series of the duration of a sound section and the duration of a silent section. It may be data listing multiple sets of data on the duration for which the amplitude of a sound signal can be considered steady (considered to be roughly constant) and the amplitude values ​​of the signal during that time in a time series. It may be data on the duration for which the frequency characteristics of a sound signal can be considered steady. It may be data listing multiple sets of data on the duration for which the frequency characteristics of a sound signal can be considered steady and the frequency characteristics during that time in a time series.

[0173] The data format may be, for example, data indicating the outline of a spectrogram. Furthermore, the volume that serves as a reference for the frequency characteristics may be used as the reference volume. Information on the reference volume and information indicating the properties of the sound data may be used to calculate the volume of the direct sound or reflected sound to be perceived by the listener, as well as in a selection process for selecting whether or not to perceive the direct sound or reflected sound. Other examples of information indicating the properties of the sound data and specific uses for the selection process will be described later.

[0174] Orientation information is typically expressed using yaw, pitch, and roll. Alternatively, the roll rotation may be omitted and the information may be expressed using azimuth (yaw) and elevation (pitch). Orientation information may change over time, and if it does change, it is transmitted to the rendering unit.

[0175] The information about the listener is information about the position and orientation of the listener in sound space. The position information is expressed as a position on the XYZ axes in Euclidean space, but it does not necessarily have to be three-dimensional information and may be two-dimensional information. The information about orientation is typically expressed using yaw, pitch, and roll. Alternatively, the roll rotation may be omitted and the information may be expressed using azimuth (yaw) and elevation (pitch). The position information and orientation information may change over time, and if they change, they are transmitted to the rendering unit.

[0176] The sensor information includes the amount of rotation or displacement detected by a sensor worn by the listener and the position and orientation of the listener. The sensor information is transmitted to a rendering unit, which updates the information on the position and orientation of the listener based on the sensor information. The sensor information may be, for example, position information obtained by a mobile terminal performing self-position estimation using a GPS, a camera, or LiDAR (Laser Imaging Detection and Ranging). Furthermore, information obtained from an external source other than a sensor via a communication module may be detected as sensor information. Information indicating the temperature of the audio signal processing device and information indicating the remaining battery level may be obtained from the sensor. Computing resources (CPU capacity, memory resources, PC performance) of the audio signal processing device and the audio signal presentation device may be obtained in real time.

[0177] The analysis unit performs the same function as the acquisition unit 111 in the above example. That is, it analyzes the input signal and acquires information required by the path calculation unit 121 and the output sound generation unit 131.

[0178] The synthesis unit performs the same functions as the output sound generation unit 131 and the signal output unit 141 in the above example. Based on the audio signal of the direct sound and information on the direct sound arrival time and volume at the time of direct sound arrival calculated by the analysis unit, the synthesis unit processes the input audio signal to generate direct sound. Also, based on information on the reflected sound arrival time and volume at the time of reflected sound arrival calculated by the analysis unit, the synthesis unit processes the input audio signal to generate reflected sound. The synthesis unit synthesizes the generated direct sound and reflected sound and outputs the synthesized sound.

[0179] The panning unit performs the same function as the generation unit 134 in the above example. That is, based on the sound source directions of the multiple sound sources (target signals) acquired by the analysis unit, panning by sound from a specific representative direction is performed by time shifting the sound source and adjusting the gain, thereby performing panning to represent the sound source. The processing performed by the above-mentioned panning unit may be executed as part of pipeline processing such as that described in International Publication No. 2021 / 180938, for example.

[0180] FIG. 16 is a block diagram showing an example of the configuration for the rendering unit 1300 to perform pipeline processing.

[0181] The rendering unit 1300 in Fig. 16 includes a reverberation processing unit 1311, an early reflection processing unit 1312, a distance attenuation processing unit 1313, a selection unit 1314, a generation unit 1315, and a binaural processing unit 1316. These multiple components may be configured from the multiple components of the rendering unit shown in Fig. 15, or may be configured from at least some of the multiple components of the acoustic signal processing device shown in Fig. 14.

[0182] Pipeline processing refers to dividing the process for applying sound effects into multiple processes and executing the multiple processes one by one in sequence. Each of the multiple processes performs, for example, signal processing on an audio signal or generation of parameters used in the signal processing.

[0183] The rendering unit 1300 may perform reverberation processing, early reflection processing, distance attenuation processing, binaural processing, and the like as pipeline processing. However, these processes are merely examples, and the pipeline processing may include other processes or may not include some of the processes. For example, the pipeline processing may include diffraction processing and occlusion processing. Furthermore, for example, reverberation processing may be omitted if it is not necessary. Furthermore, not all sounds may be processed in the binaural processing stage.

[0184] Each process may be expressed as a stage. Furthermore, audio signals such as reflected sound generated as a result of each process may be expressed as a rendering item, render item, RI (Render Item), or simply as a playback sound, sound source, or item. Information indicating the type of each render item and the position from which the sound is emitted may be linked to information indicating the audio signal of the render item. Types of render items include, for example, direct sound, reflected sound, reverberation sound, diffracted sound, and other environmental sounds. These examples are merely examples, and other types of render items may be included. The multiple stages in the pipeline processing and their order are not limited to the example shown in FIG. 16 . For example, the processing of the panning unit may be performed in a binaural processing stage, which is one of the multiple stages included in the pipeline processing. The binaural processing unit performs functions equivalent to those of the synthesis unit described above.

[0185] In the stereophonic sound reproduction system 600 described above, in order to make the user 99 perceive as if they are moving their head within a three-dimensional sound field by changing the sound presented in accordance with the movement of the user 99's head, as explained in the example of the sound reproduction system 100 above, it is necessary to detect the position and orientation of the user 99's head (orientation relative to the position of the sound source object).

[0186] At this time, it is necessary to obtain the detection results of the position and orientation of the head of the user 99 and perform information processing accordingly, and therefore, ideally, all of the components would be built into a device related to the final output portion, such as a headphone, i.e., the audio presentation device 602 in Fig. 6, as shown in Figures 1 and 2. However, due to constraints such as power supply availability, information processing performance, and housing size and weight, it is necessary to divide the processing so that the information processing portion is performed in the information processing device 601 and the audio output is performed in the audio presentation device 602. Furthermore, in recent years, it has become desirable to connect the audio presentation device 602 to the information processing device 601 via wireless communication, and therefore, when transmitting and receiving information between the audio presentation device 602 and the information processing device 601 via such wireless communication, limitations on the amount of information that can be transmitted and received simultaneously (i.e., limitations on the communication bandwidth) become a bottleneck.

[0187] For example, if the information processing device 601 outputs an output sound signal, the output sound signal must be transmitted to the audio presentation device 602 via wireless communication. This creates delays due to encoding and decoding processes conforming to wireless communication standards, as well as transmission delays due to the large amount of information in the output sound signal itself, which contains a three-dimensional audio signal. This potentially detracts from the user's experience. Furthermore, the information processing device 601 must first acquire the detection results of the user's head position and orientation before generating the output sound signal, making it difficult to instantaneously track the user's head movement. Therefore, the movement of a sound source object specified in the content is processed in the information processing device 601, generating a three-dimensional audio signal, which is then panned to compress the amount of information. The compressed, panned audio signal is then transmitted to the audio presentation device 602, thereby avoiding limitations on the transmission path bandwidth. Because the processing up to this point is a predetermined part of the content, the information can be transmitted to the audio presentation device 602 in advance and buffered.

[0188] Then, the movement of the user's 99's head is detected by the audio presentation device 602, and the detection result is used to convolve the head-related transfer function corresponding to the representative direction (after the head has moved) according to the detection result into the audio signal that has been panned on the audio presentation device 602, thereby making it possible to instantaneously track the direction of arrival of the sound to the movement of the user's head. At this time, by performing the panning process in advance in the information processing device 601, only the process of convolving the head-related transfer functions for a not-so-large number of representative directions is executed in the audio presentation device 602, which makes it difficult to increase the processing resources required on the audio presentation device 602 side, and significantly reduces the delay due to information processing on the audio presentation device 602. Because the representative direction associated with head movement is updated on the audio presentation device 602 side, no communication is required between the information processing device 601 and the audio presentation device 602 after head movement is detected and before the head-related transfer function is updated, making it possible to minimize the time required to update the direction of arrival of the sound.

[0189] In this way, in a stereophonic sound reproduction system 600 divided into an information processing device 601 and an audio presentation device 602, by making the audio signals transmitted and received between the information processing device 601 and the audio presentation device 602 audio signals that have been panned, it is possible to split the information processing between the two devices while still allowing the direction of sound arrival to instantly follow the movement of the user's 99 head.

[0190] When the information processing device 601 is the first terminal and the audio presentation device 602 is the second terminal, the above configuration allows (1) the first terminal, which is not required to be worn by the user 99 and therefore has relatively high processing performance compared to the second terminal, to perform panning processing on the audio signal, including the movement of a sound source object specified in the content, which requires relatively high processing resources. Furthermore, (2) the panning processing compresses the amount of information by consolidating sounds from, for example, 10 or fewer representative directions, reducing the communication bandwidth constraints when transmitting and receiving between the first terminal and the second terminal. Furthermore, (3) the second terminal worn by the user 99 (on the head where the ears are located) can easily detect the position and orientation of the user 99's head and use this information to output an output sound signal from the panned audio signal, thereby instantly updating the sound arrival direction in response to head movement. Furthermore, (4) the output sound signal can be output by simply convolving head-related transfer functions for a small number of representative directions when panning is performed, thereby reducing the processing resources required for the second terminal. The above benefits can be obtained.

[0191] In the configurations shown in FIGS. 2 to 5, the first terminal includes the communication module 102, some functions of the acquisition unit 111, and functions other than the convolution processing of the path calculation unit 121 and the output sound generation unit 131 shown in FIG. 2. The path calculation unit 121 calculates the relative arrival direction of the sound source object from the position of the reference position using a reference position, i.e., the coordinate position or orientation of the user 99 already acquired by the first terminal before transmitting sound information to the second terminal, the coordinate position or orientation of the user 99 at the time of initialization of the system (the first terminal and the second terminal), or a predetermined coordinate position or direction determined in advance by the system. The reference position may be determined in any manner as long as it is the same coordinate position or orientation shared by the first terminal and the second terminal. Therefore, the orientation as the reference position may be determined as a specific direction such as "north" in absolute coordinates, or as an average direction calculated from the direction in which the user 99 faced the user 99 during a predetermined period of time (e.g., one minute) in the past. In the latter case, the orientation as the reference position is updated every predetermined period. In this way, the first terminal and the second terminal share the same coordinate position or orientation using the reference position, and then perform a series of rendering processes, so long as the shared reference position does not change during the series of rendering processes. When updating the reference position, the update may be performed by including it in the information update thread in the spatial information update process mentioned in <Description of Decoder Function>, or by including it in a thread for updating other information, or by using a dedicated thread for updating the reference position.

[0192] 2 , the functions of other parts of the acquisition unit 111, the function of convolution processing of head-related transfer functions in the output sound generation unit 131, the signal output unit 141, the database 105, and the driver 104. The signal output unit 141 converts the direction to set the reference position to the determined coordinates and orientation of the user 99, and outputs an output sound signal according to the detection result by the detector 103.

[0193] In the configurations of Figures 6 to 15, the first terminal is equipped with the analysis unit and panning unit shown in Figure 15. The analysis unit calculates the relative arrival direction from the position of the sound source object to the reference position using a reference position based on first sound information, which is an input signal. The first sound information includes sound source object position information and an audio signal. The analysis unit has functions equivalent to the path calculation unit 121 described above. The panning unit determines a representative direction from a representative point to the reference position based on the arrival direction calculated by the analysis unit, and performs panning processing to distribute the audio signal to the representative point (representative direction). In other words, by performing panning processing, the first sound information is converted into second sound information including position information of the representative point and a panned audio signal. The second terminal is equipped with a synthesis unit. The synthesis unit converts the direction to match the reference position with the coordinates and orientation of the user 99 detected by the second terminal, and performs convolution processing of the head-related transfer function based on the coordinates and orientation of the user 99.

[0194] [Specific Example of Panning Processing] To reiterate, in panning processing, reproduced sounds from multiple sound source objects are represented by representative sounds from multiple representative directions. For example, two or three directions can be used as these representative directions. Specifically, in panning processing, the number of representative points is reduced to a number less than the number of sound source objects, and the reproduced sounds can be perceived as sounds coming from the direction of arrival using only the head-related transfer functions of the representative directions for these representative points.

[0195] In this case, the panning process calculates a time shift (delay) that maximizes the cross-correlation between the head-related transfer function in the direction of arrival from the sound source object and the head-related transfer function in the representative direction. The time shift obtained here, or a time shift obtained by adding a negative sign to this time shift, is applied to the reproduced sound of the sound source object, and the subsequent processing is performed assuming that the signal after the time shift is in the representative direction. Hereinafter, the panning process will be described using direct sound emitted from the sound source object as an example of reproduced sound, but it goes without saying that panning process can be applied even if the reproduced sound is not direct sound. For example, it may be direct sound, reflected sound, reverberation sound, diffracted sound, other environmental sound, etc. As described above, the reproduced sound here can be interpreted as a rendering item, render item, RI (Render Item), or simply as a sound source, item, etc.

[0196] This time shift may be a time shift shorter than the sampling frequency (a shift in which the sample position is expressed by a decimal number; hereinafter, referred to as a "decimal shift"). This decimal shift can be performed by oversampling.

[0197] Here, in the panning process, a gain is applied to the signal of a representative direction obtained by time-shifting the reproduced sound of the sound source object, and the sum of these values ​​calculated for each representative point is calculated, and the sum is convolved with the head-related transfer function at each representative point, thereby synthesizing a signal equivalent to the reproduced sound of the sound source object convolved with the head-related transfer function of the arrival direction.

[0198] On the other hand, in the panning process, when synthesizing the head-related transfer function (vector) of the arrival direction by the sum of the head-related transfer functions (vector) of the representative direction, the gain may be calculated so that the error signal vector between the synthesized head-related transfer function (vector) and the head-related transfer function (vector) of the arrival direction is orthogonal to the head-related transfer function (vector) of the representative direction. Note that the head-related transfer function (vector) is a time waveform of a head impulse response, which is an expression of the head-related transfer function in the time domain, regarded as a vector. Hereinafter, this head-related transfer function (vector) will also be simply referred to as a "head-related transfer function vector."

[0199] In the panning process, this gain is corrected so that the energy balance of the head-related transfer functions from the position of the sound source object to the left and right ears of the user 99 is maintained even in the head-related transfer functions substantially synthesized by the panning process using head-related transfer functions from multiple representative points. In other words, in the panning process, the gain may be corrected so that the energy balance of the head-related transfer functions of the left and right ears of the user 99 due to the sound source object is maintained even in the head-related transfer functions substantially synthesized by the panning process.

[0200] In this embodiment, the panning process can calculate, for each direction of arrival of the sound source object, a gain value to be multiplied by the head transfer function of the representative direction and a time shift value to be applied to the head transfer function of the representative direction, and store them in a head transfer function table, which will be described later.

[0201] Then, in the panning process, each sound source object is time-shifted by a time shift value and a gain value corresponding to the direction of arrival of each sound source object, and the time shifts and gains are multiplied and summed to generate a sum signal. In the panning process, this sum signal is treated as being present at the position of the representative point. In the panning process, the head-related transfer function at the position of the representative point is convolved with this sum signal to generate a signal at the ear of the user 99.

[0202] The panning process and the associated audio playback process will be described in detail below, step by step, with reference to the flowchart in Fig. 17. First, the process of acquiring the positions of the sound source and the user is performed (step S201). For example, the acquisition unit 111 acquires the position of the sound source object and the position of the user. Then, the path calculation unit 121 acquires the direction of the sound source object as seen by the user 99.

[0203] Specifically, the acquisition unit 111 acquires an audio signal (target signal) of the sound source object. This audio signal may have any sampling frequency and any number of quantization bits. In this embodiment, an example will be described in which an audio signal with a sampling frequency of 48 kHz and a number of quantization bits of 16 is used. Furthermore, the path calculation unit 121 acquires directional information of the sound source object added to the audio signal of the content or the audio signals of the participants in the remote call. Furthermore, the acquisition unit 111 acquires the position of the user 99 based on the detection result from the detector 103.

[0204] Then, the path calculation unit 121 grasps the spatial arrangement of the sound source object and the user 99. As described above, this arrangement may be an arrangement within a space including a virtual space set in content or the like. Then, the path calculation unit 121 calculates the direction of the sound source object as seen by the user 99, i.e., the arrival direction, according to the grasped arrangement within the space. Similarly, the path calculation unit 121 can also calculate the arrival direction for the audio signal of the content based on the arrangement of the user 99 by referring to directional information of the audio signal of the sound source object.

[0205] The path calculation unit 121 may also calculate the direction of the user 99 from the sound source object.

[0206] Next, the representative direction determination unit 133 determines a representative direction to be used in the panning process based on the acquired position of the sound source and the position of the user 99 (step S202). The representative direction determination process will be described later.

[0207] Next, the panning unit (the generation unit 134 in the functional block diagram of FIG. 5 ), which is a processing unit that executes the panning process, performs the panning process using the determined representative direction (step S203). Here, the panning unit performs the panning process on the sound source object using the direction information. In this embodiment, the panning unit performs the panning process from the perspective of how closely the sound synthesized at the ear by the panning process can be made to resemble the sound that should be at the ear.

[0208] 18, the calculation performed by the panning unit when panning a sound source object (sound source S-1) using representative points R-1 and R-2 will be described. Fig. 18 is a diagram for explaining the synthesis of head-related transfer functions in audio reproduction processing according to an embodiment. Here, the signal to be panned is sound source S-1, but in the following, to calculate the optimal shift amount and optimal gain for this, calculations are performed using head-related transfer functions from sound source S-1, representative point R-1, and representative point R-2 to the ears.

[0209] 18, the head-related transfer function with P sampling points (number of taps) from the sound source S-1 to the ear is a P-dimensional vector, which is denoted as v{x} (in the following embodiments, a vector will be represented as "v{}").

[0210] Here, the panning unit calculates the head-related transfer function from the representative point R-1 to the ear of the user 99 as v{x 01}, and the head-related transfer function from the representative point R-2 to the ear is v{x 02}. v{x} and v{x 01}, and v{x 01} is time-shifted to v{x1}. Similarly, v{x} and v{x 02}, and v{x 02} is calculated as v{x2} by shifting the time.

[0211] This v{x1} is multiplied by gain A, and v{x2} is multiplied by gain B, and v{x} is approximated by the sum of these. In other words, v{x} is approximated as the approximate value of v{x} = A × v{x1} + B × v{x2}. This makes it possible to achieve panning processing with reduced error.

[0212] The calculation of the gain and the time shift will be described in detail below. First, the calculation of the gain will be described. The error vector obtained by approximating v{x} is expressed by the following equation (1).

[0213]

[0214] In the above formula (1), the arrows on the variables indicate vectors. Here, when A and B are optimally sized, that is, when the magnitude of the error vector is minimized, the error vector v{e} is orthogonal to the plane spanned by the original vectors v{x1} and v{x2}. Therefore, the relationship in the following formula (2) holds.

[0215]

[0216] As a result, the following equation (3) is calculated.

[0217]

[0218] By modifying this equation (3), the following equation (4) is obtained.

[0219]

[0220] For the above equation (4), |v{x2}| 2 When v{x1}·v{x2} is calculated for the equation below, the following equation (5) is obtained.

[0221]

[0222] A can be calculated by subtracting the lower equation from the upper equation of equation (5) and eliminating B. This is shown in equation (6).

[0223]

[0224] Therefore, the gain A is expressed by the following equation (7).

[0225]

[0226] Similarly, by eliminating the gain A, the gain B can be calculated as shown in the following equation (8).

[0227]

[0228] In this way, the gains A and B are determined so that the error vector between the composite signal and the target signal is orthogonal to the representative direction vector used.

[0229] The gains A and B obtained by this calculation are multiplied by the waveform of the head-related transfer function of v{x1} after the time shift due to cross-correlation and the waveform of the head-related transfer function of v{x2}, and it becomes possible to synthesize the head-related transfer function to be output. In other words, these time shift amounts (time shift values) and gains A and B are applied to the sound source S-1 to perform panning processing.

[0230] Next, a specific calculation process for the time shift that maximizes the cross-correlation will be described. 01} handles the head-related transfer function with the number of samples at P points as a vector. Therefore, the subscript of the time (position of the sample point) of the head-related transfer function can be explicitly written as in the following equation (9).

[0231]

[0232] Then, the cross-correlation between the two vectors in equation (9) is defined as a function of "k" as shown in equation (10) below.

[0233]

[0234] where φ xx01 The k that gives the maximum value of (k) is k max01 The panning unit, for example, substitutes each value into k, and max01 Similarly, φ xx02 The k that gives the maximum value of (k) is k max02 The panning section writes this k max02 k max01 This k is calculated in the same way. max01 and k max02 Hereinafter, either of the above will be referred to simply as "k max " should be written.

[0235] The panning unit may, for example, use gains A, B, and k calculated for the arrival direction of each sound source object that differs every 2 degrees around 360 degrees. max01 , k max02 are stored in the head-related transfer function table as gain values ​​and time shift values, respectively, and are used in the output process described below. max01 , kmax02 It is also possible to perform only the audio output process described below using a head-related transfer function table in which the values ​​of (a) and (b) have already been calculated and stored.

[0236] Next, the panning unit and the output unit perform audio output processing (step S204). First, the panning unit obtains a gain value and a time shift value corresponding to the obtained arrival direction from the head-related transfer function table for each sound source object. Then, the panning unit multiplies each sampling point (sample) of the waveform of the sound source object by this gain value.

[0237] At this time, the panning unit may correct the gain so that the energy balance of the left and right ear objects due to the sound source object is maintained in the head-related transfer functions synthesized by the panning process. That is, each gain value may be multiplied by an adjustment coefficient that matches the energy balance between the left and right head-related transfer functions with the original head-related transfer functions.

[0238] Next, the panning unit performs a time shift on the signal multiplied by this gain value.

[0239] The details of this time shift are as follows: 01} element k max A vector v{x1} shifted by a sample is generated by the following procedure.

[0240] First, when the phase is advanced, that is, k max If ≧0, add k to the end of the vector max Only the sample is set to zero, and the length of the vector is maintained. On the other hand, if the phase is delayed, that is, k max If <0, the vector begins with k max Only the samples are set to zero and the length of the vector is maintained, that is, set as in the following equation (11).

[0241]

[0242] In this way, a time-shifted vector v{x1} is generated. The positive or negative polarity of the time shift amount value is reversed depending on which is used as the reference for calculating the cross-correlation. Also, when convolving the head-related transfer function with the sound source signal, attention must be paid to the polarity of the time shift amount.

[0243] The panning unit may perform this time shift by a decimal multiple of the number of taps, rather than by an integer multiple of the number of taps. Alternatively, the time shift may be multiplied by a gain value after the time shift.

[0244] The panning unit treats the signal that has been calculated in this way and that has undergone a gain and time shift as a representative point signal that exists at the position of representative point R. The panning unit then takes the sum of the representative point signals of the sound source objects that are grouped together at representative point R to generate a sum signal. The panning unit then convolves this sum signal with the head-related transfer function at the position of representative point R (head-related transfer function in the representative point direction) to generate a signal at the ear of the user 99.

[0245] The signals generated by the panning unit are reproduced by outputting them to the ears of the user 99. The output may be, for example, a two-channel analog audio signal corresponding to the left and right ears of the user 99.

[0246] This makes it possible to reproduce audio signals corresponding to a virtual sound field as two-channel audio signals through headphones. This completes the audio reproduction process.

[0247] The above configuration can provide the following effects.

[0248] In recent years, when content such as movies, AR, VR, MR, and games is played using VR headphones or HMDs, a rendering technology (binauralization technology) that appropriately describes and plays back the entire 3D sound field has been required. Conventional 3D stereophonic sound (binaural signals) has been generated by individually convolving a head-related transfer function (HRTF) of the corresponding direction of arrival with a plurality of sound source signals. In this way, convolving a HRTF with each individual sound source object has posed a problem in that a huge amount of calculation is required to follow human movement (6 DoF: 6 Degrees of Freedom) with a high sense of realism.

[0249] On the other hand, in the conventional speaker panning process, a sound image is created between speakers by controlling the volume balance of the speakers using the sine law, tangent law, etc. (The sound source object is localized.) However, simply controlling the volume balance is not enough to properly reproduce a stereophonic sound image through headphones.

[0250] In contrast, the above-described audio playback process is characterized by using a path calculation unit 121 that acquires the direction of arrival of a sound source object, and a panning unit that expresses the sound source object by performing panning processing using sound from a specific representative direction based on the direction of arrival acquired by the path calculation unit 121 through time shifting and gain adjustment of the sound source object.

[0251] This configuration enables more efficient and effective rendering by synthesizing sound source objects through panning of representative directions and reducing the number of arrival directions. This reduces the amount of computation compared to conventional methods that individually convolve head-related transfer functions into the signals of each sound source object. That is, the panning unit equivalently synthesizes head-related transfer functions of representative directions that approximate the arrival directions acquired by the path calculation unit 121 through panning processing, thereby generating head-related transfer functions for the arrival directions. By reducing the amount of computation in this way, the system can be applied to VR / AR applications such as games and movies as a 3D sound field playback system. Furthermore, by applying it to smartphones and home appliances, the amount of computation required to generate stereophonic sound can be reduced, thereby reducing costs. Furthermore, as a method with even lower computational complexity, the system can be applied to international standardization, etc.

[0252] In the above-described embodiment, an example has been described in which the panning unit expresses a sound source signal by panning processing using representative points in two directions, left and right, that is, an example in which a vector of a head-related transfer function in the direction of arrival is equivalently synthesized using a vector of a head-related transfer function in the left and right directions. That is, in the above-described embodiment, an example has been described in which the angular directions of the left and right of the user 99 are taken into consideration as direction information.

[0253] However, the vertical direction can also be considered as the direction of arrival. Specifically, it is also possible to equivalently synthesize the vector of the head-related transfer function of the direction of arrival by interpolation using the vectors of the head-related transfer functions in three directions. In other words, the panning unit can also perform panning processing using representative points in three directions including the elevation angle direction (or depression angle direction).

[0254] In this case, similar to the interpolation from two directions, the head related transfer functions in the representative directions are time-shifted so that the cross-correlation with v{x} is maximized, and are expressed as vectors v{x1}, v{x2}, and v{x3}. In this case, the error vector v{e} is expressed by the following equation (12).

[0255]

[0256] This is applied to the following equation (13) and solved.

[0257]

[0258] Specifically, the optimal gains A, B, and C can be calculated by the following equation (14).

[0259]

[0260] In the above equation (14), the "-1" on the right shoulder of the matrix means the inverse matrix. The time shift amount k of the HRIR in the representative direction determined so as to maximize the cross-correlation is max01 , k max02 , k max03 Similarly to the values ​​in the two directions, the values ​​are calculated before the gain values ​​are calculated.

[0261] In the above embodiment, an example in which two to four representative points R are used has been described.

[0262] However, it is of course possible to use two or more representative points R. For example, as shown in the examples described later, it is also possible to use four to six representative points R corresponding to range angles of 90° and 60°, etc. Furthermore, even in the case of four representative points, it is also possible to set the representative points R at different positions, such as diagonally (45°, 135°, 225°, and 315°) or vertically and horizontally (0°, 90°, 180°, and 270°) relative to the user 99. It is also possible to select two or three points from the four to six representative points R that are closest to the direction of arrival and use them as representative points R for synthesizing the sound source.

[0263] That is, in the audio reproduction process, the panning process may use a gain calculated so as to minimize the energy or L2 norm of the error signal vector between the synthesized HRIR vector and the HRIR vector in the sound source direction.

[0264] [Weighting filter when calculating time shift and gain] In the above example, when calculating the time shift and gain that maximize the cross-correlation, the head-related transfer function itself is used. On the other hand, the time shift and / or gain may be subjected to a weighting filter on the frequency axis and then the cross-correlation may be calculated.

[0265] That is, when calculating the time shift and gain that maximize the cross-correlation, it is possible to use a filter that has been subjected to a weighting filter on the frequency axis (hereinafter also referred to as a "frequency weighting filter").

[0266] It is preferable to use a frequency weighting filter that has a cutoff frequency near or slightly higher than the frequency band where human hearing sensitivity is high, and attenuates the higher frequency band, i.e., the frequency band where human hearing sensitivity decreases. For example, it is preferable to use a low-pass filter (LPF) with a cutoff frequency of 3000 Hz to 6000 Hz and approximately 6 dB / oct (octave) to 12 dB / oct.

[0267] Specifically, v{x} and v{x 01} handles the head-related transfer function of point P as a vector, so it is possible to explicitly write the time subscript of the head-related transfer function and write it as in the above equation (9). Here, if the impulse response w of the frequency weighting filter is added to the two vectors in the above equation (9), c (n) is convoluted and truncated to a length P, as shown in the following equation (15).

[0268]

[0269] Here, the operation "*" indicates convolution. Then, the cross-correlation of the two vectors in equation (15) is defined as a function of "k" as shown in equation (16) below.

[0270]

[0271] Here, φ according to equation (16) xx01 The k that gives the maximum value of (k) is k max The panning unit is, for example, a vector v{x 01} element k max A vector v{x1} shifted by one sample is generated in the following procedure, similar to the above equation (11).

[0272] Specifically, when the phase is advanced, that is, k max If ≧0, k maxTo ensure that the vector is a sample, zeros are added to the end of the vector to maintain its length. max If ≧0, the vector v{x1} is v{x1}=(x 01 (0+k max ), x 01 (1+k max ), x 01 (2+k max ), …… x 01 (P-1), ... 0,0,0).

[0273] Also, if the phase is delayed, that is, k max If <0, pad the beginning of the vector with zeros and max Keep the length of the vector to be k samples. max < 0, the vector v{x1} is v{x1} = (0,0,0, ..., x 01 (0), x 01 (1), x 01 (2), ……, x 01 (P-1+k max )) becomes.

[0274] In the above, the vector v{x 01w} is a vector v{x 01}. In this way, it is possible to generate a vector v{x1}, i.e., a cross-correlation can be calculated and used to calculate the time shift, similar to what was described above.

[0275] In the above-described method, when calculating the error (similarity) between the synthesized head-related transfer function and the original head-related transfer function, |v{e}| of the error signal vector (error vector) v{e} is calculated as in the above-described equation (12). 2 A, B, and C that minimize the above were calculated.

[0276] In this case, v{e} may be filtered by a frequency weighting filter. Specifically, when v{e} is waveform data on the time axis, v{e} convolved with the impulse response w(n) of the weighting filter is used as v{e}. w}, then v{e w} is expressed by the following formula (17).

[0277]

[0278] The operator "*" indicates convolution. Here, the operator "*" is used for vectors, but this is a vector representation of the sequence obtained by convolving the sequence representations of the vectors on the left and right of the operator. In other words, v{x} * v{y} is the vector representation of the result of x(n) * y(n). Hereinafter, unless otherwise specified, the operator "*" for vectors will be treated in the same way.

[0279] On this basis, v{e w} into the following equation (18) and solving it, the gains A, B, and C can be calculated.

[0280]

[0281] Or equivalently, v{e} w It is also possible to calculate

[0282]

[0283] Using the time shift and gain thus determined, it becomes possible to distribute (pan) the target signal to a representative direction.

[0284] The target signal to be panned and the head-related transfer function to be convolved may be the same as those described above. That is, the target signal and the head-related transfer function to be convolved do not need to be convolved with a weighting filter.

[0285] By introducing such frequency weighting, it is possible to set the frequency band for approximation with smaller errors (higher accuracy). In particular, since the main energy of music and voice signals is concentrated in the low frequency range, good performance can be obtained by using a weighting filter that weights the low frequency range.

[0286] Furthermore, if the convolution of a weighting filter whose impulse response is w(n) and a vector is expressed as a convolution matrix W in which each row has the impulse response w(n) of the weighting filter time-shifted by one sample, then equation (17) can be transformed into equation (20) below.

[0287]

[0288] Then, in the following formula (21), |v{e}| 2 can be calculated.

[0289]

[0290] Here, W T represents the transpose matrix of W.

[0291] Furthermore, the weighting filter used when calculating the cross-correlation and when calculating the gain may have the same characteristics, or may have different characteristics. If the same weighting filter is used, the weighting filter w may be convoluted with the entire set of original head-related transfer functions, and then the time shift amount and the gain may be calculated by the same process as described above.

[0292] In addition, when the cross-correlation and the optimum gain are calculated by weighting the low frequency band with an LPF as the weighting filter as described above, if the effective band is limited to about 3000 Hz, the decimal shift described above does not need to be performed. In this case, oversampling is also not required.

[0293] In the above-described embodiment, the audio signal is panned and distributed in a plurality of representative directions, and the head-related transfer functions of each representative direction are convolved and expressed. Specifically, the head-related transfer function of the target direction is simulated by the sum of the head-related transfer functions of the representative directions, with the approximate value of v{x} in three directions = A × v{x1} + B × v{x2} + C × v{x3}.

[0294] In such cases, the amplitude characteristics of the high frequencies of the HRTF tend to be lower in level than the original HRTF compared to the low frequencies. This is because even a slight time error caused by a slight shift in the listening point can cause a large phase rotation of the high frequency components of the HRTF, which tends to be canceled out by the addition caused by the panning process.

[0295] In contrast to this, in the audio reproduction process according to this embodiment, the tendency for high frequencies to attenuate may be compensated for by a reproduction high-frequency emphasis filter.

[0296] Specifically, it is possible to compensate for the tendency of high frequencies to attenuate by applying a high-frequency emphasis filter to a signal obtained by performing panning processing and convolving a head-related transfer function in a representative direction. Alternatively, equivalently, the head-related transfer function in the representative direction itself may be subjected to a high-frequency emphasis filter process in advance to emphasize the high frequencies. This high-frequency emphasis filter may be, for example, an impulse response weighting filter that emphasizes the high frequencies by about +1 to +1.5 dB with a turnover frequency of 5000 to 15000 Hz or more.

[0297] In this way, by performing a filter process that emphasizes the high frequencies of the synthesized sound using the panning process, it is possible to further enhance the stereoscopic effect perceived by the listener.

[0298] Even when a decimal shift similar to that described above is performed, mismatches in the high frequency components of head-transmitted signals remain with normal 8 to 16 times oversampling, so a high frequency emphasis filter may be applied.

[0299] In the panning process, the adjustment amounts in the time shift adjustment and the gain adjustment may be determined according to the head-related transfer functions included in the database 105, and the time shift adjustment and the gain adjustment may be applied to the reproduced sound using the determined adjustment amounts to convert it into a representative sound. Since the optimal values ​​of the adjustment amounts in the time shift adjustment and the gain adjustment used in the panning process change according to the head-related transfer functions, first, when the head-related transfer functions included in the database 105 are read out, the adjustment amounts in the time shift adjustment and the gain adjustment corresponding to the read out head-related transfer functions can be determined, and thereafter, the same adjustment amounts can be reused as long as these head-related transfer functions are used, which is advantageous in terms of the amount of processing.

[0300] The head-related transfer function table is an example of table data including head-related transfer functions stored in the database 105. The head-related transfer function table stores the head-related transfer functions together with the adjustment amounts in the time shift adjustment and the gain adjustment determined according to the head-related transfer functions, which are linked to each other. That is, the head-related transfer function table may be constructed by calculating the adjustment amounts in the time shift adjustment and the gain adjustment in advance for each head-related transfer function included in the database 105. In this way, table data of the head-related transfer function table linking each head-related transfer function with the adjustment amount may be stored in the database 105. In this way, the database 105 is an example of a storage unit. The calculation of the adjustment amount for each head-related transfer function may be performed by the generation unit 134 or the decoding processing unit 113. Alternatively, the calculation of the adjustment amount may be performed by an external device and stored in the memory of the external device. In this case, the memory of the external device corresponds to an example of a storage unit.

[0301] Furthermore, adjustment amounts in the time shift adjustment and the gain adjustment may be calculated in advance, and an adjustment amount table linked to each of a plurality of representative directions may be constructed and stored in the database 105. The adjustment amount table may include table data linking the head related transfer functions of each of the plurality of representative directions with the adjustment amounts in the time shift adjustment and the gain adjustment, or the head related transfer functions of each of the plurality of representative directions may be extracted from head related transfer functions of the entire celestial sphere (multiple directions) that are acquired in advance and stored in the database 105 at the time of rendering or at the time of system initialization.

[0302] Furthermore, the adjustment amount table may be a table including, for example, information as to which of a plurality of representative directions the signal is to be distributed to, for a sound signal arriving at the position of the listener from the direction of each head-related transfer function in the spherical head-related transfer function database, and information on a time shift adjustment amount and a gain adjustment amount to be multiplied by the sound signal for each representative direction when distributing.

[0303] When performing the convolution process of the head-related transfer function, the adjustment amount table stored in the database 105 is referenced, and the adjustment amounts for the time shift adjustment and gain adjustment linked to the head-related transfer function of the direction to be applied are used, which eliminates the need to calculate the adjustment amount for each convolution process and contributes to reducing the amount of processing.

[0304] Note that the embodiment of the present invention can also be applied to new head-related transfer functions that are not included in the database 105. When decoding a sound signal, when powering on the sound reproduction system 100, or when initializing the sound reproduction system 100, the head-related transfer functions of the entire three-dimensional sound field may be newly read, and the adjustment amount for each head-related transfer function may be calculated using the method disclosed in this embodiment or another method. In this case, table data linking the head-related transfer functions with the adjustment amounts may be stored in the database 105. Alternatively, the adjustment amounts may be calculated by an external device and stored in the memory of the external device. When performing convolution processing of head-related transfer functions, by referencing the adjustment amounts for the time shift adjustment and gain adjustment linked to the applied head-related transfer function, it is not necessary to calculate the adjustment amount for each convolution processing, which can contribute to reducing the amount of processing.

[0305] In this way, when a new head-related transfer function that is not stored in the database 105 is read, adjustment amounts in the time shift adjustment and gain adjustment used in the panning process may be determined for the new head-related transfer function before storing it in the database 105, and a head-related transfer function table may be constructed by linking the new head-related transfer function with the determined adjustment amounts, and the head-related transfer function table may be stored in the database 105. Then, when performing the panning process, these adjustment amounts are read from the database 105, and the shift adjustment and gain adjustment are applied based on these adjustment amounts. Note that the new head-related transfer function may be one that was previously stored in the database 105, but was temporarily removed from the database 105 when decoding a sound signal, when powering on the sound reproduction system 100, when initializing the sound reproduction system 100, or the like, and then re-stored in the database 105. Instead of the generation unit 134, a second generation unit may be provided which applies time shift adjustment and gain adjustment to the reproduced sound using adjustment amounts linked to a new head-related transfer function stored in the database to convert it into a representative sound, and generates an output sound signal by convolving a head-related transfer function corresponding to a representative direction from each position of the representative point toward the user's position into the representative sound.

[0306] [Representative Direction Determination Process] The representative direction determination process will be conceptually described below with reference to FIGS. 19 to 22. FIGS. 19 and 20 are diagrams for explaining the arrangement of representative directions in this embodiment. FIGS. 21 and 22 are diagrams for explaining the determination of representative directions according to this embodiment. FIG. 19 shows ideal representative directions R1, R2, R3, R4, R5, and R6 arranged at 30°, 90°, 150°, 210°, 270°, and 330° counterclockwise as viewed from above (the user's zenith side), so as to divide 360° horizontally surrounding the user 99 into six equal parts. For example, if a sound source object is located in area A1, an output sound signal is output using representative directions R1 and R2. That is, the reproduced sound emitted by a sound source object in area A1 is distributed to representative directions R1 and R2, and a 30° head-related transfer function is convolved with the sound distributed to representative direction R1, and a 90° head-related transfer function is convolved with the sound distributed to representative direction R2. Conversely, the reproduced sound of a sound source object within a 120° range between 330° and 90° is distributed to representative direction R1.

[0307] Although the above description has been given assuming that the ideal representative directions R1, R2, R3, R4, R5, and R6 are disposed at positions of 30°, 90°, 150°, 210°, 270°, and 330° counterclockwise when viewed from above (the user's zenith side), they can also be said to be disposed at positions of 30°, 90°, 150°, 210°, 270°, and 330° clockwise when viewed from below (the user's nadir side). Similarly, the ideal representative directions R6, R5, R4, R3, R2, and R1 can also be said to be disposed at positions of 30°, 90°, 150°, 210°, 270°, and 330° clockwise when viewed from above, or at positions of 30°, 90°, 150°, 210°, 270°, and 330° counterclockwise when viewed from below.

[0308] In this way, by dividing 360° into six equal parts, the area is divided into areas A1, A2, A3, A4, A5, and A6, and a representative direction for distributing the reproduced sound in each area is determined. The ideal representative direction is a representative direction arranged so as to divide 360° in such a horizontal plane into six equal parts. Then, as shown in FIG. 20 , in addition to these horizontal directions, an ideal representative direction R7 (elevation angle 90°) corresponding to the zenith in the median plane of the user 99 and an ideal representative direction R8 (elevation angle −90°) corresponding to the nadir in the median plane of the user 99 can be used to configure representative directions that can almost cover the three-dimensional sound field.

[0309] On the other hand, a head-related transfer function data set that can actually be used may not include head-related transfer functions corresponding to these ideal representative directions. Therefore, for example, as shown in FIG. 21 (a plan view of a horizontal plane), it may be necessary to select a representative direction to be used from representative direction candidates R5a and R5b that have the smallest deviation from the ideal representative direction R5 in the horizontal plane, instead of the ideal representative direction R5. Here, the representative direction candidates R5a and R5b have the same deviation from the ideal representative direction R5. Here, "same" means that the deviation is about the same, allowing an error of less than 10% in terms of deviation from the ideal representative direction R5, and also includes cases where the magnitude of the deviation does not exactly match.

[0310] Representative direction candidate R5a is closer to the front direction of the user 99 than representative direction candidate R5b. Human directional resolution for sound tends to be more accurate on the front side than on the rear side, so when representative direction candidates with similar deviations compete with each other, representative direction candidate R5a, which is closer to the front direction of the user 99, can be determined as the representative direction to be used.

[0311] Furthermore, as shown in FIG. 22 (a plan view of the median plane), instead of the ideal representative directions R7 and R8 within the median plane, it may be necessary to select a representative direction candidate that has the smallest deviation from the ideal representative directions R7 and R8. In particular, the representative directions corresponding to the zenith and nadir affect the distribution of all reproduced sounds in the elevation and depression angle directions, so accurate directional resolution of the sound is required. Therefore, the representative direction selected for the ideal representative directions R7 and R8 should be selected from directions within the range of the front side of the user 99 (elevation angle -90° to 90°).

[0312] In other words, representative direction candidates R7c and R8c are not selected no matter how small the deviation from the ideal representative directions R7 and R8 is. Then, among the directions (for example, representative direction candidates R7a and R7b) within the range in front of the user 99 (elevation angle 0 to 90°), representative direction candidate R7b, which is closer to the zenith (elevation angle 90°), is selected. Similarly, among the directions (for example, representative direction candidates R8a and R8b) within the range in front of the user 99 (elevation angle -90 to 0°), representative direction candidate R8b, which is closer to the nadir (elevation angle -90°), is selected.

[0313] To achieve the above, the representative direction determination unit 133 operates as follows.

[0314] First, from the provided HRIR dataset, six HRIRs on the horizontal plane whose azimuth angle is closest to k×60+30° (where k = 0, 1, 2, 3, 4, and 5) are selected. Then, if there are no HRIRs on the horizontal plane, the representative direction determination unit 133 selects the HRIR with the elevation angle closest to 0°. If there are multiple candidates for one k value, the representative direction determination unit 133 prioritizes selecting the HRIR with the smaller azimuth angle within the range of 0° or more and less than 180°, and prioritizes selecting the HRIR with the larger azimuth angle within the range of 180° or more and less than 360°. Note that the elevation angles of the six HRIRs may all be the same, or may be different. Next, the representative direction determination unit 133 selects the HRIR with the elevation angle closest to 90° within the range of 0° to 90° on the median plane.

[0315] If there is no HRIR on the median plane, the representative direction determination unit 133 selects the HRIR with the azimuth angle closest to 0°. This is used as the zenith HRIR. Finally, the representative direction determination unit 133 selects the HRIR with the elevation angle closest to -90° within the range of 0° to -90° on the median plane. If there is no HRIR on the median plane, the representative direction determination unit 133 selects the HRIR with the azimuth angle closest to 0°. This is used as the nadir HRIR. In this way, HRIRs for six representative directions, as well as HRIRs for the zenith and nadir, can be selected from the horizontal plane, resulting in a total of eight HRIRs for representative directions.

[0316] Other examples of how to select a representative point that prioritizes the front side include the following. For example, the method for setting the eight representative directions described above is just an example, and it is not necessary to set an HRIR as the representative direction in the above steps. In other words, only when an HRIR representing a predetermined ideal direction is not included in the acquired HRIR set, an HRIR that prioritizes the front side of the listener as the representative direction should be set.

[0317] Here, the front side of the listener refers to the range indicated by a hemisphere with an elevation angle of -90° to 90°, an azimuth angle of 270° to 360°, and an azimuth angle of 0° to 90°. Furthermore, within the above-mentioned range, a position closer to an azimuth angle of 0° (= 360°) is referred to as being closer to the front. Conversely, the range indicated by a hemisphere with an azimuth angle of 90° to 270°, an elevation angle of -180° to -90°, and an elevation angle of 90° to 180° may be referred to as the rear side, and a position closer to an azimuth angle of 180° is referred to as being closer to the rear.

[0318] Another example of a method for selecting a representative point that prioritizes the front side will be described below. First, a representative point to be set on a horizontal plane will be described. For example, when six positions on a horizontal plane (elevation angle 0°) with azimuth angles of k×60+30° (where k=0, 1, 2, 3, 4, and 5) are defined as ideal directions on the horizontal plane, HRIRs that fall within a predetermined range of deviation from the azimuth angle of the ideal direction may be first extracted as candidates, and then the HRIR closest to the front side from among the multiple candidates may be set as the HRIR for the representative direction.

[0319] Next, the representative points to be set on the median plane will be described. When HRIRs on the median plane (azimuth angle 0°) with elevation angles of 90° and -90° are set as the ideal representative directions of the zenith and nadir, respectively, multiple candidates close to an azimuth angle of 0° that fall within a predetermined deviation range from the ideal value of an azimuth angle of 0° may first be extracted, and HRIRs with elevation angles closest to 90° and -90° may be selected from among them and set as the zenith and nadir, respectively. Conversely, HRIRs that fall within a predetermined deviation range from the ideal value of the elevation angle may first be extracted as candidates, and then the HRIR with an azimuth angle closest to 0° from among the multiple candidates may be set as the HRIR for the representative direction.

[0320] The predetermined range of deviation may be the same for both the azimuth angle and the elevation angle, or may be weighted towards whichever is more important, the azimuth angle or the elevation angle.

[0321] By setting the representative point in the manner described above, the representative point is set with priority given to the front side of the listener rather than the side or rear side, which makes it more likely that panning processing can be performed without impairing the listener's listening experience as much as possible.

[0322] Specifically, humans have a hearing characteristic that makes it easier to accurately perceive the direction of a sound source from the front than from the side or back. By moving the representative point closer to the front, sounds from the front can be reproduced more accurately, enabling playback that is in line with the nature of hearing.

[0323] Furthermore, by prioritizing the reproduction performance of sounds arriving from the front caused by objects in front of the listener's field of view over sounds caused by objects behind the listener's field of view, it may be possible to align visual perception with auditory perception and prevent the listener from feeling uncomfortable.

[0324] In other words, by selecting a representative point using the method disclosed in this specification, it is expected that the following effects (1) to (3) will be achieved compared to setting a representative point using a method other than this method.

[0325] (1) There are cases where the acquired HRIR set is sparse and it is not possible to obtain an HRIR in an ideal position as a representative point. However, if a representative point is selected using the method described above, even if the HRIR in the representative direction used for the panning process is not included in the acquired HRIR set, there is a high possibility that the panning process can be performed without impairing the listening experience.

[0326] (2) When there are multiple HRIR candidates within the same distance range from the ideal position, it is possible to select an HRIR that is more suitable for panning processing.

[0327] (3) It is possible to select a representative point so that sound from the front (front side), which is important for perception, can be localized more accurately than when the areas in front and behind the listener are not taken into consideration.

[0328] Although the representative point selection method described above is applied to panning processing using time shift and gain adjustment as an example of an embodiment, the present invention is not limited to this example and can also be applied to other panning processing such as VBAP.

[0329] [Cross-fade processing accompanying panning processing] The necessity of cross-fade processing accompanying panning processing and specific means for performing the same will be further described below. Fig. 23 is a diagram for explaining the cross-fade processing according to the embodiment. Fig. 24 is a flowchart of the operation related to the cross-fade processing according to the embodiment.

[0330] If the positional relationship between the listener and the sound source object changes over time, the representative point R used in the panning process and the time shift and gain adjustment parameters used in the panning process linked to the representative point R will switch. If the time shift and gain adjustment parameters used in the panning process switch, the time domain signal generated by connecting the current frame and the next frame will become discontinuous, which may cause an abnormal sound to be generated near the listener's ear or sound unnatural.

[0331] In order to prevent such auditory inconvenience, it is preferable to perform cross-fade processing when the representative point R used in the panning process is switched, or when the time shift and gain adjustment parameters used in the panning process linked to the representative point R are switched.

[0332] The cross-fade process is performed in the following procedure as shown in FIG.

[0333] First, one sound signal is generated by multiplying the time domain including the time shift amount (delay1) associated with the representative point before switching by the gain value (gain1) associated with the representative point before switching (preTemp).

[0334] Next, the other sound signal is generated by multiplying the time domain including the time shift amount (delay2) associated with the representative point after the switching by the gain value (gain2) associated with the representative point after the switching (Temp).

[0335] The start point of the fade-out for one of the audio signals is determined according to the time shift amount (delay1) associated with the representative point before switching. The volume is gradually reduced from the start point of the fade-out.

[0336] The fade-in start point for the other sound signal is determined according to the time shift amount (delay2) associated with the representative point after switching. Processing is performed to gradually increase the volume from the fade-in start point.

[0337] Finally, one sound signal and the other sound signal are added together at a representative point.

[0338] If the time shift and gain adjustment parameters used in the panning process do not change, there is no auditory benefit to performing cross-fade processing, and in fact the amount of calculation increases, so it is preferable to skip cross-fade processing when these parameters do not change. Here, we will describe a method for detecting when cross-fade processing is necessary and performing cross-fade processing only.

[0339] The determination as to whether or not to execute the cross-fade process is made in the following procedure, as shown in FIG. 24A.

[0340] First, the information update thread, which is a low-frequency processing thread, determines whether there has been a change in the positional relationship between the sound source position and the listener position for each thread, and if there has been a change, turns on a change flag indicating that the panning parameters, the time shift amount and gain adjustment amount, have changed.If there has been no change in the positional relationship, the parameters have not changed either, so the change flag is turned off.Whether the sound source position relative to the listener has changed may also be determined by whether the selected SOFA index is different from that of the previous thread.

[0341] The process then proceeds to the process shown in Fig. 24A. Fig. 24A shows the processing flow of the audio output thread, which is a high-frequency processing thread. As shown in Fig. 24A, the panning parameters and change flags for each item are obtained (S401), and a loop is entered in which the same process is repeated for each item according to the number of items (S402).

[0342] In step S402, the process is performed when the parameters are switched and the change flag is on. Specifically, it is determined whether the acquired change flag for the selected item is on (S402a). If the change flag is on (Yes in S402a), the previous parameters and the new current parameters are multiplied by a time shift and a gain value to perform a crossfade, and then added to the representative point position (S402b).

[0343] Then, for the next processing, the current parameters are transferred to the previous parameters, and the current parameters are replaced so that they are treated as the previous parameters in the next processing (S402c). If the change flag is not on (No in S402a), steps S402b and S402c are skipped and the process proceeds to the next loop. When the loop processing for the number of items is completed, a crossfade is performed for each representative point that groups together items whose change flags are on (S403). A representative point that groups together items whose change flags are set to on refers to a representative point that includes items whose change flags are on among the multiple items that are the subject of collectively adding at the representative point position.

[0344] Thereafter, a loop is entered in which the same process is repeated for each item according to the number of items (S404).

[0345] In step S404, processing is performed for the case where there is no parameter switching, the change flag is off, and crossfading is not performed. In this loop processing, when there is no parameter switching, there is no need to change the time shift amount and gain adjustment amount, and there is no need to perform crossfading on audio signals that are continuous on the time axis, so processing is performed to prevent crossfading.

[0346] Specifically, it is determined whether the acquired change flag for the selected target item is off (S404a). If the change flag is off (Yes in S404a), the previously set panning parameters are multiplied by the time shift and gain value, and then added to the representative point position (S404b). If the change flag is not off (No in S404a), step S404b is skipped and the process proceeds to the next step S404c.

[0347] In step S404c, the change flag for the selected item is turned off (S404c). Once the change flag is turned off, it remains off until a change in positional relationship is detected in the information update thread. After that, the signal data for the current frame is copied to the delay buffer so that it can be referenced in the next processing. That is, the current data is copied to the previous data (specifically, as shown in FIG. 23, "curr256smpls" is moved to "previous256smple"; S404d). When the next frame is processed, cross-fade processing is performed using the signal data stored in the delay buffer and the signal data for the new frame. After completing the loop processing for the number of items, the audio signals grouped at the representative point by the panning processing are convolved with the HRIR based on the azimuth and elevation angles of the representative point (S405). Since this convolution is performed on the frequency axis, it is converted to the time axis (S406) and added to the output buffer.

[0348] In FIG. 24A , an example was described in which the determination of whether the change flag is on and the determination of whether the change flag is off are performed in separate loop processes. However, if the change flag is not on, it can be considered that the change flag is off. Therefore, the determination of whether the change flag is on and the determination of whether the change flag is off can be performed simultaneously. The two loop processes (steps S402 and S404) shown in FIG. 24A can be combined into a single loop process. FIG. 24B shows the processing flow of the audio output thread, which is a high-frequency processing thread, when the two loop processes are combined into a single loop process. Note that in FIG. 24B , the same processes as in FIG. 24A are denoted by the same reference numerals. The specific content of the processes in FIG. 24B may be omitted by referring to the description of the processes denoted by the same reference numerals in FIG. 24A .

[0349] As shown in FIG. 24B, the loop processing of step S402 and the loop processing of step S404 in FIG. 24A are combined into a single loop processing S410. Specifically, after step S401, a loop is executed in step S410 in which the same processing is repeated for each item according to the number of items. In step S410, it is first determined whether the change flag is on (S402a). If the change flag is on (Yes in S402a), the process proceeds to step S402b as in FIG. 24A, and then to step S402c. On the other hand, if the change flag is not on (No in S402a), the change flag can be said to be off, and the process proceeds to step S404b as in the case where step S404a in FIG. 24A is Yes, and then to step S404c. Note that in this example, the process proceeds to step S404c after step S402c. When the loop processing for the number of items is completed, a crossfade is performed for each representative point that groups together items whose change flags are on (S403). Then, in this example, step S404d, in which the current data is copied to the previous data for the next processing, is performed after step S403, not within the loop processing of step S410. Then, steps S405 and S406 are executed sequentially. This allows a single loop processing to determine whether the change flag is on and whether the change flag is off, and then perform the subsequent processing.

[0350] The relationship between the information update thread and the audio output thread can be summarized from the processing up to this point. The information update thread (not shown) detects changes in the relative position between the listener and the sound source, sets panning parameters, and determines whether to switch the change flag. The audio output thread shown in Figure 24 references the change flag and panning parameters set in the information update thread and executes audio output processing, including cross-fade processing.

[0351] In the above-described method, the relative positions of the listener and sound source need only be determined in the thread that occurs infrequently, while the audio output thread that occurs frequently only needs to detect whether the flag is on or off. This makes it possible to switch between performing cross-fade processing and not performing it while preventing any adverse effects on the audio output processed by the audio output thread.

[0352] The above threads will be described below. When the spatial information management unit 1201 and the rendering unit 1202 execute the spatial information update process (information update thread) and the audio signal output process (audio thread) with acoustic processing added in different threads, the startup frequency of the threads may be set individually, or the processes may be executed in parallel.

[0353] When the spatial information management unit 1201 and the rendering unit 1202 execute processes in different independent threads, it is possible to allocate computing resources preferentially to the rendering unit 1202. This makes it possible to safely execute sound output processing in which even the slightest delay is unacceptable, for example, in which a delay of one sample (0.02 msec) would cause a popping noise.

[0354] In this case, the allocation of computational resources to the spatial information management unit 1201 is limited. However, because updating of spatial information is a process that occurs less frequently than output processing of audio signals (for example, a process such as updating the direction of the listener's face), it does not necessarily have to be performed instantaneously like output processing of audio signals. Therefore, even if the allocation of computational resources is limited, it does not have a significant impact on acoustic quality.

[0355] The spatial information may be updated periodically at preset times or intervals, or when preset conditions are met. The spatial information may also be updated manually by a listener or a sound space manager, or may be updated in response to a change in an external system.

[0356] For example, the spatial information may be updated when a listener operates a controller to instantly warp the position of his / her avatar or instantly advance or reverse the time. Alternatively, the spatial information may be updated when an administrator of the virtual space suddenly changes the environment of the space. In these cases, the thread for updating the spatial information managed by the spatial information management unit 1201 may be started as a one-off interrupt process in addition to being started periodically.

[0357] For example, these processes may be performed when the virtual space is created (when the software is created), when processing of the virtual space starts (when the software is launched or rendering starts), or when an information update thread that occurs periodically in processing of the virtual space occurs, etc. Furthermore, when the virtual space is created may be when the virtual space is constructed before the start of acoustic processing, or when information about the virtual space (spatial information) is acquired, or when the software is acquired.

[0358] Here, in the information update thread, processing is performed to update the spatial information managed by the spatial information management units (1101, 1201).

[0359] The role of the information update thread is, for example, to update the position and orientation of the listener's avatar placed in the virtual space based on the position and orientation of the VR goggles worn by the listener, or to update the position of an object moving in the virtual space, etc. Such processing is handled within a processing thread that runs at a relatively low frequency of about several tens of Hz.

[0360] The process of updating information indicating the characteristics of the direct sound may be performed in such a processing thread that occurs less frequently. This is because the characteristics of the direct sound change less frequently than the frequency with which audio processing frames for audio output occur. This makes it possible to relatively reduce the computational load of this process. Furthermore, updating information at an unnecessarily high frequency poses a risk of generating pulsive noise. Updating information at a low frequency makes it possible to avoid such a risk.

[0361] [Switching to Original Processing Without Panning Processing] In the above-described embodiment, a method for setting two or more predetermined representative points R from multiple HRIRs included in an HRIR set has been described. However, if the acquired HRIR set contains a small number of HRIRs, setting a predetermined number of representative points R may make it impossible to solve the formula for calculating the gain value used in the panning processing, and an appropriate gain value may not be obtained. This may make it difficult to output the audio signal itself. Furthermore, even if the formula for calculating the gain value can be solved, an abnormally large value may be calculated. In this case, even if it is possible to output an audio signal, the sound quality may deteriorate, making it difficult to output an appropriate audio signal.

[0362] For example, the gain calculation formula used in the panning process refers to the above-mentioned formulas (1) to (8) and formulas (9) to (11). A case where a formula cannot be solved occurs, for example, when a preset number of representative points R are selected from a sparse HRIR set, and the same HRIR is set as another representative point R. In other words, for example, among v{x1}, v{x2}, and v{x3} in formulas (9) to (11), v{x1} and v{x2} are the same vector. In such a case, the matrix equation is not regular, making it impossible to solve the inverse matrix, and as a result, it may be impossible to obtain an appropriate gain value.

[0363] In order to avoid such a case, an example will be described in which, when there is a high possibility that the gain calculation formula cannot be solved, the panning process is not performed and the convolution process is performed using the HRIRs included in the HRIR set. For example, the cases when there is a high possibility that the gain calculation formula cannot be solved are determined according to the number and distribution of HRIRs included in the HRIR set.

[0364] When a sparse HRIR set is obtained, if a convolution process is performed using the HRIRs included in the HRIR set without performing a panning process, even if the spatial resolution of the output sound signal is reduced, it is possible to avoid the above-mentioned errors occurring and making it impossible to output a sound signal.

[0365] If the distribution of HRIRs in the HRIR set is, for example, (1) or (2) below, it is possible to perform convolution processing of head-related transfer functions according to the direction of arrival based on the listener position and sound source position without performing panning processing.

[0366] (1) When the HRIR is only in front or behind the horizontal plane. (2) When the HRIR is only at an elevation angle of 0 degrees or more.

[0367] If the HRIRs included in the HRIR set are distributed only behind the horizontal plane in the celestial sphere, attempting to perform the above-described panning process is likely to result in failure to properly set the required representative point on the horizontal plane, and therefore, in such cases, it is preferable not to perform the panning process.

[0368] If the HRIRs included in the HRIR set are distributed only at elevation angles of 0° or greater or at elevation angles of 0° or less across the entire celestial sphere, attempting to perform the above-described panning process is likely to result in an inability to properly set the representative point of the zenith or nadir, and therefore, in such cases, it is preferable not to perform the panning process.

[0369] Although the decision on whether to execute panning processing may be made at any time after the HRIR set is acquired, it is preferable to make the decision at software startup or initialization, for example, because the HRIR set is acquired when the software is started or initialized. By making the decision only at such times, it becomes unnecessary to make the decision in the periodically generated information update thread, and it becomes possible to reduce the amount of calculation in the periodically generated thread.

[0370] Other Embodiments Although the embodiments have been described above, the present disclosure is not limited to the above-described embodiments.

[0371] For example, the sound reproduction system described in the above embodiment may be realized as a single device including all of the components, or may be realized by allocating each function to multiple devices and coordinating these multiple devices. In the latter case, the device corresponding to the information processing device may be an information processing device such as a smartphone, tablet terminal, or PC. For example, in the sound reproduction system 100 having a function as a renderer that generates an audio signal with added sound effects, all or part of the renderer's functions may be performed by a server. That is, all or part of the acquisition unit 111, path calculation unit 121, output sound generation unit 131, and signal output unit 141 may be located on a server (not shown). In this case, the sound reproduction system 100 is realized by combining, for example, an information processing device such as a computer or smartphone, a sound presentation device such as a head-mounted display (HMD) or earphones worn by the user 99, and a server (not shown). Note that the computer, sound presentation device, and server may be connected to each other so as to be able to communicate with each other via the same network, or may be connected via different networks. When the sound reproduction system 100 is connected via different networks, the possibility of communication delays increases, so processing on the server may be permitted only when the computer, sound presentation device, and server are connected so as to be able to communicate with each other via the same network. Also, depending on the amount of bitstream data received by the sound reproduction system 100, it may be determined whether the server will take on all or part of the functions of the renderer.

[0372] The sound reproduction system of the present disclosure can also be realized as an information processing device that is connected to a reproduction device having only a driver and that simply reproduces an output sound signal generated based on acquired sound information to the reproduction device. In this case, the information processing device may be realized as hardware having a dedicated circuit, or as software that causes a general-purpose processor to execute specific processing.

[0373] In the above-described embodiment, the processing performed by a specific processing unit may be performed by another processing unit. The order of multiple processing operations may be changed, or multiple processing operations may be performed in parallel.

[0374] In the above-described embodiments, each component may be realized by executing a software program suitable for that component, or by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0375] Furthermore, each component may be realized by hardware. For example, each component may be a circuit (or integrated circuit). These circuits may form a single circuit as a whole, or each may be a separate circuit. Furthermore, each of these circuits may be a general-purpose circuit or a dedicated circuit.

[0376] Furthermore, the general or specific aspects of the present disclosure may be realized as an apparatus, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, etc. Furthermore, the general or specific aspects of the present disclosure may be realized as any combination of an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0377] For example, the present disclosure may be realized as an audio signal reproducing method executed by a computer, or as a program for causing a computer to execute the audio signal reproducing method. The present disclosure may also be realized as a computer-readable non-transitory recording medium on which such a program is recorded.

[0378] In addition, this disclosure also includes forms obtained by applying various modifications to each embodiment that a person skilled in the art would think of, or forms realized by arbitrarily combining the components and functions of each embodiment within the scope that does not deviate from the intent of this disclosure.

[0379] In the present disclosure, the encoded sound information can be rephrased as a bitstream containing a sound signal, which is information about a predetermined sound to be reproduced by the sound reproduction system 100, and metadata, which is information about the localization position when the sound image of the predetermined sound is localized at a predetermined position within a three-dimensional sound field. For example, the sound information may be acquired by the sound reproduction system 100 as a bitstream encoded in a predetermined format such as MPEG-H 3D Audio (ISO / IEC 23008-3). As an example, the encoded sound signal includes information about the predetermined sound to be reproduced by the sound reproduction system 100. The predetermined sound here refers to a sound emitted by a sound source object present in the three-dimensional sound field or a natural environmental sound, and may include, for example, a mechanical sound or the sounds of animals, including humans. When multiple sound source objects are present in the three-dimensional sound field, the sound reproduction system 100 acquires multiple sound signals corresponding to the multiple sound source objects.

[0380] On the other hand, metadata is, for example, information used to control acoustic processing of a sound signal in the sound reproduction system 100. Metadata may be information used to describe a scene expressed in a virtual space (three-dimensional sound field). Here, a scene is a term that refers to a collection of all elements representing three-dimensional images and acoustic events in a virtual space, modeled by the sound reproduction system 100 using metadata. In other words, the metadata here may include not only information for controlling acoustic processing, but also information for controlling video processing. Of course, the metadata may include information for controlling only either audio processing or video processing, or may include information used to control both. In the present disclosure, the bitstream acquired by the sound reproduction system 100 may include such metadata. Alternatively, the sound reproduction system 100 may acquire the metadata separately from the bitstream, as described below.

[0381] The sound reproduction system 100 generates virtual sound effects by performing sound processing on the sound signal using metadata included in the bitstream and additionally acquired position information of the interactive user 99. For example, sound effects such as early reflection sound generation, late reverberation sound generation, diffraction sound generation, distance attenuation effect, localization, sound image localization processing, and Doppler effect may be added. Information for switching on and off all or part of the sound effects may also be added as metadata.

[0382] All or part of the metadata may be obtained from sources other than the bitstream of audio information. For example, either the metadata controlling audio or the metadata controlling video may be obtained from sources other than the bitstream, or both may be obtained from sources other than the bitstream.

[0383] Furthermore, if metadata for controlling the video is included in the bitstream acquired by the audio reproduction system 100, the audio reproduction system 100 may have a function for outputting the metadata that can be used for controlling the video to a display device that displays images or a 3D video reproduction device that reproduces 3D video.

[0384] As an example, the encoded metadata includes information about a three-dimensional sound field including a sound source object emitting a sound and an obstacle object, and information about a localization position when a sound image of the sound is localized at a predetermined position in the three-dimensional sound field (i.e., the sound is perceived as arriving from a predetermined direction), i.e., information about the predetermined direction. Here, an obstacle object is an object that can affect the sound perceived by the user 99 by, for example, blocking or reflecting the sound emitted by the sound source object before it reaches the user 99. Obstacle objects may include not only stationary objects but also animals such as people or moving objects such as machines. Furthermore, when multiple sound source objects exist in a three-dimensional sound field, the other sound source objects may be obstacle objects for any one sound source object. Furthermore, both non-sound-source objects such as building materials or inanimate objects and sound-emitting sound objects may be obstacle objects.

[0385] The spatial information constituting the metadata may include not only the shape of the three-dimensional sound field, but also information representing the shape and position of obstacle objects present in the three-dimensional sound field and the shape and position of sound source objects present in the three-dimensional sound field. The three-dimensional sound field may be either a closed space or an open space, and the metadata may include information representing the reflectance of structures that can reflect sound in the three-dimensional sound field, such as floors, walls, or ceilings, and the reflectance of obstacle objects present in the three-dimensional sound field. Here, the reflectance is the ratio of the energy of reflected sound to incident sound, and is set for each frequency band of sound. Of course, the reflectance may be set uniformly regardless of the frequency band of the sound. Furthermore, when the three-dimensional sound field is an open space, parameters such as a uniform attenuation rate, diffracted sound, or early reflection sound may be used.

[0386] In the above description, reflectance is cited as a parameter related to an obstacle object or a sound source object included in the metadata, but the metadata may include information other than reflectance. For example, information about the material of the object may be included as metadata related to both the sound source object and the non-sound source object. Specifically, the metadata may include parameters such as diffusion rate, transmittance, or sound absorption rate.

[0387] Information about the sound source object may include information such as volume, radiation characteristics (directivity), playback conditions, the number and type of sound sources emitted from a single object, or information specifying a sound source area within the object. The playback conditions may, for example, determine whether the sound is a continuous sound or an event-triggering sound. The sound source area within the object may be determined based on the relative relationship between the position of the user 99 and the position of the object, or may be determined based on the object. When the sound source area is determined based on the relative relationship between the position of the user 99 and the position of the object, the surface from which the user 99 is viewing the object is used as the reference, and the user 99 can be made to perceive sound X as emanating from the right side of the object and sound Y as emanating from the left side of the object as viewed from the user 99. When the sound source area is determined based on the object, the surface from which the user 99 is viewing the object is used as the reference, and the sound emitted from which area of ​​the object can be fixed regardless of the direction the user 99 is viewing. For example, the user 99 can be made to perceive a high-pitched sound coming from the right side and a low-pitched sound coming from the left side when viewing the object from the front. In this case, when the user 99 goes around to the back of the object, the user 99 can be made to perceive a low-pitched sound coming from the right side and a high-pitched sound coming from the left side as viewed from the back.

[0388] The spatial metadata may include the time to early reflections, the reverberation time, or the ratio of direct sound to diffuse sound. If the ratio of direct sound to diffuse sound is zero, the user 99 will perceive only direct sound.

[0389] Furthermore, information indicating the position and orientation of the user 99 in the three-dimensional sound field may be included in the bitstream as metadata as an initial setting, or may not be included in the bitstream. If the information indicating the position and orientation of the user 99 is not included in the bitstream, the information indicating the position and orientation of the user 99 is acquired from information other than the bitstream. For example, the position information of the user 99 in the VR space may be acquired from an app that provides VR content. The position information of the user 99 for presenting sound as AR may be position information obtained by a mobile terminal performing self-position estimation using GPS, a camera, LiDAR (Laser Imaging Detection and Ranging), or the like. Note that the sound signal and metadata may be stored in a single bitstream or may be stored separately in multiple bitstreams. Similarly, the sound signal and metadata may be stored in a single file or may be stored separately in multiple files.

[0390] When an audio signal and metadata are stored separately in multiple bitstreams, information indicating other related bitstreams may be included in one or some of the multiple bitstreams in which the audio signal and metadata are stored. Also, information indicating other related bitstreams may be included in metadata or control information of each of the multiple bitstreams in which the audio signal and metadata are stored. When an audio signal and metadata are stored separately in multiple files, information indicating other related bitstreams or files may be included in one or some of the multiple files in which the audio signal and metadata are stored. Also, information indicating other related bitstreams or files may be included in metadata or control information of each of the multiple bitstreams in which the audio signal and metadata are stored.

[0391] Here, the related bitstreams or files are, for example, bitstreams or files that may be used simultaneously during audio processing. Furthermore, information indicating other related bitstreams may be collectively described in the metadata or control information of one bitstream among multiple bitstreams storing audio signals and metadata, or may be separately described in the metadata or control information of two or more bitstreams among the multiple bitstreams storing audio signals and metadata. Similarly, information indicating other related bitstreams or files may be collectively described in the metadata or control information of one file among multiple files storing audio signals and metadata, or may be separately described in the metadata or control information of two or more files among the multiple files storing audio signals and metadata. Furthermore, a control file collectively describing information indicating other related bitstreams or files may be generated separately from the multiple files storing audio signals and metadata. In this case, the control file does not need to store the audio signal and metadata.

[0392] Here, the information indicating the other related bitstream or file may be, for example, an identifier indicating the other bitstream, a file name indicating the other file, a URL (Uniform Resource Locator), or a URI (Uniform Resource Identifier). In this case, the acquisition unit identifies or acquires the bitstream or file based on the information indicating the other related bitstream or file. Furthermore, the information indicating the other related bitstream may be included in metadata or control information of at least some of the bitstreams among a plurality of bitstreams storing audio signals and metadata, and the information indicating the other related file may be included in metadata or control information of at least some of the files among a plurality of files storing audio signals and metadata. Here, the file including information indicating the related bitstream or file may be, for example, a control file such as a manifest file used for content distribution.

[0393] The present disclosure is useful in reproducing sound, for example, by allowing a user to perceive stereoscopic sound.

[0394] 99 User 100 Sound reproduction system 101 Information processing device 102 Communication module 103 Detector 104 Driver 105 Database 111 Acquisition unit 112 Encoded sound information input unit 113 Decode processing unit 114 Sensing information input unit 121 Path calculation unit 131 Output sound generation unit 133 Representative direction determination unit 134 Generation unit 135 Synthesis unit 141 Signal output unit 300 3D video reproduction device

Claims

1. An acoustic information processing method executed by an information processing terminal, comprising the steps of: acquiring position information of a sound source object in a three-dimensional sound field based on an acoustic signal; acquiring position information of a user in the three-dimensional sound field; determining a plurality of representative directions; and performing a panning process to distribute the signal of the sound source object to the plurality of representative directions based on the user's position information and the sound source object's position information, wherein in the determining step, one or more candidates are selected from among selectable representative direction candidates based on the smallness of deviation from an ideal direction in a horizontal plane, and if only one candidate is selected, the one candidate is determined to be one of the plurality of representative directions, and if two or more candidates are selected, the one candidate closest to a front direction of the user is determined to be one of the plurality of representative directions.

2. The acoustic information processing method of claim 1, further comprising the steps of: selecting, from among the selectable representative direction candidates, one candidate that has the smallest angle with respect to the user's median plane and is closest to an elevation angle of 90° from among direction candidates within a range of 0 to 90° when the horizontal plane direction of the user is defined as an elevation angle of 0°, and determining the one candidate as one of the plurality of representative directions.

3. The acoustic information processing method according to claim 1 or 2, wherein the determining step further comprises selecting, from among the selectable representative direction candidates, one candidate that has the smallest angle with respect to the user's median plane and is closest to an elevation angle of -90° from among direction candidates within a range of elevation angles of -90° to 0° when the horizontal plane direction of the user is an elevation angle of 0°, and determining the one candidate as one of the plurality of representative directions.

4. The acoustic information processing method according to claim 1 or 2, wherein the ideal directions in the horizontal plane are six ideal directions that divide 360° of the horizontal plane into six equal parts, and in the determining step, one of the plurality of representative directions is determined for each of the six ideal directions.

5. The acoustic information processing method according to claim 4, wherein the ideal directions in the horizontal plane are six ideal directions corresponding to 30, 90, 150, 210, 270, and 330° on the horizontal plane when the front direction of the user is set to 0°, and in the determining step, one of the plurality of representative directions is determined for each of the six ideal directions.

6. An information processing device comprising: a first acquisition unit that acquires position information of a sound source object in a three-dimensional sound field based on an acoustic signal; a second acquisition unit that acquires position information of a user in the three-dimensional sound field; a representative direction determination unit that determines a plurality of representative directions; and a panning unit that performs a panning process to distribute the signal of the sound source object to the plurality of representative directions based on the user's position information and the sound source object's position information, wherein the representative direction determination unit selects one or more candidates from among selectable representative direction candidates based on the smallness of deviation from an ideal direction in a horizontal plane, and if there is only one selected candidate, determines the one candidate as one of the plurality of representative directions, and if there are two or more selected candidates, determines the one candidate from the two or more candidates that is closest to a front direction of the user as one of the plurality of representative directions.

7. A program for causing a computer to execute the acoustic information processing method according to claim 1.

Citation Information

Patent Citations

  • Voice generation program in virtual space, generation method of quadtree, and voice generation device

    JP2020018620A

  • Sound generation apparatus, sound reproducing apparatus, sound generation method, and sound signal processing program

    JP2023164284A

  • Reproduction apparatus, reproduction method, information processing apparatus, information processing method, and program

    WO2022124084A1