Information processing device, information processing method, and program

By synthesizing sound signals using head-related transfer functions and selective panning processing, the method addresses computational challenges in virtual three-dimensional sound reproduction, maintaining sound quality and realism.

WO2025205328A1PCT designated stage Publication Date: 2025-10-02AKITA PREFECTURAL UNIVERSITY +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/010710
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-25
Filing Date
2025-03-19
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing sound reproduction technologies for virtual three-dimensional spaces face challenges in processing large amounts of data, leading to increased computational requirements and potential sound quality degradation when applying panning processing to reduce processing load.

Method used

A method that generates an output sound signal by synthesizing a first sound signal convolved with a head-related transfer function based on the direction of arrival and a second sound signal convolved with a representative direction, selectively applying panning processing to reduce processing load while minimizing sound degradation.

Benefits of technology

This approach balances processing load reduction with sound quality preservation, enabling more realistic sound reproduction in virtual three-dimensional environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025010710_02102025_PF_FP_ABST
    Figure JP2025010710_02102025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device (101) uses an information processing method which is executed by a computer, the method including: a step (S101) for acquiring sound information that includes an acoustic signal and information as to the position of a sound source object in a three-dimensional sound field; a step (S107) for generating a first sound signal by using the acoustic signal and a head transfer function corresponding to an arrival direction based on the position of the sound source object and the position of a user in the three-dimensional sound field; a step (S106) for generating a second sound signal by using the acoustic signal and a head transfer function corresponding to a representative direction based on the position of a representative point set in the three-dimensional sound field and the position of the user; and a step (S109) for synthesizing the generated first sound signal and second sound signal to generate an output sound signal.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and program

[0001] The present disclosure relates to an information processing device, an information processing method, and a program.

[0002] Conventionally, there are known techniques related to sound reproduction that allow a user (also referred to as a listener of the reproduced sound) to perceive three-dimensional sound in a virtual three-dimensional space (see, for example, Patent Document 1). Furthermore, in order to perceive sound as if it is coming from a sound source object to the user in such a three-dimensional space, processing is required to generate output sound information from original sound information. In particular, reproducing three-dimensional sound in response to the user's body movements in a virtual space requires extensive processing. Advances in computer graphics (CG) have made it relatively easy to create visually complex virtual environments, making technology that realizes corresponding auditory information important. In addition, when processing from sound information to output sound information is performed in advance, a large memory area is required to store the pre-calculated processing results. Furthermore, transmitting such large amounts of processed data may require a wide communication bandwidth.

[0003] To realize a more realistic sound environment, the number of objects that emit sound in the virtual three-dimensional space increases, and secondary sounds based on acoustic effects such as reflected sound, diffracted sound, and reverberation also increase, and these secondary sounds must be appropriately changed in response to the user's movements, requiring a large amount of processing.To reduce this large amount of processing, a conversion technique called panning processing is known, which represents sounds in a three-dimensional space using sounds from several representative points set in advance within the three-dimensional space.

[0004] Japanese Patent Application Laid-Open No. 2020-18620

[0005] However, in a conversion process such as panning, a reduction in the amount of processing may result in a degradation of sound quality. Therefore, an object of the present disclosure is to provide an information processing method for appropriately performing the conversion process.

[0006] An information processing method according to one aspect of the present disclosure is an information processing method executed by a computer, and includes the steps of acquiring sound information including an acoustic signal and information on the position of a sound source object within a three-dimensional sound field; generating a first sound signal using the acoustic signal and a head-related transfer function corresponding to a direction of arrival based on the position of the sound source object and the position of a user within the three-dimensional sound field; generating a second sound signal using the acoustic signal and a head-related transfer function corresponding to a representative direction based on the position of a representative point set within the three-dimensional sound field and the position of the user; and synthesizing the generated first sound signal and second sound signal to generate an output sound signal.

[0007] In addition, an information processing device according to one aspect of the present disclosure includes an acquisition unit that acquires sound information including an acoustic signal and information on the position of a sound source object within a three-dimensional sound field; a first generation unit that generates a first sound signal using a head-related transfer function according to an arrival direction based on the position of the sound source object and the position of a user within the three-dimensional sound field and the acoustic signal; a second generation unit that generates a second sound signal using a head-related transfer function according to a representative direction based on the position of a representative point set within the three-dimensional sound field and the position of the user and the acoustic signal; and a synthesis unit that synthesizes the generated first sound signal and second sound signal to generate an output sound signal.

[0008] Furthermore, one aspect of the present disclosure can also be realized as a program for causing a computer to execute the information processing method described above.

[0009] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0010] According to the present disclosure, it is possible to appropriately perform conversion processing.

[0011] FIG. 1 is a schematic diagram illustrating a use example of a sound reproduction system according to an embodiment. FIG. 2 is a block diagram illustrating a functional configuration of the sound reproduction system according to an embodiment. FIG. 3 is a diagram illustrating an example of an audio signal according to an embodiment. FIG. 4 is a block diagram illustrating a functional configuration of an acquisition unit according to an embodiment. FIG. 5 is a block diagram illustrating a functional configuration of an output sound generation unit according to an embodiment. FIG. 6 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 7 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 8 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 9 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 10 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 11 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 12 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 13 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 14 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 15 is a diagram illustrating another example of a sound reproduction system according to an embodiment. FIG. 16 is a flowchart illustrating an example of the operation of an information processing device according to an embodiment.

[0012] (Knowledge that forms the basis of the disclosure) Conventionally, a technology related to sound reproduction that allows a user to perceive stereoscopic sound in a virtual three-dimensional space (hereinafter sometimes referred to as a three-dimensional sound field) has been known (see, for example, Patent Document 1). Using this technology, a user can perceive sound as if a sound source object exists at a predetermined position in the virtual space and the sound is coming from that direction. In order to localize a sound image at a predetermined position in the virtual three-dimensional space in this way, for example, calculation processing is required for a sound signal emitted by a sound source object (also referred to as a sound emitted from the sound source object or a reproduced sound) to generate a sound arrival time difference between the two ears and a sound level difference (or sound pressure difference) between the two ears that causes the sound to be perceived as stereoscopic sound. Such calculation processing is performed by applying a stereophonic filter. A stereophonic filter is an information processing filter that, when an output sound signal obtained by applying the filter to original sound information is reproduced, causes the position (such as the direction and distance) of the sound, the size of the sound source, the width of the space, and the like to be perceived with a three-dimensional effect.

[0013] As an example of the computational process for applying such a stereophonic filter, a process is known in which a head-related transfer function (HRTF) is convolved with a target sound signal to make the sound perceived as coming from a predetermined direction. By performing this HRTF convolution process at a sufficiently fine angle with respect to the direction of arrival of the reproduced sound from the position of the sound source object to the user's position, the sense of realism experienced by the user is improved.

[0014] In recent years, there has been active development of technologies related to virtual reality (VR). Virtual reality focuses on appropriately changing the position of a sound source object in a virtual three-dimensional space in response to a user's movements, allowing the user to experience the sensation of moving within the virtual space. To achieve this, it is necessary to move the localization position of a sound image in the virtual space relative to the user's movements. This processing has been performed by applying a stereophonic filter, such as the head-related transfer function described above, to the original sound information. However, when a user moves within a three-dimensional space, the sound transmission path changes from moment to moment due to changes in the positional relationship between the sound source object and the user, such as due to sound reverberation and interference. This requires determining the sound transmission path from the sound source object based on the positional relationship between the sound source object and the user, and convolving the transfer function to take into account sound reverberation and interference. However, such information processing requires a huge amount of processing, and an improvement in the sense of realism may not be achieved without a large-scale processing device.

[0015] Therefore, in order to reduce such an enormous amount of processing, attempts have been made to apply a panning process to the reproduced sound to reduce the amount of convolution of the head-related transfer function. Specifically, instead of convolving the reproduced sound with a head-related transfer function for each of a number of sound source objects in a three-dimensional space, the reproduced sound from the sound source object is re-expressed using sounds (representative sounds) from several representative points pre-set in the three-dimensional space. Then, by simply convolving the representative sounds with the head-related transfer functions from the representative points to the user's position, it becomes possible to allow the user to perceive a three-dimensional sound that is comparable to that of the original sound source objects. If the number of representative points is smaller than the number of original sound source objects, the number of targets for convolution of the head-related transfer function will naturally be reduced, which is advantageous in terms of processing amount.

[0016] On the other hand, when applying such panning processing, the application of the panning processing may be inappropriate. Specifically, applying the panning processing changes the sound compared to the original reproduced sound, which may lead to a decrease in sound quality (sound degradation). Therefore, in the present disclosure, panning processing is applied partially to some acoustic signals included in the sound information, and for the remaining parts, convolution of a transfer function based on the positional relationship between the sound source object and the user is performed without applying panning processing. This makes it possible to reduce the amount of processing for the partial acoustic signals to which panning processing has been applied, while suppressing sound degradation in the remaining parts. In other words, the conversion processing (panning processing) can be performed more appropriately from the perspective of suppressing sound degradation.

[0017] A more specific outline of the present disclosure is as follows.

[0018] An information processing method according to a first aspect of the present disclosure is an information processing method executed by a computer, and includes the steps of acquiring sound information including an acoustic signal and information on the position of a sound source object within a three-dimensional sound field, generating a first sound signal using the acoustic signal and a head-related transfer function corresponding to a direction of arrival based on the position of the sound source object and the position of a user within the three-dimensional sound field, generating a second sound signal using the acoustic signal and a head-related transfer function corresponding to a representative direction based on the position of a representative point set in the three-dimensional sound field and the position of the user, and synthesizing the generated first sound signal and second sound signal to generate an output sound signal.

[0019] According to this information processing device, it is possible to generate a first sound signal using a head-related transfer function corresponding to the direction of arrival calculated using a first generation unit, generate a second sound signal using a head-related transfer function corresponding to a representative direction using a second generation unit, and generate an output sound signal by combining the first sound signal and the second sound signal. For example, in generating an output sound signal from one piece of sound information, it is possible to reduce the amount of processing by partially convolving a head-related transfer function corresponding to the representative direction, while suppressing sound degradation by partially convolving a head-related transfer function corresponding to the direction of arrival. In this way, it is possible to achieve a balance between reducing the amount of processing and suppressing sound degradation, making it possible to appropriately perform conversion processing.

[0020] An information processing method according to a second aspect is the information processing method according to the first aspect, wherein in the step of generating the first sound signal, the first sound signal is generated by convolving a head-related transfer function corresponding to the arrival direction with a reproduced sound emitted from the sound source object by the acoustic signal, and in the step of generating the second sound signal, a conversion process is performed to convert the reproduced sound into a representative sound arriving from the representative point, and the second sound signal is generated by convolving a head-related transfer function corresponding to the representative direction.

[0021] According to this, the first generation unit generates a first output sound signal by convolving a head-related transfer function according to the direction of arrival with a reproduced sound, and the second generation unit generates a second sound signal that expresses the sound from the direction of arrival using representative sounds arriving from respective representative points set in a three-dimensional sound field through a conversion process such as a panning process, and the output sound signal is generated by combining these. For example, the second generation unit can be used in a portion where applying the conversion process effectively reduces the amount of processing, and the first generation unit can be used in a portion where applying the conversion process is likely to lead to deterioration of sound.

[0022] An information processing method according to a third aspect is the information processing method according to the second aspect, wherein the conversion process applies time shift adjustment and gain adjustment to the reproduced sound to convert it into the representative sound.

[0023] According to this, in the conversion process, the playback sound can be converted into a representative sound by applying time shift adjustment and gain adjustment, which reduces the sense of incongruity even when the conversion process is applied, and enables the generation of an output sound signal with a higher sense of realism.

[0024] An information processing method according to a fourth aspect is an information processing method according to any one of the first to third aspects, wherein the sound information includes a plurality of the acoustic signals and a plurality of positions of the sound source objects corresponding one-to-one to the plurality of the acoustic signals, and the information processing method further includes a step of determining, for each of the plurality of acoustic signals included in the sound information, whether to use the individual acoustic signal to generate the first sound signal or to use the individual acoustic signal to generate the second sound signal.

[0025] This makes it possible to determine, for each audio signal corresponding to each of the multiple sound source objects contained in the sound information, i.e., for each sound emitted by each sound source object, whether to convolve a head-related transfer function according to the representative direction to reduce the amount of processing, or to convolve a head-related transfer function according to the direction of arrival to suppress sound degradation.

[0026] An information processing method according to a fifth aspect is an information processing method according to any one of the first to third aspects, wherein the sound information includes a plurality of the acoustic signals and the position of one of the sound source objects that corresponds many-to-one to the plurality of the acoustic signals, and the information processing method further includes a step of determining, for each of the plurality of acoustic signals included in the sound information, whether to use the individual acoustic signal to generate the first sound signal or to use the individual acoustic signal to generate the second sound signal.

[0027] This makes it possible to determine whether to convolve a head-related transfer function according to the representative direction to reduce the amount of processing, or to convolve a head-related transfer function according to the direction of arrival to suppress sound degradation, for each of the multiple audio signals corresponding to one sound source object contained in the sound information, i.e., for each of the multiple sounds emitted by the sound source object.

[0028] Furthermore, an information processing method according to a sixth aspect is the information processing method according to the fifth aspect, further including a step of converting the acquired sound information, in which the acquired sound information is converted into sound information including a plurality of the acoustic signals.

[0029] According to this, by converting the acquired sound information, sound information including a plurality of audio signals can be obtained.

[0030] Furthermore, an information processing method according to a seventh aspect is an information processing method according to any one of the fourth to sixth aspects, wherein the plurality of acoustic signals include an acoustic signal relating to a direct sound that arrives directly from the position of the corresponding sound source object to the position of the user, and an acoustic signal relating to a secondary sound that occurs in conjunction with the direct sound and arrives via a path different from that of the direct sound, and in the determining step, it is determined that the acoustic signal relating to the direct sound will be used to generate the first sound signal, and it is determined that the acoustic signal relating to the secondary sound will be used to generate the second sound signal.

[0031] This allows the amount of processing to be reduced by convolving the secondary sound with a head-related transfer function according to the representative direction, while suppressing sound degradation by convolving the direct sound with a head-related transfer function according to the direction of arrival.

[0032] Furthermore, an information processing method according to an eighth aspect is an information processing method according to any one of the fourth to seventh aspects, wherein the plurality of acoustic signals include acoustic signals relating to direct sounds that arrive directly from the position of the corresponding sound source object to the position of the user, and acoustic signals relating to secondary sounds that occur in association with the direct sounds and arrive via a path different from that of the direct sounds, and the secondary sounds include other secondary sounds that are further generated in association with one of the secondary sounds, and in the determining step, it is determined whether each of the acoustic signals relating to the secondary sounds will be used to generate the first sound signal or the second sound signal based on information relating to the generation system from the direct sound for each of the secondary sounds.

[0033] This makes it possible to determine whether each audio signal related to a secondary sound is to be used to generate a first sound signal or a second sound signal based on information regarding the generation system from the direct sound for each secondary sound.

[0034] An information processing method according to a ninth aspect is an information processing method according to any one of the fourth to eighth aspects, wherein in the determining step, it is determined whether each of the plurality of acoustic signals will be used to generate the first sound signal or the second sound signal based on the direction of the corresponding sound source object relative to the front direction of the user.

[0035] This makes it possible to determine whether each of the multiple audio signals is to be used to generate a first sound signal or a second sound signal based on the direction of the corresponding sound source object relative to the front direction of the user.

[0036] An information processing method according to a tenth aspect is an information processing method according to any one of the fourth to ninth aspects, wherein in the determining step, it is determined whether each of the plurality of acoustic signals will be used to generate the first sound signal or the second sound signal based on the distance from the user's position to the position of the corresponding sound source object.

[0037] This allows for determining whether each of the multiple audio signals is to be used to generate a first sound signal or a second sound signal based on the distance from the user's position to the position of the corresponding sound source object.

[0038] An information processing method according to an eleventh aspect is an information processing method according to any one of the fourth to tenth aspects, in which the determining step determines whether to use each of the plurality of acoustic signals to generate the first sound signal or the second sound signal based on characteristics related to sound quality degradation when the second sound signal is generated for each of the plurality of acoustic signals.

[0039] This makes it possible to determine whether to use a plurality of audio signals to generate a first sound signal or a second sound signal based on the characteristics of the sound quality degradation that occurs when a second sound signal is generated for each of the plurality of audio signals.

[0040] An information processing method according to a twelfth aspect is an information processing method according to any one of the first to eleventh aspects, in which the sound information includes an acoustic signal relating to a direct sound that arrives directly from the position of a corresponding sound source object to the position of the user, and an acoustic signal relating to a secondary sound that occurs in conjunction with the direct sound and arrives via a path different from that of the direct sound, and if the acoustic signal relating to the secondary sound is identified as a reverberant sound based on information relating to the generation system from the direct sound, it is decided not to use the acoustic signal for generating the second sound signal.

[0041] According to this, for each audio signal related to the secondary sound, it can be determined that the secondary sound identified as a reverberation sound is not used to generate the second sound signal.

[0042] In addition, an information processing device according to a thirteenth aspect includes an acquisition unit that acquires sound information including an acoustic signal and information on the position of a sound source object within a three-dimensional sound field; a first generation unit that generates a first sound signal using a head-related transfer function according to an arrival direction based on the position of the sound source object and the position of a user within the three-dimensional sound field and the acoustic signal; a second generation unit that generates a second sound signal using a head-related transfer function according to a representative direction based on the position of a representative point set within the three-dimensional sound field and the position of the user and the acoustic signal; and a synthesis unit that synthesizes the generated first sound signal and second sound signal to generate an output sound signal.

[0043] This can achieve the same effects as the information processing method described above.

[0044] A program according to a fourteenth aspect is a program for causing a computer to execute the information processing method described above.

[0045] This makes it possible to achieve the same effects as the information processing method described above using a computer.

[0046] Furthermore, these comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0047] Hereinafter, embodiments will be described in detail with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not recited in independent claims will be described as optional components. Note that each figure is a schematic diagram and is not necessarily an exact illustration. Furthermore, in each figure, substantially identical components are assigned the same reference numerals, and duplicated descriptions may be omitted or simplified.

[0048] In the following description, elements may be assigned ordinal numbers such as first, second, and third. These ordinal numbers are assigned to elements in order to identify them and do not necessarily correspond to a meaningful order. These ordinal numbers may be rearranged, newly added, or removed as appropriate.

[0049] In the following description, an acoustic signal included in sound information may be described, but the acoustic signal may also be referred to as a voice signal or a sound signal. In other words, in the present disclosure, the acoustic signal has the same meaning as the voice signal or the sound signal.

[0050] (Embodiment) [Overview] First, an overview of an audio reproduction system according to an embodiment will be described. Fig. 1 is a schematic diagram showing an example of use of an audio reproduction system according to an embodiment. Fig. 1 shows a user 99 using an audio reproduction system 100.

[0051] The sound reproduction system 100 shown in FIG. 1 is used simultaneously with, for example, a three-dimensional video reproduction device 300. By simultaneously viewing three-dimensional images and three-dimensional sound, the image enhances the auditory sense of realism, and the sound enhances the visual sense of realism, allowing the user to experience the image and sound as if they were actually at the scene where they were captured. For example, when an image (moving image) of people having a conversation is displayed, it is known that even if the localization of the sound image (sound source object) of the conversation sound is not aligned with the person's mouth, the user 99 will perceive the conversation sound as coming from the person's mouth. In this way, the visual information can correct the position of the sound image, and the image and sound can be combined to enhance the sense of realism.

[0052] The three-dimensional video reproduction device 300 is an image display device worn on the head of the user 99. Therefore, the three-dimensional video reproduction device 300 moves integrally with the head of the user 99. For example, as shown in the figure, the three-dimensional video reproduction device 300 is a glasses-type device that is supported by the ears and nose of the user 99.

[0053] The three-dimensional video reproduction device 300 changes the displayed image in accordance with the movement of the user 99's head, thereby making the user 99 perceive the movement of his or her head in the three-dimensional image space. In other words, when an object in the three-dimensional image space is located in front of the user 99, if the user 99 turns to the right, the object moves to the left of the user 99, and if the user 99 turns to the left, the object moves to the right of the user 99. In this way, the three-dimensional video reproduction device 300 moves the three-dimensional image space in the opposite direction to the movement of the user 99.

[0054] The 3D video playback device 300 displays two images with a parallax difference to each of the user's 99's left and right eyes. The user 99 can perceive the three-dimensional position of an object on the image based on the parallax difference between the displayed images. Note that when the user 99 uses the audio playback system 100 with their eyes closed, for example, when using it to play healing sounds for sleep induction, the 3D video playback device 300 does not need to be used at the same time. In other words, the 3D video playback device 300 is not an essential component of the present disclosure. In addition to dedicated video display devices, the 3D video playback device 300 may also be a general-purpose mobile terminal owned by the user 99, such as a smartphone or tablet device.

[0055] Such general-purpose mobile terminals are equipped with not only a display for displaying images but also various sensors for detecting the terminal's posture and movement. Furthermore, they are also equipped with a processor for information processing, and are capable of connecting to a network to transmit and receive information to and from a server device such as a cloud server. In other words, the 3D video playback device 300 and the audio playback system 100 can be realized by combining a smartphone with general-purpose headphones or the like that do not have an information processing function.

[0056] As in this example, the head movement detection function, the video presentation function, the video information processing function for presentation, the sound presentation function, and the sound information processing function for presentation may be appropriately arranged in one or more devices to realize the 3D video reproduction device 300 and the sound reproduction system 100. If the 3D video reproduction device 300 is not required, it is sufficient to appropriately arrange the head movement detection function, the sound presentation function, and the sound information processing function for presentation in one or more devices. For example, the sound reproduction system 100 can be realized by a processing device such as a computer or smartphone having a sound information processing function for presentation, and headphones or the like having a head movement detection function and a sound presentation function.

[0057] The sound reproduction system 100 is a sound presentation device that is worn on the head of the user 99. Therefore, the sound reproduction system 100 moves integrally with the head of the user 99. For example, the sound reproduction system 100 in this embodiment is a so-called over-ear headphone type device. Note that there are no particular limitations on the form of the sound reproduction system 100, and it may be, for example, two earplug-type devices that are worn independently on the left and right ears of the user 99.

[0058] The sound reproduction system 100 changes the sound presented in accordance with the movement of the head of the user 99, thereby making the user 99 perceive as if he or she is moving his or her head within the three-dimensional sound field. For this reason, as described above, the sound reproduction system 100 moves the three-dimensional sound field in the opposite direction to the movement of the user 99.

[0059] Here, when the user 99 moves within the three-dimensional sound field, the position of the sound source object changes relative to the user 99's position within the three-dimensional sound field. As a result, each time the user 99 moves, it is necessary to perform calculation processing based on the positions of the sound source object and the user 99 to generate an output sound signal for playback. Since such processing typically requires a huge amount of processing, in the present disclosure, a panning process is applied as a conversion process to reduce the amount of processing. The reproduced sound is expressed as a representative sound from a representative point. As a result, simply by convolving a head-related transfer function with the representative sound, the user 99 can perceive the reproduced sound from the sound source object. On the other hand, when panning processing is applied, the sense of sound localization may be impaired, making it difficult to perceive the position of the sound image (blurring). Therefore, when panning processing is applied to sounds that are significantly affected by the loss of the sense of localization, there is a disadvantage that the original sense of realism may be impaired. Therefore, in the present disclosure, a three-dimensional sound field is formed by applying panning processing to a portion of the audio signal and generating an output sound signal without applying panning processing to the other portion. In the following, in this embodiment, a case where panning processing is used as an example of conversion processing will be described, but the conversion processing is not limited to panning processing, and any conversion processing can be applied as long as the conversion is expected to reduce the amount of processing depending on the conditions. Also, while panning processing will be described using specific examples, the panning processing is not limited to the specific example method described below, and existing panning processing methods such as VBAP (Vector Based Amplitude Panning), DBAP (Distance Based Amplitude Panning), and Ambisonics can also be applied.

[0060] [Configuration] Next, the configuration of the sound reproduction system 100 according to this embodiment will be described with reference to Fig. 2. Fig. 2 is a block diagram showing the functional configuration of the sound reproduction system according to this embodiment.

[0061] As shown in FIG. 2, the sound reproduction system 100 according to this embodiment includes an information processing device 101, a communication module 102, a detector 103, a driver 104, and a database 105.

[0062] The information processing device 101 is an arithmetic device for performing various signal processing in the sound reproduction system 100. The information processing device 101 includes a processor and a memory, such as a computer, and is realized by the processor executing a program stored in the memory. Execution of this program provides functions related to each functional unit described below.

[0063] The information processing device 101 includes an acquisition unit 111, a path calculation unit 121, an output sound generation unit 131, and a signal output unit 141. Details of each functional unit included in the information processing device 101 will be described below together with details of the configuration other than the information processing device 101.

[0064] The communication module 102 is an interface device for accepting input of sound information to the sound reproduction system 100. The communication module 102 includes, for example, an antenna and a signal converter, and receives sound information from an external device via wireless communication. More specifically, the communication module 102 receives, using the antenna, a wireless signal indicating sound information converted into a format for wireless communication, and reconverts the wireless signal into sound information using the signal converter. In this way, the sound reproduction system 100 acquires sound information from the external device via wireless communication. The sound information acquired by the communication module 102 is acquired by the acquisition unit 111. In this way, the acquisition unit 111 is an example of a sound acquisition unit. The sound information is input to the information processing device 101 in the above manner. Note that communication between the sound reproduction system 100 and the external device may be performed via wired communication.

[0065] The sound information acquired by the sound reproduction system 100 is encoded in a predetermined format, such as MPEG-H 3D Audio (ISO / IEC 23008-3). As an example, the encoded sound information includes information about the sound reproduced by the sound reproduction system 100 and information about the localization position when the sound image of the sound is localized at a predetermined position in a three-dimensional sound field (i.e., perceived as sound coming from a predetermined direction). The sound information can also be interpreted as information about a sound source object. In other words, the sound information includes the position of the sound source object in the three-dimensional sound field and the sound produced by the sound source object. The sound information may also include a flag for determining whether or not to apply panning processing. This flag will be described later.

[0066] As described above, sound information is obtained as input data, and includes an audio signal (acoustic signal), which is information about the reproduced sound, and other information, such as information about the position of a sound source object in a three-dimensional sound field. The other information may also include information for defining a three-dimensional sound field. Therefore, the other information may be collectively referred to as information about space (spatial information), including information about the position of a sound source object and information for defining a three-dimensional sound field. When the audio signal is viewed as the main focus, the input data can be said to be sound information in which other information (metadata) is attached to the audio signal. When the spatial information is viewed as the main focus, the input data can be said to be information in which the audio signal is attached to the spatial information. Alternatively, since the input data has both aspects, the input data can also be considered as sound spatial information.

[0067] As a specific example, the sound information includes information about a plurality of sounds including a first reproduced sound and a second reproduced sound, and the sound images of the respective sounds when reproduced are localized so that they are perceived as coming from different positions in a three-dimensional sound field. Therefore, the sound source object of the first reproduced sound is localized at a first position in the three-dimensional sound field, and the sound source object of the second reproduced sound is localized at a second position in the three-dimensional sound field. Thus, the sound information may include a plurality of sounds. In other words, the sound information may include a plurality of audio signals corresponding to the first reproduced sound and the second reproduced sound, respectively, and the positions of a plurality of sound source objects at first and second positions that correspond one-to-one to the plurality of audio signals.

[0068] 3 is a diagram illustrating an example of an audio signal according to an embodiment. For example, as shown in (a) of FIG. 3, the sound information may include an audio signal of a first direct sound arriving from a first position (from a first direction) to the position of the user 99 and an audio signal of a second direct sound arriving from a second position (from a second direction) to the position of the user 99. The sound information immediately after acquisition may include only information about the reproduced sound. In this case, information about a predetermined position may be acquired separately, and the subsequent processing may be performed once the information is collected. As described above, the sound information includes first sound information about the first reproduced sound and second sound information about the second reproduced sound. However, multiple pieces of sound information each including these may be acquired separately and played back simultaneously (i.e., treated as a single piece of sound information), thereby localizing sound images at different positions in a three-dimensional sound field and causing the reproduced sounds to arrive from different directions.

[0069] Alternatively, the sound information may include multiple audio signals and the position of a single sound source object that corresponds to the multiple audio signals in a many-to-one relationship. For example, such sound information is used in a situation where multiple sounds are reproduced from a sound source object. For example, each of the multiple audio signals corresponds to a direct sound that arrives directly from the position of the sound source object to the position of the user 99, and a secondary sound (sound generated by indirect propagation) that accompanies the direct sound and arrives via a path different from that of the direct sound.

[0070] For example, as shown in (b) of FIG. 3 , the sound information immediately after acquisition includes an audio signal related to the direct sound, and is converted into sound information including audio signals such as reverberation, primary reflection, and diffraction through a conversion process that calculates secondary sounds. This conversion process that calculates the secondary sounds uses information on the spatial environment conditions of the three-dimensional sound field (e.g., the position, reflection, and diffraction characteristics of objects in the three-dimensional sound field). Thus, secondary sounds are computationally generated from sound information related to a single reproduced sound based on the spatial environment conditions of the three-dimensional sound field. Therefore, the secondary sounds are not included in the sound information immediately after acquisition, and sound information including these secondary sounds is generated through the conversion process that calculates the secondary sounds. Another secondary sound may be generated from one secondary sound through its propagation. Note that the information on the spatial environment conditions is part of the spatial information and may be acquired together with the audio signal from the input sound information. Alternatively, the audio signal and the spatial information may be acquired separately. That is, the sound information may be acquired from a single file or bitstream, or may be acquired separately by dividing it into multiple files or bitstreams. For example, the audio signal and the spatial information may be acquired from separate files or bitstreams, or the audio signal and the spatial information may each be acquired from a plurality of files or bitstreams.

[0071] The example in FIG. 3B illustrates that a secondary reflected sound is generated from a primary reflected sound. As shown in the figure, these secondary sounds are assigned tags that identify their relationships, such as parent, child, and grandchild, as information regarding the genealogy of their generation from the direct sound (in other words, the generation system). Here, the information regarding the generation system can be rephrased as information assigned to each sound to identify the type of sound based on the generation system. For example, if the direct sound is the parent, the primary reflected sound is the child, and the secondary reflected sound is the grandchild. Conversely, if the secondary reflected sound has a grandchild relationship, the parent-related sound contains the direct sound, and the child-related sound contains the primary reflected sound. In other words, the information regarding the generation system can be used to identify the type of sound. In other words, the information regarding the generation system identifies which of multiple types of sound the sound belongs to, such as direct sound, primary reflected sound, secondary reflected sound, reverberant sound, and diffracted sound. Alternatively, when the direct sound is defined as generation 0, the information may be quantified as a generation number, such as the first generation to which the first reflected sound belongs, the second generation to which the second reflected sound belongs, etc. Such information on the generation system is used when determining whether or not to apply panning processing to the audio signal. Note that the number of generations of the direct sound that are allowed to be generated may be set according to the scale of the calculation resources.

[0072] As described above, there is no particular limitation on the form of the input sound information, as long as the acquisition unit 111 corresponding to various forms of sound information is provided in the sound reproduction system 100. Before or after being input, the sound information is in a state including a plurality of sound signals in order to distinguish at least sound signals to which panning processing is applied from sound signals to which no panning processing is applied.

[0073] An example of the acquisition unit 111 will now be described with reference to Fig. 4. Fig. 4 is a block diagram showing the functional configuration of the acquisition unit according to the embodiment. As shown in Fig. 4, the acquisition unit 111 in this embodiment includes, for example, an encoded sound information input unit 112, a decoding processing unit 113, and a sensing information input unit 114.

[0074] The encoded sound information input unit 112 is a processing unit to which the coded (in other words, encoded) sound information acquired by the acquisition unit 111 is input. The encoded sound information input unit 112 outputs the input sound information to the decoding processing unit 113. The decoding processing unit 113 is a processing unit that decodes (in other words, decodes) the sound information output from the encoded sound information input unit 112 to generate a playback sound and the position of a sound source object included in the sound information in a format that is used for subsequent processing. The sensing information input unit 114 will be described below together with the function of the detector 103.

[0075] The detector 103 is a device for detecting the speed of movement of the head of the user 99. The detector 103 is configured by combining various sensors used for detecting movement, such as a gyro sensor and an acceleration sensor. In this embodiment, the detector 103 is built into the sound reproduction system 100. However, the detector 103 may also be built into an external device, such as a three-dimensional image reproduction device 300 that moves in accordance with the movement of the head of the user 99, similar to the sound reproduction system 100. In this case, the detector 103 does not need to be included in the sound reproduction system 100. Alternatively, the detector 103 may be an external imaging device or the like that captures an image of the moving user 99 and detects the movement of the user 99 by processing the captured image.

[0076] The detector 103 is, for example, fixed integrally to the housing of the sound reproduction system 100 and detects the speed of movement of the housing. After the sound reproduction system 100 including the housing is worn by the user 99, it moves integrally with the head of the user 99, and as a result, the detector 103 can detect the speed of movement of the head of the user 99.

[0077] The detector 103 may detect, for example, the amount of head movement of the user 99 as the amount of rotation about at least one of three axes that are orthogonal to each other in three-dimensional space as the rotation axis, or may detect the amount of displacement about at least one of the three axes as the displacement direction. Furthermore, the detector 103 may detect both the amount of rotation and the amount of displacement as the amount of head movement of the user 99.

[0078] The sensing information input unit 114 acquires the movement speed of the user 99's head from the detector 103. More specifically, the sensing information input unit 114 acquires the amount of head movement of the user 99 detected by the detector 103 per unit time as the movement speed. In this way, the sensing information input unit 114 acquires at least one of the rotation speed and the displacement speed from the detector 103. The amount of head movement of the user 99 acquired here is used to determine the position and posture (in other words, coordinates and orientation) of the user 99 in the three-dimensional sound field. Therefore, the acquisition unit 111 also functions as a position acquisition unit via the sensing information input unit 114. In the sound reproduction system 100, the relative position of the sound source object with respect to the user 99 is determined based on the determined coordinates and orientation of the user 99, and sound is reproduced. Specifically, the above functions are realized by the path calculation unit 121 and the output sound generation unit 131.

[0079] The path calculation unit 121 includes an arrival direction calculation function that calculates the relative arrival direction of the reproduced sound from the position of the sound source object to the position of the user 99 based on the determined coordinates and orientation of the user 99, and a function that performs the conversion process to calculate the secondary sound described above. Therefore, the path calculation unit 121 includes a function that calculates a propagation path from the sound source object and calculates the secondary sound and the arrival direction of the secondary sound that arrives at the position of the user 99 through indirect propagation of the reproduced sound according to the calculated propagation path of the reproduced sound. The arrival direction of the secondary sound includes additional information, such as what object the sound will be reflected from in the case of a reflected sound and the attenuation rate upon reflection. The additional information is included in the arrival direction of the secondary sound calculated from the input sound information. In other words, the additional information is computationally generated and acquired from the sound information.

[0080] To summarize the spatial information, the spatial information includes further information such as the spatial position of a sound source object in a space (three-dimensional sound field) (information on the position of the sound source object), sound reflection and diffraction characteristics at the sound source object (together with information on the conditions of the spatial environment), and the size of the three-dimensional sound field. Based on the spatial information, the path calculation unit 121 generates secondary sounds depending on which sound source object the reproduced sound is reflected or diffracted by, and calculates, as additional information, the direction from which the secondary sounds arrive and the volume of the secondary sounds after attenuation by reflection or diffraction. The sound information (input data) includes spatial information in the form of audio signals and accompanying metadata. As described above, the spatial information includes, as information other than the audio signal, information necessary to turn the sound into stereophonic sound and position the sound source object in a three-dimensional sound field, and / or information used to calculate information necessary to turn the sound into stereophonic sound and position the sound source object in a three-dimensional sound field.

[0081] The path calculation unit 121 may be realized by any processing as long as it can calculate the arrival direction of the reproduced sound when the reproduced sound reaches the user as a direct sound and can calculate the arrival direction of a secondary sound that arrives at the position of the user 99 due to secondary propagation of the reproduced sound. The path calculation unit 121 determines from which direction in the three-dimensional sound field the reproduced sound and the secondary sound are to be perceived by the user 99 as coming from, based on the coordinates and orientation of the user 99, and processes the sound information so that the output sound signal is perceived as such a sound when it is reproduced.

[0082] The output sound generating unit 131 is a processing unit that processes information about the reproduced sound included in the sound information to generate an output sound signal.

[0083] An example of the output sound generation unit 131 will now be described with reference to Fig. 5. Fig. 5 is a block diagram showing the functional configuration of the output sound generation unit according to an embodiment. As shown in Fig. 5, the output sound generation unit 131 in this embodiment includes, for example, a target determination unit 132, a first generation unit 133, a second generation unit 134, and a synthesis unit 135.

[0084] The target determination unit 132 is a processing unit for determining, for each of a plurality of audio signals included in the sound information, whether to use the first generation unit 133 or the second generation unit 134. In other words, the target determination unit 132 determines, for each individual audio signal, whether to use the audio signal to generate a first sound signal by the first generation unit 133 or to generate a second sound signal by the second generation unit 134. For this reason, the target determination unit 132 has a function of acquiring information for determining whether to use the first generation unit 133 or the second generation unit 134.

[0085] The first generation unit 133 is a processing unit used when a head-related transfer function is directly convolved with a reproduced sound without applying panning processing. The first generation unit 133 is a processing unit used when generating an output sound signal in a so-called "normal mode." The first generation unit 133 acquires a reproduced sound and a head-related transfer function corresponding to the direction from which the reproduced sound arrives, and performs convolution processing of the acquired head-related transfer function with the reproduced sound to generate a first sound signal.

[0086] The second generation unit 134 is a processing unit used when applying a panning process to perform a conversion process of converting a playback sound into a representative sound, and then convolving a head-related transfer function with the converted representative sound. The second generation unit 134 is a processing unit used when generating an output sound signal in a so-called "low processing mode". The second generation unit 134 acquires the playback sound and the position of the representative point, and performs a conversion process to a representative sound for reproducing the playback sound using sound from the representative point.

[0087] For example, if a sound source object is located midway between two representative points, a sound is generated so that the same sound as the playback sound is emitted from each of the two representative points. In other words, the playback sound is distributed to the two representative points. Then, a representative sound can be generated by adjusting the gain of the generated sound to match the position of the sound source object. Conversion of the playback sound to a representative sound is not limited to this example. For example, as described below, conversion of the playback sound to a representative sound may be performed by time shift adjustment and gain adjustment, or any other existing conversion may be used as long as the playback sound can be converted to a representative sound that is reproduced as a sound from a representative point. Furthermore, in this specification, the conversion process of the playback sound to a representative sound may be interpreted as a process of distributing the playback sound to representative points (representative directions). Specifically, sound signals of the playback sound associated with the positions of each sound source object are distributed to the positions of the representative points, and a representative sound arriving from the representative points (representative directions) to the listener is generated. Here, the representative direction refers to the direction of the representative point as seen from the listener, or the direction of the listener as seen from the representative point. An example of the conversion that performs the time shift adjustment and the gain adjustment will be described later. The second generation unit 134 acquires the same number of representative sounds as the number of representative points obtained by the conversion and head-related transfer functions corresponding to the representative directions from each representative point to the position of the user 99, and performs a convolution process of the acquired head-related transfer functions on the representative sounds to generate a second sound signal.

[0088] The synthesis unit 135 synthesizes the generated first sound signal and second sound signal by adding them together to generate an output sound signal. The synthesis unit 135 may also perform EQ adjustment on the second sound signal before synthesizing it with the first sound signal. Specifically, the synthesis unit 135 may perform EQ adjustment to increase the gain of a high-frequency domain that is likely to be attenuated in the panning process, thereby emphasizing this high-frequency domain. Therefore, the synthesis unit 135 also functions as an EQ adjustment unit. When there are multiple second sound signals, the EQ adjustment performed by the synthesis unit 135 may be performed on only some or all of the multiple second sound signals.

[0089] Referring again to FIG. 2 , the output sound generation unit 131 acquires a head-related transfer function used to generate an output sound signal from the database 105. The database 105 is an information storage device that functions both as a storage device for storing information and as a storage controller that reads out the stored information and outputs it to an external component. The database 105 stores a head-related transfer function for each direction of arrival of the sound from the user 99. The head-related transfer functions included in the database 105 are a set of general-purpose head-related transfer functions that can be used by everyone, a set of head-related transfer functions optimized for each individual user 99, or a set of publicly available head-related transfer functions. The database 105 receives an inquiry from the output sound generation unit 131 using the direction of arrival as a query, and outputs a head-related transfer function corresponding to the direction of arrival to the output sound generation unit 131.

[0090] The signal output unit 141 is a functional unit that outputs the generated output sound signal to the driver 104. The signal output unit 141 generates a waveform signal by performing signal conversion from a digital signal to an analog signal based on the output sound signal, and then causes the driver 104 to generate sound waves based on the waveform signal, thereby presenting sound to the user 99. The driver 104 includes, for example, a diaphragm and a drive mechanism such as a magnet and a voice coil. The driver 104 operates the drive mechanism in response to the waveform signal, causing the drive mechanism to vibrate the diaphragm. In this way, the driver 104 generates sound waves by vibrating the diaphragm in response to the output sound signal (this means "reproducing" the output sound signal; in other words, "reproducing" does not include perception by the user 99). The sound waves propagate through the air and reach the ears of the user 99, and the user 99 perceives the sound.

[0091] [Another Configuration Example] In the above example, the sound reproduction system 100 according to the present embodiment is a sound presentation device, and has been described as including an information processing device 101, a communication module 102, a detector 103, a database 105, and a driver 104. However, the functions of the sound reproduction system 100 may be realized by a plurality of devices or by a single device. Specific examples will be described using Figures 6 to 15. Figures 6 to 15 are diagrams for explaining other examples of the sound reproduction system according to the embodiment.

[0092] For example, the information processing device 601 may be included in the audio presentation device 602, and the audio presentation device 602 may perform both acoustic processing and sound presentation. Alternatively, the information processing device 601 and the audio presentation device 602 may share the acoustic processing described in the present disclosure, or a server connected to the information processing device 601 or the audio presentation device 602 via a network may perform part or all of the acoustic processing described in the present disclosure.

[0093] In the above explanation, the information processing device 601 is referred to as such, but when the information processing device 601 performs acoustic processing by decoding a bitstream generated by encoding at least a portion of the data of an audio signal or spatial information used for acoustic processing, the information processing device 601 may be referred to as a decoding device, and the acoustic reproduction system 100 (i.e., the stereophonic sound reproduction system 600 in the figure) may be referred to as a decoding processing system.

[0094] Here, an example will be described in which the sound reproduction system 100 functions as a decoding processing system.

[0095] <Example of Encoding Device> FIG. 7 is a functional block diagram showing the configuration of an encoding device 700 that is an example of an encoding device according to the present disclosure.

[0096] Input data 701 is data to be coded, including spatial information and / or an audio signal, that is input to an encoder 702. Details of the spatial information will be described later.

[0097] The encoder 702 encodes the input data 701 to generate encoded data 703. The encoded data 703 is, for example, a bit stream generated by the encoding process.

[0098] The memory 704 stores the encoded data 703. The memory 704 may be, for example, a hard disk or a solid-state drive (SSD), or may be any other storage device.

[0099] In the above description, a bitstream generated by an encoding process is given as an example of the encoded data 703 stored in memory 704. However, data other than a bitstream may also be used. For example, the encoding device 700 may convert a bitstream into a predetermined data format and store the converted data in memory 704. The converted data may be, for example, a file or multiplexed stream storing one or more bitstreams. Here, the file may have a file format such as ISOBMFF (ISO Base Media File Format). The encoded data 703 may also be in the form of multiple packets generated by dividing the bitstream or file. When converting the bitstream generated by the encoder 702 into data other than the bitstream, the encoding device 700 may include a conversion unit (not shown), or the conversion process may be performed by a CPU (Central Processing Unit).

[0100] <Example of Decoding Device> FIG. 8 is a functional block diagram showing the configuration of a decoding device 800 that is an example of a decoding device according to the present disclosure.

[0101] The memory 804 stores, for example, the same data as the coded data 703 generated by the coding device 700. The memory 804 reads out the stored data and inputs it as input data 803 to the decoder 802. The input data 803 is, for example, a bitstream to be decoded. The memory 804 may be, for example, a hard disk or an SSD, or may be another storage device.

[0102] Note that the decoding device 800 may not use the data stored in the memory 804 as input data 803 as is, but may convert the read data and generate converted data as input data 803. The data before conversion may be, for example, multiplexed data storing one or more bitstreams. Here, the multiplexed data may be, for example, a file having a file format such as ISOBMFF. The data before conversion may also be in the form of multiple packets generated by dividing the bitstream or file. When converting data different from the bitstream read from the memory 804 into a bitstream, the decoding device 800 may be provided with a conversion unit (not shown), or the conversion process may be performed by a CPU.

[0103] Decoder 802 decodes input data 803 to generate audio signal 801 that is presented to a listener.

[0104] <Another Example of Encoding Device> Fig. 9 is a functional block diagram showing the configuration of an encoding device 900, which is another example of an encoding device of the present disclosure. In Fig. 9, components having the same functions as those in Fig. 7 are assigned the same reference numerals as those in Fig. 7, and descriptions of these components will be omitted.

[0105] The coding device 900 differs from the coding device 700 in that the coding device 700 includes a memory 704 for storing coded data 703, whereas the coding device 900 includes a transmitting unit 901 for transmitting coded data 703 to the outside.

[0106] The transmitter 901 transmits a transmission signal 902 to another device or a server based on the encoded data 703 or data in another data format generated by converting the encoded data 703. The data used to generate the transmission signal 902 is, for example, the bit stream, multiplexed data, file, or packet described in the encoding device 700.

[0107] 10 is a functional block diagram showing the configuration of a decoding device 1000, which is another example of a decoding device according to the present disclosure. In Fig. 10, components having the same functions as those in Fig. 8 are assigned the same reference numerals, and description of these components will be omitted.

[0108] The decoding device 1000 differs from the decoding device 800 in that the decoding device 800 includes a memory 804 for reading out input data 803, whereas the decoding device 1000 includes a receiving unit 1001 for receiving input data 803 from the outside.

[0109] The receiving unit 1001 receives a received signal 1002, acquires received data, and outputs input data 803 to be input to the decoder 802. The received data may be the same as the input data 803 to be input to the decoder 802, or may be data in a data format different from that of the input data 803. If the received data is data in a data format different from that of the input data 803, the receiving unit 1001 may convert the received data into the input data 803, or a conversion unit or CPU (not shown) included in the decoding device 1000 may convert the received data into the input data 803. The received data is, for example, a bit stream, multiplexed data, a file, or a packet, as described in the encoding device 900.

[0110] <Functional Description of Decoder> FIG. 11 is a functional block diagram showing the configuration of a decoder 1100, which is an example of the decoder 802 in FIG. 8 or 10.

[0111] The input data 803 is an encoded bitstream, and includes encoded audio data, which is an encoded audio signal, and metadata used for acoustic processing.

[0112] The spatial information management unit 1101 acquires metadata included in the input data 803 and analyzes the metadata. The metadata includes information describing elements that act on sounds arranged in a sound space. The spatial information management unit 1101 manages spatial information necessary for acoustic processing obtained by analyzing the metadata and provides the spatial information to the rendering unit 1103. Note that, although the information used for acoustic processing is referred to as spatial information in this disclosure, it may be called by other names. The information used for acoustic processing may be called, for example, sound space information or scene information. Furthermore, when the information used for acoustic processing changes over time, the spatial information input to the rendering unit 1103 may be called a space state, a sound space state, a scene state, or the like.

[0113] Furthermore, spatial information may be managed for each sound space or for each scene. For example, when different rooms are represented as virtual spaces, each room may be managed as a scene of a different sound space, or even if the same space is represented, spatial information may be managed as different scenes depending on the situation being represented. In managing spatial information, an identifier for identifying each piece of spatial information may be assigned. The spatial information data may be included in a bitstream, which is one form of input data 803, or the bitstream may include an identifier for the spatial information, and the spatial information data may be acquired from a source other than the bitstream. When the bitstream includes only the identifier for the spatial information, the identifier for the spatial information may be used during rendering to acquire the spatial information data stored in the memory of the acoustic signal processing device or an external server as input data.

[0114] Note that the information managed by the spatial information management unit 1101 is not limited to information included in the bitstream. For example, the input data 803 may include data indicating the characteristics or structure of a space acquired from a software application or server providing VR or AR, as data not included in the bitstream. Furthermore, for example, the input data 803 may include data indicating the characteristics or position of a listener or object, as data not included in the bitstream. Furthermore, the input data 803 may include, as information indicating the position of the listener, information acquired by a sensor provided in a terminal including a decoding device, or information indicating the position of the terminal estimated based on information acquired by the sensor. In other words, the spatial information management unit 1101 may communicate with an external system or server to acquire spatial information and the position of the listener. Furthermore, the spatial information management unit 1101 may acquire clock synchronization information from an external system and execute a process of synchronizing with the clock of the rendering unit 1103. Note that the space in the above description may be a virtually formed space, i.e., a VR space, or may be a real space or a virtual space corresponding to a real space, i.e., an AR space or an MR (Mixed Reality) space. The virtual space may also be called a sound field or a sound space. Furthermore, the information indicating a position in the above description may be information such as coordinate values ​​indicating a position within a space, information indicating a relative position with respect to a predetermined reference position, or information indicating the movement or acceleration of a position within a space.

[0115] The audio data decoder 1102 decodes the encoded audio data included in the input data 803 to obtain an audio signal.

[0116] The encoded audio data acquired by the stereophonic sound reproduction system 600 is a bitstream encoded in a predetermined format such as MPEG-H 3D Audio (ISO / IEC 23008-3). Note that MPEG-H 3D Audio is merely one example of an encoding method that can be used to generate the encoded audio data included in the bitstream, and the encoded audio data may be included in a bitstream encoded in another encoding method. For example, the encoding method used may be a lossy codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), or Vorbis, or a lossless codec such as ALAC (Apple Lossless Audio Codec) or FLAC (Free Lossless Audio Codec), or any other encoding method may be used. For example, PCM (Pulse Code Modulation) data may be a type of encoded audio data. In this case, the decoding process may be, for example, a process of converting an N-bit binary number into a number format (e.g., floating-point format) that can be processed by the rendering unit 1103, where the number of quantization bits of the PCM data is N.

[0117] The rendering unit 1103 receives an audio signal and spatial information as input, performs acoustic processing on the audio signal using the spatial information, and outputs an audio signal 801 after acoustic processing.

[0118] Before starting rendering, the spatial information management unit 1101 reads metadata of the input signal, detects rendering items such as objects or sounds defined in the spatial information, and transmits the detected items to the rendering unit 1103. After starting rendering, the spatial information management unit 1101 grasps changes over time in the spatial information and the position of the listener, and updates and manages the spatial information. The spatial information management unit 1101 then transmits the updated spatial information to the rendering unit 1103. The rendering unit 1103 generates and outputs an audio signal to which acoustic processing has been applied based on the audio signal included in the input data and the spatial information received from the spatial information management unit 1101.

[0119] The spatial information update process and the audio signal output process with added acoustic processing may be executed in the same thread, or the spatial information management unit 1101 and the rendering unit 1103 may be allocated to independent threads. When the spatial information update process and the audio signal output process with added acoustic processing are executed in different threads, the thread startup frequency may be set individually, or the processes may be executed in parallel.

[0120] By having the spatial information management unit 1101 and the rendering unit 1103 execute their processes in different, independent threads, computational resources can be preferentially allocated to the rendering unit 1103. Therefore, in the case of sound output processing in which even a slight delay cannot be tolerated, for example, sound output processing in which a delay of even one sample (0.02 msec) would cause a popping noise, can be safely performed. In this case, the allocation of computational resources to the spatial information management unit 1101 is limited. However, compared to audio signal output processing, updating spatial information is a less frequent process (e.g., a process such as updating the listener's facial orientation). Therefore, unlike audio signal output processing, updating spatial information does not necessarily require an instantaneous response, and therefore limiting the allocation of computational resources does not significantly affect the acoustic quality provided to the listener.

[0121] The space information may be updated periodically at preset times or intervals, or when preset conditions are met. The space information may be updated manually by a listener or a sound space manager, or may be triggered by a change in an external system. For example, when a listener operates a controller to instantly warp the position of their avatar, instantly advance or rewind the time, or when a virtual space manager suddenly changes the environment of the space, the thread in which the space information management unit 1101 is located may be started as a one-off interrupt process in addition to being started periodically.

[0122] The role of the information update thread that executes the spatial information update process is, for example, to update the position or orientation of the listener's avatar placed in the virtual space based on the position or orientation of the VR goggles worn by the listener, and to update the position of objects moving in the virtual space. These tasks are handled within a processing thread that runs relatively infrequently, on the order of several tens of Hz. Processing to reflect the properties of direct sound may be performed in such an infrequently occurring processing thread. This is because the properties of direct sound change less frequently than the frequency of audio processing frames for audio output. Doing so can relatively reduce the computational load of the process and also avoid the risk of pulsive noise occurring when information is updated at an unnecessarily fast frequency.

[0123] FIG. 12 is a functional block diagram showing the configuration of a decoder 1200, which is another example of the decoder 802 in FIG. 8 or 10.

[0124] Figure 12 differs from Figure 11 in that the input data 803 includes an unencoded audio signal rather than encoded audio data. The input data 803 includes a bitstream including metadata and an audio signal.

[0125] The spatial information management unit 1201 is the same as the spatial information management unit 1101 in FIG. 11, and therefore a description thereof will be omitted.

[0126] The rendering unit 1202 is the same as the rendering unit 1103 in FIG. 11, and therefore a description thereof will be omitted.

[0127] In the above description, the configuration in Fig. 12 is called a decoder, but it may also be called an audio processing unit that performs audio processing. Furthermore, a device including an audio processing unit may also be called an audio processing device rather than a decoding device. Furthermore, the audio signal processing device (information processing device 601) may also be called an audio processing device.

[0128] <Physical Configuration of Encoding Apparatus> Fig. 13 is a diagram showing an example of the physical configuration of an encoding apparatus. The encoding apparatus shown in Fig. 13 is an example of the encoding apparatuses 700 and 900 described above.

[0129] The encoding device of FIG. 13 includes a processor, a memory, and a communication IF.

[0130] The processor may be, for example, a CPU (Central Processing Unit), a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit), and the CPU, DSP, or GPU may execute a program stored in a memory to perform the encoding process of the present disclosure. Alternatively, the processor may be a dedicated circuit that performs signal processing on an audio signal, including the encoding process of the present disclosure.

[0131] The memory may be configured, for example, by RAM (Random Access Memory) or ROM (Read Only Memory). The memory may also include a magnetic storage medium such as a hard disk or a semiconductor memory such as an SSD (Solid State Drive). The term "memory" may also refer to an internal memory built into a CPU or GPU.

[0132] The communication IF (Interface) is a communication module compatible with a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The encoding device has a function of communicating with other communication devices via the communication IF and transmits an encoded bitstream.

[0133] The communication module is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. In the above example, Bluetooth (registered trademark) or WIGIG (registered trademark) is used as the communication method, but other communication methods such as LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark) may also be supported. Furthermore, the communication IF may be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface) instead of the wireless communication method described above.

[0134] <Physical configuration of audio signal processing device> Fig. 14 is a diagram showing an example of the physical configuration of an audio signal processing device. Note that the audio signal processing device in Fig. 14 may be a decoding device. Furthermore, part of the configuration described here may be provided in the audio presentation device 602. Furthermore, the audio signal processing device shown in Fig. 14 is an example of the audio signal processing device 601 described above.

[0135] The acoustic signal processing device of FIG. 14 includes a processor, a memory, a communication IF, a sensor, and a speaker.

[0136] The processor may be, for example, a CPU (Central Processing Unit), a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit), and the CPU, DSP, or GPU may execute a program stored in a memory to perform the audio processing or decoding processing of the present disclosure. Alternatively, the processor may be a dedicated circuit that performs signal processing on audio signals, including the audio processing of the present disclosure.

[0137] The memory may be configured, for example, by RAM (Random Access Memory) or ROM (Read Only Memory). The memory may also include a magnetic storage medium such as a hard disk or a semiconductor memory such as an SSD (Solid State Drive). The term "memory" may also refer to an internal memory built into a CPU or GPU.

[0138] The communication IF (Interface) is a communication module compatible with a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The acoustic signal processing device shown in Fig. 14 has a function of communicating with other communication devices via the communication IF and acquires a bitstream to be decoded. The acquired bitstream is stored in a memory, for example.

[0139] The communication module is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. In the above example, Bluetooth (registered trademark) or WIGIG (registered trademark) is used as the communication method, but other communication methods such as LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark) may also be supported. Furthermore, the communication IF may be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface) instead of the wireless communication method described above.

[0140] The sensor performs sensing to estimate the position or orientation of the listener. Specifically, the sensor estimates the position and / or orientation of the listener based on one or more detection results of the position, orientation, movement, velocity, angular velocity, or acceleration of a part or the entire body of the listener, such as the head, and generates position information indicating the position and / or orientation of the listener. Note that the position information may be information indicating the position and / or orientation of the listener in real space, or information indicating a displacement of the position and / or orientation of the listener based on the position and / or orientation of the listener at a predetermined time. Furthermore, the position information may be information indicating the position and / or orientation relative to the stereophonic sound reproduction system or an external device equipped with the sensor.

[0141] The sensor may be, for example, an imaging device such as a camera or a ranging device such as LiDAR (Light Detection and Ranging), and may capture an image of the listener's head movement and detect the movement of the listener's head by processing the captured image. Alternatively, the sensor may be a device that performs position estimation using wireless signals of any frequency band, such as millimeter waves.

[0142] The audio signal processing device shown in Fig. 14 may acquire position information from an external device equipped with a sensor via a communication IF. In this case, the audio signal processing device does not need to include a sensor. Here, the external device is, for example, the audio presentation device 602 described in Fig. 6 or a 3D video playback device worn on the listener's head. In this case, the sensor is configured by combining various sensors such as a gyro sensor and an acceleration sensor.

[0143] The sensor may, for example, detect the angular velocity of rotation around at least one of three mutually perpendicular axes in the sound space as the axis of rotation as the speed of movement of the listener's head, or may detect the acceleration of displacement with at least one of the three axes as the direction of displacement.

[0144] For example, the sensor may detect the amount of rotation about at least one of three mutually orthogonal axes in the sound space as the rotation axis, or the amount of displacement about at least one of the three axes as the displacement direction, as the amount of movement of the listener's head. Specifically, the sensor detects 6 DoF (position (x, y, z) and angle (yaw, pitch, roll)) as the position of the listener. The sensor is configured by combining various sensors used for detecting movement, such as a gyro sensor and an acceleration sensor.

[0145] The sensor may be any device capable of detecting the position of the listener, such as a camera or a GPS (Global Positioning System) receiver. Alternatively, the sensor may use location information obtained by performing self-position estimation using LiDAR (Laser Imaging Detection and Ranging). For example, when the audio signal reproduction system is implemented by a smartphone, the sensor is built into the smartphone.

[0146] The sensor may also include a temperature sensor such as a thermocouple that detects the temperature of the acoustic signal processing device shown in Figure 14, and a sensor that detects the remaining charge of a battery provided in or connected to the acoustic signal processing device.

[0147] A speaker has, for example, a diaphragm, a drive mechanism such as a magnet or a voice coil, and an amplifier, and presents an audio signal after acoustic processing to a listener as sound. The speaker operates the drive mechanism in response to the audio signal (more specifically, a waveform signal indicating the waveform of the sound) amplified by the amplifier, and the drive mechanism vibrates the diaphragm. In this way, the diaphragm vibrating in response to the audio signal generates sound waves, which propagate through the air and reach the listener's ears, causing the listener to perceive the sound.

[0148] Note that, although the description has been given here of an example in which the acoustic signal processing device shown in FIG. 14 includes a speaker and presents an audio signal after acoustic processing via the speaker, the means for presenting the audio signal is not limited to the above configuration. For example, the audio signal after acoustic processing may be output to an external audio presentation device 602 connected via a communication module. Communication via the communication module may be wired or wireless. As another example, the acoustic signal processing device shown in FIG. 14 may include a terminal for outputting an analog audio signal, and a cable such as earphones may be connected to the terminal to present the audio signal from the earphones. In the above case, the audio signal is reproduced by headphones, earphones, a head-mounted display, a neck speaker, a wearable speaker, a surround speaker composed of multiple fixed speakers, or the like, which are worn on the head or part of the body of the listener, which is the audio presentation device 602.

[0149] <Functional Description of Rendering Unit> FIG. 15 is a functional block diagram showing an example of the detailed configuration of the rendering units 1103 and 1202 in FIGS.

[0150] The rendering unit is composed of an analysis unit and a synthesis unit (different from the synthesis unit 135), and applies acoustic processing to the sound data contained in the input signal and outputs the result.

[0151] The information contained in the input signal will now be described.

[0152] The input signal may be composed of, for example, spatial information, sensor information, and sound data. The input signal may also include a bitstream composed of sound data and metadata (control information), in which case the metadata may include spatial information.

[0153] Spatial information is information about the sound space (three-dimensional sound field) created by the stereophonic playback system, and is composed of information about the objects included in the sound space and information about the listener. Objects include sound source objects that emit sound and act as sound sources, and non-sound-emitting objects that do not emit sound. Non-sound-emitting objects function as obstacle objects that reflect sounds emitted by sound source objects, but sound source objects may also function as obstacle objects that reflect sounds emitted by other sound source objects.

[0154] Information commonly assigned to sound source objects and non-sound generating objects includes position information, shape information, and the rate of attenuation of the volume when the object reflects sound.

[0155] The position information is expressed as coordinate values ​​on three axes, for example, the X-axis, Y-axis, and Z-axis, in Euclidean space, but it does not necessarily have to be three-dimensional information. For example, it may be two-dimensional information expressed as coordinate values ​​on two axes, the X-axis and the Y-axis. The position information of an object is determined by a representative position of a shape expressed by a mesh or voxels.

[0156] The shape information may include information about the surface material.

[0157] The information may also include information indicating whether the object belongs to a living thing, information indicating whether the object is a moving object, etc. If the object is a moving object, the position information may change over time, and the changed position information or the amount of change is transmitted to the rendering unit.

[0158] The information about the sound source object includes the information commonly given to the sound source object and the non-sound generating object, as well as sound data and information required to radiate the sound data into the sound space.

[0159] The sound data is data that represents the sound perceived by a listener, including information about the frequency and intensity of the sound. The sound data is typically a PCM signal, but may also be data compressed using an encoding method such as MP3. In this case, the signal must be decoded at least before it reaches the synthesis unit, so the rendering unit may include a decoding unit (not shown). Alternatively, the signal may be decoded by the audio data decoder 1102.

[0160] At least one piece of sound data may be set for one sound source object, but multiple pieces of sound data may also be set. Furthermore, identification information for identifying each piece of sound data may be assigned, and the identification information for the sound data may be stored as information about the sound source object.

[0161] The information necessary for radiating sound data into a sound space may include, for example, information on a reference volume that serves as a reference when playing back sound data, information indicating the properties (also called characteristics) of the sound data, information on the position of the sound source object, information on the orientation of the sound source object, information on the directivity of the sound emitted by the sound source object, etc. The information on the reference volume is, for example, the effective value of the amplitude value of the sound data at the sound source position when radiating the sound data into a sound space, and may be expressed as a floating-point decibel (dB) value.

[0162] For example, if the reference volume is 0 dB, it may indicate that sound is to be emitted into the sound space from the position indicated by the information about the position at the same volume as the signal level indicated by the sound data, without increasing or decreasing the volume, or if it is -6 dB, it may indicate that sound is to be emitted into the sound space from the position indicated by the information about the position with the volume of the signal level indicated by the sound data reduced to about half. These pieces of information are assigned to one piece of sound data or to multiple pieces of sound data collectively.

[0163] The information indicating the properties of the sound data may be, for example, information regarding the volume of the sound source, and may be information indicating time-series fluctuations. For example, if the sound space is a virtual conference room and the sound source is a speaker, the volume will transition intermittently over a short period of time. To put it more simply, this can be said to be alternating between sound and silence.

[0164] If the sound space is a concert hall and the sound source is a performer, the volume is maintained for a certain period of time. If the sound space is a battlefield and the sound source is an explosive, the volume of the explosion increases for a moment and then remains silent. In this way, the volume information of the sound source includes not only information about the volume of the sound but also information about the transition of the volume of the sound, and such information may be used as information indicating the properties of the sound data.

[0165] Here, the information on the transition in loudness of a sound may be data showing frequency characteristics in a time series. It may be data showing the duration of a sound section. It may be data showing a time series of the duration of a sound section and the duration of a silent section. It may be data listing multiple sets of data on the duration for which the amplitude of a sound signal can be considered steady (considered to be roughly constant) and the amplitude values ​​of the signal during that time in a time series. It may be data on the duration for which the frequency characteristics of a sound signal can be considered steady. It may be data listing multiple sets of data on the duration for which the frequency characteristics of a sound signal can be considered steady and the frequency characteristics during that time in a time series.

[0166] The data format may be, for example, data indicating the outline of a spectrogram. Furthermore, the volume that serves as a reference for the frequency characteristics may be used as the reference volume. Information on the reference volume and information indicating the properties of the sound data may be used to calculate the volume of the direct sound or reflected sound to be perceived by the listener, as well as in a selection process for selecting whether or not to perceive the direct sound or reflected sound. Other examples of information indicating the properties of the sound data and specific uses for the selection process will be described later.

[0167] Orientation information is typically expressed using yaw, pitch, and roll. Alternatively, the roll rotation may be omitted and the information may be expressed using azimuth (yaw) and elevation (pitch). Orientation information may change over time, and if it does change, it is transmitted to the rendering unit.

[0168] The information about the listener is information about the position and orientation of the listener in sound space. The position information is expressed as a position on the XYZ axes in Euclidean space, but it does not necessarily have to be three-dimensional information and may be two-dimensional information. The information about orientation is typically expressed using yaw, pitch, and roll. Alternatively, the roll rotation may be omitted and the information may be expressed using azimuth (yaw) and elevation (pitch). The position information and orientation information may change over time, and if they change, they are transmitted to the rendering unit.

[0169] The sensor information includes the amount of rotation or displacement detected by a sensor worn by the listener and the position and orientation of the listener. The sensor information is transmitted to a rendering unit, which updates the information on the position and orientation of the listener based on the sensor information. The sensor information may be, for example, position information obtained by a mobile terminal performing self-position estimation using a GPS, a camera, or LiDAR (Laser Imaging Detection and Ranging). Furthermore, information obtained from an external source other than a sensor via a communication module may be detected as sensor information. Information indicating the temperature of the audio signal processing device and information indicating the remaining battery level may be obtained from the sensor. Computing resources (CPU capacity, memory resources, PC performance) of the audio signal processing device and the audio signal presentation device may be obtained in real time.

[0170] The analysis unit performs the same function as the acquisition unit 111 in the above example. That is, it analyzes the input signal and acquires information required by the path calculation unit 121 and the output sound generation unit 131.

[0171] The synthesis unit performs functions equivalent to those of the path calculation unit 121, output sound generation unit 131, and signal output unit 141 in the above example. Based on the audio signal of the direct sound and information on the direct sound arrival time and volume at the time of direct sound arrival calculated by the analysis unit, the synthesis unit processes the input audio signal to generate direct sound. Also, based on information on the reflected sound arrival time and volume at the time of reflected sound arrival calculated by the analysis unit, the synthesis unit processes the input audio signal to generate reflected sound. The synthesis unit synthesizes the generated direct sound and reflected sound and outputs the synthesized sound.

[0172] [Operation] Next, with reference to FIG. 16 , the operation of the sound reproduction system 100, particularly the information processing device 101, described above will be described. FIG. 16 is a flowchart showing an example of the operation of the information processing device according to the embodiment. In the example of operation shown in the figure, the acquisition unit 111 acquires sound information via the communication module 102 (step S101). The sound information is decoded by the decoding processing unit 113 into information about the reproduced sound and information about the position of the sound source object, and generation of an output sound signal begins. First, the path calculation unit 121 uses the decoded sound information to convert, in particular, the sound information into multiple audio signals, and converts the sound information into sound information containing these multiple audio signals (step S102). Each of the multiple audio signals corresponds, for example, to a direct sound and one or more secondary sounds. Note that if sound information that contains multiple audio signals without conversion is acquired, this conversion process may be skipped. The sensing information input unit 114 acquires information (position information) about the position of the user 99 (step S103).

[0173] Next, the target determination unit 132 selects one audio signal from the multiple audio signals (step S104). The audio signals are all selected sequentially, so the order of selection is not particularly limited. The target determination unit 132 determines whether the selected audio signal satisfies a predetermined condition (step S105). If it is determined that the predetermined condition is satisfied (Yes in step S105), the audio signal is provided to the second generation unit 134, where it is subjected to panning processing, and a second sound signal is generated to be perceived by the user 99 as a representative sound from a representative point (step S106). That is, in step S106, the second sound signal is generated by converting the audio signal into a representative sound and convolving a head-related transfer function corresponding to the representative direction.

[0174] If it is determined that the predetermined condition is not satisfied (No in step S105), the audio signal is provided to the first generation unit 133, where the audio signal is not subjected to panning processing, and a first sound signal is generated to make the user perceive the audio signal as coming from the position of the original sound source object (step S107). That is, in step S107, the first sound signal is generated by convolving a head-related transfer function according to the arrival direction with a playback sound emitted from the sound source object in the audio signal.

[0175] Here, the predetermined conditions will be described. The predetermined conditions are conditions used to distinguish whether or not panning processing is performed, as explained above. In other words, by using the predetermined conditions as conditions for distinguishing between audio signals suitable for panning processing and audio signals not suitable for panning processing, it is possible to perform panning processing, i.e., conversion processing, appropriately. An audio signal suitable for panning processing is an audio signal related to a sound that is unlikely to cause discomfort to the user 99 when panning processing is performed. Conversely, an audio signal not suitable for panning processing is an audio signal related to a sound that is likely to cause discomfort to the user 99 when panning processing is performed.

[0176] For example, between a direct sound and a secondary sound generated by conversion as described above, the direct sound is more susceptible to sound degradation, while the secondary sound is less susceptible to sound degradation than the direct sound. Therefore, a predetermined condition may be set so that the audio signal related to the direct sound is not subjected to panning processing, and the audio signal related to the secondary sound is subjected to panning processing. In this case, the predetermined condition is "(1) that it corresponds to a secondary sound." To determine whether or not a sound corresponds to a secondary sound, for example, information related to the generation system may be used. Specifically, it may be determined that audio information of a generation newer than the "parent" by one or more generations corresponds to a secondary sound.

[0177] Furthermore, some secondary sounds are susceptible to sound degradation due to panning processing, while others are not. For example, sounds in which the high-frequency domain is important, such as reverberation, are susceptible to sound degradation due to the panning processing characteristic that the high-frequency domain is easily attenuated. In this way, a predetermined condition can be set so that audio signals related to secondary sounds, such as reverberation, whose sound quality degradation characteristics are susceptible to panning processing are not subjected to panning processing. The predetermined condition in this case is "(2) the sound quality degradation characteristics are not susceptible to panning processing." To determine whether a sound quality degradation characteristic is susceptible or not susceptible to panning processing, for example, a sound quality degradation characteristic flag can be assigned to each audio signal, and if the flag is ON (the characteristic is susceptible to panning processing), the characteristic is determined to be susceptible to panning processing. The above flag may be added during the conversion process that calculates the secondary sound, or if the sound information originally contains multiple audio signals, it may be added when the sound information is generated (by the side that transmits the sound information to the communication module 102).

[0178] Furthermore, a predetermined condition may be set based on the secondary sound generation system. Specifically, information on whether or not panning is performed may be passed on to the next generation in the secondary sound system. In other words, if the audio signal related to the secondary sound of the previous generation is panned, the audio signal related to the secondary sound of the current generation and the audio signal related to the secondary sound of the next generation will all be panned. Conversely, if the audio signal related to the secondary sound of the previous generation is not panned, the audio signal related to the secondary sound of the current generation and the audio signal related to the secondary sound of the next generation will not be panned. In this case, the predetermined condition is "(3) The audio signal related to the secondary sound of the previous generation is determined to be subject to panning processing." Note that, instead of the predetermined condition "(3) The audio signal related to the secondary sound of the previous generation is determined to be subject to panning processing," a predetermined condition "(3') The audio signal related to the secondary sound of the previous generation is not determined to be subject to panning processing" may be used. This means that if the previous generation was not panned, the current generation will be panned.

[0179] Furthermore, depending on the constraints of computational resources, it may be necessary to perform panning processing not only on direct sound, secondary sound, etc., but also on direct sound. In such cases, a predetermined condition may be set depending on the relative direction of the position of the sound source object with respect to the front of the user 99, such as an audio signal related to the sound reproduced in front of the user 99 (e.g., −45 degrees to 45 degrees when directly in front is 0 degrees), an audio signal related to the sound reproduced to the side of the user 99 (e.g., −135 degrees to −45 degrees, 45 degrees to 135 degrees when directly in front is 0 degrees), and an audio signal related to the sound reproduced behind the user 99 (e.g., −180 degrees to −135 degrees, 135 degrees to 180 degrees when directly in front is 0 degrees). The predetermined condition in this case is "(4) the direction of the corresponding sound source object with respect to the front direction of the user 99 is outside a threshold range."

[0180] For example, in the example of the angle range described above, if it is desired to suppress sound degradation in front (when prioritizing the sound of the sound source object on the front side), if the direction of the sound source object related to the audio signal is within the range of -45 degrees to 45 degrees, it is within the threshold range and panning processing is not performed. On the other hand, if the direction of the sound source object related to the audio signal is within the range of -135 degrees to -45 degrees or 45 degrees to 135 degrees, it is outside the threshold range and panning processing is performed. Furthermore, for example, if it is desired to suppress sound degradation extending to the sides (when prioritizing the sound of the sound source object on the side), if the direction of the sound source object related to the audio signal is within the range of -135 degrees to -45 degrees or 45 degrees to 135 degrees, it is within the threshold range and panning processing is not performed. On the other hand, if the direction of the sound source object related to the audio signal is within the range of -180 degrees to -135 degrees or 135 degrees to 180 degrees, it is outside the threshold range and panning processing is performed. It is also possible to prioritize the sound of both the front side and the side side sound source objects.

[0181] Furthermore, some sound degradation may be acceptable for playback sounds emitted by sound source objects that are relatively far from the user 99, secondary sounds generated in relation to relatively far objects, or playback sounds emitted by sound source objects that the user 99 cannot approach due to the settings, or secondary sounds generated in relation to objects that the user 99 cannot approach. Therefore, a predetermined condition may be set so that panning is performed on audio signals related to playback sounds or secondary sounds when the distance between the sound source object or related object and the user 99 is equal to or greater than a threshold. In this case, the predetermined condition is "(5) the distance between the sound source object or related object is equal to or greater than a threshold." Note that the setting of this threshold depends on external factors such as the hearing ability of the user 99, and therefore may be set empirically or experimentally for each user 99. Alternatively, instead of such distance, the sound pressure, gain, or level of the sound may be used, and a predetermined condition may be set so that panning is performed if the sound pressure, gain, or level of the sound is compared to a threshold and is equal to or less than the threshold. In this case, the predetermined condition is "(6) the sound pressure, gain, or level of the sound is equal to or less than a threshold."

[0182] As long as the predetermined conditions (1) to (6) described above are not in a contradictory relationship, two or more predetermined conditions can be combined and used for the determination in step S105. For example, if the predetermined condition (1) or the predetermined condition (2) is satisfied, the result in step S105 will be Yes, or if the predetermined condition (1) and the predetermined condition (2) are satisfied, the result in step S105 will be Yes.

[0183] Returning to the description of FIG. 16 , the target determination unit 132 determines whether all audio signals have been selected (step S108). If all audio signals have not been selected (No in step S108), the process returns to step S104, where a new audio signal is selected and the same process is repeated. If all audio signals have been selected (Yes in step S108), the synthesis unit 135 synthesizes (adds) the generated first and second sound signals (step S109), generating an output sound signal. Then, the signal output unit 141 outputs the output sound signal (step S110).

[0184] [Specific Example of Panning Processing] In the panning processing of the present disclosure, reproduced sounds from multiple sound source objects are represented by representative sounds from multiple representative directions. For example, two or three directions can be used as these representative directions. Specifically, in the panning processing, the number of representative points is reduced to a number less than the number of sound source objects, and the reproduced sounds can be perceived as sounds coming from the direction of arrival using only the head-related transfer functions of the representative directions for these representative points.

[0185] In this case, the panning process calculates a time shift (delay) that maximizes the cross-correlation between the head-related transfer function of the sound source object in the arrival direction and the head-related transfer function of the representative direction. The time shift obtained here, or a time shift obtained by adding a negative sign to this time shift, is applied to the reproduced sound from the sound source object, and the subsequent processing is performed assuming that the signal after the time shift is in the representative direction.

[0186] This time shift may be a time shift shorter than the sampling frequency (a shift in which the sample position is expressed by a decimal number; hereinafter, referred to as a "decimal shift"). This decimal shift can be performed by oversampling.

[0187] Here, in the panning process, a gain is applied to the signal of a representative direction obtained by time-shifting the reproduced sound of the sound source object, and the sum of these values ​​calculated for each representative point is calculated, and the sum is convolved with the head-related transfer function at each representative point, thereby synthesizing a signal equivalent to the reproduced sound of the sound source object convolved with the head-related transfer function of the arrival direction.

[0188] On the other hand, in the panning process, when synthesizing the head-related transfer function (vector) of the arrival direction by the sum of the head-related transfer functions (vector) of the representative direction, the gain may be calculated by orthogonalizing the error signal vector between the synthesized head-related transfer function (vector) and the head-related transfer function (vector) of the arrival direction to the head-related transfer function (vector) of the representative direction. Note that the head-related transfer function (vector) is a time waveform of a head impulse response, which is an expression of the head-related transfer function in the time domain, as a vector. Hereinafter, this head-related transfer function (vector) will also be simply referred to as a "head-related transfer function vector."

[0189] In the panning process, this gain is corrected so that the energy balance of the head-related transfer functions from the position of the sound source object to the left and right ears of the user 99 is maintained even in the head-related transfer functions substantially synthesized by the panning process using head-related transfer functions from multiple representative points. In other words, in the panning process, the gain may be corrected so that the energy balance of the head-related transfer functions of the left and right ears of the user 99 due to the sound source object is maintained even in the head-related transfer functions substantially synthesized by the panning process.

[0190] In this embodiment, the panning process can calculate, for each direction of arrival of the sound source object, a gain value to be multiplied by the head transfer function of the representative direction and a time shift value to be applied to the head transfer function of the representative direction, and store them in a head transfer function table, which will be described later.

[0191] Then, in the panning process, each sound source object is time-shifted by a time shift value and a gain value corresponding to the direction of arrival of each sound source object, and the time shifts and gains are multiplied and summed to generate a sum signal. In the panning process, this sum signal is treated as being present at the position of the representative point. In the panning process, the head-related transfer function at the position of the representative point is convolved with this sum signal to generate a signal at the ear of the user 99.

[0192] In addition, the panning process may use a gain calculated so as to minimize the energy or L2 norm of the error signal vector between the synthesized HRIR vector and the HRIR vector in the sound source direction.

[0193] In the above example, the time shift and the gain that maximize the cross-correlation are calculated by using the head-related transfer function itself. On the other hand, the time shift and / or the gain may be calculated by applying a weighting filter on the frequency axis.

[0194] That is, when calculating the time shift and gain that maximize the cross-correlation, it is possible to use a filter that has been subjected to a weighting filter on the frequency axis (hereinafter also referred to as a "frequency weighting filter").

[0195] It is preferable to use a frequency weighting filter that has a cutoff frequency near or slightly higher than the frequency band where human hearing sensitivity is high, and attenuates the higher frequency band, i.e., the frequency band where human hearing sensitivity decreases. For example, it is preferable to use a low-pass filter (LPF) with a cutoff frequency of 3000 Hz to 6000 Hz and approximately 6 dB / oct (octave) to 12 dB / oct.

[0196] In the panning process, the adjustment amounts in the time shift adjustment and the gain adjustment may be determined according to the head-related transfer functions included in the database 105, and the time shift adjustment and the gain adjustment may be applied to the reproduced sound using the determined adjustment amounts to convert it into a representative sound. Since the optimal values ​​of the adjustment amounts in the time shift adjustment and the gain adjustment used in the panning process change according to the head-related transfer functions, first, when the head-related transfer functions included in the database 105 are read out, the adjustment amounts in the time shift adjustment and the gain adjustment corresponding to the read out head-related transfer functions can be determined, and thereafter, the same adjustment amounts can be reused as long as these head-related transfer functions are used, which is advantageous in terms of the amount of processing.

[0197] The head-related transfer function table is an example of table data including head-related transfer functions stored in the database 105. The head-related transfer function table stores the head-related transfer functions together with the adjustment amounts in the time shift adjustment and the gain adjustment determined according to the head-related transfer functions, which are linked to each other. That is, a head-related transfer function table may be constructed by calculating the adjustment amounts in the time shift adjustment and the gain adjustment in advance for each head-related transfer function included in the database 105. In this way, table data of the head-related transfer function table linking each head-related transfer function with the adjustment amount may be stored in the database 105. In this way, the database 105 is an example of a storage unit. The calculation of the adjustment amount for each head-related transfer function may be performed by the second generation unit 134 or the decoding processing unit 113. Alternatively, the calculation of the adjustment amount may be performed by an external device and stored in the memory of the external device. In this case, the memory of the external device corresponds to an example of a storage unit.

[0198] Furthermore, adjustment amounts in the time shift adjustment and the gain adjustment may be calculated in advance, and an adjustment amount table linked to each of a plurality of representative directions may be constructed and stored in the database 105. The adjustment amount table may include table data linking the head related transfer functions of each of the plurality of representative directions with the adjustment amounts in the time shift adjustment and the gain adjustment, or the head related transfer functions of each of the plurality of representative directions may be extracted from head related transfer functions of the entire celestial sphere (multiple directions) that are acquired in advance and stored in the database 105 at the time of rendering or at the time of system initialization.

[0199] Furthermore, the adjustment amount table may be a table including, for example, information as to which of a plurality of representative directions the signal is to be distributed to, for a sound signal arriving at the position of the listener from the direction of each head-related transfer function in the spherical head-related transfer function database, and information on a time shift adjustment amount and a gain adjustment amount to be multiplied by the sound signal for each representative direction when distributing.

[0200] When performing the convolution process of the head-related transfer function, the adjustment amount table stored in the database 105 is referenced, and the adjustment amounts for the time shift adjustment and gain adjustment linked to the head-related transfer function of the direction to be applied are used, which eliminates the need to calculate the adjustment amount for each convolution process and contributes to reducing the amount of processing.

[0201] Note that the embodiment of the present invention can also be applied to new head-related transfer functions that are not included in the database 105. When decoding a sound signal, when powering on the sound reproduction system 100, or when initializing the sound reproduction system 100, the head-related transfer functions of the entire three-dimensional sound field may be newly read, and the adjustment amount for each head-related transfer function may be calculated using the method disclosed in this embodiment or another method. In this case, table data linking the head-related transfer functions with the adjustment amounts may be stored in the database 105. Alternatively, the adjustment amounts may be calculated by an external device and stored in the memory of the external device. When performing convolution processing of head-related transfer functions, by referencing the adjustment amounts for the time shift adjustment and gain adjustment linked to the applied head-related transfer function, it is not necessary to calculate the adjustment amount for each convolution processing, which can contribute to reducing the amount of processing.

[0202] In this way, when a new head-related transfer function not stored in the database 105 is read, adjustment amounts for the time shift adjustment and gain adjustment used in the panning process may be determined for the new head-related transfer function before storing it in the database 105, and a head-related transfer function table may be constructed by linking the new head-related transfer function with the determined adjustment amounts, and the head-related transfer function table may be stored in the database 105. Then, when performing panning process, these adjustment amounts are read from the database 105, and shift adjustment and gain adjustment are applied based on these adjustment amounts. Note that the new head-related transfer function may be one that was previously stored in the database 105, but was temporarily removed from the database 105 when decoding a sound signal, when powering on the sound reproduction system 100, when initializing the sound reproduction system 100, or the like, and then re-stored in the database 105. Furthermore, such determination of the adjustment amounts is effective without switching whether or not to perform panning process. That is, instead of the first generation unit 133 and the second generation unit 134, a third generation unit different from the first generation unit 133 and the second generation unit 134 may be provided, which applies time shift adjustment and gain adjustment to the reproduced sound with an adjustment amount linked to a new head-related transfer function stored in the database to convert it into a representative sound, and generates an output sound signal by convolving a head-related transfer function corresponding to a representative direction from each position of the representative point toward the user's position into the representative sound.

[0203] Other Embodiments Although the embodiments have been described above, the present disclosure is not limited to the above-described embodiments.

[0204] For example, the sound reproduction system described in the above embodiment may be realized as a single device including all of the components, or may be realized by allocating each function to multiple devices and coordinating these multiple devices. In the latter case, the device corresponding to the information processing device may be an information processing device such as a smartphone, tablet terminal, or PC. For example, in the sound reproduction system 100 having a function as a renderer that generates an audio signal with added sound effects, all or part of the renderer's functions may be performed by a server. That is, all or part of the acquisition unit 111, path calculation unit 121, output sound generation unit 131, and signal output unit 141 may be located on a server (not shown). In this case, the sound reproduction system 100 is realized by combining, for example, an information processing device such as a computer or smartphone, a sound presentation device such as a head-mounted display (HMD) or earphones worn by the user 99, and a server (not shown). Note that the computer, sound presentation device, and server may be connected to each other so as to be able to communicate with each other via the same network, or may be connected via different networks. When the sound reproduction system 100 is connected via different networks, the possibility of communication delays increases, so processing on the server may be permitted only when the computer, sound presentation device, and server are connected so as to be able to communicate with each other via the same network. Also, depending on the amount of bitstream data received by the sound reproduction system 100, it may be determined whether the server will take on all or part of the functions of the renderer.

[0205] The sound reproduction system of the present disclosure can also be realized as an information processing device that is connected to a reproduction device having only a driver and that simply reproduces an output sound signal generated based on acquired sound information to the reproduction device. In this case, the information processing device may be realized as hardware having a dedicated circuit, or as software that causes a general-purpose processor to execute specific processing.

[0206] In the above-described embodiment, the processing performed by a specific processing unit may be performed by another processing unit. The order of multiple processing operations may be changed, or multiple processing operations may be performed in parallel.

[0207] In the above-described embodiments, each component may be realized by executing a software program suitable for that component, or by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0208] Furthermore, each component may be realized by hardware. For example, each component may be a circuit (or integrated circuit). These circuits may form a single circuit as a whole, or each may be a separate circuit. Furthermore, each of these circuits may be a general-purpose circuit or a dedicated circuit.

[0209] Furthermore, the general or specific aspects of the present disclosure may be realized as an apparatus, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, etc. Furthermore, the general or specific aspects of the present disclosure may be realized as any combination of an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0210] For example, the present disclosure may be realized as an audio signal reproducing method executed by a computer, or as a program for causing a computer to execute the audio signal reproducing method. The present disclosure may also be realized as a computer-readable non-transitory recording medium on which such a program is recorded.

[0211] In addition, this disclosure also includes forms obtained by applying various modifications to each embodiment that a person skilled in the art would think of, or forms realized by arbitrarily combining the components and functions of each embodiment within the scope that does not deviate from the intent of this disclosure.

[0212] In the present disclosure, the encoded sound information can be rephrased as a bitstream containing a sound signal, which is information about a predetermined sound to be reproduced by the sound reproduction system 100, and metadata, which is information about the localization position when the sound image of the predetermined sound is localized at a predetermined position within a three-dimensional sound field. For example, the sound information may be acquired by the sound reproduction system 100 as a bitstream encoded in a predetermined format such as MPEG-H 3D Audio (ISO / IEC 23008-3). As an example, the encoded sound signal includes information about the predetermined sound to be reproduced by the sound reproduction system 100. The predetermined sound here refers to a sound emitted by a sound source object present in the three-dimensional sound field or a natural environmental sound, and may include, for example, a mechanical sound or the sounds of animals, including humans. When multiple sound source objects are present in the three-dimensional sound field, the sound reproduction system 100 acquires multiple sound signals corresponding to the multiple sound source objects.

[0213] On the other hand, metadata is, for example, information used to control acoustic processing of a sound signal in the sound reproduction system 100. Metadata may be information used to describe a scene expressed in a virtual space (three-dimensional sound field). Here, a scene is a term that refers to a collection of all elements representing three-dimensional images and acoustic events in a virtual space, modeled by the sound reproduction system 100 using metadata. In other words, the metadata here may include not only information for controlling acoustic processing, but also information for controlling video processing. Of course, the metadata may include information for controlling only either audio processing or video processing, or may include information used to control both. In the present disclosure, the bitstream acquired by the sound reproduction system 100 may include such metadata. Alternatively, the sound reproduction system 100 may acquire the metadata separately from the bitstream, as described below.

[0214] The sound reproduction system 100 generates virtual sound effects by performing sound processing on the sound signal using metadata included in the bitstream and additionally acquired position information of the interactive user 99. For example, sound effects such as early reflection sound generation, late reverberation sound generation, diffraction sound generation, distance attenuation effect, localization, sound image localization processing, and Doppler effect may be added. Information for switching on and off all or part of the sound effects may also be added as metadata.

[0215] All or part of the metadata may be obtained from sources other than the bitstream of audio information. For example, either the metadata controlling audio or the metadata controlling video may be obtained from sources other than the bitstream, or both may be obtained from sources other than the bitstream.

[0216] Furthermore, if metadata for controlling the video is included in the bitstream acquired by the audio reproduction system 100, the audio reproduction system 100 may have a function for outputting the metadata that can be used for controlling the video to a display device that displays images or a 3D video reproduction device that reproduces 3D video.

[0217] As an example, the encoded metadata includes information about a three-dimensional sound field including a sound source object emitting a sound and an obstacle object, and information about a localization position when a sound image of the sound is localized at a predetermined position in the three-dimensional sound field (i.e., the sound is perceived as arriving from a predetermined direction), i.e., information about the predetermined direction. Here, an obstacle object is an object that can affect the sound perceived by the user 99 by, for example, blocking or reflecting the sound emitted by the sound source object before it reaches the user 99. Obstacle objects may include not only stationary objects but also animals such as people or moving objects such as machines. Furthermore, when multiple sound source objects exist in a three-dimensional sound field, the other sound source objects may be obstacle objects for any one sound source object. Furthermore, both non-sound-source objects such as building materials or inanimate objects and sound-emitting sound objects may be obstacle objects.

[0218] The spatial information constituting the metadata may include not only the shape of the three-dimensional sound field, but also information representing the shape and position of obstacle objects present in the three-dimensional sound field and the shape and position of sound source objects present in the three-dimensional sound field. The three-dimensional sound field may be either a closed space or an open space, and the metadata may include information representing the reflectance of structures that can reflect sound in the three-dimensional sound field, such as floors, walls, or ceilings, and the reflectance of obstacle objects present in the three-dimensional sound field. Here, the reflectance is the ratio of the energy of reflected sound to incident sound, and is set for each frequency band of sound. Of course, the reflectance may be set uniformly regardless of the frequency band of the sound. Furthermore, when the three-dimensional sound field is an open space, parameters such as a uniform attenuation rate, diffracted sound, or early reflection sound may be used.

[0219] In the above description, reflectance is cited as a parameter related to an obstacle object or a sound source object included in the metadata, but the metadata may include information other than reflectance. For example, information about the material of the object may be included as metadata related to both the sound source object and the non-sound source object. Specifically, the metadata may include parameters such as diffusion rate, transmittance, or sound absorption rate.

[0220] Information about the sound source object may include information such as volume, radiation characteristics (directivity), playback conditions, the number and type of sound sources emitted from a single object, or information specifying a sound source area within the object. The playback conditions may, for example, determine whether the sound is a continuous sound or an event-triggering sound. The sound source area within the object may be determined based on the relative relationship between the position of the user 99 and the position of the object, or may be determined based on the object. When the sound source area is determined based on the relative relationship between the position of the user 99 and the position of the object, the surface from which the user 99 is viewing the object is used as the reference, and the user 99 can be made to perceive sound X as emanating from the right side of the object and sound Y as emanating from the left side of the object as viewed from the user 99. When the sound source area is determined based on the object, the surface from which the user 99 is viewing the object is used as the reference, and the sound emitted from which area of ​​the object can be fixed regardless of the direction the user 99 is viewing. For example, the user 99 can be made to perceive a high-pitched sound coming from the right side and a low-pitched sound coming from the left side when viewing the object from the front. In this case, when the user 99 goes around to the back of the object, the user 99 can be made to perceive a low-pitched sound coming from the right side and a high-pitched sound coming from the left side as viewed from the back.

[0221] The spatial metadata may include the time to early reflections, the reverberation time, or the ratio of direct sound to diffuse sound. If the ratio of direct sound to diffuse sound is zero, the user 99 will perceive only direct sound.

[0222] Furthermore, information indicating the position and orientation of the user 99 in the three-dimensional sound field may be included in the bitstream as metadata as an initial setting, or may not be included in the bitstream. If the information indicating the position and orientation of the user 99 is not included in the bitstream, the information indicating the position and orientation of the user 99 is acquired from information other than the bitstream. For example, the position information of the user 99 in the VR space may be acquired from an app that provides VR content. The position information of the user 99 for presenting sound as AR may be position information obtained by a mobile terminal performing self-position estimation using GPS, a camera, LiDAR (Laser Imaging Detection and Ranging), or the like. Note that the sound signal and metadata may be stored in a single bitstream or may be stored separately in multiple bitstreams. Similarly, the sound signal and metadata may be stored in a single file or may be stored separately in multiple files.

[0223] When an audio signal and metadata are stored separately in multiple bitstreams, information indicating other related bitstreams may be included in one or some of the multiple bitstreams in which the audio signal and metadata are stored. Also, information indicating other related bitstreams may be included in metadata or control information of each of the multiple bitstreams in which the audio signal and metadata are stored. When an audio signal and metadata are stored separately in multiple files, information indicating other related bitstreams or files may be included in one or some of the multiple files in which the audio signal and metadata are stored. Also, information indicating other related bitstreams or files may be included in metadata or control information of each of the multiple bitstreams in which the audio signal and metadata are stored.

[0224] Here, the related bitstreams or files are, for example, bitstreams or files that may be used simultaneously during audio processing. Furthermore, information indicating other related bitstreams may be collectively described in the metadata or control information of one bitstream among multiple bitstreams storing audio signals and metadata, or may be separately described in the metadata or control information of two or more bitstreams among the multiple bitstreams storing audio signals and metadata. Similarly, information indicating other related bitstreams or files may be collectively described in the metadata or control information of one file among multiple files storing audio signals and metadata, or may be separately described in the metadata or control information of two or more files among the multiple files storing audio signals and metadata. Furthermore, a control file collectively describing information indicating other related bitstreams or files may be generated separately from the multiple files storing audio signals and metadata. In this case, the control file does not need to store the audio signal and metadata.

[0225] Here, the information indicating the other related bitstream or file may be, for example, an identifier indicating the other bitstream, a file name indicating the other file, a URL (Uniform Resource Locator), or a URI (Uniform Resource Identifier). In this case, the acquisition unit identifies or acquires the bitstream or file based on the information indicating the other related bitstream or file. Furthermore, the information indicating the other related bitstream may be included in metadata or control information of at least some of the bitstreams among a plurality of bitstreams storing audio signals and metadata, and the information indicating the other related file may be included in metadata or control information of at least some of the files among a plurality of files storing audio signals and metadata. Here, the file including information indicating the related bitstream or file may be, for example, a control file such as a manifest file used for content distribution.

[0226] The present disclosure is useful in reproducing sound, for example, by allowing a user to perceive stereoscopic sound.

[0227] 99 User 100 Sound reproduction system 101 Information processing device 102 Communication module 103 Detector 104 Driver 105 Database 111 Acquisition unit 112 Encoded sound information input unit 113 Decode processing unit 114 Sensing information input unit 121 Path calculation unit 131 Output sound generation unit 132 Object determination unit 133 First generation unit 134 Second generation unit 135 Synthesis unit 141 Signal output unit 300 3D video reproduction device

Claims

1. An information processing method executed by a computer, comprising: a step of acquiring sound information including an acoustic signal and information on the position of a sound source object in a three-dimensional sound field; a step of generating a first sound signal using the acoustic signal and a head-related transfer function corresponding to an arrival direction based on the position of the sound source object and the position of a user in the three-dimensional sound field; a step of generating a second sound signal using the acoustic signal and a head-related transfer function corresponding to a representative direction based on the position of a representative point set in the three-dimensional sound field and the position of the user; and a step of synthesizing the generated first sound signal and second sound signal to generate an output sound signal.

2. The information processing method of claim 1, wherein in the step of generating the first sound signal, the first sound signal is generated by convolving a head-related transfer function corresponding to the arrival direction with a reproduced sound emitted from the sound source object by the acoustic signal, and in the step of generating the second sound signal, a conversion process is performed to convert the reproduced sound into a representative sound arriving from the representative point, and the second sound signal is generated by convolving a head-related transfer function corresponding to the representative direction.

3. The information processing method according to claim 2, wherein the conversion process applies time shift adjustment and gain adjustment to the reproduced sound to convert it into the representative sound.

4. The information processing method according to claim 1, wherein the sound information includes a plurality of the sound signals and a plurality of positions of the sound source objects that correspond one-to-one to the plurality of the sound signals, and the information processing method further includes a step of determining whether each of the plurality of sound signals included in the sound information is to be used to generate the first sound signal or the second sound signal.

5. The information processing method according to claim 1, wherein the sound information includes a plurality of the sound signals and the position of one of the sound source objects that corresponds many-to-one to the plurality of the sound signals, and the information processing method further includes a step of determining whether each of the plurality of sound signals included in the sound information is to be used to generate the first sound signal or the second sound signal.

6. The information processing method according to claim 5, further comprising a step of converting the acquired sound information, wherein the acquired sound information is converted into sound information including a plurality of the acoustic signals.

7. An information processing method according to any one of claims 4 to 6, wherein the plurality of sound signals include a sound signal relating to a direct sound that arrives directly from the position of the corresponding sound source object to the position of the user, and a sound signal relating to a secondary sound that occurs in conjunction with the direct sound and arrives via a path different from that of the direct sound, and wherein the determining step determines that the sound signal relating to the direct sound will be used to generate the first sound signal, and determines that the sound signal relating to the secondary sound will be used to generate the second sound signal.

8. The information processing method according to any one of claims 4 to 6, wherein the plurality of sound signals include sound signals relating to direct sounds that arrive directly from the position of the corresponding sound source object to the position of the user, and sound signals relating to secondary sounds that occur in association with the direct sounds and arrive via a path different from that of the direct sounds, and the secondary sounds include other secondary sounds that occur in association with one of the secondary sounds, and in the determining step, it is determined whether each of the sound signals relating to the secondary sounds will be used to generate the second sound signal based on information relating to the generation system of each of the secondary sounds from the direct sound.

9. An information processing method according to any one of claims 4 to 6, wherein in the determining step, it is determined whether each of the plurality of acoustic signals will be used to generate the first sound signal or the second sound signal based on the direction of the corresponding sound source object relative to the front direction of the user.

10. An information processing method according to any one of claims 4 to 6, wherein in the determining step, it is determined whether each of the plurality of acoustic signals will be used to generate the first sound signal or the second sound signal based on the distance from the user's position to the corresponding position of the sound source object.

11. An information processing method according to any one of claims 4 to 6, wherein in the determining step, it is determined whether to use a plurality of acoustic signals to generate the first sound signal or the second sound signal based on characteristics of sound quality degradation when the second sound signal is generated for each of the plurality of acoustic signals.

12. The information processing method described in claim 1, wherein the sound information includes the acoustic signal related to a direct sound that arrives directly from the position of the corresponding sound source object to the position of the user, and the acoustic signal related to a secondary sound that occurs in conjunction with the direct sound and arrives via a path different from that of the direct sound, and when the acoustic signal related to the secondary sound is identified as a reverberant sound based on information related to the generation system from the direct sound, it is determined not to use the acoustic signal related to the secondary sound in generating the second sound signal.

13. An information processing device comprising: an acquisition unit that acquires sound information including an acoustic signal and information on the position of a sound source object within a three-dimensional sound field; a first generation unit that generates a first sound signal using the acoustic signal and a head-related transfer function corresponding to an arrival direction based on the position of the sound source object and the position of a user within the three-dimensional sound field; a second generation unit that generates a second sound signal using the acoustic signal and a head-related transfer function corresponding to a representative direction based on the position of a representative point set within the three-dimensional sound field and the position of the user; and a synthesis unit that synthesizes the generated first sound signal and second sound signal to generate an output sound signal.

14. A program for causing the computer to execute the information processing method according to claim 1.

Citation Information

Patent Citations

  • Signal processing device and method, and program

    WO2019116890A1

  • Information processing method, program, and acoustic reproduction device

    WO2022038929A1

  • Sound generation device, sound reproduction device, sound generation method, and sound signal processing program

    WO2023210699A1