Information processing device, information processing method, and program

JPWO2024214799A5Pending Publication Date: 2026-01-20
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025514020
Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2025-09-29
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Conventional techniques for generating three-dimensional sound in virtual spaces require significant processing and storage due to the need for precise sound localization, which becomes computationally intensive with multiple sound sources and complex acoustic effects, and panning processing may not effectively reduce processing loads.

Method used

An information processing device that acquires sound information and generates output sound signals using head-related transfer functions based on the user's position and sound source positions, allowing for dynamic adjustment of processing methods between direct convolution and panning processing to optimize processing efficiency.

Benefits of technology

This approach enables efficient generation of three-dimensional sound by dynamically adjusting processing methods, reducing computational load and improving the sense of realism while maintaining processing efficiency, even with multiple sound sources and complex acoustic effects.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

An information processing device (101) comprises: an acquisition unit (111) that acquires sound information including a sound signal and information relating to the position of a sound source object in a three-dimensional sound field; a first generation unit (133) that generates an output sound signal using a head-related transfer function corresponding to an arrival direction based on the position of the sound source object and the position of a user in the three-dimensional sound field, and the sound signal; and a second generation unit (134) that generates an output sound signal using a head-related transfer function corresponding to a representative direction based on the position of a representative point set in the three-dimensional sound field and the position of the user, and the sound signal.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and program

[0001] The present disclosure relates to an information processing device, an information processing method, and a program.

[0002] Conventionally, there are known techniques for reproducing sound in a virtual three-dimensional space to allow a user to perceive stereoscopic sound (see, for example, Patent Document 1). Furthermore, in order to perceive sound as if it is coming from a sound source object to the user in such a three-dimensional space, processing is required to generate output sound information from original sound information. In particular, reproducing stereoscopic sound in response to the user's body movements in a virtual space requires extensive processing. Advances in computer graphics (CG) have made it relatively easy to create visually complex virtual environments, making technology for realizing corresponding auditory information important. Additionally, when processing from sound information to output sound information is performed in advance, a large memory area is required to store the pre-calculated processing results. Furthermore, transmitting such large amounts of processed data may require a wide communication bandwidth.

[0003] To realize a more realistic sound environment, the number of objects that emit sound in the virtual three-dimensional space increases, sound effects such as reflected sound, diffracted sound, and reverberation increase, and these sound effects must be appropriately changed in response to the user's movements, requiring a large amount of processing.To reduce this large amount of processing, a conversion technique called panning processing is known, which represents sounds in a three-dimensional space by sounds from several representative points set in advance within the three-dimensional space.

[0004] Japanese Patent Application Laid-Open No. 2020-18620

[0005] However, conversion processes such as panning may not be effective in reducing the amount of processing. Therefore, an object of the present disclosure is to provide an information processing device or the like for effectively applying conversion processes.

[0006] An information processing device according to one aspect of the present disclosure includes an acquisition unit that acquires sound information including an audio signal and information on the position of a sound source object within a three-dimensional sound field, a first generation unit that generates an output sound signal using a head-related transfer function corresponding to an arrival direction based on the position of the sound source object and the position of a user within the three-dimensional sound field and the audio signal, and a second generation unit that generates an output sound signal using a head-related transfer function corresponding to a representative direction based on the position of a representative point set within the three-dimensional sound field and the position of the user and the audio signal.

[0007] In addition, an information processing device according to another aspect of the present disclosure includes a memory unit that stores a time shift adjustment amount and a gain adjustment amount in association with each of a plurality of directions; an acquisition unit that acquires an audio signal and information on the position of a sound source object within a three-dimensional sound field; and a second generation unit that generates an output sound signal as sound arriving at the user's position from a second direction using the audio signal and the time shift adjustment amount and gain adjustment amount corresponding to a first direction based on the position of the sound source object and the position of the user within the three-dimensional sound field.

[0008] Furthermore, an information processing method according to one aspect of the present disclosure is an information processing method executed by a computer that processes sound information to generate an output sound signal as sound arriving from a sound source object in a virtual three-dimensional sound field, and includes the steps of acquiring the position of the sound source object and an audio signal, the audio signal including a reproduced sound emitted from the sound source object by the audio signal; acquiring the position of a user in the three-dimensional sound field; calculating the arrival direction of the reproduced sound arriving from the position of the sound source object to the user's position; generating the output sound signal using a head-related transfer function corresponding to the calculated arrival direction and the reproduced sound; and generating the output sound signal using a head-related transfer function corresponding to a representative direction based on the position of a representative point set in the three-dimensional sound field and the position of the user, and the audio signal.

[0009] Furthermore, one aspect of the present disclosure can also be realized as a program for causing a computer to execute the information processing method described above.

[0010] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0011] According to the present disclosure, it is possible to effectively apply conversion processing.

[0012] FIG. 1 is a schematic diagram showing a use example of an audio reproduction system according to an embodiment. FIG. 2 is a block diagram showing a functional configuration of an audio reproduction system according to an embodiment. FIG. 3 is a block diagram showing a functional configuration of an acquisition unit according to an embodiment. FIG. 4 is a block diagram showing a functional configuration of an output sound generation unit according to an embodiment. FIG. 5 is a flowchart showing a first operation example of an information processing device according to an embodiment. FIG. 6 is a flowchart showing a second operation example of an information processing device according to an embodiment. FIG. 7 is a diagram for explaining a processing target of panning processing according to an embodiment. FIG. 8 is a flowchart showing a third operation example of an information processing device according to an embodiment.

[0013] (Knowledge that forms the basis of the disclosure) Conventionally, a technology related to sound reproduction that allows a user to perceive stereoscopic sound in a virtual three-dimensional space (hereinafter sometimes referred to as a three-dimensional sound field) has been known (see, for example, Patent Document 1). Using this technology, a user can perceive sound as if a sound source object exists at a predetermined position in the virtual space and the sound is coming from that direction. In order to localize a sound image at a predetermined position in the virtual three-dimensional space in this way, for example, calculation processing is required for a sound signal emitted by a sound source object (also referred to as a sound emitted from the sound source object or a reproduced sound) to generate a sound arrival time difference between the two ears and a sound level difference (or sound pressure difference) between the two ears that causes the sound to be perceived as stereoscopic sound. Such calculation processing is performed by applying a stereophonic filter. A stereophonic filter is an information processing filter that, when an output sound signal obtained by applying the filter to original sound information is reproduced, causes the position (such as the direction and distance) of the sound, the size of the sound source, the width of the space, and the like to be perceived with a three-dimensional effect.

[0014] As an example of the computational process for applying such a stereophonic filter, a process is known in which a head-related transfer function (HRTF) is convolved with a target sound signal to make the sound perceived as coming from a predetermined direction. By performing this HRTF convolution process at a sufficiently fine angle with respect to the direction of arrival of the reproduced sound from the position of the sound source object to the user's position, the sense of realism experienced by the user is improved.

[0015] In recent years, there has been active development of technologies related to virtual reality (VR). Virtual reality focuses on appropriately changing the position of a sound object in a virtual three-dimensional space in response to a user's movements, allowing the user to experience the sensation of moving within the virtual space. To achieve this, it is necessary to move the localization position of a sound image in the virtual space relative to the user's movements. This processing has been performed by applying a stereophonic filter, such as the head-related transfer function described above, to the original sound information. However, when a user moves within a three-dimensional space, the sound transmission path changes from moment to moment depending on the positional relationship between the sound source object and the user, due to factors such as sound reverberation and interference. Therefore, if the sound transmission path from the sound source object is determined based on the positional relationship between the sound source object and the user each time and the transfer function is convolved taking into account sound reverberation and interference, the amount of information processing becomes enormous, and an improvement in the sense of realism may not be achieved without a large-scale processing device.

[0016] Therefore, in order to reduce such an enormous amount of processing, attempts have been made to apply a panning process to the reproduced sound to reduce the amount of convolution of the head-related transfer function. Specifically, instead of convolving the reproduced sound with a head-related transfer function for each of a number of sound source objects in a three-dimensional space, the reproduced sound from the sound source object is re-expressed using sounds (representative sounds) from several representative points pre-set in the three-dimensional space. Then, by simply convolving the representative sounds with the head-related transfer functions from the representative points to the user's position, it becomes possible to allow the user to perceive a three-dimensional sound that is comparable to that of the original sound source objects. If the number of representative points is smaller than the number of original sound source objects, the number of targets for convolution of the head-related transfer function will naturally be reduced, which is advantageous in terms of processing amount.

[0017] On the other hand, when such panning processing is applied, under certain conditions, such as when the number of original sound source objects is small, the processing amount of the panning processing itself increases, and therefore the overall processing amount reduction effect may not be achieved. Therefore, the present disclosure provides an information processing device that includes a processing unit for generating two types of output sound signals, so that panning processing can be applied and not applied. This makes it possible to generate an output sound signal with panning processing applied if panning processing is effective in reducing the processing amount, and to generate an output sound signal without applying panning processing if panning processing is not effective in reducing the processing amount. In other words, it is possible to effectively apply conversion processing such as panning processing.

[0018] A more specific outline of the present disclosure is as follows.

[0019] An information processing device according to a first aspect of the present disclosure includes an acquisition unit that acquires sound information including an audio signal and information on the position of a sound source object within a three-dimensional sound field, a first generation unit that generates an output sound signal using the audio signal and a head-related transfer function corresponding to an arrival direction based on the position of the sound source object and the position of a user within the three-dimensional sound field, and a second generation unit that generates the output sound signal using the audio signal and a head-related transfer function corresponding to a representative direction based on the position of a representative point set within the three-dimensional sound field and the position of the user.

[0020] According to such an information processing device, it is possible to generate an output sound signal using a head-related transfer function corresponding to the arrival direction calculated using the first generation unit, and to generate an output sound signal using a head-related transfer function corresponding to the representative direction using the second generation unit. For example, if using the second generation unit effectively reduces the amount of processing, the second generation unit can be used, and if not, the first generation unit can be used. In other words, from the perspective of the amount of processing, it is possible to effectively apply the conversion process by, for example, classifying conditions.

[0021] In addition, the information processing device according to the second aspect is the information processing device described in the first aspect, in which the first generation unit generates an output sound signal by convolving a head-related transfer function corresponding to the arrival direction with a playback sound emitted from a sound source object by an audio signal, and the second generation unit performs a conversion process to convert the playback sound into a representative sound arriving from a representative point, and generates the output sound signal by convolving a head-related transfer function corresponding to the representative direction.

[0022] According to this, the first generation unit generates an output sound signal by convolving a head-related transfer function corresponding to the direction of arrival with the reproduced sound, and the second generation unit generates an output sound signal by performing a conversion process such as a panning process to represent the sound from the direction of arrival using representative sounds arriving from each of the representative points set in a three-dimensional sound field. For example, if applying the conversion process effectively reduces the amount of processing, the second generation unit can be used, and if not, the first generation unit can be used. In other words, from the perspective of the amount of processing, it is possible to effectively apply the conversion process by, for example, classifying conditions.

[0023] An information processing device according to a third aspect is the information processing device according to the second aspect, wherein the conversion process applies time shift adjustment and gain adjustment to the reproduced sound to convert it into the representative sound.

[0024] According to this, in the conversion process, the playback sound can be converted into a representative sound by applying time shift adjustment and gain adjustment, which reduces the sense of incongruity even when the conversion process is applied, and enables the generation of an output sound signal with a higher sense of realism.

[0025] Furthermore, an information processing device according to a fourth aspect is an information processing device according to any one of the first to third aspects, in which the sound information includes the positions of each of the multiple sound source objects and the reproduced sounds emitted from each of the multiple sound source objects by the audio signal, and the number of representative points is determined based on the number of sound source objects.

[0026] This allows the number of representative points to be dynamically changed based on the number of sound source objects, and appropriate conversion processing to be performed each time.

[0027] An information processing device according to a fifth aspect is the information processing device according to the fourth aspect, in which the number of representative points is smaller than the number of sound source objects.

[0028] This allows the number of representative points to be dynamically changed based on the number of sound source objects, and appropriate conversion processing can be performed each time. In particular, since the number of representative points can be reduced relative to the number of sound source objects, there is an advantage in that it is easy to increase the effect of reducing the amount of processing required by the conversion processing.

[0029] In addition, an information processing device according to a sixth aspect is the information processing device according to the third aspect, in which, in the time shift adjustment of the conversion process, a time shift is performed on the reproduced sound that is calculated so as to maximize the cross-correlation between the head-related transfer function corresponding to the arrival direction and the head-related transfer function corresponding to the representative direction, or a time shift with a negative sign added to the time shift.

[0030] This allows time shift adjustment to be made to the reproduced sound by performing a time shift calculated to maximize the cross-correlation between the head transfer function of the arrival direction and the head transfer function of the representative direction, or by performing a time shift with a negative sign added to the time shift.

[0031] In addition, an information processing device according to a seventh aspect is the information processing device according to the sixth aspect, wherein in the conversion process, at least one of the time shift adjustment and the gain adjustment is performed by applying a weighting filter on the frequency axis and then performing a time shift calculated to maximize the cross-correlation, or a time shift with a negative sign added to the time shift.

[0032] According to this, by applying a weighting filter on the frequency axis and then performing a time shift calculated to maximize the cross-correlation, or by performing a time shift with a negative sign added to the time shift, it is possible to perform at least one of a time shift adjustment and a gain adjustment.

[0033] In addition, an information processing device according to an eighth aspect is the information processing device according to the sixth aspect, in which, in the conversion process, for each of two or more representative points, the time-shifted playback sound is multiplied by a gain set for the playback sound and the representative direction.

[0034] According to this, for each of the two or more representative points, the conversion process can be performed by applying a gain set for each of the reproduced sound arrival direction and each representative direction to the time-shifted reproduced sound.

[0035] Furthermore, an information processing device according to a ninth aspect is the information processing device according to the eighth aspect, in which, in the conversion process, when synthesizing a head-related transfer function vector according to the direction of arrival by adding up head-related transfer function vectors according to the representative direction, a gain calculated so that an error signal vector between the synthesized head-related transfer function vector and the head-related transfer function vector according to the direction of arrival is orthogonal to the head-related transfer function vector according to the representative direction is used.

[0036] According to this, when synthesizing a head transfer function vector of the arrival direction by adding up head transfer function vectors of the representative direction, a conversion process can be performed using a gain calculated so that the error signal vector between the synthesized head transfer function vector and the head transfer function vector of the arrival direction is orthogonal to the head transfer function vector of the representative direction.

[0037] Furthermore, an information processing device according to a tenth aspect is the information processing device according to the eighth aspect, in which the conversion process uses a gain calculated so as to minimize the energy or L2 norm of an error signal vector between the synthesized head-related transfer function vector and a head-related transfer function vector according to the direction of arrival.

[0038] This allows the conversion process to be performed using a gain calculated so as to minimize the energy or L2 norm of the error signal vector between the synthesized head-related transfer function vector and the head-related transfer function vector of the arrival direction.

[0039] An information processing device according to an eleventh aspect is the information processing device according to the tenth aspect, wherein the error signal vector is subjected to a weighting filter on the frequency axis.

[0040] According to this, an error signal vector that has been subjected to a weighting filter on the frequency axis can be used.

[0041] In addition, an information processing device according to a certain aspect is the information processing device described in the third aspect, in which, when the information processing device reads a new head-related transfer function that is not stored in a memory unit for storing head-related transfer functions, the information processing device determines adjustment amounts in the time shift adjustment and gain adjustment to be used in the conversion process for the new head-related transfer function, links the read new head-related transfer function with the determined adjustment amounts and stores them in a database, and in the conversion process applies time shift adjustment and gain adjustment to the reproduced sound using the adjustment amounts linked to the new head-related transfer function stored in the memory unit to convert it into a representative sound.

[0042] According to this, when a new head-related transfer function that is not stored in the storage unit for storing head-related transfer functions is read, adjustment amounts in the time shift adjustment and gain adjustment used in the conversion process for the new head-related transfer function are determined, and the read new head-related transfer function and the determined adjustment amounts are linked and stored in the storage unit, and can be used in the conversion process. The new head-related transfer function has an adjustment amount appropriate for the head-related transfer function, and by determining such adjustment amount before starting the conversion process (for example, when decoding the sound signal, when turning on the power of the sound reproduction system, or when initializing the sound reproduction system), the conversion process with an appropriate adjustment amount can be performed while suppressing an increase in the processing amount.

[0043] Furthermore, the information processing device according to the twelfth aspect is the information processing device according to the third aspect, which, at the time of initialization, stores in a memory an adjustment amount table in which the head related transfer function of a representative direction and the adjustment amounts in the time shift adjustment and gain adjustment used in the conversion process are linked for each direction of the head related transfer function, and in the conversion process, converts the reproduced sound into a representative sound by applying the time shift adjustment and gain adjustment with the adjustment amounts linked for each direction of the head related transfer function according to the representative direction in the adjustment amount table stored in the memory.

[0044] According to this, in the conversion process, it is possible to convert into a representative sound by applying time shift adjustment and gain adjustment using adjustment amounts linked to each direction of the head related transfer function according to the representative direction from the adjustment amount table stored in the memory unit at the time of initialization.

[0045] In addition, an information processing device according to a thirteenth aspect is the information processing device according to the twelfth aspect, in which a plurality of representative directions are determined at the time of initialization, and the adjustment amount table is created based on the head-related transfer functions of the determined plurality of representative directions.

[0046] According to this, in the conversion process, time shift adjustment and gain adjustment can be applied to convert into a representative sound using adjustment amounts linked to each direction of the head related transfer function created based on the head related transfer functions of the determined multiple representative directions.

[0047] Furthermore, an information processing device according to a fourteenth aspect is an information processing device according to any one of the first to thirteenth aspects, in which the sound information includes a flag specifying whether to generate an output sound signal using the first generation unit or the second generation unit, and the information processing device generates the output sound signal using either the first generation unit or the second generation unit, as specified in the flag included in the acquired sound information.

[0048] According to this, the output sound signal can be generated using a designated one of the first generation unit and the second generation unit, depending on the flag included in the sound information. In other words, the flag can specify which of the first generation unit and the second generation unit to use.

[0049] In addition, an information processing device according to a fifteenth aspect is an information processing device according to any one of the first to fourteenth aspects, which is provided with a switching unit that switches between generating an output sound signal using the first generation unit and generating an output sound signal using the second generation unit.

[0050] This makes it possible to switch between generating an output sound signal using the first generation unit and generating an output sound signal using the second generation unit.

[0051] In addition, an information processing device according to a 16th aspect is the information processing device according to the 15th aspect, in which the switching unit compares the number of sound source objects included in the sound information with the number of representative points set in the three-dimensional sound field, and switches between generating an output sound signal using the first generation unit and generating an output sound signal using the second generation unit depending on the comparison result.

[0052] According to this, the switching unit can compare the number of sound source objects contained in the sound information with the number of representative points set in the three-dimensional sound field, and appropriately switch between generating an output sound signal using the first generation unit or generating an output sound signal using the second generation unit.

[0053] Further, an information processing device according to a seventeenth aspect is the information processing device according to the fifteenth aspect, in which the switching unit switches to generate an output sound signal using the first generation unit when the head-related transfer function stored in the memory unit for storing the head-related transfer function does not satisfy a predetermined condition.

[0054] According to this, when the head-related transfer function in the storage unit does not satisfy a predetermined condition, the switching unit can switch to generating an output sound signal using the first generation unit.

[0055] In addition, an information processing device according to an 18th aspect is an information processing device according to any one of the 1st to 17th aspects, which is equipped with a path calculation unit that calculates a propagation path of a reproduced sound emitted from a sound source object by an audio signal, and calculates a synthesized sound that arrives at the user's position by indirect propagation of the reproduced sound according to the calculated propagation path of the reproduced sound, and the arrival direction of the synthesized sound.

[0056] According to this, the path calculation unit calculates the propagation path of the reproduced sound from the sound source object, and can calculate the synthetic sound that arrives at the user's position through indirect propagation of the reproduced sound and the direction of arrival of the synthetic sound according to the calculated propagation path of the reproduced sound.

[0057] In addition, an information processing device according to a 19th aspect is the information processing method described in the 18th aspect, which includes a switching unit that switches between generating an output sound signal using the first generation unit and generating an output sound signal using the second generation unit, and the switching unit switches between generating an output sound signal using the first generation unit and generating an output sound signal using the second generation unit, individually for each of the reproduced sound and the synthesized sound.

[0058] This makes it possible to switch between generating an output sound signal using the first generation unit and generating an output sound signal using the second generation unit individually for each of the reproduced sound and the synthesized sound.

[0059] An information processing device according to a twentieth aspect is the information processing method according to the eighteenth aspect, further comprising a switching unit that switches between generating an output sound signal using the first generation unit and generating an output sound signal using the second generation unit, wherein the path calculation unit calculates two or more synthetic sounds that arrive at the user's position by different indirect propagations and the directions of arrival of each of the two or more synthetic sounds, and the switching unit switches individually for each of the two or more synthetic sounds between generating an output sound signal using the first generation unit and generating an output sound signal using the second generation unit.

[0060] According to this, the path calculation unit calculates two or more synthetic sounds that arrive at the user's position through different indirect propagations and the directions of arrival of each of the two or more synthetic sounds, and it is possible to switch individually for each of the two or more synthetic sounds between generating an output sound signal using the first generation unit or generating an output sound signal using the second generation unit.

[0061] In addition, the information processing device of the 21st aspect is the information processing method described in the 18th aspect, which includes a switching unit that switches between generating an output sound signal using the first generation unit and generating an output sound signal using the second generation unit, and the switching unit compares the total number of reproduced sounds and synthesized sounds with the number of representative points set in the three-dimensional sound field, and switches between generating an output sound signal using the first generation unit and generating an output sound signal using the second generation unit depending on the comparison result.

[0062] This makes it possible to compare the total number of reproduced sounds and synthesized sounds with the number of representative points set in the three-dimensional sound field, and switch between generating an output sound signal using the first generation unit or generating an output sound signal using the second generation unit.

[0063] Furthermore, the information processing method according to the 22nd aspect is an information processing method executed by a computer that processes sound information to generate an output sound signal as sound arriving from a sound source object in a virtual three-dimensional sound field, and includes the steps of acquiring the position of the sound source object and an audio signal, the audio signal including a reproduced sound emitted from the sound source object by the audio signal; acquiring the position of a user in the three-dimensional sound field; calculating the arrival direction of the reproduced sound arriving from the position of the sound source object to the user's position; generating an output sound signal using a head-related transfer function corresponding to the calculated arrival direction and the reproduced sound; and generating the output sound signal using a head-related transfer function corresponding to a representative direction based on the position of a representative point set in the three-dimensional sound field and the position of the user, and the audio signal.

[0064] This can provide the same effects as the information processing device described above.

[0065] A program according to a twenty-third aspect is a program for causing a computer to execute the information processing method described above.

[0066] This makes it possible to achieve the same effects as the information processing method described above using a computer.

[0067] Further, an information processing device according to another aspect is an information processing device that processes sound information using a head-related transfer function to generate an output sound signal as a sound arriving from a sound source object in a virtual three-dimensional sound field, and includes a sound acquisition unit that acquires sound information including the position of the sound source object and a reproduced sound emitted from the sound source object, a position acquisition unit that acquires the position of a user in the three-dimensional sound field, an arrival direction calculation unit that calculates the relative arrival direction of the reproduced sound arriving from the position of the sound source object to the position of the user, and a third generation unit, and the head-related transfer function is stored in a storage unit for storing the head-related transfer function. When a new head-related transfer function is read, the adjustment amounts in the time shift adjustment and gain adjustment to be used in the conversion process are determined for the new head-related transfer function before storing it in the memory unit, and the read new head-related transfer function and the determined adjustment amounts are linked and stored in the memory unit, and a third generation unit applies time shift adjustment and gain adjustment to the reproduced sound using the adjustment amounts linked to the new head-related transfer function stored in the memory unit to convert it into a representative sound, and generates an output sound signal by convolving the head-related transfer function corresponding to the representative direction from each position of the representative point toward the user's position into the representative sound.

[0068] According to this, when a new head-related transfer function that is not stored in the storage unit for storing head-related transfer functions is read, adjustment amounts in time shift adjustment and gain adjustment used in conversion processing such as panning processing are determined for the new head-related transfer function, and the read new head-related transfer function and the determined adjustment amounts are linked and stored in the storage unit, and can be used for conversion processing. The new head-related transfer function has an adjustment amount appropriate for the head-related transfer function, and by determining such adjustment amount before starting the conversion processing (for example, when decoding the sound signal, when turning on the power of the sound reproduction system, or when initializing the sound reproduction system), it is possible to perform conversion processing with an appropriate adjustment amount while suppressing an increase in the processing amount.

[0069] In addition, the information processing device according to the 24th aspect is an information processing device that includes a memory unit that stores a time shift adjustment amount and a gain adjustment amount in association with each of a plurality of directions, an acquisition unit that acquires an audio signal and information on the position of a sound source object within a three-dimensional sound field, and a second generation unit that generates an output sound signal as sound arriving at the user's position from a second direction using the audio signal and the time shift adjustment amount and gain adjustment amount corresponding to a first direction based on the position of the sound source object and the position of the user within the three-dimensional sound field.

[0070] According to this, by storing a correspondence between each of a plurality of directions and a time shift adjustment amount and a gain adjustment amount, an output sound signal can be generated as sound arriving at the user's position from the second direction using the acquired sound signal and a time shift adjustment amount and gain adjustment amount corresponding to a first direction based on the position of the sound source object in the three-dimensional sound field and the position of the user, while suppressing an increase in processing volume with an appropriate adjustment amount.

[0071] Furthermore, the information processing device according to the 25th aspect is the information processing device described in the 24th aspect, in which the memory unit further stores a head-related transfer function corresponding to a second direction, and the second generation unit uses the audio signal, the time shift adjustment amount and gain adjustment amount corresponding to the first direction, and the head-related transfer function corresponding to the second direction to generate an output sound signal as sound arriving at the user's position from the second direction.

[0072] According to this, by further retaining an auxiliary memory unit or the like that includes a head-related transfer function corresponding to the second direction and storing such information, an output sound signal can be generated as sound arriving at the user's position from the second direction using the time shift adjustment amount and gain adjustment amount corresponding to the first direction and the head-related transfer function corresponding to the second direction.

[0073] In addition, an information processing device according to a 26th aspect is the information processing device described in the 24th aspect, in which the memory unit further stores head-related transfer functions corresponding to the second direction and directions other than the second direction, the second generation unit uses the audio signal, a time shift adjustment amount and a gain adjustment amount corresponding to the first direction, and the head-related transfer function corresponding to the second direction to generate an output sound signal as sound arriving at the user's position from the second direction, and the information processing device further includes a first generation unit, and the first generation unit uses the audio signal and the head-related transfer function corresponding to the first direction to generate an audio signal as sound arriving at the user's position from the first direction.

[0074] According to this, the storage unit further holds an auxiliary storage unit or the like including head-related transfer functions corresponding to the second direction and directions other than the second direction, and stores such information. The second generation unit uses the audio signal, the time shift adjustment amount and gain adjustment amount corresponding to the first direction, and the head-related transfer function corresponding to the second direction to generate an output sound signal as sound arriving at the user's position from the second direction. The information processing device further includes a first generation unit, and the first generation unit uses the audio signal and the head-related transfer function corresponding to the first direction to generate an audio signal as sound arriving at the user's position from the first direction. For example, if performing processing using the time shift adjustment amount and the gain adjustment amount effectively reduces the amount of processing, the second generation unit can be used; otherwise, the first generation unit can be used. In other words, from the perspective of processing amount, it is possible to effectively apply the conversion process by, for example, classifying conditions.

[0075] Furthermore, an information processing method according to yet another aspect is an information processing method that includes an auxiliary memory unit that stores a plurality of directions in association with a time shift adjustment amount and a gain adjustment amount, acquires an audio signal and information on the position of a sound source object within a three-dimensional sound field, and generates an output sound signal as sound arriving at the user's position from a second direction using the audio signal and a time shift adjustment amount and a gain adjustment amount corresponding to a first direction based on the position of the sound source object and the position of the user.

[0076] This can provide the same effects as the information processing device described in the twenty-third aspect.

[0077] A program according to yet another aspect is a program for causing a computer to execute the information processing method according to the above-described yet another aspect.

[0078] According to this, it is possible to achieve the same effect as the information processing method described in the above-mentioned further aspect by using a computer.

[0079] Furthermore, these comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0080] Hereinafter, embodiments will be described in detail with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not recited in independent claims will be described as optional components. Note that each figure is a schematic diagram and is not necessarily an exact illustration. Furthermore, in each figure, substantially identical components are assigned the same reference numerals, and duplicated descriptions may be omitted or simplified.

[0081] In the following description, elements may be assigned ordinal numbers such as first, second, and third. These ordinal numbers are assigned to elements in order to identify them and do not necessarily correspond to a meaningful order. These ordinal numbers may be rearranged, newly added, or removed as appropriate.

[0082] (Embodiment) [Overview] First, an overview of an audio reproduction system according to an embodiment will be described. Fig. 1 is a schematic diagram showing an example of use of an audio reproduction system according to an embodiment. Fig. 1 shows a user 99 using an audio reproduction system 100.

[0083] The audio reproduction system 100 shown in FIG. 1 is used simultaneously with a stereoscopic video reproduction device 300. By simultaneously viewing stereoscopic images and stereoscopic sound, the images enhance the auditory sense of realism, and the sounds enhance the visual sense of realism, allowing the user to experience the image and sound as if they were actually at the scene where they were captured. For example, when an image (moving image) of people having a conversation is displayed, it is known that even if the localization of the sound image (sound source object) of the conversation sound is not aligned with the person's mouth, the user 99 will perceive the conversation sound as coming from the person's mouth. In this way, the visual information can correct the position of the sound image, and the combination of the image and sound can enhance the sense of realism.

[0084] The three-dimensional video reproduction device 300 is an image display device worn on the head of the user 99. Therefore, the three-dimensional video reproduction device 300 moves integrally with the head of the user 99. For example, as shown in the figure, the three-dimensional video reproduction device 300 is a glasses-type device that is supported by the ears and nose of the user 99.

[0085] The three-dimensional video reproduction device 300 changes the displayed image in accordance with the movement of the user 99's head, thereby making the user 99 perceive the movement of his or her head in the three-dimensional image space. In other words, when an object in the three-dimensional image space is located in front of the user 99, if the user 99 turns to the right, the object moves to the left of the user 99, and if the user 99 turns to the left, the object moves to the right of the user 99. In this way, the three-dimensional video reproduction device 300 moves the three-dimensional image space in the opposite direction to the movement of the user 99.

[0086] The 3D video playback device 300 displays two images with a parallax difference to each of the user's 99's left and right eyes. The user 99 can perceive the three-dimensional position of an object on the image based on the parallax difference between the displayed images. Note that when the user 99 uses the audio playback system 100 with their eyes closed, for example, when using it to play healing sounds for sleep induction, the 3D video playback device 300 does not need to be used at the same time. In other words, the 3D video playback device 300 is not an essential component of the present disclosure. In addition to dedicated video display devices, the 3D video playback device 300 may also be a general-purpose mobile terminal owned by the user 99, such as a smartphone or tablet device.

[0087] Such general-purpose mobile terminals are equipped with not only a display for displaying images but also various sensors for detecting the terminal's posture and movement. Furthermore, they are also equipped with a processor for information processing, and are capable of connecting to a network to transmit and receive information to and from a server device such as a cloud server. In other words, the 3D video playback device 300 and the audio playback system 100 can be realized by combining a smartphone with general-purpose headphones or the like that do not have an information processing function.

[0088] As in this example, the head movement detection function, the video presentation function, the video information processing function for presentation, the sound presentation function, and the sound information processing function for presentation may be appropriately arranged in one or more devices to realize the 3D video reproduction device 300 and the sound reproduction system 100. If the 3D video reproduction device 300 is not required, it is sufficient to appropriately arrange the head movement detection function, the sound presentation function, and the sound information processing function for presentation in one or more devices. For example, the sound reproduction system 100 can be realized by a processing device such as a computer or smartphone having a sound information processing function for presentation, and headphones or the like having a head movement detection function and a sound presentation function.

[0089] The sound reproduction system 100 is a sound presentation device that is worn on the head of the user 99. Therefore, the sound reproduction system 100 moves integrally with the head of the user 99. For example, the sound reproduction system 100 in this embodiment is a so-called over-ear headphone type device. Note that there are no particular limitations on the form of the sound reproduction system 100, and it may be, for example, two earplug-type devices that are worn independently on the left and right ears of the user 99.

[0090] The sound reproduction system 100 changes the sound presented in accordance with the movement of the head of the user 99, thereby making the user 99 perceive as if he or she is moving his or her head within the three-dimensional sound field. For this reason, as described above, the sound reproduction system 100 moves the three-dimensional sound field in the opposite direction to the movement of the user 99.

[0091] Here, when the user 99 moves within the three-dimensional sound field, the position of the sound source object relative to the position of the user 99 within the three-dimensional sound field changes. As a result, each time the user 99 moves, it is necessary to perform calculation processing based on the positions of the sound source object and the user 99 to generate an output sound signal for playback. Since such processing typically requires a huge amount of processing, in the present disclosure, a panning process is applied as one type of conversion processing to represent the reproduced sound as a representative sound from a representative point, in order to reduce the amount of processing. As a result, it is possible to allow the user 99 to perceive the reproduced sound from the sound source object simply by convolving a head-related transfer function with the representative sound. Hereinafter, in the present embodiment, a case will be described in which panning processing is used as an example of conversion processing. However, the conversion processing is not limited to panning processing, and any conversion processing can be applied as long as the conversion processing is expected to reduce the amount of processing, depending on the conditions.

[0092] If the number of representative points preset in the three-dimensional sound field is fewer than the number of sound source objects, the amount of convolution of head-related transfer functions is reduced, which can contribute to a reduction in the amount of processing. However, since the panning process itself requires processing that is not required when convolving head-related transfer functions with the original reproduced sound, the reduction in processing amount can only be achieved when the number of sound source objects is several times greater than the number of representative points. In addition, there are several conditions for achieving the reduction in processing amount, so in the present disclosure, when a reduction in processing amount is not expected, the output sound signal is generated in normal mode, in which head-related transfer functions are convolved with the reproduced sound.

[0093] [Configuration] Next, the configuration of the sound reproduction system 100 according to this embodiment will be described with reference to Fig. 2. Fig. 2 is a block diagram showing the functional configuration of the sound reproduction system according to this embodiment.

[0094] As shown in FIG. 2, the sound reproduction system 100 according to this embodiment includes an information processing device 101, a communication module 102, a detector 103, and a driver 104.

[0095] The information processing device 101 is an arithmetic device for performing various signal processing in the sound reproduction system 100. The information processing device 101 includes a processor and a memory, such as a computer, and is realized by the processor executing a program stored in the memory. Execution of this program provides functions related to each functional unit described below.

[0096] The information processing device 101 includes an acquisition unit 111, a path calculation unit 121, an output sound generation unit 131, a signal output unit 141, and a storage unit 105. Details of each functional unit included in the information processing device 101 will be described below together with details of the configuration other than the information processing device 101.

[0097] The communication module 102 is an interface device for accepting input of sound information to the sound reproduction system 100. The communication module 102 includes, for example, an antenna and a signal converter, and receives sound information from an external device via wireless communication. The communication module 102 may also receive a set of head-related transfer functions, such as an SOFA file, from the external device. More specifically, the communication module 102 receives, using an antenna, a wireless signal representing sound information converted into a format for wireless communication, and reconverts the wireless signal into sound information using a signal converter. In this way, the sound reproduction system 100 acquires sound information and a set of head-related transfer functions from the external device via wireless communication. The sound information and the set of head-related transfer functions acquired by the communication module 102 are acquired by the acquisition unit 111. In this way, the acquisition unit 111 is an example of a sound acquisition unit. The sound information is input to the information processing device 101 in the above manner. Note that communication between the sound reproduction system 100 and the external device may also be performed via wired communication.

[0098] The sound information acquired by the sound reproduction system 100 is composed of information about the sound (sound signal) reproduced by the sound reproduction system 100 and information about the localization position when the sound image of the sound is localized at a predetermined position within a three-dimensional sound field (i.e., perceived as sound coming from a predetermined direction). The information about the reproduced sound may be, for example, a sound signal encoded in a predetermined format such as MPEG-H 3D Audio (ISO / IEC 23008-3), or an unencoded PCM signal. The information about the localization position can also be interpreted as information about the sound source object. In other words, the sound information includes the position of the sound source object within the three-dimensional sound field and the sound produced by the sound source object. The sound information may also include a flag for determining whether or not to apply panning processing. This flag will be described later.

[0099] As described above, sound information is obtained as input data, and includes an audio signal (acoustic signal), which is information about the reproduced sound, and other information, such as information about the position of a sound source object in a three-dimensional sound field. The other information may also include information for defining a three-dimensional sound field. Therefore, the other information may be collectively referred to as information about space (spatial information), including information about the position of a sound source object and information for defining a three-dimensional sound field. When the audio signal is viewed as the main focus, the input data can be said to be sound information in which other information (metadata) is attached to the audio signal. When the spatial information is viewed as the main focus, the input data can be said to be information in which the audio signal is attached to the spatial information. Alternatively, since the input data has both aspects, the input data can also be considered as sound spatial information.

[0100] As a specific example, the sound information includes information about a plurality of sounds including a first reproduced sound and a second reproduced sound, and the sound images of the respective reproduced sounds are localized so that they are perceived as coming from different positions in a three-dimensional sound field. Therefore, the sound source object of the first reproduced sound is localized at a first position in the three-dimensional sound field, and the sound source object of the second reproduced sound is localized at a second position in the three-dimensional sound field. In this way, the sound information may include a plurality of sounds.

[0101] The stereoscopic sound, together with the image viewed using the stereoscopic video playback device 300, can enhance the sense of realism of the content being viewed. The sound information may include only information about the playback sound. In this case, information about a predetermined position may be acquired separately. As described above, the sound information includes first sound information about a first playback sound and second sound information about a second playback sound. However, sound images may be localized at different positions in a three-dimensional sound field by acquiring and simultaneously playing multiple pieces of sound information that include these separately. As described above, the form of the input sound information is not particularly limited, and the sound playback system 100 may be provided with an acquisition unit 111 that can accommodate various forms of sound information.

[0102] The sound information immediately after acquisition includes an audio signal related to the direct sound, and is converted into sound information including audio signals such as reverberation, primary reflection, and diffraction through a conversion process that calculates the secondary sound. Alternatively, sound information including audio signals related to such secondary sounds may be acquired in addition to sound information including audio signals related to the direct sound. The conversion process that adds secondary sounds to sound information through calculation uses information on the spatial environment conditions of the three-dimensional sound field (e.g., the position, reflection, and diffraction characteristics of objects in the three-dimensional sound field). In this way, secondary sounds are computationally generated from sound information related to one reproduced sound based on the spatial environment conditions of the three-dimensional sound field. One secondary sound may generate another secondary sound through propagation of that secondary sound. The information on the spatial environment conditions is part of the spatial information and is acquired together with the audio signal from the input sound information.

[0103] The arrival direction of the secondary sound includes additional information such as what object the sound is reflected from and the rate of attenuation at the time of reflection. The additional information is included in the arrival direction of the secondary sound calculated from the input sound information. In other words, the additional information is computationally generated and obtained from the sound information.

[0104] To summarize the spatial information, the spatial information includes further information such as the spatial position of a sound source object in a space (three-dimensional sound field) (information on the position of the sound source object), sound reflection and diffraction characteristics at the sound source object (together with information on the conditions of the spatial environment), and the size of the three-dimensional sound field. Based on the spatial information, the path calculation unit 121 generates secondary sounds depending on which sound source object the reproduced sound is reflected or diffracted by, and calculates, as additional information, the direction from which the secondary sounds arrive and the volume of the secondary sounds after attenuation by reflection or diffraction. The sound information (input data) includes spatial information in the form of audio signals and accompanying metadata. As described above, the spatial information includes, as information other than the audio signal, information necessary to turn the sound into stereophonic sound and position the sound source object in a three-dimensional sound field, and / or information used to calculate information necessary to turn the sound into stereophonic sound and position the sound source object in a three-dimensional sound field.

[0105] An example of the acquisition unit 111 will now be described with reference to Fig. 3. The acquisition unit 111 is a processing unit that acquires information necessary for generating an output sound, and the information necessary for generating an output sound includes sound information, a set of head-related transfer functions, sensing information, etc. Fig. 3 is a block diagram showing the functional configuration of the acquisition unit according to an embodiment. As shown in Fig. 3, the acquisition unit 111 in this embodiment includes, for example, an encoded sound information input unit 112, a decoding processing unit 113, and a sensing information input unit 114.

[0106] The encoded sound information input unit 112 is a processing unit to which the encoded (in other words, encoded) sound information acquired by the acquisition unit 111 is input. The encoded sound information includes a sound signal encoded in a predetermined format, such as MPEG-H 3D Audio (ISO / IEC 23008-3). The encoded sound information input unit 112 outputs the input sound information to the decoding processing unit 113. The decoding processing unit 113 is a processing unit that decodes (in other words, decodes) the sound information output from the encoded sound information input unit 112 to generate a reproduced sound (sound signal), the position of a sound source object, and a flag contained in the sound information in a format used for subsequent processing. The sensing information input unit 114 will be described below along with the function of the detector 103.

[0107] The processing performed by the encoded sound information input unit 112 and the decoding processing unit 113 may be executed by a device external to the information processing device 101. In other words, the acquisition unit 111 only needs to acquire sound information, and may acquire sound information that has been decoded by an external device via the communication module 102. Also, although an example in which the sound information is encoded has been described, the sound information does not have to be encoded. For example, information on the reproduced sound may be acquired as an unencoded sound signal such as a PCM signal.

[0108] The sound signal and spatial information included in the sound information may be acquired in separate streams or files, or may be acquired in the same stream or file.

[0109] In addition, the acquisition unit 111 may be equipped with a head-related transfer function input unit (not shown), and may acquire a set of head-related transfer functions acquired from the outside via the communication module 102 and output them to the memory unit 105.

[0110] The detector 103 is a device for detecting the speed of movement of the head of the user 99. The detector 103 is configured by combining various sensors used for detecting movement, such as a gyro sensor and an acceleration sensor. In this embodiment, the detector 103 is built into the sound reproduction system 100. However, the detector 103 may be built into an external device, such as a three-dimensional image reproduction device 300 that operates in response to the movement of the head of the user 99 in the same way as the sound reproduction system 100. In this case, the detector 103 does not need to be included in the sound reproduction system 100. Alternatively, the detector 103 may be an external imaging device or the like that captures the movement of the head of the user 99 and detects the movement of the user 99 by processing the captured image.

[0111] The detector 103 is, for example, fixed integrally to the housing of the sound reproduction system 100 and detects the speed of movement of the housing. After the sound reproduction system 100 including the housing is worn by the user 99, it moves integrally with the head of the user 99, and as a result, the detector 103 can detect the speed of movement of the head of the user 99.

[0112] The detector 103 may detect, for example, the amount of head movement of the user 99 as the amount of rotation about at least one of three axes that are orthogonal to each other in three-dimensional space as the rotation axis, or may detect the amount of displacement about at least one of the three axes as the displacement direction. Furthermore, the detector 103 may detect both the amount of rotation and the amount of displacement as the amount of head movement of the user 99.

[0113] The sensing information input unit 114 acquires the movement speed of the user 99's head from the detector 103. More specifically, the sensing information input unit 114 acquires the amount of head movement of the user 99 detected by the detector 103 per unit time as the movement speed. In this way, the sensing information input unit 114 acquires at least one of the rotation speed and the displacement speed from the detector 103. The amount of head movement of the user 99 acquired here is used to determine the position and posture (in other words, coordinates and orientation) of the user 99 in the three-dimensional sound field. Therefore, the acquisition unit 111 also functions as a position acquisition unit via the sensing information input unit 114. In the sound reproduction system 100, the relative position of the sound image object with respect to the user 99 is determined based on the determined coordinates and orientation of the user 99, and sound is reproduced. Specifically, the above functions are realized by the path calculation unit 121 and the output sound generation unit 131.

[0114] The path calculation unit 121 includes an arrival direction calculation function that calculates the relative arrival direction of the reproduced sound from the position of the sound source object to the position of the user 99 based on the determined coordinates and orientation of the user 99, and a synthetic sound calculation function that calculates a propagation path from the sound source object and calculates a synthetic sound that arrives at the position of the user 99 by indirect propagation of the reproduced sound according to the calculated propagation path of the reproduced sound, and the arrival direction of the synthetic sound. In other words, the path calculation unit 121 is also an example of an arrival direction calculation unit.

[0115] The path calculation unit 121 may be realized by any processing as long as it can calculate the arrival direction of the reproduced sound when the reproduced sound reaches the user as direct sound, and can calculate the arrival direction of a synthesized sound (for example, a reflected sound, a diffracted sound, a reverberant sound, etc.) that arrives at the position of the user 99 due to indirect propagation of the reproduced sound. The path calculation unit 121 determines from which direction in the three-dimensional sound field the reproduced sound and the synthesized sound are to be perceived by the user 99 as coming from, based on the coordinates and orientation of the user 99, and processes the sound information so that the output sound signal is perceived as such a sound when it is reproduced.

[0116] The output sound generating unit 131 is a processing unit that processes information about the reproduced sound included in the sound information to generate an output sound signal.

[0117] An example of the output sound generation unit 131 will now be described with reference to FIG. 4. FIG. 4 is a block diagram showing the functional configuration of the output sound generation unit according to the embodiment. As shown in FIG. 4, the output sound generation unit 131 according to the present embodiment includes, for example, a switching unit 132, a first generation unit 133, and a second generation unit 134. The switching unit 132 is a processing unit for switching whether to use the first generation unit 133 or the second generation unit 134 when generating an output sound signal. Therefore, the switching unit 132 has a function of acquiring information for determining whether to use the first generation unit 133 or the second generation unit 134.

[0118] The first generation unit 133 is a processing unit used when a head-related transfer function is directly convolved with a reproduced sound without applying panning processing. The first generation unit 133 is a processing unit used when generating an output sound signal in a so-called "normal mode". The first generation unit 133 acquires a reproduced sound and a head-related transfer function corresponding to the direction from which the reproduced sound arrives, and performs convolution processing of the acquired head-related transfer function with the reproduced sound to generate an output sound signal.

[0119] The second generation unit 134 is a processing unit used when applying a panning process to perform a conversion process of converting a playback sound into a representative sound, and then convolving a head-related transfer function with the converted representative sound. The second generation unit 134 is a processing unit used when generating an output sound signal in a so-called "low processing mode". The second generation unit 134 acquires the playback sound and the position of the representative point, and performs a conversion process to a representative sound for reproducing the playback sound using sound from the representative point.

[0120] For example, if a sound source object is located midway between two representative points, a sound is generated so that the same sound as the playback sound is emitted from each of the two representative points. Then, a representative sound can be generated by adjusting the gain of the generated sound to match the position of the sound source object. The conversion from the playback sound to the representative sound is not limited to this example. For example, the conversion from the playback sound to the representative sound may be performed by performing time shift adjustment and gain adjustment, as described below, or any other existing conversion may be used as long as the conversion to the representative sound is performed to reproduce the playback sound using sounds from the representative points. Examples of conversions that perform time shift adjustment and gain adjustment will be described later. The second generation unit 134 acquires the same number of representative sounds as the number of representative points obtained by the conversion and head-related transfer functions corresponding to the representative directions from each representative point to the position of the user 99, and performs a convolution process of the acquired head-related transfer functions on the representative sounds to generate an output sound signal.

[0121] Referring again to FIG. 2 , the output sound generation unit 131 acquires a head-related transfer function used to generate an output sound signal from the storage unit 105. The storage unit 105 is an information storage device that functions both as a storage device for storing information and as a storage controller that reads out the stored information and outputs it to other processing units included in the information processing device. The storage unit 105 may be interpreted as a memory included in the information processing device 101. The storage unit 105 stores the head-related transfer functions acquired by the acquisition unit 111 for each direction of arrival of the sound to the user 99. The head-related transfer functions included in the storage unit 105 are a set of general-purpose head-related transfer functions that can be used by everyone, a set of head-related transfer functions optimized for each individual user 99, or a set of publicly available head-related transfer functions. The storage unit 105 receives an inquiry from the output sound generation unit 131 using the direction of arrival as a query, and outputs the head-related transfer function corresponding to the direction of arrival to the output sound generation unit 131. Furthermore, the output sound generation unit 131 may output the entire set of head-related transfer functions or output the characteristics of the set of head-related transfer functions themselves in response to an inquiry from the switching unit 132. The set of head-related transfer functions may be acquired from the outside by the acquisition unit 111 in the format of, for example, an SOFA file, and then stored in the storage unit 105.

[0122] The signal output unit 141 is a functional unit that outputs the generated output sound signal to the driver 104. The signal output unit 141 generates a waveform signal by performing signal conversion from a digital signal to an analog signal based on the output sound signal, and then causes the driver 104 to generate sound waves based on the waveform signal, thereby presenting sound to the user 99. The driver 104 includes, for example, a diaphragm and a drive mechanism such as a magnet and a voice coil. The driver 104 operates the drive mechanism in response to the waveform signal, causing the drive mechanism to vibrate the diaphragm. In this way, the driver 104 generates sound waves by vibrating the diaphragm in response to the output sound signal (this means "reproducing" the output sound signal; in other words, "reproducing" does not include perception by the user 99). The sound waves propagate through the air and reach the ears of the user 99, and the user 99 perceives the sound.

[0123] [Operation] Next, the operation of the above-described sound reproduction system 100, particularly the information processing device 101, will be described with reference to FIGS.

[0124] 5 is a flowchart showing a first operation example of the information processing device according to the embodiment. In the first operation example of the information processing device 101, the acquisition unit 111 acquires sound information via the communication module 102 (step S11). The sound information is decoded by the decoding processing unit 113 into information about the reproduced sound, information about the position of the sound source object, and a flag, and generation of an output sound signal is initiated.

[0125] The sensing information input unit 114 acquires information about the position of the user 99 (step S12). The path calculation unit 121 calculates the arrival direction of the reproduced sound based on the positions of the sound source objects and the user 99 (step S13). Here, the flag included in the sound information is a flag added by the creator when creating the sound information. This flag is a flag for specifying whether to generate an output sound signal by the first generation unit 133 or the second generation unit 134. Since the creator knows what sound source objects are included in the original sound information, it is possible to add a flag that causes the first generation unit 133 to generate an output sound signal, for example, because the number of sound source objects included in the sound information is quite small.

[0126] Alternatively, the creator may add a flag to cause the second generation unit 134 to generate an output sound signal because, for example, the sound information contains a large number of sound source objects. If the flag specifies that the first generation unit 133 is to generate an output sound signal, it may be treated as equivalent to a flag specifying that the second generation unit 134 is not to generate an output sound signal. Furthermore, if the flag specifies that the second generation unit 134 is to generate an output sound signal, it may be treated as equivalent to a flag specifying that the first generation unit 133 is not to generate an output sound signal.

[0127] In the operation example of FIG. 5 , it is determined whether the flag specifies that the first generation unit 133 is to be used (step S14). If the flag specifies that the first generation unit 133 is to be used (Yes in S14), the first generation unit 133 generates an output sound signal (step S15). On the other hand, if the flag does not specify that the first generation unit 133 is to be used (No in S14), the second generation unit 134 generates an output sound signal (step S16). Note that the output sound generation unit 131 may use the switching unit 132 to make the determination in step S14 and switch between generating an output sound signal using the first generation unit 133 and generating an output sound signal using the second generation unit 134. Alternatively, a flag determination unit (not shown) may make the determination in step S14, and depending on the determination result, the acquisition unit 111 may directly input sound information to the first generation unit 133 or the second generation unit 134. In other words, it is not essential to switch between generating an output sound signal using the first generating unit 133 and generating an output sound signal using the second generating unit 134 .

[0128] Next, Fig. 6 is a flowchart showing a second operation example of the information processing device according to the embodiment. The operation example shown in Fig. 6 is the same as Fig. 5 except that step S24 is executed instead of step S14, and therefore description thereof will be omitted. In the operation example of Fig. 6, after step S13, the number of sound source objects included in the sound information is compared with the number of representative points set in the three-dimensional sound field, and whether to execute step S15 or step S16 is switched depending on whether the comparison result satisfies a predetermined condition. Specifically, the switching unit 132 acquires the sound information and counts the number of sound source objects.

[0129] The switching unit 132 also acquires the number of representative points set in the three-dimensional sound field (the number of representative points is stored as setting information in a storage unit (not shown) or the like). The switching unit 132 then compares the number of sound source objects with the number of representative points. The switching unit 132 determines whether the comparison result satisfies a predetermined condition, for example, by determining whether the number of sound source objects is less than a coefficient multiplied by the number of representative points (step S24). If the predetermined condition is satisfied (the number of sound source objects is less than a coefficient multiplied by the number of representative points) (Yes in S24), the switching unit 132 switches to execute step S15. If the predetermined condition is not satisfied (the number of sound source objects is equal to or more than a coefficient multiplied by the number of representative points) (No in S24), the switching unit 132 switches to execute step S16. The coefficient multiplier here is set assuming that generating an output sound signal in the normal mode is equivalent to or more advantageous than generating an output sound signal in the low-processing mode in terms of processing amount. As explained above, the panning process has its own processing load, and therefore the coefficient varies depending on the panning process implemented, within a range from several times, such as 1x, 3x, or 5x, to several tens of times, such as 10x, 30x, or 50x. In other words, the coefficient may be set to a value appropriate to the mode of panning process.

[0130] Furthermore, as shown in FIG. 7 , in a three-dimensional sound field (the outermost rectangle in the figure), when a reproduced sound from a sound source object represented by a hollow circle arrives at the position of a user 99 represented by a filled circle, in addition to the direct sound that the reproduced sound arrives directly, reflected sound, diffracted sound (affected by objects in the space represented by inverted triangles), reverberation (not shown), and other sounds are also generated due to indirect propagation. At this time, the reproduced sound arrives from all directions due to different indirect propagations such as reflected sound, diffracted sound, and reverberation. Therefore, generating an output sound signal for the reproduced sound may sound unnatural. Therefore, in this embodiment, the reproduced sound that arrives due to indirect propagation is generated as a synthesized sound. Such a synthesized sound must also be perceived as a sound coming from the appropriate direction and must be included in the output sound signal, just like the reproduced sound from the sound source object. In other words, it is necessary to determine whether or not panning should be performed on a pair of a reproduced sound and a synthesized sound, or on a pair of synthesized sounds resulting from different indirect propagations. 6, the total number of reproduced sounds and synthesized sounds may be used as the object to be compared with the number of representative points in step S24. In this case, too, whether to execute step S15 or step S16 is switched depending on whether a predetermined condition such as a coefficient multiplication is satisfied.

[0131] Next, FIG. 8 is a flowchart showing a third operation example of the information processing device according to the embodiment. The operation example shown in FIG. 8 is the same as that shown in FIG. 5 except that step S34 is executed instead of step S14, and therefore description thereof will be omitted. In the operation example shown in FIG. 8 , after step S13, whether to execute step S15 or step S16 is switched depending on whether the head-related transfer functions stored in the storage unit 105 are dense enough to fully achieve the effect of reducing the amount of processing by applying the panning process. Specifically, the switching unit 132 queries the storage unit 105 and reads out a set of head-related transfer functions or characteristic information related to the density of the head-related transfer functions. Then, the switching unit 132 determines whether the head-related transfer functions stored in the storage unit 105 are sparser or denser than a preset threshold.

[0132] That is, the switching unit 132 determines whether the density characteristics of the head-related transfer functions satisfy a predetermined condition based on whether they are denser than a density threshold (step S34). If the predetermined condition is not satisfied (the density characteristics of the head-related transfer functions are sparser than the threshold) (Yes in S34), the switching unit 132 switches to execute step S15. If the predetermined condition is satisfied (the density characteristics of the head-related transfer functions are denser than the threshold) (No in S34), the switching unit 132 switches to execute step S16. This density threshold is set, for example, depending on whether a head-related transfer function is included that is denser than a step angle in at least one direction, such as 5 degrees, 10 degrees, or 15 degrees in the horizontal direction and 5 degrees, 10 degrees, or 15 degrees in the vertical direction. The density threshold depends on the arrival direction of the reproduced sound from the sound source object included in the sound information and also on the representative direction from the set representative point. Therefore, it may be set appropriately depending on the arrival direction of the reproduced sound from the sound source object included in the sound information and the representative direction from the representative point.

[0133] [Specific Example of Panning Processing] In panning processing, sounds reproduced from multiple sound source objects are represented by representative sounds from multiple representative directions. For example, two or three directions can be used as the representative directions. Specifically, in panning processing, the number of representative points is reduced to a number less than the number of sound source objects, and the reproduced sounds can be perceived as sounds coming from the direction of arrival using only the head-related transfer functions of the representative directions for these representative points. Panning processing may also be interpreted as a process of distributing reproduced sounds to representative points (representative directions). Specifically, sound signals of the reproduced sounds associated with the positions of the sound source objects are distributed to the positions of the representative points, and representative sounds arriving at the listener from the representative points (representative directions) are generated. Here, the representative direction is a direction determined by the relationship between the listener's head direction and the position of the representative point. For example, it refers to the direction of the representative point as seen from the front of the listener. It may also be rephrased as the direction of the representative point when the direction in which the listener's face is facing is used as a reference, or the direction of the representative point as seen from the listener's eyes.

[0134] In this case, the panning process calculates a time shift (delay) that maximizes the cross-correlation between the head-related transfer function of the sound source object in the arrival direction and the head-related transfer function of the representative direction. The time shift obtained here, or a time shift obtained by adding a negative sign to this time shift, is applied to the reproduced sound from the sound source object, and the subsequent processing is performed assuming that the signal after the time shift is in the representative direction.

[0135] This time shift may be a time shift shorter than the sampling period (a shift in which the sample position is expressed by a decimal number; hereinafter, referred to as a "decimal shift"). This decimal shift can be performed by oversampling.

[0136] Here, in the panning process, a gain is applied to the signal of a representative direction obtained by time-shifting the reproduced sound of the sound source object, and the sum of the signals calculated for each representative point convolved with the head transfer function at each representative point is calculated, thereby synthesizing a signal equivalent to the reproduced sound of the sound source object convolved with the head transfer function of the arrival direction.

[0137] On the other hand, in the panning process, when synthesizing the head-related transfer function (vector) of the arrival direction by the sum of the head-related transfer functions (vector) of the representative direction, the gain may be calculated by orthogonalizing the error signal vector between the synthesized head-related transfer function (vector) and the head-related transfer function (vector) of the arrival direction to the head-related transfer function (vector) of the representative direction. Note that the head-related transfer function (vector) is a time waveform of a head impulse response, which is an expression of the head-related transfer function in the time domain, as a vector. Hereinafter, this head-related transfer function (vector) will also be simply referred to as a "head-related transfer function vector."

[0138] In the panning process, this gain is corrected so that the energy balance of the head-related transfer functions from the position of the sound source object to the left and right ears of the user 99 is maintained in the head-related transfer function synthesized by the panning process using head-related transfer functions from multiple representative points. In other words, in the panning process, the gain may be corrected so that the energy balance of the head-related transfer functions of the left and right ears of the user 99 due to the sound source object is maintained in the head-related transfer function synthesized by the panning process.

[0139] In this embodiment, in the panning process, for each arrival direction of the sound source object, a gain value to be multiplied by the head transfer function of the representative direction and a time shift value to be applied to the head transfer function of the representative direction can be calculated and stored in table data (head transfer function table or adjustment amount table) described later.

[0140] Then, in the panning process, each sound source object is time-shifted by a time shift value and a gain value corresponding to the direction of arrival of each sound source object, and the time shifts and gains are multiplied and summed to generate a sum signal. In the panning process, this sum signal is treated as being present at the position of the representative point. In the panning process, a head-related transfer function in the direction of the representative point is convolved with this sum signal to generate a signal at the ear of the user 99.

[0141] In addition, the panning process may use a gain calculated so as to minimize the energy or L2 norm of the error signal vector between the synthesized HRIR vector and the HRIR vector of the sound source direction. The HRIR vector is a vector whose elements are values ​​obtained by sampling the time domain waveform of the head-related transfer function at a sampling frequency of 48 kHz.

[0142] In the above example, the time shift and the gain that maximize the cross-correlation are calculated by using the head-related transfer function itself. On the other hand, the time shift and / or the gain may be calculated by applying a weighting filter on the frequency axis.

[0143] That is, when calculating the time shift and gain that maximize the cross-correlation, it is possible to use a filter that has been subjected to a weighting filter on the frequency axis (hereinafter also referred to as a "frequency weighting filter").

[0144] It is preferable to use a frequency weighting filter that has a cutoff frequency near or slightly higher than the frequency band where human hearing sensitivity is high, and attenuates the higher frequency band, i.e., the frequency band where human hearing sensitivity decreases. For example, it is preferable to use a low-pass filter (LPF) with a cutoff frequency of 3000 Hz to 6000 Hz and approximately 6 dB / oct (octave) to 12 dB / oct.

[0145] In the panning process, adjustment amounts in the time shift adjustment and the gain adjustment may be determined according to a set of head-related transfer functions included in the storage unit 105, and the time shift adjustment and the gain adjustment may be applied to the reproduced sound using the determined adjustment amounts to convert it into a representative sound. Since the optimal values ​​of the adjustment amounts in the time shift adjustment and the gain adjustment used in the panning process change depending on the head-related transfer functions, first, when a set of head-related transfer functions such as a SOFA file is acquired or when a set of head-related transfer functions included in the storage unit 105 is read, the adjustment amounts in the time shift adjustment and the gain adjustment corresponding to each head-related transfer function included in the set of head-related transfer functions are determined. Thereafter, the same adjustment amounts can be reused as long as this set of head-related transfer functions is used, which is advantageous in terms of processing amount. More specifically, when, for example, three representative directions are used in the panning process, first, when a set of head-related transfer functions is acquired (e.g., at the time of initialization), a plurality of representative direction candidates (e.g., eight directions) are selected from the directions of the celestial sphere included in the set of head-related transfer functions. Next, for each of the head-related transfer functions of the celestial sphere included in the set of head-related transfer functions, it is determined which three directions among the plurality of representative direction candidates will be used as the representative direction. Next, for each of the celestial sphere directions included in the set of head-related transfer functions, the adjustment amounts for time shift adjustment and gain adjustment for allocating signals to the identified three representative directions are calculated, and the calculated adjustment amounts are determined as the adjustment amounts associated with each of the celestial sphere directions included in the set of head-related transfer functions.

[0146] The head-related transfer function table is an example of table data including head-related transfer functions stored in the storage unit 105. The head-related transfer function table stores the head-related transfer functions together with the adjustment amounts in the time shift adjustment and the gain adjustment determined according to the head-related transfer functions, which are linked to each other. That is, the head-related transfer function table may be constructed by calculating the adjustment amounts in the time shift adjustment and the gain adjustment in advance for each head-related transfer function included in the storage unit 105. In this way, table data of the head-related transfer function table linking each head-related transfer function with the adjustment amount may be stored in the storage unit 105. Note that the calculation of the adjustment amount for each head-related transfer function may be performed by the second generation unit 134 or the decoding processing unit 113. Alternatively, the calculation of the adjustment amount may be performed by an external device and stored in the memory of the external device. In this case, the memory of the external device corresponds to an example of a storage unit.

[0147] Alternatively, the adjustment amounts in the time shift adjustment and the gain adjustment may be calculated in advance for each direction of a plurality of head-related transfer functions included in the set of head-related transfer functions, and an adjustment amount table linking each of a plurality of representative directions with the adjustment amount for each direction of a plurality of head-related transfer functions included in the set of head-related transfer functions may be constructed and stored in the storage unit 105. In this case, since the plurality of representative directions are representative directions (e.g., three directions) selected from a plurality of representative direction candidates (e.g., eight directions), the adjustment amount table includes information on which representative directions (e.g., three directions) have been selected from the plurality of representative direction candidates (e.g., eight directions) for each direction of the celestial sphere included in the set of head-related transfer functions. The adjustment amount table may include table data linking the head-related transfer functions for each of the plurality of representative directions with the adjustment amounts in the time shift adjustment and the gain adjustment for each direction of a plurality of head-related transfer functions included in the set of head-related transfer functions, or the head-related transfer functions for each of the plurality of representative directions may be extracted at the time of rendering or at the time of system initialization from a set of head-related transfer functions for the celestial sphere (multiple directions) that has been acquired in advance and stored in the storage unit 105. Furthermore, a set of head-related transfer functions may be acquired from outside when the system is initialized, and an adjustment amount table may be constructed at the time of initialization and then stored in the storage unit 105. In this case, the adjustment amount table stored in the storage unit 105 may be read and used when processing the output of an audio signal.

[0148] The initialization and output processing in this disclosure will be described below.

[0149] For example, in an embodiment, the spatial information update process (information update thread) and the audio signal output process (audio thread) with added acoustic processing may be executed in a single thread, or may be executed in different threads. When these two processes are executed in different threads, the thread startup frequency may be set individually, or the processes may be executed in parallel.

[0150] In particular, when these two processes are executed in separate threads, it is possible to allocate computing resources preferentially to the output process of the audio signal to which acoustic processing has been added. This makes it possible to safely execute sound output processes that cannot tolerate even the slightest delay, such as when a delay of one sample (0.02 msec) would cause a popping noise.

[0151] In this case, the allocation of computational resources to the spatial information update process is limited. However, because the spatial information update is a less frequent process compared to the audio signal output process (for example, a process such as updating the listener's facial orientation), it does not necessarily need to be performed in near real time with no delay, as is the case with the audio signal output process. Therefore, even if the allocation of computational resources is limited, it does not have a significant impact on acoustic quality.

[0152] The spatial information may be updated periodically at preset times or intervals, or when preset conditions are met. The spatial information may also be updated manually by a listener or a sound space manager, or may be updated in response to a change in an external system.

[0153] For example, the spatial information may be updated when a listener operates a controller to instantly warp the position of the listener's avatar or instantly advance or reverse the time. Alternatively, the spatial information may be updated when an administrator of the virtual space suddenly changes the environment of the space. In these cases, the thread for updating the spatial information may be started as a one-off interrupt process in addition to being started periodically.

[0154] For example, the space information update process may be performed when the virtual space is created (when the software is created), when virtual space information (scene information) is read, when virtual space processing starts (when the software is launched or when rendering starts), or when an information update thread that occurs periodically in virtual space processing occurs, etc. Furthermore, the virtual space may be created when the virtual space is constructed before acoustic processing starts, when virtual space information (spatial information) is acquired, or when software is acquired.

[0155] As described above, in the present disclosure, there are three processing threads (in other words, workflows) that occur with different frequencies: a processing thread that occurs irregularly, a processing thread that occurs regularly but infrequently, such as updating the listener's facial orientation, and a processing thread that occurs regularly but frequently, such as sound output processing. The initialization process in the present disclosure corresponds to the irregularly occurring processing thread among the above.

[0156] Furthermore, the adjustment amount table may be a table including, for example, information on which of a plurality of representative directions to distribute a sound signal arriving at the position of a listener from the direction of each head-related transfer function of the set of head-related transfer functions of the celestial sphere, and information on a time shift adjustment amount and a gain adjustment amount to be multiplied by the sound signal for each representative direction when distributing the sound signal.

[0157] When performing the convolution process of the head-related transfer function, the adjustment amount table stored in the memory unit 105 is referenced, and the adjustment amounts for the time shift adjustment and gain adjustment linked to the head-related transfer function of the direction to be applied are used, thereby eliminating the need to calculate the adjustment amount for each convolution process and contributing to a reduction in the amount of processing.

[0158] Note that the embodiment of the present invention can also be applied to a new set of head-related transfer functions (e.g., a SOFA file) that is not included in the storage unit 105. When decoding a sound signal, when powering on the sound reproduction system 100, or when initializing the sound reproduction system 100, the head-related transfer functions of the entire three-dimensional sound field may be newly read, the representative direction may be determined anew using the method disclosed in this embodiment or another method, and an adjustment amount may be calculated for each head-related transfer function included in the new set of head-related transfer functions. In this case, table data linking the head-related transfer functions with the adjustment amounts may be stored in the storage unit 105. Alternatively, the adjustment amounts may be calculated by an external device and stored in the memory of the external device. When performing convolution processing of head-related transfer functions, by referencing the adjustment amounts for the time shift adjustment and gain adjustment linked to the applied head-related transfer functions, it is not necessary to calculate the adjustment amount for each convolution processing, which contributes to reducing the amount of processing.

[0159] In this way, when a new set of head-related transfer functions not stored in the storage unit 105 is loaded, adjustment amounts for the time shift adjustment and gain adjustment used in the panning process may be determined for the new set of head-related transfer functions before storing the new set in the storage unit 105, and a head-related transfer function table may be constructed by linking the new set of head-related transfer functions with the determined adjustment amounts, and the head-related transfer function table may be stored in the storage unit 105. Then, when performing the panning process, these adjustment amounts are read from the storage unit 105, and the shift adjustment and gain adjustment are applied based on these adjustment amounts. Note that the new head-related transfer functions may be previously stored in the storage unit 105, but may be temporarily removed from the storage unit 105 when decoding a sound signal, when powering on the sound reproduction system 100, when initializing the sound reproduction system 100, or the like, and then stored again in the storage unit 105. Alternatively, table data for the new set of head-related transfer functions and table data previously stored in the storage unit 105 may each be stored in the storage unit 105 as table data corresponding to a different set of head-related transfer functions. It goes without saying that the table data stored in the storage unit 105 may be a head-related transfer function table or an adjustment amount table that is a part of the head-related transfer function table. Furthermore, determining such an adjustment amount is effective without switching whether or not to perform panning processing. That is, instead of the first generation unit 133 and the second generation unit 134, a third generation unit different from the first generation unit 133 and the second generation unit 134 may be provided. The third generation unit applies time shift adjustment and gain adjustment to the playback sound using adjustment amounts associated with the new head-related transfer functions stored in the storage unit 105 to convert the playback sound into a representative sound, and generates an output sound signal by convolving a head-related transfer function corresponding to a representative direction from each position of the representative point toward the user's position into the representative sound.

[0160] Other Embodiments Although the embodiments have been described above, the present disclosure is not limited to the above-described embodiments.

[0161] For example, the sound reproduction system described in the above embodiment may be realized as a single device including all of the components, or may be realized by allocating each function to multiple devices and coordinating these multiple devices. In the latter case, the device corresponding to the information processing device may be an information processing device such as a smartphone, tablet terminal, or PC. For example, in the sound reproduction system 100 having a function as a renderer that generates an audio signal with added sound effects, all or part of the renderer's functions may be performed by a server. That is, all or part of the acquisition unit 111, path calculation unit 121, output sound generation unit 131, and signal output unit 141 may be located on a server (not shown). In this case, the sound reproduction system 100 is realized by combining, for example, an information processing device such as a computer or smartphone, a sound presentation device such as a head-mounted display (HMD) or earphones worn by the user 99, and a server (not shown). Note that the computer, sound presentation device, and server may be connected to each other so as to be able to communicate with each other via the same network, or may be connected via different networks. When the sound reproduction system 100 is connected via different networks, the possibility of communication delays increases, so processing on the server may be permitted only when the computer, sound presentation device, and server are connected so as to be able to communicate with each other via the same network. Also, depending on the amount of bitstream data received by the sound reproduction system 100, it may be determined whether the server will take on all or part of the functions of the renderer.

[0162] The sound reproduction system of the present disclosure can also be realized as an information processing device that is connected to a reproduction device having only a driver and that simply reproduces an output sound signal generated based on acquired sound information to the reproduction device. In this case, the information processing device may be realized as hardware having a dedicated circuit, or as software that causes a general-purpose processor to execute specific processing.

[0163] In the above-described embodiment, the processing performed by a specific processing unit may be performed by another processing unit. The order of multiple processing operations may be changed, or multiple processing operations may be performed in parallel.

[0164] In the above-described embodiments, each component may be realized by executing a software program suitable for that component, or by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0165] Furthermore, each component may be realized by hardware. For example, each component may be a circuit (or integrated circuit). These circuits may form a single circuit as a whole, or each may be a separate circuit. Furthermore, each of these circuits may be a general-purpose circuit or a dedicated circuit.

[0166] Furthermore, the general or specific aspects of the present disclosure may be realized as an apparatus, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, etc. Furthermore, the general or specific aspects of the present disclosure may be realized as any combination of an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0167] For example, the present disclosure may be realized as an audio signal reproducing method executed by a computer, or as a program for causing a computer to execute the audio signal reproducing method. The present disclosure may also be realized as a computer-readable non-transitory recording medium on which such a program is recorded.

[0168] In addition, this disclosure also includes forms obtained by applying various modifications to each embodiment that a person skilled in the art would think of, or forms realized by arbitrarily combining the components and functions of each embodiment within the scope that does not deviate from the intent of this disclosure.

[0169] In the present disclosure, the encoded sound information can be rephrased as a bitstream containing a sound signal, which is information about a predetermined sound to be reproduced by the sound reproduction system 100, and metadata, which is information about the localization position when the sound image of the predetermined sound is localized at a predetermined position within a three-dimensional sound field. For example, the sound information may be acquired by the sound reproduction system 100 as a bitstream encoded in a predetermined format such as MPEG-H 3D Audio (ISO / IEC 23008-3). As an example, the encoded sound signal includes information about the predetermined sound to be reproduced by the sound reproduction system 100. The predetermined sound here refers to a sound emitted by a sound source object present in the three-dimensional sound field or a natural environmental sound, and may include, for example, a mechanical sound or the sounds of animals, including humans. When multiple sound source objects are present in the three-dimensional sound field, the sound reproduction system 100 acquires multiple sound signals corresponding to the multiple sound source objects.

[0170] On the other hand, metadata is, for example, information used to control acoustic processing of a sound signal in the sound reproduction system 100. Metadata may be information used to describe a scene expressed in a virtual space (three-dimensional sound field). Here, a scene is a term that refers to a collection of all elements representing three-dimensional images and acoustic events in a virtual space, modeled by the sound reproduction system 100 using metadata. In other words, the metadata here may include not only information for controlling acoustic processing, but also information for controlling video processing. Of course, the metadata may include information for controlling only either audio processing or video processing, or may include information used to control both. In the present disclosure, the bitstream acquired by the sound reproduction system 100 may include such metadata. Alternatively, the sound reproduction system 100 may acquire the metadata separately from the bitstream, as described below.

[0171] The sound reproduction system 100 generates virtual sound effects by performing sound processing on the sound signal using metadata included in the bitstream and additionally acquired position information of the interactive user 99. For example, sound effects such as early reflection sound generation, late reverberation sound generation, diffraction sound generation, distance attenuation effect, localization, sound image localization processing, and Doppler effect may be added. Information for switching on and off all or part of the sound effects may also be added as metadata.

[0172] All or part of the metadata may be obtained from sources other than the bitstream of audio information. For example, either the metadata controlling audio or the metadata controlling video may be obtained from sources other than the bitstream, or both may be obtained from sources other than the bitstream.

[0173] Furthermore, if metadata for controlling the video is included in the bitstream acquired by the audio reproduction system 100, the audio reproduction system 100 may have a function for outputting the metadata that can be used for controlling the video to a display device that displays images or a 3D video reproduction device that reproduces 3D video.

[0174] As an example, the encoded metadata includes information about a three-dimensional sound field including a sound source object emitting a sound and an obstacle object, and information about a localization position when a sound image of the sound is localized at a predetermined position in the three-dimensional sound field (i.e., the sound is perceived as arriving from a predetermined direction), i.e., information about the predetermined direction. Here, an obstacle object is an object that can affect the sound perceived by the user 99 by, for example, blocking or reflecting the sound emitted by the sound source object before it reaches the user 99. Obstacle objects may include not only stationary objects but also animals such as people or moving objects such as machines. Furthermore, when multiple sound source objects exist in a three-dimensional sound field, the other sound source objects may be obstacle objects for any one sound source object. Furthermore, both non-sound-source objects such as building materials or inanimate objects and sound-emitting sound objects may be obstacle objects.

[0175] The spatial information constituting the metadata may include not only the shape of the three-dimensional sound field, but also information representing the shape and position of obstacle objects present in the three-dimensional sound field and the shape and position of sound source objects present in the three-dimensional sound field. The three-dimensional sound field may be either a closed space or an open space, and the metadata may include information representing the reflectance of structures that can reflect sound in the three-dimensional sound field, such as floors, walls, or ceilings, and the reflectance of obstacle objects present in the three-dimensional sound field. Here, the reflectance is the ratio of the energy of reflected sound to incident sound, and is set for each frequency band of sound. Of course, the reflectance may be set uniformly regardless of the frequency band of the sound. Furthermore, when the three-dimensional sound field is an open space, parameters such as a uniform attenuation rate, diffracted sound, or early reflection sound may be used.

[0176] In the above description, reflectance is cited as a parameter related to an obstacle object or a sound source object included in the metadata, but the metadata may include information other than reflectance. For example, information about the material of the object may be included as metadata related to both the sound source object and the non-sound source object. Specifically, the metadata may include parameters such as diffusion rate, transmittance, or sound absorption rate.

[0177] Information about the sound source object may include information such as volume, radiation characteristics (directivity), playback conditions, the number and type of sound sources emitted from a single object, or information specifying a sound source area within the object. The playback conditions may, for example, determine whether the sound is a continuous sound or an event-triggering sound. The sound source area within the object may be determined based on the relative relationship between the position of the user 99 and the position of the object, or may be determined based on the object. When the sound source area is determined based on the relative relationship between the position of the user 99 and the position of the object, the surface from which the user 99 is viewing the object is used as the reference, and the user 99 can be made to perceive sound X as emanating from the right side of the object and sound Y as emanating from the left side of the object as viewed from the user 99. When the sound source area is determined based on the object, the surface from which the user 99 is viewing the object is used as the reference, and the sound emitted from which area of ​​the object can be fixed regardless of the direction the user 99 is viewing. For example, the user 99 can be made to perceive a high-pitched sound coming from the right side and a low-pitched sound coming from the left side when viewing the object from the front. In this case, when the user 99 goes around to the back of the object, the user 99 can be made to perceive a low-pitched sound coming from the right side and a high-pitched sound coming from the left side as viewed from the back.

[0178] The spatial metadata may include the time to early reflections, the reverberation time, or the ratio of direct sound to diffuse sound. If the ratio of direct sound to diffuse sound is zero, the user 99 will perceive only direct sound.

[0179] Furthermore, information indicating the position and orientation of the user 99 in the three-dimensional sound field may be included in the bitstream as metadata as an initial setting, or may not be included in the bitstream. If the information indicating the position and orientation of the user 99 is not included in the bitstream, the information indicating the position and orientation of the user 99 is acquired from information other than the bitstream. For example, the position information of the user 99 in the VR space may be acquired from an app that provides VR content. The position information of the user 99 for presenting sound as AR may be position information obtained by a mobile terminal performing self-position estimation using GPS, a camera, LiDAR (Laser Imaging Detection and Ranging), or the like. Note that the sound signal and metadata may be stored in a single bitstream or may be stored separately in multiple bitstreams. Similarly, the sound signal and metadata may be stored in a single file or may be stored separately in multiple files.

[0180] When an audio signal and metadata are stored separately in multiple bitstreams, information indicating other related bitstreams may be included in one or some of the multiple bitstreams in which the audio signal and metadata are stored. Also, information indicating other related bitstreams may be included in metadata or control information of each of the multiple bitstreams in which the audio signal and metadata are stored. When an audio signal and metadata are stored separately in multiple files, information indicating other related bitstreams or files may be included in one or some of the multiple files in which the audio signal and metadata are stored. Also, information indicating other related bitstreams or files may be included in metadata or control information of each of the multiple bitstreams in which the audio signal and metadata are stored.

[0181] Here, the related bitstreams or files are, for example, bitstreams or files that may be used simultaneously during audio processing. Furthermore, information indicating other related bitstreams may be collectively described in the metadata or control information of one bitstream among multiple bitstreams storing audio signals and metadata, or may be separately described in the metadata or control information of two or more bitstreams among the multiple bitstreams storing audio signals and metadata. Similarly, information indicating other related bitstreams or files may be collectively described in the metadata or control information of one file among multiple files storing audio signals and metadata, or may be separately described in the metadata or control information of two or more files among the multiple files storing audio signals and metadata. Furthermore, a control file collectively describing information indicating other related bitstreams or files may be generated separately from the multiple files storing audio signals and metadata. In this case, the control file does not need to store the audio signal and metadata.

[0182] Here, the information indicating the other related bitstream or file may be, for example, an identifier indicating the other bitstream, a file name indicating the other file, a URL (Uniform Resource Locator), or a URI (Uniform Resource Identifier). In this case, the acquisition unit 111 identifies or acquires the bitstream or file based on the information indicating the other related bitstream or file. Furthermore, the information indicating the other related bitstream may be included in the metadata or control information of at least some of the bitstreams among a plurality of bitstreams storing audio signals and metadata, and the information indicating the other related file may be included in the metadata or control information of at least some of the files among a plurality of files storing audio signals and metadata. Here, the file including the information indicating the related bitstream or file may be, for example, a control file such as a manifest file used for content distribution.

[0183] The present disclosure is useful in reproducing sound, for example, by allowing a user to perceive stereoscopic sound.

[0184] 99 User 100 Sound reproduction system 101 Information processing device 102 Communication module 103 Detector 104 Driver 105 Storage unit 111 Acquisition unit 112 Encoded sound information input unit 113 Decode processing unit 114 Sensing information input unit 121 Path calculation unit 131 Output sound generation unit 132 Switching unit 133 First generation unit 134 Second generation unit 141 Signal output unit 300 3D video reproduction device

Claims

1. an acquisition unit that acquires sound information including an audio signal and information about the position of a sound source object in a three-dimensional sound field; a first generation unit that generates an output sound signal using a head-related transfer function according to an arrival direction based on a position of the sound source object and a position of a user in the three-dimensional sound field, and the audio signal; a second generation unit that generates an output sound signal using a head-related transfer function according to a representative direction based on the position of a representative point set in the three-dimensional sound field and the position of the user, and the audio signal. Information processing device.

2. the first generation unit generates the output sound signal by convolving a head-related transfer function according to the arrival direction with a playback sound emitted from the sound source object in response to the audio signal; The second generation unit performs a conversion process to convert the reproduced sound into a representative sound arriving from the representative point, and generates the output sound signal by convolving a head related transfer function according to the representative direction. The information processing device according to claim 1 .

3. In the conversion process, a time shift adjustment and a gain adjustment are applied to the reproduced sound to convert it into the representative sound. The information processing device according to claim 2 .

4. In the time shift adjustment of the conversion process, a time shift calculated so as to maximize the cross-correlation between the head related transfer function corresponding to the arrival direction and the head related transfer function corresponding to the representative direction, or a time shift obtained by adding a negative sign to the time shift, is performed on the reproduced sound. The information processing device according to claim 3 .

5. In the conversion process, at least one of the time shift adjustment and the gain adjustment is performed by applying a weighting filter on the frequency axis and then performing a time shift calculated so as to maximize the cross-correlation, or a time shift obtained by adding a negative sign to the time shift. The information processing device according to claim 4 .

6. In the conversion process, for each of the two or more representative points, a gain set for each of the reproduced sound and the representative direction is applied to the time-shifted reproduced sound. The information processing device according to claim 4 .

7. In the conversion process, when synthesizing a head-related transfer function vector according to the arrival direction by the sum of head-related transfer function vectors according to the representative direction, a gain calculated so that an error signal vector between the synthesized head-related transfer function vector and the head-related transfer function vector according to the arrival direction is orthogonal to the head-related transfer function vector according to the representative direction is used. The information processing device according to claim 6 .

8. In the conversion process, a gain calculated so as to minimize the energy or L2 norm of an error signal vector between the synthesized head-related transfer function vector and the head-related transfer function vector according to the direction of arrival is used. The information processing device according to claim 6 .

9. The error signal vector is filtered by a weighting filter on the frequency axis. The information processing device according to claim 8 .

10. At the time of initialization, the information processing device stores in a storage unit an adjustment amount table in which head related transfer functions in a representative direction and adjustment amounts in time shift adjustment and gain adjustment used in the conversion process are linked for each direction of the head related transfer functions; In the conversion process, The playback sound is converted into the representative sound by applying a time shift adjustment and a gain adjustment to the playback sound using the adjustment amounts associated with each direction of a head related transfer function according to the representative direction in the adjustment amount table stored in the storage unit. The information processing device according to claim 3 .

11. the information processing device determines a plurality of the representative directions at the time of the initialization; The adjustment amount table is created based on the determined head-related transfer functions of the representative directions. The information processing device according to claim 10.

12. the sound information includes a flag that specifies whether the output sound signal is generated using the first generation unit or the second generation unit, The information processing device generates the output sound signal using one of the first generation unit and the second generation unit, which is specified by a flag included in the acquired sound information. The information processing device according to claim 1 .

13. a switching unit that switches between generating the output sound signal using the first generating unit and generating the output sound signal using the second generating unit; The information processing device according to claim 1 .

14. a path calculation unit that calculates a propagation path of a reproduced sound emitted from the sound source object in response to the audio signal, and calculates a synthesized sound that arrives at the user's position by indirect propagation of the reproduced sound according to the calculated propagation path of the reproduced sound, and an arrival direction of the synthesized sound; The information processing device according to any one of claims 1 to 13.

15. 1. A computer-implemented information processing method for generating an output sound signal as sound coming from a sound source object in a virtual three-dimensional sound field by processing sound information, comprising: acquiring a position of the sound source object and an audio signal, the audio signal including a reproduced sound emitted from the sound source object by the audio signal; obtaining a user's position within the three-dimensional sound field; calculating an arrival direction of the reproduced sound arriving at the user's position from the position of the sound source object; generating the output sound signal using the calculated head-related transfer function according to the arrival direction and the reproduced sound; generating the output sound signal using a head-related transfer function according to a representative direction based on the position of a representative point set in the three-dimensional sound field and the position of the user, and the audio signal. Information processing methods.

16. A method for causing a computer to execute the information processing method according to claim 15. program.

17. a storage unit that stores a plurality of directions, a time shift adjustment amount, and a gain adjustment amount in association with each other; an acquisition unit that acquires an audio signal and information about the position of a sound source object in a three-dimensional sound field; a second generation unit that generates an output sound signal as a sound arriving at the user's position from a second direction, using the audio signal and the time shift adjustment amount and gain adjustment amount corresponding to a first direction based on the position of the sound source object and the user's position in the three-dimensional sound field. Information processing device.

18. the storage unit further stores a head-related transfer function corresponding to the second direction, The second generation unit generates an output sound signal as a sound arriving at the user's position from the second direction by using the audio signal, the time shift adjustment amount and the gain adjustment amount corresponding to the first direction, and a head-related transfer function corresponding to the second direction. The information processing device according to claim 17.

19. the storage unit further stores head-related transfer functions corresponding to the second direction and directions other than the second direction, the second generation unit generates an output sound signal as sound arriving at a position of the user from the second direction by using the audio signal, the time shift adjustment amount and the gain adjustment amount corresponding to the first direction, and a head-related transfer function corresponding to the second direction; The information processing device further includes a first generation unit, The first generation unit generates an audio signal as a sound arriving at a position of the user from the first direction, using the audio signal and a head-related transfer function corresponding to the first direction. The information processing device according to claim 17.